Everybody knows that a study of twenty people is weaker than a study of two thousand. That much is intuitive. What is less intuitive, and considerably more useful, is that small studies do not merely err randomly. Under the conditions of real research, they err upwards.
Understanding why takes about five minutes and permanently improves how you read a headline.
Sampling variation
Suppose there is a real difference between two groups in a population, and it is modest. You take a sample and measure it.
With a small sample, chance plays a large role. A few unusual participants shift the average substantially. Repeat the exercise many times and the estimates scatter widely: some far above the truth, some far below, occasionally reversed in sign.
With a large sample, the scatter narrows. Individual oddities cancel. Estimates cluster near the truth.
So far this is symmetric: small studies are wrong in both directions equally. If that were the end of it, small studies would be merely imprecise.
The filter that makes it asymmetric
Now add the way research actually reaches the world.
For a result to be published, and certainly for it to be reported in a newspaper, it generally has to reach statistical significance. In a small sample, that requires the observed effect to be large. A small study that happens to draw a sample showing the true, modest effect will produce a non significant result and is more likely to end up unpublished.
So the small studies you see are disproportionately the ones that overestimated. The filter selects for exaggeration.
This is why the first study on a topic so often reports a large effect and subsequent larger studies report a smaller one. Nothing has changed about the world. The selection has changed, because a large study can detect a modest effect and therefore gets published while reporting the modest truth.
Power, and the word that hides the problem
Statistical power is the probability that a study will detect an effect of a given size if it exists. It depends on sample size, on how big the effect is, and on how noisy the measurement is.
Studies in psychology and education have historically been badly underpowered: too small to reliably detect the effects they were looking for. This was known for decades and largely ignored, and it is one of the structural causes of the replication problems described in replication in plain words.
An underpowered study that finds nothing has told you almost nothing, since it probably could not have found anything. An underpowered study that finds something has told you that either an effect exists and this sample overstated it, or nothing exists and this is a false positive. Neither reading supports the confident version.
Why this bites hardest in research on children
Several features of the subject make samples small and noisy.
Recruiting young children is expensive and slow. Testing them takes an adult per child, and sessions fail because children are tired, shy or uninterested. Attrition in longitudinal work is substantial and not random, since families who drop out differ from those who stay.
Measurement is also noisier than in adult research. A four year old's performance varies enormously between one morning and another, which inflates the variability against which any effect must be detected.
The result is a literature with many small studies, exactly the conditions under which the upward bias operates most strongly. This is not a criticism of developmental researchers, who are working within real constraints. It is a reason to read their field with the constraint in mind.
Large is not automatically good
The obvious corrective, run bigger studies, brings its own difficulties, and it would be a poor magazine that omitted them.
A very large sample detects very small effects, and detection is not importance. Large administrative datasets routinely produce statistically significant relationships too small to matter to anybody, which is the subject of effect size in plain words.
Large samples also tend to use cruder measures, because detailed assessment does not scale. A study of fifty thousand children using a four item questionnaire may be measuring something less well than a study of eighty using a proper assessment.
And size does nothing about bias. A large biased sample gives a precise wrong answer, which is worse than an imprecise one because it inspires confidence.
What to do with this as a reader
When a study is reported, look for the number of participants, which good reporting includes and poor reporting omits.
If it is small and the effect described is dramatic, apply the correction: the true effect, if any, is probably considerably smaller. If it is small and no effect was found, conclude very little.
If it is large, ask what was measured rather than how many were measured, and ask how big the difference was in terms you can picture.
And in every case, ask whether anybody has done it again, which remains the most informative question available and the one least often answered.
The shape of the thing
Small studies are not blurry photographs of the truth. They are photographs selected for being dramatic, from a pile in which the dull ones were discarded.
That is a specific and correctable error. Once you know the direction of the bias, you can adjust for it in your head, which is the closest thing to a superpower this section can offer. For the general framework, see what one study can tell you.
