Two questions get asked of any research finding, and they are routinely collapsed into one.
The first: did anything happen, or could this be chance? The second: how much happened?
Statistical significance answers the first. Effect size answers the second. Coverage of research on children answers the first and reports it as though it had answered the second, and that single substitution accounts for most of the misleading things you will read.
What significance actually means
A significance test asks how surprising the observed data would be if there were no real effect. If sufficiently surprising by a conventional threshold, the result is called significant.
Note what that does not say. It does not say the effect is large. It does not say it is important. It does not say the finding is true, and it does not give the probability that the hypothesis is correct, though it is constantly reported as though it did.
It says: this pattern would be unusual if nothing were going on. That is all, and it is much less than the word significant suggests in ordinary English. The word is a historical accident and it has done enormous damage.
The size problem in one paragraph
Significance depends on both the size of the effect and the size of the sample. A large effect in a small sample can be significant. A minuscule effect in a very large sample will also be significant.
Modern research on children often uses very large administrative or cohort datasets. In such datasets, almost everything is significantly related to almost everything else, because with enough participants any tiny relationship is detectable.
This is why headlines about large studies of children so often describe alarming links that turn out, on inspection, to correspond to differences no parent could observe. The relationship is real and it is negligible, and both halves of that sentence are true simultaneously.
Ways to picture an effect without arithmetic
Effect sizes are usually reported in standardised units that mean nothing to a general reader. Several translations help.
Overlap. If two groups differ, how much do they overlap? For most effects reported in education and developmental research, the overlap is enormous. Knowing which group a child is in tells you very little about the child.
The pick two test. If you picked one child at random from each group, how often would the one from the higher scoring group actually score higher? For a typical educational effect the answer is a little over half the time. Barely better than a coin.
In the units of the thing. Months of additional progress, words of vocabulary, minutes of sleep. Concrete units are always more informative than standardised ones, which is why the Education Endowment Foundation reports its toolkit in months of progress. That has its own critics, and it is a better default than an abstract coefficient.
Small does not mean nothing
The opposite error deserves equal attention.
A small effect applied across an entire population can matter a great deal. A tiny average shift in a national measure represents a large number of children. Public health and education policy operate at that scale, and dismissing small effects as trivial misunderstands what policy is for.
A small effect can also be cheap. An intervention costing almost nothing and producing a small gain may be excellent value, while an expensive one producing a slightly larger gain is not. Effect size without cost is only half the decision.
What a small effect will not do is transform an individual child, which is how such findings are usually sold to parents.
The average conceals the spread
One more layer. An effect size is an average across participants, and averages hide variation.
A programme with a modest average effect may work well for some children and not at all for others. Or it may work slightly for everybody. Those are entirely different situations with the same summary number, and most studies are not powered to distinguish them.
This matters for a parent reading about an intervention. The average result does not tell you what would happen to your child, and no study of averages ever can. It is the same problem as reading a milestone age backwards onto one child, described in the range of normal.
Reading coverage with this in hand
When you meet a claim about research on children, ask how big.
If the answer is not there, that is information. Coverage that reports significance without size is reporting the less interesting half. If the answer is given in standardised units, translate it: how much overlap, how many months, how many words.
And notice the vocabulary. Words like linked, associated and raises the risk of describe direction without magnitude and are perfectly compatible with an effect too small to notice.
The one thing to remember
Significant means probably not chance. It does not mean big, and it does not mean important. Everything else in this article follows from holding those apart.
For what to do with any single result, see what one study can tell you, and for how coverage manages to lose this distinction so reliably, why the parenting press gets it wrong.
