Every week a study appears in the news about something a parent does. It has a number in it. It has an implication. And it is very often the only study that has ever been done on the question, or the only one that produced a result interesting enough to report.
This piece is about what such a study can actually support. It is the founding article of this magazine's evidence section, and everything else in the section elaborates on it.
Signal in noise
Start with what a study is doing. There is something you want to know about a population. You cannot examine the population, so you take a sample. You cannot measure the thing directly, so you use a proxy. You then look for a pattern.
Every one of those steps introduces noise. A different sample would have given a different answer. A different measure would have captured a slightly different construct. A different analyst would have made different defensible decisions about outliers, covariates and exclusions.
The result you read is one draw from a distribution of results that the same question could have produced. Sometimes it lands near the truth. Sometimes it does not. Nothing about the paper tells you which.
Why the published one is likely to be the extreme one
Here is the part that changes how you read everything.
Studies that find something are more likely to be written up, submitted, accepted and covered than studies that find nothing. This is publication bias, it is well documented, and it is not a conspiracy: it emerges from thousands of individually reasonable decisions.
The consequence is that the published record is not a random sample of the research conducted. It is enriched for larger effects. If ten teams investigate something real but modest, the one that happened to draw a sample producing a striking result is the one you hear about.
So the expected size of a published effect is larger than the truth, and the first replication is expected to be smaller. This is not cynicism. It is arithmetic, and it is why replication matters so much.
The questions worth asking
Not a scoring system. A set of habits.
Who was in it? Twenty undergraduates, or four thousand children followed for a decade. Both can be informative and they support very different claims.
What was actually measured? This is the most useful question and the least asked. A study of screen time measured a parent's estimate of hours. A study of curiosity measured a teacher's rating. The conclusion is written in the language of the concept; the finding is in the language of the measure.
Was anything assigned, or only observed? If nobody assigned anything, causal language in the write up is decoration. See correlation and cause.
Compared with what? An intervention compared with nothing at all tells you less than one compared with a plausible alternative, because doing something is not the same as doing this thing.
How big was it? Not whether it reached significance, but how much difference it made. See effect size in plain words.
Has anyone found it again? The single most informative question, and the one newspapers never answer.
Why new studies get more coverage than better ones
A systematic review synthesising forty trials is more informative than any of them and gets a fraction of the attention, because it is not news. It has no single striking number and no counterintuitive angle.
The result is an inversion: coverage is strongest where evidence is weakest. A first surprising finding is maximally newsworthy and minimally reliable. A well replicated dull one is the reverse.
The practical implication is that the reliability of what you read about children's development is roughly inversely proportional to how interesting the headline is. Institutions that synthesise rather than generate, such as Cochrane and the Education Endowment Foundation, are duller and considerably more useful.
The translation problem
Between the paper and the reader sit several stages, each of which loses caveats: the abstract, the press release, the wire copy, the article, the headline, and the version repeated at a party.
Losses accumulate in one direction. Nobody adds uncertainty at any stage. A finding described as an association in the paper becomes a link in the release, an effect in the article and a cause in the headline, without any single step being an outright lie.
Knowing that this ratchet exists is most of the defence against it. We describe it in more detail in why the parenting press gets it wrong.
What to do with a striking study
Be interested. Do not change anything.
That sounds passive and is the correct response nearly every time. A single result is a reason to watch a question, not to reorganise a household. If the finding is real it will still be there in three years, better established and less exciting. If it is not, you will have saved yourself the trouble.
The exception is when a finding is consistent with a large existing body of evidence rather than surprising. A study confirming something already well supported is much more likely to be right and much less likely to be reported.
The disposition worth having
Not scepticism as a pose. Scepticism as a pose is as lazy as credulity and considerably more smug.
What is worth having is a sense of how much weight a piece of evidence can bear. A single study bears a little. A body of consistent evidence bears more. A finding that has survived people trying hard to break it bears a great deal.
Most of what reaches parents is in the first category and is presented as though it were in the third. That gap is the subject of this entire section.
