Little Minds

How children learn, for people who like thinking

Little Minds Magazine
Published by Northbank Media
August 2026
Reading the research

Replication in plain words, and what happened when psychology checked itself

A finding that only exists once is not yet knowledge. The past decade of attempting to reproduce well known results has been uncomfortable and enormously useful.

Evidence · Replication · Method
The short answer

Replication means doing a study again and seeing whether the same thing happens. Large coordinated attempts to replicate published psychology findings have succeeded considerably less often than expected, with replicated effects typically much smaller than the originals. This is a story about how a field improved, not about research being worthless.

Pressed leaves from one tree, laid side by side to be compared.
Pressed leaves from one tree, laid side by side to be compared.

Replication is the least glamorous activity in science and the one that does the most work. It means doing a study again, ideally by different people, and seeing whether the same thing happens.

A finding that has only ever happened once is not yet knowledge. It is a candidate.

What the replication projects found

Over the past decade or so, coordinated efforts have taken sets of published findings and attempted to reproduce them, usually with larger samples and preregistered plans.

Success rates have been considerably lower than most researchers expected. In several large projects, roughly half or fewer of the attempted replications produced a result in the same direction and reaching significance, and effects that did replicate were typically substantially smaller than the originals.

Similar exercises in other fields, including parts of biomedical research and economics, produced comparable discomfort. This is not a psychology problem, though psychology examined itself earliest and most publicly, which is to its credit and has damaged its reputation accordingly.

Why it happened, without villains

Almost none of this required misconduct. It emerged from ordinary incentives.

Underpowered studies. Small samples, as set out in sample size in plain words, produce inflated estimates when filtered for significance.

Analytical flexibility. A dataset permits many defensible analyses: which participants to exclude, which variables to control, how to handle outliers, which of several outcomes to emphasise. Making those choices after seeing the data, without intending to cheat, reliably produces significant results from noise.

Publication incentives. Careers were built on novel positive findings. Replications and null results were hard to publish, so they were not attempted.

The file drawer. Studies that found nothing stayed unpublished, so the visible literature was not the actual literature.

Each of those is individually reasonable and collectively corrosive, which is the characteristic shape of a systemic problem.

What this means for research on children

Several widely repeated claims about child development have had a hard decade. Effects that entered popular guidance on the strength of one striking demonstration have failed to appear at the same size when tested properly.

Specific examples appear elsewhere in this magazine: the mindset literature in praise and feedback, working memory training in working memory in plain words, and the bilingual executive function advantage in bilingual households.

Developmental research is also unusually exposed, because its samples are small, its measurement is noisy and its participants are hard to recruit. The field has responded seriously, including through multi laboratory collaborations that run the same infant protocol across many sites, which is exactly the right structural answer.

What the crisis does not license

Two bad inferences are common and should be resisted.

The first is nihilism: that since some findings failed, none can be trusted, and personal intuition is as good as evidence. This does not follow. Many findings replicate robustly, and the ones this magazine leans on are chosen partly for that. Spacing and retrieval effects in memory, described in why repetition works, have survived a century of attempts to break them.

The second is selective scepticism: applying the replication argument only to findings you dislike. This is common in education debates in every direction, and it is intellectually worthless.

The correct posture is to ask, of any claim, how many independent groups have found it, how large the effect is, and whether the studies were preregistered.

What has actually improved

More than the gloomy coverage suggests.

Preregistration, stating hypotheses and analysis plans before collecting data, has become common and removes most of the flexibility problem. Registered reports, where journals accept a study on the strength of its design before results exist, remove publication bias entirely for those papers. Data and materials sharing has become normal. Sample sizes have risen substantially. Multi site collaborations have become routine in infancy research.

Ten years is a short time for a field to change its methods this much. A discipline that publicly examined its own reliability and then altered its practices is behaving better than one that never looked.

What a reader should take from it

A calibration, not a dismissal.

Findings published before about the middle of the last decade, based on small samples, striking and counterintuitive, deserve considerable caution. Findings that have been reproduced by independent groups, preregistered, with adequate samples, deserve substantially more confidence.

And the single most useful question about any claim remains: has anyone else found this? A great deal of what circulates in parenting advice was demonstrated once, in a small sample, decades ago, and has never been successfully found again. That is a fact about the claim, and it is discoverable.

For how such claims travel regardless, see why the parenting press gets it wrong.

Questions we are asked

Does the replication crisis mean psychology is worthless?

No. It means a portion of the published literature was less reliable than assumed, that the causes are structural and understood, and that the field has changed its methods substantially in response.

How do I know whether a finding has replicated?

Look for whether independent groups have reported it, whether syntheses such as Cochrane reviews or Education Endowment Foundation summaries cover it, and whether the studies were preregistered. Coverage of a single new paper will not tell you.

What is preregistration?

Publishing your hypotheses and analysis plan before collecting data, so that the analysis cannot be shaped by the results. It is one of the more effective reforms of the past decade.

Why were small samples tolerated for so long?

The statistical arguments were available for decades and the incentives pointed the other way. Publishing more papers with smaller samples was rewarded, and nobody was checking.

Is medical research affected too?

Yes, and medicine developed its responses earlier: trial registration, reporting requirements and systematic review infrastructure. Psychology and education have adopted similar measures more recently.

Institutions and frameworks referred to

Links to public bodies, statutory frameworks and research databases. They are cited because they are public and checkable, not as endorsement of anything written here. All external links on this site are marked nofollow.

No commercial links. This article contains no affiliate links, no sponsored mentions and no links to any commercial client. Nothing in it has been paid for, and no organisation has seen it before publication. Little Minds Magazine recommends no products to parents. Published by Northbank Media. Our two declared revenue lines, paid listings and a labelled newsletter sponsor, are described in full on the about page and cannot influence any evidence framing on this site.

The Little Minds letter

One email a fortnight. What we have published, one thing we changed our mind about, and one piece of research worth reading properly. No advice, no products and no guilt.

We use your address for the letter and nothing else, and we never pass it to a sponsor. Details in privacy.

Sponsor lineEach issue carries one clearly labelled sponsor line, sold at a published rate. It sits below the editorial and can buy nothing else.