OnCo
key papersKey paper

Reproducibility Project: Cancer Biology found that landmark preclinical results mostly shrank or vanished on replication

An eight-year effort to repeat 50 experiments from 23 high-impact cancer biology papers found that replication effect sizes were on average 85% smaller than the originals, and fewer than half of the effects replicated by most criteria.

The Center for Open Science and Science Exchange set out in 2013 to replicate key experiments from 53 high-impact cancer biology papers published 2010-2012, with registered reports and original-author consultation. Only 50 experiments from 23 papers could be completed; the companion paper (Errington et al., eLife 2021, 'Challenges for assessing replicability') documents that no original paper contained enough methodological detail to design a replication without contacting the authors, and that about a third of authors were unhelpful or unresponsive.

Across 158 measured effects, the median replication effect size was 85% smaller than the original; 92% of replication effects were smaller than the originals. Using five criteria, 46% of effects replicated on more criteria than they failed; for original positive results, about 40% replicated, while null results replicated at 80%.

The project quantified the preclinical reproducibility problem that pharmaceutical groups (Begley and Ellis 2012; Prinz 2011) had reported anecdotally, and shaped funder requirements for rigour and data sharing.

Meta-analysisHas not changed practice yet
Authors
Errington TM, Mathur M, Soderberg CK, et al.
Published
eLife, 2021
What it found
  • 50 experiments from 23 papers completed out of 193 planned from 53 papers
  • Median replication effect size 85% smaller than original; 92% of replication effects smaller than originals
  • 46% of 158 effects replicated on more criteria than they failed; original positive results replicated about 40% of the time, null results 80%
  • No original paper described methods in enough detail to replicate without author contact; 32% of authors were minimally helpful or unresponsive
What it means

Many exciting laboratory findings that motivate drug programmes are weaker or less reliable than published, which helps explain the high failure rate of drugs entering clinical trials. It argues for pre-registration, detailed methods, data sharing and independent replication before major translational investment.

Be careful
  • Replications used the original protocols where possible, but reagents, animals and laboratories differ; some failures may reflect context sensitivity rather than error
  • Selection of high-impact papers may not represent the field
  • Under-powered replications could miss true effects; effect-size shrinkage is the more robust finding
  • Only about a quarter of planned experiments were completed, which limits generalisation

Connected

9top