Why trials fail: underpowered, wrong endpoint, control arm drift, subgroup fishing, crossover
Trials fail for a short list of reasons that recur: too few patients for the real effect, an endpoint that does not track what matters, a control arm that did better than the planners assumed, a benefit that was only ever a subgroup illusion, and control patients receiving the experimental drug anyway.
Overview
Most negative trials are negative because the drug does not work. But a substantial minority fail for reasons that have nothing to do with the treatment, and reading the failure modes is part of reading the result. Underpowering: the sample size was built on an effect estimated from a small early study, which by selection tends to be inflated (the winner's curse), and the true, smaller effect could not be detected. DREAM3R, sized on encouraging single-arm and phase 2 data, was stopped early without meeting its survival endpoint. Wrong endpoint: the primary endpoint was a surrogate that does not predict benefit in that setting, or was too early. UKCTOCS found ovarian cancers earlier without reducing deaths; TROPION-Breast01 met progression-free survival and missed overall survival, and was approved on the former. Wrong dose, a cousin: RTOG 0617 tested 74 Gy against the standard 60 Gy and patients on the higher dose lived shorter.
Control arm drift: the standard of care improves during a five-year trial, or the control arm is delivered better inside a trial than in the historical data used to plan it, and the assumed gap closes. ACT IV found both its vaccine and control arms outperformed historical expectations, and CONVERT's arms both beat historical controls, which is why single-arm results compared with history should be trusted less than they usually are and why ANNOUNCE and ATLANTIS, confirmatory trials of drugs approved on small early studies, showed no survival benefit. Subgroup fishing: a trial that misses overall reports a positive subgroup, and the more subgroups examined the more certain it is that one will be positive by chance; the preoperative progesterone trial at Tata Memorial found a benefit only in node-positive women and its authors rightly called for confirmation rather than claiming a result. Crossover contamination: control patients receive the experimental drug at progression, so a real survival difference is diluted; VISION, PSMAfore, TheraP and CodeBreaK 200 all show clear progression gains with little or no survival difference after crossover.
There are quieter failure modes too. Non-adherence and dropout pull any comparison towards no difference, which flatters a non-inferiority trial and sinks a superiority one. Protocol choices can decide the verdict: BELINDA's long manufacturing interval and strict week-twelve event definition erased an effect that two similar CAR-T trials found. Informative censoring biases progression-free survival when patients leave for reasons linked to their outcome. And a trial can fail to fail: a result that is statistically significant but clinically trivial, a technically negative trial whose confidence interval nonetheless excludes any meaningful harm, or a positive result in a population that no longer exists because the standard has moved on. The failure museum on this site collects the trials in the corpus that did not work and the lesson each taught.
Similar pages
not linked directly; found by shared links- TermEstimands and intercurrent events (ICH E9(R1))
Shares Intention-to-treat (ITT) and per-protocol analysis, Pre-specified vs post-hoc analysis, Trial protocol and statistical analysis plan, Kaplan-Meier curve, censoring and proportional hazards.
- TermPrimary, secondary and co-primary endpoints
Shares Futility analysis (stopped for futility), Pre-specified vs post-hoc analysis, Statistical significance (P values, alpha, multiplicity), Trial registration and results reporting (ClinicalTrials.gov, EU CTR).
- TermGroup sequential design, stopping rules and alpha spending
Shares TROPION-Breast01, Futility analysis (stopped for futility), Statistical significance (P values, alpha, multiplicity), Statistical power, sample size and re-estimation.
- TermStepped-wedge design
Shares Non-inferiority trial, Non-inferiority margin and equivalence trials, Pragmatic trial, Randomised trial.
- TermTrial lifecycle: from protocol to label
Shares ATLANTIS, ANNOUNCE, Trial registration and results reporting (ClinicalTrials.gov, EU CTR), Confirmatory trial.
- TermRegistry-based randomised trial
Shares External and synthetic control arms, Pragmatic trial, Randomised trial, Seamless, adaptive and Bayesian trial designs.
- TermUmbrella trial
Shares Trial protocol and statistical analysis plan, Enrichment and biomarker-stratified designs, Single-arm trial, Seamless, adaptive and Bayesian trial designs.
- TermConfidence interval
Shares Statistical significance (P values, alpha, multiplicity), Statistical power, sample size and re-estimation, Kaplan-Meier curve, censoring and proportional hazards, Non-inferiority margin and equivalence trials.