Statistical power, sample size and re-estimation
Before a trial starts, statisticians work out how many patients, or how many deaths or relapses, are needed to detect the benefit they hope for; an adaptive trial can check that guess part way through and enlarge itself if the guess was wrong.
Overview
Every trial is sized around a guess. The sponsor states the smallest benefit worth detecting (a hazard ratio of 0.75, say), the false-positive rate it will accept (usually 5 percent two-sided) and the chance it wants of detecting the benefit if it is real (the power, usually 80 or 90 percent). From those three numbers and an estimate of how quickly events will occur comes the sample size, or for survival endpoints the number of events the trial must wait for before the analysis. If the guess about the benefit is too optimistic, the trial is underpowered: a real but smaller effect can produce a hazard ratio below 1 whose confidence interval still crosses 1, and the trial is declared negative. Underpowering is the commonest single reason a genuinely useful treatment fails a trial.
Sample size re-estimation is the adaptive fix. At a planned interim look, the trial re-examines the assumptions and, within rules fixed in advance, may increase the number of patients or events. Blinded re-estimation uses only pooled data (the overall event rate, the overall variance) and costs nothing statistically. Unblinded re-estimation looks at the emerging treatment effect, typically enlarging the trial only when the interim result falls in a promising zone, and must be paired with methods that keep the overall false-positive rate at its nominal level. Regulators accept both when the rules are pre-specified and the interim data stay behind the data monitoring committee's firewall.
The corpus shows why sizing matters. DREAM3R, the phase 3 of durvalumab with chemotherapy in mesothelioma, was stopped early and did not meet its overall survival endpoint despite an encouraging single-arm predecessor, a pattern in which a small early study inflates the assumed effect and the confirmatory trial is sized for a benefit that was never real. Add-Aspirin took the opposite approach, powering four tumour-specific cohorts separately for their own recurrence endpoints inside one 11,000-patient trial, so that a modest effect of a cheap drug could be detected in each cancer. Event-driven trials such as ADAURA report when a pre-set number of events has accrued, which is why readout dates slip when patients do better than expected.
Similar pages
not linked directly; found by shared links- TermData monitoring committee (DSMB, IDMC)
Shares DREAM3R, Futility analysis (stopped for futility), Group sequential design, stopping rules and alpha spending, Interim analysis, readout and data cut-off.
- TermKaplan-Meier curve, censoring and proportional hazards
Shares Data maturity (immature vs mature survival data), Confidence interval, ADAURA, Why trials fail: underpowered, wrong endpoint, control arm drift, subgroup fishing, crossover.
- TermPrimary, secondary and co-primary endpoints
Shares Data maturity (immature vs mature survival data), Futility analysis (stopped for futility), Statistical significance (P values, alpha, multiplicity), Group sequential design, stopping rules and alpha spending.
- TermP-value
Shares Confidence interval, Statistical significance (P values, alpha, multiplicity), Group sequential design, stopping rules and alpha spending, Hazard ratio (HR).
- TermBayesian trial design
Shares Confidence interval, Statistical significance (P values, alpha, multiplicity), Seamless, adaptive and Bayesian trial designs, Phase 1, 2 and 3 trials.
- TrialTrilynX
Shares Futility analysis (stopped for futility), Group sequential design, stopping rules and alpha spending, Interim analysis, readout and data cut-off.
- TermLandmark and milestone survival (5-year survival, median follow-up)
Shares Data maturity (immature vs mature survival data), Interim analysis, readout and data cut-off, Hazard ratio (HR).
- TermAbsolute versus relative benefit (number needed to treat)
Shares Confidence interval, ADAURA, Hazard ratio (HR), Phase 1, 2 and 3 trials.