Group sequential design, stopping rules and alpha spending
A group sequential trial plans in advance how many times it will peek at the data and how strong the evidence must be at each peek to stop early, so that looking several times does not inflate the chance of a false positive.
Overview
Every look at accumulating trial data is another chance to be fooled by noise. If a trial were analysed at 5 percent significance five times, the chance of at least one false positive would be far above 5 percent. A group sequential design fixes this by deciding in advance how many interim analyses there will be, at what fractions of the total information (events), and how much of the total 5 percent false-positive budget, the alpha, is spent at each. The O'Brien-Fleming boundary spends almost nothing early, requiring an overwhelming result to stop at the first look, and saves most of the alpha for the final analysis; the Lan-DeMets spending function generalises this so that the timing of looks can shift with the actual event rate. A result reported as significant at an interim analysis has crossed a pre-specified boundary at a nominal threshold much stricter than 0.05, which is why trial papers quote odd-looking p-value thresholds.
ADAURA is the worked example. The trial of three years of adjuvant osimertinib after resection of EGFR-mutant lung cancer was unblinded early on the recommendation of its independent data monitoring committee because the disease-free survival difference was far beyond the boundary, and the overall survival benefit (hazard ratio 0.49) followed at five years. Early stopping for efficacy is a mixed blessing: it gets an effective drug to patients sooner and spares control patients, but a trial stopped at the first boundary crossing tends to overestimate the effect (the truncation bias), and the secondary endpoints and long-term safety data are cut short. Futility boundaries work in the other direction: JAVELIN Head and Neck 100 and TrilynX were stopped when the interim data made success implausible, and MAGNITUDE closed one of its two cohorts for futility while the other continued.
The rules only protect the trial if they are followed. Boundaries must be written into the statistical analysis plan before the first look, the interim results must be seen only by the data monitoring committee, and the trial should not be stopped on a secondary endpoint or an unplanned look. When a trial with two primary endpoints, such as progression-free and overall survival, splits its alpha between them, a win on one at a small allocated alpha and a miss on the other is a common and confusing outcome, as in TROPION-Breast01 where progression-free survival was met and overall survival was not.
Similar pages
not linked directly; found by shared links- TermPre-specified vs post-hoc analysis
Shares P-value, Statistical significance (P values, alpha, multiplicity), MAGNITUDE, Trial protocol and statistical analysis plan.
- TermData maturity (immature vs mature survival data)
Shares Statistical power, sample size and re-estimation, Primary, secondary and co-primary endpoints, Interim analysis, readout and data cut-off, ADAURA.
- TrialDREAM3R
Shares Futility analysis (stopped for futility), Statistical power, sample size and re-estimation, Data monitoring committee (DSMB, IDMC).
- TermTrial lifecycle: from protocol to label
Shares Data monitoring committee (DSMB, IDMC), Trial protocol and statistical analysis plan, Interim analysis, readout and data cut-off, ADAURA.
- TermBayesian trial design
Shares P-value, Statistical significance (P values, alpha, multiplicity), Seamless, adaptive and Bayesian trial designs, Phase 1, 2 and 3 trials.
- TermResponse-adaptive randomisation
Shares Futility analysis (stopped for futility), Data monitoring committee (DSMB, IDMC), Seamless, adaptive and Bayesian trial designs, Phase 1, 2 and 3 trials.
- TermWhy trials fail: underpowered, wrong endpoint, control arm drift, subgroup fishing, crossover
Shares TROPION-Breast01, Futility analysis (stopped for futility), Statistical significance (P values, alpha, multiplicity), Statistical power, sample size and re-estimation.
- TermConfidence interval
Shares P-value, Statistical significance (P values, alpha, multiplicity), Statistical power, sample size and re-estimation.