OnCo
ideasIdea

Sequestered, prospectively collected benchmark datasets that no one can train on

Keep test datasets locked away and collect them going forward, so AI claims are checked on data the developers have never seen and could not have memorised.

Public benchmarks leak into training sets and go stale; retrospective validation flatters models. The proposal is a set of sequestered evaluation datasets for key cancer AI tasks (mammography, lung nodules, prostate biopsy, HER2 scoring, ctDNA calls), collected prospectively from multiple sites and countries, held by a neutral body, with evaluation only via submission of the model or an API, and results published. NIST's face recognition testing and the MICCAI challenge model are precedents.

Hypothesis
Performance on sequestered prospective data will be materially lower than published performance for most models, and public reporting will shift developers toward robust training and honest claims.
Rationale
In face recognition, NIST's sequestered testing became the de facto standard buyers rely on; in medical imaging, external test sets consistently reveal performance drops that published papers omit.
What would test it
Stand up two sequestered benchmarks; evaluate all willing vendors; publish results alongside their published claims; repeat annually to measure whether the gap narrows.
Maturity
early clinical
Who has to act
research
Cost to try
Medium ($1M to $50M)
Years to first evidence
2
Bottlenecks it attacks

Connected

8top