OnCo
ideasIdea

A registry of external validation datasets for cancer AI models, with mandatory reporting

Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.

Most published cancer AI models lack external validation, and performance drops sharply on data from other institutions. A curated registry of held-out datasets across modalities (pathology, radiology, genomics) hosted by neutral custodians, with a submission protocol that returns performance metrics without releasing the data, would make external validation routine. Journals and regulators would require a registry validation for any clinical claim.

Hypothesis
Models validated through the registry will show a median performance drop of at least ten percentage points from internal to external validation, and the requirement will improve the external performance of subsequently published models.
Rationale
Held-out evaluation servers (as in machine learning benchmarks) prevent overfitting to the test set; medicine has the datasets but not the shared infrastructure.
What would test it
Establish registry datasets for three tasks (HER2 scoring, lung nodule malignancy, ctDNA variant calling); validate 50 published models; report the distribution of performance changes.
Maturity
early clinical
Who has to act
data
Cost to try
Medium ($1M to $50M)
Years to first evidence
2
Bottlenecks it attacks

Connected

6top

Pages like this

not linked directly; found by shared links