OnCo
ideasIdea

Public gold-standard datasets for validating every cancer biomarker test

Anyone building a new test for HER2, PD-L1 or tumour DNA should be able to check it against the same public reference set. Today each developer validates on private data nobody can inspect.

Analytical validation of biomarker tests uses proprietary sample sets, making performance claims incomparable. A public repository of consensus-annotated cases per biomarker (whole slide images with multi-pathologist scores, sequencing data with orthogonally confirmed variants, plasma samples with defined variant allele fractions) would allow head-to-head benchmarking and independent verification of any test or algorithm, including AI-based ones.

Hypothesis
Tests benchmarked on the public sets will show performance differences hidden by developer-reported validations, and regulators will begin requiring performance on the reference sets within three years.
Rationale
ImageNet-style benchmarks drove progress and honesty in machine learning; the Genome in a Bottle reference genomes did the same for variant calling.
What would test it
Release reference sets for HER2, PD-L1 and ctDNA; invite all commercial and academic tests to report; publish a leaderboard.
Maturity
early clinical
Who has to act
data
Cost to try
Medium ($1M to $50M)
Years to first evidence
2
Bottlenecks it attacks

Connected

9top

Pages like this

not linked directly; found by shared links