# Score every model system on how well it predicted real trial results

Source: https://onco.cc/ideas/idea-bio1-model-predictivity-benchmark/  
OnCo record `idea-bio1-model-predictivity-benchmark` (Idea). Data CC BY-NC 4.0, attribute "Data from OnCo (onco.cc)"; commercial use needs a licence.

## TL;DR

No one keeps score of which laboratory models actually predicted what happened in patients. A public scoreboard would show which models to trust.

## Summary

Protein structure prediction improved rapidly once CASP created a blinded, periodic benchmark. An oncology equivalent would take drugs with known but embargoed clinical outcomes, ask model owners (organoids, PDX, chips, in silico) to submit blinded predictions of response rate or ranking, and publish accuracy by model class. Over time this creates evidence for which systems merit regulatory and investment weight.

## Fields

- Kind: Idea
- Last checked: 2026-09-08
- Hypothesis: Blinded benchmarking reveals large and reproducible differences between model classes in predicting clinical response rates, and participation improves accuracy across rounds.
- Rationale: Community benchmarks with held-out truth transformed structural biology and machine learning; oncology model validation is currently self-reported and non-comparable.
- Proposed test: Run a first round with ten agents whose phase 2 results are complete but unpublished or paywalled, and publish accuracy metrics per submitted model class.
- Maturity: speculative
- Actor: data

## Sources

- Bottleneck evidence (Lab models that fail to predict what happens in patients): Wong, Siah & Lo, Estimation of clinical trial success rates (Biostatistics 2019): https://doi.org/10.1093/biostatistics/kxx069

## Connected records

- technologies: [AI-driven drug & target discovery](https://onco.cc/technologies/ai-drug-design/), [Patient-derived organoids](https://onco.cc/technologies/organoids/), [Patient-derived xenografts](https://onco.cc/technologies/pdx-models/)
- institutions: [Broad Institute of MIT and Harvard](https://onco.cc/institutions/broad-institute/)
- bottlenecks: [AI that is built but not validated or deployed](https://onco.cc/bottlenecks/b-ai-validation/), [Lab models that fail to predict what happens in patients](https://onco.cc/bottlenecks/b-preclinical-models/), [Preclinical results do not reproduce](https://onco.cc/bottlenecks/b-reproducibility/)
- key papers: [Estimation of clinical trial success rates and related parameters](https://onco.cc/key-papers/paper-wong-biostatistics/)

---
JSON: https://onco.cc/api/v1/entities/idea-bio1-model-predictivity-benchmark.json