# A monthly-updated benchmark for AI answers to oncology questions with citation accuracy

Source: https://onco.cc/ideas/idea-data-living-llm-oncology-benchmark/  
OnCo record `idea-data-living-llm-oncology-benchmark` (Idea). Data CC BY-NC 4.0, attribute "Data from OnCo (onco.cc)"; commercial use needs a licence.

## TL;DR

Test the large language models doctors and patients are already using against a continually refreshed set of cancer questions, scoring not just correct answers but whether the sources they cite are real and support the claim.

## Summary

Clinicians and patients use general-purpose language models for oncology questions; evaluations are static, quickly outdated and rarely check citations. The proposal is a living benchmark: new questions each month drawn from recent practice changes, expert-graded answers, and scoring of citation validity and support, with public leaderboards and per-cancer breakdowns, run by an independent academic consortium.

## Fields

- Kind: Idea
- Last checked: 2026-09-08
- Hypothesis: Public, living evaluation will drive measurable improvement in citation accuracy and currency of oncology answers across models within a year, and will identify failure modes (outdated standards, hallucinated trials) that static benchmarks miss.
- Rationale: Public benchmarks have driven progress in every area of machine learning; medical question benchmarks exist but are static and do not test currency, which is the key oncology failure.
- Proposed test: Run the benchmark monthly for a year on the major models; publish trends; check whether model releases show improvement on the citation and currency metrics.
- Maturity: early-clinical
- Actor: research

## Sources

- Bottleneck evidence (Knowledge reaches practice too slowly): Morris, Wooding & Grant, The answer is 17 years, what is the question (JRSM 2011): https://doi.org/10.1258/jrsm.2011.110180

## Connected records

- fronts: [AI & Computation](https://onco.cc/fronts/ai-computation/)
- bottlenecks: [AI that is built but not validated or deployed](https://onco.cc/bottlenecks/b-ai-validation/), [Knowledge reaches practice too slowly](https://onco.cc/bottlenecks/b-knowledge-diffusion/), [Misinformation and unproven therapies](https://onco.cc/bottlenecks/b-misinformation/)
- key papers: [The answer is 17 years, what is the question: understanding time lags in translational research](https://onco.cc/key-papers/paper-morris-j-r-soc-med/)

---
JSON: https://onco.cc/api/v1/entities/idea-data-living-llm-oncology-benchmark.json