OnCo
ideasIdea

A public benchmark and audit of chatbot answers to cancer questions

Patients now ask AI assistants about their cancer. Test those assistants regularly on real questions, publish the scores, and certify the ones that meet the bar.

Large language models are becoming a primary source of medical information; studies show variable accuracy, occasional dangerous advice and inconsistent sourcing on oncology questions. The proposal is an independent, continuously updated benchmark of patient-style cancer questions (treatment options, side-effects, alternative therapies, prognosis) scored by oncologists and patients for accuracy, safety, sourcing and readability, with public leaderboards and a certification mark for assistants that meet thresholds and disclose limitations.

Hypothesis
Public auditing raises the accuracy and safety of assistant answers to cancer questions across vendors within a year, and certified assistants are preferred by patient organisations.
Rationale
Benchmarks drive model behaviour in AI development; making the oncology benchmark public and patient-facing turns that pressure toward safety.
What would test it
Run the benchmark quarterly on major assistants for a year; measure score trajectories and adoption of the certification by patient organisations and health systems.
Maturity
early clinical
Who has to act
data
Cost to try
Small (under $1M)
Years to first evidence
1
Bottlenecks it attacks

Connected

2top