ideasIdea
Mandatory subgroup performance reporting for cancer AI
Every AI tool would have to report how well it works for women and men, different ethnic groups, ages, scanner types and hospitals, not just an overall score.
Cancer AI is often validated on populations that do not match deployment populations; performance gaps by skin tone (dermatology), breast density, ethnicity and scanner vendor are documented. The proposal requires, for clearance and in the model registry, performance reporting across a standard set of subgroups with minimum sample sizes and confidence intervals, and labelling restrictions where performance is unknown or inadequate.
Hypothesis
Mandatory subgroup reporting will reveal clinically meaningful performance disparities in a substantial share of cleared cancer AI and lead to label restrictions or retraining for those models.
Rationale
Pulse oximetry's racial bias went unrecognised for decades because subgroup performance was not required; AI will repeat this at scale unless reporting is mandatory.
What would test it
Evaluate ten cleared cancer AI devices on the standard subgroup set using sequestered data; publish disparities; track subsequent label changes.
Maturity
speculative
Who has to act
regulator
Cost to try
Small (under $1M)
Years to first evidence
2
Bottlenecks it attacks
- AI that is built but not validated or deployed · Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.
- Trials do not represent the people who get cancer · Older, Black, Hispanic, Asian, rural, poor and multimorbid patients are under-represented, so results may not apply to them.