OnCo
ideasIdea

A standard for monitoring AI performance drift with pause thresholds

Set common rules for how hospitals check that an AI tool still works as the scanners, patients and practices around it change, and when it must be switched off.

Model performance shifts when scanners, staining protocols, populations or clinical practice change. Few deployments monitor this. The proposal is a technical standard: a per-site reference dataset re-scored monthly, input distribution monitoring, calibration and subgroup checks, pre-specified thresholds for alert and pause, and a documented recalibration or retraining pathway, integrated with the vendor's change control plan and reported to the registry.

Hypothesis
Sites following the drift standard will detect degradation months earlier than unmonitored sites and avoid patient harm events attributable to silent drift.
Rationale
Industrial machine learning monitors drift as routine engineering practice; healthcare deployments largely do not, and the few audits done have found drift within a year or two of deployment.
What would test it
Implement the standard at ten sites running the same pathology or radiology model; compare detected drift events and time to detection with ten unmonitored sites over two years.
Maturity
early clinical
Who has to act
engineering
Cost to try
Small (under $1M)
Years to first evidence
2
Bottlenecks it attacks

Connected

5top