A calibrated model's predicted 30 percent risk really happens about 30 percent of the time; a reliability diagram plots predicted against observed, and the Brier score measures the squared gap.
In statistics, calibration in this sense refers to whether predicted probabilities match observed frequencies (Wikipedia). The Brier score is a strictly proper scoring rule equal to the mean squared error of predicted probabilities (Wikipedia), and Graf and colleagues extended it to censored survival data as the integrated Brier score over time. Discrimination (C-index, AUC) and calibration are independent: a model can rank patients well and still overstate every risk, and calibration is the property that breaks first under domain shift, which is why a model's reliability on a new platform has to be re-measured.
Shares Conformal prediction, Uncertainty quantification and confidence gates, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares ROC-AUC, PR-AUC and time-dependent AUC, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Analytical versus clinical validation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Uncertainty quantification and confidence gates, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Analytical versus clinical validation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Uncertainty quantification and confidence gates, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares C-index (concordance index), Harrell's and Uno's, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares C-index (concordance index), Harrell's and Uno's, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.