How a drug-response dataset is split decides what a model's accuracy means: hold out cell lines to test personalised prediction, hold out drugs to test drug design, hold out tissues to test repurposing; holding out random pairs only tests imputation.
Bernett and colleagues' DrEval framework evaluates drug-response prediction models under explicit protocols, holding out cell lines, drugs or tissues so that a model is scored on the generalisation it claims, and finds that many published gains shrink under the harder splits. Training on one screen and testing on another (cross-study validation) is the sternest bar because assays, cell-line panels and summary statistics differ. Random pair splits leak both the cell line and the drug into training and should only support imputation claims.
Shares Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA, Data leakage in model evaluation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Data leakage in model evaluation, External validation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Data leakage in model evaluation, External validation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Drug response and drug sensitivity (IC50, AUC), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Drug response and drug sensitivity (IC50, AUC), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares External validation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.