A benchmark is a fixed dataset and task on which models are compared, and a leaderboard ranks them; both mislead once the test data has been seen in training or the metric no longer tracks the goal.
A benchmark runs a standard set of tests to assess relative performance (Wikipedia). Cancer AI has few: TCGA tasks, drug-response splits from DrEval and IMPROVE, the Virtual Cell Challenge. Because TCGA slides and profiles are in the pretraining data of most foundation models, downstream TCGA scores are at risk of contamination, and leaderboard rank on a saturated task says little about clinical use. Publishing the exact split and seed is what makes a head-to-head comparison mean anything.
Shares Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA, Data leakage in model evaluation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Reproducibility and negative results, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Reproducibility and negative results, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Virtual cell models and in-silico perturbation screens, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Data leakage in model evaluation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Data leakage in model evaluation, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.