{"entity":{"id":"leaderboard-benchmark","kind":"term","name":"Benchmarks, leaderboards and contamination","aka":["benchmark","benchmarks","benchmark dataset","leaderboard","leaderboards","benchmark contamination","public benchmark","head-to-head comparison"],"tldr":"A benchmark is a fixed dataset and task on which models are compared, and a leaderboard ranks them; both mislead once the test data has been seen in training or the metric no longer tracks the goal.","summary":"A benchmark runs a standard set of tests to assess relative performance (Wikipedia). Cancer AI has few: TCGA tasks, drug-response splits from DrEval and IMPROVE, the Virtual Cell Challenge. Because TCGA slides and profiles are in the pretraining data of most foundation models, downstream TCGA scores are at risk of contamination, and leaderboard rank on a saturated task says little about clinical use. Publishing the exact split and seed is what makes a head-to-head comparison mean anything.","asOf":"2026-09-24","wikipedia":"https://en.wikipedia.org/wiki/Benchmark_(computing)","links":[{"label":"Wikipedia","url":"https://en.wikipedia.org/wiki/Benchmark_(computing)"}],"tags":["cansim-terms"],"related":["cancer-ai-vocabulary"],"cancers":[],"sections":[],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":["data-leakage","drug-response-baselines","virtual-cell-models","reproducibility"],"trials":[],"people":[],"bottlenecks":[],"keyPapers":[],"journals":[],"dependsOn":[],"notes":["Listed in the CanSim terms map 1.0.0 (docs/onco/terms.json, generated 2026-09-24), CC BY 4.0, attribution: CanSim project, an open, public-data-first cancer foundation-model programme; CanSim page path /terms/leaderboard."],"provenance":{"editedBy":"OnCo CanSim terms wave (Wikipedia summaries, standards and project pages, GDC and FDA pages, Europe PMC)","editedOn":"2026-09-24","note":"CanSim terms map 1.0.0 (docs/onco/terms.json, generated 2026-09-24), CC BY 4.0, attribution: CanSim project, an open, public-data-first cancer foundation-model programme"},"category":"Methods and models"},"route":"/terms/leaderboard-benchmark/","neighbours":{"term":[{"id":"cancer-ai-vocabulary","kind":"term","name":"Cancer AI vocabulary (CanSim terms map)","route":"/terms/cancer-ai-vocabulary/"},{"id":"data-leakage","kind":"term","name":"Data leakage in model evaluation","route":"/terms/data-leakage/"},{"id":"drug-response-baselines","kind":"term","name":"Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA","route":"/terms/drug-response-baselines/"},{"id":"reproducibility","kind":"term","name":"Reproducibility and negative results","route":"/terms/reproducibility/"},{"id":"virtual-cell-models","kind":"term","name":"Virtual cell models and in-silico perturbation screens","route":"/terms/virtual-cell-models/"}]}}