{"entity":{"id":"data-leakage","kind":"term","name":"Data leakage in model evaluation","aka":["data leakage","leakage","train-test leakage","cross-validation leakage","leaky preprocessing","fit on train fold only","benchmark contamination","test set contamination"],"tldr":"Leakage is when information from the test data reaches the model during training, through a scaler fitted on all samples, a patient split across folds or a benchmark that appeared in the pretraining set, making the model look better than it is.","summary":"In statistics and machine learning, leakage is the use of information during training that would not be available at prediction time, which produces overly optimistic performance estimates (Wikipedia). Common cancer-data forms are normalising or selecting genes on the full cohort before splitting, duplicated patients across TCGA and cBioPortal studies, and benchmark contamination, where a foundation model's pretraining corpus contained the public test set (TCGA slides or profiles), so that its downstream score is partly recall.","asOf":"2026-09-24","wikipedia":"https://en.wikipedia.org/wiki/Leakage_(machine_learning)","links":[{"label":"Wikipedia","url":"https://en.wikipedia.org/wiki/Leakage_(machine_learning)"}],"tags":["cansim-terms"],"related":["cancer-ai-vocabulary"],"cancers":[],"sections":[],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":["cross-validation","train-test-discipline","batch-effects","leaderboard-benchmark"],"trials":[],"people":[],"bottlenecks":[],"keyPapers":[],"journals":[],"dependsOn":[],"notes":["Listed in the CanSim terms map 1.0.0 (docs/onco/terms.json, generated 2026-09-24), CC BY 4.0, attribution: CanSim project, an open, public-data-first cancer foundation-model programme; CanSim page path /terms/cross-validation-leakage."],"provenance":{"editedBy":"OnCo CanSim terms wave (Wikipedia summaries, standards and project pages, GDC and FDA pages, Europe PMC)","editedOn":"2026-09-24","note":"CanSim terms map 1.0.0 (docs/onco/terms.json, generated 2026-09-24), CC BY 4.0, attribution: CanSim project, an open, public-data-first cancer foundation-model programme"},"category":"Methods and models"},"route":"/terms/data-leakage/","neighbours":{"term":[{"id":"batch-effects","kind":"term","name":"Batch effects and harmonisation","route":"/terms/batch-effects/"},{"id":"leaderboard-benchmark","kind":"term","name":"Benchmarks, leaderboards and contamination","route":"/terms/leaderboard-benchmark/"},{"id":"cancer-ai-vocabulary","kind":"term","name":"Cancer AI vocabulary (CanSim terms map)","route":"/terms/cancer-ai-vocabulary/"},{"id":"cross-validation","kind":"term","name":"Cross-validation and stratified k-fold","route":"/terms/cross-validation/"},{"id":"drug-response-splits","kind":"term","name":"Drug-response data splits: leave-cell-line-out, leave-drug-out, leave-tissue-out","route":"/terms/drug-response-splits/"},{"id":"pca","kind":"term","name":"Principal component analysis (PCA) as a feature compressor","route":"/terms/pca/"},{"id":"train-test-discipline","kind":"term","name":"Train, validation and test split discipline","route":"/terms/train-test-discipline/"}]}}