{"slug":"cansim-terms","tag":"cansim-terms","variants":["cansim-terms"],"description":"Glossary terms, gene targets and the hub added on 24 September 2026 from the CanSim terms map (CC BY 4.0): the vocabulary of cancer AI, each record paraphrasing the page it links.","count":140,"kinds":{"term":129,"target":11},"related":[{"slug":"hub","tag":"hub","shared":1}],"records":[{"id":"tumour-board","kind":"term","name":"Tumour board (multidisciplinary team meeting)","route":"/terms/tumour-board/","tldr":"A tumour board is a meeting where doctors from different specialties review one patient's cancer together and agree a treatment plan."},{"id":"grade-vs-stage","kind":"term","name":"Grade versus stage","route":"/terms/grade-vs-stage/","tldr":"Grade describes how abnormal the cancer cells look under the microscope; stage describes how far the cancer has grown and spread."},{"id":"tumour-purity","kind":"term","name":"Tumour purity","route":"/terms/tumour-purity/","tldr":"Tumour purity is the fraction of cells in a sample that are actually cancer cells rather than normal, immune or stromal cells."},{"id":"cell-composition-confound","kind":"term","name":"Cell composition confound (tumour versus stroma and immune cells)","route":"/terms/cell-composition-confound/","tldr":"Because a tumour sample is a mix of cancer, stromal and immune cells, a difference between two samples may only reflect a different mix of cells."},{"id":"intra-tumour-heterogeneity","kind":"term","name":"Intra-tumour heterogeneity","route":"/terms/intra-tumour-heterogeneity/","tldr":"Cells within one tumour can differ in their genes, appearance and behaviour, so a single biopsy may not represent the whole cancer."},{"id":"tumour-evolution","kind":"term","name":"Tumour evolution (somatic evolution)","route":"/terms/tumour-evolution/","tldr":"A tumour changes over time as its cells acquire mutations and the fittest clones take over, which is why cancers relapse and resist treatment."},{"id":"censoring-and-events","kind":"term","name":"Censoring and events in survival data","route":"/terms/censoring-and-events/","tldr":"A patient is censored when the study ends or they leave before the event (death or progression) happens, so we only know they survived at least that long."},{"id":"prognosis-risk-percentile","kind":"term","name":"Risk score and risk percentile","route":"/terms/prognosis-risk-percentile/","tldr":"A risk score is a model's estimate of how likely a bad outcome is; a risk percentile says where that estimate sits compared with a reference group of patients."},{"id":"clinical-covariates","kind":"term","name":"Clinical covariates (age, stage, nodes, treatment flags)","route":"/terms/clinical-covariates/","tldr":"Clinical covariates are the ordinary facts about a patient, such as age, stage, node count and whether they had chemotherapy or radiotherapy, that any prognostic model must beat or build on."},{"id":"endocrine-therapy-resistance","kind":"term","name":"Endocrine therapy and endocrine resistance","route":"/terms/endocrine-therapy-resistance/","tldr":"Endocrine therapy blocks the hormones that feed some breast and prostate cancers; endocrine resistance is when the cancer learns to grow without them."},{"id":"immunotherapy-response","kind":"term","name":"Immunotherapy response and its prediction","route":"/terms/immunotherapy-response/","tldr":"Immunotherapy response means whether a patient's cancer shrinks or stays controlled when the immune system is unleashed; predicting who will respond is one of the hardest open problems in cancer AI."},{"id":"tertiary-lymphoid-structures","kind":"term","name":"Tertiary lymphoid structures (TLS)","route":"/terms/tertiary-lymphoid-structures/","tldr":"Tertiary lymphoid structures are organised clusters of immune cells that form inside tumours and resemble small lymph nodes; their presence is linked to better immunotherapy outcomes."},{"id":"pharmacogenomics-term","kind":"term","name":"Pharmacogenomics","route":"/terms/pharmacogenomics-term/","tldr":"Pharmacogenomics studies how a person's genes, and a tumour's genes, change the way drugs work or cause harm."},{"id":"drug-response-sensitivity","kind":"term","name":"Drug response and drug sensitivity (IC50, AUC)","route":"/terms/drug-response-sensitivity/","tldr":"Drug sensitivity is how strongly a tumour or cell line is held back by a compound; it is summarised by the concentration that halves growth (IC50) or by the area under the whole dose-response curve."},{"id":"cell-lines-as-proxy","kind":"term","name":"Cell lines as a proxy for patients","route":"/terms/cell-lines-as-proxy/","tldr":"Cancer cell lines are tumour cells grown indefinitely in the laboratory; they are convenient for drug testing but only imperfectly stand in for a patient's tumour."},{"id":"co-amplification","kind":"term","name":"Co-amplification and the 17q12 HER2 amplicon","route":"/terms/co-amplification/","tldr":"When a cancer copies the HER2 gene many times over, its neighbours on chromosome 17 are copied with it; those co-amplified genes are passengers, not the driver."},{"id":"copy-number-variation-term","kind":"term","name":"Copy number alteration (CNA)","route":"/terms/copy-number-variation-term/","tldr":"A copy number alteration is a stretch of DNA that a tumour has gained extra copies of or lost, from a single gene to a whole chromosome arm."},{"id":"cancer-drivers-vs-actionable","kind":"term","name":"Actionable genomic biomarkers","route":"/terms/cancer-drivers-vs-actionable/","tldr":"An actionable alteration is a change in a tumour's DNA that maps directly to an approved or investigational drug, so finding it changes what the patient is offered."},{"id":"mrna-protein-concordance","kind":"term","name":"mRNA to protein concordance","route":"/terms/mrna-protein-concordance/","tldr":"mRNA to protein concordance is how well the amount of a gene's messenger RNA tracks the amount of the protein it codes for; for many genes it tracks poorly."},{"id":"post-transcriptional-regulation-term","kind":"term","name":"Post-transcriptional and post-translational regulation","route":"/terms/post-transcriptional-regulation-term/","tldr":"Cells control genes after the messenger RNA is made, by editing, transporting and degrading it and by modifying the finished protein, which is why RNA levels and protein activity can disagree."},{"id":"pathway-activation-state","kind":"term","name":"Pathway activation state (phosphosignalling)","route":"/terms/pathway-activation-state/","tldr":"Whether a signalling pathway is switched on is set by phosphorylation of its proteins, not by how much of them is present, so it has to be measured at the protein level."},{"id":"gene-co-expression","kind":"term","name":"Gene co-expression structure","route":"/terms/gene-co-expression/","tldr":"Genes that rise and fall together across samples form co-expression modules; this correlation structure is what expression models mostly learn."},{"id":"organ-of-origin-signal","kind":"term","name":"Tissue-of-origin signal in tumour data","route":"/terms/organ-of-origin-signal/","tldr":"A tumour's molecular profile is dominated by the organ it came from, so a model can look impressive by recognising the organ and must be judged against a baseline that knows only that."},{"id":"spatial-autocorrelation","kind":"term","name":"Spatial autocorrelation (Moran's I)","route":"/terms/spatial-autocorrelation/","tldr":"Spatial autocorrelation means that nearby spots in a tissue section tend to have similar values; Moran's I is the standard number for how strong that tendency is."},{"id":"variant-effect-prediction","kind":"term","name":"Variant effect prediction","route":"/terms/variant-effect-prediction/","tldr":"Variant effect prediction uses computation to guess whether a DNA change damages a protein or matters clinically, before or instead of laboratory evidence."},{"id":"bulk-rna-seq","kind":"term","name":"Bulk RNA sequencing (RNA-seq) and the full transcriptome","route":"/terms/bulk-rna-seq/","tldr":"Bulk RNA sequencing reads all the messenger RNA in a piece of tumour at once, giving one averaged expression value per gene for the whole sample."},{"id":"tpm-fpkm-counts","kind":"term","name":"TPM, FPKM and raw counts (expression units)","route":"/terms/tpm-fpkm-counts/","tldr":"Raw counts are how many sequencing reads hit each gene; TPM and FPKM rescale them for gene length and sequencing depth so genes and samples can be compared, and log2 TPM is the usual model input."},{"id":"star-salmon","kind":"term","name":"STAR and Salmon (RNA-seq alignment and quantification)","route":"/terms/star-salmon/","tldr":"STAR lines sequencing reads up against the genome and Salmon estimates how much of each transcript is present; they are the two most common first steps of an RNA-seq pipeline."},{"id":"variant-calling","kind":"term","name":"Variant calling","route":"/terms/variant-calling/","tldr":"Variant calling is the computational step that turns raw sequencing reads into a list of the DNA changes present in a tumour."},{"id":"somatic-mutations-wxs-wgs","kind":"term","name":"Somatic mutations from exome and genome sequencing (WXS, WGS)","route":"/terms/somatic-mutations-wxs-wgs/","tldr":"Somatic mutations are the DNA changes a tumour acquired during life; they are read from exome (protein-coding) or whole-genome sequencing of tumour and matched normal tissue."},{"id":"targeted-panel-sequencing","kind":"term","name":"Targeted panel sequencing","route":"/terms/targeted-panel-sequencing/","tldr":"A targeted panel sequences only a chosen set of a few hundred cancer genes, deeply and cheaply, which is what most hospitals run on tumours today."},{"id":"gistic","kind":"term","name":"GISTIC (copy number driver detection)","route":"/terms/gistic/","tldr":"GISTIC scans copy number data across many tumours to find the regions that are amplified or deleted far more often than chance, the likely targets of selection."},{"id":"mutsig","kind":"term","name":"MutSig (significantly mutated gene detection)","route":"/terms/mutsig/","tldr":"MutSig asks which genes are mutated more often than the background mutation rate would predict, separating likely drivers from long or late-replicating genes that collect passengers."},{"id":"batch-effects","kind":"term","name":"Batch effects and harmonisation","route":"/terms/batch-effects/","tldr":"A batch effect is a difference in the data caused by when, where or how samples were processed rather than by biology; if it lines up with the outcome, a model learns the batch instead of the disease."},{"id":"gsea","kind":"term","name":"Gene set enrichment analysis (GSEA and ssGSEA)","route":"/terms/gsea/","tldr":"Gene set enrichment analysis asks whether the genes that changed in an experiment cluster in a known pathway or signature, turning a list of genes into a biological story."},{"id":"single-cell-rna-seq","kind":"term","name":"Single-cell RNA sequencing (scRNA-seq, 10x Chromium)","route":"/terms/single-cell-rna-seq/","tldr":"Single-cell RNA sequencing measures gene expression one cell at a time, revealing the cell types inside a tumour that bulk sequencing averages away."},{"id":"spatial-transcriptomics-platforms","kind":"term","name":"Spatial transcriptomics platforms (Visium HD, Xenium, MERFISH, CosMx, CODEX)","route":"/terms/spatial-transcriptomics-platforms/","tldr":"Spatial platforms measure RNA or protein while keeping each measurement's position on the tissue slide, so cell types and gene programmes can be mapped onto the tumour's architecture."},{"id":"h-and-e-staining","kind":"term","name":"H&E staining (haematoxylin and eosin)","route":"/terms/h-and-e-staining/","tldr":"H&E is the purple and pink stain on almost every pathology slide: haematoxylin colours cell nuclei blue-purple and eosin colours the cytoplasm and connective tissue pink."},{"id":"digital-pathology-wsi","kind":"term","name":"Digital pathology and whole-slide images (WSI)","route":"/terms/digital-pathology-wsi/","tldr":"Digital pathology scans glass slides into gigapixel images that can be viewed, shared and analysed by software, which is what makes AI on pathology possible."},{"id":"tile-patch-encoding","kind":"term","name":"Tile and patch encoding of slides","route":"/terms/tile-patch-encoding/","tldr":"A whole slide is cut into thousands of small square tiles, each tile is turned into a vector by an image model, and the vectors are pooled to describe the slide."},{"id":"magnification","kind":"term","name":"Magnification (20x, 40x) and microns per pixel","route":"/terms/magnification/","tldr":"Slides are scanned at 20x or 40x objective magnification, about 0.5 or 0.25 microns per pixel; a model trained at one scale can misread tissue at another."},{"id":"methylation-arrays","kind":"term","name":"DNA methylation arrays and beta values (450k, EPIC)","route":"/terms/methylation-arrays/","tldr":"Methylation arrays measure, at hundreds of thousands of CpG sites, the fraction of DNA copies carrying a methyl mark; that fraction, between 0 and 1, is the beta value."},{"id":"microarray-expression","kind":"term","name":"Microarray expression data","route":"/terms/microarray-expression/","tldr":"Microarrays measured gene expression by hybridising labelled RNA to spots of known DNA on a chip; they preceded RNA sequencing and still hold the longest-followed cohorts."},{"id":"rppa","kind":"term","name":"Reverse-phase protein array (RPPA)","route":"/terms/rppa/","tldr":"RPPA spots tiny amounts of protein extract from many tumours onto a slide and probes them with antibodies, measuring a few hundred proteins and phosphoproteins in each sample at once."},{"id":"mass-spec-proteome","kind":"term","name":"Deep mass-spectrometry proteome (CPTAC)","route":"/terms/mass-spec-proteome/","tldr":"Mass spectrometry weighs fragments of every protein in a sample to identify and quantify thousands of proteins and their phosphorylation sites, the closest measurement to what a cell is actually doing."},{"id":"ehr-text-pathology-reports","kind":"term","name":"Clinical text: EHR notes and pathology reports","route":"/terms/ehr-text-pathology-reports/","tldr":"Much of what is known about a patient sits in free text, in clinic notes and pathology reports; language models are now used to read it and to pair it with slides."},{"id":"radiology-imaging-modality","kind":"term","name":"Radiology imaging as a data modality (CT, MRI, TCIA)","route":"/terms/radiology-imaging-modality/","tldr":"CT, MRI and PET scans are a data modality in their own right; public archives such as TCIA hold de-identified scans, some linked to TCGA tumours, for training imaging models."},{"id":"os-pfs-time-event","kind":"term","name":"Survival outcomes as time plus event (OS, PFS)","route":"/terms/os-pfs-time-event/","tldr":"A survival outcome is recorded as two numbers per patient: how long they were followed, and whether the event (death, or progression) happened by then."},{"id":"controlled-access-data","kind":"term","name":"Controlled-access genomic data (dbGaP, EGA)","route":"/terms/controlled-access-data/","tldr":"Controlled-access data can identify a person (raw sequence reads, germline variants), so it is held in archives like dbGaP and EGA and released only to approved researchers under an agreement."},{"id":"tcga-tiers","kind":"term","name":"TCGA open versus controlled data tiers","route":"/terms/tcga-tiers/","tldr":"TCGA data comes in two tiers: open files (gene expression, somatic mutations, clinical tables, slide images) anyone can download, and controlled files (raw reads, germline variants) that need dbGaP approval."},{"id":"data-use-agreements","kind":"term","name":"Data use agreements and consent for research use","route":"/terms/data-use-agreements/","tldr":"A data use agreement is the contract a researcher signs to receive controlled data, promising to protect it and use it only as the participants consented."},{"id":"samd","kind":"term","name":"Software as a medical device (SaMD)","route":"/terms/samd/","tldr":"Software as a medical device is software that is itself the medical device, such as a program that predicts a diagnosis or a treatment response, and is regulated like one."},{"id":"research-use-only","kind":"term","name":"Research use only (RUO)","route":"/terms/research-use-only/","tldr":"A research use only label means a test, reagent or model has not been validated or cleared for making decisions about a patient and must not be used that way."},{"id":"analytical-vs-clinical-validation","kind":"term","name":"Analytical versus clinical validation","route":"/terms/analytical-vs-clinical-validation/","tldr":"Analytical validation shows a test measures what it claims, reliably; clinical validation shows the measurement actually predicts the patient outcome it is meant to."},{"id":"hgnc-symbol","kind":"term","name":"HGNC gene symbol","route":"/terms/hgnc-symbol/","tldr":"The HGNC symbol is the one official short name for each human gene (ERBB2, not HER2 or NEU), so that datasets can be joined without guessing."},{"id":"ensembl-gene-id","kind":"term","name":"Ensembl gene ID","route":"/terms/ensembl-gene-id/","tldr":"An Ensembl gene ID such as ENSG00000141736 is a stable machine identifier for a gene that survives symbol changes, so pipelines join on it rather than on names."},{"id":"genome-builds","kind":"term","name":"Genome builds: GRCh38 versus hg19 (GRCh37)","route":"/terms/genome-builds/","tldr":"A genome build is the version of the human reference sequence coordinates are measured against; mixing GRCh38 and the older hg19 puts variants at the wrong positions."},{"id":"hgvs","kind":"term","name":"HGVS variant nomenclature","route":"/terms/hgvs/","tldr":"HGVS is the standard way to write a DNA or protein change, such as EGFR c.2573T>G or p.Leu858Arg, so that the same variant is named the same way everywhere."},{"id":"icd-o-3","kind":"term","name":"ICD-O-3 (International Classification of Diseases for Oncology)","route":"/terms/icd-o-3/","tldr":"ICD-O-3 codes every tumour twice: a topography code for where it arose and a morphology code for what kind of cells it is made of and how it behaves."},{"id":"oncotree-term","kind":"term","name":"OncoTree cancer classification","route":"/terms/oncotree-term/","tldr":"OncoTree is a hierarchy of cancer types from tissue down to subtype, with short codes such as LUAD and BRCA, built for precision oncology and used by cBioPortal and AACR GENIE."},{"id":"ncit","kind":"term","name":"NCI Thesaurus (NCIt)","route":"/terms/ncit/","tldr":"The NCI Thesaurus is the US National Cancer Institute's reference vocabulary of diseases, drugs, anatomy and findings, each with a stable code such as C4872 for breast carcinoma."},{"id":"mondo","kind":"term","name":"Mondo disease ontology","route":"/terms/mondo/","tldr":"Mondo is a unified disease ontology that merges the disease vocabularies (OMIM, Orphanet, NCIt, ICD and others) into one hierarchy with cross-references, so a disease named in one system can be found in the others."},{"id":"hpo","kind":"term","name":"Human Phenotype Ontology (HPO)","route":"/terms/hpo/","tldr":"The Human Phenotype Ontology is a standard vocabulary of clinical features (signs, symptoms, findings) with codes, used to describe patients in a way computers can compare."},{"id":"uberon","kind":"term","name":"Uberon anatomy ontology and the Cell Ontology","route":"/terms/uberon/","tldr":"Uberon names anatomical structures (organs, tissues) and the Cell Ontology names cell types, each with a code, so that a tissue label or a cell-type annotation means the same thing across datasets."},{"id":"units-ontology","kind":"term","name":"Units of measurement ontology (UO)","route":"/terms/units-ontology/","tldr":"The Units of measurement ontology gives every unit (milligram per square metre, months, TPM) a code, so a number in a dataset says what it is a number of."},{"id":"rxnorm","kind":"term","name":"RxNorm drug nomenclature","route":"/terms/rxnorm/","tldr":"RxNorm is the US National Library of Medicine's standard list of drug names and codes, linking brand, generic and dose forms so that the same medicine is recognised across records."},{"id":"ajcc-stage","kind":"term","name":"AJCC stage (TNM staging manual)","route":"/terms/ajcc-stage/","tldr":"AJCC stage is the stage group (I to IV) assigned from the TNM manual published by the American Joint Committee on Cancer; it is the staging most clinical tables record."},{"id":"tcga-barcode","kind":"term","name":"TCGA barcode","route":"/terms/tcga-barcode/","tldr":"A TCGA barcode such as TCGA-A1-A0SB-01A-11R-A144-07 encodes the project, tissue source site, patient, sample type, vial, portion, analyte, plate and centre, and is the key that joins one patient's data files."},{"id":"provenance-fields","kind":"term","name":"Provenance fields for research data","route":"/terms/provenance-fields/","tldr":"Provenance fields record where each data file came from and how it was processed (source, accession, pipeline version, genome build, date, checksum), so a result can be traced and reproduced."},{"id":"desmoplastic-stroma-rich","kind":"term","name":"Stroma-rich and desmoplastic tumours in molecular data","route":"/terms/desmoplastic-stroma-rich/","tldr":"Some cancers, pancreatic cancer above all, are mostly dense scar-like stroma with few tumour cells; their bulk molecular profiles are dominated by that stroma."},{"id":"cancer-ai-vocabulary","kind":"term","name":"Cancer AI vocabulary (CanSim terms map)","route":"/terms/cancer-ai-vocabulary/","tldr":"A hub for the vocabulary of cancer AI: the assays and cohorts models train on, the machine-learning and statistics terms in their papers, the standards their data must follow and the licences that govern reuse."},{"id":"foundation-model","kind":"term","name":"Foundation model","route":"/terms/foundation-model/","tldr":"A foundation model is a large model trained on a vast amount of data without a specific task in mind, then adapted to many downstream uses."},{"id":"self-supervised-pretraining","kind":"term","name":"Self-supervised pretraining (SSL)","route":"/terms/self-supervised-pretraining/","tldr":"Self-supervised pretraining teaches a model from unlabelled data by hiding part of each example and asking it to predict the hidden part, so no expert labels are needed."},{"id":"masked-modelling","kind":"term","name":"Masked autoencoders and masked gene modelling","route":"/terms/masked-modelling/","tldr":"A masked autoencoder hides a random part of the input (image patches, or half the genes in a profile) and learns to reconstruct it from the rest."},{"id":"contrastive-learning","kind":"term","name":"Contrastive learning (InfoNCE)","route":"/terms/contrastive-learning/","tldr":"Contrastive learning trains a model to pull matching pairs (two views of one slide, or one patient's RNA and protein) close together in embedding space and push non-matching pairs apart."},{"id":"transformer-architecture","kind":"term","name":"Transformer and attention","route":"/terms/transformer-architecture/","tldr":"A transformer is a neural network that turns its input into a sequence of tokens and lets every token weigh every other through attention, the architecture behind language models and most new biology models."},{"id":"tokenisation","kind":"term","name":"Tokenisation (genes, tiles and sequence as tokens)","route":"/terms/tokenisation/","tldr":"Tokenisation is how raw input is chopped into the discrete pieces a transformer reads: words into subwords, a slide into tiles, an expression profile into genes ranked or binned by level."},{"id":"embedding","kind":"term","name":"Embedding (learned representation)","route":"/terms/embedding/","tldr":"An embedding is a list of numbers a model produces to stand for an input (a tile, a gene, a patient), placed so that similar inputs land near each other."},{"id":"fine-tuning-vs-frozen","kind":"term","name":"Fine-tuning versus frozen features, and LoRA","route":"/terms/fine-tuning-vs-frozen/","tldr":"Fine-tuning updates a pretrained model's weights on the new task; using it frozen keeps the weights fixed and trains only a small head on its features; LoRA is a cheap middle way that trains small low-rank updates."},{"id":"linear-probe","kind":"term","name":"Linear probe","route":"/terms/linear-probe/","tldr":"A linear probe is a simple linear model (logistic or ridge regression) trained on a frozen model's embeddings to test how much useful information those embeddings hold."},{"id":"transfer-learning","kind":"term","name":"Transfer learning and the low-label regime","route":"/terms/transfer-learning/","tldr":"Transfer learning reuses what a model learned on one task or dataset to do better on a related task with few labels, the situation for almost every cancer outcome."},{"id":"zero-shot","kind":"term","name":"Zero-shot prediction","route":"/terms/zero-shot/","tldr":"Zero-shot prediction applies a model to a task or class it was never trained on, with no task-specific labels at all."},{"id":"autoregressive-modelling","kind":"term","name":"Autoregressive (next-token) modelling","route":"/terms/autoregressive-modelling/","tldr":"An autoregressive model predicts the next element of a sequence from the ones before it; trained on DNA it learns to continue a genome, which is how Evo 2 was built."},{"id":"multimodal-fusion","kind":"term","name":"Multimodal fusion (early, late, modality dropout)","route":"/terms/multimodal-fusion/","tldr":"Multimodal fusion combines several kinds of data about one patient (slides, expression, mutations, clinical variables) in one model; modality dropout randomly hides modalities during training so the model still works when some are missing."},{"id":"abmil","kind":"term","name":"Attention-based multiple-instance learning (ABMIL, CLAM)","route":"/terms/abmil/","tldr":"Attention-based multiple-instance learning gives each tile of a slide a learned weight and sums the weighted tile vectors into one slide vector, so a slide-level label can train the model and the weights show which regions mattered."},{"id":"pathology-foundation-models","kind":"term","name":"Pathology foundation models: UNI, UNI2, Virchow2, CTransPath, CONCH, TITAN","route":"/terms/pathology-foundation-models/","tldr":"Pathology foundation models are image encoders pretrained without labels on millions of slide tiles; UNI and CONCH come from the Mahmood Lab at Harvard, Virchow from Paige, and CTransPath was an early transformer version."},{"id":"single-cell-foundation-models","kind":"term","name":"Single-cell and transcriptome foundation models: UCE, GeneCompass, BulkFormer, BulkRNABert","route":"/terms/single-cell-foundation-models/","tldr":"These are transformer models pretrained on expression profiles: UCE and GeneCompass on tens of millions of single cells, BulkFormer and BulkRNABert on bulk tumour and tissue transcriptomes."},{"id":"genomic-and-protein-language-models","kind":"term","name":"Genomic and protein language models: Evo 2, Enformer, ESM","route":"/terms/genomic-and-protein-language-models/","tldr":"Sequence models read DNA or protein letters the way language models read text: ESM predicts protein structure and function, Enformer predicts gene expression from DNA sequence, and Evo 2 models whole genomes with a million-base context."},{"id":"spagcn","kind":"term","name":"Spatially aware clustering (SpaGCN, KNN smoothing)","route":"/terms/spagcn/","tldr":"Spatially aware clustering groups the spots of a tissue section into domains using both what they express and where they sit, so the domains are contiguous regions rather than scattered spots."},{"id":"virtual-cell-models","kind":"term","name":"Virtual cell models and in-silico perturbation screens","route":"/terms/virtual-cell-models/","tldr":"A virtual cell is a computer model of a cell that predicts what happens when a gene is knocked out or a drug is added, letting researchers run perturbation experiments in software before the laboratory."},{"id":"drug-response-splits","kind":"term","name":"Drug-response data splits: leave-cell-line-out, leave-drug-out, leave-tissue-out","route":"/terms/drug-response-splits/","tldr":"How a drug-response dataset is split decides what a model's accuracy means: hold out cell lines to test personalised prediction, hold out drugs to test drug design, hold out tissues to test repurposing; holding out random pairs only tests imputation."},{"id":"drug-response-baselines","kind":"term","name":"Drug-response baselines and frameworks: mean-drug floor, LightGBM, DrEval, IMPROVE, DeepTTA","route":"/terms/drug-response-baselines/","tldr":"The floor for any drug-response model is predicting each drug's average effect across cell lines; tuned gradient-boosted trees such as LightGBM often tie or beat deep models under honest splits, which is why evaluation frameworks now exist."},{"id":"domain-adaptation","kind":"term","name":"Domain shift and domain adaptation (cell line to patient)","route":"/terms/domain-adaptation/","tldr":"Domain shift is the mismatch between the data a model was trained on and the data it meets later (cell lines versus tumours, one sequencing platform versus another); domain adaptation is the family of methods that try to bridge it."},{"id":"mechanism-of-action-recovery","kind":"term","name":"Mechanism-of-action recovery and known-biology probes","route":"/terms/mechanism-of-action-recovery/","tldr":"A model that claims to have learned biology should, without being told, point at the known mechanism of a known drug or the known marker of a known cell type; recovering such known biology is an orthogonal check that a correlation is not an artefact."},{"id":"cross-validation","kind":"term","name":"Cross-validation and stratified k-fold","route":"/terms/cross-validation/","tldr":"Cross-validation splits the data into k folds, trains on k minus one and tests on the last, rotating so every sample is tested once; stratified folds keep the class balance the same in each fold."},{"id":"train-test-discipline","kind":"term","name":"Train, validation and test split discipline","route":"/terms/train-test-discipline/","tldr":"Data is divided into a training set the model learns from, a validation set used to choose settings, and a test set touched once at the end; using the test set to make choices turns it into a second validation set."},{"id":"data-leakage","kind":"term","name":"Data leakage in model evaluation","route":"/terms/data-leakage/","tldr":"Leakage is when information from the test data reaches the model during training, through a scaler fitted on all samples, a patient split across folds or a benchmark that appeared in the pretraining set, making the model look better than it is."},{"id":"external-validation","kind":"term","name":"External validation","route":"/terms/external-validation/","tldr":"External validation tests a model on patients from a different hospital, country or platform than it was trained on; it is the only evidence that a result travels."},{"id":"concordance-index","kind":"term","name":"C-index (concordance index), Harrell's and Uno's","route":"/terms/concordance-index/","tldr":"The C-index is the fraction of patient pairs in which the model ranked the patient who had the event sooner as higher risk; 0.5 is a coin toss and 1.0 is perfect ranking."},{"id":"calibration","kind":"term","name":"Calibration: reliability diagrams and the Brier score","route":"/terms/calibration/","tldr":"A calibrated model's predicted 30 percent risk really happens about 30 percent of the time; a reliability diagram plots predicted against observed, and the Brier score measures the squared gap."},{"id":"roc-auc","kind":"term","name":"ROC-AUC, PR-AUC and time-dependent AUC","route":"/terms/roc-auc/","tldr":"ROC-AUC is the chance a classifier scores a random positive above a random negative; PR-AUC focuses on the positives and is the better summary when they are rare; time-dependent AUC applies the idea to survival at a chosen horizon."},{"id":"accuracy-f1","kind":"term","name":"Accuracy, macro-F1 and confusion matrices","route":"/terms/accuracy-f1/","tldr":"Accuracy is the share of predictions that were right; macro-F1 averages the F1 score of each class equally, so a model cannot look good by getting only the common cancer types right."},{"id":"time-dependent-auc","kind":"term","name":"Univariate Cox scores and time-dependent metrics","route":"/terms/time-dependent-auc/","tldr":"A univariate Cox model fits one gene at a time against survival and its z-score ranks genes by prognostic strength; the Breslow approximation is how the fit handles patients whose events fall on the same day."},{"id":"bootstrap","kind":"term","name":"Bootstrap resampling","route":"/terms/bootstrap/","tldr":"Bootstrapping refits a model or recomputes a statistic on many resamples of the data drawn with replacement, and the spread of the results is a confidence interval that needs no formula."},{"id":"permutation-test","kind":"term","name":"Permutation test and the Mann-Whitney U test","route":"/terms/permutation-test/","tldr":"A permutation test builds the null distribution by shuffling labels (or spot positions) many times and asks how often chance does as well as the real result; the Mann-Whitney U test compares two groups by ranks without assuming a normal distribution."},{"id":"spearman-correlation","kind":"term","name":"Spearman rank correlation","route":"/terms/spearman-correlation/","tldr":"Spearman's rho measures how well the ranking of one variable matches the ranking of another, from minus one to plus one, ignoring the actual values."},{"id":"logistic-regression-term","kind":"term","name":"Logistic regression and nearest-centroid classifiers","route":"/terms/logistic-regression-term/","tldr":"Logistic regression is the plain linear classifier used as a baseline for subtype prediction; a nearest-centroid classifier instead assigns each sample to the subtype whose average profile it most resembles, which is how PAM50 works."},{"id":"ridge-regression","kind":"term","name":"Ridge regression and the multilayer perceptron","route":"/terms/ridge-regression/","tldr":"Ridge regression is linear regression with a penalty that shrinks coefficients, which keeps it stable when there are more genes than samples; a multilayer perceptron is the simplest neural network, a few fully connected layers."},{"id":"pca","kind":"term","name":"Principal component analysis (PCA) as a feature compressor","route":"/terms/pca/","tldr":"PCA rotates the data onto the directions of greatest variance and keeps the top few, compressing twenty thousand genes into a few hundred numbers before a model sees them."},{"id":"quantile-normalisation","kind":"term","name":"Quantile normalisation, rank transforms and z-scores","route":"/terms/quantile-normalisation/","tldr":"Quantile normalisation forces every sample's expression values onto the same distribution by rank, and a z-score expresses each gene as standard deviations from its mean; both are ways to make data from different platforms comparable."},{"id":"cross-entropy-mse","kind":"term","name":"Loss functions: cross-entropy and mean squared error","route":"/terms/cross-entropy-mse/","tldr":"A loss function is the number training tries to make small: cross-entropy for classification (how surprised the model was by the true class) and mean squared error for regression and reconstruction (how far off the predicted values were)."},{"id":"ablation-study","kind":"term","name":"Ablation study and multi-task heads","route":"/terms/ablation-study/","tldr":"An ablation removes one component or modality at a time and re-measures performance, which is the only way to know what each part contributes; multi-task heads let one shared backbone serve several outputs."},{"id":"ood-detection","kind":"term","name":"Out-of-distribution detection (Mahalanobis guard)","route":"/terms/ood-detection/","tldr":"Out-of-distribution detection flags an input that does not look like anything the model was trained on, so the model can refuse to predict instead of guessing; the Mahalanobis distance from the training cloud is the simplest such guard."},{"id":"uncertainty-quantification","kind":"term","name":"Uncertainty quantification and confidence gates","route":"/terms/uncertainty-quantification/","tldr":"Uncertainty quantification attaches to each prediction an estimate of how much to trust it, so a system can report high confidence, low confidence or refuse."},{"id":"conformal-prediction","kind":"term","name":"Conformal prediction","route":"/terms/conformal-prediction/","tldr":"Conformal prediction wraps any model so that, instead of one answer, it returns a set or interval guaranteed to contain the truth a chosen fraction of the time (say 90 percent), assuming new patients resemble the calibration patients."},{"id":"model-card","kind":"term","name":"Model cards and datasheets for datasets","route":"/terms/model-card/","tldr":"A model card is a short standard document shipped with a model stating what it was trained on, how it was evaluated, its intended use and its known failures; a datasheet does the same for a dataset."},{"id":"leaderboard-benchmark","kind":"term","name":"Benchmarks, leaderboards and contamination","route":"/terms/leaderboard-benchmark/","tldr":"A benchmark is a fixed dataset and task on which models are compared, and a leaderboard ranks them; both mislead once the test data has been seen in training or the metric no longer tracks the goal."},{"id":"reproducibility","kind":"term","name":"Reproducibility and negative results","route":"/terms/reproducibility/","tldr":"A result is reproducible when someone else can get it again from the same data and code; a negative result is an experiment that did not show the hoped-for effect, and reporting it is how a field stops repeating dead ends."},{"id":"pre-registered-experiment","kind":"term","name":"Pre-registered experiment","route":"/terms/pre-registered-experiment/","tldr":"Pre-registration means writing down the hypothesis, the analysis and the success criterion before running the experiment, so the result cannot be quietly redefined afterwards."},{"id":"open-weights","kind":"term","name":"Open weights, open code and gated models","route":"/terms/open-weights/","tldr":"Open weights means a model's trained parameters are published so anyone can run and adapt it; that is less than open source, which also needs the training code and data, and more than gated weights, which require a request and a licence click."},{"id":"open-licences","kind":"term","name":"Open licences: Apache-2.0, MIT and CC BY 4.0","route":"/terms/open-licences/","tldr":"MIT and Apache-2.0 are permissive software licences that let anyone use, change and redistribute code (Apache adds an explicit patent grant); CC BY 4.0 is the equivalent for data and text, requiring only attribution."},{"id":"hugging-face-hub","kind":"term","name":"Hugging Face Hub","route":"/terms/hugging-face-hub/","tldr":"The Hugging Face Hub is the website where most machine-learning models and many datasets are published, with a model card, version history and, for gated models, an access request form."},{"id":"zenodo-doi","kind":"term","name":"Zenodo DOIs for data and code","route":"/terms/zenodo-doi/","tldr":"Zenodo is CERN's free open repository where researchers deposit datasets, code and papers and receive a permanent DOI for each version, so a result can cite exactly the files it used."},{"id":"citation-cff","kind":"term","name":"CITATION.cff (Citation File Format)","route":"/terms/citation-cff/","tldr":"CITATION.cff is a small file in a code repository that says, in a machine-readable way, how to cite the software; GitHub and Zenodo read it to generate citations."},{"id":"anndata-h5ad","kind":"term","name":"AnnData and h5ad files","route":"/terms/anndata-h5ad/","tldr":"AnnData is the Python object, saved as an .h5ad file, that holds an expression matrix together with its per-cell and per-gene annotations; it is the lingua franca of single-cell analysis."},{"id":"hdf5-zarr-parquet","kind":"term","name":"Array and table formats: HDF5, Zarr, OME-Zarr, Parquet","route":"/terms/hdf5-zarr-parquet/","tldr":"HDF5 and Zarr store large numerical arrays in chunks (Zarr is the cloud-friendly one, and OME-Zarr its bio-imaging profile for slide tiles); Parquet stores tables column by column for fast analytical queries."},{"id":"duckdb-catalog","kind":"term","name":"Analytical databases as a data catalogue (DuckDB, PostgreSQL)","route":"/terms/duckdb-catalog/","tldr":"DuckDB is an in-process analytical database that queries Parquet files with SQL; PostgreSQL is the standard client-server database; either can hold the metadata spine that joins patients, specimens and assays."},{"id":"mixed-precision-gpu","kind":"term","name":"GPU training and mixed precision","route":"/terms/mixed-precision-gpu/","tldr":"Foundation models are trained on graphics processors, and mixed precision stores most numbers in 16-bit floats to halve memory and double speed; a laptop-class chip can fine-tune small models but pretraining at scale waits for a data-centre GPU."},{"id":"local-llm-reasoning-layer","kind":"term","name":"Local-only language models over patient data (privacy by architecture)","route":"/terms/local-llm-reasoning-layer/","tldr":"Running an open language model on the same machine as a patient's data, with code that blocks any network call, lets a system explain model outputs in words without the data ever leaving the room."},{"id":"grb7","kind":"target","name":"GRB7","route":"/targets/grb7/","tldr":"GRB7 is an adaptor protein gene that sits next to HER2 on chromosome 17 and is copied along with it in HER2-positive cancers."},{"id":"stard3","kind":"target","name":"STARD3","route":"/targets/stard3/","tldr":"STARD3 (MLN64) moves cholesterol between cell compartments and is co-amplified with HER2 in breast cancer."},{"id":"pgap3","kind":"target","name":"PGAP3","route":"/targets/pgap3/","tldr":"PGAP3 is an enzyme gene in the HER2 amplicon that remodels the lipid anchors of cell-surface proteins."},{"id":"mien1","kind":"target","name":"MIEN1","route":"/targets/mien1/","tldr":"MIEN1 is a small HER2-amplicon gene whose protein promotes cell migration and is strongly raised in breast and prostate cancers."},{"id":"pnmt","kind":"target","name":"PNMT","route":"/targets/pnmt/","tldr":"PNMT makes adrenaline from noradrenaline in the adrenal gland; in breast cancer it matters only because it lies inside the HER2 amplicon."},{"id":"spry4","kind":"target","name":"SPRY4","route":"/targets/spry4/","tldr":"SPRY4 is a feedback brake on growth-factor signalling that the MAPK pathway switches on, so its expression is a read-out of how active that pathway is."},{"id":"phlpp2","kind":"target","name":"PHLPP2","route":"/targets/phlpp2/","tldr":"PHLPP2 is a phosphatase that switches AKT off; it is lost in most colorectal cancers, which leaves the survival pathway on."},{"id":"phlda1","kind":"target","name":"PHLDA1","route":"/targets/phlda1/","tldr":"PHLDA1 is a growth-factor-responsive gene involved in cell death; it is high in benign moles and falls as melanoma progresses."},{"id":"e2f7","kind":"target","name":"E2F7","route":"/targets/e2f7/","tldr":"E2F7 is an unusual member of the E2F family that represses genes rather than activating them, helping cells pause after DNA damage."},{"id":"dclk1","kind":"target","name":"DCLK1","route":"/targets/dclk1/","tldr":"DCLK1 is a kinase gene from neuronal development that also marks a rare stem-like cell population in the gut, studied as a colorectal and pancreatic cancer stem-cell marker."},{"id":"dnajc12","kind":"target","name":"DNAJC12","route":"/targets/dnajc12/","tldr":"DNAJC12 is a co-chaperone gene that is strongly expressed in oestrogen receptor-positive (luminal) breast cancers and is used as a luminal marker."}]}