{"entity":{"id":"tokenisation","kind":"term","name":"Tokenisation (genes, tiles and sequence as tokens)","aka":["tokenisation","tokenization","gene tokenisation","gene tokens","rank-based tokenisation","value binning","byte-pair encoding","universal tokenisation","semantic token grounding"],"tldr":"Tokenisation is how raw input is chopped into the discrete pieces a transformer reads: words into subwords, a slide into tiles, an expression profile into genes ranked or binned by level.","summary":"Lexical tokenisation converts text into meaningful tokens of defined categories (Wikipedia); language models use subword schemes such as byte-pair encoding. Biology models must invent their tokens: expression can be tokenised by ranking the genes of a cell (Geneformer), keeping the top N, or binning values; DNA by nucleotides or k-mers. Grounding a gene token in its sequence or protein embedding (as UCE does with ESM2) rather than an arbitrary ID lets a model generalise to unseen genes and species.","asOf":"2026-09-24","wikipedia":"https://en.wikipedia.org/wiki/Lexical_analysis","links":[{"label":"Wikipedia: byte-pair encoding","url":"https://en.wikipedia.org/wiki/Byte-pair_encoding"},{"label":"Wikipedia","url":"https://en.wikipedia.org/wiki/Lexical_analysis"}],"tags":["cansim-terms"],"related":["cancer-ai-vocabulary"],"cancers":[],"sections":[],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":["transformer-architecture","embedding","masked-modelling"],"trials":[],"people":[],"bottlenecks":[],"keyPapers":[],"journals":[],"dependsOn":[],"notes":["Listed in the CanSim terms map 1.0.0 (docs/onco/terms.json, generated 2026-09-24), CC BY 4.0, attribution: CanSim project, an open, public-data-first cancer foundation-model programme; CanSim page path /terms/tokenisation."],"provenance":{"editedBy":"OnCo CanSim terms wave (Wikipedia summaries, standards and project pages, GDC and FDA pages, Europe PMC)","editedOn":"2026-09-24","note":"CanSim terms map 1.0.0 (docs/onco/terms.json, generated 2026-09-24), CC BY 4.0, attribution: CanSim project, an open, public-data-first cancer foundation-model programme"},"category":"Methods and models"},"route":"/terms/tokenisation/","neighbours":{"term":[{"id":"cancer-ai-vocabulary","kind":"term","name":"Cancer AI vocabulary (CanSim terms map)","route":"/terms/cancer-ai-vocabulary/"},{"id":"embedding","kind":"term","name":"Embedding (learned representation)","route":"/terms/embedding/"},{"id":"genomic-and-protein-language-models","kind":"term","name":"Genomic and protein language models: Evo 2, Enformer, ESM","route":"/terms/genomic-and-protein-language-models/"},{"id":"masked-modelling","kind":"term","name":"Masked autoencoders and masked gene modelling","route":"/terms/masked-modelling/"},{"id":"single-cell-foundation-models","kind":"term","name":"Single-cell and transcriptome foundation models: UCE, GeneCompass, BulkFormer, BulkRNABert","route":"/terms/single-cell-foundation-models/"},{"id":"transformer-architecture","kind":"term","name":"Transformer and attention","route":"/terms/transformer-architecture/"}]}}