Tokenisation is how raw input is chopped into the discrete pieces a transformer reads: words into subwords, a slide into tiles, an expression profile into genes ranked or binned by level.
Lexical tokenisation converts text into meaningful tokens of defined categories (Wikipedia); language models use subword schemes such as byte-pair encoding. Biology models must invent their tokens: expression can be tokenised by ranking the genes of a cell (Geneformer), keeping the top N, or binning values; DNA by nucleotides or k-mers. Grounding a gene token in its sequence or protein embedding (as UCE does with ESM2) rather than an arbitrary ID lets a model generalise to unseen genes and species.
Shares Genomic and protein language models: Evo 2, Enformer, ESM, Embedding (learned representation), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Genomic and protein language models: Evo 2, Enformer, ESM, Transformer and attention, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Embedding (learned representation), Single-cell and transcriptome foundation models: UCE, GeneCompass, BulkFormer, BulkRNABert, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Masked autoencoders and masked gene modelling, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Masked autoencoders and masked gene modelling, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Transformer and attention, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Embedding (learned representation), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Masked autoencoders and masked gene modelling, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.