Sequence models read DNA or protein letters the way language models read text: ESM predicts protein structure and function, Enformer predicts gene expression from DNA sequence, and Evo 2 models whole genomes with a million-base context.
ESM-2, from Lin and colleagues at Meta, is a protein language model whose embeddings predict atomic-level structure and are reused as gene tokens by cell models such as UCE. Enformer, from Avsec and colleagues at DeepMind, predicts expression and chromatin signals from 200 kilobases of DNA by combining convolutions with transformer attention. Evo 2 is a 40-billion-parameter genomic model on the StripedHyena 2 architecture with a one-million-base context; its zero-shot variant-effect claims are the kind of result that must be checked on clinical variant sets, and at least one independent check found them wanting.
Showing the technology this term belongs to: Nucleotide Transformer (InstaDeep).
Shares Tokenisation (genes, tiles and sequence as tokens), Autoregressive (next-token) modelling, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Tokenisation (genes, tiles and sequence as tokens), Zero-shot prediction, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Zero-shot prediction, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Tokenisation (genes, tiles and sequence as tokens), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Variant effect prediction, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Tokenisation (genes, tiles and sequence as tokens), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Variant effect prediction, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.