OnCo

Model and dataset registry

44 foundation and risk models and 21 datasets from the corpus, with the fields that let you compare them: parameters, modality, training data, whether the weights can be downloaded (27 are open), licence, a reported benchmark and the paper. Figures are the developers' own; blanks are unverified, not zero.

65 models and datasets
Aidoc CARE (clinical radiology foundation model)
Aidoc CARE is a single foundation model behind many FDA-cleared triage alerts in emergency radiology.
none
AlphaGenome
Reads a million letters of DNA at once and predicts how a mutation changes gene regulation, splicing and chromatin.
none
Arc Virtual Cell Atlas
Arc's growing library of cell data, the fuel for virtual cell models.
none
Atlas (Aignostics, Mayo Clinic, Charité)
Atlas is a pathology foundation model trained on 1.2 million slides from two of the world's largest hospitals.
none
BioEmu (Microsoft)
Predicts the many shapes a protein moves between, not just one, thousands of times faster than simulation.
none
Cell2Sentence / C2S-Scale (Yale, Google)
Turns a cell's gene expression into a sentence so a normal language model can reason about it; a 27-billion-parameter version proposed a cancer immunotherapy idea that was confirmed in the lab.
27 B
CT-FM (whole-body CT foundation model)
A model pretrained on 148,000 CT scans to segment organs and triage findings.
none
ESM3 (EvolutionaryScale)
ESM3 is a generative protein model that designed a working fluorescent protein far from any natural sequence.
98 B
Evo 2 (Arc Institute, NVIDIA)
A DNA language model trained on 9.3 trillion bases that can flag cancer-causing BRCA1 variants without being told about them.
40 B
Midnight (kaiko.ai)
Midnight is a pathology model that matched the leaders while training on far fewer slides.
none
MUSK (Stanford, vision-language pathology)
A model that reads slides and clinical text together to predict who will respond to immunotherapy.
none
State (Arc Institute perturbation model)
Predicts how cells will respond to a drug or gene knockout, trained on over 100 million perturbed cells.
none
Tahoe-100M
Tahoe-100M is the biggest single-cell dataset ever released, built to teach AI how cancer cells respond to drugs.
none
TranscriptFormer and rBio (CZI virtual cell models)
CZI's open cross-species cell models and a reasoning model trained on them.
none
AlphaFold 3
Predicts the 3D shape of proteins together with DNA, RNA, small molecules and antibodies, the starting point for much modern drug design.
none
Boltz-1 / Boltz-2 (MIT, open)
Open-source structure models that match AlphaFold 3, with Boltz-2 also predicting how strongly a drug binds.
none
CellFM
CellFM is an 800-million-parameter single-cell model trained on 100 million human cells.
800 M
Chai-1 / Chai-2
Structure and antibody-design models from Chai Discovery, with Chai-2 reporting high zero-shot antibody hit rates.
none
CHIEF (Harvard, Yu Lab)
A pathology model trained across 19 cancer types that predicts survival and mutations from slides.
none
Foresight (generative EHR model)
A model trained on millions of hospital records that forecasts a patient's next diagnoses.
none
H-optimus (Bioptimus)
An open 1.1-billion-parameter pathology model from a French startup, among the strongest on public benchmarks.
1.1 B
Hibou (HistAI)
Hibou is a family of open pathology foundation models under a permissive licence.
307 M
Med-Gemini and MedLM (Google)
Google's medical versions of its Gemini models, able to reason over text, images, and long records.
none
MedSAM / SAM-Med3D (segment anything for medicine)
Adaptations of Meta's Segment Anything model that outline tumours and organs on any scan with a click.
none
Merlin (Stanford abdominal CT vision-language model)
Merlin is a model trained on 15,000 CT scans with their reports that can find and describe hundreds of findings.
none
Nicheformer (spatial single-cell)
Nicheformer is a model trained on both dissociated and spatial data so it learns how a cell's neighbourhood shapes it.
none
Nucleotide Transformer (InstaDeep)
DNA language models trained on thousands of genomes for variant and regulatory prediction.
2.5 B
Phenom-2 and Recursion OS
A model trained on billions of cell microscopy images to read what a drug or gene knockout does to a cell.
1.9 B
Phikon / Phikon-v2 (Owkin)
Owkin's open pathology models trained on TCGA and its federated hospital network.
307 M
PLUTO (PathAI)
PathAI's compact pathology foundation model designed to run across many tasks and resolutions.
none
Prov-GigaPath (Microsoft, Providence)
An open pathology model trained on 1.3 billion image tiles from a US health system, modelling whole slides at gigapixel scale.
1.3 B
scFoundation (BioMap)
scFoundation is a 100-million-parameter model trained on 50 million cells, from China's BioMap.
100 M
scGPT
A GPT-style model for single-cell data that predicts cell types, perturbation responses, and gene networks.
none
Tempus multimodal models
Models trained on Tempus's paired genomic, pathology, imaging and outcome data to predict response and prognosis.
none
TITAN (whole-slide multimodal model)
TITAN is a model that summarises a whole slide, not just tiles, and can write a draft pathology report.
none
UNI and CONCH (Harvard, Mahmood Lab)
Two open academic pathology models: UNI reads tissue images, CONCH links images with pathology text.
307 M
Virchow / Virchow2 (Paige, MSK)
A pathology foundation model trained on millions of slides that can detect cancer and predict biomarkers from an ordinary H&E slide.
1.9 B
AlphaMissense
Scored all 71 million possible single-letter protein changes in humans as likely harmful or benign.
none
Geneformer
The first widely used transformer trained on millions of single cells, able to predict which genes matter in a disease.
none
GenePT
Uses text embeddings of gene descriptions from a general LLM to represent cells, and performs surprisingly well.
none
RadFM (generalist radiology foundation model)
An open generalist model that answers questions about 2D and 3D scans.
14 B
RFdiffusion / RFdiffusion2 and ProteinMPNN (Baker Lab)
The tools that design entirely new proteins to bind a chosen target, now used for cancer binders and antibodies.
none
Sybil (MIT/MGH lung cancer risk from CT)
Predicts a person's six-year lung cancer risk from one low-dose CT, even when no nodule is visible.
none
Universal Cell Embedding (UCE)
Universal Cell Embedding maps any cell from any species into one shared space without retraining.
none
Enformer and Borzoi (DeepMind, Calico)
Models that predict how DNA sequence controls gene activity, used to interpret non-coding cancer mutations.
none
Mirai (MIT breast cancer risk from mammograms)
Reads a mammogram to estimate five-year breast cancer risk, consistently across races and devices.
none
All of Us Research Program
America's answer to UK Biobank, built for diversity.
none
AACR Project GENIE
AACR Project GENIE is real-world tumour sequencing data shared by leading cancer centres.
none
DepMap (Cancer Dependency Map)
Which genes each cancer cell line cannot live without. The map of synthetic-lethal targets.
none
The Cancer Imaging Archive (TCIA)
The public archive of cancer scans that most radiology AI is trained and tested on.
none
TCGA / NCI Genomic Data Commons
The reference atlas of cancer genomes that most cancer biology since 2008 is built on.
none
UK Biobank
UK Biobank is the richest population cohort for linking genes, blood and imaging to who later develops cancer.
none
Cancer Models (PDCM Finder) & HCMI
Find a mouse or dish model that matches a tumour type or mutation.
none
cBioPortal for Cancer Genomics
cBioPortal lets you browse the mutations, copy number, and expression of tens of thousands of tumours without writing code.
none
CIViC
CIViC is an open, Wikipedia-style database of what cancer mutations mean for treatment.
none
COSMIC (Catalogue of Somatic Mutations in Cancer)
COSMIC is the encyclopaedia of cancer mutations, including the definitive list of cancer genes.
none
CPTAC (Clinical Proteomic Tumor Analysis Consortium)
CPTAC measures the proteins, not just the genes, of thousands of tumours.
none
CZ CELLxGENE / Human Cell Atlas
CELLxGENE and the Human Cell Atlas hold single-cell data from healthy and diseased tissue, browsable and downloadable.
none
Flatiron Health–Foundation Medicine Clinico-Genomic Database
Real-world evidence at scale: what happened to patients with a given genomic profile on a given treatment.
none
Human Protein Atlas
The Human Protein Atlas shows where in the body each protein is found, which tells you whether an ADC target is safe.
none
NCI CDAS: NLST and PLCO screening trial data
NCI CDAS holds the lung screening trial images that trained Sybil and most lung-nodule AI.
none
NCI Imaging Data Commons (IDC)
The Imaging Data Commons is TCIA in the cloud, ready for large-scale model training.
none
OncoKB
Tells you, for a given mutation, whether an approved or investigational drug exists and how strong the evidence is.
none
Open Targets Platform
Open Targets scores how strongly each gene is linked to each disease, with the evidence behind it.
none
Pathology AI benchmarks (CAMELYON, PANDA, TCGA slide tasks)
The exam papers every pathology model is graded on.
none

Weights status and licence were read from model cards and repositories on the date in each technology record and change often; check the linked page before relying on them. Benchmark results are the developers' own claims on their own test sets and are not comparable across rows. Prospective clinical validation is the exception, not the rule; see the open question on that point.