Model and dataset registry
44 foundation and risk models and 21 datasets from the corpus, with the fields that let you compare them: parameters, modality, training data, whether the weights can be downloaded (27 are open), licence, a reported benchmark and the paper. Figures are the developers' own; blanks are unverified, not zero.
65 models and datasets
| Training data / size | Licence | Reported result / consent | Paper / weights | Trained on / used by | ||||
|---|---|---|---|---|---|---|---|---|
Aidoc CARE (clinical radiology foundation model) Aidoc CARE is a single foundation model behind many FDA-cleared triage alerts in emergency radiology. | none | Not disclosed; underpins FDA-cleared triage products. | none | none | none | none | 2025 | |
AlphaGenome Reads a million letters of DNA at once and predicts how a mutation changes gene regulation, splicing and chromatin. | none | Human and mouse reference genomes with thousands of functional genomics tracks; 1 megabase input at single-base resolution. | API terms (non-commercial preview) | Matched or beat specialist models on 22 of 24 sequence prediction and 24 of 26 variant-effect tasks (DeepMind). | 10.1038/s41586-025-09937-1, weights | none | 2025 | |
Arc Virtual Cell Atlas Arc's growing library of cell data, the fuel for virtual cell models. | none | Hundreds of millions of cells (observational and perturbational) | Open | Harmonised public datasets; per-dataset licences apply. | portal | State | 2025 | |
Atlas (Aignostics, Mayo Clinic, Charité) Atlas is a pathology foundation model trained on 1.2 million slides from two of the world's largest hospitals. | none | 1.2 million slides from Mayo Clinic and Charité across scanners and stains. | none | Top scores on public benchmark tasks reported in the arXiv paper. | arXiv:2501.05409 | none | 2025 | |
BioEmu (Microsoft) Predicts the many shapes a protein moves between, not just one, thousands of times faster than simulation. | none | Molecular dynamics ensembles and experimental folding free energies. | MIT | none | 10.1126/science.adv9817, weights | none | 2025 | |
Cell2Sentence / C2S-Scale (Yale, Google) Turns a cell's gene expression into a sentence so a normal language model can reason about it; a 27-billion-parameter version proposed a cancer immunotherapy idea that was confirmed in the lab. | 27 B | Cell sentences from more than 50 million cells plus biological text, on Gemma-2 backbones up to 27B. | Gemma terms of use (model card) | Predicted that silmitasertib raises antigen presentation under low interferon; validated in vitro (paper). | bioRxiv 10.1101/2025.04.14.648850, weights | CZ CELLxGENE / Human Cell Atlas | 2025 | |
CT-FM (whole-body CT foundation model) A model pretrained on 148,000 CT scans to segment organs and triage findings. | none | 148,000 whole-body CT scans, self-supervised. | none | none | arXiv:2501.09001, weights | The Cancer Imaging Archive, NCI Imaging Data Commons | 2025 | |
ESM3 (EvolutionaryScale) ESM3 is a generative protein model that designed a working fluorescent protein far from any natural sequence. | 98 B | 2.78 billion protein sequences with structure and function tokens. | EvolutionaryScale Cambrian licence (non-commercial) for the open 1.4B model; larger models via API | Generated esmGFP, a functional fluorescent protein 58% identical to its nearest known relative (Science 2025). | 10.1126/science.ads0018, weights | none | 2025 | |
Evo 2 (Arc Institute, NVIDIA) A DNA language model trained on 9.3 trillion bases that can flag cancer-causing BRCA1 variants without being told about them. | 40 B | 9.3 trillion nucleotides across all domains of life (OpenGenome2); 7B and 40B parameter models, 1 million token context. | Apache 2.0 | Zero-shot BRCA1 variant pathogenicity prediction competitive with supervised methods (paper). | bioRxiv 10.1101/2025.02.18.638918, weights | none | 2025 | |
Midnight (kaiko.ai) Midnight is a pathology model that matched the leaders while training on far fewer slides. | none | 12,000 TCGA slides (Midnight-12k) with a combined DINOv2 and high-resolution objective. | MIT (model card) | Leading scores on the eva pathology benchmark suite with a fraction of competitors' training slides (kaiko.ai). | weights or model card | TCGA / NCI Genomic Data Commons, Pathology AI benchmarks | 2025 | |
MUSK (Stanford, vision-language pathology) A model that reads slides and clinical text together to predict who will respond to immunotherapy. | none | 50 million pathology images and 1 billion pathology-related text tokens. | CC BY-NC-ND 4.0 | Better prediction of immunotherapy response and prognosis than single-modality models across cancers (paper). | 10.1038/s41586-024-08378-w, weights | none | 2025 | |
State (Arc Institute perturbation model) Predicts how cells will respond to a drug or gene knockout, trained on over 100 million perturbed cells. | none | State transition model: over 100 million perturbed cells including Tahoe-100M; embedding model: 167 million human cells. | Apache 2.0 (repository) | Reference entry for the Arc Virtual Cell Challenge (2025). | bioRxiv 10.1101/2025.06.26.661135, weights | Tahoe-100M, Arc Virtual Cell Atlas | 2025 | |
Tahoe-100M Tahoe-100M is the biggest single-cell dataset ever released, built to teach AI how cancer cells respond to drugs. | none | 100 million single-cell profiles; 1,100 drugs; 50 cell lines | Open (CC BY) | Cell lines only; CC BY release. | portal | State | 2025 | |
TranscriptFormer and rBio (CZI virtual cell models) CZI's open cross-species cell models and a reasoning model trained on them. | none | 112 million cells across 12 species. | none | none | weights or model card | CZ CELLxGENE / Human Cell Atlas | 2025 | |
AlphaFold 3 Predicts the 3D shape of proteins together with DNA, RNA, small molecules and antibodies, the starting point for much modern drug design. | none | Protein Data Bank structures and distillation sets across proteins, nucleic acids, ligands and ions. | Code CC BY-NC-SA 4.0; weights for non-commercial academic use on request | At least 50% better ligand-interaction accuracy than prior methods on PoseBusters (paper). | 10.1038/s41586-024-07487-w, weights | none | 2024 | |
Boltz-1 / Boltz-2 (MIT, open) Open-source structure models that match AlphaFold 3, with Boltz-2 also predicting how strongly a drug binds. | none | PDB and distillation data; Boltz-2 adds binding-affinity training. | MIT | Boltz-1 matched AlphaFold 3 accuracy; Boltz-2 affinity prediction approached FEP+ at about 1,000 times lower cost (preprint). | bioRxiv 10.1101/2024.11.19.624167, weights | none | 2024 | |
CellFM CellFM is an 800-million-parameter single-cell model trained on 100 million human cells. | 800 M | About 100 million human cells. | none | none | bioRxiv 10.1101/2024.06.04.597369, weights | none | 2024 | |
Chai-1 / Chai-2 Structure and antibody-design models from Chai Discovery, with Chai-2 reporting high zero-shot antibody hit rates. | none | PDB-derived structures; Chai-2 trained for antibody and binder design. | Chai-1 code Apache 2.0; weights under Chai Discovery licence (non-commercial) | Chai-2: about 16% zero-shot binder hit rate across dozens of targets in wet-lab tests (company report). | bioRxiv 10.1101/2024.10.10.615955, weights | none | 2024 | |
CHIEF (Harvard, Yu Lab) A pathology model trained across 19 cancer types that predicts survival and mutations from slides. | none | Pretrained on 15 million tiles, then 60,530 slides; validated on 19,400 slides from 24 hospitals. | AGPL-3.0 (code repository) | Cancer detection accuracy up to 94% and improvement of about 36% over prior deep-learning methods on external cohorts (paper). | 10.1038/s41586-024-07894-z, weights | TCGA / NCI Genomic Data Commons, CPTAC | 2024 | |
Foresight (generative EHR model) A model trained on millions of hospital records that forecasts a patient's next diagnoses. | none | Coded EHR timelines from King's College Hospital and South London and Maudsley; Foresight 2 extended to more than 5 million patients. | none | none | 10.1016/S2589-7500(24)00025-6 | none | 2024 | |
H-optimus (Bioptimus) An open 1.1-billion-parameter pathology model from a French startup, among the strongest on public benchmarks. | 1.1 B | Hundreds of millions of tiles from 500,000 slides (H-optimus-0). | Apache 2.0 | none | weights or model card | none | 2024 | |
Hibou (HistAI) Hibou is a family of open pathology foundation models under a permissive licence. | 307 M | Over 1 million slides (Hibou-B 86M parameters; Hibou-L 307M). | Apache 2.0 | none | arXiv:2406.05074, weights | none | 2024 | |
Med-Gemini and MedLM (Google) Google's medical versions of its Gemini models, able to reason over text, images, and long records. | none | Gemini models fine-tuned on medical text, images and genomics. | none | 91.1% on MedQA (USMLE) in the arXiv report. | arXiv:2404.18416 | none | 2024 | |
MedSAM / SAM-Med3D (segment anything for medicine) Adaptations of Meta's Segment Anything model that outline tumours and organs on any scan with a click. | none | 1.57 million image-mask pairs across 10 imaging modalities and more than 30 cancer types. | Apache 2.0 | Median Dice above specialist models on internal and external validation (paper). | 10.1038/s41467-024-44824-z, weights | The Cancer Imaging Archive | 2024 | |
Merlin (Stanford abdominal CT vision-language model) Merlin is a model trained on 15,000 CT scans with their reports that can find and describe hundreds of findings. | none | 6 million images from 15,331 abdominal CTs with 6 million EHR diagnosis codes and 1.8 million reports. | none | none | arXiv:2406.06512, weights | none | 2024 | |
Nicheformer (spatial single-cell) Nicheformer is a model trained on both dissociated and spatial data so it learns how a cell's neighbourhood shapes it. | none | 110 million cells: 57 million dissociated and 53 million spatially resolved. | none | none | bioRxiv 10.1101/2024.04.15.589472, weights | none | 2024 | |
Nucleotide Transformer (InstaDeep) DNA language models trained on thousands of genomes for variant and regulatory prediction. | 2.5 B | 3,202 human genomes and 850 species genomes. | CC BY-NC-SA 4.0 | none | 10.1038/s41592-024-02523-z, weights | none | 2024 | |
Phenom-2 and Recursion OS A model trained on billions of cell microscopy images to read what a drug or gene knockout does to a cell. | 1.9 B | 8 billion cell-painting images (Recursion). | none | none | none | none | 2024 | |
Phikon / Phikon-v2 (Owkin) Owkin's open pathology models trained on TCGA and its federated hospital network. | 307 M | Phikon: 43 million tiles from TCGA. Phikon-v2: 460 million tiles from 58 million slides of 30 cancer types (PANCAN-XL). | Owkin non-commercial licence (model card) | none | arXiv:2409.09173, weights | TCGA / NCI Genomic Data Commons | 2024 | |
PLUTO (PathAI) PathAI's compact pathology foundation model designed to run across many tasks and resolutions. | none | 195 million tiles from 158,000 slides across more than 50 sites, multi-scale. | none | Powers PathAI AISight and biomarker quantification products; results in the arXiv report. | arXiv:2407.07033 | none | 2024 | |
Prov-GigaPath (Microsoft, Providence) An open pathology model trained on 1.3 billion image tiles from a US health system, modelling whole slides at gigapixel scale. | 1.3 B | 1.3 billion tiles from 171,189 whole slides from more than 30,000 Providence patients. | Prov-GigaPath licence (research use; see model card) | Best on 25 of 26 tasks in the paper, including mutation prediction and cancer subtyping. | 10.1038/s41586-024-07441-w, weights | TCGA / NCI Genomic Data Commons | 2024 | |
scFoundation (BioMap) scFoundation is a 100-million-parameter model trained on 50 million cells, from China's BioMap. | 100 M | Over 50 million cells across all about 19,000 human genes. | none | none | 10.1038/s41592-024-02305-7, weights | none | 2024 | |
scGPT A GPT-style model for single-cell data that predicts cell types, perturbation responses, and gene networks. | none | 33 million cells (CELLxGENE). | MIT | none | 10.1038/s41592-024-02201-0, weights | CZ CELLxGENE / Human Cell Atlas | 2024 | |
Tempus multimodal models Models trained on Tempus's paired genomic, pathology, imaging and outcome data to predict response and prognosis. | none | Tempus clinico-genomic and imaging data; details in company publications. | none | none | none | none | 2024 | |
TITAN (whole-slide multimodal model) TITAN is a model that summarises a whole slide, not just tiles, and can write a draft pathology report. | none | 335,645 whole slides (Mass-340K) with vision-only and vision-language pretraining; slide-level embeddings. | CC BY-NC-ND 4.0 | Slide-level retrieval, prognosis and report generation; outperformed patch-based aggregation in the paper. | arXiv:2411.19666, weights | none | 2024 | |
UNI and CONCH (Harvard, Mahmood Lab) Two open academic pathology models: UNI reads tissue images, CONCH links images with pathology text. | 307 M | UNI: 100 million tiles from 100,000 slides across 20 tissue types (Mass-100K). CONCH: 1.17 million image-caption pairs. | CC BY-NC-ND 4.0 | UNI: best or tied-best on 34 clinical tasks in the paper; CONCH: zero-shot classification and retrieval across 14 tasks. | 10.1038/s41591-024-02857-3, weights | TCGA / NCI Genomic Data Commons, Pathology AI benchmarks | 2024 | |
Virchow / Virchow2 (Paige, MSK) A pathology foundation model trained on millions of slides that can detect cancer and predict biomarkers from an ordinary H&E slide. | 1.9 B | Virchow: 1.5 million slides from MSK (632M parameters). Virchow2 and Virchow2G: 3.1 million slides from MSK and global sites at mixed magnification. | CC BY-NC-ND 4.0 (Virchow2 model card) | Pan-cancer detection AUC 0.95 across 17 cancer types, including rare cancers, in the Virchow paper. | 10.1038/s41591-024-03141-0, weights | none | 2024 | |
AlphaMissense Scored all 71 million possible single-letter protein changes in humans as likely harmful or benign. | none | Fine-tuned from AlphaFold on population frequency data; scored all 71 million possible human missense variants. | CC BY 4.0 (predictions); code CC BY-NC-SA | Classified 89% of missense variants as likely benign or likely pathogenic (Science 2023). | 10.1126/science.adg7492, weights | none | 2023 | |
Geneformer The first widely used transformer trained on millions of single cells, able to predict which genes matter in a disease. | none | Genecorpus-30M (about 30 million cells), later 95 million cells. | Apache 2.0 | none | 10.1038/s41586-023-06139-9, weights | CZ CELLxGENE / Human Cell Atlas | 2023 | |
GenePT Uses text embeddings of gene descriptions from a general LLM to represent cells, and performs surprisingly well. | none | GPT-3.5 embeddings of NCBI gene summaries; no single-cell pretraining. | none | none | bioRxiv 10.1101/2023.10.16.562533, weights | none | 2023 | |
RadFM (generalist radiology foundation model) An open generalist model that answers questions about 2D and 3D scans. | 14 B | MedMD: 16 million 2D and 3D scans with text. | none | none | arXiv:2308.02463, weights | none | 2023 | |
RFdiffusion / RFdiffusion2 and ProteinMPNN (Baker Lab) The tools that design entirely new proteins to bind a chosen target, now used for cancer binders and antibodies. | none | PDB structures; fine-tuned from RoseTTAFold. | BSD (code) | Experimentally validated binders, symmetric assemblies and enzyme active sites across the paper's design tasks. | 10.1038/s41586-023-06415-8, weights | none | 2023 | |
Sybil (MIT/MGH lung cancer risk from CT) Predicts a person's six-year lung cancer risk from one low-dose CT, even when no nodule is visible. | none | Low-dose CT scans from the National Lung Screening Trial; validated at MGH and in Taiwan. | MIT | One-year lung cancer risk AUC 0.86 to 0.94 across validation sets (JCO 2023). | 10.1200/JCO.22.01345, weights | NCI CDAS: NLST and PLCO screening trial data | 2023 | |
Universal Cell Embedding (UCE) Universal Cell Embedding maps any cell from any species into one shared space without retraining. | none | 36 million cells across 8 species, using protein embeddings of genes. | MIT | none | bioRxiv 10.1101/2023.11.28.568918, weights | CZ CELLxGENE / Human Cell Atlas | 2023 | |
Enformer and Borzoi (DeepMind, Calico) Models that predict how DNA sequence controls gene activity, used to interpret non-coding cancer mutations. | none | Enformer: 200 kb input, thousands of epigenomic tracks. Borzoi: 524 kb input, RNA-seq coverage across human and mouse. | Apache 2.0 | none | 10.1038/s41592-021-01252-x, weights | none | 2021 | |
Mirai (MIT breast cancer risk from mammograms) Reads a mammogram to estimate five-year breast cancer risk, consistently across races and devices. | none | Mammograms from MGH; validated across seven hospitals in several countries. | none | Five-year breast cancer risk C-index 0.76 to 0.81 across sites (Sci Transl Med 2021). | 10.1126/scitranslmed.aba4373, weights | none | 2021 | |
All of Us Research Program America's answer to UK Biobank, built for diversity. | none | More than 250,000 whole genomes; target one million participants | Registered tier via Researcher Workbench | Registered and controlled tiers via the Researcher Workbench; emphasis on under-represented groups. | portal | none | 2018 | |
AACR Project GENIE AACR Project GENIE is real-world tumour sequencing data shared by leading cancer centres. | none | More than 200,000 sequenced tumours from 19 institutions | Open (via cBioPortal and Synapse) | De-identified clinico-genomic data via cBioPortal and Synapse; Biopharma Collaborative adds outcomes. | portal | none | 2017 | |
DepMap (Cancer Dependency Map) Which genes each cancer cell line cannot live without. The map of synthetic-lethal targets. | none | Genome-wide CRISPR and drug screens on more than 1,000 cell lines | CC BY 4.0 | Cell lines only; CC BY 4.0. | portal | none | 2017 | |
The Cancer Imaging Archive (TCIA) The public archive of cancer scans that most radiology AI is trained and tested on. | none | More than 200 imaging collections (CT, MRI, PET, pathology) | Mostly CC BY; per-collection | De-identified; mostly CC BY, some collections restricted. | portal | MedSAM / SAM-Med3D, CT-FM | 2011 | |
TCGA / NCI Genomic Data Commons The reference atlas of cancer genomes that most cancer biology since 2008 is built on. | none | More than 11,000 tumours across 33 cancer types (TCGA) plus TARGET and CPTAC | Open (controlled access for raw sequence) | Open for processed data; dbGaP authorisation for raw sequence and germline. | portal | UNI and CONCH, Prov-GigaPath, CHIEF, Phikon / Phikon-v2, Midnight | 2006 | |
UK Biobank UK Biobank is the richest population cohort for linking genes, blood and imaging to who later develops cancer. | none | 500,000 participants; whole genomes, imaging, proteomics, linked cancer registry | Approved researchers | Approved researchers only; broad consent with linkage; no return of individual results. | portal | none | 2006 | |
Cancer Models (PDCM Finder) & HCMI Find a mouse or dish model that matches a tumour type or mutation. | none | Thousands of PDX, organoid and cell-line models across providers | Open | Model metadata open; models by request from providers. | portal | none | none | |
cBioPortal for Cancer Genomics cBioPortal lets you browse the mutations, copy number, and expression of tens of thousands of tumours without writing code. | none | More than 300 cancer genomics studies | Open source (AGPL); data per study | Per-study; public studies are de-identified. | portal | none | none | |
CIViC CIViC is an open, Wikipedia-style database of what cancer mutations mean for treatment. | none | Thousands of curated clinical variant interpretations | CC0 | CC0. | portal | none | none | |
COSMIC (Catalogue of Somatic Mutations in Cancer) COSMIC is the encyclopaedia of cancer mutations, including the definitive list of cancer genes. | none | Curated somatic mutations across millions of samples; Cancer Gene Census of about 750 genes | Free academic; commercial licence | Free academic registration; commercial licence. | portal | none | none | |
CPTAC (Clinical Proteomic Tumor Analysis Consortium) CPTAC measures the proteins, not just the genes, of thousands of tumours. | none | Proteogenomic profiles on more than 1,000 TCGA-linked tumours | Open | Open processed data; controlled raw data through dbGaP. | portal | CHIEF | none | |
CZ CELLxGENE / Human Cell Atlas CELLxGENE and the Human Cell Atlas hold single-cell data from healthy and diseased tissue, browsable and downloadable. | none | Tens of millions of annotated single cells | CC BY | CC BY; donor consent handled by contributing studies. | portal | scGPT, Geneformer, Universal Cell Embedding, Cell2Sentence / C2S-Scale, TranscriptFormer and rBio | none | |
Flatiron Health–Foundation Medicine Clinico-Genomic Database Real-world evidence at scale: what happened to patients with a given genomic profile on a given treatment. | none | More than 100,000 US patients with linked EHR outcomes and genomic profiles | Commercial / research access | De-identified under HIPAA; commercial and research licences. | portal | none | none | |
Human Protein Atlas The Human Protein Atlas shows where in the body each protein is found, which tells you whether an ADC target is safe. | none | Protein expression across tissues, cancers and single cells | CC BY-SA | CC BY-SA 4.0. | portal | none | none | |
NCI CDAS: NLST and PLCO screening trial data NCI CDAS holds the lung screening trial images that trained Sybil and most lung-nodule AI. | none | NLST: 53,000 participants with low-dose CT; PLCO screening data | Data access request | Data access request to NCI CDAS. | portal | Sybil | none | |
NCI Imaging Data Commons (IDC) The Imaging Data Commons is TCIA in the cloud, ready for large-scale model training. | none | Cloud copy of TCIA and other collections in DICOM | Per collection | Per collection; queryable with BigQuery. | portal | CT-FM | none | |
OncoKB Tells you, for a given mutation, whether an approved or investigational drug exists and how strong the evidence is. | none | Precision oncology knowledge base with FDA-recognised levels of evidence | Free for academic use; commercial licence | Free academic licence; commercial licence. | portal | none | none | |
Open Targets Platform Open Targets scores how strongly each gene is linked to each disease, with the evidence behind it. | none | Target-disease evidence for about 60,000 targets | CC0 / open | CC0 and per-source licences. | portal | none | none | |
Pathology AI benchmarks (CAMELYON, PANDA, TCGA slide tasks) The exam papers every pathology model is graded on. | none | CAMELYON16/17, PANDA (11,000 biopsies) and TCGA slide tasks | CC BY-NC-SA and per challenge | CC BY-NC-SA and per-challenge terms. | portal | UNI and CONCH, Midnight | none |
Weights status and licence were read from model cards and repositories on the date in each technology record and change often; check the linked page before relying on them. Benchmark results are the developers' own claims on their own test sets and are not comparable across rows. Prospective clinical validation is the exception, not the rule; see the open question on that point.