Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell
The attempt to build a computer model of a cell good enough to predict what a drug or mutation will do before anyone runs the experiment.
A virtual cell would let researchers test thousands of drug ideas in silico and personalise treatment from a patient's own tumour profile. The field moved from static atlases to perturbation-trained models in five years; the honest status is that current models generalise poorly to unseen contexts and barely beat simple baselines on rigorous benchmarks, while data generation has begun to scale to the size the problem needs.
- 2008-2018historic
Atlases and bulk omics
TCGA catalogues the genomes of 11,000 tumours; single-cell RNA-seq matures; the Human Cell Atlas begins. Models are statistical, per-dataset, and descriptive.
- 2019-2022historic
Pooled perturbation screens meet single cells
Perturb-seq and genome-wide CRISPR screens (DepMap) give causal training data; GEARS shows graph models can predict some unseen knockouts.
- 2023-2024current
First single-cell foundation models
Geneformer, scGPT, UCE, scFoundation and others pretrain on tens of millions of cells. Benchmarks reveal that perturbation prediction often does not beat linear or mean baselines, forcing better evaluation.
- 2025-2026current
Data at scale and context-aware models
Tahoe-100M (100M cells, 1,100 drugs, 50 cancer lines), Arc's Virtual Cell Atlas and Challenge, State trained on 100M+ perturbed cells, C2S-Scale's lab-validated hypothesis, CZI's cross-species models. The problem becomes one of held-out generalisation across cell contexts.
- 2027-2029emerging
Patient-derived contexts and spatial niches
Models trained on perturbations in patient-derived organoids and spatial data (tumour niches, immune contexts) rather than cell lines alone; coupling with structure models for mechanism; prospective use to rank drug combinations for organoid confirmation.
- 2030+speculative
Speculative: in silico trials and digital twins
A tumour's multi-omic profile seeds a patient-specific virtual cell population; treatment sequences are simulated before the first cycle; models are updated from ctDNA and imaging during care. Requires validation standards that do not yet exist.
Probability ranges are named estimates that the claim is borne out on roughly a five-year horizon. They are meant to be argued with: propose a revision with your name and reasoning via a pull request to src/data/confidence.ts.
Story
topAtlases and bulk omics
TCGA catalogues the genomes of 11,000 tumours; single-cell RNA-seq matures; the Human Cell Atlas begins. Models are statistical, per-dataset, and descriptive.
The reference atlas of cancer genomes that most cancer biology since 2008 is built on.
CELLxGENE and the Human Cell Atlas hold single-cell data from healthy and diseased tissue, browsable and downloadable.
Reading the genes of each individual cell, and mapping where each cell sits in the tumour.
Pooled perturbation screens meet single cells
Perturb-seq and genome-wide CRISPR screens (DepMap) give causal training data; GEARS shows graph models can predict some unseen knockouts.
Knocking out every gene one at a time in cancer cells to find which ones they cannot live without.
Which genes each cancer cell line cannot live without. The map of synthetic-lethal targets.
GEARS is a graph model predicting the effect of gene knockouts; the perturbation benchmarks around it showed how hard the problem is.
First single-cell foundation models
Geneformer, scGPT, UCE, scFoundation and others pretrain on tens of millions of cells. Benchmarks reveal that perturbation prediction often does not beat linear or mean baselines, forcing better evaluation.
The first widely used transformer trained on millions of single cells, able to predict which genes matter in a disease.
A GPT-style model for single-cell data that predicts cell types, perturbation responses, and gene networks.
Universal Cell Embedding maps any cell from any species into one shared space without retraining.
scFoundation is a 100-million-parameter model trained on 50 million cells, from China's BioMap.
GEARS is a graph model predicting the effect of gene knockouts; the perturbation benchmarks around it showed how hard the problem is.
Data at scale and context-aware models
Tahoe-100M (100M cells, 1,100 drugs, 50 cancer lines), Arc's Virtual Cell Atlas and Challenge, State trained on 100M+ perturbed cells, C2S-Scale's lab-validated hypothesis, CZI's cross-species models. The problem becomes one of held-out generalisation across cell contexts.
Tahoe-100M is the biggest single-cell dataset ever released, built to teach AI how cancer cells respond to drugs.
Arc's growing library of cell data, the fuel for virtual cell models.
Predicts how cells will respond to a drug or gene knockout, trained on over 100 million perturbed cells.
Turns a cell's gene expression into a sentence so a normal language model can reason about it; a 27-billion-parameter version proposed a cancer immunotherapy idea that was confirmed in the lab.
CZI's open cross-species cell models and a reasoning model trained on them.
Nicheformer is a model trained on both dissociated and spatial data so it learns how a cell's neighbourhood shapes it.
Patient-derived contexts and spatial niches
Models trained on perturbations in patient-derived organoids and spatial data (tumour niches, immune contexts) rather than cell lines alone; coupling with structure models for mechanism; prospective use to rank drug combinations for organoid confirmation.
Patient-derived organoids are miniature 3D versions of a patient's tumour grown in the lab.
Growing a patient's own cancer cells in a dish and testing drugs on them directly, instead of guessing from genetics.
Nicheformer is a model trained on both dissociated and spatial data so it learns how a cell's neighbourhood shapes it.
Open-source structure models that match AlphaFold 3, with Boltz-2 also predicting how strongly a drug binds.
Speculative: in silico trials and digital twins
A tumour's multi-omic profile seeds a patient-specific virtual cell population; treatment sequences are simulated before the first cycle; models are updated from ctDNA and imaging during care. Requires validation standards that do not yet exist.
Train one AI on scans, slides, genomics, and outcomes from many patients so it can predict, for a new patient, which treatment will work.
An ultra-sensitive blood test after surgery that detects leftover cancer months before a scan would.