Drug discovery roadmap: screening in mice → maps of dependency → designing in silico
Finding the next cancer drug used to mean testing compounds on mice and cell lines and hoping. It now means mapping which genes each cancer cannot live without, growing a patient's tumour in a dish, and designing molecules on a computer; the job is making those tools predict what happens in people.
Overview
Nine in ten cancer drugs that work in mice fail in humans, and most of the history of drug discovery is the attempt to close that gap. Natural-product screening found vincristine and paclitaxel; target-based discovery, structural biology and high-throughput screening produced the kinase inhibitors. What they could not do was predict which patients a drug would help or which combination would hold.
The present toolkit attacks that directly. Genome-wide CRISPR screens (DepMap) map the dependencies of a thousand cancer cell lines and expose synthetic-lethal targets; patient-derived organoids and xenografts keep a tumour's biology closer to the patient's; functional testing of drugs on a patient's own cells is being run alongside trials. Structure prediction (AlphaFold 3, Boltz, Chai) and generative design have put the first AI-designed molecules into oncology trials, and perturbation datasets of a hundred million cells are training models that try to predict a drug's effect before the experiment.
The pace is set by the predictive validity of models, by the half of landmark findings that do not reproduce, by the valley between an academic discovery and a funded programme, and by the secrecy that keeps compound libraries and negative results locked up.
- 1950s-1990shistoric
Screening nature and hoping
The NCI screened tens of thousands of compounds in mouse leukaemias and, from 1990, in a panel of sixty human cell lines. It found the periwinkle alkaloid vincristine, the yew-bark taxane paclitaxel and the antibiotic dactinomycin, and it established the pipeline everyone still uses: cells, then mice, then people. What it could not do was say which people.
- 1990s-2015historic
Targets, structures and high-throughput screens
Cloning the oncogenes gave discovery a target; crystal structures gave it a shape to fit; robotic screening of millions of compounds and later DNA-encoded libraries gave it throughput. Imatinib was the proof. The cost was a generation of drugs that hit their target and did nothing for patients, because the cell-line and xenograft models that selected them did not represent human tumours. Nine in ten oncology drugs entering trials still fail.
- 2015-2024current
Maps of dependency and models closer to the patient
Genome-wide CRISPR knockout screens across a thousand cell lines (DepMap) list which genes each cancer cannot live without, and expose synthetic-lethal pairs such as PRMT5 in MTAP-deleted tumours and WRN in mismatch-repair-deficient ones. Patient-derived organoids keep a tumour's architecture and drug response in a dish; xenograft banks keep it in a mouse; TCGA, GENIE and CPTAC supply the genomes and proteomes to interpret them. The first drugs found this way are now in trials.
CRISPR functional genomicsDepMap (Cancer Dependency Map)Synthetic lethality approachesPatient-derived organoidsHUB OrganoidsPatient-derived xenograftsChampions OncologyCancer Models (PDCM Finder) & HCMITCGA / NCI Genomic Data CommonsAACR Project GENIECPTAC (Clinical Proteomic Tumor Analysis Consortium)Proteomics & phosphoproteomics - 2020-2026current
Structure prediction and generative design
AlphaFold turned protein structure into a lookup, and AlphaFold 3 (2024), Boltz and Chai extended it to drug and antibody complexes; RFdiffusion and ESM3 design proteins from scratch. Insilico's generative chemistry produced the first AI-discovered drug to reach phase 2; Isomorphic's first oncology candidate entered trials; Recursion merged with Exscientia to pair image-based biology with design; Xaira launched with over a billion dollars to build discovery around these models. The honest scorecard: faster hit-to-candidate, no approved cancer drug yet.
AlphaFold 3Boltz-1 / Boltz-2 (MIT, open)Chai-1 / Chai-2Chai DiscoveryRFdiffusion / RFdiffusion2 and ProteinMPNN (Baker Lab)ESM3 (EvolutionaryScale)AI-driven drug & target discoveryChemistry42 and Pharma.AI (Insilico)Insilico MedicineIsomorphic LabsGoogle DeepMind (and Google Research)Recursion PharmaceuticalsExscientiaPhenom-2 and Recursion OSXaira TherapeuticsGenerate:Biomedicines - 2015-2026current
New modalities as platforms
Discovery is no longer only about small molecules. Degrader and molecular glue platforms remove proteins that cannot be inhibited; ADC linker chemistry (Araris) and bicyclic peptide conjugates (Bicycle) turn a payload into a targeted drug; oligonucleotides silence genes; chemoproteomics finds covalent handles on KRAS. Each platform generates candidates faster than trials can test them, which moves the constraint downstream.
PROTACs & molecular glues (targeted protein degradation)Molecular glue discovery platformsDegrader-antibody conjugate (DAC)Araris Biotech (Taiho)Bicycle TherapeuticsPeptide-drug & small-molecule-drug conjugatesOligonucleotide therapeuticsFrontier MedicinesSotorasibADC roadmap: from Mylotarg to bispecific and dual-payload ADCs - 2026-2030emerging
Functional precision medicine
Instead of inferring drug response from genotype, test the drug on the patient's own cells: organoid pharmacotyping in pancreatic cancer, the PARIS organoid screen, tumour fragments kept alive with their vessels, BH3 profiling in leukaemia. The proposal that would make it a field is to grow each trial patient's tumour as organoids and let the results decide which platform arm opens next, with shared reference organoid and xenograft panels so every laboratory tests against the same models.
Functional (ex vivo) drug testingOrganoid-guided therapy at scalePDAC organoid pharmacotypingSEngine Precision MedicineCuresponseBH3 profiling (functional apoptosis testing)Grow each trial patient's tumour as organoids to decide which platform arm opens nextShared reference organoid and PDX panels that every lab can test against - 2026-2030emerging
Perturbation data at the scale the problem needs
Tahoe-100M measured a hundred million single cells across 1,100 drugs and fifty cancer lines; Arc's Virtual Cell Atlas and State model, Geneformer and scGPT are the attempts to learn from that scale how a cell will respond to a perturbation it has never seen. Rigorous benchmarks show current models barely beat simple baselines on unseen contexts, which is the right kind of bad news: the problem is now measurable. The virtual cell roadmap follows this in detail.
- 2030+speculative
In silico first
The long-range bet is that a candidate is designed, its dose chosen and its toxicity screened in silico and on linked human organ chips before the first mouse, and that digital twins reduce the size of the trials that follow. That requires models that generalise, which requires data that reproduce. A funded replication in every cancer biology PhD and a registry for preclinical experiments that did not work are the unglamorous prerequisites.
In silico trials to choose the dose before the first patientLinked human organ chips to predict side effects before people are dosedDigital twins and virtual control armsDe novo designed protein bindersEvery cancer biology PhD begins with a funded replication of a published findingA registry for preclinical experiments that did not workAI in oncology roadmap: pattern readers → foundation models → agents in the workflow - What sets the pacecurrent
Models, reproducibility and the valley of death
Preclinical models still do not predict people, fewer than half of landmark findings reproduce, and most academic discoveries die before anyone tests them in humans because no one funds the step between. Companies hold compound libraries and negative results that would save others years. A shared compound pool for rare cancer researchers and a guaranteed purchase prize for the first drug against a named hard target are two proposals that attack the incentive problem directly.
Preclinical models that do not predict peoplePreclinical results do not reproduceThe valley of death between lab and productSecrecy and intellectual property block collaborationFailures are hiddenThe undruggable driversA shared compound library that rare cancer researchers can actually useA guaranteed purchase prize for the first drug against a named hard target
Probability ranges are named estimates that the claim is borne out on roughly a five-year horizon. They are meant to be argued with: propose a revision with your name and reasoning via a pull request to src/data/confidence.ts.
Story
topScreening nature and hoping
The NCI screened tens of thousands of compounds in mouse leukaemias and, from 1990, in a panel of sixty human cell lines. It found the periwinkle alkaloid vincristine, the yew-bark taxane paclitaxel and the antibiotic dactinomycin, and it established the pipeline everyone still uses: cells, then mice, then people. What it could not do was say which people.
The NCI is the US government's cancer research agency, spending ~$7B a year and running the Cancer Centers Program, TCGA, and Rosenberg's cell therapy lab.
A plant-derived chemotherapy from the Madagascar periwinkle that has been in almost every childhood leukaemia and lymphoma regimen since the 1960s.
A microtubule poison discovered in the Pacific yew tree, among the most used chemotherapies in breast, lung, and ovarian cancer.
The first antibiotic used as an anticancer drug (1954) and still the core of chemotherapy for Wilms tumour, rhabdomyosarcoma and gestational trophoblastic disease.
A patient-derived xenograft is a patient's tumour grown in a mouse, used to test drugs before they reach people.
Targets, structures and high-throughput screens
Cloning the oncogenes gave discovery a target; crystal structures gave it a shape to fit; robotic screening of millions of compounds and later DNA-encoded libraries gave it throughput. Imatinib was the proof. The cost was a generation of drugs that hit their target and did nothing for patients, because the cell-line and xenograft models that selected them did not represent human tumours. Nine in ten oncology drugs entering trials still fail.
The drug that started the targeted therapy era in 2001, turning chronic myeloid leukaemia into a manageable condition with near-normal life expectancy.
Testing millions or billions of chemical compounds against a cancer target automatically to find starting points for new drugs.
Structural biology infrastructure is the microscopes, X-ray sources, and prediction models that show what a cancer protein looks like so chemists can design a drug to fit it.
Nine in ten cancer drugs that work in mice fail in humans. Our models are the reason.
A patient-derived xenograft is a patient's tumour grown in a mouse, used to test drugs before they reach people.
Maps of dependency and models closer to the patient
Genome-wide CRISPR knockout screens across a thousand cell lines (DepMap) list which genes each cancer cannot live without, and expose synthetic-lethal pairs such as PRMT5 in MTAP-deleted tumours and WRN in mismatch-repair-deficient ones. Patient-derived organoids keep a tumour's architecture and drug response in a dish; xenograft banks keep it in a mouse; TCGA, GENIE and CPTAC supply the genomes and proteomes to interpret them. The first drugs found this way are now in trials.
Knocking out every gene one at a time in cancer cells to find which ones they cannot live without.
Which genes each cancer cell line cannot live without. The map of synthetic-lethal targets.
Finding a second gene that a cancer needs only because its first gene is broken, then hitting the second one.
Patient-derived organoids are miniature 3D versions of a patient's tumour grown in the lab.
The Clevers-lab spin-out that licenses patient-derived organoid technology and maintains a living biobank.
A patient-derived xenograft is a patient's tumour grown in a mouse, used to test drugs before they reach people.
Oncology CRO with the largest bank of patient-derived xenograft models (TumorGraft), now adding 3D organoid screening and radiopharmaceutical services.
Find a mouse or dish model that matches a tumour type or mutation.
The reference atlas of cancer genomes that most cancer biology since 2008 is built on.
AACR Project GENIE is real-world tumour sequencing data shared by leading cancer centres.
CPTAC measures the proteins, not just the genes, of thousands of tumours.
Measuring the proteins in a tumour, which is what drugs actually hit, rather than the genes that encode them.
Structure prediction and generative design
AlphaFold turned protein structure into a lookup, and AlphaFold 3 (2024), Boltz and Chai extended it to drug and antibody complexes; RFdiffusion and ESM3 design proteins from scratch. Insilico's generative chemistry produced the first AI-discovered drug to reach phase 2; Isomorphic's first oncology candidate entered trials; Recursion merged with Exscientia to pair image-based biology with design; Xaira launched with over a billion dollars to build discovery around these models. The honest scorecard: faster hit-to-candidate, no approved cancer drug yet.
Predicts the 3D shape of proteins together with DNA, RNA, small molecules and antibodies, the starting point for much modern drug design.
Open-source structure models that match AlphaFold 3, with Boltz-2 also predicting how strongly a drug binds.
Structure and antibody-design models from Chai Discovery, with Chai-2 reporting high zero-shot antibody hit rates.
Makes Chai-1 and Chai-2, open-weight structure models used for antibody and binder design.
The tools that design entirely new proteins to bind a chosen target, now used for cancer binders and antibodies.
ESM3 is a generative protein model that designed a working fluorescent protein far from any natural sequence.
Using machine learning to pick targets, design molecules and antibodies, and predict which ADC will work.
Generative chemistry platform behind the first AI-discovered drug to reach phase 2, plus oncology candidates.
Generative-AI drug discovery company, listed in Hong Kong in December 2025, with a pan-KRAS candidate and a pan-TEAD inhibitor in the clinic.
Alphabet's AlphaFold-derived drug design company; its first AI-designed oncology candidate was cleared for human trials in January 2026 after a $2.1B raise.
Google DeepMind built AlphaFold, AlphaMissense, AlphaGenome and Med-Gemini, the reference models for structure, variants, and medical multimodal reasoning.
Recursion is an AI-first biotech (merged with Exscientia in 2024) with a clinical oncology pipeline that includes an RBM39 degrader and a MEK inhibitor for familial adenomatous polyposis.
Exscientia was a British pioneer in designing drugs with artificial intelligence, including cancer drugs that reached early clinical trials. It has been folded into the US company Recursion.
A model trained on billions of cell microscopy images to read what a drug or gene knockout does to a cell.
Launched in 2024 with over $1 billion to build AI-native drug discovery from Baker-lab protein design.
Generative-AI protein design company (Flagship) that went public in 2026 and is testing an antibody that neutralises leaked ADC payload to reduce side effects.
New modalities as platforms
Discovery is no longer only about small molecules. Degrader and molecular glue platforms remove proteins that cannot be inhibited; ADC linker chemistry (Araris) and bicyclic peptide conjugates (Bicycle) turn a payload into a targeted drug; oligonucleotides silence genes; chemoproteomics finds covalent handles on KRAS. Each platform generates candidates faster than trials can test them, which moves the constraint downstream.
Instead of blocking a protein, these drugs tag it for the cell's own garbage disposal, removing it entirely.
Molecular glues are small molecules that stick two proteins together so the cell destroys one of them. They are smaller and more drug-like than bifunctional degraders.
An ADC that delivers a protein-destroying molecule instead of chemotherapy, hitting targets inside the cell that were previously unreachable.
Swiss linker-technology company (AraLinQ) acquired by Taiho for up to $1.14B; first clinical ADC ARC-02 (CD79b) dosed in June 2026.
Inventor of bicyclic peptide drug conjugates; its lead Nectin-4 conjugate was deprioritised in 2026 after regulatory feedback.
Like an ADC but with a small targeting peptide instead of an antibody, so it penetrates tumours faster and is cheaper to make.
Oligonucleotide therapeutics are short synthetic strands of genetic code that silence a specific cancer gene.
Chemoproteomics company with FMC-376, a KRAS G12C inhibitor that hits both the ON and OFF states of the protein, in phase 1/2.
Sotorasib (Lumakras) was the first drug to hit KRAS, approved in 2021 after four decades of failure.
Twenty-five years of trying to make chemotherapy hit only cancer cells, from the unstable first ADC to today's third-generation blockbusters and the fourth generation now in trials.
Functional precision medicine
Instead of inferring drug response from genotype, test the drug on the patient's own cells: organoid pharmacotyping in pancreatic cancer, the PARIS organoid screen, tumour fragments kept alive with their vessels, BH3 profiling in leukaemia. The proposal that would make it a field is to grow each trial patient's tumour as organoids and let the results decide which platform arm opens next, with shared reference organoid and xenograft panels so every laboratory tests against the same models.
Growing a patient's own cancer cells in a dish and testing drugs on them directly, instead of guessing from genetics.
Organoid-guided therapy means routinely growing a piece of each patient's tumour and testing drugs on it before choosing, rather than relying on genetics alone.
Growing a patient's pancreatic tumour as mini-organs in a dish and testing chemotherapies on them to pick the regimen most likely to work.
Runs the PARIS test: a patient's tumour grown as organoids and screened against 240+ drugs to find options sequencing cannot see.
Israeli company whose cResponse test keeps a patient's tumour fragment alive, with its vessels and immune cells, to test which treatments it responds to.
A lab test that measures how close a leukaemia cell is to self-destructing, and which survival protein is holding it back, to predict response to venetoclax-type drugs.
While patients are treated in a platform trial, their tumour cells grow in a dish and are tested against dozens of drug pairs. The pairs that win in the dish become the next arms.
If every lab had access to the same set of well-characterised tumour models, results could be compared directly instead of each lab using its own private models.
Perturbation data at the scale the problem needs
Tahoe-100M measured a hundred million single cells across 1,100 drugs and fifty cancer lines; Arc's Virtual Cell Atlas and State model, Geneformer and scGPT are the attempts to learn from that scale how a cell will respond to a perturbation it has never seen. Rigorous benchmarks show current models barely beat simple baselines on unseen contexts, which is the right kind of bad news: the problem is now measurable. The virtual cell roadmap follows this in detail.
Tahoe-100M is the biggest single-cell dataset ever released, built to teach AI how cancer cells respond to drugs.
Produced Tahoe-100M, the largest single-cell drug-perturbation atlas, and trains models on it.
Arc's growing library of cell data, the fuel for virtual cell models.
The Arc Institute is a well-funded nonprofit research institute building 'virtual cell' AI models and the datasets to train them.
Predicts how cells will respond to a drug or gene knockout, trained on over 100 million perturbed cells.
The first widely used transformer trained on millions of single cells, able to predict which genes matter in a disease.
A GPT-style model for single-cell data that predicts cell types, perturbation responses, and gene networks.
The attempt to build a computer model of a cell good enough to predict what a drug or mutation will do before anyone runs the experiment.
In silico first
The long-range bet is that a candidate is designed, its dose chosen and its toxicity screened in silico and on linked human organ chips before the first mouse, and that digital twins reduce the size of the trials that follow. That requires models that generalise, which requires data that reproduce. A funded replication in every cancer biology PhD and a registry for preclinical experiments that did not work are the unglamorous prerequisites.
Simulating thousands of virtual patients on a computer can suggest which dose and schedule to test, so fewer real patients receive doses that are too high or too low.
Damage to the lungs, heart or liver is a common reason cancer drugs fail. Connected chips of human tissue may spot this earlier than animal tests.
Using a model of what would have happened to a patient on standard treatment, so fewer people have to be randomised to it.
Designing a protein from scratch on a computer to grip a chosen target, instead of finding one in an animal or a library.
Make the first project of every doctoral student a careful, published attempt to repeat an important result. Students learn rigour, and the field gets thousands of replications a year.
Most lab experiments that fail are never written up, so other labs repeat them. A simple, structured registry with a citable record for each failed experiment would stop the waste.
Artificial intelligence in cancer started as software that flagged spots on a mammogram. It now designs molecules, reads slides better than any single pathologist for some tasks, and is beginning to match patients to trials and draft the tumour board summary; the question is which of it will be proven to help.
Models, reproducibility and the valley of death
Preclinical models still do not predict people, fewer than half of landmark findings reproduce, and most academic discoveries die before anyone tests them in humans because no one funds the step between. Companies hold compound libraries and negative results that would save others years. A shared compound pool for rare cancer researchers and a guaranteed purchase prize for the first drug against a named hard target are two proposals that attack the incentive problem directly.
Nine in ten cancer drugs that work in mice fail in humans. Our models are the reason.
Fewer than half of landmark cancer biology findings reproduce when someone else tries.
Most academic discoveries die before anyone tests them in people because nobody funds the middle step.
Companies with complementary drugs rarely test them together, and data that could answer questions stays locked up.
Negative trials, failed drugs and abandoned programmes are rarely published, so the same mistakes are repeated.
The proteins that drive most cancers, such as MYC, mutant p53 and most RAS variants, still have no good drug.
Companies hold thousands of well-characterised drugs that could help rare cancers, but each request takes a year of legal negotiation. One standing agreement would unblock it.
Governments promised in advance to buy vaccines that did not yet exist, and they got made. The same promise could be made for a drug against a target everyone has given up on.
Pages like this
not linked directly; found by shared links- IdeaTest drugs on the patient's own cancer cells when there is no trial to join
Shares Champions Oncology, Curesponse, SEngine Precision Medicine, BH3 profiling (functional apoptosis testing).
- Key paperDefining a Cancer Dependency Map: which genes each cancer cell line cannot live without
Shares High-throughput screening and DNA-encoded libraries, DepMap (Cancer Dependency Map), Synthetic lethality approaches, Functional (ex vivo) drug testing.
- Key paperAlphaFold 2: predicting protein structures to near-experimental accuracy
Shares AlphaFold 3, Structural biology infrastructure (cryo-EM, synchrotrons, AlphaFold), De novo designed protein binders, AI-driven drug & target discovery.
- IdeaMake in vivo metastasis screens a required step in drug discovery
Shares Cancer Models (PDCM Finder) & HCMI, DepMap (Cancer Dependency Map), Patient-derived xenografts, CRISPR functional genomics.
- IdeaAn open model bank for the rare tumours nobody has models for
Shares Cancer Models (PDCM Finder) & HCMI, DepMap (Cancer Dependency Map), Patient-derived xenografts, Functional (ex vivo) drug testing.
- IdeaA public atlas of drug-pair responses across a thousand patient-derived organoids
Shares Shared reference organoid and PDX panels that every lab can test against, Cancer Models (PDCM Finder) & HCMI, DepMap (Cancer Dependency Map), Functional (ex vivo) drug testing.
- IdeaHold organoid drug tests to the same standard as a diagnostic test
Shares Curesponse, SEngine Precision Medicine, Functional (ex vivo) drug testing, Patient-derived organoids.
- IdeaSelf-driving laboratories that run the cancer biology hypothesis loop autonomously
Shares Recursion Pharmaceuticals, CRISPR functional genomics, Patient-derived organoids, Preclinical results do not reproduce.