AI in oncology roadmap: pattern readers → foundation models → agents in the workflow
Artificial intelligence in cancer started as software that flagged spots on a mammogram. It now designs molecules, reads slides better than any single pathologist for some tasks, and is beginning to match patients to trials and draft the tumour board summary; the question is which of it will be proven to help.
Overview
Three strands of AI are converging on oncology. In discovery, structure prediction (AlphaFold 3, Boltz, Chai) and generative chemistry have produced the first AI-designed candidates in trials, and perturbation-scale single-cell datasets are training models that try to predict what a drug will do to a cell. In diagnosis, foundation models trained on millions of slides and scans (Virchow, Prov-GigaPath, UNI, TITAN, CT-FM) underpin the first AI tests cleared to predict treatment benefit (ArteraAI Prostate 2025, ArteraAI Breast 2026) and the first randomised evidence that AI reading improves screening (MASAI). In the clinic, language models are entering trial matching, documentation and tumour-board support, with radiotherapy auto-contouring as the most mature deployed use.
The gap between the thousands of published models and the handful in clinical use is the defining feature of the field. Prospective, ideally randomised, evidence that an AI-guided decision improves an outcome exists for a few tools; a regulatory route for models that keep updating, payment codes for AI-derived biomarkers, and data that can be shared or federated across hospitals are all unsettled.
This roadmap covers the whole stack from molecule to clinic; the companion roadmaps go deeper on the AI-assisted clinic and on the virtual cell.
- 1998-2016historic
Computer-aided detection
The first cleared cancer AI was computer-aided detection for mammography in 1998, which marked suspicious regions for the radiologist and, in large observational studies, did not improve accuracy. Rule-based decision support for treatment recommendations was tried and mostly abandoned. The lesson that survived: an algorithm has to be evaluated on the decision it changes, not on the pattern it finds.
- 2017-2022historic
Deep learning reaches cleared devices
Convolutional networks trained on labelled images matched specialists on narrow tasks. Paige Prostate (2021) became the first FDA-authorised AI for reading pathology slides; radiology triage tools for haemorrhage and embolism were cleared by the dozen; Sybil and Mirai predicted future lung and breast cancer from today's scan. Whole-slide scanning became routine in large centres, which made slide-level AI possible at all.
- 2020-2026current
Structure prediction and generative design
AlphaFold made protein structure a lookup rather than a two-year experiment; AlphaFold 3 (2024), Boltz and Chai extended it to drug-protein and antibody complexes, and RFdiffusion and ESM3 design proteins that never existed. Insilico's generative chemistry produced the first AI-discovered drug to reach phase 2, Isomorphic's first oncology candidate was cleared for trials, and Recursion and Xaira are betting that image and perturbation data can find targets no hypothesis would. None has yet produced an approved cancer drug, which is the honest benchmark.
AlphaFold 3Boltz-1 / Boltz-2 (MIT, open)Chai-1 / Chai-2RFdiffusion / RFdiffusion2 and ProteinMPNN (Baker Lab)ESM3 (EvolutionaryScale)AI-driven drug & target discoveryChemistry42 and Pharma.AI (Insilico)Insilico MedicineIsomorphic LabsRecursion PharmaceuticalsPhenom-2 and Recursion OSXaira TherapeuticsGoogle DeepMind (and Google Research) - 2023-2026current
Foundation models and the first predictive tests
Pathology models pretrained on millions of slides (Virchow, Prov-GigaPath, UNI and CONCH, H-optimus, TITAN) predict mutations, biomarkers and outcomes from a routine stain; MUSK adds clinical text. ArteraAI Prostate (2025) was the first AI test cleared to predict benefit from a treatment, and ArteraAI Breast followed in 2026. MASAI gave the first randomised evidence that AI-supported screening finds more cancers with less workload; Aidoc CARE (January 2026) was the first foundation-model triage platform cleared. Single-cell models (Geneformer, scGPT, State) and the Tahoe-100M dataset began the same arc for biology.
Pathology & radiology foundation modelsVirchow / Virchow2 (Paige, MSK)Prov-GigaPath (Microsoft, Providence)UNI and CONCH (Harvard, Mahmood Lab)H-optimus (Bioptimus)TITAN (whole-slide multimodal model)MUSK (Stanford, vision-language pathology)CHIEF (Harvard, Yu Lab)ArteraAI ProstateArteraAI BreastArteraMASAI (Mammography Screening with Artificial Intelligence)Aidoc CARE (clinical radiology foundation model)CT-FM (whole-body CT foundation model)Merlin (Stanford abdominal CT vision-language model)GeneformerscGPTState (Arc Institute perturbation model)Tahoe-100MPathology AI benchmarks (CAMELYON, PANDA, TCGA slide tasks) - 2025-2028emerging
Language models enter the workflow
The first widely deployed AI in cancer care is not a diagnosis but a time-saver: auto-contouring of organs and tumours for radiotherapy planning now runs in hundreds of centres. Language models are being tested to match patients to trials from the record at the moment a treatment is chosen, to draft tumour-board summaries and pathology reports, and to answer patient questions under supervision. Federated learning lets models train across hospitals without moving data. The evidence standard for each is still being written.
AI auto-contouring and adaptive planningLimbus AITheraPanaceaAI trial matching & clinical decision supportTrial LibraryMassive BioMed-Gemini and MedLM (Google)Foresight (generative EHR model)Federated learning and privacy-preserving AIOwkinTempus AIMultidisciplinary tumour boardsTrial matching inside the electronic record at the moment a treatment is chosen - 2027-2032emerging
From prediction to prospective proof
The field has thousands of retrospective models and a handful of prospective trials. The infrastructure being proposed: a registry of external validation datasets with mandatory reporting, AI-first reading for high-volume common diagnoses with pathologists handling exceptions, every routine CT checked opportunistically for early cancer signs with a tracked pathway, AI central reads to cut trial endpoint cost, and digital twins as virtual control arms where a randomised control is unethical. Regulators are building predetermined change control plans so that models can update without re-clearance.
A registry of external validation datasets for cancer AI models, with mandatory reportingAI-first reading for high-volume common cancer diagnoses, pathologist for the exceptionsEvery routine CT scan checked by AI for early cancer signs, with a tracked follow-up pathwayAI-assisted central imaging reads to cut endpoint cost and variabilityDigital twins and virtual control armsNCI Imaging Data Commons (IDC)Flatiron Health–Foundation Medicine Clinico-Genomic DatabaseAI that is built but not validated or deployed - 2030+speculative
Patient-level models and the virtual cell
The two long-range bets are a multimodal model that reads slides, scans, genomics and the record to recommend and monitor treatment, and a virtual cell accurate enough to run a drug experiment in silico before it is run in a dish. Both depend on data at a scale no single institution holds, on validation standards that do not yet exist, and on liability and consent questions that are open today. The companion roadmaps on the AI clinic and the virtual cell follow each in detail.
Patient-level multimodal foundation models for treatment selectionTempus multimodal modelsPathos AINoetikArc Virtual Cell AtlasIn silico trials to choose the dose before the first patientAI in the oncology clinic: from narrow cleared tools to multimodal decision supportVirtual cell roadmap: from bulk omics to a predictive model of a cancer cell - What sets the pacecurrent
Validation, data and compute
The bottleneck is not model quality but the path from a published model to a deployed one: prospective evidence, external validation, regulatory status for updating models, payment, and data that can be shared. Records, scans and genomes sit in silos; real-world outcomes are weakly recorded, so there is little to learn from; and the workforce that would supervise AI is already short. Compute and model platforms are the one input that is not scarce.
Probability ranges are named estimates that the claim is borne out on roughly a five-year horizon. They are meant to be argued with: propose a revision with your name and reasoning via a pull request to src/data/confidence.ts.
Story
topComputer-aided detection
The first cleared cancer AI was computer-aided detection for mammography in 1998, which marked suspicious regions for the radiologist and, in large observational studies, did not improve accuracy. Rule-based decision support for treatment recommendations was tried and mostly abandoned. The lesson that survived: an algorithm has to be evaluated on the decision it changes, not on the pattern it finds.
Deep learning reaches cleared devices
Convolutional networks trained on labelled images matched specialists on narrow tasks. Paige Prostate (2021) became the first FDA-authorised AI for reading pathology slides; radiology triage tools for haemorrhage and embolism were cleared by the dozen; Sybil and Mirai predicted future lung and breast cancer from today's scan. Whole-slide scanning became routine in large centres, which made slide-level AI possible at all.
The first AI for reading biopsy slides authorised by the FDA, which points pathologists to prostate cancer they might otherwise miss.
MSK spin-out with the first FDA-cleared AI pathology product and the Virchow foundation model.
Scanning microscope slides and letting software measure things a pathologist cannot see, including predictions of who will benefit from a treatment.
The scanners that turn glass slides into gigapixel images, and the software that stores and serves them, without which pathology AI cannot run.
Predicts a person's six-year lung cancer risk from one low-dose CT, even when no nodule is visible.
Reads a mammogram to estimate five-year breast cancer risk, consistently across races and devices.
Radiology AI company with the first FDA-cleared foundation-model triage platform (CARE, January 2026); oncology-relevant for incidental findings and workflow.
Structure prediction and generative design
AlphaFold made protein structure a lookup rather than a two-year experiment; AlphaFold 3 (2024), Boltz and Chai extended it to drug-protein and antibody complexes, and RFdiffusion and ESM3 design proteins that never existed. Insilico's generative chemistry produced the first AI-discovered drug to reach phase 2, Isomorphic's first oncology candidate was cleared for trials, and Recursion and Xaira are betting that image and perturbation data can find targets no hypothesis would. None has yet produced an approved cancer drug, which is the honest benchmark.
Predicts the 3D shape of proteins together with DNA, RNA, small molecules and antibodies, the starting point for much modern drug design.
Open-source structure models that match AlphaFold 3, with Boltz-2 also predicting how strongly a drug binds.
Structure and antibody-design models from Chai Discovery, with Chai-2 reporting high zero-shot antibody hit rates.
The tools that design entirely new proteins to bind a chosen target, now used for cancer binders and antibodies.
ESM3 is a generative protein model that designed a working fluorescent protein far from any natural sequence.
Using machine learning to pick targets, design molecules and antibodies, and predict which ADC will work.
Generative chemistry platform behind the first AI-discovered drug to reach phase 2, plus oncology candidates.
Generative-AI drug discovery company, listed in Hong Kong in December 2025, with a pan-KRAS candidate and a pan-TEAD inhibitor in the clinic.
Alphabet's AlphaFold-derived drug design company; its first AI-designed oncology candidate was cleared for human trials in January 2026 after a $2.1B raise.
Recursion is an AI-first biotech (merged with Exscientia in 2024) with a clinical oncology pipeline that includes an RBM39 degrader and a MEK inhibitor for familial adenomatous polyposis.
A model trained on billions of cell microscopy images to read what a drug or gene knockout does to a cell.
Launched in 2024 with over $1 billion to build AI-native drug discovery from Baker-lab protein design.
Google DeepMind built AlphaFold, AlphaMissense, AlphaGenome and Med-Gemini, the reference models for structure, variants, and medical multimodal reasoning.
Foundation models and the first predictive tests
Pathology models pretrained on millions of slides (Virchow, Prov-GigaPath, UNI and CONCH, H-optimus, TITAN) predict mutations, biomarkers and outcomes from a routine stain; MUSK adds clinical text. ArteraAI Prostate (2025) was the first AI test cleared to predict benefit from a treatment, and ArteraAI Breast followed in 2026. MASAI gave the first randomised evidence that AI-supported screening finds more cancers with less workload; Aidoc CARE (January 2026) was the first foundation-model triage platform cleared. Single-cell models (Geneformer, scGPT, State) and the Tahoe-100M dataset began the same arc for biology.
Very large AI models trained on millions of slides or scans that can be adapted to almost any diagnostic question.
A pathology foundation model trained on millions of slides that can detect cancer and predict biomarkers from an ordinary H&E slide.
An open pathology model trained on 1.3 billion image tiles from a US health system, modelling whole slides at gigapixel scale.
Two open academic pathology models: UNI reads tissue images, CONCH links images with pathology text.
An open 1.1-billion-parameter pathology model from a French startup, among the strongest on public benchmarks.
TITAN is a model that summarises a whole slide, not just tiles, and can write a draft pathology report.
A model that reads slides and clinical text together to predict who will respond to immunotherapy.
A pathology model trained across 19 cancer types that predicts survival and mutations from slides.
The first AI tool cleared by the FDA to predict both prognosis and treatment benefit from a routine biopsy slide, in prostate cancer.
An FDA-cleared AI test (May 2026) that reads breast cancer slides to estimate recurrence risk in early hormone-positive disease.
First company with FDA-cleared AI pathology tests that predict treatment benefit (prostate 2025, breast 2026).
The first randomised trial of AI in breast screening found more cancers and cut radiologists' reading work almost in half without more false alarms.
Aidoc CARE is a single foundation model behind many FDA-cleared triage alerts in emergency radiology.
A model pretrained on 148,000 CT scans to segment organs and triage findings.
Merlin is a model trained on 15,000 CT scans with their reports that can find and describe hundreds of findings.
The first widely used transformer trained on millions of single cells, able to predict which genes matter in a disease.
A GPT-style model for single-cell data that predicts cell types, perturbation responses, and gene networks.
Predicts how cells will respond to a drug or gene knockout, trained on over 100 million perturbed cells.
Tahoe-100M is the biggest single-cell dataset ever released, built to teach AI how cancer cells respond to drugs.
The exam papers every pathology model is graded on.
Language models enter the workflow
The first widely deployed AI in cancer care is not a diagnosis but a time-saver: auto-contouring of organs and tumours for radiotherapy planning now runs in hundreds of centres. Language models are being tested to match patients to trials from the record at the moment a treatment is chosen, to draft tumour-board summaries and pathology reports, and to answer patient questions under supervision. Federated learning lets models train across hospitals without moving data. The evidence standard for each is still being written.
Software that draws organs and tumours on scans automatically, saving hours per patient and making daily plan adaptation practical.
Limbus AI provides AI auto-contouring for radiotherapy, FDA-cleared and used across hundreds of centres.
TheraPanacea is a Paris AI company whose ART-Plan does auto-contouring and synthetic CT in radiotherapy.
Software, increasingly LLM-based, that reads a patient's record and finds trials or guideline options they qualify for.
Trial Library is software that helps everyday cancer clinics spot which patients might qualify for a clinical trial, refer them, and remove practical barriers like transport so more people, especially in under served communities, can join trials.
Massive Bio uses artificial intelligence to read a cancer patient's medical records and match them to clinical trials they may be eligible for, anywhere in the world, with doctors checking the results.
Google's medical versions of its Gemini models, able to reason over text, images, and long records.
A model trained on millions of hospital records that forecasts a patient's next diagnoses.
Training AI models across many hospitals without moving patient data, so the model learns from everyone while the data stay put.
French AI biotech using federated learning across hospitals; first CE-marked AI for MSI prediction from H&E.
Genomic testing plus one of the largest multimodal clinical datasets, used for AI models and trial matching.
Regular meetings where surgeons, oncologists, radiologists, pathologists and others review each patient's case together and agree a plan; mandatory in many countries and associated with more guideline-concordant care.
When an oncologist opens the order screen to prescribe a new line of treatment, the record would show the trials this patient may fit, with the nearest open site and a one-click referral.
From prediction to prospective proof
The field has thousands of retrospective models and a handful of prospective trials. The infrastructure being proposed: a registry of external validation datasets with mandatory reporting, AI-first reading for high-volume common diagnoses with pathologists handling exceptions, every routine CT checked opportunistically for early cancer signs with a tracked pathway, AI central reads to cut trial endpoint cost, and digital twins as virtual control arms where a randomised control is unethical. Regulators are building predetermined change control plans so that models can update without re-clearance.
Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.
Let validated AI make the first read on routine, high-volume samples like cervical smears and standard breast biopsy stains, so scarce pathologists spend their time on the difficult cases.
Hundreds of millions of CT scans are done each year for other reasons. Software could check each one for early lung, kidney, liver and pancreas changes, but only if a follow-up system exists.
Measuring tumours on scans for trials is slow, expensive and inconsistent between readers. Software that measures lesions and flags changes, checked by a radiologist, could make trial endpoints cheaper and more reliable.
Using a model of what would have happened to a patient on standard treatment, so fewer people have to be randomised to it.
The Imaging Data Commons is TCIA in the cloud, ready for large-scale model training.
Real-world evidence at scale: what happened to patients with a given genomic profile on a given treatment.
Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.
Patient-level models and the virtual cell
The two long-range bets are a multimodal model that reads slides, scans, genomics and the record to recommend and monitor treatment, and a virtual cell accurate enough to run a drug experiment in silico before it is run in a dish. Both depend on data at a scale no single institution holds, on validation standards that do not yet exist, and on liability and consent questions that are open today. The companion roadmaps on the AI clinic and the virtual cell follow each in detail.
Train one AI on scans, slides, genomics, and outcomes from many patients so it can predict, for a new patient, which treatment will work.
Models trained on Tempus's paired genomic, pathology, imaging and outcome data to predict response and prognosis.
Pathos AI uses very large artificial intelligence models trained on millions of cancer patients' records, scans and genetic data to pick which experimental cancer drugs to develop and which patients to test them in, and is building its own pipeline of such drugs.
Noetik builds artificial intelligence models of tumours from huge sets of tissue images and molecular data, to predict which patients will respond to a cancer drug and to find new targets.
Arc's growing library of cell data, the fuel for virtual cell models.
Simulating thousands of virtual patients on a computer can suggest which dose and schedule to test, so fewer real patients receive doses that are too high or too low.
How AI is moving from single-task readers of scans and slides towards systems that weigh everything about a patient, and what regulators and evidence still require.
The attempt to build a computer model of a cell good enough to predict what a drug or mutation will do before anyone runs the experiment.
Validation, data and compute
The bottleneck is not model quality but the path from a published model to a deployed one: prospective evidence, external validation, regulatory status for updating models, payment, and data that can be shared. Records, scans and genomes sit in silos; real-world outcomes are weakly recorded, so there is little to learn from; and the workforce that would supervise AI is already short. Compute and model platforms are the one input that is not scarce.
Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.
Records, scans, genomes and outcomes sit in separate systems that cannot talk. Every patient's experience is lost to the next.
We do not reliably know what happens to patients after approval, so we cannot tell which drugs deliver in practice.
The number of people with cancer is rising faster than the workforce trained to treat them.
Fewer than half of landmark cancer biology findings reproduce when someone else tries.
AI compute platforms are the GPUs, model libraries, and cloud services that pathology, radiology, and drug-design AI run on.
Supplies the GPUs and the BioNeMo framework most biological foundation models are trained on, and co-developed Evo 2 with Arc.
BioNeMo is the software stack many biology foundation models are trained and served with.
Pages like this
not linked directly; found by shared links- RoadmapDiagnostics roadmap: stains → gene panels → blood tests that decide treatment
Shares Prov-GigaPath (Microsoft, Providence), MUSK (Stanford, vision-language pathology), Virchow / Virchow2 (Paige, MSK), ArteraAI Breast.
- IdeaA pre-competitive consortium to train a shared multimodal cancer foundation model
Shares PathAI, Owkin, Paige AI, Patient-level multimodal foundation models for treatment selection.
- RoadmapTrial modernisation roadmap: the randomised trial → platforms and adaptive designs → decentralised, pragmatic and always-on
Shares Massive Bio, Flatiron Health–Foundation Medicine Clinico-Genomic Database, Trial Library, Flatiron Health (Roche).
- InstitutionArc Institute
Shares Arc Virtual Cell Atlas, Tahoe-100M, State (Arc Institute perturbation model), Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell.
- IdeaA federated learning consortium of cancer centres that jointly own the models
Shares Owkin, Patient-level multimodal foundation models for treatment selection, Pathology & radiology foundation models, AI in radiology.
- CompanyLunit
Shares AI-assisted central imaging reads to cut endpoint cost and variability, AI-first reading for high-volume common cancer diagnoses, pathologist for the exceptions, Mammography & tomosynthesis, AI in radiology.
- TechnologyCT (computed tomography)
Shares CT-FM (whole-body CT foundation model), Merlin (Stanford abdominal CT vision-language model), Sybil (MIT/MGH lung cancer risk from CT), AI-assisted central imaging reads to cut endpoint cost and variability.
- CompanyFoundation Medicine (Roche)
Shares Flatiron Health–Foundation Medicine Clinico-Genomic Database, Flatiron Health (Roche), Digital twins and virtual control arms, Weak real-world evidence and registries.