An open foundation model of the cancer cell trained on perturbation data
Build a shared, openly available AI model that has learned how cancer cells respond to genetic and drug perturbations, so any lab can predict what a new drug or combination might do.
Single-cell perturbation atlases, CRISPR screens (DepMap), drug-response datasets and proteomics now exist at scale, but models trained on them are mostly proprietary or single-lab. The proposal is a pre-competitive, openly licensed foundation model of the cancer cell (transcriptomic and proteomic state under perturbation) trained on pooled public and consortium data with open weights, evaluated on held-out perturbations and prospective wet-lab validation, in the way AlphaFold became shared infrastructure for structure. The Chan Zuckerberg Initiative's virtual cell work and the Arc Institute's efforts are precedents.
- AI that is built but not validated or deployed · Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.
- Preclinical models that do not predict people · Nine in ten cancer drugs that work in mice fail in humans. Our models are the reason.
- Too many combinations to test · There are thousands of possible drug pairs and sequences. Trials can test a few dozen a year.