OnCo

Open source in oncology

379 open-source projects the war on cancer runs on, in fourteen categories: 349 repositories read from the GitHub API on 2026-09-23 (licence as declared, stars, last push) and 30 project pages, package indexes and model cards fetched the same day. 277 were pushed to in the last two years. Filter by category, licence family, language, openness, last activity, technology, data source, cancer or maintainer; every chip is a filter and every name opens the project.

379 projects197,668 stars between them184 under permissive licences, 33 with no licence file143 maintained by an institution or company in OnCo179 name a paper in their README34 looked for and not recorded

How oncology is done in the open

Most of what happens between a tumour sample and a treatment decision now runs on code anyone can read. Reads are trimmed with fastp, aligned and called with GATK, Strelka or hmftools inside nf-core/sarek; variants are annotated by Ensembl VEP and interpreted against CIViC and OncoKB; signatures come from SigProfiler, clones from PyClone and PhyloWGS, and the whole cohort is explored in cBioPortal. Imaging is read in OHIF and 3D Slicer and segmented by nnU-Net and TotalSegmentator; radiotherapy plans are researched in matRad and OpenTPS and checked against Monte Carlo from Geant4, TOPAS and GATE; slides are analysed in QuPath and increasingly by foundation models whose weights are, sometimes, released.

The pipeline engine hub walks that path step by step and says which of these projects sits at each step. This page is the inventory: what exists, who maintains it, under which licence, and how open it really is. Read the licence column before reuse: a quarter of the repositories declare custom terms or none at all, and several of the best-known models release weights only for non-commercial use.

Contribute. Missing a project, or a fact is stale? The list is one file in the repository: add an entry with the repository and what it relates to, run npx tsx scripts/fetch-open-source.ts, and the licence, stars and dates are read from the source, never typed in. Or suggest it and a maintainer will. Add the project to the Open Medical Registry as well, so the wider catalogue of open medicine has it.

379 projects
AlphaFold 2
Google DeepMind (and Google Research) · since 2021
DeepMind's protein structure prediction system, code and weights released; the structures fill the AlphaFold Database.
14.9k
nnU-Net
German Cancer Research Center (DKFZ) · since 2019
The self-configuring segmentation method from DKFZ that wins most medical segmentation challenges out of the box, including brain, kidney and liver tumour tasks.
8.9k
MONAI
NVIDIA · since 2019
The PyTorch framework for deep learning in medical imaging, co-founded by NVIDIA and King's College London, used for tumour segmentation and detection research and products.
8.7k
AlphaFold 3
Google DeepMind (and Google Research) · since 2024
DeepMind's model of proteins with ligands, nucleic acids and modifications; code is released for non-commercial use and weights by request.
8.6k
DeepChem
since 2015
A Python library that democratises deep learning for drug discovery, materials and biology.
7k
MedSAM
University of Toronto, Bo Wang lab · since 2023
Segment Anything adapted to medical images across modalities, with released weights.
4.4k
OHIF Viewer
Open Health Imaging Foundation · since 2015
The Open Health Imaging Foundation's zero-footprint web DICOM viewer, the front end of many research imaging platforms and the NCI Imaging Data Commons.
4.3k
Evo 2
Arc Institute · since 2025
Arc Institute's genomic foundation model across all domains of life, with open weights.
4.2k
Boltz
MIT and Recursion · since 2024
An open biomolecular structure and affinity prediction model from MIT and Recursion, released under MIT with weights.
4.2k
RDKit
since 2013
The open cheminformatics toolkit that nearly all open drug discovery code depends on for molecules, fingerprints and descriptors.
3.6k
OpenFold
Herbert Irving Comprehensive Cancer Center, Columbia University · since 2021
A trainable, open reproduction of AlphaFold 2 from the AlQuraishi lab, with training data released.
3.4k
RFdiffusion
University of Washington, Institute for Protein Design · since 2023
The Baker lab's diffusion model for designing new proteins and binders, with released weights.
3.1k
TotalSegmentator
University Hospital Basel · since 2022
Segments over a hundred anatomical structures in CT and MRI in one command, widely used for organs at risk and body composition.
3k
ESM3 and ESM C
EvolutionaryScale · since 2024
EvolutionaryScale's protein language models; small models have open weights, larger ones are under a non-commercial licence.
3k
Seurat
New York Genome Center and NYU, Satija lab · since 2015
The R toolkit for single-cell genomics from the Satija lab, with integration, clustering and spatial support.
2.8k
3D Slicer
Slicer community, led from Brigham and Women's Hospital and Kitware · since 2020
The desktop platform for medical image analysis, segmentation, registration and image-guided therapy, with hundreds of extensions and a large research community.
2.6k
Scanpy
scverse · since 2017
The Python toolkit for single-cell gene expression analysis, the centre of the scverse ecosystem used across tumour single-cell studies.
2.6k
Chemprop
MIT · since 2019
MIT's message-passing neural networks for molecular property prediction, used in antibiotic and oncology screening papers.
2.5k
fastp
since 2017
An all-in-one preprocessor for sequencing reads (quality control, trimming, UMI handling) used at the head of many cancer pipelines.
2.4k
Cellpose
since 2020
A generalist deep-learning cell segmentation model used widely on histology, multiplex and cell-culture images.
2.4k
pydicom
since 2013
The Python library for reading and writing DICOM files that most imaging research code depends on.
2.2k
AlphaGenome
Google DeepMind (and Google Research) · since 2024
DeepMind's model predicting regulatory effects of DNA variants; API access is free for non-commercial use.
2.2k
Chai-1
Chai Discovery · since 2024
Chai Discovery's multi-modal structure prediction model; code and weights are released, with commercial use permitted under its terms.
2k
GATK (with Mutect2)
Broad Institute of MIT and Harvard · since 2014
The Broad Institute's Genome Analysis Toolkit: variant discovery for germline and somatic DNA, including the Mutect2 somatic caller and copy-number tools most cancer pipelines start from.
2k
OpenMM
since 2013
A high-performance molecular dynamics toolkit used for binding free energy and protein simulation in drug discovery.
2k
Galaxy
Galaxy Project · since 2015
A web platform for accessible, reproducible data analysis with thousands of tools, including full somatic variant and cancer workflows on public servers.
1.9k
ProteinMPNN
University of Washington, Institute for Protein Design · since 2022
Fast sequence design for a given protein backbone, from the Baker lab.
1.9k
CLAM
Mahmood lab, Brigham and Women's Hospital and Harvard Medical School · since 2020
Clustering-constrained attention multiple instance learning: the weakly supervised whole-slide classification pipeline from the Mahmood lab that many pathology AI papers build on.
1.7k
scvi-tools
scverse · since 2017
Deep probabilistic models for single-cell omics: integration, annotation and differential expression.
1.7k
ITK
Insight Software Consortium, Kitware · since 2010
The Insight Toolkit for image segmentation and registration, the foundation under 3D Slicer, ANTs and much of medical image computing.
1.7k
(showing the first 30)

Looked for and not recorded

Closed source, source available only under a restrictive licence, not about cancer, or a fetch that could not verify the licence. Each is on record so the gap is a decision, not an oversight.

  • Sentieon Closed-source commercial software; excluded by design.
  • REDCap Source available to consortium members only under a licence agreement; not open source.
  • COSMIC Data under a commercial licence for non-academic use and no open code; recorded on /data-sources/ instead.
  • Cancer Genome Interpreter (CGI) Web service from the Barcelona Biomedical Genomics Lab; the interpreter's own code is not published. Its open building blocks (IntOGen, boostDM, OncodriveFML, OncodriveCLUSTL) are recorded.
  • ClinGen No single open code base identified that could be verified as ClinGen's; the Allele Registry and curation interfaces are hosted services. Recorded via GA4GH VRS, which ClinGen co-develops.
  • GISTIC2 Distributed from the Broad's software page under a research licence, not an open repository.
  • ABSOLUTE Distributed from the Broad's software page under a research licence, not an open repository.
  • MutSigCV Distributed under a research licence from the Broad, not an open repository.
  • MiXCR Source available under a custom licence that restricts commercial use; not OSI open source.
  • CIBERSORT Web tool under a non-commercial academic licence; source not open.
  • Heidelberg CNS methylation classifier (production) Served as a web service; only the training code is public (recorded).
  • PLUTO (PathAI), Atlas (Aignostics), PHENOM-2 (Phenomic AI) Closed foundation models without released weights or code.
  • Med-Gemini No released weights or code.
  • NetMHCpan Academic licence, source not open.
  • Cancer Commons A non-profit service; the site states no open licence for its content and publishes no code.
  • CanRisk (BOADICEA) Free web tool under a non-commercial licence; code not open.
  • Tyrer-Cuzick (IBIS) Free for non-commercial use; code not open.
  • LIFEx and PRIMO Free-to-use radiomics and Monte Carlo software; source not open.
  • SEER*Stat Free software from the NCI, source not published.
  • UK Biobank, All of Us Access-controlled research resources rather than open-source projects; UK Biobank is recorded as a collection.
  • Open Source Ventilator projects Not about cancer; open hardware for respiratory support only.
  • GDSC, CTRP, PRISM Datasets rather than software; GDSC data are under a non-commercial licence.
  • Nextflow, Cromwell, Snakemake General workflow engines; oncology-specific pipelines that run on them are recorded instead.
  • SimpleITK, ANTs, fastMRI, lungmask General medical imaging tools with no oncology focus in their own description.
  • OpenMRS, Bahmni General electronic health records; the oncology modules used at AMPATH are deployment-specific.
  • Plastimatch Hosted on GitLab; the project page could not be verified to state a licence at the time of the fetch.
  • OpenREGGUI Registration on the project site is needed to reach the source; not verifiable without an account.
  • TOPAS (topasmc/topas) The pre-2024 TOPAS was source-available by registration; OpenTOPAS is the open release and is recorded.
  • Sybil (lung cancer risk from CT) The official repository could not be found on GitHub at the time of the fetch (only third-party wrappers); not recorded until it can be verified.
  • MOQUI (GPU proton Monte Carlo) No public repository found under the expected organisations.
  • CMSclassifier (colorectal consensus subtypes) The Sage Bionetworks repository named in the paper was not found; CMScaller is recorded instead.
  • CancerModels.Org / PDCM Finder The portal's code base could not be located under a verifiable organisation; the resource is recorded on /data-sources/.
  • HemOnc.org The site answered the fetch with a 403 bot challenge, so its stated licence could not be read.
  • Silicone breast phantoms (HardwareX 2025) The publisher page needs JavaScript to reach the article, so the licence could not be read from the fetch.

Method. The list is hand-curated in scripts/open-source-curated.ts (one record per project, category, openness and what in OnCo it relates to). scripts/fetch-open-source.ts then reads each repository from the GitHub REST API and each non-repository project from its own page, and records the licence as declared (SPDX id, or "custom" where GitHub cannot map the terms, "none stated" where there is no licence file, "not stated" where a page named none), the primary language, stars, the year the repository was created, the date of the last push and the first non-Zenodo DOI the README names. Summaries are OnCo's plain-English glosses; the project's own description travels with the record in the API as /api/v1/open-source.json. Stars measure attention, not quality; a licence read by machine is a starting point, not legal advice.

Missing a project or a link to a technology? Suggest an edit. Builders: the recipes show how to read this table from the API.