TCGA data comes in two tiers: open files (gene expression, somatic mutations, clinical tables, slide images) anyone can download, and controlled files (raw reads, germline variants) that need dbGaP approval.
The GDC's access documentation divides its data into open access, which requires no authorisation and carries only a ban on attempting to re-identify participants, and controlled access, which requires an approved dbGaP application because the files could identify a person. Expression matrices, somatic MAFs, copy number segments, methylation betas, RPPA and diagnostic slides are open; BAM files and germline calls are controlled. Most public cancer models are trained entirely on the open tier.
Shares TCGA barcode, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Controlled-access genomic data (dbGaP, EGA), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares TCGA / NCI Genomic Data Commons, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares TCGA barcode, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.