Controlled-access data can identify a person (raw sequence reads, germline variants), so it is held in archives like dbGaP and EGA and released only to approved researchers under an agreement.
The European Genome-phenome Archive is a repository for potentially identifiable genetic, phenotypic and clinical data from biomedical research, held under secure storage (Wikipedia); dbGaP at the NCBI is the US counterpart. The GDC page on controlled data explains that TCGA's raw sequence and germline files require dbGaP authorisation through an institutional signing official, while processed, non-identifying data are open. Models trained on controlled data cannot release the data, and sometimes not the weights, without the same review.
Shares TCGA open versus controlled data tiers, TCGA / NCI Genomic Data Commons, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Local-only language models over patient data (privacy by architecture), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Local-only language models over patient data (privacy by architecture), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares TCGA / NCI Genomic Data Commons, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares HIPAA (Health Insurance Portability and Accountability Act), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.