Provenance fields record where each data file came from and how it was processed (source, accession, pipeline version, genome build, date, checksum), so a result can be traced and reproduced.
Data lineage is the tracking of how data is generated, transformed and used across systems, documenting its origins and movements so errors can be traced (Wikipedia). For genomic data the minimum fields are the source dataset and accession, the pipeline and its version, the genome build and annotation release, the download date and a checksum. Content-addressable storage takes the last step further by keying each raw file on the hash of its contents, so the same bytes are stored once and any change is detected (Wikipedia).
Shares Ensembl gene ID, Genome builds: GRCh38 versus hg19 (GRCh37), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Zenodo DOIs for data and code, Reproducibility and negative results, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Analytical databases as a data catalogue (DuckDB, PostgreSQL), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Units of measurement ontology (UO), Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Zenodo DOIs for data and code, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Zenodo DOIs for data and code, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Reproducibility and negative results, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares TCGA barcode, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.