DuckDB is an in-process analytical database that queries Parquet files with SQL; PostgreSQL is the standard client-server database; either can hold the metadata spine that joins patients, specimens and assays.
DuckDB is an open-source column-oriented relational database designed for high performance on analytical queries in embedded configurations (Wikipedia); PostgreSQL is a free relational database emphasising extensibility and SQL compliance (Wikipedia). A research data layer typically keeps the large arrays in HDF5, Zarr or Parquet and a small catalogue in one of these, modelling the hierarchy patient to specimen to assay with provenance on every row so that a query can say which files went into a result.
Shares Array and table formats: HDF5, Zarr, OME-Zarr, Parquet, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares TCGA barcode, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Provenance fields for research data, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Provenance fields for research data, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Array and table formats: HDF5, Zarr, OME-Zarr, Parquet, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Provenance fields for research data, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Provenance fields for research data, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.
Shares Provenance fields for research data, Cancer AI vocabulary (CanSim terms map) and the tag cansim-terms.