🔎 Data Models Explorer: https://mc2-center.github.io/data-models/
This project contains the released versions of the JSON-LD schemas for the
Cancer Complexity Knowledge Portal (CCKP), and more broadly, MC2 Center.
You can learn more about the schemas/data models and other aspects of this
project in our Data Models Explorer. The MC2 Center data model
is in both CSV and JSON-LD format, and individual entity schemas are also
exported as standalone JSON Schemas in ./json_schemas.
Beyond the original entity types (Dataset, Study, Publication, Grant, Educational Resource, File, Tool, Person), the model also covers Biospecimen, Individual, Model (non-human organism/cell line), and Data Catalog entities, as well as assay-level metadata for imaging (multiplexed/single-channel imaging), NanoString GeoMx Digital Spatial Profiler (DSP) spatial transcriptomics, bulk/single-cell sequencing, and 10x Genomics Visium spatial transcriptomics.
A companion kg-pipeline project builds an RDF knowledge
graph from this model and live CCKP portal data - see
Knowledge Graph Pipeline below.
Requires Python 3.10+.
pip install -r requirements.txt
# Full build: update valid values in all modules -> collate -> generate JSON Schemas
make allSteps individually:
python update_valid_values.py # reads modules/mapping.yaml, rewrites annotationProperty.csv Valid Values columns
make collate # concatenates all modules/*/annotationProperty.csv -> mc2.model.csv
make convert # converts mc2.model.csv -> mc2.model.jsonld
make generate-json # python create_json_from_model.py <data types> -> json_schemas/
# Generate JSON schemas for specific data types only
python create_json_from_model.py Biospecimen Study DatasetTo build and preview the docs site locally:
mkdocs serve # http://localhost:8000See contributing guidelines for the full development and release process.
kg-pipeline/ is a separate, self-contained pipeline that
converts this model plus live CCKP portal data (pulled from Synapse) into a
queryable RDF knowledge graph: an extract → harmonize → map-to-RDF →
validate build producing a LinkML/OWL schema
and per-entity RDF instance data, with real ontology IRI mappings (NCIT,
MONDO, EFO, OBI, ...) sourced from this model's own controlled vocabularies.
Instance data is dual-typed against Biolink
alongside the portal's own classes, and addressed by Synapse's own canonical
IRI wherever a row already has one, rather than minting a second identifier.
It also links the graph to related efforts - Data Catalog (native Synapse
Dataset annotations), SCDM (Sage Common Data Model) federation, and
sagebrain-model
interoperability for assay-level metadata.
Built graphs are published to Synapse, and optionally to a SageBrain Neptune
S3 bucket (a local-runnable equivalent of
nf-osi/kg-pipeline's own S3 upload
workflow) - each publish is paired with a small manifest.ttl PROV/VOID
statement about the build, used as a lightweight trigger file by a
downstream auto-loader.
It has its own README, Makefile, and Python environment (isolated from
this repo's root requirements.txt). See
kg-pipeline/README.md for setup, the full
command reference, and design rationale, and the
knowledge graph design page
for the layer-by-layer build and validation design.
.
├── docs/
├── json_schemas/
├── kg-pipeline/
├── modules/
├── scripts/
└── templates/
All docs are located in the ./docs directory and are written in Markdown
format. Some docs are generated before the site is built, which is handled
by the hooks.py script in ./scripts.
Valid values are separated into modules (located in ./modules), where
various pieces of the data model can be updated, including the standard
terms/valid values.
When a new valid value needs to be added to the data model:
-
Research the term and make sure we do not already have a synonym for it that exists. Using NCIt is excellent for this, though sometimes looking outside of NCIt is necessary. If we do currently have a synonym in use, add the valid value as a "non preferred term" in the applicable attribute CSV in
./modules. If not: -
Add the valid value in the "attribute" column of the applicable csv in the appropriate module folder. E.g. if a new tumor type needs to be added go to
tumorType.csvand add the new term in the attribute column). Fill out the rest of the columns as completely as possible, this includes the description, the required column, parent column, source column, non-preferred terms column, the ontology identifier, url, NCIt Code, and any notes. Please make a note of who added it and the date. -
Be sure to look up any synonyms and add to the "non preferred terms" column. This will make annotating easier in the future.
Please open a ticket and let the MC2 Center internal data team know the reasoning behind why a valid value should be updated/removed.
A collection of ready-for-use templates are available in ./templates, for
curating and submitting metadata manifests to add/update entities on the
CCKP.
Thank you helping us continuously improve the MC2 Center data models! To contribute, please read our contributing guidelines on the docs site.