Skip to content

Latest commit

 

History

560 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MC2 Center Data Models

Data models and standard terms used by MC2 Center


GitHub release (latest by date) GitHub Release Date GitHub


🔎 Data Models Explorer: https://mc2-center.github.io/data-models/


Overview

This project contains the released versions of the JSON-LD schemas for the Cancer Complexity Knowledge Portal (CCKP), and more broadly, MC2 Center. You can learn more about the schemas/data models and other aspects of this project in our Data Models Explorer. The MC2 Center data model is in both CSV and JSON-LD format, and individual entity schemas are also exported as standalone JSON Schemas in ./json_schemas.

Beyond the original entity types (Dataset, Study, Publication, Grant, Educational Resource, File, Tool, Person), the model also covers Biospecimen, Individual, Model (non-human organism/cell line), and Data Catalog entities, as well as assay-level metadata for imaging (multiplexed/single-channel imaging), NanoString GeoMx Digital Spatial Profiler (DSP) spatial transcriptomics, bulk/single-cell sequencing, and 10x Genomics Visium spatial transcriptomics.

A companion kg-pipeline project builds an RDF knowledge graph from this model and live CCKP portal data - see Knowledge Graph Pipeline below.

Usage

Requires Python 3.10+.

pip install -r requirements.txt

# Full build: update valid values in all modules -> collate -> generate JSON Schemas
make all

Steps individually:

python update_valid_values.py   # reads modules/mapping.yaml, rewrites annotationProperty.csv Valid Values columns
make collate                    # concatenates all modules/*/annotationProperty.csv -> mc2.model.csv
make convert                    # converts mc2.model.csv -> mc2.model.jsonld
make generate-json              # python create_json_from_model.py <data types> -> json_schemas/

# Generate JSON schemas for specific data types only
python create_json_from_model.py Biospecimen Study Dataset

To build and preview the docs site locally:

mkdocs serve   # http://localhost:8000

See contributing guidelines for the full development and release process.

Knowledge Graph Pipeline

kg-pipeline/ is a separate, self-contained pipeline that converts this model plus live CCKP portal data (pulled from Synapse) into a queryable RDF knowledge graph: an extract → harmonize → map-to-RDF → validate build producing a LinkML/OWL schema and per-entity RDF instance data, with real ontology IRI mappings (NCIT, MONDO, EFO, OBI, ...) sourced from this model's own controlled vocabularies. Instance data is dual-typed against Biolink alongside the portal's own classes, and addressed by Synapse's own canonical IRI wherever a row already has one, rather than minting a second identifier. It also links the graph to related efforts - Data Catalog (native Synapse Dataset annotations), SCDM (Sage Common Data Model) federation, and sagebrain-model interoperability for assay-level metadata.

Built graphs are published to Synapse, and optionally to a SageBrain Neptune S3 bucket (a local-runnable equivalent of nf-osi/kg-pipeline's own S3 upload workflow) - each publish is paired with a small manifest.ttl PROV/VOID statement about the build, used as a lightweight trigger file by a downstream auto-loader.

It has its own README, Makefile, and Python environment (isolated from this repo's root requirements.txt). See kg-pipeline/README.md for setup, the full command reference, and design rationale, and the knowledge graph design page for the layer-by-layer build and validation design.

Folder Structure

.
├── docs/
├── json_schemas/
├── kg-pipeline/
├── modules/
├── scripts/
└── templates/

Documentation

All docs are located in the ./docs directory and are written in Markdown format. Some docs are generated before the site is built, which is handled by the hooks.py script in ./scripts.

Valid Values

Valid values are separated into modules (located in ./modules), where various pieces of the data model can be updated, including the standard terms/valid values.

Add a new valid value

When a new valid value needs to be added to the data model:

  1. Research the term and make sure we do not already have a synonym for it that exists. Using NCIt is excellent for this, though sometimes looking outside of NCIt is necessary. If we do currently have a synonym in use, add the valid value as a "non preferred term" in the applicable attribute CSV in ./modules. If not:

  2. Add the valid value in the "attribute" column of the applicable csv in the appropriate module folder. E.g. if a new tumor type needs to be added go to tumorType.csv and add the new term in the attribute column). Fill out the rest of the columns as completely as possible, this includes the description, the required column, parent column, source column, non-preferred terms column, the ontology identifier, url, NCIt Code, and any notes. Please make a note of who added it and the date.

  3. Be sure to look up any synonyms and add to the "non preferred terms" column. This will make annotating easier in the future.

Update a valid value

Please open a ticket and let the MC2 Center internal data team know the reasoning behind why a valid value should be updated/removed.

Annotation Templates

A collection of ready-for-use templates are available in ./templates, for curating and submitting metadata manifests to add/update entities on the CCKP.

How to Contribute

Thank you helping us continuously improve the MC2 Center data models! To contribute, please read our contributing guidelines on the docs site.

About

Versioned history of the MC2 Center data model

Resources

Contributing

Stars

6 stars

Watchers

4 watching

Forks

Releases

Used by

Contributors

Languages