Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
227 commits
Select commit Hold shift + click to select a range
bc2021d
Move cv files to shared
Bankso Apr 2, 2026
745fe28
Update duo from governance model
Bankso Apr 2, 2026
46218e4
Add governance-focused models
Bankso Apr 2, 2026
affe79d
Add updated governance attributes
Bankso Apr 2, 2026
9eb952f
Update mapping.yaml
Bankso Apr 2, 2026
51255dd
Update valid values
Bankso Apr 2, 2026
fa7e429
Update all_valid_values.csv
Bankso Apr 2, 2026
3eea5f5
Build model CSV
Bankso Apr 2, 2026
8522b0d
Build JSON-LD
Bankso Apr 2, 2026
1cd13bc
Add script to build JSON with curator
Bankso Apr 2, 2026
13ca6f2
Use curator-based JSON build tool
Bankso Apr 2, 2026
229f765
Remove unused schemas
Bankso Apr 2, 2026
0a30990
Remove outdated schemas
Bankso Apr 2, 2026
b8a4128
Delete cckp_datacatalog.jsonld
Bankso Apr 2, 2026
b0ba90b
Add json schemas from curator
Bankso Apr 2, 2026
0334e96
Rebuild JSON schemas with synapseclient 4.12
Bankso Apr 7, 2026
b4ce174
Add base NGS Level 1 attributes to RNA Level 1
Bankso Apr 24, 2026
617fbfe
Build schemas from CSV, not JSON-LD
Bankso Apr 24, 2026
2bd316f
Remove unused functions
Bankso Apr 24, 2026
f3858c3
Make model
Bankso Apr 24, 2026
c9f3de3
Integrate FileView attributes for curator adoption
Bankso Apr 25, 2026
dbcea81
Update mc2.model.csv
Bankso Apr 25, 2026
ec6da3b
Use file-specific data use codes attribute
Bankso Apr 27, 2026
faaffa3
Use File Data Use Codes attribute
Bankso Apr 27, 2026
a42905d
Update mc2.model.csv
Bankso Apr 27, 2026
b5a8b6c
Create SequencingRNALevel1.json
Bankso Apr 27, 2026
09cee5b
Use catalog-specific DUO attribute
Bankso Apr 27, 2026
8483e25
Use DatasetView-specific DUO attribute
Bankso Apr 27, 2026
7ca2f30
Update docs link
Bankso Apr 27, 2026
9857dda
Use Study specific DUO attribute
Bankso Apr 27, 2026
06dfe44
Use catalog-specific DUO attribute
Bankso Apr 27, 2026
b41f9e7
Update annotationProperty.csv
Bankso Apr 27, 2026
8813490
Only use DUO valid values for Study
Bankso Apr 27, 2026
55653da
Collate model CSV
Bankso Apr 27, 2026
d8da15b
Don't use conditional DUO elements in DSP
Bankso Apr 27, 2026
a03bcda
Pass JSON schema types via makefile
Bankso Apr 27, 2026
484b71b
Pull data types from CLI instead
Bankso Apr 27, 2026
2d45873
Update all_valid_values.csv
Bankso Apr 27, 2026
ff5eff7
Update annotationProperty.csv
Bankso Apr 27, 2026
2de8787
Update mc2.model.csv
Bankso Apr 27, 2026
1ecde30
Fix link
Bankso Apr 27, 2026
6a947ed
rebuild model
Bankso Apr 27, 2026
8ddb580
Rebuild JSON schemas
Bankso Apr 27, 2026
e4970ba
Update link
Bankso Apr 27, 2026
8073f5a
Re-separate fileview attributes
Bankso Apr 27, 2026
357cc38
Remove specimen info from file attributes
Bankso Apr 28, 2026
79fb283
Remove specimen info from file schema
Bankso Apr 28, 2026
6df50a7
Combine current file attributes with RNA seq assay
Bankso Apr 28, 2026
d2b3044
Collate model
Bankso Apr 28, 2026
93825b1
Update SequencingRNALevel1.json
Bankso Apr 28, 2026
7b46e5e
Update FileView.json
Bankso Apr 28, 2026
8cc3124
Update DataDSP.json
Bankso Apr 29, 2026
33987ad
Update primary key descriptions
Bankso May 12, 2026
e2c5908
Update to current curator CSV format
Bankso May 13, 2026
d5dee2c
Update to current curator CSV format
Bankso May 13, 2026
8253958
Update to current curator CSV format
Bankso May 13, 2026
f2d531f
Update to current curator CSV format
Bankso May 13, 2026
9d93ecb
Update to current curator CSV format
Bankso May 13, 2026
f86b0c7
Update to current curator CSV format
Bankso May 13, 2026
dcf70b1
Update to current curator CSV format
Bankso May 13, 2026
c7b9485
Pattern rule is uri, not url
Bankso May 13, 2026
eb10f71
Reformatting for curator model compliance
Bankso May 28, 2026
6b8820e
Reformatting for curator model compliance (part 2)
Bankso Jun 10, 2026
7476142
Update collection attributes
Bankso Jun 10, 2026
dfd813e
Add primary_key property label and regex patterns
Bankso Jun 11, 2026
e59f143
Add common file attributes to file-based models
Bankso Jun 11, 2026
bb04519
Collate CSVs
Bankso Jun 11, 2026
1cf5460
Fix copy/paste error
Bankso Jun 11, 2026
378c28c
Remove old schema
Bankso Jun 11, 2026
2d744bc
Add update curator JSON schemas
Bankso Jun 11, 2026
cc8c7ca
Update id pattern
Bankso Jun 22, 2026
001b0d4
Update id pattern
Bankso Jun 22, 2026
0c91cdd
Update id pattern
Bankso Jun 22, 2026
70b3059
Update key patterns to align with ids
Bankso Jun 22, 2026
645fb33
Collate model
Bankso Jun 22, 2026
5cd2589
Add updated JSON
Bankso Jun 22, 2026
1c49d0f
Update File Design description
Bankso Jun 24, 2026
c293346
Add CLAUDE.md
Bankso Jun 25, 2026
2042e02
Use OLS MCP claude skill to prototype ontology mapping tools
Bankso Jun 27, 2026
984fef4
Add results of CDE matching by claude skill
Bankso Jul 9, 2026
8a3af11
Update annotationProperty CSVs and add review docs
Bankso Jul 21, 2026
33b5f94
Remove CC1.0 license
vpchung Jul 1, 2026
5840772
Add Apache 2.0 License (#254)
vpchung Jul 1, 2026
f423f07
Governance model integration and curator updates (#247)
Bankso Jul 20, 2026
feecc80
Fix duplicate DependsOn key in VisiumRNALevel2 and resync DUO codes
Bankso Jul 22, 2026
4ec7a52
Document previously-undocumented modules and rebuild Full Field Refer…
Bankso Jul 22, 2026
d4b467b
Add/fix CDE mappings surfaced by mapping-description-update's CDE review
Bankso Jul 22, 2026
384910b
Regenerate mc2.model.csv after CDE mapping fixes
Bankso Jul 22, 2026
e27a240
Move mapping artifacts to worklists folder
Bankso Jul 22, 2026
1e01fb7
Surface key/CDE/ontology mappings in the docs site
Bankso Jul 22, 2026
aac37bc
Backfill descriptions/ontology terms for 15 small controlled-vocab files
Bankso Jul 22, 2026
eca8195
Fill in DUO ontology identifiers/URLs mechanically in shared/duo.csv
Bankso Jul 22, 2026
8aa9d16
Backfill descriptions/ontology terms for sequencing small CV files
Bankso Jul 22, 2026
90e83c8
Backfill descriptions/ontology terms for remaining Wave 1 CV files
Bankso Jul 22, 2026
8c5a746
Normalize all controlled-vocabulary CSVs to LF line endings
Bankso Jul 23, 2026
5212bd2
Backfill descriptions/ontology terms for Wave 2 medium CV files
Bankso Jul 23, 2026
dcd8856
Add descriptions and ontology mappings for toolFormat
Bankso Jul 23, 2026
3b09b27
Don't list non-preferred terms in enum tables
Bankso Jul 23, 2026
f17a42b
Add ontology identifiers and clean CSV headers
Bankso Jul 24, 2026
6fcb053
Add ontology IDs and metadata to shared CSVs
Bankso Jul 31, 2026
cecd8da
Add CCKP knowledge-graph pipeline
Bankso Aug 14, 2026
960db95
Normalize mapping URIs and update CSVs
Bankso Aug 17, 2026
4aa31ab
Fill in ontology mappings for modules/tool/ and lymphStage N-stage codes
Bankso Aug 17, 2026
5625633
Vendor csv-to-linkml converter for reproducibility
Bankso Aug 18, 2026
f58b0fb
Normalize Study deidentification property names
Bankso Aug 18, 2026
30dd0a2
Add suggest-mappings and update-baseline targets
Bankso Aug 18, 2026
0408a05
Add suggest-mappings workflow and SSSOM files
Bankso Aug 18, 2026
970176d
Add ROR and ontology IDs to CSVs
Bankso Aug 18, 2026
7c2da66
Align kg-pipeline with sagebrain-model conventions (Part A)
Bankso Aug 18, 2026
e9d8196
Add MC2 assay-metadata KG, linked into sagebrain-model (Part B)
Bankso Aug 18, 2026
3923961
Add .gitkeep placeholders for mc2_assay data dirs
Bankso Aug 18, 2026
1c5fd6f
Add Synapse secure-publish target for both kg-pipeline outputs
Bankso Aug 19, 2026
c0dddfb
Merge branch 'main' into database-model-kg
Bankso Aug 19, 2026
c91d04c
Correct stale malformed_cv_terms.csv claim in kg-pipeline README
Bankso Sep 1, 2026
c8f64c9
Attach File Tissue/File Tumor Type to every File-attribute class
Bankso Sep 1, 2026
a904146
Promote grantNumber to a shared top-level slot in cckp_portal schema
Bankso Sep 1, 2026
7ab7232
Remove redundant Tissue/Tumor Type harmonization in link_sagebrain.py
Bankso Sep 1, 2026
fe877e9
Add plans/ for SCDM alignment and kg-pipeline README-fix proposals
Bankso Sep 1, 2026
890520c
Vendor SageCommonDataModel + add Makefile targets for SCDM interop
Bankso Sep 1, 2026
e379d01
Add institution/consortium crosswalks to SCDM entities
Bankso Sep 1, 2026
879c822
Link the CCKP graph to SCDM Organization/Program/Person entities
Bankso Sep 1, 2026
5717fef
Document SCDM interop in kg-pipeline README
Bankso Sep 1, 2026
e2f7caf
Record SCDM alignment steps 1-6 Implementation Report
Bankso Sep 1, 2026
cb2346a
Add Institution_id/Consortium_id + shared Institution Key attribute
Bankso Sep 1, 2026
6a0edeb
Wire Institution/Consortium/PersonView Key into Grant/Project/Person …
Bankso Sep 1, 2026
ba90b9f
Regenerate mc2.model.csv + kg-pipeline schema for Institution/Consort…
Bankso Sep 1, 2026
6e86cac
Record SCDM alignment step 7 Implementation Report
Bankso Sep 1, 2026
5734464
Align modules/dataCatalog with real Synapse Dataset entity annotations
Bankso Sep 2, 2026
c651408
Extend existing CVs for dataCatalog fields, preferring reuse over new…
Bankso Sep 2, 2026
8e2d7f4
Update modules/mapping.yaml for dataCatalog's realigned CVs
Bankso Sep 2, 2026
b3b68fa
Regenerate mc2.model.csv + kg-pipeline schema for dataCatalog changes
Bankso Sep 2, 2026
09058e6
Add Data Catalog pipeline stage: extract, triples, merge + tests
Bankso Sep 2, 2026
158f49c
Document the Data Catalog stage in kg-pipeline README
Bankso Sep 2, 2026
f840848
Record Data Catalog integration Implementation Report
Bankso Sep 2, 2026
cd942ec
Add CRDC CDE reference data, mapping report, and integration plan
Bankso Sep 3, 2026
7fb45a6
CRDC CDE integration for Individual/Model/Biospecimen clinical modules
Bankso Sep 3, 2026
a57a593
CRDC CDE integration + attribute consolidation for remaining modules
Bankso Sep 3, 2026
a3b0c8b
Update modules/mapping.yaml for the CRDC integration + consolidation
Bankso Sep 3, 2026
f0bcf2b
Regenerate mc2.model.csv + all_valid_values.csv
Bankso Sep 3, 2026
6552e4d
Retire redundant FK-shadowing attributes + naming fixes
Bankso Sep 3, 2026
1755c05
Consolidate Grant/Project/Study Investigator into shared Investigator
Bankso Sep 3, 2026
eaf3ddd
Legacy CDE cleanup, description/columnType fixes
Bankso Sep 3, 2026
58fed0e
Add Study Number of Samples, resolving the long-standing dangling ref…
Bankso Sep 3, 2026
dcb6982
Document round 7 in the Implementation Report
Bankso Sep 3, 2026
872023a
Update modules/mapping.yaml for round 7's attribute changes
Bankso Sep 3, 2026
8db1ad2
Regenerate mc2.model.csv + all_valid_values.csv
Bankso Sep 3, 2026
07918f9
Backfill real Valid Values from live caDSR data
Bankso Sep 3, 2026
8b53811
Add live CDE reference data, update mapping report with real CV findings
Bankso Sep 3, 2026
55bacdd
Document Round 8 in the Implementation Report
Bankso Sep 3, 2026
6c8ea2c
Regenerate mc2.model.csv + all_valid_values.csv
Bankso Sep 3, 2026
3f77119
Re-scope Biospecimen Type Category from CDE 11253427 to CDE 12445832
Bankso Sep 3, 2026
9780fdb
Un-map Tool Entity Role from CDE 2201713
Bankso Sep 3, 2026
7eb2172
Reformat Tumor Grade to match CDE 11325685's real serialization
Bankso Sep 3, 2026
77ca6e0
Update crdc_cde_mapping_report.csv for the three Round 9 decisions
Bankso Sep 3, 2026
51e62c0
Document Round 9 in the Implementation Report
Bankso Sep 3, 2026
dba153e
Regenerate mc2.model.csv + all_valid_values.csv
Bankso Sep 3, 2026
4110108
Remove Tool Entity Role's legacy CDE:2201713 tag entirely
Bankso Sep 3, 2026
9d42c41
Remove redundant Biospecimen Analyte Type; document 2 final decisions
Bankso Sep 3, 2026
5b6f54f
Fix Biospecimen Taxonomy ID: columnType number -> string
Bankso Sep 3, 2026
a378f83
Switch make convert to synapseclient's curator extension
Bankso Sep 3, 2026
12fa7f6
Remove schematicpy from requirements.txt; bump synapseclient to 4.13.0
Bankso Sep 3, 2026
14a36af
Regenerate kg-pipeline mc2_model LinkML schema from this session's mo…
Bankso Sep 4, 2026
debfa7e
Fix kg-pipeline scripts for renamed/consolidated MC2 attributes
Bankso Sep 4, 2026
4b6946e
Update kg-pipeline tests for consolidated attributes and the retired …
Bankso Sep 4, 2026
caa4941
Refresh license/species SSSOM crosswalks after enum consolidation
Bankso Sep 4, 2026
e3348ad
Consolidate per-entity SSSOM crosswalks into per-enum files
Bankso Sep 4, 2026
9f5ff05
Fix stale mc2_enum references in cckp_portal.linkml.yaml, drop the re…
Bankso Sep 4, 2026
bf7ee2b
Update kg-pipeline coverage baseline after consolidation
Bankso Sep 4, 2026
ac514f2
Document the kg-pipeline propagation round in plans/crdc_cde_integrat…
Bankso Sep 4, 2026
52f0396
Register a real CRDC_CDE: prefix, resolving the unregistered-prefix w…
Bankso Sep 4, 2026
01b1a3c
Verify and document: Grant.consortium repoint rejected, link_sagebrai…
Bankso Sep 4, 2026
a6beddb
Document Round 11: kg-pipeline follow-ups verified and resolved
Bankso Sep 4, 2026
dc5ad59
Complete the Grant.consortium human review: 11 SCDM Programs sourced
Bankso Sep 4, 2026
4901a69
Add combined-kg/full-kg targets so DataCatalog/SCDM merges can't go s…
Bankso Sep 4, 2026
707fe7e
Fix DataCatalog license/dataUseModifiers field-name drift in kg-pipeline
Bankso Sep 4, 2026
ade1326
Surface caDSR CDE mappings in the docs site's field reference tables
Bankso Sep 4, 2026
25c3e73
Add DataCatalog to the docs site
Bankso Sep 4, 2026
f1e03cb
Remove QC model
Bankso Sep 4, 2026
cd901f2
Add a Usage section to the root README
Bankso Sep 8, 2026
40dabb9
Separate kg-pipeline's README into usage docs and decision rationale
Bankso Sep 8, 2026
6130f03
Fix malformed rows in shared/assay.csv breaking strict CSV parsing
Bankso Sep 9, 2026
1d07453
Finish Biospecimen Acquisition Method's CRDC CDE alignment; remove or…
Bankso Sep 9, 2026
063338a
Backfill CV descriptions from existing ontology mappings (OLS + SPDX)
Bankso Sep 9, 2026
62e074a
Synthesize descriptions for CV terms with no real ontology equivalent
Bankso Sep 9, 2026
6577de7
Docs: render "No description provided" instead of nan; link CDE tags
Bankso Sep 9, 2026
cefd07f
Add cv_quality_pass plan + implementation report; regenerate model
Bankso Sep 9, 2026
2dbf9a7
File module: free up Longitudinal timepoint fields, add Batch Identifier
Bankso Sep 9, 2026
019798d
Regenerate mc2.model.csv/jsonld + JSON schemas for File module changes
Bankso Sep 9, 2026
e70e569
Move File Batch Identifier's ontology ref to Properties; propagate to…
Bankso Sep 9, 2026
3ef833e
Regenerate model + JSON schemas for File Batch Identifier propagation
Bankso Sep 9, 2026
0b1a2b5
kg-pipeline: register the MIxS prefix for File Batch Identifier
Bankso Sep 9, 2026
4f948a6
Propagate model changes to kg-pipeline schema artifacts
Bankso Sep 9, 2026
aae3cba
harmonize.py/build_triples.py: split comma-crammed scalar CV values
Bankso Sep 9, 2026
2c43af4
Update kg-pipeline provenance/coverage baseline for live-data re-run
Bankso Sep 9, 2026
e203812
suggest_mappings.py: pluggable registry backends (add SPDX)
Bankso Sep 9, 2026
729beca
suggest_mappings.py: make registry selection prefix-driven, not path-…
Bankso Sep 9, 2026
fe7433f
suggest_mappings.py: skip confirmed-unmappable values, fix short-code…
Bankso Sep 9, 2026
fa9d63c
cckp_portal.linkml.yaml: align 5 CCKP classes to schema.org classes
Bankso Sep 10, 2026
fda0bd7
Document schema.org class alignment in kg-pipeline README + decisions…
Bankso Sep 10, 2026
2449408
plans: add schema class alignment plan (implemented) + MONDO/UBERON f…
Bankso Sep 10, 2026
5a2eb91
crosswalk_ontology.py: gate MONDO/UBERON crosswalk rows behind a revi…
Bankso Sep 10, 2026
4e1d8a5
Add link_ontology_crosswalk.py: promote reviewed MONDO/UBERON edges
Bankso Sep 10, 2026
1038223
Document MONDO/UBERON crosswalk promotion in README + decisions log
Bankso Sep 10, 2026
d3fec49
plans: append Implementation Report to mondo_uberon_federation_promot…
Bankso Sep 10, 2026
9fd071e
Update .gitignore
Bankso Sep 10, 2026
ec48c14
Update README to remove DCA config and link
Bankso Sep 10, 2026
191c8f1
crosswalk_scdm.py: fix broken consortium crosswalk, point at consorti…
Bankso Sep 10, 2026
5585deb
Register Biospecimen Type Category CV, fix CDE-drift Analyte row
Bankso Sep 10, 2026
9da0482
Model-wide cleanup: remove process/historical narration from Descript…
Bankso Sep 10, 2026
30ec13c
Regenerate derived artifacts for Biospecimen Type Category fix + desc…
Bankso Sep 10, 2026
ceb1093
Biospecimen Type Category: achieve full CDE 12445832 parity (19/19)
Bankso Sep 10, 2026
9a83080
Model-wide CDE alignment audit: fix 4 confirmed missing-value gaps
Bankso Sep 10, 2026
53376ec
Regenerate derived artifacts for full-parity + CDE-alignment fixes
Bankso Sep 10, 2026
e187a64
Split File Assay Category out of Assay/DSP Dataset Assay for CRDC_CDE…
Bankso Sep 11, 2026
03d6437
Resolve Image Assay Type / GeoMx DSP Assay Type CDE 7789196 conflict
Bankso Sep 11, 2026
6dfef5a
Remove wrong CDE:8037927 tag from Biospecimen Embedding Medium
Bankso Sep 11, 2026
2597356
Biospecimen Preservation Method: achieve full CDE/CRDC_CDE 8028962 pa…
Bankso Sep 11, 2026
a6c5bf0
Biospecimen Composition: trim CV to exact CDE/CRDC_CDE 12922545 parit…
Bankso Sep 11, 2026
26a5205
Split File Format / Dataset File Formats CVs for CDE/CRDC_CDE 1141692…
Bankso Sep 11, 2026
73e102d
NGS Library Strategy: add 4 missing CDE/CRDC_CDE 6273393 values
Bankso Sep 11, 2026
325e449
Sex: trim CV to exact CDE/CRDC_CDE 7572817 parity (3/3, approved)
Bankso Sep 11, 2026
7baf373
Remove unresolvable CDE tags from Study_id and DSP Data Use Codes
Bankso Sep 11, 2026
ac37599
plans/cde_alignment_audit.md: finalize Implementation Report
Bankso Sep 11, 2026
81ca56c
Regenerate derived artifacts for full CDE/CRDC_CDE alignment pass
Bankso Sep 11, 2026
b9a1c39
kg-pipeline: skip rows with a blank declared identifier instead of cr…
Bankso Sep 11, 2026
d487af8
kg-pipeline: regenerate schema + re-extract/harmonize provenance for …
Bankso Sep 11, 2026
f7d63ec
qc_model/qc_attribute_mapping.csv: align with current Publication/Dat…
Bankso Sep 11, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
3 changes: 3 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
* text=auto eol=lf

*.csv text eol=lf
3 changes: 2 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -78,4 +78,5 @@ kg-pipeline/data/mc2_assay/rdf/*
kg-pipeline/__pycache__/
kg-pipeline/**/__pycache__/
kg-pipeline/.venv/
.pytest_cache/
.pytest_cache/
/.vscode
64 changes: 64 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## What this repo is

Data models and controlled vocabularies for the [Cancer Complexity Knowledge Portal](https://cancercomplexity.synapse.org/) (CCKP). The model is maintained as CSV files in domain-specific `modules/`, collated into `mc2.model.csv`, and converted to JSON-LD/JSON Schema via `synapseclient.extensions.curator` (the actively-maintained successor to [schematicpy](https://pypi.org/project/schematicpy/), which Sage Bionetworks has announced will be retired by end of 2026) for use by the [Data Curator App](https://dca.app.sagebionetworks.org/).

## Commands

```bash
# Install dependencies (Python 3.10+)
pip install -r requirements.txt

# Full build: update valid values in all modules → collate → generate JSON Schemas
make all

# Steps individually:
python update_valid_values.py # reads modules/mapping.yaml, rewrites annotationProperty.csv Valid Values columns
make collate # concatenates all modules/*/annotationProperty.csv → mc2.model.csv
make convert # python convert_model_to_jsonld.py → mc2.model.jsonld
make generate-json # python create_json_from_model.py <data types> → json_schemas/

# Generate JSON schemas for specific data types only
python create_json_from_model.py Biospecimen Study Dataset

# QC model build (NOTE: qc_convert still calls the old, non-functional
# `schematic schema convert` against qc_model/mc2_qc.model.csv, a separately
# hand-formatted copy of the model with schematicpy's older required columns
# -- unlike `make convert`, this was not migrated; needs a follow-up decision)
make qc

# Docs dev server
mkdocs serve # http://localhost:8000
```

## Architecture

### The module + collation pattern

Each domain lives in `modules/<domain>/`:
- `annotationProperty.csv` — attribute definitions for that domain (type, description, valid values, validation rules)
- One CSV per controlled vocabulary (e.g., `specimenType.csv`, `fixative.csv`) — the actual enumerated terms

`modules/mapping.yaml` is the central registry: it maps each attribute name to the CV CSV that provides its valid values. `update_valid_values.py` reads this file and rewrites the `Valid Values` column in each `annotationProperty.csv`.

`make collate` then concatenates all `annotationProperty.csv` files into `mc2.model.csv` (the header comes from the consortium module, body from all others via `tail -n +2`).

### CSV conventions
- All columns read/written as `dtype=str` — `TRUE`/`FALSE` must not become `True`/`False`
- No NaN — use empty strings (`keep_default_na=False`)
- Index on `Attribute` column when updating via pandas

### PR requirements
PRs to main must have exactly one semantic label: `major`, `minor`, `patch`, or `non-release`. The `pr-check.yml` workflow enforces this.

### CI workflows (`.github/workflows/`)
| Workflow | Trigger | What it does |
|----------|---------|--------------|
| `build-jsonld.yml` | PR to main (module changes) | `make all` to validate collation + schema conversion |
| `build-docs.yml` | Push to main | Builds MkDocs site → GitHub Pages |
| `pr-check.yml` | PR events | Validates semantic label |
| `google-sheet-sync.yml` | Scheduled/manual | Syncs RFC Google Sheets to a CSV branch |
| `create-release.yml` | Manual trigger | Creates GitHub release with version bump |
10 changes: 2 additions & 8 deletions Makefile
Original file line number Diff line number Diff line change
@@ -1,10 +1,7 @@
CSV := mc2.model.csv
QC := ./qc_model/mc2_qc.model.csv
DATA := DataDSP Study FileView PublicationView GrantView ToolView EducationalResource DatasetView DataCatalog Biospecimen Model Individual SequencingLevel1 SequencingLevel2 SequencingLevel3 SequencingRNALevel1 ImagingLevel1 ImagingLevel2 ImagingLevel3Image ImagingLevel3Segments ImagingLevel4 NanoStringGeoMxAuxiliaryFiles NanoStringGeoMxDSPImaging NanoStringGeoMxDSPLevel1 NanoStringGeoMxDSPLevel2 NanoStringGeoMxDSPLevel3 NanoStringGeoMXROISegmentAnnotation 10xVisiumAuxiliaryFiles 10xVisiumRNALevel1 10xVisiumRNALevel2 10xVisiumRNALevel3 10xVisiumRNALevel4

all: collate generate-json

qc: collate qc_convert
all: collate convert generate-json

collate:
@echo "Collating module components..."
Expand All @@ -13,10 +10,7 @@ collate:
tail -n +2 -q modules/*/annotationProperty.csv >> ${CSV}

convert:
schematic schema convert ${CSV}

qc_convert:
schematic schema convert ${QC}
python convert_model_to_jsonld.py

generate-json:
python create_json_from_model.py ${DATA}
40 changes: 32 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,16 +26,14 @@

🔎 **Data Models Explorer**: https://mc2-center.github.io/data-models/

📊 **Data Curator App**: https://dca.app.sagebionetworks.org/

---

## Overview

This project contains the released versions of the JSON-LD schemas for the
[Cancer Complexity Knowledge Portal] (CCKP), and more broadly, MC2 Center.
You can learn more about the schemas/data models and other aspects of this
project in our portal documentation - coming soon! The MC2 Center data model
project in our Data Models Explorer. The MC2 Center data model
is in both CSV and JSON-LD format, and individual entity schemas are also
exported as standalone JSON Schemas in `./json_schemas`.

Expand All @@ -46,21 +44,47 @@ assay-level metadata for imaging (multiplexed/single-channel imaging),
NanoString GeoMx Digital Spatial Profiler (DSP) spatial transcriptomics,
bulk/single-cell sequencing, and 10x Genomics Visium spatial transcriptomics.

## Usage

Requires Python 3.10+.

```bash
pip install -r requirements.txt

# Full build: update valid values in all modules -> collate -> generate JSON Schemas
make all
```

Steps individually:

```bash
python update_valid_values.py # reads modules/mapping.yaml, rewrites annotationProperty.csv Valid Values columns
make collate # concatenates all modules/*/annotationProperty.csv -> mc2.model.csv
make convert # converts mc2.model.csv -> mc2.model.jsonld
make generate-json # python create_json_from_model.py <data types> -> json_schemas/

# Generate JSON schemas for specific data types only
python create_json_from_model.py Biospecimen Study Dataset
```

To build and preview the docs site locally:

```bash
mkdocs serve # http://localhost:8000
```

See [contributing guidelines] for the full development and release process.

## Folder Structure

```
.
├── dca_config/
├── docs/
├── modules/
├── scripts/
└── templates/
```

### DCA Configuration

MC2 Center's configurations for the DCA is located in `./dca_config`.

### Documentation

All docs are located in the `./docs` directory and are written in Markdown
Expand Down
Loading