Create and use mapping files for secondary (retired/withdrawn) biological database identifiers and labels to primary (current) identifiers and labels.
Outputs mappings in SSSOM format by default. Subject
ids and labels (subject_id, subject_label) are secondary, objects are
primary.
uv pip install pysec2priOr install from source:
uv pip install git+https://github.com/jmillanacosta/pysec2pri.gitMost sources have two commands. ids maps retired identifiers to current ones,
labels maps old labels to current ones:
pysec2pri hgnc ids
pysec2pri hgnc labelsRun pysec2pri --help to see every source, and pysec2pri <source> ids --help
for one source's options. Input files are downloaded unless you pass them:
pysec2pri hgnc ids --withdrawn withdrawn.txt --complete hgnc_complete_set.txtIn Python there are two functions, one per kind:
from pysec2pri import generate_ids, generate_labels, sources
sources() # every source
sources("labels") # sources with labels
hgnc = generate_ids("hgnc")
chebi = generate_labels("chebi", subset="3star", version="350")
ensembl = generate_ids("ensembl", version="115", species="9606")Both return an SSSOM MappingSet: IdMappingSet or LabelMappingSet. Subjects
are secondary, objects are primary. Options that a source does not have are
ignored, so species on HGNC does nothing.
The default output is SSSOM TSV.
Use a mapping set to update your own data. Labels:
from pysec2pri import generate_labels, resolve_labels
chebi = generate_labels("chebi")
resolve_labels(["Glucose", "ATP", "Guanine"], chebi)Identifiers in a dataframe:
from pysec2pri import generate_ids, update_ids
ensembl = generate_ids("ensembl", version="115", species="9606")
df_with_new_column = update_ids(mapping_set=ensembl, ids=df, at="Ensembl_id") # `at` is the column nameOr from the command line, given a TSV file gene_ex.tsv:
gene data
HGNC:131 3.5
Resolve the gene column to primary HGNC IDs (a new _primary column is
added):
pysec2pri update-ids gene_ex.tsv hgnc --at gene -o gene_ex_primary.tsv
# gene data gene_primary
# HGNC:131 3.5 HGNC:145The same pattern works for labels with update-labels, and multiple columns can
be resolved by repeating --at:
pysec2pri update-ids data.tsv hgnc --at gene_id --at related_gene_idTo skip regenerating the mapping set, pass a pre-built mapping file:
pysec2pri hgnc ids # outputs hgnc_ids_{version}.sssom.tsv
pysec2pri update-ids gene_ex.tsv hgnc --at gene --mapping hgnc_ids_{version}.sssom.tsvIn Python, load_mapping reads a written SSSOM file back in, so you can
generate once and reuse the file:
from pysec2pri import load_mapping, update_ids
hgnc = load_mapping("hgnc_ids_115.sssom.tsv")
df_with_new_column = update_ids(mapping_set=hgnc, ids=df, at="gene")Use load_label_mapping for a label mapping set.
Ambiguous mappings (where a deprecated ID or label serves as a recommended for another entity) are not resolved, but flagged for users to solve them manually. If the input file has a column of known aliases or synonyms for each row, pass it as a hint to resolve ambiguous names automatically:
pysec2pri update-ids data.tsv hgnc --at gene_id --synonyms gene_aliases
# Pairs gene_aliases hints with gene_id; repeat --at X--synonyms Y for more columns.A subset with ambiguous mappings only can be generated like:
pysec2pri ambiguous hgnc-labelsEvery row is one secondary (subject) and one primary (object). Which
predicate joins them says what happened to it.
A retired ID either has a replacement or does not:
flowchart LR
S["subject_id (retired)"]
P["object_id (current)"]
N["sssom:NoTermFound"]
S -->|"IAO:0100001 (term replaced by)"| P
S -->|"oboInOwl:consider (no replacement)"| N
mapping_cardinality says how the two sides line up: 1:1, n:1 when several
retired IDs were merged into one, 1:0 for a withdrawal with no replacement.
One label mapping set holds both of a source's label changes, told apart by predicate:
flowchart LR
PREV["subject_label (previous symbol)"]
ALIAS["subject_label (alias / synonym)"]
CUR["object_label / object_id (current)"]
PREV -->|"IAO:0100001 (term replaced by)"| CUR
ALIAS -->|"oboInOwl:hasExactSynonym"| CUR
A previous symbol is one the entity used to have. An alias is another name it still goes by. Only the first is a rename; the second is what the resolver uses as evidence below.
--consolidate reads all of the source's past releases, finds mappings the
current release no longer mentions, and gives every mapping the release it first
appeared in:
pysec2pri hgnc ids --consolidate -o hgnc.sssom.tsvfrom pysec2pri import generate_ids, supports_consolidate
supports_consolidate("hgnc", "ids")
generate_ids("hgnc", consolidate=True)A value is ambiguous when it is retired in one row and current in another. This is not resolved: the row is flagged and left alone.
flowchart LR
C["C (retired)"] -->|term replaced by| A["A (retired, and current for C)"]
A -->|term replaced by| B["B (current)"]
The same holds for labels: a symbol can be a subject_label (someone's old
name) and an object_label (someone else's current name).
When a name is ambiguous, alias mappings are used as evidence. For each candidate interpretation the resolver checks whether any user-supplied hint matches a known alias of that candidate's primary entity. A hit on the secondary candidate's target confirms the name is being used as a previous name; a hit on the primary candidate's aliases confirms it is already current.
flowchart TD
Name["ambiguous name"]
Hint["Alias hint"]
Check{"Hint matches alias of…"}
SecPath["Replacement target: replace"]
PriPath["Name itself: keep"]
Blank["Neither: flag for manual review"]
Name --> Check
Hint -.-> Check
Check -->|secondary candidate| SecPath
Check -->|primary candidate| PriPath
Check -->|no match| Blank
Alias hints are one kind of context: a per-row piece of independent evidence
that helps decide which entity an ambiguous name actually means. update_ids
and update_labels support three kinds, via ContextSpec:
label-- an alias/synonym string (thesynonyms=/--synonymsshown above).id-- a related/foreign identifier string, matched the same way.xref-- a cross-reference token (e.g. an Ensembl ID) resolved through an independent crosswalk table (XrefMapping).
All three only ever touch cells already flagged ambiguous, and every attempt can be written to an auditable decision log:
from pysec2pri import generate_labels, load_xref_mapping, update_labels
label_ms = generate_labels("hgnc")
ensembl_to_hgnc = load_xref_mapping("ensembl_to_hgnc.tsv") # subject_id/object_id/object_label
resolved = update_labels(
df, label_ms, at="gene_name",
xref="ensembl", # column with each row's Ensembl ID
xref_mapping=ensembl_to_hgnc,
report_path="decisions.tsv", # stage, token, predicate_id, candidate, accepted, reason
)The same options are available on the CLI:
pysec2pri update-labels genes.tsv hgnc --at gene_name \
--xref ensembl --xref-source hgnc_custom --xref-on ensembl \
--report decisions.tsv--xref-source names a table listed in the source's config. hgnc_custom is
HGNC download, one row per gene:
| HGNC ID | Approved symbol | Status | Previous symbols | NCBI Gene ID | Ensembl ID | UniProt ID |
|---|---|---|---|---|---|---|
| HGNC:5 | A1BG | Approved | 1 | ENSG00000121410 | P04217 | |
| HGNC:37133 | A1BG-AS1 | Approved | NCRNA00181, A1BGAS, A1BG-AS | 503538 | ENSG00000268895 | |
| HGNC:6 | A1S9T | Symbol Withdrawn | ||||
| HGNC:7 | A2M | Approved | 2 | ENSG00000175899 | P01023 |
Two columns of it are already a crosswalk: pick Ensembl ID and HGNC ID and
you can map one to the other with
--xref ensembl --xref-source hgnc_custom --xref-on ensembl
Pass any table with --xref-file. It needs three columns: subject_id (what
you key on), object_id (this source's identifier), and object_label (its
label):
subject_id object_id object_label
ENSG00000121410 HGNC:5 A1BG
ENSG00000175899 HGNC:7 A2M
pysec2pri update-ids genes.tsv hgnc --at gene_id --xref ensembl \
--xref-file my_crosswalk.tsvIn Python you can point at the columns instead of renaming them, so a file like HGNC download works like:
from pysec2pri import generate_ids, load_xref_mapping, update_ids
xref = load_xref_mapping(
"hgnc_custom.tsv",
subject_col="Ensembl ID(supplied by Ensembl)",
object_col="HGNC ID",
object_label_col="Approved symbol",
)
update_ids(df, generate_ids("hgnc"), at="gene_id", xref="ensembl", xref_mapping=xref)diff compares two SSSOM files (e.g. two releases of the same mapping set) and
reports added/removed/changed rows:
pysec2pri diff old.sssom.tsv new.sssom.tsv --datasource hgnc -o diff.tsvFull documentation: https://pysec2pri.readthedocs.io/
| Datasource | license | citation |
|---|---|---|
| ChEBI | CC BY 4.0. | Hastings J, Owen G, Dekker A, et al. ChEBI in 2016: Improved services and an expanding collection of metabolites. Nucleic Acids Research. 2016 Jan;44(D1):D1214-9. DOI: 10.1093/nar/gkv1031. PMID: 26467479; PMCID: PMC4702775. |
| Ensembl | link | Martin FJ, Amode MR, Aneja A, et al. Ensembl 2023. Nucleic Acids Res. 2023 Jan 6;51(D1):D933-D941. doi: 10.1093/nar/gkac958. PMID: 36318249; PMCID: PMC9825606. |
| HMDB | CC BY 4.0 | Wishart DS, Guo A, Oler E, Wang F, Anjum A, Peters H, Dizon R, Sayeeda Z, Tian S, Lee BL, Berjanskii M, Mah R, Yamamoto M, Jovel J, Torres-Calzada C, Hiebert-Giesbrecht M, Lui VW, Varshavi D, Varshavi D, Allen D, Arndt D, Khetarpal N, Sivakumaran A, Harford K, Sanford S, Yee K, Cao X, Budinski Z, Liigand J, Zhang L, Zheng J, Mandal R, Karu N, Dambrova M, Schiöth HB, Greiner R, Gautam V. HMDB 5.0: the Human Metabolome Database for 2022. Nucleic Acids Res. 2022 Jan 7;50(D1):D622-D631. doi: 10.1093/nar/gkab1062. PMID: 34986597; PMCID: PMC8728138. |
| HGNC | link | Seal RL, Braschi B, Gray K, Jones TEM, Tweedie S, Haim-Vilmovsky L, Bruford EA. Genenames.org: the HGNC resources in 2023. Nucleic Acids Res. 2023 Jan 6;51(D1):D1003-D1009. doi: 10.1093/nar/gkac888. PMID: 36243972; PMCID: PMC9825485. |
| NCBI | link | Sayers EW, Bolton EE, Brister JR, Canese K, Chan J, Comeau DC, Connor R, Funk K, Kelly C, Kim S, Madej T, Marchler-Bauer A, Lanczycki C, Lathrop S, Lu Z, Thibaud-Nissen F, Murphy T, Phan L, Skripchenko Y, Tse T, Wang J, Williams R, Trawick BW, Pruitt KD, Sherry ST. Database resources of the national center for biotechnology information. Nucleic Acids Res. 2022 Jan 7;50(D1):D20-D26. doi: 10.1093/nar/gkab1112. PMID: 34850941; PMCID: PMC8728269. |
| UniProt | CC BY 4.0 | UniProt Consortium. UniProt: the universal protein knowledgebase in 2021. Nucleic Acids Res. 2021 Jan 8;49(D1):D480-D489. doi: 10.1093/nar/gkaa1100. PMID: 33237286; PMCID: PMC7778908. |
| VGNC | link | Tweedie S, Braschi B, Gray KA, Jones TEM, Seal RL, Yates B, Bruford EA. Genenames.org: the HGNC and VGNC resources in 2021. Nucleic Acids Res. 2021 Jan 8;49(D1):D939-D946. doi: 10.1093/nar/gkaa980. PMID: 33152070; PMCID: PMC7779007. |
| Wikidata | Vrandecic, D., Krotzsch, M. Wikidata: a free collaborative knowledgebase. Communications of the ACM. 2014. doi: 10.1145/2629489. |
MIT License. See LICENSE for details.