Small, reviewable functions for turning measured protein-group signal into functional profiles and comparing acquisitions with an explicit sampling design.
This is a standalone extraction and numerical review of FastaLake's experimental metaproteomics helpers. It is independent of AlphaPeptTools: use that package for its broader QC/AnnData workflow. Source attribution, the exact upstream revision, and changes are recorded in UPSTREAM.json and LICENSE.
git clone https://github.com/treitpeter/metaproteomics-tools.git
cd metaproteomics-tools
python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[test]'
python -m pytest -q
python benchmarks/run_evidence.py --output evidence/local.jsonimport pandas as pd
from metaproteomics_tools import aggregate_terms, relative_signal
signal = pd.DataFrame({'sample_1': [12., 3., 5.]}, index=['g1', 'g2', 'g3'])
annotations = {'g1': ['K1', 'K2'], 'g2': ['K1']}
profile, coverage = aggregate_terms(signal, annotations)
print(profile) # K1: 9, K2: 6, unannotated: 5
print(relative_signal(profile)) # 0.45, 0.30, 0.25The missing annotation keeps its signal in the denominator. Splitting the first group's signal between two terms conserves total signal; it does not measure how much that protein contributes to either biological function.
| Function | Contract |
|---|---|
aggregate_terms |
Split signal across distinct declared terms at one annotation level; retain unannotated signal and report coverage |
relative_signal |
Divide by observed column totals; empty acquisitions remain undefined |
diversity_summary |
Feature richness, Shannon entropy/effective diversity and Simpson concentration |
bray_curtis |
Compare composition by default, or raw observed signal explicitly; return excluded acquisitions |
clr_complete |
Natural-log CLR on features positive in every selected acquisition; return exclusions |
permanova |
One-way pseudo-F and seeded label permutations, optionally restricted within blocks |
Inputs are real numeric DataFrames with features in rows and acquisitions in columns. Identifiers must be flat, unique and nonmissing. Missing signal is NaN; negative values, infinity and total-signal overflow are rejected. PERMANOVA takes a square distance DataFrame and explicitly indexed metadata Series instead.
Install .[jit] to enable the optional Numba PERMANOVA sum. The default NumPy
implementation needs no compiler. Both use the same seeded label permutations.
- Why these choices: reasoning, alternatives and what the evidence does not establish.
- Graph fundamentals: peptide–protein evidence, ambiguity and the difference between assignment and cover.
- Empirical examples: generated measurements and their reproducible commands.
- External review brief: concrete questions for scientific and software reviewers.
The examples are synthetic and distributable. They test arithmetic, invariants, input handling and permutation mechanics. They do not establish real-cohort error control, biomarker validity, species attribution, biomass or metabolic flux. PERMANOVA here is one-way; it is not a replacement for a model with covariates, interactions, dispersion diagnostics or a study-specific permutation design.
There are no private study tables, machine paths, credentials, or institution-specific launchers in this repository. Copyright notices identify source ownership and remain as required by the source license.