A reproducible pipeline that turns one AI incident's official record into a fact ledger and measures how far each fact travelled: across press circles and jurisdictions, over the phases of the incident, and in the words the press chose.
Built in three days for the Apart Research × CeSIA AI Incident Response Sprint (September 2026) and applied to the July 2026 intrusion of OpenAI evaluation agents into Hugging Face infrastructure. This repository is the pipeline; the run artefacts (fetched articles, model outputs, figures) are not included.
- Seed ledger (
watershed seed-ledger). Hand-registered official statements (inputs/statements.yaml) are segmented; a model extracts key events and key facts that cite segment numbers, and code resolves every citation to verbatim text. Three extraction runs are consolidated into a narrative chain of key events. The responsible party's statements are coded with a grounded, audited three-step reading of responsibility framing and response postures (adapted from Coombs' Situational Crisis Communication Theory), with Krippendorff's alpha across runs. - Coverage pool (
watershed coverage). Google News discovery in twelve language and country tracks; canonical-URL dedupe; fetch; per-outlet classification into one press circle and a country; per-article coding against the ledger (key events reported, key facts carried with verified excerpts, parties cited, verbatim characterisations); syndication detection by body-text similarity; aggregation of reach, lag and silence. - Discovery (
watershed discovery). Clio-style bottom-up structure: verbatim press characterisations and official stance passages are translated to English, embedded, clustered and named by a model, with no predefined codebook of frames. - Analysis (
watershed analysis figures). Static figures and tables: volume by circle and by region across the incident's phases, and where each circle's and each jurisdiction's vocabulary concentrates on the map of characterisations.
Design rules that hold throughout: models cite segment numbers and never write quotations; every human-readable label comes from one codebook that the code checks against the schemas; every stage boundary is a typed contract; every model call is traced.
uv sync
cp .env.example .env # model endpoints, SerpAPI key, Langfuse keys
uv run watershed seed-ledger run --runs 3 --scct-runs 3
uv run watershed coverage search --coverage-id <id> ; fetch ; outlets ; code ; aggregate
uv run watershed discovery phrases --coverage-id <id> --translate
uv run watershed analysis figures --run-id <run> --coverage-id <id>Specifications are in docs/ (in Chinese): seed-ledger.md, coverage.md, analysis.md.
This is a sprint prototype. It works end to end on one incident, and its reliability estimates are honest about where it is weak: model-assigned outlet circles and cluster names have not been human-reviewed, article-level framing codes did not reach acceptable inter-run agreement and are not used, and the coverage pool is bounded by search caps and fetch failures. The intent is to develop it into a rigorous, reusable instrument for incident-response measurement: human-validated outlet and cluster labels, a second model family and a human-coded validation set for every coded variable, social-media diffusion alongside press coverage, and application to further incidents so that findings can be compared rather than only described.
Sprint report: Press circles and jurisdictions across the OpenAI–Hugging Face agent intrusion (September 2026)
Applied to the OpenAI–Hugging Face intrusion: the cybersecurity press wrote a quarter of the coverage in the five days before OpenAI acknowledged that its own models were responsible, and 1% to 6% of the coverage in every later period; general news and business outlets wrote most of the coverage of the two technical incident reports and described the event as autonomous, uncontrolled AI. American and European outlets described a rogue system; mainland Chinese outlets described a race between AI companies. Outlet categories and phrase clusters are model-assigned and unreviewed, and the Chinese sample is small. The full report, Press circles and jurisdictions across the OpenAI–Hugging Face agent intrusion, was submitted to the Apart Research × CeSIA AI Incident Response Sprint.
Dexter Yao. Implementation and analysis were carried out with a coding agent under the author's direction; model calls inside the pipeline use GPT-5.6 Terra and text-embedding-3-large.