Skip to content

Latest commit

 

History

History
790 lines (634 loc) · 73.8 KB

File metadata and controls

790 lines (634 loc) · 73.8 KB

Workflow Protocol

Session Start

  1. Declare session intent (one sentence: what is this session for?)
  2. Determine mode:
    • Quick: Bug fix, config change, single small component (<2 hours, no design decisions)
    • Full: New feature, design exploration, multi-component work, anything requiring decisions
  3. Escalation triggers: second feature touched, design decision emerges, >2hr estimate, contradiction detected. See .claude/PROTOCOL.md for full criteria and protocol definitions.

Review routing

Read the dated record below in full, including the backlog obligations. Explicit current founder instructions take precedence over repository copies. Historical quota notes are not evidence of current account availability, and a changed review route does not waive or close existing backlog obligations.

Dated founder routing record (preserved)

Founder override 2026-09-01:

  • Do not invoke Claude as a second opinion or reviewer for Codex-authored work. This supersedes older Strale plans, backlogs, and injected/static routing copies that require a Claude or cross-provider follow-up.
  • Use a fresh, separate Codex task for critical plan, exact-commit, and closeout verification. Pin gpt-5.6-sol at xhigh, give it a bounded read-only scope, and let the task finish after its verdict.
  • Claude Code may still be the primary environment when Petter chooses it; its critical outward transitions are reviewed by a fresh Codex task. Prior Claude verdicts remain valid historical evidence, but no new Claude-review backlog is created.

Amendment 2026-09-02 (founder, in session): Codex quota is exhausted. Until it returns, the independent review in the batch loop is performed by a fresh read-only Claude agent that did not author the batch; PRs say so, and docs/programs/cto-readiness/PROGRAM.md batch-loop step 6 carries the re-review obligation once Codex returns. The 2026-09-01 override otherwise stands.

Amendment 2026-09-03 (founder, in session) — DEC-20260903-A. Waiting for Codex costs more than proceeding. Work does not stop for the quota, and no track may be blocked on the Codex review path alone. The fresh read-only Claude agent remains the independent review; every batch that would otherwise have gone to Codex is recorded in docs/programs/codex-review-backlog.yaml. A row closes only by a Codex verdict archived under archive/ and cited by path, or by a founder waiver naming its decision. npm run codex:check enforces that against git, not prose: compared with the merge-base on main, a row may not be deleted, move backward, reopen once closed, change its commit or its recorded verdict, or have policy.review_by pushed later without a founder decision; on the file itself it refuses a commit the repository lacks, a reviewed row whose evidence file is not a verdict for it, a waiver by anyone but Petter, a decision id this repository does not record, and a row still pending past review_by. Three things it cannot see, and the pull-request review guards each: a batch never added; a drain that reaches main through review (the base moves forward on merge, as for receipts); and a decision entry added to this file in the same change that cites it — the checker proves a decision is recorded in a reviewed file, not that the founder made it. When Codex returns, drain the register before starting work that adds to it. A Codex FAIL on merged work opens a remediation batch; it does not revert anything automatically.

Founder review policy 2026-09-07, and DEC-20260910-A. On 2026-09-07 the founder removed the mandatory cross-provider review requirement: a required independent review may use the same provider in a separate review context, a different provider is optional, and work is not blocked solely because another provider is unavailable or has not reviewed it. The independent-review requirement itself, required tests, and authorization requirements for sending, publishing, deploying, spending and destructive actions all remain. This supersedes the provider-diversity parts of the 2026-09-01 override and its amendments above.

That left docs/programs/codex-review-backlog.yaml tracking an obligation that no longer existed, and on 2026-09-07 every one of its 36 rows passed policy.review_by, so codex:check failed on main and blocked every PR in the repository from 2026-09-07 until this waiver (the last merge before it was on 2026-09-06). On 2026-09-10 the founder directed, in session, that all 36 be waived (DEC-20260910-A). They are closed as waived, waived_by: petter, citing that decision. No new batches are added to the register: an independent same-provider review in a separate context satisfies the review requirement, and the PR says which review ran. The register and codex:check stay in place, so the waived history cannot be edited, deleted or reopened.

Repo-native migration continuation — pre-cutover

When a session is asked to continue the repo-native operating-model migration, start in an isolated worktree from current origin/main and read the Current continuation checkpoint in docs/strategy/2026-08-31-repo-native-operating-model-migration.md before choosing work. Follow the exact handoff and next bounded task named there. Do not work from the dirty shared checkout or infer the next step from an older dated handoff.

This is a navigation pointer, not M4 cutover. Candidate project documents remain inactive and Notion-backed workflows remain authoritative until the explicit atomic cutover.

Program register — where multi-batch work resumes

Long-running work is tracked in docs/programs/ (index: docs/programs/README.md). Each program has a PROGRAM.md with a Resume here section and a machine-checked tracks.yaml (npm run programs:check). A session continuing a program starts with those two files and follows their pointers: the active track's resume_file names anything else that batch needs. The migration checkpoint above is the M2-through-M7 detail behind the cto-readiness program. Programs are execution records, not project truth.

Research and ideas — where each one lives

Research goes to docs/research/ in the dated, front-matter template (docs/research/research.schema.json), checked by npm run research:check: one current file per topic, reciprocal and acyclic supersession, resolvable links and sources, and no active decision record citing research that is not current. Research is evidence, never authority — a finding that changes direction produces a decision record, not a rewrite of the research file. Ideas go to docs/company/IDEAS.md, nowhere else — not a Journal entry, not a prose aside in a plan or handoff.

Design tokens — where design values live

Design values are data, in design/tokens/, not prose or hardcoded literals in code. design/tokens/active.json is what production runs, per surface, with provenance and the decision that adopted it. A direction under consideration is a candidate — design/tokens/candidates/*.json — and carries its own status (exploring → proposed → adopted | rejected). Promotion is a decision record plus a file swap, never an edit to active.json values in place. If a value the tokens don't have is needed anywhere a surface's design is consumed, add the token first — never reach for a literal. npm run design:check refuses off-token colours, fonts, and off-scale spacing/radii in the surfaces it covers. See design/README.md and design/PROVENANCE.md.

Cheap extras — env manifest, model registry, claims register (T14)

Three more small "values are data" registers, same shape as tokens/research/ programs. config/env-manifest.yaml (schema config/env-manifest.schema.json) is one row per distinct process.env.NAME read under apps/api/src, apps/api/scripts, packages, and scripts — purpose, provider, holder, cost class, where it's required and where it's actually set. npm run env:check fails on an undocumented read or a dead row; npm run env:example regenerates .env.example and apps/api/.env.example from it. apps/api/src/lib/models.ts is the only place a Claude/Voyage/GPT model id may live — every capability imports a role (MODELS.capability_default.id, etc.), never a literal. npm run models:check fails on a model-id literal anywhere else, or a registry entry missing pinned_at/decision. docs/company/claims.yaml (schema docs/company/claims.schema.json, writing rules in docs/company/VOICE.md) rules every public claim allowed | needs_evidence | forbidden | retired. npm run claims:check scans README.md, package READMEs, manifest descriptions, and platform-facts.ts (plus, read-only, the sibling frontend's llms.txt when present) and fails on a forbidden claim. All three are wired into CI after design:check / design:test.

Evidence receipts and the migration ledger (T15)

Evidence is a receipt file under archive/receipts/, cited by path — never a bare test count in prose. A receipt (archive/receipts/receipt.schema.json; naming rule YYYY-MM-DD-<kind>-<topic>.json) is written once, by the tool that produced it, and never edited afterward; npm run receipts:check enforces this as a git fact (a tracked receipt's blob at HEAD must match the blob at the commit that first added it), validates the schema, checks that every evidence: / production_evidence: path cited from a decision record, a program track, or a remediation package resolves, and warns on a post-2026-09-02 handoff stating a test count with no receipt link. Write one with npm run receipt -- --kind <kind> --topic <topic> --from <file|->. Migration blocks in apps/api/src/lib/startup-migrations.ts are append-only and ledgered: apps/api/src/lib/startup-migrations.ledger.json carries a content hash and columns_written per block, and npm run migrations:check fails on an edited block (the fix is a new block, never an in-place edit), an unledgered block, or two blocks writing the same column unless it's allowlisted in known_overlaps — the 2026-08-21 incident class, where two blocks derived one column and fought every boot. Both wired into CI after docs:test / archive:index:test.

Session contract — both tools, every session

  1. Orient first. Read docs/programs/README.md, then the active track's next_action and resume_file in docs/programs/<program>/tracks.yaml. Claude Code prints this at SessionStart; Codex runs npm run handoff:orient. Work happens in one batch worktree (strale-wt-<track>) on a feature branch cut from origin/main; the trunk stays on main and clean (WORKTREES.md).
  2. Gate before stopping. npm run handoff:check must pass before a session ends: no uncommitted paths, the branch pushed to its upstream, at most one batch worktree, no merged branch left locally or on the remote, and every code change accompanied by an updated next_action in the program register or a new handoff/_general/from-code/ file. The check prints one fix per finding; apply them and rerun. Claude Code's Stop hook (.claude/settings.json) requests continuation on failure, with a six-block escape for repeated finding codes. That escape records failure; it does not mean the gate passed. Codex's notify wrapper (scripts/handoff/codex-notify.mjs, chained in ~/.codex/config.toml by its trunk path, which must exist on main) records the result in .claude/state/handoff/last-codex.json, which the next session's orientation shows, and .codex/hooks.json blocks the stop the same way when the hook is enabled, trusted, discovered, and successfully invoked. Hook commands resolve their script from the current Git worktree root, including when the session starts in a subdirectory. Notify records a result after the turn; it does not block stopping. A Codex session runs the check itself before its final turn and fixes what it lists.
  3. Git hooks come with npm ci (prepare runs the same installer as npm run hooks:install, setting core.hooksPath=.githooks for every worktree of the clone): pre-commit refuses a commit on main and an inventory-target edit without npm run context:generate; pre-push refuses a direct push to main. main changes only through reviewed PRs merged on GitHub; pushing the working branch is routine backup and needs no approval. Worktrees and branches recorded in scripts/handoff/baseline.json wait for a founder decision and are never deleted by a session.

Notion Access (REQUIRED)

Notion Workspace Structure (8 sections under Project Home)

  1. 🏠 Start Here — overview + navigation
  2. 🎯 Strategy — what Strale is, the problem, opportunity, competitive landscape, business model
  3. 🛠️ Products — SQS, Audit Trail, Discovery, Capabilities & solutions, Feature Registry DB
  4. ✅ To-do & Build Plan — THE ONLY task list (To-do DB + Deferred DB)
  5. 📣 Go-to-market — distribution surfaces, activation funnel, brand & voice, social media, Social Media Posts DB
  6. 🔧 Internals — testing system, testing rules, onboarding pipeline, bug fix framework, tech stack
  7. 📓 Journal — session logs, brainstorms, analyses (Journal DB)
  8. ⚙️ How we work — working rules, governance, Decisions DB, Glossary DB

Notion Governance Rules (enforced)

  • Check before creating — look at the page directory before creating any new page
  • ONE page per topic — never create v2, update existing subpage
  • Brainstorms go to Journal DB, not as standalone pages
  • To-do DB is THE ONLY task list — action items never live in prose
  • Superseded pages archived same session (prefix + move to archive)
  • Search existing pages before creating new ones

GitHub Access (REQUIRED)

  • Repo: strale (local)
  • Main branch: main
  • Feature branch pattern: type/kebab-description

Project Spec

The original MVP spec files have been removed from this repo (archived to Notion). For current build plan, priorities, and architecture, see Notion Project Home: https://www.notion.so/31167c87-082c-81fb-96da-d3188d34aa72

Tech Stack

  • Runtime: Node.js + TypeScript
  • Framework: Hono
  • Database: PostgreSQL
  • ORM: Drizzle
  • Payments: Stripe Checkout (wallet top-ups only, no Connect)
  • Hosting: Railway (US East / Virginia, project: desirable-serenity)
  • Headless browser: Browserless.io (managed, NOT self-hosted Puppeteer)
  • SDKs: TypeScript first, then Python

Project Structure

strale/
├── apps/
│   └── api/                    # Hono API server
│       ├── src/
│       │   ├── routes/         # API route handlers
│       │   ├── capabilities/   # Capability executor functions
│       │   ├── db/             # Drizzle schema + queries + seed
│       │   ├── lib/            # Stripe, matching, auth, quality helpers
│       │   └── index.ts        # Entry point
│       ├── drizzle/            # Migration files (0001–0006+)
│       └── package.json
├── packages/
│   ├── mcp-server/             # strale-mcp (npm published)
│   ├── sdk-typescript/         # @strale/sdk (npm published)
│   ├── sdk-python/             # straleio (PyPI published)
│   ├── semantic-kernel-strale/ # strale-semantic-kernel (npm)
│   ├── langchain-strale/       # langchain-strale (PyPI published)
│   └── crewai-strale/          # crewai-strale (PyPI published)
├── package.json                # Monorepo root
└── CLAUDE.md

Active Decisions

MVP Decisions (Feb 2026)

  • DEC-1: Scope reduced to 4-week MVP proving developers will let agents buy capabilities
  • DEC-2: Prepaid wallet via Stripe Checkout — internal ledger for micropayments, zero per-transaction cost
  • DEC-3: No bidding/auction — fixed pricing, instant routing, keyword matching for 5 capabilities
  • DEC-4: Founder is the only provider for first 3 months
  • DEC-5: TypeScript backend (Hono + Drizzle + PostgreSQL)
  • DEC-6: EU/Nordic data wedge — 5 seed capabilities
  • DEC-7: Use Browserless.io instead of self-hosted Puppeteer (unanimous reviewer feedback)
  • DEC-8: SELECT FOR UPDATE row-level locking on wallet debits (unanimous)
  • DEC-9: Idempotency-Key header on POST /v1/do (unanimous)
  • DEC-10: €2.00 trial credits on signup, no card required (unanimous)
  • DEC-11: Rating endpoint removed from MVP (unanimous)
  • DEC-12: screenshot-url and eu-address-validate dropped; replaced by vat-validate and annual-report-extract
  • DEC-13: Invoice extraction price raised to €0.50
  • DEC-14: Don't charge before execution succeeds — lock → execute → deduct on success
  • DEC-15: Add capability_slug override to POST /v1/do
  • DEC-16: Add dry_run mode to POST /v1/do
  • DEC-17: Return wallet_balance_cents in /v1/do response
  • DEC-18: Dashboard scope reduced to: register, API key, balance, top-up, transaction list
  • DEC-19: Structured error responses with stable error_code enum
  • DEC-20: Hash API keys in DB, store key_prefix for lookup
  • DEC-21: Rate limiting: 10 req/sec per key + €100/hour spend cap
  • DEC-22: Hybrid sync/async execution — sync for <5s, async+poll for longer capabilities
  • DEC-23: TypeScript SDK ships before Python SDK
  • DEC-20260225-P-c5d6: 6th table — failed_requests (id, user_id, task, category, max_price_cents, created_at) logs every no_matching_capability response
  • DEC-20260225-P-m5n6: swedish-company-data accepts fuzzy natural-language input; cheap LLM call resolves to org number before registry lookup

Current Decisions (March 2026)

  • DEC-20260302-A: Capability Pricing Framework (€0.02–€1.00 per call)
  • DEC-20260302-B: Capability QA Framework (tiered scheduling: smoke/daily/weekly)
  • DEC-20260302-C: Historical homepage prescription; superseded for the apps/web redesign by DEC-20260905-A. Its outcome-before-plumbing rationale is preserved.
  • DEC-20260303-D: Search input uses query completions, not result dropdown
  • DEC-20260303-E: POST /v1/suggest uses Voyage AI embeddings + Claude Haiku re-ranking
  • DEC-20260303-G: Historical eleven-section homepage order; superseded for the apps/web redesign by DEC-20260905-A. Evidence still belongs near the claim it supports.
  • DEC-20260305-A through G: Trust display centralization, test infrastructure, security hardening
  • DEC-20260306-A through F: Test run audit log, metric consistency, capability detail audit
  • DEC-20260307: SQS Constitution adopted as authoritative scoring spec; Notion Governance Protocol established

Current Decisions (April 2026)

  • DEC-20260428-A (global, active): Third-party scraping doctrine — three-tier framework. Tier 1: Strale itself never operates scrapers (absolute). Tier 2: may consume vendor-scraped data when underlying data is public records by statute, vendor has documented redistribution rights + indemnification, vendor provides primary-source provenance per fact, and Strale discloses sourcing via provenance.upstream_vendor / acquisition_method / primary_source_reference. Tier 3: prefer licensed-bulk over scraping-derived when both are available at compatible economics. Anchored on Meta v. Bright Data (NDCal Jan 2024) and hiQ v. LinkedIn (settled Dec 2022, $500k judgment). Supersedes the implicit absolute no-scraping rule. Full doctrine: Notion Decisions DB (page id 35067c87-082c-810d-b6a4-edf9f14b4446).
  • DEC-20260428-B (global, active): Engineering bar for Strale-built data services (sanctions/PEP, UBO, adverse media, future registry self-builds). Codifies regulatory-grade requirements: versioned dataset with stale-data circuit breaker, source-list manifest per response, Merkle-rooted ingest, match explainability, confidence buckets, dispute endpoint with disposition tracking, replay capability, golden test suite, canary deploys, per-list kill switches, GDPR Art. 22 compliance, threat-model document and public methodology page mandatory before production. AI synthesis steps (e.g. risk-narrative-generate) must require per-flag source citation, "screening checks found" framing, and never assert facts not present in input. Pairs with DEC-20260428-A.

Current Decisions (September 2026)

  • DEC-20260910-A (global, active): The Codex review backlog is waived under the founder's 2026-09-07 review policy. Directed by Petter in session on 2026-09-10 ("go with option 1, waive all 36"). The 2026-09-07 policy made cross-provider review optional — an independent review may be same-provider in a separate context — which removed the obligation docs/programs/codex-review-backlog.yaml existed to track. All 36 pending rows (CX-1 to CX-36; 23 high, 13 medium) passed policy.review_by of 2026-09-07 and made codex:check fail on main, blocking every PR from 2026-09-07 to 2026-09-10. They are closed waived, waived_by: petter, citing this decision. No new batches are added to the register; codex:check remains so the waived history stays immutable. Amends DEC-20260903-A (does not delete its register). Full text: the Review routing section above.
  • DEC-20260905-A (global, active): Benefit-first positioning for the redesign. Founder approved the reviewed positioning brief on 5 September: tools and data for AI agents; useful recurring agent work first, shared access/integration benefit next, customer-visible execution evidence with route-specific limits. Marketing uses tools; data services remains explanatory and technical identifiers remain in API contexts. Quiet Material is the control for refinement, not newly adopted production tokens or final artwork. Broad-library strategy, x402 priority and claim/publication gates remain. Supersedes DEC-20260302-C and DEC-20260303-G homepage composition prescriptions for the redesign while preserving their outcome-first and evidence-near-claim rationale. Adoption and narrow VOICE.md reconciliation: docs/strategy/2026-09-05-brand-direction-adoption.md; execution resumes at docs/programs/brand-website/PROGRAM.md.
  • DEC-20260904-C (global, active): Capabilities labelled Unverified are listed on the website with the label, not hidden. Directed by Petter 2026-09-04 on an M2 batch-9 finding whose premise turned out stale: the website's isSQSUnqualified filter would have hidden every capability labelled Unverified, but it has had no callers since the 2026-08 audit follow-up, and strale.dev already lists such capabilities dimmed with an "Awaiting traffic" badge. Affirms DEC-20260313-C and the current behaviour; pending and Building-track-record states keep their behaviour. The dead filter was aligned with the decision in strale-frontend PR #24 (kept in maintenance under DEC-20260902-A) so a revived caller cannot reintroduce hiding. Lesson: a comment is not evidence of behaviour; check the callers. Notion Decisions DB entry filed 2026-09-04; repo-native record follows through the M2 closure path.
  • DEC-20260904-B (operational, active): Cross-surface identity mechanism for the M2 closure register (git-qualified record keys). The record-key grammar gains a second source qualifier, symmetric to --notion-<32 hex page id>: --git-<7 to 40 lowercase hex>, naming the commit that introduced the claim directly in Git (^DEC-[A-Za-z0-9]+(?:-[A-Za-z0-9]+)*(?:--notion-[0-9a-f]{32}|--git-[0-9a-f]{7,40})?$). A git-qualified record must have id equal to the key with the qualifier removed, source_kind: git-native/source_rows: [], and git_provenance equal to its own first evidence entry, a full-sha https://github.com/strale-io/strale/commit/<sha> URL whose prefix matches and which is an ancestor of HEAD (findings RECORD_GIT_KEY_ID_MISMATCH/_SOURCE_KIND/_PROVENANCE_MISMATCH/_NOT_ANCESTOR; COMMIT_UNVERIFIABLE when git is unreachable). A bare collided id — now including a cross-surface collision id — is never a record key (RECORD_KEY_BARE_CROSS_SURFACE_ID). A cross-surface row may resolve to resolved_collision/documented_only only when a git-qualified record exists for the collision id AND a gap report cited in the row's own evidence names its page id; row_disposition: formal_record stays unsupported on a cross-surface row this stage (CROSS_SURFACE_FORMAL_RECORD_UNSUPPORTED); any other combination is DECISION_ROW_CROSS_SURFACE_STATE_INVALID. Stage 1 only (lands and verifies the mechanism): creates no DEC-20260422-A.md record in either meaning and does not change that row's disposition. Full text in the formal candidate record for this id (M2 candidate path, inactive); mechanism gap: archive/sessions/2026-09-01-m2-enforcement-protocol-source-gaps.md.
  • DEC-20260904-A (operational, active): Pre-readiness feature-scoped M2 decision rows are evidence-only. A preserved M2 closure-register Decision row is classified intentionally_historical (evidence-only) rather than left pending migration when it is active, historical_scope: feature, decided before 2026-08-12 (the DEC-20260812-A readiness-program adoption date), still not_yet_reconciled, and not a collision-registry row, a Git-native protocol label, or an existing formal-record id. 76 of 216 private rows matched at archive commit 995cece3; the remaining 129 not_yet_reconciled rows (128 global, 1 temporary) are unaffected and stay G1 batch work. Full predicate, exclusions, and rationale in the formal candidate record for this id (M2 candidate path, inactive); row list: archive/sessions/2026-09-04-m2-g1-pre-readiness-feature-rows-gaps.md.
  • DEC-20260903-A (global, active): Work does not stop for the Codex quota; the review debt is a checked register. Directed by Petter 2026-09-03. Amends the 2026-09-01 review-routing override and its 2026-09-02 amendment: the fresh read-only Claude agent remains the independent review, no track may be blocked on the Codex review path alone, and every batch that would otherwise have gone to Codex is recorded in docs/programs/codex-review-backlog.yaml until a Codex verdict archived under archive/ closes it or Petter waives it naming a decision. Enforced against git by npm run codex:check. Notion Decisions DB entry filed 2026-09-03; full text in the Review routing section above.
  • DEC-20260902-A (global, active): The website redesign is built inside this repository as apps/web (monorepo). Directed by Petter 2026-09-02. Preserve first, then build: strale-frontend was swept and its design material preserved (release preserve-2026-09-02, archive tags, tracked candidates) and is kept, not extended, until the apps/web site serves production. Design-token work lands in apps/web. Reversal: a new record if Cloudflare Pages cannot build from a monorepo subdirectory or the site source must stay private. Notion Decisions DB entry filed 2026-09-02; repo-native record follows through the M2 closure path.

Current Decisions (August 2026)

  • DEC-20260813-A (global, active): DEC-20260518-F affirmed as the operative interpretation of DEC-20260428-A. Tier 1 ("Strale never operates scrapers") targets bulk collection and scraping infrastructure, NOT targeted per-call parsing. Per-call HTML/PDF parsing of statutorily-public registry pages is permitted when ALL four constraints hold: (a) statutorily public, (b) registry ToS permits per-call automated access — verified and recorded in the capability's manifest before launch, (c) per-entity/per-customer-request, never bulk, (d) attribution + provenance preserved. Still absolute: bulk crawling, ToS-prohibited targets (DEC-20260420-H social platforms, DEC-20260427-H-4 Google), robots.txt evasion, CAPTCHA solving, proxy rotation, login-wall circumvention. Preference order: official API > licensed bulk > Tier-2 vendor > per-call parsing — F is the floor, not the default. Opens the Greece per-call path; supersedes the absolutist reading in the 2026-05-18 MT/HU partials. Full text: Notion Decisions DB 3bb67c87-082c-8101-a08a-ddaf92ffb5df.
  • DEC-20260815-A (global, active): Operating charter — Claude runs day-to-day operations. Full text docs/company/CHARTER.md. Amends (does not merely extend) DEC-20260812-A's escalation contract. Principle: the tier of risk stays the same, the width expands. No technical question goes to Petter — architecture, implementation, what to measure, what to build and in what order, testing, tooling, vendor-API choice are all Claude's; asking him to arbitrate a technical choice is a failure of the role. Claude also decides-then-tells on: turning services on/off, pricing inside the existing €0.02–€1.00 band, quality gates, quarantine/promote, refunds, retries, delisting, merging its own work once repo gates pass, dispatching agents, scheduling sessions, spend inside €50/week. Petter alone decides: spend beyond the envelope, anything legally binding Moonlighter AB (accounts, terms, vendor contact), one-way public acts, pricing outside the band, and regulator-facing claims. Shipping is never Petter's decision — the session that opens a PR merges it and reports afterwards in plain English; the morning check-in sweeps every open PR and dirty branch daily; a merge is not "shipped" until the served artifact is verified. Customer-data boundary: no outreach derived from transaction evidence; telemetry yields anonymous product insight only; named prospects come from public research or from customers who registered; a payment is not a relationship; content is redacted at 90 days on every capability; widening any of this is Petter's explicit call. Companion docs: GOALS.md (revenue ladder, EUR-denominated), DECISION-QUEUE.md, BUDGET.md, WORKFORCE.md, MEASUREMENT.md, DESIGN-SYSTEM.md. Notion Decisions DB entry: 3be67c87-082c-8143-b70d-c6503893ba73 (filed 2026-08-15).
  • DEC-20260822-A (global, active): Daily-run reform — two artifacts, wider autonomy, systematic self-improvement. Directed by Petter 2026-08-22 on reviewing a week of daily runs. Amends DEC-20260815-A (does not supersede it). Three parts. (1) Two artifacts per daily run. An internal operating record (handoff/_general/from-code/) carrying the full technical evidence, and a CEO morning brief (docs/company/briefs/YYYY-MM-DD.md) written only after the operating work is complete: ~300–600 words of non-technical English (a target; the gate warns above it and fails only past a 900-word ceiling), five fixed sections (business performance · what materially changed · fixed automatically · working on now · needs your decision), no filenames/commit ids/queries/branches/test counts/jargon. The brief is a synthesis, never a work log. (2) Wider autonomy, unchanged risk ceiling. Claude now acts without asking on: obvious reversible evidence-backed code errors; false monitoring/instrumentation signals; demonstrably inaccurate public copy (narrowing only — a stronger replacement claim stays founder-gated); routine internal-account/data cleanup, quarantine/promotion, refunds, retries, delisting where policy already determines the answer; and investigating factual/technical uncertainty rather than escalating it. Every escalation must first fail the test "could further code inspection, production measurement, experimentation, or an existing decision resolve this?", and must carry five fields: the choice, what is established, the options, my recommendation, the concrete consequence of each. No unresolved technical question ever reaches Petter. (3) Failure families and the three-strike rule. docs/company/LESSONS.md tracks recurring mistakes by family (F1 false quality attribution, F2 wrong denominator, F3 billing/economic judgement, F4 misleading metric, F5 hollow test, F6 stale public claim, F7 state drift, F8 duplicated authority, F9 incorrect escalation, F10 approval-boundary breach). Counts live in LESSONS.md and are not restated here — three documents carried three different figures for F1 before this rule. A third materially similar incident in one family automatically becomes a root-cause investigation — identify the shared authority, measure the full affected population, falsify the hypothesis, repair the mechanism, add a discriminating guard, replay history, verify in production. F1 (quality/failure attribution) and F5 (hollow tests) are both past threshold and their investigations are open now; F10 was opened at two incidents because one breach of an approval gate costs the gate its meaning. Authority files: docs/company/DAILY-RUN.md (the run itself; the strale-checkin-morning scheduled task is a pointer to it, never a copy), docs/company/LESSONS.md, amended CHARTER.md / MEASUREMENT.md / WORKFORCE.md. (4) Authorization is bound to code, not prose. Autonomy is limited by hard authorization boundaries: being right about an action is never authority to take it, an approval-gated item leaves Petter's queue only when he moves it, and a permission not held is a stop rather than an obstacle. Daily-run items carry one of three statuses — SYSTEM_ACTING (decided and done, inside delegated authority), FOUNDER_DECISION (judgement is his), AUTHORIZATION_UNAVAILABLE (decision settled, execution permission absent — never authority to act, and never used for something already done). These are names for shapes produced by apps/api/src/lib/production-authority.ts (DEC-20260822-B / PR #361), which the charter binds to by symbol; charter-authorization-binding.test.ts fails if the charter names a symbol that module does not export. Machinery: apps/api/src/lib/metrics/commercial.ts (+ tests), scripts/commercial-brief.ts, src/lib/ceo-brief-lint.ts + scripts/check-ceo-brief.ts wired into CI.
  • DEC-20260812-A (global, active): Readiness program adopted. The 2026-08-05 Direction Plan Part One (library-as-product, x402 primary rail) plus the Platform Readiness & Self-Operation Program (docs/strategy/2026-08-12-platform-readiness-program.md) are the operating strategy. Supersedes DEC-20260502-A (Counterparty Assurance rename/ICP) and DEC-20260503-A (dual-domain architecture); the Counterparty Assurance framing is retired as primary product — compliance is a separate track gated on customer discovery. Confirmed defaults: €25 external-cost cap per full-catalog prod sweep (denylist honored); escalation contract (platform acts alone on quarantine/promote, fixture refresh, retries, delisting, refunds, draft PRs — humans decide spend above cap, vendor/license, pricing, deactivating revenue earners, DEC-20260428-B-grade builds, new external claims); quality floor quarantine <70% / deactivate <30% on ≥10 real calls/30d, auto-promote on recovery; factory may dark-launch zero-maintenance-class capabilities (invisible + non-x402 until first green week). Full text: Notion Decisions DB 3ba67c87-082c-8129-86c6-c35d82bc986f.

Capabilities & Quality

290+ capabilities across 7 verticals (company-data, compliance, developer-tools, finance, data-processing, web-scraping, monitoring) plus 100+ bundled solutions across 6 categories. Full catalog: GET /v1/capabilities. Solutions: GET /v1/solutions. Counts grow frequently — check manifests/*.yaml and recent git log for exact current numbers.

x402 Payment Gateway (March 2026): All capabilities and solutions available via x402 pay-per-use USDC payments on Base mainnet. No signup or API key needed — payment IS the auth. DB-driven: adding capabilities to x402 requires only UPDATE capabilities SET x402_enabled = true. Catalog: GET /x402/catalog. Discovery: GET /.well-known/x402.json. Wildcard handler: GET/POST /x402/:slug.

New capabilities (March 2026):

  • pep-check — Dilisense consolidated PEP database (230+ territories, EU C/2023/724-aligned, RCAs included). Category: compliance. Price: €0.05. Transparency: algorithmic. Uses DILISENSE_API_KEY. (OpenSanctions previously primary with Dilisense fallback; OS dropped 2026-04-27 commit 16ca790 — single-vendor on Dilisense per DEC-20260429-A.)
  • adverse-media-check — Dilisense Adverse Media (235k+ news sources, FATF-categorized) primary; Serper.dev (Google) fallback with deterministic keyword classification. No LLM. Category: compliance. Price: €0.20. Transparency: algorithmic. Uses DILISENSE_API_KEY (primary) + SERPER_API_KEY (fallback). Risk-level rule documented in output via risk_level_thresholds.
  • risk-narrative-generate — AI synthesis of structured check results into plain-language risk narrative. Category: agent-tooling. Price: €0.05. Transparency: ai_generated. Uses ANTHROPIC_API_KEY.
  • au-company-data — Australian Business Register (ABR) lookup by ABN. Category: company-data. Price: €0.05. Transparency: algorithmic. Uses ABN_LOOKUP_GUID.

New solutions (March 2026):

  • KYB Essentials (×20 countries) — Quick company verification. 3-4 checks, €1.50. Slug: kyb-essentials-{cc}
  • KYB Complete (×20 countries) — Full compliance check with risk narrative. 11-14 checks, €2.50. Slug: kyb-complete-{cc}
  • Invoice Verify (×20 countries) — Invoice fraud detection with risk narrative. 12-14 checks, €2.50. Slug: invoice-verify-{cc}
  • Countries: SE, NO, DK, FI, UK, DE, FR, NL, BE, AT, IE, ES, IT, CH, PL, PT, US, CA, AU, SG
  • Predecessors that overlap the KYB families: kyc-sweden, kyc-norway, kyc-denmark, kyc-finland, verify-us-company. This file previously recorded them as deprecated with isActive: false; production contradicted that on 2026-08-14. Status is deliberately not restated here — read is_active / x402_enabled from the solutions table (per the drift-prevention rule below) before acting on them, and do not deactivate any of them on the strength of a line in this file.

SQS scoring engine deleted per DEC-20260503-B (PR1 shipped 2026-05-05). The dual-profile model (QP + RP + 5×5 matrix), the min_sqs request parameter on POST /v1/do, the platform floor SQS gate, the floor-aware solution SQS rule, the public /v1/quality/:slug endpoint, and the automatic lifecycle transitions (probation→active, active→degraded, degraded→active, degraded→suspended) are all gone. PR2 will drop the residual schema columns (qp_score, rp_score, matrix_sqs, matrix_sqs_raw, trend, guidance_*) and the sqs_daily_snapshot table, and rename capability_health → source_health. Test scheduling now filters on test_suites.scheduled_testing_eligible = TRUE (DEC-20260503-B), with external_cost_cents = 0 as the underlying source of truth. A startup migration rewrites the flag on every boot: SET scheduled_testing_eligible = (external_cost_cents = 0). Consequence: hand-editing scheduled_testing_eligible is silently reverted at the next deploy; external_cost_cents is the only durable knob even though it appears billing-only. Paid capabilities are not proactively tested; quality signals come from production observability, piggyback test suites, and any zero-cost auth-less probes the vendor permits. Circuit-breaker logic on capability_health survives. Fixture and canary test modes survive.

Free-tier: 11 capabilities as of 2026-08 (email-validate, dns-lookup, json-repair, url-to-markdown, iban-validate, plus 6 crypto address validators: bitcoin/eth/solana/tron/dogecoin/xrp-address-validate) require no auth/signup. Canonical list is is_free_tier = true in the capabilities table, surfaced via GET /v1/platform/facts (free_tier_slugs) — check there before quoting a count. IP-based daily rate limit (10/day, enforced via DB counter in do.ts using rateLimitByIp). Authenticated users calling free-tier capabilities get normal rate limits and no wallet debit.

Testing: test_suites table has test_mode column: live (calls real API), fixture (uses saved data, €0 external cost), canary (periodic live check at reduced frequency). external_cost_cents tracks estimated external API cost per test execution.

Stripe is LIVE in production (sk_live_ key on Railway). Local .env uses sk_test_ for development.

Adding New Capabilities (MANDATORY PIPELINE)

All new capabilities MUST go through the manifest-driven pipeline. (The historical seed.ts file was deleted in PR #79; it duplicated manifest content and only generated 2 of 5 required test types via its onboarding hook. The canonical pipeline is apps/api/scripts/onboard.ts — it generates all 5 test types and is the only sanctioned path for capability creation.)

Recommended workflow (--discover):

  1. Write the executor at apps/api/src/capabilities/{slug}.ts

    • Register via registerCapability(slug, handler)
    • Handler returns { output, provenance: { source, fetched_at } }
    • All external calls must have AbortSignal.timeout()
    • Errors must be structured, never raw HTML or stack traces
  2. Auto-registered — executors are auto-imported at startup by src/capabilities/auto-register.ts. No manual import in app.ts needed.

  3. Create minimal manifest at manifests/{slug}.yaml Required fields:

    • slug, name, description, category, price_cents
    • data_source, data_source_type, transparency_tag, freshness_category
    • test_fixtures.health_check_input — simple input that always works
    • limitations — at least 1 (every capability has limitations) No need to write expected_fields or output_field_reliability — the pipeline generates them.
  4. Run the pipeline with --discover: cd apps/api && npx tsx scripts/onboard.ts --discover --manifest ../../manifests/{slug}.yaml The pipeline:

    • Executes the capability with health_check_input
    • Auto-generates expected_fields from the actual output
    • Auto-generates output_field_reliability (all fields marked guaranteed initially)
    • Writes the updated manifest back to disk
    • Generates all 5 test types (known_answer, schema_check, negative, edge_case, dependency_health)
    • Verifies the known_answer test passes against live output
  5. Review: Check the auto-generated expected_fields in the manifest. Adjust reliability levels (guaranteed/common/rare) as needed.

  6. Verify: npx tsx scripts/smoke-test.ts --slug {slug}

Pipeline flags:

--manifest <path>    Path to YAML manifest (required)
--dry-run            Preview without inserting to DB
--backfill           Update existing capability (add missing tests, update fixtures)
--discover           Auto-generate expected_fields from live execution output
--fix                Auto-correct high-confidence fixture mismatches (field name typos, case, type coercion)
--strict             Abort if execute-and-verify fails

Combine flags for existing capabilities: --backfill --discover --fix

For backfilling existing capabilities:

cd apps/api && npx tsx scripts/onboard.ts --manifest ../../manifests/{slug}.yaml --backfill Skips capability creation, adds only missing test types, updates field reliability + limitations.

Use --backfill --discover --fix to auto-correct fixture mismatches on existing capabilities.

Field reliability rules:

  • guaranteed — always present in successful responses. Safe to assert on.
  • common — usually present, may be absent for some inputs. Type-checked only.
  • rare — only present for specific inputs. Never asserted on.

Only guaranteed fields are used in known_answer test assertions. This prevents the "expected non-null on optional field" problem that broke 8 EU registries.

What the pipeline does NOT do (human must provide):

  • The known_answer test input (a real entity you've verified works)
  • Field reliability annotations (which fields are truly guaranteed)
  • Limitations (honest assessment of coverage gaps)
  • The executor code itself

Everything else is auto-generated. This is how the platform scales to third-party providers.

Quick reference — manifest template:

slug: "example-capability"
name: "Example Capability"
description: "What it does (50-160 chars for SEO)"
category: "validation"
price_cents: 5
data_source: "Example API"
data_source_type: "api"  # api | scrape | computed | reference
transparency_tag: "algorithmic"  # algorithmic | ai_generated | mixed
freshness_category: "live-fetch"  # live-fetch | reference-data | computed

# Bucket C — GDPR Art. 22 classification (optional; default 'data_lookup').
# Set to 'screening_signal' for capabilities that produce matches/findings
# the customer uses to decide (e.g. sanctions-check). Set to
# 'risk_synthesis' for AI synthesis producing a recommendation
# (e.g. risk-narrative-generate). Surfaced in the audit body's gdpr block
# along with the dispute_endpoint URL. Validated at authoring time by
# validateCapabilityStructure (gate 15) against the canonical enum.
gdpr_art_22_classification: "data_lookup"  # data_lookup | screening_signal | risk_synthesis

test_fixtures:
  known_answer:
    input:
      field_name: "real_value"
    expected_fields:
      - { field: "output_field", operator: "not_null", reliability: "guaranteed" }
  health_check_input:
    field_name: "real_value"

output_field_reliability:
  output_field: "guaranteed"
  optional_field: "common"
  rare_field: "rare"

limitations:
  - title: "Coverage limitation"
    text: "Description of the limitation"
    category: "coverage"
    severity: "info"

Scoring Integrity (retired with the SQS engine — DEC-20260503-B)

The Scoring Integrity Protocol has been retired. The SQS engine, sqs.ts, EXTERNAL_SERVICE_PATTERNS, isExternalServiceFailure, and computeFromRows no longer exist (PR1 deletion 2026-05-05). When a capability scores poorly under a future routing-engine signal, the same root-cause discipline still applies: diagnose the underlying issue (missing credential, bad fixture, real bug) and fix it; never bend the substrate to mask a specific capability's behaviour.

See also: Capability Onboarding Protocol (DEC-20260320-B).

Test Infrastructure Cost Principles (always enforce)

Principle A — Zero-cost health probes: Health probes in dependency-manifest.ts must never consume billable API calls. Use skipAuth: true on the health probe for paid APIs so the probe sends no auth header — a 401 proves connectivity without consuming quota. Probes run ~4×/day per provider; authenticated probes waste 120+ API calls/month.

Principle B — Input validation before paid APIs: Every capability that calls a paid external API must validate input and throw an error for empty, null, or sub-2-character input BEFORE making the API call. This protects both test budget and customer-traffic budget.

Principle C — Piggyback suites never scheduled: Piggyback test suites (test_type = 'piggyback') receive data exclusively from real customer traffic via recordPiggybackResult(). The test scheduler excludes them from all runs (test-runner.ts line 117). They are never executed proactively.

Distribution PR Integrity Protocol (DEC-20260422-A)

MANDATORY — applies to ANY session that touches a PR on a repo outside strale-io/* OR that publishes / modifies a *-strale package.

Trigger: the session prompt mentions a PR on a framework repo (Pipedream, LangFlow, Flowise, pydantic-ai, langchain, crewAI, agno, composio, semantic-kernel, awesome-list, etc.), OR modifies files under packages/*-strale/, OR edits PyPI/npm publication metadata.

Background: on 2026-04-21 the pydantic-ai maintainer DouweM closed pydantic/pydantic-ai#4866 with "Shame on you" after finding that the published pydantic-ai-strale package contained zero pydantic-ai-specific code. An audit found two more packages with the same gap (google-adk-strale, openai-agents-strale). A prior agent session (2026-04-18) had edited the PR to trim promotional prose but did not verify the code example's imports. See archive/sessions/CONTAINMENT_REPORT.md for the full incident.

Required steps (non-negotiable):

  1. Verify every imported symbol before touching a distribution PR. Before editing the PR body, inline comments, code examples, or any description-style text on a repo outside strale-io/*, fetch the referenced Strale package's __init__.py / entry point via gh api repos/strale-io/strale/contents/packages/<pkg>/... and grep for every symbol the PR imports. If any symbol is not in __all__ or not exported, STOP and flag for Petter. Do not trim prose, do not fix bot findings, do not rebase, do not reply — nothing — while the PR contains a fabricated import.

  2. Run the distribution PR pre-flight checklist. See docs/governance/protocols/DISTRIBUTION_PR_PREFLIGHT.md. The four verifications (imports resolve, package on approved list, description matches code, tone matches neighbors) are the standard. All four must pass before the session opens or edits a distribution PR.

  3. Run the framework-package integrity check locally before shipping a new or modified *-strale package:

    node apps/api/scripts/check-framework-packages.mjs
    

    If the check fails, the package does not match its name. Either make the package live up to the name, rename it, or deprecate — do not publish.

  4. Never batch-create framework packages. One framework package per PR, each including (a) real framework-interface code importing from the framework, (b) at least one test exercising the framework's own primitives, (c) README content that only references what's in the module.

  5. Polishing is not a substitute for verification. A cleaner-looking PR containing a fabricated import is worse than the original. If a bot finding points at prose but the code example has an import problem underneath, fix the import first and the prose second, or stop and flag.

  6. Run the production contract smoke test after every publish. Install the published artefact the way a stranger would (npx -y <pkg>@<version>), run it against production, and record the observed startup values -- not "looks fine". Look specifically for errors the program logs and then continues from: a caught, logged, non-fatal degradation is what CI cannot see and users never report. CI tests source against doubles; it never runs the published artefact against production, and that gap is where this class of defect lives. On 2026-08-22 this check found strale-mcp had shipped 0 cap trust, 0 sol trust for ~3.5 months (routes deleted in May; the admin wall on their prefix answered 401 instead of 404; the client logged to invisible stderr and carried on). Procedure and the per-package contract: docs/release/npm-publishing.md.

At session end, report:

  • Every distribution PR touched, with the verification result for each.
  • Every package published, with the observed production smoke-test values.
  • Every *-strale package modified, with the check-framework-packages output.
  • Anything that couldn't be verified and why.

Do NOT mark a distribution task as done if the pre-flight checklist didn't pass. Report what's missing.

This rule does NOT override:

  • The Capability Onboarding Protocol (DEC-20260320-B).
  • Any PR-closure or code-change authorization that requires explicit Petter approval.

Capability Onboarding Protocol (DEC-20260320-B)

MANDATORY — applies to ANY session that creates, modifies, or onboards a capability.

Trigger: Claude Code detects that the session involves any of: new executor file in src/capabilities/, new or modified DB row in capabilities table, new capability slug, manifest file, seed entry, or the prompt mentions adding/creating a capability.

Rule: The Capability Onboarding Pipeline spec is the authority on HOW capabilities enter the system. The prompt describes WHAT to build. These are separate concerns. A prompt that says "add pep-check capability" without mentioning manifests, field reliability, or validation does NOT mean those steps are optional.

Required steps (non-negotiable):

  1. Read the spec first. Before writing any code, read the Capability Onboarding Pipeline design spec. If Notion is accessible, fetch page 32467c87-082c-819a-a731-d8a5f7237b33. If not, the key requirements are listed below.
  2. Create/update onboarding manifest (YAML file in repo) with: slug, name, description, category, schemas, pricing, data_source, transparency_tag, test_fixtures (known_answer + health_check_input), output_field_reliability for ALL output fields, and at least 1 limitation.
  3. Declare output_field_reliability for every output field: guaranteed (always present), common (usually present), or rare (sometimes present). Only guaranteed fields get not_null test assertions.
  4. Set avg_latency_ms — measure from test execution or estimate from transparency_tag (algorithmic=20ms, ai_generated=3000ms, mixed=2000ms, external API=check similar capabilities).
  5. Run structural validation: npx tsx scripts/validate-capability.ts --slug <slug>
  6. Run readiness check: Verify checkReadiness(slug) returns ready: true with zero issues.
  7. Run smoke test (if available): npx tsx scripts/smoke-test.ts --slug <slug>

At session end, report:

  • Readiness check result (pass/fail + any issues)
  • Steps completed that the prompt didn't mention
  • Steps that couldn't be completed and why

Do NOT mark a capability task as done if the readiness check fails. Report what's missing.

This rule does NOT override:

  • The prompt's specification of what the capability does (slug, schemas, pricing, implementation logic)
  • The DEACTIVATED list in src/capabilities/auto-register.ts

Audit-Follow-up Test Coverage Protocol (DEC-20260504-A)

MANDATORY — applies to ANY commit that introduces or substantially modifies a code path in response to a cert-audit finding (Y-, A-, B-, RED-, MED-, CRIT-, F-AUDIT- numbered findings).

Trigger: the commit message references a cert-audit finding code, OR the change adds/modifies a function that runs inside a wallet transaction, an audit-trail builder, a chain-integrity primitive, a spend-cap check, an idempotency check, or any other money/compliance-critical path.

Background: PR #43 (2026-05-04) fixed a Date-in-sql-template encoding bug in spendCapWouldExceed that shipped as part of cert-audit A-7 (commit 6613bd7, 2026-04-30). The same audit batch shipped a structurally identical bug in db-retention.ts (commit 968bc82); together they had been silently 500-ing paid /v1/do calls for capped users for 4 days and silently failing all retention pruning. Neither had test coverage. The audit pattern was producing new code paths without tests, and the bugs slipped past because the audit reviewers focused on the architectural correctness of the fix and not on its bind-encoder shape.

Required steps (non-negotiable):

  1. Every cert-audit follow-up commit that introduces a new code path must include at least one regression test that exercises the new path. The test does not need to be a full integration test — a unit test that captures the structural shape of the fix (e.g. "no Date instance reaches tx.execute(sql\`)` queryChunks") is sufficient.
  2. The test must fail against the un-applied fix and pass against the applied fix. Verify both directions during PR review. A test that passes regardless of the fix is not a regression test.
  3. If the new path runs DB writes, the test must assert on the bind-parameter shape — specifically, walk every tx.execute(sql\`)anddb.execute(sql``)call and verify no parameter is aDate, Buffer`, or other shape postgres-js's encoder cannot serialize. The PR-43 incident is the case study here.
  4. For commits that touch error-swallowing catch blocks, the test must include a "swallow visibility" assertion: when the swallowed error fires, the surrounding summary log must surface the failure (not silently report total: 0 or similar). The db-retention.ts fix is the case study — pre-fix every tick logged "pruned successfully" while every rule errored.

At session end, report:

  • Every cert-audit-finding-numbered commit landed in this session, with the test name(s) covering it.
  • Anything that landed without a test, and the explicit reason (e.g. "no test harness for /v1/do route-level integration; deferred per CLAUDE.md test-harness exemption").

Do NOT mark a cert-audit follow-up as done if the regression test gap is open. Report what's missing.

This rule does NOT override:

  • The Capability Onboarding Protocol (DEC-20260320-B).
  • The Distribution PR Integrity Protocol (DEC-20260422-A).

Test-harness exemption: if the relevant test harness doesn't exist in the repo (e.g. route-level integration testing for /v1/do requires a Postgres-backed harness that hasn't been built), the commit may ship with a unit-level regression test that captures the structural shape of the bug, plus an explicit reference to the missing harness in the PR description and Journal entry. The exemption does not apply to cases where the harness exists and the author skipped writing a test.

Bulk-Operation Deploy Protocol (DEC-20260504-B)

MANDATORY — applies to ANY deploy that fixes a long-silent bulk operation (retention, archival, reconciliation, batch processing, periodic cleanup).

Rule: When fixing a long-silent bulk operation, the deploy must include either: (a) a pre-fix backlog drain plan, or (b) a self-throttling fix that bounds resource usage per tick (e.g., LIMIT-paginated DELETE). Do NOT treat this as a normal bug fix. The first successful run after fixing a long-silent bulk operation is a workload-resumption event, not a routine execution. Audit accumulated workload BEFORE deploying. If accumulated workload could exceed infrastructure capacity (disk, memory, connection pool, rate limits, WAL volume), pre-cleaning or throttling is required.

Background: 2026-05-04 Postgres crash incident (Journal entry 35667c87-082c-8148-ae24-faee34f01c1d). PR #44's retention fix re-enabled bulk DELETE on accumulated rows that had been silently failing for the prior outage window. The first successful run filled the volume; Postgres crash-looped for 28 minutes. The fix itself was correct in isolation — the failure mode was treating "first successful execution after a long silent failure" as a routine deploy.

Required steps (non-negotiable):

  1. Identify the latency. If a bulk operation has been silently failing or unscheduled for >24h, the next successful run is a workload-resumption event. Treat it as one.
  2. Audit accumulated workload before merge. Query the table(s) the operation touches and estimate the row count, byte volume, and downstream effects (WAL bytes, replication lag, rate-limit consumption, connection-pool occupancy). Document the estimate in the PR body.
  3. Pick a deploy strategy explicitly:
    • (a) Pre-drain: ship a one-shot script that processes the backlog under operator supervision, then deploy the fix. The fix's first run sees a clean state.
    • (b) Self-throttle: ship a fix that bounds per-tick resource usage (LIMIT-paginated DELETE, batch-size cap, time-budgeted loop with early exit). The fix processes the backlog over many ticks.
  4. Reject the third option. "Just deploy the fix and let it run" is the failure mode this protocol exists to prevent. If the deploy strategy is "ship and pray," stop and pick (a) or (b).

At session end, report:

  • The accumulated-workload estimate and the chosen deploy strategy.
  • The first-successful-run outcome (rows processed, duration, peak resource usage).

Do NOT mark a bulk-operation fix as deployed if the accumulated-workload audit was skipped. Report what's missing.

This rule does NOT override:

  • The Capability Onboarding Protocol (DEC-20260320-B).
  • The Distribution PR Integrity Protocol (DEC-20260422-A).
  • The Audit-Follow-up Test Coverage Protocol (DEC-20260504-A).

Deploy Mechanism Verification Protocol (DEC-20260504-C)

MANDATORY — applies to ANY PR that adds a code path which depends on a deploy-pipeline behavior (migrations running, env vars read, build steps, startup hooks, scheduled jobs, cron triggers).

Rule: When a PR adds a code path that depends on a deploy-pipeline behavior, the pre-merge audit MUST verify the deploy mechanism actually does what's expected. "Code is correct" is not the same as "deploy of code will produce expected effect on prod." Verification means reading the actual deploy mechanism (Dockerfile, build config, startup wiring, package.json scripts, CI workflows) and confirming the new code path will execute as intended. Do NOT assume historical patterns hold without verification. A clean post-deploy log is not verification — query prod for the expected effect of the change.

Background: 2026-05-04 PR #42 outage (28-minute customer-facing 500s). apps/api/scripts/apply-migrations.ts had been a dead file — never compiled (excluded by tsconfig.json rootDir), never invoked by Dockerfile CMD. PR #42's pre-merge audit was thorough on PR contents but didn't verify the deploy mechanism actually ran the script. Compounded by PR #49's UPDATE blocks also silently never running for 7 hours under the same dead-file mechanism. The structural fix shipped in PR #51 + PR #52 (runStartupMigrations() wired into index.ts:69, single source of truth between admin endpoint and startup wiring).

Required steps (non-negotiable):

  1. Identify the deploy-pipeline dependency. If the new code path depends on something other than its own module being imported and called from request-handler code, name the dependency: which deploy step has to fire? Which env var has to be read? Which startup hook has to invoke it?
  2. Read the actual deploy mechanism. Open the Dockerfile, the CMD line, package.json scripts, index.ts startup wiring, the relevant CI workflow, the cron config — whatever produces the dependency. Confirm the new code path is reached.
  3. Confirm reach by file path, not by historical pattern. "Migrations have always run" or "env vars are always loaded" is not verification. The verification is: this specific file is on the import graph from index.ts (or whichever entry point fires at deploy time), or this specific script is invoked by a specific line in the Dockerfile / package.json.
  4. Post-deploy: query prod for the expected effect. A clean log line proves the line was emitted. It does not prove the schema changed, the row was written, the env var was read, the cron fired. Query the actual artifact: \d table for a column add, SELECT COUNT(*) for a backfill, GET /health/version for a build SHA, etc.

At session end, report:

  • The deploy-pipeline dependency identified, and the file/line that proves the dependency is satisfied.
  • The post-deploy prod query and its result.

Do NOT mark a deploy-mechanism-dependent change as done if the verification step was skipped. Report what's missing.

This rule does NOT override:

  • The Capability Onboarding Protocol (DEC-20260320-B).
  • The Distribution PR Integrity Protocol (DEC-20260422-A).
  • The Audit-Follow-up Test Coverage Protocol (DEC-20260504-A).
  • The Bulk-Operation Deploy Protocol (DEC-20260504-B).

Quick Session Checklist

  1. Declare session intent
  2. Connectivity check (Git + handoff; Notion if needed). Log failures.
  3. Read handoff/from-chat/ for pending items (if empty, proceed)
  4. Do the work
  5. Code-review gate (before /end-session): if any code was modified this session and /go was not run on it, halt and run /go (or escalate to Petter if a hard refusal blocks /go). Never run /end-session over unreviewed code. Docs / CLAUDE.md / Notion-only sessions are exempt.
  6. Move completed To-do items to Archive > Completed To-dos (page ID: 34067c87-082c-814e-a45c-fa8d851c8f12)
  7. Write handoff file to handoff/_general/from-code/. Even one-liner, starts with Intent:
  8. Create Journal entry in Notion (even one line)

Full Session Checklist

  1. Declare session intent
  2. Run full Pre-Build Connectivity Checklist. Log failures.
  3. Read Project Home → current focus
  4. Read last 5 relevant Journal entries filtered by feature
  5. Read active Decisions — global always, feature-scope when relevant
  6. Read handoff/from-chat/ for pending specs or feedback
  7. Do the work
  8. Code-review gate (before /end-session): if any code was modified this session and /go was not run on it, halt and run /go (or escalate to Petter if a hard refusal blocks /go). Never run /end-session over unreviewed code. Docs / CLAUDE.md / Notion-only sessions are exempt.
  9. Move completed To-do items to Archive > Completed To-dos (page ID: 34067c87-082c-814e-a45c-fa8d851c8f12)
  10. Create Journal entry (full format)
  11. Log decisions made (respect authority thresholds)
  12. Save session summary to handoff/_general/from-code/
  13. Contradiction check if decisions were made

Shared-Checkout Rule (concurrency safety)

This checkout is shared. Several Claude Code sessions and background agents run against the same working tree at once. Git branch switching is not concurrency-safe here.

The failure mode: an agent runs git checkout <branch> in the main tree while another process holds file locks (tsc, vitest, npm, an editor). On Windows git's delete-then-rewrite sequence fails partway — the old files are unlinked, the new ones never written. ~1,000 tracked files vanish from disk while the index still lists them, always in apps/api/** and packages/** where node holds handles. Hit three times on 2026-08-14.

Rules:

  1. Any agent that edits files MUST be launched with isolation: "worktree". An agent working in the shared checkout will eventually collide with the main loop or another agent. This is the actual prevention.
  2. Agents must never git checkout a branch in the main tree. If an agent without worktree isolation needs a branch, it should create its own worktree (git worktree add) rather than moving the shared one.
  3. Before branch-switching in the main tree, check git status for another session's uncommitted work. Uncommitted changes travel across branch switches and can end up staged onto the wrong branch.
  4. Never "fix" phantom breakage. If files that are committed suddenly ENOENT, that is this bug, not a real deletion. Run node scripts/guard-tree-integrity.mjs (or just any Bash command, if the PostToolUse hook is wired) and re-check before diagnosing further.
  5. Never use git stash in any worktree of this clone. refs/stash is repo-wide, shared across ALL worktrees — concurrent sessions' stash push/pop interleave, and a pop in one worktree can consume (and on conflict, destroy) another session's stashed work. Hit on 2026-08-16: one agent's stash pop returned a sibling agent's quality-floor changes. For fail-before verification or temporary reverts, use git checkout <base-sha> -- <paths> + git checkout <branch> -- <paths> to restore, or a temporary WIP commit. If a stash accident happens, recover via git fsck --dangling (stash commits survive as dangling commits) and save the foreign diff to a patch file — never discard it.

The guard at scripts/guard-tree-integrity.mjs auto-repairs the damage and is wired as a PostToolUse/Bash hook in .claude/settings.json (tracked since T3 alongside the session hooks; machine-local additions go in the ignored settings.local.json). It only ever restores tracked-and-deleted paths, so it cannot discard work. It is a safety net, not a substitute for rule 1.

Worktree node_modules Hazard

Second failure mode: creating a Windows directory junction from a temporary worktree to the main checkout's node_modules, then later removing that worktree with rm -rf, follows the junction and deletes the real node_modules in the main tree. Symptom: every command fails with ERR_MODULE_NOT_FOUND for packages that are definitely installed. No source is lost — it is generated — but recovery requires npm install at the repo root followed by npm --workspace=packages/mcp-server run build, or src/routes/mcp.ts shows phantom type errors. The rule: run npm install inside each worktree instead of linking, and remove worktrees with git worktree remove, never rm -rf.

Report Filing Convention

Large investigative/audit/session reports (AUDIT-, FIX_PHASE_, SESSION_, RESOLUTION_REPORT, REVIEW_FINDINGS_, *_INVENTORY, *_RESEARCH, checklists tied to a specific incident, etc.) route to archive/sessions/ — flat layout for individual reports (see existing entries for naming convention); directory sweeps imported wholesale (Phase 3, 2026-08-17: audit/, audit-output/, audit-reports/, a2a-sample/, tasks/, capability-sources/, distribution/) keep their internal structure as archive/sessions/<dirname>/. Never to the repo root.

Root contains exactly: README.md, CLAUDE.md, AGENTS.md (+ .agents/, .codex/), WORKTREES.md, LICENSE; the monorepo build files (package.json, package-lock.json, tsconfig.json); and four MCP/package-registry discovery manifests each required at repo root by the registry that reads it — context7.json (Context7), glama.json (Glama MCP directory), server.json (the official MCP registry, io.github.strale-io/strale), smithery.yaml (Smithery stdio install config). None of the four can move: each registry's crawler looks for its file at the repository root by convention, not at a configurable path. (T5, 2026-09-02: the two review/pre-flight checklists that previously sat at root — DISTRIBUTION_PR_PREFLIGHT.md, REVIEW_TEMPLATE.md — moved to docs/governance/protocols/; nothing reads them from the root path.) AGENTS.md is a condensed derivative of CLAUDE.md for Codex-CLI sessions — CLAUDE.md is canon; AGENTS.md points at CLAUDE.md sections for anything that can drift rather than restating it. Refresh AGENTS.md against CLAUDE.md whenever drift is noticed (new decisions, new mandatory protocols, Shared-Checkout Rule changes) rather than letting it go stale again.

Workflow Invariants (Non-Negotiable)

  • NEVER edit Journal entries, Decision content, or Deferred content
  • NEVER delete anything in Notion
  • Corrections → new Journal entry, type = course-correction
  • Global decisions → ALWAYS get confirmation
  • Supersessions → ALWAYS use Contradiction Protocol (including CLAUDE.md update)

Conflict duty: If the human's request would contradict an active Decision, state the conflict before proceeding. Quote the specific Decision being violated and ask the human to confirm, supersede, or revise.

Degraded Mode

If Notion unavailable: work continues, log to handoff files with [BACKFILL] prefix. If Git unavailable: STOP. Fix before proceeding.

Cross-Repo Updates

When making changes to the backend repo, check if these files in the frontend repo (strale-frontend) need updating:

  • public/llms.txt — Update when: adding/removing/renaming capability categories, adding new SDKs or integrations, changing API endpoints or auth flow, changing pricing model. This file is what LLMs read when someone shares strale.dev — it must stay accurate.
  • public/sitemap.xml — Regenerate when: adding new pages or routes. Run: npx tsx scripts/generate-sitemap.ts in strale-frontend.
  • src/lib/compliance-types.ts AuditRecord interface — Update when: any change to AuditRecord in apps/api/src/routes/audit.ts. The two declarations must match shape (field names + types). The CI check at apps/api/scripts/check-shape-contracts.mjs (registry-driven; run with --list to see registered contracts) enforces this and runs in the weekly drift cron with both repos checked out. If you add/remove/rename a field on the backend without updating the frontend, the check fails and an issue is auto-opened. Add new shared interfaces to the CONTRACTS array in that script when applicable.

Wire-shape rule for /v1/public/ops/trust/* endpoints

The trust endpoints (and any future endpoint surfacing money values, scores, or anything formattable) MUST emit canonical machine-readable values, NEVER pre-formatted display strings. Specifically:

  • Money is always integer cents (*_cents), never a formatted string like "€0.02".
  • Scores are 0-100 integers or 0-1 decimals (be consistent within an endpoint).
  • Dates are ISO 8601 strings.
  • If you need to ship a pre-rendered display value alongside the canonical one, add a *_formatted field — additive, never replacing.

Why: a 2026-04-30 cert-audit finding traced the empty "SQS 0 / Price unavailable" fallback card on capability detail pages to a serializer that emitted fallback_price: "€0.02" (string) and dropped the integer. The frontend normalizer either had to regex-parse currency or read a fictitious *_cents field that defaulted to 0. Removing the formatted strings entirely was cheaper than maintaining a deprecated lossy field forever. Display formatting is the consumer's responsibility.

For wire-shape ↔ consumer-shape contracts (where backend uses one set of names and the frontend normalizer maps to different names — e.g. backend fallback_capability → frontend capability_slug), the shape-check script can't help (the names are different by design). Use a frozen-fixture contract test instead. Pattern: strale-frontend/src/lib/api.contract.test.ts with the fixture in src/lib/__fixtures__/. Re-capture the fixture when the wire shape legitimately changes (the test file's docstring has the curl command); the test fails loudly when a normalizer change drops a field or reads a wrong name.

Drift-prevention surfaces

When changing facts that appear on multiple surfaces (capability count, country count, retention period, vendor names, free-tier list, processing region), update only the canonical source and let consumers read from it:

  • Backend canonical source: apps/api/src/lib/platform-facts.ts — STATIC_FACTS for fixed values, computePlatformFacts() for live-DB values. Exposed via GET /v1/platform/facts (cached 5 min).
  • Frontend consumer: usePlatformFacts() hook in strale-frontend/src/hooks/use-platform-facts.ts. Component pages read from this; never hardcode the displayed value.
  • Static frontend files that can't reach the hook (public/llms.txt, public/.well-known/*.json): use phrasing that doesn't bake in counts, with a pointer to /v1/platform/facts.

The apps/api/scripts/check-platform-facts-drift.ts guard (run via npx tsx, wired in weekly-drift.yml) catches new hardcoded values introduced into surface files. The weekly cron runs the same sweep across both repos and opens a tracking issue on any drift.

For vendor switches specifically, invoke the vendor-switch skill (in .claude/skills/vendor-switch/SKILL.md) — it codifies the full surface-update + DEC-entry checklist.