Skip to content

v1.1.0: sync security and correctness stabilization, payload-markdown 1.6 peer, upgrade migration helper - #10

Merged
cooperlarson merged 61 commits into
mainfrom
stabilization/2026-10
Oct 2, 2026
Merged

cooperlarson merged 61 commits into
mainfrom
stabilization/2026-10

Conversation

@cooperlarson

Copy link
Copy Markdown
Contributor

Summary

1.1.0 is the docs plugin's stabilization release, the counterpart of payload-markdown 1.6.0. It fixes the sync, security and CLI findings from the joint audit, and makes @valkyrianlabs/payload-markdown (now ^1.6.0) and @payloadcms/plugin-seo peer dependencies. CI now tests a real consumer install.

This release needs a database migration. The new schema is additive, but one unique index fails on duplicate nonce rows that older versions allowed. The release ships a helper for that, plus an upgrade guide: docs/reference/upgrade-1-1.md.

Upgrading (for users)

  1. Install the peers. Run pnpm add @valkyrianlabs/payload-markdown@^1.6.0 @payloadcms/plugin-seo, then payload generate:importmap.
  2. Migrate. Run payload migrate:create, then add one line at the top of up, then run payload migrate:
    import { prepareDocsSyncMigration } from '@valkyrianlabs/payload-markdown-docs/migrations'
    
    export async function up({ db, payload, req }: MigrateUpArgs) {
      await prepareDocsSyncMigration({ payload, req }) // removes expired + duplicate nonces
      await db.execute(sql`...generated...`)
    }
    • Dev push mode: if startup fails on keyId_nonce_idx, run the two DELETE statements in the guide, then start again.
    • MongoDB: a payload run script with the same helper.
  3. Review the changed defaults:
    • plugin collections are admin-only;
    • asset changes from draft syncs apply on the next --publish sync;
    • forwarded headers are opt-in.

Schema delta:

  • new tables docs_sets_repositories and docs_access_rels;
  • new columns docs_sets.allow_tag_refs (existing rows get true) and docs.sync_fields_hash_at_last_sync, plus their version-table copies;
  • unique index keyId_nonce_idx on docs_sync_nonces (key_id, nonce).

Upgrade validated on Postgres. I built a 1.0.5 schema from origin/main via migrations and seeded it with expired and duplicate nonces. Then I upgraded to this branch:

Upgrade path Result
migrate:create without the helper migrate fails on the index and rolls back cleanly
migrate:create with the helper Migrates; removes 2 expired and 1 duplicate nonce, keeping the oldest; the existing docs set gets allow_tag_refs = true; a follow-up migrate:create --skip-empty finds no drift
Dev push mode with duplicates Startup fails; after the documented DELETEs, it starts and finishes the schema

Security

  • Access: plugin collections are admin-only by default (config.admin.user) and can be overridden. Sync runs and nonces are read-only for humans.

  • Replay protection:

    • nonces are inserted first, under a unique (keyId, nonce) index;
    • concurrent duplicates are rejected;
    • nonce and jti TTLs are bounded by the skew window;
    • expired nonces are pruned.

    If the index is missing (failed Mongo index build, migration not run), a post-insert check still rejects replays and fails closed.

  • Credential scope:

    • Ed25519 keys and OIDC owner records can be limited to specific docs sets (unscoped keys still work, with a warning);
    • OIDC tokens can be bound to repositories;
    • allowTagRefs is on by default, so tag-triggered publishing keeps working;
    • pull request tokens are checked against base_ref.
  • Pre-auth surface:

    • the request body read is bounded;
    • source.id must be slug-shaped;
    • authentication runs before docs-set lookup;
    • unauthenticated callers get a uniform 401;
    • JWKS fetches are cached with a timeout, and an unknown kid triggers a rate-limited refetch.
  • Assets:

    • a shared content-type allowlist;
    • nosniff and a CSP sandbox;
    • attachments for non-text types;
    • asset routes are confined to their prefix.
  • Origins: the configured origin wins; X-Forwarded-* headers are trusted only with endpoint.trustForwardedHeaders.

Sync correctness

  • Visibility: one module owns public visibility (drafts, archived records, unpublished sets) for pages, nav, sitemap, llms.txt, and CTA/hero blocks. Unpublished docs no longer leak through any of them.
  • Atomic apply:
    • route claims are resolved before any write;
    • archives tombstone their routes;
    • apply runs in one transaction where the adapter supports it, and a failed sync leaves no partial state.
  • Removals:
    • removals reach the published record;
    • assets: [] archives existing assets;
    • the shared TS and C++ planners skip records that are already archived.
  • Admin drafts: docs-set bookkeeping no longer publishes admin drafts.
  • Manual-edit protection: the field hash covers title, description, navTitle and order.
  • Errors: an error.issues contract, no raw DB messages, and correct sync-run codes.

CLI (pmdocs) and TS↔C++ contract

  • Shared rules: shared contract vectors (contracts/vectors) pin frontmatter, order, titles, routes, body limits, content types and UTF-8 handling to identical behavior in TS and C++. pmdocs matches the server on all 112 valid corpus files.
  • Nothing dropped silently:
    • hidden files are published with a warning (--skip-hidden);
    • build, dist and .next are skipped only at the docs root, with a warning.
  • New output and flags:
    • prints issues, routeCollisions and conflicts;
    • validate detects route collisions;
    • new --route-base, --asset-route-base, --publish and --existing-assets.
  • Security:
    • refuses plain http except on loopback (--allow-insecure-http);
    • writes keys and output atomically with 0600/0700 permissions.
  • Packaging:
    • the Homebrew formula bundles skill data;
    • doctor reports degraded when skill data is missing;
    • requires Meson ≥ 1.3.

Architecture

  • Sync endpoint: split into stages (src/endpoints/sync/*: parse → authenticate → policy → plan → apply).
  • Shared modules: record helpers (src/shared/records.ts), one route derivation (src/routing/docsSetRoutes.ts, 7 copies removed), one docs-set resolver.
  • Performance: root llms.txt runs 10 queries instead of 2+2N.
  • Packaging: core, plugin-seo, next, react and payload ^3.0.0 are peers; the new ./migrations export holds the upgrade helper.
  • Fixes:
    • apps without a media collection no longer crash;
    • the llms-full link rewriter ignores code spans and fences and handles linked images.

CI and release

  • Build and Test:
    • actionlint, lint, build, dev-app typecheck and unit tests;
    • the Postgres-backed suite and Playwright e2e against postgres:16 service containers;
    • native CLI build, tests and install smoke;
    • release tooling pytest and tools.release check.
  • Consumer smoke: packs this plugin and payload-markdown main into a fresh Next + Payload app with pnpm, runs next build, does a signed pmdocs sync and checks the rendered output (13 checks).
  • Release:
    • DeepSeek changelog provider support (org-level VL_DEEPSEEK_API_KEY / VL_AI_RELEASE_PROFILE);
    • least-privilege permissions and SHA-pinned actions;
    • pull requests run on GitHub-hosted runners only.
  • Self-hosted runner: e2e installs Playwright's system packages with plain apt-get, which the runner's sudoers allows.
  • Version: 1.0.5 → 1.1.0 via tools.release bump minor (VERSION, package.json, meson, debian, Homebrew).

Validation

All gates were run locally on this branch's final tree, on fresh scratch Postgres databases:

  • Static checks: pnpm install --frozen-lockfile, pnpm build, pnpm lint, dev-app typecheck: pass.
  • vitest unit: 716 passed, 82 DB-gated skipped.
  • vitest against Postgres: 798/798, 27 files. Includes the helper and concurrency regressions.
  • Playwright e2e: 2/2.
  • Native CLI: meson test 13/13; pmdocs --version prints 1.1.0.
  • Release tooling: pytest 346 passed; tools.release check OK.
  • Workflows: actionlint clean.
  • Consumer smoke with the published @valkyrianlabs/payload-markdown@1.6.0 tarball, this package's tarball and the 1.1.0 pmdocs: 13/13. That covers no peer warnings, a single core copy, the import map resolving, next build, a signed sync and rendered pages.
  • Upgrade from 1.0.5: validated on Postgres as described above.

Known follow-ups (not in this PR)

  • X-4: docs pages render two <h1> elements and a nested <article>. Removing one changes appearance and needs an owner decision.
  • DOCS-11, doc-level: a --publish sync still publishes CMS draft edits to fields the sync does not manage (e.g. hero image).
  • X-5: the llms-full link rewriter does not handle indented code blocks; this needs an AST URL hook.
  • Identifier length: Payload 3.90 warns that the _docs_sets_v_version_advanced_security_allowed_workflow_refs_* index names exceed Postgres's 63-character limit. This predates the PR, and no schema drift was observed. Shortening the names with dbName would itself need a migration.
  • Debian changelog: debian/changelog still has the skeleton entry text (the version header is bumped).
  • ESLint 10: deferred; it is a major upgrade together with @payloadcms/eslint-config.

Port the provider-agnostic release AI pipeline from vaulthalla's release
tooling and add the ds-flash profile, keeping payload-markdown-docs'
own release behavior (Homebrew, major-release framing, previous-tag
policy, empty-tree first-release diff).

Provider support
- deepseek provider kind over the OpenAI Responses transport
  (https://api.deepseek.com/v1), per-stage provider/base_url so profiles
  can mix providers, `max` reasoning effort (mapped to xhigh for OpenAI)
- strict-compatible JSON schemas for strict_json_schema requests and
  null-optional cleanup of strict responses (providers/strict_schema.py)
- provider-aware key gating and bounded per-request timeouts
  (RELEASE_AI_PROVIDER_TIMEOUT_SECONDS / ..._EMERGENCY_TRIAGE_...)
- emergency triage fallback evidence, zero-items guard and context
  artifact; stale debian/changelog guard; explicit --since-tag equal to
  the current tag is rejected; agent tooling dirs scored as noise

Environment
- DeepSeek key resolves from VL_DEEPSEEK_API_KEY (valkyrianlabs org
  secret / host shell), then DEEPSEEK_API_KEY; errors name the
  variables, never values
- profile resolves from VL_AI_RELEASE_PROFILE (org variable), then
  RELEASE_AI_PROFILE, then legacy RELEASE_AI_PROFILE_OPENAI
- default profile is now openai-premium-deep-triage, which exists in
  ai.yml (the old openai-balanced default did not)
- release.yml passes the org-level VL_* secret/variable to the
  changelog step; no hard-coded profile fallback
- /.bashrc (repo-local, untracked) is gitignored for mapping host
  variables without committing secrets

ai.yml gains ds-flash verbatim from vaulthalla. Tests: 309 -> 342,
including ported provider/strict-schema/stale-guard/noise-path tests,
org env-name precedence tests, and test isolation so the CLI changelog
tests no longer write .changelog_scratch/ into the checkout.

Live smoke: ds-flash draft stage via VL_DEEPSEEK_API_KEY, strict
json_schema round trip succeeded.
Fixes the only pnpm lint error (perfectionist/sort-named-exports) so the
repository lints clean; lint was not run in CI, so this went unnoticed.
`pnpm update` (repository policy, no --latest): payload/@payloadcms/*
and @payloadcms/plugin-seo 3.85.1 -> 3.90.2, next/eslint-config-next
16.2.9 -> 16.3.8, tailwind/postcss/prettier/eslint patch bumps.
@valkyrianlabs/payload-markdown stays at 1.5.0 (latest on npm).

No compatibility fixes were required.

Validation: frozen install, build, lint, test:int 288/288 (with the
Postgres-backed suite), e2e 2/2, native meson test 12/12, release
tooling pytest 342, tools.release check, npm pack; docs also builds and
passes 288/288 against a locally packed core build of the updated
payload-markdown.
Payload 3 `find({ draft: false })` still returns never-published documents
(main-table rows with `_status: 'draft'`). llms.txt / llms-full.txt (per docs
set and root) checked only `sync.archived`, so non-publish syncs leaked full
unreleased markdown publicly (DOCS-1). The sitemap docs query did not select
`_status`, so isVisibleDocsRecord treated drafts as visible (DOCS-12).

- add src/payload/visibility.ts: one predicate set for raw public records
  (docs: published and not archived; docs sets: published; assets: not archived)
- llms doc/skill records and the public docs-set lookups (findAllDocsSets,
  findDocsSetByRoutePrefix) use it
- sitemap selects `_status`
- next.spec.ts mock now mirrors real Payload semantics (draft:false returns
  draft-only rows; `select` drops unselected keys) instead of hiding the bug
- add a real-Postgres regression suite (dev/regressions.spec.ts) and harness;
  DB runs execute spec files serially because Payload's dev schema push races
… removals

Archived docs kept their unique `route`, so moving guide.md to
guide/index.md, renaming a file that keeps `slug:`, or swapping slugs hit the
unique index mid-apply. Writes were not transactional, so earlier writes
stayed committed, revalidation was skipped, and every retry failed the same
way (DOCS-2). Removals in non-publish syncs only wrote a draft version, so a
doc deleted from Git kept being served (DOCS-3).

- archiving (and deleteBehavior 'draft') writes the main record with a
  released tombstone route `archived:<id>:<route>`; reactivation restores it
- plan-time route claim resolution against main-table rows: same-set holders
  that are archived, draft-only, or moving in this sync are released (route-only
  db.updateOne, no new version); other owners and published docs this sync
  does not publish are rejected with 409 route_collision before any write
  (dry-runs report it too)
- write order: removals, releases, updates, reactivations, creates
- docs, assets, and docs-set bookkeeping run in one Payload transaction when
  the adapter supports it; failures roll back
- failed applies mark the sync run `failed` with the endpoint error code
  (was recorded as `invalid_manifest`, and in-apply conflicts left runs
  `pending`); raw DB/validation messages are logged server-side and no
  longer echoed to clients (DOCS-17 partial)
- real-Postgres regressions: move, swap, rename-with-slug, draft swap then
  publish, pre-fix archived holders, rollback leaves no partial state,
  non-publish removal takes the published doc offline
…, lift 1000-row caps

- assets: existing assets are always loaded when the assets collection is
  enabled, so `assets: []` archives every live asset instead of leaving
  removed skills/llms files public forever (DOCS-9). A missing assets table
  still does not fail docs-only manifests.
- docs and assets that are already archived are dropped from the archive and
  draft removal lists server-side: no repeated `archive: N`, no new version
  per sync, `archivedAt` keeps the original date (DOCS-10). The shared
  planner still proposes them; the planner change is requested from the
  protocol workstream.
- write-side and public lookups use `pagination: false` instead of a silent
  `limit: 1000` (existing docs/assets, route collisions, trusted OIDC
  sources, docs sets/groups, llms, skill assets) (DOCS-19).
None of the plugin collections defined `access`, so Payload's default
(`Boolean(req.user)`) let any authenticated user of any auth collection
(customers, members) register their own Ed25519 sync key through REST or
GraphQL, delete replay nonces, and edit docs and sync runs (DOCS-6).

- docs, docs sets, docs groups, docs assets, and Access records: admin-only
  (users of `config.admin.user`) for every operation
- sync runs and nonces: read-only for admins; written only by the sync
  endpoint through the Local API
- `access.admin` customizes the admin rule; `collections.<key>.access`
  overrides single operations
- the sync endpoint and `/next` read helpers keep using overrideAccess, and
  existing admin-collection users keep full access
…ction

A consumer app without a `media` upload collection crashed Payload config
sanitization with `InvalidFieldRelationship: Field Meta Image has invalid
relationship 'media'` from the docs-set SEO group; opting into heroes or the
docsCTA block crashed the same way on their background media fields (X-17).

- when `media` (or a listed additionalMediaCollections slug) is not in the
  incoming config, omit the docs-set SEO meta image, the docs heroImage field,
  and upload fields inside installed heroes/blocks that only reference it, and
  log one warning
- apps that define `media` get an identical collections config (verified by
  serializing the plugin output before and after)
- wiring tests that assert media fields now declare the media collections
… key handling

Replay protection was check-then-insert without a unique index, so four
concurrent requests with one nonce were all accepted, and the nonce was only
stored after planning, so requests rejected after authentication could be
replayed (DOCS-8). OIDC jti rows expired at `exp` although tokens are
accepted until `exp + maxSkew` (DOCS-7). The docs set was looked up before
authentication (an existence oracle and pre-auth DB work), `source.id` of any
JSON type reached the query, the body was buffered before the size check,
and OIDC discovery was fetched on every request (DOCS-15). An unknown `kid`
after GitHub key rotation failed for up to 5 minutes (DOCS-16).

- nonces: unique (keyId, nonce) compound index; consumeNonce inserts right
  after authentication and treats an insert failure with an unexpired row as
  replay; expired rows are pruned at most every 10 minutes (DOCS-20)
- Ed25519 nonce expiry covers the signed timestamp + maxSkew (+1s); OIDC jti
  expiry is exp + maxSkew; endpoint maxSkewSeconds is passed to OIDC
- request order: body read with a Content-Length/stream limit, `source.id`
  slug check (no DB), authentication, then docs-set lookup; OIDC is split into
  identity (signature, iss, aud = source.id, time, trusted owner/repo) and
  docs-set policy (refs, workflow refs, pull requests)
- OIDC discovery cached for an hour; fetches time out after 5s; an unknown
  kid forces one JWKS refetch, rate-limited to once a minute per URL
- `sync.auditDryRuns` (default true) makes dry-run sync-run rows optional
- unit mock now fails duplicate nonce inserts like the unique index does;
  real-Postgres tests cover the race, post-auth burn, and oracle
…docs set

Authorization was global: any registered Ed25519 key could sync, publish,
and archive any docs set, and any repository under a trusted OIDC owner
could mint a token for any docs set slug, with any `refs/tags/*` ref
bypassing the docs set branch (DOCS-5, CLI-8). The docs described
docs-set-owned keys and allowlists that did not exist.

- Access records gain an optional `Allowed docs sets` relationship (Ed25519
  and GitHub OIDC). Empty keeps today's behavior; an unscoped Ed25519 key logs
  a one-time warning. Out-of-scope requests get 403.
- docs sets gain `Allowed repositories` (OIDC repository binding; empty =
  any trusted) and `Allow tag refs` (default on, preserving release-tag
  publishing; existing records without the field count as on)
- docs (sync config, GitHub OIDC, security model) describe the actual
  semantics and how to narrow them
`updateDocsSetAfterSync` updated the docs set with `_status` and
`draft: !publish`. Payload merges an update onto the latest version, so a
`--publish` sync published whatever unpublished admin draft the docs set had
(hero, SEO, description), and every non-publish sync wrote a new docs-set
draft version (DOCS-11).

- sync bookkeeping (`sync.lastStatus`, `sync.lastSyncedAt`) is written to the
  main record with a route-only style `db.updateOne`: no version, `_status`
  untouched
- the first `--publish` sync of a never-published docs set still publishes it
  (the existing way synced docs become reachable)
- runs inside the sync transaction

Generated docs records are not changed here; see the report for the remaining
doc-level case (non-synced fields edited as drafts).
…torage errors

Rejected syncs gave the operator no actionable detail: `invalid_manifest`
discarded the validator's issues and warnings, and an intra-manifest route
duplicate was reported as a collision "with an existing route reservation"
(DOCS-17, server side of CLI-6). Any error mentioning the assets collection,
including validation and unique-constraint errors, was reported as "assets
schema missing".

- every non-2xx sync response keeps its shape and may carry
  `error.issues: { code, message, path?, severity: 'error' | 'warning' }[]`:
  validator issues (error) and warnings (warning) for `invalid_manifest`,
  one entry per route for `route_collision` (duplicate manifest routes name
  the files that produce them), one per doc for `manual_edit_conflict`;
  `routeCollisions` / `conflicts` stay
- assets storage is "unavailable" only for a missing table (Postgres
  `relation ... does not exist`, SQLite `no such table`), checked through the
  error cause chain
- troubleshooting docs describe the response shape
…fer draft assets

The manifest's asset `contentType` was stored as given and replayed verbatim,
without nosniff, CSP, or Content-Disposition, so any sync principal could
serve `text/html` with script on the site origin (stored XSS against admins
browsing the site). Assets also went live immediately from syncs without
`--publish`, while the same sync's docs stayed drafts (DOCS-4, CLI-22).

- server-side allowlist per asset kind (text/markdown, text/plain; skills and
  static also JSON/YAML; static also CSV), matching what pmdocs emits;
  violations return `invalid_manifest` with `error.issues`
- all asset, llms, skill, zip, and 404 responses send
  `X-Content-Type-Options: nosniff` and `Content-Security-Policy:
  default-src 'none'; sandbox`; non-text types are attachments; stored rows
  with a disallowed type are served as text/plain
- with drafts enabled, non-publish syncs defer asset creates/updates to the
  next `--publish` sync (warning `assets_deferred_until_publish`); removals
  apply; `sync.applyAssetsOnDraftSync: true` restores the old behavior
Generated llms URLs took `X-Forwarded-Host` and `Host` before Payload
`serverURL`, so a client header changed canonical URLs in llms.txt and
llms-full.txt (DOCS-18).

- configured origins (NEXT_PUBLIC_SERVER_URL, NEXT_PUBLIC_SITE_URL, SITE_URL,
  Vercel URLs, then `serverURL`) always win
- forwarded headers are used only with `endpoint.trustForwardedHeaders: true`
  and only when no origin is configured; otherwise Host, then the request URL
Local runs of `python -m tools.release` write `.changelog_scratch/`
(cached drafts, AI failure artifacts) and `release/` (staged packages,
changelog outputs) into the checkout. Neither is ever committed, so
ignore both to keep stray drafts out of commits (CLI-19).
`cpp_std=c++23,c++20` uses the comma-separated standard fallback syntax
that Meson only understands from 1.3.0. Meson 1.2.x accepted the
declared `>=1.2.0` minimum and then failed with a confusing combo-option
error. Declare the real minimum in meson.build and the Debian
Build-Depends so older Meson fails fast with a version message (CLI-15).

Verified: Meson 1.2.3 now stops with "project requires >=1.3.0";
Meson 1.3.2 configures, builds and passes `meson test`.
Add `contracts/vectors/*.json` (contract version 1) and a vitest runner
(`src/sync/contracts.spec.ts`) so the server and the native CLI can be
checked against the same cases, and fix the server-side defects the
audit found while writing them:

- frontmatter (DOCS-13, CLI-20): strip a leading BOM; treat CRLF, CR and
  LF as line endings; ignore full-line comments; accept `_`/`-` in keys
  (unknown keys warn instead of failing the sync); ignore everything
  nested below unknown keys instead of letting `seo: { title }` override
  top-level fields; report nested/multi-line values and block scalars on
  known fields with a clear message; support single-line flow lists;
  keep a lone `"` literal; warn on empty `order`/`slug`/`title`
- titles (X-9, CLI-3): infer from the first top-level H1 only (ATX or
  setext), skip fenced/indented code, blockquotes and lists, strip
  inline markup (code, emphasis, links, images, autolinks, escapes);
  an empty frontmatter title falls back to inference; an empty filename
  title becomes "Untitled"
- routes (DOCS-22): reject route segments with `?`, `#` or control
  characters (`invalid_route`), warn on whitespace (`route_whitespace`);
  add `findManifestRouteCollisions` for exact and ASCII case-insensitive
  in-manifest collisions (X-8, CLI-6)
- assets (DOCS-14): derive skill routes from the docs set (client routes
  ignored with `asset_route_ignored`), confine llms/llms-full/static
  routes and reject `.`/`..`, `%`, `?`, `#` and control characters
  (`invalid_asset_route`)
- export the asset content-type allowlist (CLI-22) and the body-size
  helpers (CLI-5) used by the vectors

Route oracle: all 113 Markdown files in the core docs, these docs,
examples/docs and dev/docs-fixtures derive identical routes, titles,
frontmatter and issues before and after this change.
Move the protocol semantics out of docs.cpp into
`cli/src/sync_contract.cpp` (`pmdocs::contract`), port them line for
line from `src/sync/*`, and run `contracts/vectors/*.json` from a new
doctest binary (`pmdocs contract vectors` in `meson test`; the vectors
directory reaches the test through PMDOCS_CONTRACT_VECTORS_DIR).

Behaviour changes in pmdocs, all matching what the server accepts:

- `index` stripping is ASCII case-insensitive, so `Index.md` and
  `sub/Index.md` derive `/<base>` and `/<base>/sub` (X-8, CLI-3); core's
  `docs/Index.md` now plans to `/payload-markdown` like the server
- `order` follows ECMAScript Number(): empty is 0 (warning), 0b/0o/0x
  integers are accepted, `0x1p3` is rejected (CLI-2)
- `slug: ""` is ignored with a warning instead of failing (CLI-2)
- titles: `# C#`, `#<tab>Title`, setext headings, inline markup, fenced
  and indented code, non-ASCII filename capitalisation (generated JS
  toUpperCase table), lone-quote titles and lone CR now match the
  server (CLI-3, X-9)
- frontmatter BOM/comments/flow lists/nested keys/block scalars follow
  the shared subset (DOCS-13, CLI-20)
- unservable route segments, asset route confinement and `version: 1.0`
  manifests match the server (DOCS-22, DOCS-14)

Route oracle: pmdocs validate now derives the same route, title and
frontmatter as the server for all 112 valid corpus files (before: 1
difference, core `Index.md`).

Also drop the stale npm parity-harness instructions from the Debian and
Homebrew READMEs (CLI-16).
…-8 files

The docs and skills walks silently dropped content and could archive
previously synced docs on the next push (CLI-4):

- `build`, `dist` and `.next` are now skipped only directly below the
  walked root; `docs/guides/build/index.md` is published again
- `.git`/`node_modules` (any depth), root build directories, symlinks
  and Markdown-like files that are not lowercase `.md` (`README.MD`,
  `.markdown`, `.mdx`) are reported as `skipped_path` warnings
- hidden files stay included (unchanged behaviour) but each one is
  reported as `hidden_path`; the new `--skip-hidden` flag leaves them out
- skill files with unsupported extensions are reported instead of
  silently dropped

Files whose content or name is not valid UTF-8 now fail `validate`,
`manifest`, `plan` and `push` with an `invalid_encoding` error naming the
file, instead of passing validate and crashing later with a raw
nlohmann type_error (CLI-11). JSON output uses a replacing UTF-8 handler.

A trailing separator no longer breaks path arguments (CLI-12):
`validate ./mydocs/` derives the source id `mydocs`, an empty derived id
is an explicit error (it no longer walks every skills package), and
`install skill --out dir/`, `--out .` and
`install routes --payload-app "app/(payload)/"` pass the containment
check.

Warnings from the walk are printed with validate output and on stderr
for manifest, plan and push.
The server rejects sync request bodies above `maxBodyBytes` (5,000,000 by
default) with HTTP 413 before it validates the manifest, but pmdocs only
checked the sum of content bytes. JSON escaping, paths and hashes made a
4.6 MB docs set a 5.6 MB body that passed `validate` and failed `push`
(CLI-5).

`validate`, `manifest` and `plan` now check the largest body `push` could
send (mode "dry-run", deleteBehavior "archive", publish false) and `push`
checks the exact body, reporting `body_too_large`. The new
`--max-body-bytes` flag matches a server configured with a different
limit; `validate --json` reports `requestBodyBytes`. The shared
serialization vectors (contracts/vectors/limits.json) guarantee the
measured bytes equal JSON.stringify output on the server.
`pmdocs validate` passed manifests the server rejects with a 409 because
it never checked routes across files, and `push` printed only
`error.message`, dropping the server's `routeCollisions` and `conflicts`
(CLI-6, X-8):

- validate/manifest/plan now run the shared
  `find_route_collisions` contract: exact collisions (`index.md` +
  `Index.md`, `guide.md` + `guide/index.md`, slug clashes) are
  `route_collision` errors on every involved file, ASCII case-only
  collisions are `route_case_collision` warnings; push keeps them as
  warnings because only the server knows the real route base
- failed pushes print `error.issues` (path, message, code, severity),
  `routeCollisions` (route, reason, optional paths) and `conflicts`
  (sourcePath, reason, route), accepting each array under `error` or at
  the top level and tolerating its absence on older servers
- `push --json` adds a normalized `failure` object with the same data
`push` accepted `http://` endpoints on any host, so a GitHub OIDC bearer
token (not bound to the request body) or a signed manifest could cross
a network in clear text and be captured and replayed (CLI-9).

`push` now requires `https://` unless the host is loopback (`localhost`,
an IPv4 literal in 127.0.0.0/8, or `::1`); `127.example.com`-style DNS
names do not count. The new `--allow-insecure-http` flag opts in for a
trusted private network. The check runs before any token is requested
or any key is read.
`pmdocs keygen --out <dir>` wrote docs-sync-private.pem through a plain
ofstream, so the key was 0664 or 0644 depending on the umask (CLI-10).

The private key is now written to a mkstemp() temporary file (created
0600 with O_EXCL) and renamed over the target, so it is never visible
with a wider mode, including when `--force` replaces an existing 0644
key. A newly created `--out` directory is created 0700. `push` warns on
stderr when `--private-key-file` is accessible by group or other users
(warning only, so existing CI key files keep working). Doctest
regression tests run keygen under umask 0.
…assets

`pmdocs plan` hard-coded the route base `/<source>`, could not plan a
`--publish` push (every published record showed up as an update), and
never compared assets with existing records (CLI-13):

- `plan` and `validate` accept `--route-base` (grouped:
  `/<group>/<source>`, product-nested: `/<group>/<source>/docs`) and
  `--asset-route-base` (defaults to the route base)
- `plan --publish` plans published output like `push --publish`
- `plan --existing-assets <file>` compares assets (hash, route, content
  type, kind, archived) with the same rules as the server
- plan output reports the route base and publish state

The defaults are unchanged; `push --dry-run` stays the authoritative
plan, which the docs now say explicitly.
`pmdocs doctor` printed `status: ok` and exited 0 even when the bundled
skill data was missing, so `pmdocs install skill` failed for every
Homebrew user while doctor said all was well (CLI-14, CLI-21). Doctor
now prints `status: degraded` and exits 1 when the payload-markdown-docs
skill data is missing (with how to fix it), and notes when only the
companion payload-markdown skill is absent.

Help and reference fixes (CLI-21):
- `--source` defaults to the GITHUB_REPOSITORY repository name, else the
  docs root directory name (`local-docs` for `docs`), not "local-docs"
- `--max-files`, `--max-file-bytes` and `--max-total-bytes` say they
  cover docs and assets and give their defaults
- the CLI reference no longer lists `--strict-routes` as a common flag
  (it is push-only; validate rejects it with exit 109)
The formula built with Meson defaults (`install_skill_data=false`), so
`pmdocs install skill` failed for every brew user, and enabling the
option alone would have failed the build because the tag archive has no
node_modules for the companion payload-markdown skill (CLI-14).

- new `companion_skill_data` feature option: `enabled` (default, keeps
  the Debian build strict) fails without the companion skill, `auto`
  installs it only when node_modules provides it
- the formula builds with `-Dinstall_skill_data=true
  -Dcompanion_skill_data=auto` and `brew test` now runs `pmdocs doctor`
  and `pmdocs install skill --agent codex --dry-run`
- the release workflow's Homebrew parity job builds with the formula's
  options, stages `meson install` and runs the same smoke commands
- packaging contract tests cover the formula options and test block

Verified with a source build of the working tree without node_modules
(`meson install --destdir`): doctor reports `status: ok`, install skill
lists payload-markdown-docs files. The formula parses cleanly with the
ruby/prism parser (no Ruby interpreter on this host; CI runs `ruby -c`).
…ted runners

Harden the GitHub workflows (CLI-17):

- release.yml: the workflow default is `contents: read`; only
  `assemble-release-artifacts` (GitHub release upload) gets
  `contents: write`, and only `publish-npm` (npm trusted publishing) and
  `publish-docs` (pmdocs --github-oidc) get `id-token: write`. Jobs that
  run dependency code (npm-package, native-debian, ...) no longer get
  repository write access or OIDC minting rights. The org-level
  VL_DEEPSEEK_API_KEY / VL_AI_RELEASE_PROFILE wiring is unchanged.
- deploy.yml: read-only permissions, and `pull_request` jobs run on
  GitHub-hosted `ubuntu-latest`; the persistent self-hosted runners that
  also execute Production release jobs only build pushes to main.
- every action (workflows and the publish-docs example) is pinned to the
  commit SHA of the major tag in use, with the resolved version in a
  comment; no action is upgraded.
- the publish-docs example plans locally on pull requests and makes the
  server-side OIDC dry-run opt-in (`DOCS_SYNC_PR_DRY_RUN`) and same-repo
  only, because PR tokens carry `refs/pull/<n>/merge` and fork PRs get
  no token (CLI-7 client side); the CI guide explains why.

Workflow contract tests assert the permission scoping, SHA pins and PR
runner routing.
…ors and agents

Bring the CLI-facing docs and the bundled skill references in line with
the contract vectors:

- skill frontmatter references: accepted order forms, flow lists,
  one-line values, lowercase `.md` names and route-safe file names
- skill troubleshooting: in-manifest route collisions and skipped or
  hidden files
- manifest reference: the fixed asset content types pmdocs emits (CLI-22)
- cli/README and PHASE3 notes: where the shared vectors live and how
  `meson test` runs them
…se branch

Real pull_request tokens carry `ref: refs/pull/<n>/merge`. The ref allowlist
ran before the pull request rule, so `allowPullRequests: true` never took
effect and the documented PR dry-run job always failed with
`oidc_ref_not_allowed`; the unit fixture used `refs/heads/main` and hid it
(CLI-7, requested by the protocol workstream).

- pull request tokens (event_name or merge ref) are rejected unless the docs
  set allows pull requests; when allowed, `base_ref` must match the docs set
  branch
- fixture uses the real token shape
…collisions

Follow-up to the protocol workstream merge (requests from docs-c):

- asset content types are checked with the shared
  `isAllowedDocsAssetContentType` (pinned by contracts/vectors) instead of a
  parallel server list; served types are normalized to `charset=utf-8`
- in-manifest route duplicates come from `findManifestRouteCollisions`;
  `routeCollisions` entries carry `paths` and the message no longer claims an
  "existing route reservation"
- case-insensitive route collisions are returned as success warnings
- 413 responses carry `error.issues: [{ code: 'body_too_large', ... }]` with
  the measured size and the limit
Public surfaces now hide draft docs sets (DOCS-1), and the dev seed created
its docs set without `_status`, so it stayed a draft and the e2e llms.txt
checks for it returned 404. The seed publishes it; e2e passes again
(`CI= PORT=3922 ... playwright test`, on a fresh database).
…piles

`pnpm dev -- --port N` forwarded a literal `--`, so Next could ignore the
port and Playwright waited on the wrong one until the 60s webServer
timeout. Call `next dev` directly with --port (and PORT), allow 180s for
a cold compile plus the Payload schema push on an empty database, and
give each test 120s because the first request to each dynamic route
compiles it in dev mode (measured 15-20s cold, <0.4s warm).
The server worked around re-archiving by filtering the plan after the
fact (src/payload/planNormalization.ts), so `pmdocs plan` and the
server disagreed and the rule lived in two places. Make planDocsSync /
planDocsAssetsSync and the C++ planner skip records that are already
archived for archive/draft removal (hard delete still applies), and drop
the server-side filter.
payload-markdown-docs registers admin components from
@valkyrianlabs/payload-markdown and @payloadcms/plugin-seo in the host
app's import map, but depended on both as floating "latest" runtime
dependencies. Under pnpm a README install left those specifiers
unresolvable from the app root (X-2), and an app that pinned its own
payload-markdown got a second copy whose plugin settings docs pages never
saw, so themes/code/icon options silently disappeared on docs pages
(X-3). plugin-seo "latest" also pulled a second @payloadcms/ui (X-11).

- peerDependencies: @valkyrianlabs/payload-markdown ^1.5.0,
  @payloadcms/plugin-seo ^3.0.0, payload ^3.0.0 (was >=3.0.0, which
  admitted Payload 4), next ^15.2.0 || ^16.0.0, react/react-dom ^19
  (X-15); both packages stay devDependencies for local work
- README/installation: install the peers explicitly; document the
  Tailwind @source lines and theme tokens the components need (X-16)
- scripts/consumer-smoke: installs both packed tarballs into a fresh
  Next.js + Payload app the way the README says and checks one core copy,
  import-map resolution, `next build` without a media collection (X-17),
  a pmdocs-signed sync with a docs-set-scoped key, and rendering through
  the real core with the app's own payloadMarkdown() options (13 checks;
  before this change the copy and config-context checks failed)
…onsumer smoke

The audit's worst bugs were invisible to CI because it only ran the
build and mocked unit tests. Build and Test now also runs:

- lint and the dev-app typecheck
- the Postgres-backed suite (incl. dev/regressions.spec.ts) against a
  postgres:16 service on a random host port
- Playwright e2e against its own fresh Postgres service
- release tooling pytest + `python -m tools.release check`
- a consumer smoke test that packs this plugin and payload-markdown main,
  installs both into a fresh Next.js + Payload app per the README,
  builds it, syncs docs with the CI-built pmdocs and renders them through
  the real core renderer

Fix the native install smoke: `pmdocs doctor` now exits 1 when bundled
skill data is missing (CLI-14), and the job installed a build without
it. The smoke now installs a build with skill data, like the packages.

Pull-request jobs stay on GitHub-hosted runners; actions stay SHA-pinned.
…ed Postgres DB

Pins the exact output of root and per-docs-set llms.txt / llms-full.txt, sitemap
entries, nav and header nav items, route resolution with sidebars (both route modes,
nested and custom groups, drafts, archived docs, product-nested aliases), skill
endpoints, docs-set manager data, resolveDocsSetSkills, and sync dry-run / error
responses. Recorded before the read-side and sync-endpoint restructuring so those
refactors are checked for byte-identical behavior. DB-gated; the suite clears the
plugin collections first and normalizes record ids and timestamps.
Adds src/shared/records.ts and replaces the private copies in the payload
readers (docsSets, existingDocs, existingAssets, routeCollisions, routeClaims,
docsAccess, syncRuns, visibility), the route adapter (next/records and its
importers), the admin manager data, the assets/llms/sync endpoints, and
normalizeShared. Semantics are unchanged: monomorphic relationship ids only;
the polymorphic-aware helpers in normalizeShared and fields/docsReferences keep
their own unwrapping.
The group route-path walk was implemented six times (payload/docsSets,
next/route, next/links, next/sitemap, admin/docsSetManagerData,
utilities/normalizeShared x2) with two different cycle rules. They now all
use src/routing/docsSetRoutes.ts: resolveDocsGroupRoutePath (one walker,
pluggable lookup for id maps or populated relationships),
getDocsGroupRoutePath, indexDocsGroupsById, parseDocsSetRouteMode and
resolveDocsSetRoutes (slug + group + route mode -> productRoute/routeBase).

Output is unchanged for every well-formed group graph (characterization
snapshots and DB suite identical). Only malformed data differs, now
consistently across sync and read: a parent cycle stops at the first
revisited group everywhere (the sync-side lookup and the marketing-href
walk used to append the revisited slug again, so sync could validate
routes the site never served), and a whitespace-only slug no longer
counts as a group/docs-set slug in the read paths that previously
accepted it.
…te adapter

Docs sets were resolved twice: payload/docsSets (sync, llms, assets) and
next/records toResolvedDocsSet + route.ts withComputedDocsSetRoute (route
adapter, nav). Both projections now take the docs groups and get their routes
from routing/docsSetRoutes, so withComputedDocsSetRoute is gone and a group
always carries its walked route (toResolvedDocsGroup takes groupsById).

payload/docsSets: the four finders share one loader (findResolvedDocsSets) and
one public listing (findPublicDocsSets); the unused, visibility-less
findDocsSetByRouteBase is removed.

next/route: the four docs-set queries (by route base, by route prefix,
product-nested aliases, group index) and the per-call group loads go through a
per-request lookup that loads docs sets and groups once; findById keeps its
own query. The unreachable getRelatedDocsSet fallback is removed (it only ran
when the doc's docsSet relationship had no id, in which case it always
returned undefined).

Characterization snapshots and the DB suite are unchanged.
…ility

payload/visibility now owns the draft/archived rules for all readers:
- raw-record predicates (existing) plus getPayloadDraftStatus,
  isVisibleToReader (resolved records, includeDrafts-aware),
  getDocsRecordLifecycleStatus (admin labels), and fresh notArchivedWhere /
  publishedWhere query constraints;
- route adapter, sidebar and nav (next/records isVisibleDocsSet /
  isVisibleDocsRecord delegate to isVisibleToReader), sitemap (published
  docs-set filter, archived-asset filter and query), llms and asset/skill
  endpoints (archived filters and queries), normalizeSkills, the admin
  manager status, and the sync-side route checks (routeCollisions,
  routeClaims, existingDocs) no longer inline `sync.archived` / `_status`
  checks.

Behavior is unchanged (characterization snapshots and DB suite identical).
Root llms.txt and llms-full.txt issued two queries per published docs set
(docs and skill assets), e.g. 602 queries for 300 docs sets. Docs and skill
assets are now loaded with `in` constraints for batches of at most
LLMS_DOCS_SET_BATCH_SIZE (100) docs sets and partitioned per docs set in
memory with the same ownership rules (docs: docsSet relationship; skills:
docsSet relationship, sourceId or sync.sourceId). The root index renders
links only, so its docs query selects index fields and skips markdown
bodies. Per-docs-set llms files use the same loader with bodies.

Output is byte-identical: characterization snapshots unchanged, and the
sha256 of root llms.txt / llms-full.txt matches the previous build on
seeded databases with 301 and 1001 docs sets. Local timings (5 runs):
1001 sets 5.0-9.2s -> 1.4-2.8s; 301 sets 0.65-1.5s -> 0.35-0.63s.

The assets spec mock now understands `in`, and new tests cover batch
boundaries (205 docs sets), skill ownership across batches, and that docs
set llms-full.txt still loads bodies.
…int split

Snapshots the raw JSON text (key order included) and content type of a
dry-run plan, an in-manifest route collision, and an invalid source id,
recorded on the pre-split sync endpoint.
src/endpoints/sync.ts (2,000 lines) is now the public factory only;
createSyncEndpoint, CreateSyncEndpointOptions, DocsSyncEndpointErrorCode and
SyncErrorIssue keep their names and shapes. The pipeline lives in
src/endpoints/sync/:

- context.ts      CreateSyncEndpointOptions, SyncRequestContext (req, payload
                  typed once, options, clock, limits) and SyncContext (+ identity,
                  docs set source, validated manifest, policy)
- request.ts      method, bounded body read, JSON manifest, source id
- authenticate.ts Ed25519 and GitHub OIDC, nonce consumption -> SyncIdentity
- policy.ts       SyncPolicy (apply, publish, deleteBehavior, draftsEnabled,
                  writesMainForUpdates, deferAssetWrites, recordAudit) decided
                  once; credential scope, lifecycle and apply gates
- validate.ts     docs set resolution, manifest + asset content-type policy,
                  route collisions, route claims
- plan.ts         existing records, planners, asset deferral, warnings,
                  summary, manual-edit conflicts
- apply.ts        transactional apply (runInSyncTransaction)
- record.ts       sync-run audit start / success / failure
- revalidate.ts   Next cache revalidation
- respond.ts      error contract (error.issues, conflicts, routeCollisions),
                  SyncRequestError, apply-failure classification, the
                  no-raw-messages error boundary, success body
- handler.ts      the ordered pipeline

Stages reject by throwing SyncRequestError at the same point the old code
returned early, so status codes, bodies (raw JSON key order included, see
the characterization snapshots), audit records and DB writes are unchanged.
policy.spec.ts covers the SyncPolicy decision table and gates.
The four public asset routes (root llms, docs-set llms, skill files, skill
ZIPs) each repeated the same try/catch for a missing docs assets table.
assetsStorage.ts now owns the whole storage-unavailable contract: the
message, the classifier, the plain-text 500 response and
withDocsAssetsStorageGuard, which the asset routes use; the sync endpoint
answers the same condition with its JSON error from the same message and
classifier. Responses are unchanged.
…query

getPayloadMarkdownDocsSidebar loaded every doc of the docs set with its full
markdown body (up to 1,000 docs per page render). It now selects only the
fields the sidebar builder and visibility check read (_status, sync, route,
sourcePath, title, navTitle, order, overrides). Sidebar output is unchanged
(characterization snapshots for published and includeDrafts routes).
The marketing-block afterRead hydrator fetched referenced docs sets and
docs pages with findByID + overrideAccess and no visibility check. In
Payload 3 that returns never-published (draft-only) records from the
main table, so a public page with a docs CTA or hero referencing an
unreleased docs set or page rendered its title, description and link,
and archived docs pages were linked too: the DOCS-1 failure class on a
different surface.

Public reads now hydrate only publicly visible records (shared
visibility module); unpublished references stay unresolved, exactly as a
missing record renders. Draft-aware reads (?draft=true, or a draft host
document such as an admin preview) still hydrate drafts. Archived docs
pages are never hydrated.
… parentheses (X-5)

The docs link rewriter works line by line on raw Markdown before core
renders it. It toggled fence state on any ```/~~~ line, so a shorter
inner fence inside a longer one flipped the state and links in code
examples were rewritten; link syntax inside inline code spans was
rewritten too, mutating displayed code. Linked images
([![img](./i.png)](./page.md)) and destinations with parentheses kept
their raw .md href.

- fences follow CommonMark: opening run of 3+ backticks/tildes, closed
  only by the same character with an equal or longer run
- backtick code spans (N backticks closed by exactly N) are left as is
- the link label may contain one nested image, and destinations may
  contain one level of balanced parentheses

Output for ordinary links is unchanged (existing rewrite tests and the
read-side characterization snapshots pass untouched). Links inside
indented code blocks are still rewritten: telling them apart from
indented list content needs a Markdown parser; that is the AST URL hook
planned for the feature phase.
The Postgres-backed jobs read `job.services.postgres.ports` in job-level
`env`, where the `job` context is not available, so GitHub rejected the
whole workflow ("workflow file issue": a failed run with no jobs).
Export DATABASE_URL from a step instead, and run actionlint (pinned
1.7.7, self-hosted label declared in .github/actionlint.yaml) as the
first gate so invalid workflow files fail visibly in CI.
`playwright install --with-deps` runs `sudo -- sh -c "apt-get ..."`, which
the self-hosted runner's sudoers policy (apt-get only) does not allow.
Read the package list from `playwright install-deps --dry-run` and call
the allowed apt-get commands directly (same fix as payload-markdown).
Peer and dev dependency move to ^1.6.0, so docs installs get core's
1.6.0 security and correctness fixes (card href XSS, directive openers
that swallowed page content, structured icon sanitizing). The dev app's
import map gains core's new MarkdownBlockParamsEnableField.
…ndex

1.1 adds a unique (keyId, nonce) index to the sync nonces collection. Creating
it fails on installs that hold duplicate pairs, which earlier versions allowed.

The new ./migrations export removes expired nonces and keeps the oldest row of
each remaining duplicate pair through payload.db, so it works on every adapter
and runs inside the migration's transaction when given its req. Call it at the
top of the generated migration's up(), before the SQL.

Covered by a Postgres-backed regression test that drops the index, seeds
duplicates, runs the helper and recreates the index.
consumeNonce relied on the unique (keyId, nonce) index alone: a successful
insert meant a new nonce. When the index does not exist (a MongoDB background
index build that failed on old duplicates, or a migration that never ran),
every replay would be accepted.

After inserting, look for another live row for the same pair and treat it as a
replay. Concurrent duplicates then fail closed instead of all passing.
Covers the peer dependency install and import map, the schema changes, the
migrate:create + prepareDocsSyncMigration flow, dev schema push recovery,
MongoDB, changed defaults (admin-only access, draft-sync assets, forwarded
headers, tag refs, unscoped keys) and the CLI walk changes. Documents the
/migrations export in the public API reference.
The default source id prefers GITHUB_REPOSITORY, which GitHub Actions always
sets, so the test saw the repository name instead of the docs directory.
@cooperlarson cooperlarson self-assigned this Oct 2, 2026
@cooperlarson
cooperlarson marked this pull request as ready for review October 2, 2026 21:48
@cooperlarson
cooperlarson merged commit ff97dc9 into main Oct 2, 2026
6 checks passed
@cooperlarson
cooperlarson deleted the stabilization/2026-10 branch October 2, 2026 21:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant