@AGENTS.md
Everything above is imported from AGENTS.md: the start-here reading list, the twelve critical rules, the working conventions, verification and the canon ids. They live there once, tool-neutral, so no second copy can drift. Change a rule in AGENTS.md, never here. This file carries only what changes release to release.
v0.21.1 (a patch: test speed) is the current release (CHANGELOG.md has every release since 0.4.0;
each story's own changes are in games/<slug>/CHANGELOG.md).
Test times (v0.21.1, scripts/run_tests.py, measured 2026-10-08):
full (the hybrid) 16.7 min, exit 0, with 4 flaky tests passing their
solo re-run (parallel phase 4.7 min + serial phase 11.8 + re-run 0.2); full
serial (one process) 28m50s, green earlier; fast about 5.8 min. Not
re-measured at the release.
Windows: 5977 passed, 24 skipped in 54m48s (v0.21.0, measured
2026-10-07 at the release; one more test failed in that run,
tests/test_screenshot_runs.py::test_serve_refuses_a_port_that_already_answers,
a loopback stall of this machine that passes alone, and one test errored on
a port clash fixed before the release, tests/conftest.py::hold_guarded_ports;
in a checkout holding the gitignored
Design_files/, which runs the Design_files-only Garden test; a fresh
worktree skips it too), no expected failures. The 24 skips: tests/test_llm_live.py
(unless CLOCKWORK_LIVE_LLM names the configured model server); the ten
tests of tests/test_docker_smoke.py (opt-in, CLOCKWORK_DOCKER_SMOKE=1 and
a built image); the stamina soft-lock test's three stories (it covers only
stories with no rest verb); THE LONG CON's word-limit check;
tests/test_constraints.py's installed-version check for faster-whisper
and gunicorn (neither installed on Windows); six POSIX-only tests (file
modes in accounts, login and metrics, two exact-name command lookups, POSIX
process groups); and one bus test skipped by design (a reply to nothing
after hello). tests/test_simulate_thief.py alone takes about four
minutes. Plus 462
client tests under ui/tests/ (npm test --prefix ui; vitest is a
devDependency, so npm install --prefix ui once first). Re-measure and
restate these at every release rather than trusting this line -- it has
been stale before, in the very sentence that warned about it.
A bare pytest is the fast tier since v0.21.1 (slow and process left
out); the numbers above are --full.
Linux (measured mid-release, not re-run at the release: T20 started no
container): in python:3.11-slim-bookworm (pinned by digest, docs/HOSTING.md
§ Linux) on a clone in a container volume, installed with -c constraints.txt, 4641 passed, 8 skipped, 0 failed in 25m43s (T4,
2026-09-30; the skips the owner's set, the Design_files-only Garden test and
two Windows-only PATHEXT tests); and under gunicorn 23.0.0 the hosted
test files, 178 passed, 2 skipped (T13, 2026-10-05), every real front
door and worker started as -m gunicorn -c deploy/gunicorn.conf.py <wsgi>:app. The container's Node build matches the committed dist whole
(ui/index.html and its built copy are LF by .gitattributes). CI
(.github/workflows/ci.yml: suite, client and image on
ubuntu-latest) first ran on the v0.20.0 push (run 37420365350,
2026-10-06): client passed on Node 24 in 20 s; the suite took 40m01s,
5790 passed, 22 skipped, 1 failed (a drain-window timing margin in
tests/test_admin_model.py); image built but its container refused CI's
own random hex cookie key, fixed in v0.20.1. v0.20.1's run passed image
and client; its suite failed the same drain test on Linux alone, a test
that read gunicorn's refused respawn as the restarted worker (fixed in
v0.20.2, reproduced and re-run in python:3.11-slim-bookworm under gunicorn
23.0.0: tests/test_admin_model.py 32 passed). v0.20.2's run (37429512117,
2026-10-06) was fully green: image 33 s, client 19 s, the suite 40m22s.
The README carries the workflow's badge.
Docker (T18, 2026-10-06): docker build -t clockwork-dark . built a
327 MB image (110 MB of it the committed art) from the pinned base;
tests/test_docker_smoke.py (the real compose file, two stories, a stub
model server inside the container) 10 passed in 72 s: uid 10001 under
tini, a read-only root with no capability, no secret in the image, every
child under gunicorn, two accounts each playing a streamed turn through the
front door's WebSocket relay without reaching each other's run, the admin
panel seeing both, a killed gunicorn master's orphan reaped, and docker compose stop draining a turn in flight to its player and exiting 0.
host.docker.internal reached a server on the Windows host's loopback
(Docker Desktop, engine 29.8.1). No live model was called.
vLLM (T19, 2026-10-06): vllm/vllm-openai:v0.31.0 under Docker
Desktop on the RTX 2060 (sm_75: the V1 engine, TRITON_ATTN, --dtype half), Qwen/Qwen3-1.7B (Qwen3-4B's download declined),
--gpu-memory-utilization 0.90 measured from nvidia-smi.
tests/test_llm_live.py with CLOCKWORK_LIVE_LLM=vllm: 6 of 6 with
--reasoning-parser qwen3; without it 5 of 6, the HUE & CRY turn looping on
whitespace under the grammar to the cap once, then 2 of 2 on a re-run. The
vllm row is verified against vLLM 0.31.0 (bar mcp_integrations),
its fixtures recorded. Not run: the game container narrating through
http://vllm:8000/v1.
Hosted mode (v0.20.0; docs/HOSTING.md): hosting.enabled (off by
default, and then engine/hosting/ is never imported) turns the game into a
small group's server. Operator-made accounts (scripts/users.py; the first
admin with add <name> --admin), a signed-cookie login, one HTTP gate and
one socket wrapper (FlaskScene.on) on every route and event, every run and
save owned by an account (another's id answers as a missing one), generic
errors with a logged reference, and the §6.7 refusals (the studio,
llm.mcp.enabled, the Settings save). The supervisor (python -m engine.hosting.supervisor, or launcher.py with hosting on) runs a worker
process per hosting.stories slug and the front door, over a loopback bus
with a single-use token per child start: health checks, crash restarts with
backoff and a hold-down, drained start/stop/restart operations, per-process
logs, a bounded shutdown that stops the front door last. One model-server
queue for every story (engine/hosting/supervisor/queue.py): FIFO, a turn
admitted before any mechanic runs (busy = refused untouched), one live run
and one narration ticket per account, a wall-clock deadline per turn, a
reclaim backstop. The front door (engine/hosting/frontdoor/): one port,
one login, the story picker, HTTP proxied and the game's WebSocket relayed
to the chosen worker, the cookie's one writer, forwarded headers trusted by
the proxy token alone, long holds and turn slots capped so logins always
have threads. gunicorn on POSIX (deploy/gunicorn.conf.py, one
gthread worker), the development server elsewhere. The admin panel
(/admin, engine/hosting/admin/): Users, Sessions, Saves, Stories, Model
server (a drained, rolled-back apply to the admin layer), Queue, Metrics,
Errors and Audit, behind a role, a re-auth and CSRF; metadata only, never
play text (tests/test_admin_no_play_text.py); every change written first
to the audit log (engine/hosting/audit.py). Metrics in a closed
schema (engine/hosting/metrics_schema.py), kept by the supervisor in
SQLite. Docker: Dockerfile, config/docker.yaml, docker-compose.yml
(published on the host's loopback, a vllm profile).
Six stories ship, and every one can be played to an ending
(tests/test_finales.py): clockwork-dark (the flagship), wicked-garden
(the deck exemplar), neon-city (NEON CITY: THE CROSSING), the-long-con (THE
LONG CON, the first graph/deck hybrid), dev-story (the annotated bench) and
hue-and-cry (HUE & CRY, now with jobs — burgle a house, stage by stage,
with flashbacks and a guild contract, on top of the Lantern Watch — and
agendas: a seed-chosen thief robs the city by night under the Magpie's name,
Captain Ardane hunts, Silas Crook works the guild, all of it authored and
deterministic, never a model plan — and, since v0.14, a living city: beds and
bread, scrounging, honest work, night streets, seven factions and the city's
lore, with its three secret places findable — and, since v0.15, a guild
economy: the craft verb at the Porters' Hall bench, the Magpie's Hoard,
four blackmail squeezes and the fences' credit — and, since v0.16, Acts I
and II: the barge opening's three choices made real, the Honest Company's
initiation deck, the Lantern House's interrogation deck on every arrest, the
Magpie's trail laid through the houses, and the front desk's alibi and
accusation — and, since v0.17, Act III: the Hanging Fair on Gallows Green
(days 10–12), the gallows for a thief held since before it, the jailbreak,
and death.yaml (hp 0 respawns in the Snuffs; dying held at the fair is The
Rope), with all eight endings, each through its own door and each earned:
Cleared, A Lantern, Honest After All, Partners, The Legend, Guildmaster, The
Dapper's City and The Rope, the fail-forward — and, since v0.18, burglary
pays (the owner's decision: the fences pay about half a hot haul's value)
and scripts/simulate.py --game hue-and-cry runs its thief policies.
HUE & CRY is finishable; its screens, art and live play are still to come). Pick one with
launcher.py --game <slug>.
Since v0.19.0 the engine is model-server agnostic (engine/llm/, the
llm: config block): LM Studio stays the default, and llama.cpp's
llama-server, Ollama, vLLM and any other OpenAI-compatible server can
narrate every story, each through its row of engine/llm/providers.py.
Which of those facts were verified live, and how to run each server, is
docs/MODEL_SERVERS.md.
HUE & CRY, approved 2026-09-23, re-cut per feature after v0.9.0 shipped so nothing sits unpushed for weeks; re-cut again (owner, 2026-09-25) once the four engine features shipped, so v1.0.0 no longer waits at the end of the roadmap as one release -- it ships as its own run of point releases, tagged only once the last of them lands:
| Release | What | State |
|---|---|---|
| v0.8 | The audit release: presence in every story, outcome-aware evaluator, the "world moved" journal, leaks, dead code | shipped |
| v0.9.0 | Premises, plus the HUE & CRY skeleton | shipped |
| v0.10.0 | The Law | shipped |
| v0.11.0 | Jobs & flashbacks | shipped |
| v0.12.0 | Agendas | shipped |
| v0.13.0 | Engine seams for HUE & CRY's finish: secret places, custody + jailbreak, forced/repeatable decks, a terminal death, the clarity word, generate_art --game |
shipped |
| v0.14.0 | Living city: survival, forage + Rooftop Road's hidden paths, labour, boons, night encounters, factions, city lore | shipped |
| v0.15.0 | Guild economy: crafting, the Magpie's Hoard, blackmail and fence-credit threads, Brask's gate | shipped |
| v0.16.0 | Acts I–II: arcs, the opening, initiation deck, interrogation deck, the Magpie's trail, the reveal and the alibi beat | shipped |
| v0.17.0 | Act III + eight endings: the Hanging Fair event and fair-day deck, the jailbreak, The Rope via death.yaml, per-ending tests |
shipped |
| v0.18.0 | simulate.py's thief policy, agenda collisions measured, welshing's cost for a fencing burglar and the careful pickpocket's deaths measured; the fences made to pay (owner decision) and the lockpick money loop closed |
shipped |
| v0.19.0 | Model-server agnostic: LM Studio plus vLLM, the llama.cpp server, Ollama and other OpenAI-compatible backends | shipped |
| v0.20.0 | Linux as a first-class platform, and a hosted/web-served mode: auth, per-user sessions and saves, a production server, Docker -- with the supervisor and front door, the admin panel and its audit log, metrics, and vLLM run live | shipped |
| v0.21.0 | UI/UX overhaul, together with HUE & CRY's screens: the wanted poster, job panel and casing board as generic engine panels, portraits | shipped |
| v0.21.1 | Patch: test speed -- tiers, pytest-xdist hybrid runs, bounded runs (scripts/run_tests.py), libyaml, a shared HTTP client |
shipped |
| v0.22.0 | A new story: a dating simulation played through a phone of apps (dating apps, texts, instant messages, voice and video messages, two-player games), a populated cast the engine runs, no endgame (owner's brief: docs/superpowers/briefs/2026-09-30-dating-sim-brief.md) | next |
| v0.23.0 | The Clockwork Dark overhaul | queued |
| v0.24.0 | The Wicked Garden overhaul | queued |
| v0.25.0 | NEON CITY overhaul | queued |
| v0.26.0 | THE LONG CON overhaul | queued |
| v1.0.0 | Every story finished, each with its art and live play -- hue-and-cry's being a thief mistaken for "the Magpie" in the candle-port of Tallowmere, eight endings, a ~55-plate Grok art pack -- tagged only once this lands. dev-story is the engine's test bench, not a story, and is not overhauled |
queued |
Re-cut once more by the owner on 2026-09-26: the platform releases (v0.19.0 backends, v0.20.0 Linux and hosting) land before v1.0.0, and the UI/UX overhaul merges with what was HUE & CRY's own UI-plugin release into v0.21.0, so the shared surfaces are built once, as engine panels.
Re-cut again by the owner on 2026-09-26 (during v0.16.0): after the UI/UX overhaul, each of the other five stories gets the full HUE & CRY treatment, one a release (now v0.23.0–v0.26.0, Dev Story excepted): a design spec, its story and characters, the engine systems apt to it, every ending reachable and tested, measured balance, reviews, UI screens and art. A large story may take two minors, which shifts the later numbers. v1.0.0 now means all six stories finished, with art and live play for each. The owner does not approve each overhaul's design: write the spec, have it reviewed (opus), build it, and keep going.
Re-cut again by the owner on 2026-09-30 (during v0.20.0): v0.22.0 is a new story, a dating simulation, and the overhauls move up one, to v0.23.0-v0.26.0. Dev Story leaves the overhaul list: it was only ever a testing ground, not a story. The new story's brief is the owner's own, kept verbatim in docs/superpowers/briefs/2026-09-30-dating-sim-brief.md; its design follows the same rule (spec, opus review, build).
The README is kept current at every release through v1.0.0 (owner instruction, 2026-09-26): status, features and roadmap each release; new screenshots at v0.21.0's release (its T16); the backends (v0.19.0) and hosting (v0.20.0) documented when they land.
Every change updates the docs it makes stale: the story's CHANGELOG/README, the root CHANGELOG/README, this file and AGENTS.md (AGENTS.md "Docs move with the change").
Spec: docs/superpowers/specs/2026-09-23-hue-and-cry-design.md.
Plans live in docs/superpowers/plans/; each release is executed
subagent-driven, one fresh subagent per task with review between (v0.14.0's
last three tasks and its release were finished inline, at the owner's word).
Recorded rather than fixed, so nobody mistakes them for forgotten work:
-
neon-city ships zero art plates against 75 subjects, its entry location included.
-
The Wicked Garden's deck walker never reaches 7 of its 23 endings in 1000 runs (E2b, E2c, E3a, E3b, E3c, E4d, E5a; 9 in 200), and deals every card (
scripts/simulate_decks.py --game wicked-garden, measured 2026-09-26; this line said "11 unreachable, 4 orphan cards" before). -
The Garden's bargains work only as written:
bargainstrikes,dischargesettles, falling due breaks. Its renegotiations, the knife's cut, the gift auto-thread (accept_gift) and two undeclared templates (briar_witness,three_nights_or_truths) are NOT WIRED (docs/GOVERNANCE.md); the v0.24.0 Wicked Garden overhaul takes them up. -
mortal_threshold, the Garden's entry location, has no plate on purpose (it hosts the ten-card prologue), so a new player sees no scene art until the prologue ends. -
The studio review queue can keep one draft but does not draft from the browser.
-
HUE & CRY's Temple of the Everflame is moved since v0.17 T6 (the heart taken off the palace steps, -10) and READ by nothing yet but the save and the
reputationpredicate (data/world/factions.yaml's header): no ending, price or card asks the Temple's opinion. -
A generated premise's secret, carried out of a job, is HELD (the
secret_held:<premise>:<secret>flag and asecretledger fact) and opens no thread: only the four anchors' secrets name a blackmail (thread:). Nothing reads a generated secret: none of v0.16's decks does, and v0.17's endings read only an anchor's (Guildmaster, Mother Gannet's IOUs). -
survival.sleep_untillooks only for a rest entry namedsleep_bed, so in HUE & CRY (whose beds have their own names) it always sleeps rough. -
A HUE & CRY save from before v0.14 keeps the world it was generated with, so it never gains the rooftop and grating hidden paths; only the arrest reveal of the Undercroft reaches it.
-
Likewise a HUE & CRY save from before v0.16 has no Magpie's trail: its premises carry no
clue, so casing never hints and no job finds one (engine/world/clues.py::row_for; loads and plays, asserted). -
The Magpie's trail reads slowly, measured and left (v0.16 T8,
scripts/simulate_acts.py): a house gives its clue up only on the LAST watch, so a clue costs ~3 houses cased to the end and a deliberate investigator carries out ~2.5 by day 12. 40% of such runs have not unmasked the Magpie by day 12, and a retry after a wrong naming almost never lands in time (1 run in 40). The evidence bar was set to the lead -- two clues that agree (agree, not necessarily true: two herrings can open a wrong naming) -- not to a count. v0.17's endings harness kept it against the fair (scripts/simulate_endings.py: Cleared open for 68-73% of investigators, locked on 27-29 of 40); the levers, should the reveal need to come sooner, are the hint's place in casing (premises._ordered_ids) and the trail's density (clues.yamltrail/herrings). -
Hired hands (spec §4): explicitly optional there and not built. A job is walked solo, start to getaway.
-
An engine quirk a job's own measurement ran into and left alone (
jobs.yaml's header, AUTHORING §3.12): the watch's delay counts whole in-game hours (jobs.now_hourfloors), so a stage of fractional hours can bring it up to an hour late. -
Some agenda clock beats set flags no scene or ending reads yet:
ardane_warrant_sworn,ardane_doubles_the_watch,magpie_spree_fullandmagpie_emboldened(clocks.yaml).silas_splits_the_companyis read since v0.17 T6 (The Dapper's City, Guildmaster, the fair's F4). (The reveal is no longer on this list: since v0.16 the Lantern House front desk's accusation setsmagpie_unmasked; catching the thief in the act is not built.) -
Ardane's
takes_a_statementmove can only file her OWN report at a fixed deed (pickpocket): a move has no way to name the deed a witness actually saw, so it cannot upgrade that row directly -- it adds a second, lesser one instead. -
Agenda moves are never posted to the notice board, and no
fence {most: hot_goods}selector exists (fences hold no stock to count) -- both rows in docs/GOVERNANCE.md's NOT WIRED table. -
The Magpie's Hoard's
magpies_hoard_completeis read since v0.17 T6 only by The Legend's closenessscoreand its Seal beat's text (data/rules/endings.yaml): it opens no door and gates no ending, so a thief who carries the Hoard but never lifts the heart gets the Company's +8 and the reward line, and nothing else. -
The fences' credit (v0.15,
pell_advance,marrow_slate) is repaid in coin, not in goods: nothing in the condition grammar can say "carrying hot goods worth V" and no effect can hand over unnamed goods, so the debt is counted in crowns (threads.yaml's header). Selling the fence the goods is how a thief raises it. -
A craft roll's
crit_failurepays the recipe's full output (engine/skills/builtin/mechanics.py::_craft_yieldreads onlyfailureas a failed batch; pre-existing, found in v0.15's review). Latent: no shipped story'sskills.yamldegree table has acrit_failurerow, so a story that adds one would pay a fumbled batch in full. -
A recipe that goes illegal between the menu and its execution (the station left, an input spent) gets the dispatcher's generic "not a legal craft target" refusal (
engine/agents/tool_dispatcher.py), which lists raw recipe ids, rather than_craft_refusal's own reason. Still a refusal that reaches the prose (rule 1), only a less specific one. -
A
spoilers.yamlrow'slocation:naming a place that is not secret (no hidden path, known from the start) lifts the row on turn one, so it masks nothing;check_spoilers(engine/games/validation.py) refuses an unknown place but gives no warning for a known, non-secret one. -
Every intent verb shows at most eight options (
intents._MAX_OPTIONS).buycuts stock in id order andsellin inventory order, so a counter with more than eight rows, or a pack with more than eight saleable things, hides the rest. HUE & CRY's counters are held to eight bytest_every_counter_offers_all_of_its_stock; nothing guards other stories, the validator gives no advisory, and a thief carrying bench makings can crowd loot out of Marrow'sselllist. -
Some of the reveal's and the alibi's flags still wait for endings:
wrongly_accused_<suspect>andalibi_provenare read by nothing but the desk, interrogation and (since v0.17 T6) confrontation decks' own gates. (Since v0.17 T5 Cleared and A Lantern readmagpie_unmasked, and Clearedmagpie_named_wrongly.) The interrogation's reservedQ3_the_evidenceslot was released unfilled (the accusation lives at the front desk). -
The Lantern House front desk (
lantern_house_desk.yaml) is a repeatable deck, re-armed only when itswhen:is seen false: a card that becomes eligible while the thief is already standing in the Lantern House (the captain coming on duty, say, or A Lantern's badge right after a right naming at the same desk) waits until they walk out and back in (director.rearm). -
HUE & CRY's Act III arc, D4 (Cleared after the fair, v0.17 T5), Guildmaster's
the_fair_has_comeand Partners'the_last_job(v0.17 T8 fix round 1) readevent_seen: hanging_fair, which the quests pass records only at a turn that ends while the fair is on. A single action that began before day 10 and ended after day 12 would miss it; none exists today (the longest, a served sentence, is stopped by the gallows at nine on day 12). -
clues_favour {excluding: [...]}accepts ids that are not candidates of the role and ignores them, andclues.yaml'sfresh_flagaccepts any flag name but aclue_found:one (engine/world/clues.py); neither is cross-checked against the story. -
The casing board's
ofcount (premisescasing receipt,"of") is one higher on a house holding a clue, as it is on one holding a secret: it shows that a house holds something more, never whose (accepted in v0.16 T6's review). -
A bed's
requires/cost/fallback/refusals(asurvival.yamlrest entry) and an arc'snarrate:are read at run time and checked by neithervalidate_content.pynordoctor.py: a misspelt fallback or a malformed refusal row is skipped, andnarrateis read truthily. -
Every story's opening payload now carries
frame: "opening"(default_state.opening, the authored-choice gate): its turns are unchanged, but the opening payload is not byte-identical to v0.15's. -
scripts/simulate_decks.pybuilds a bareGameState, so its walks start on the flagship's phantomquiet_lifearc, not the story's default arcs (quests.seed_default_arcsruns only inprocgen.new_game_state). Judged harmless in v0.16 T1's review. -
HUE & CRY's careful pickpocket still starves, by the owner's v0.14 decision (deliberate pressure, not tuned). Measured in v0.18 T3 at 40 seeds x 14 days (
simulate.py --game hue-and-cry --policy living): it dies 2.90 times a run, and every death is hunger -- none in the street, the cells or at the fair -- at 05:00 in bed or at 21:00 waiting for the night's purse, first on day 6.2, never before day 5. It keeps 4.33 of 14 days. The fences' new pay (below) barely reaches it: its marks' goods are cheap, and it banks the extra coin rather than eating it. A careful thief who takes a porter's shift when hungry (careful_porter) keeps 11.9 (the honest porter 13.1) and dies once in 40 runs, so a living exists for one who adapts. Measured and kept. -
Welshing on a fence's credit still nets a purses-only pickpocket kept days: +3.1 on Pell's line and +1.2 on Marrow's over 10 days, unchanged by the fences' new pay. The owner accepted that in v0.15 because the cost falls on a burglar. v0.18 T3 measured it there, after the fences were made to pay (v0.18 T3 fix round 1,
data/tables/trade.yaml: a hot haul now fetches about 0.53 of its value at Pell's and 0.63 at Marrow's, not 0.25 and 0.28). Over 14 days the fencing burglar who welshes:- keeps fewer days (-0.62, -0.60);
- dies 0.7 more a run;
- ends 9.5-11 crowns poorer;
- fences about 30 crowns less, because neither fence buys again.
Over 10 days the advance still buys it +0.4-0.5 kept days. The credit is unchanged (CHANGELOG [0.18.0]).
-
The fencing burglar still dies 0.7 times a run, all hunger, while keeping 7.25 of 14 days (the careful pickpocket 4.33, the porter 13.05) and ending with 13 crowns on average. Its death log (gold and place, v0.18 T3 fix round 2) says why: 17 of its 28 deaths over 40 runs came with no coin in hand, and 17 came at 04:00-06:00 in its Snuffs bed before the quay's breakfast. The 11 with coin held 1-12 crowns, 9 of them in that same pre-dawn bed, when its two carried meals (
STOCK, every labour policy's rule) were gone and no counter was open. So it mostly starves on the lean nights between hauls. Recorded, no policy changed. -
A fence pays slightly more than an honest counter for a CLEAN thing: Pell 0.75 and Marrow 0.7 of value, against Dock Mag's and the city's 0.5. The engine has one sell spread per vendor and no separate clean rate, and v0.18 T3 raised the fences' spreads so that stolen goods pay (
data/tables/trade.yaml). It is at most a crown more on a scrounged find, and no policy exploits it. -
A death in the cells during HUE & CRY's Hanging Fair is The Rope, whatever took hp to 0. The hanging itself has a scene since v0.17 T3 (the gallows deck,
data/scenes/the_gallows.yaml, dealt at nine on the fair's last morning to a thief held since before it), but hunger inside a stretch of hours still gives that turn's prose no death receipt (a card beat's death does get one, v0.17 T1), so a thief who starves in the cell goes straight to the ending module's beats (AUTHORING §3.5). Narrowed, not closed: a prisoner serving a sentence is fed, so only one who sat unfed in the cell reaches it; flagged indeath.yaml's header. -
The bounders still CLAMP rather than reject (right for a model's output mid-turn): text past its cap and an outcome's effects past four are cut at load. Authored text has its own cap (
spec.MAX_AUTHORED_TEXT, 1500, v0.17 T8 fix round 3; model text keepsMAX_TEXT, 400). Every cut is recorded in the bounder's adjustments andvalidation.check_truncated_contentreports it as an ERROR for decks (card title and text, beats, outcomes), ending modules, thread templates, opening choices and set-pieces. Not covered: quest, encounter, epilogue and lore text, which pass through no bounder. -
A thief who names the Magpie to the captain (
magpie_unmasked) is never dealt HUE & CRY'sF4_silas_makes_his_move(it is The Dapper's City's door, and that ending shuts on the naming), so cannot stand against Silas at the fair either: if Silas's rise has won, Guildmaster stays shut for that thief (v0.17 T6,data/rules/endings.yaml). -
A locked run is dealt no door -- but only POOL doors (
deck.eligible_cards, v0.17 T6): a REQUIRED card that locks an ending is still dealt, because the Wicked Garden's recorded walk re-deals its required finale lock after the run is locked. HUE & CRY'sthe_gallowsG1_the_last_morningis required and locksthe_rope, so it could still be dealt to a locked run (its lock refused). And a deck whosewhen:reads an ending's eligibility (HUE & CRY'sporters_hall, the desk's badge branch) still comes due after the lock and deals nothing (spent, "no cards were eligible"). Both reachable only by an API or harness caller that plays on past an ending. -
validation.check_law_effectschecks areport's deed, guise and jurisdiction, awantedcondition and acommitted_deedonly against a law file: in a story with nopaths.lawit checks none of them, though at runtime such a report is refused ("no watch to report to") and such a condition is false. Only the engine-only effects are reported there. -
The Wicked Garden deals
day_09_finaletwice (pre-existing). -
Survival's hunger/death is not cut-invariant (pre-existing).
-
Two tabs can drive one local session (spec finding 7, recorded, not fixed in local mode):
session_idis persisted in the save, so a secondresumeof the same save rebuilds a session under the SAME id and replaces the_sessionsentry (engine/session/store.py::_build); a turn already in flight on the old engine can then autosave over the new session's state, and both tabs sit in the room and drive the new engine. Returning the live session instead would change the frame a reconnecting local player gets. Hosted mode closes it per account: one live run per account, the other released holding its turn lock so the old engine never autosaves over a resume (SessionStore, spec §5.4). -
A Settings panel save rewrites
config/local.yamlwhole throughyaml.safe_dumpand drops its comments (pre-existing;engine/api/settings.py::apply_settings). Keeping them needs a round-trip YAML parser. Since v0.19.0 a file that does not parse is refused rather than overwritten. -
Provider cells verified live: LM Studio's (the golden, recorded on v0.18.0, and
tests/test_llm_live.pyre-run against it at the v0.19.0 release onnvidia/nemotron-3-nano-4b), llama-server's (v0.19.0 T8,llama.cpp server b7966, the Windows Vulkan build) and Ollama's (T8,Ollama 0.34.4, the portable Windows build), with their fixturesrecorded-- barmcp_integrationson both and Ollama's proxy pass-throughauth, which no run had. Ollama's authored three-model discovery server, aformat-ignoring probe answer, a proxy's 401 and a leading inline<think>stayauthored: no live Ollama could give them (PROVENANCE.yamlnotes say why). -
vLLM was verified live in v0.20.0 T19 on ONE card and ONE small model: vLLM 0.31.0's Docker image on an RTX 2060 (sm_75) with Qwen/Qwen3-1.7B. Not run: Qwen3-4B (declined), a newer GPU, a bare-metal Linux
pipinstall,mcp_integrations, and the game container narrating through the compose network (http://vllm:8000/v1). Three fixtures stayauthored, each with aPROVENANCE.yamlnote (a 200 error body on the list, the[IMAGE:]-in-thinking stream). A generic OpenAI-compatible server has no one server to verify against. -
vLLM's structured outputs allow any whitespace between JSON tokens (
disable_any_whitespace=False, its default): once in six live narration turns Qwen3-1.7B wrote its narration and then newlines to the 4400-token cap (T19, HUE & CRY, no--reasoning-parser); the engine salvaged the narration and offered the fallback choices. Recorded, not tuned: the lever is the server's (--structured-outputs-config '{"disable_any_whitespace": true}') or a larger model (docs/MODEL_SERVERS.md § vLLM). -
Ollama's thinking models think BEFORE its
formatgrammar binds (measured, T8):qwen3:4bspent 4881 tokens on the probe's one-line question and over 16,000 characters on a narration turn, starving thebigprofile's 4400-token cap; the turn is recovered bythink: falseon the sameformat, which binds at once, so each such turn pays one wasted think. llama-server binds the grammar from the first token instead. Recorded, not tuned:llm.profiles.big.reasoning: "off"is the owner's lever (docs/MODEL_SERVERS.md § Ollama). -
backend._retry_with_roomstands down on servers that report no reasoning-token count: llama-server (usagehas nocompletion_tokens_details, b7966) and Ollama (eval_countis thinking and answer together, 0.34.4), measured in T8. A request that starved with the reasoning-off patch on stops after one request there: measured room needs a count, and the patch is the whole net. -
llama-server --reasoning-format nonewith a model whose template opens<think>in the prompt (Qwen3-4B-Thinking-2507) sends the thinking with no opening tag. A whole answer is split at the orphan</think>; a STREAM is not (holding it would show the player nothing until it ended), unless it carried an untrusted patch, so its thinking reaches the player. Documented as a flag not to run (docs/MODEL_SERVERS.md § Inline<think>); the default reasoning format splits it server-side. -
llm.mcp(Phase A) is LM Studio-only: on any other providerllm.mcp.enabledturns it OFF, with one ERROR and a doctor FAIL row (engine/agents/mechanics.py::mechanics_enabled). No engine-side tool loop exists (a GOVERNANCE NOT WIRED row), nor more than one model server per process (another). -
The golden's legacy scenario 23 differs from its v0.18 recording by one URL, sanctioned (v0.19.0 T6): v0.18 health-checked a fixed
localhost:1234whateverlmstudio.base_urlsaid, and the probe is now derived fromllm.base_url(tests/test_llm_golden_lmstudio.py,SANCTIONED_URL). The shipped variant is byte-identical. -
Ollama's gpt-oss-style models (
think: "low" | "medium" | "high", ignoringtrue/false) are handled only when declared (llm.declared_models.<id>.reasoning, docs/MODEL_SERVERS.md § Ollama). Not live-verified: the owner declined the 14 GBgpt-oss:20bdownload in v0.19.0 T8, so the effort strings andreasoning_off_trusted: falsestay as Ollama's docs describe them. -
An UNTRUSTED reasoning-off patch keeps the full cap (spec §5.2): on vLLM, llama-server and (since T8) Ollama, an
offrequest for a model with nodeclared_modelsreasoninglistingoffsends the patch (enable_thinking: false,think: false) but still pays the reasoning budget, so turns are slower than they need be on a template that honours it, until the owner declares it. Chosen over starving every turn on one that does not. Measured on both servers (T8): a template that ignores it (Qwen3-4B-Thinking-2507, Ollama's ownqwen3:4b) thinks anyway and the thinking arrives inside the answer, closed by an orphan</think>; the client moves it to the reasoning channel (client.InlineThinkSplitter(forced_open=True), both transports), reads an unclosed answer cut at the cap as starved (unless the request carried a grammar and it starts as JSON: with no grammar a cut answer, prose or JSON, from a model that honoured the patch is lost, the accepted trade-off, T9), and warns once per model. Such a stream is held until it is decided only when the request carried no grammar (fix round 1): an ungrammaredoffstream to an undeclared model -- rung 3 narration underreasoning: off-- reaches the player all at once, not as it streams. A verifiedreasoning_offcell, or Ollama'sthinkingcapability, no longer makes the patch trusted. Declaringreasoning: ["on"]for such a model is worse, not better (measured on Ollama: the probe starves, rung 3), and the docs say so. -
LM Studio gains the format block only under
structured_output: off: itsjson_objectrung and the turn after a failedautoprobe carry no shape in the prompt, because the golden pins both requests to v0.18's bytes (scenarios 05 and 08;backend.format_block_due). Every other provider gets the block on rungs 2 and 3. The reply is conformed on every rung, LM Studio's included. -
A choice whose TEXT promises an action but whose reply carries no
intentsurvives on every rung, rung 1 included:intentis optional in the turn schema, so "Follow the smoke toward Edgewood" can be offered with no mechanic behind it, and picking it moves nobody and refuses nothing -- the prose may then narrate a walk that did not happen. Pre-existing (v0.3.0's intent design), not introduced by v0.19.0'sconform, which can only judge an intent that is there. A GOVERNANCE NOT WIRED row names the gap. -
llama-server's multi-model router mode (per-model
statusin/v1/models,/props?model=) is not supported: discovery treats every listed model as loaded and sizes them all from one/props(engine/llm/discovery.py; docs/MODEL_SERVERS.md). One server per model. -
LM Studio passes inline
<think>through untouched when its reasoning split (the Developer setting that sendsreasoning_content) is off, so an ungrammared compat reply (a tools request, or the native route unavailable) can carry the model's thinking -- and any[IMAGE:]tag written in it -- into the prose. Pre-existing since v0.18. Thelmstudiorow'sinline_thinkstayspass(engine/llm/providers.py) because stripping would change LM Studio's parsed responses, which the golden pins; the owner's lever is LM Studio's reasoning-split setting (docs/MODEL_SERVERS.md § Inline<think>). -
Local mode's Socket.IO still RECORDS
cors_allowed_origins="*"(engine/scenes/flask_scene.py), but no longer honours it for another site: the guard in front of it (engine/scenes/host_guard.py, v0.20.0) refuses a DNS-rebinding page's Host, and any request on the Socket.IO path or WebSocket upgrade, whatever its method, whoseOrigin(or, with noOrigin, itsReferer, as JSONP polling sends) names another host -- the polling handshake, its POSTs and the upgrade, on the path read off the constructed server -- as well as any other state-changing request whoseOriginorRefererdoes. So cross-site WebSocket hijacking by a page the local player visits is refused, and the loopback bind keeps everyone else out. What is left is only the recorded"*"value itself: setting it to same-origin changes the local-mode golden's recordedcors_allowed_origins, so it waits for a sanctioned change (hosted mode sets its own: same-origin, or[hosting.public_origin]). The guard compares hosts, not ports, so a page on loopback at another port counts as the same site (documented in the module). -
CI runs the suite as ONE job,
run_tests.py full --workers 2(the hybrid, every tier; unmeasured on the 2-vCPU runner until it first runs; budget 150 minutes). -
Locally the suite is xdist-safe since v0.21.1:
run_tests.py fullandfastare hybrids -- every test but theprocess,mcp_serverandloopbackones on 6 workers (--dist loadgroup), then those serially, by design (the expressions aretests/tiers.py's;fastlimits both phases to its tier).loopback(an in-process server bound on loopback) is enforced bytests/tier_plugin.py's bind recorder. Under xdist the controller runs no test and is not sandboxed (it drops the marker before the workers start), holds the guarded model ports and compares the owner's storage once at the end; each worker is a whole suite with its own marker, basetemppopen-gwN, storage root, audit hook and nested Job Object. Left open:- a storage change seen by the controller names only SUSPECTS (the tests running when each changed file's mtime fell), not the writer;
- this workstation's loopback stalls now fail tests in the SERIAL phase,
a single process with nothing beside it. Measured 2026-10-08 (fix
round 2, final marks): the parallel phase green 4 of 4 (
full4.7 min twice,fast2.1 min twice); the serial phase green 1 of 4 (full13.8 and 14.2 min with 10 and 3 failures,fast3.7 min green then 10.7 min with 6). Wholefull18.5-18.9 min,fast5.8-12.8 min, against 28.9 serially. Every failure passed when its files were re-run alone (262 of 262), but fortest_hosting_secrets.py's front-door crawl, which failed 1 of 3 solo runs the same way (the bus's story table "timeout"). The fully parallel runs before the split lost 5-30 hosted tests each; the stalls are the machine's, not xdist's; test_survival.py::test_travel_alone_drains_stamina_to_zerofailed once in a parallel phase (fix round 1: its first leg refused, 0 legs, not 5) and passed alone, after every later file, and with its own file reversed; not seen again in six later runs. Not root-caused.
-
Change-based test selection (pytest-testmon) was not built (v0.21.1, an owner scope cut);
run_tests.py files <paths>runs the affected files. -
The front door's WebSocket relay:
_Link.abortdoes not wake a blocked send on Windows, so a stuck client makes the front door wait twiceRELAY_JOIN_SECONDSbefore it lets go (slow hosted shutdown of a stuck client; found in v0.21.1). -
The CI workflow (
.github/workflows/ci.yml) is held to its shape bytests/test_ci_workflow.py(parsed; noactionlinton this machine), and has run: green on v0.20.2 (above). Node 24 on Linux is therefore measured (client, 19 s). A run on a later push is the check for anything this release changed,ui/included. -
engine/hosting/boot.py::stop_master's re-parented branch (never signal a master whose worker was re-parented) is unit-tested but was not reached under real gunicorn: in T18's smoke test a SIGKILLed master's worker was ended both times by gunicorn's own parent check ("Parent changed, shutting down"), once even when frozen until the supervisor had restarted the story. No process was signalled in the dead master's name either way. -
"The front door last" holds for a shutdown the supervisor runs (a signal, Ctrl+C, a fatal error in
__main__): a supervisor killed outright (SIGKILL, a crash past__main__) takes the front door and every worker down at once through the bus lifeline, with no order. And a worker that overruns its stop by up to a second eats into the front door's 10 s reserve (engine/hosting/supervisor/server.py::shutdown), shortening only the door's graceful stop, never pastshutdown_seconds. Accepted (T18 re-review). -
The LM Studio skills server's own uvicorn (v0.20.0 T9) is not live-verified with a model loaded.
llm.mcpnow starts its ownuvicorn.Serverover fastmcp's SSE app soSkillsServer.stopcan end it; the suite covers start, stop and themcp.jsonentries against the temp root, but at the release (2026-10-06) LM Studio was not running and no model was loaded (the check never loads one), so no flagship turn has called a tool through it. The check, when a model is already loaded:llm.mcp.enabledin a temporary config layer only, one flagship turn using a tool, the owner'smcp.jsonbyte-identical to its pre-run copy, the server stopped cleanly. -
The Docker image is 327 MB, 110 MB of it the committed art under
content/, which every story's container carries whether or not it serves that story. Measured, left. -
The session snapshot of the owner's storage (
tests/conftest.py::_real_storage_is_untouched) cannot tell the suite from two other writers: the owner's own game left running (an autosave, a generated plate), and LM Studio itself rewriting itsmcp.jsonwhen the owner edits its servers. Either fails the session's last test; the message says so and lists each change's size and mtime before and after. It also proves only "net unchanged": a file made and removed within the session leaves no trace (the in-process audit hook covers "never touched"). -
The child sandbox's residual gap (
tests/conftest.py::pytest_configure,engine/config.py::child_sandbox): a child started by a route that bypasses thePopenwrapper (os.execve,os.spawnve,posix_spawn,_winapi.CreateProcess, ctypes) AND with an explicit environment that drops the marker is not sandboxed. Every other route inherits the marker (it lives in the suite's ownos.environfrom the top of the conftest) or has it re-injected by the wrapper whatever itsenv=. The AST scan (tests/test_subprocess_sandbox.py) bans theos.*spellings of those routes in first-party code but does not seegetattr(os, ...), which matters only together with a dropped marker (agetattr(os, "system")child still inherits it). -
A sandboxed child's
llm.base_urlis the discard port unless the SANDBOX LAYER's ownllm.base_urlis loopback on the port the test registered (tests/conftest.py::sandbox_model_stub(port), which sets both, viaCLOCKWORK_TEST_MODEL_STUB_PORT;engine/config.py::_sandbox_model_url). A supervisor child reaches its stub model server that way, not through aCLOCKWORK_CONFIGof the test's own, whosellm.base_urlis always overridden (spec §3.5's CURRENT note). -
scripts/start.ps1was not givenstart.sh's dangling-.venv-link refusal (v0.20.0 T5): PowerShell'sTest-Pathon a broken link was not measured here, so the Windows twin is unchanged. -
The owner's Windows
.venvcarries a stalefastmcp3.4.7 /fastmcp-slim3.4.7 record beside thefastmcp3.2.4 whose files are installed (and which imports).constraints.txtpins 3.2.4 and leavesfastmcp-slimout; the record itself is left alone, because uninstallingfastmcp-slimwould delete files the two share. A fresh.venvinstalled with-c constraints.txthas no such record. -
engine/games/caches.py::warm_all_caches()(v0.20.0 T6) is called by hosted mode's startup (engine.hosting.install, T7) and never in local mode. (The loaders a config reset raced -- everyNULLED_ATTRIBUTESsite -- read their cache once into a local since T6 fix round 2, pinned bytests/test_cache_reset_race.py; a store a reset drops shares its save folder's one index lock with its replacement since T8,saves.index_lock_for.) -
Turn payload opening choices may label a present NPC the player has not met (an intent's label or
npc_id). The opening's prose introduces those present, and withholding the label would move the turn goldens (controller ruling, v0.21.0). -
The summarizer model's input (
engine/memory/summarizer.py::_render_turns) still sends[day N, <location_id>], so a model may echo an id into the recap. Unchanged because it would move the LM Studio golden (controller ruling, v0.21.0; the fallback summary no longer prints ids). -
THE LONG CON has no stage, so on desktop the roll card's one line covers one log line for its 6 s (v0.21.0; Dev Story's engine skin has a stage, and its card sits on the plate, re-measured in the final fix wave).
-
The flagship's and NEON CITY's overlays (barter, item use, crafting, posted work, a thread's paper) send the player's words as typed text, which carries no intent, so no skill runs from them and the prose can tell of a trade the save never made (a GOVERNANCE NOT WIRED row names the five files). The overlay-to-intent path is the v0.23.0 (flagship) and v0.25.0 (NEON CITY) overhauls' work, not a fix (v0.21.0 final review, controller ruling).
-
Under 640px tall, with the people strip in the stage, the scene plate gives way to the strip entirely (HUE & CRY at 900x600 and 844x390), so the log keeps three lines and the choices two rows (v0.21.0 final fix wave).
-
At 844x390, scrolled to the choices, the roll card covers the job panel's chevron for its 6 s (no clicks are taken there; v0.21.0).
-
Under 900px the shelf is drawn above the log while the log comes first in reading order (visual and DOM order differ; v0.21.0).
-
Core's panels print a few words of their own ("Wanted", "Held", "Casing", "Prep") and a story of another register cannot rename them: a label override is not built (
ui/src/core/panels/; docs/GOVERNANCE.md). -
Not drawn or told, v0.21.0, each a docs/GOVERNANCE.md row: the companion in the people strip (
PeopleStrip.jsx), which jurisdiction the poster stands in (_law_block), the prose of a turn finished across a server restart, a story switched in another tab then a reconnect, a player's place in the queue, and the flagship's and NEON CITY's encounter look (theirs to restyle in v0.23.0 and v0.25.0). -
The
processguard (tests/tier_plugin.py, v0.21.1) sees only the suite's directsubprocess.Popenchildren. A grandchild interpreter (ansh, or on Windows acmd /c, that starts python) and a start that bypasses Popen (os.system,multiprocessingspawn,_winapi.CreateProcess, and the sandbox test's repatched-Popen route) are not recorded, so such a test is marked by hand (test_subprocess_sandbox.py's any-route test is). A child started from a background thread is charged to whichever test is running then. Closing it needs the child to report itself through the sandbox marker, which is its own work. -
This workstation's loopback fails in bursts while the hosted tests run (measured in v0.21.1 by a monitor beside
tests/test_admin_model.py: none in 1424 fresh connects at idle; during runs, spells of one to three minutes in which most fresh connects time out -- never refused, so the bus connect's refused-means-no-supervisor rule stands -- and a kept bus link can be aborted, while MsMpEng and WmiPrvSE carry heavy CPU). The owner's Windows Defender exclusions (2026-10-07 20:29: the repository, the Claude temp root, bothpython.exe) did NOT end them: the monitor, started after, still saw spells at 20:36-20:39, 20:46-20:47 and 20:50-20:51, shorter than before. The bus connect, the test relay and the hosted test clients now ride out short spells (5 of 5 runs passed across two of them); a long one can still fail a hosted test. Cause unknown and outside the repository. -
The NOT WIRED tables: docs/GOVERNANCE.md, docs/STATE.md, docs/AGENTS.md.
The full accounts -- measured, with the reasoning -- are in docs/DESIGN_REVIEW.md § Findings after the overhaul. The one-line index, because each is a mistake worth not repeating:
- A test IS a caller. Three subsystems and eleven skills had no production
caller and full test coverage;
tests/test_reachability.pynow walks the engine's own call graph, constants included. - Reached is not narrated. Threads, clocks and -- in four of five stories -- presence itself were enforced and invisible to the prose.
- Narrated is not agreed. The evaluator asked whether a roll existed, never
whether the prose matched it;
workreported how a shift went under the key meaning whether it happened. - Tests were talking to LM Studio while believing they were mocked; the conftest guard asserts at teardown because the pipeline's forgiveness swallowed a raise.
- State outlived its test twice: a module memo (
_DOOM_DECLARED) and a story activation. Both have autouse fixtures now. - A veneer hid a template. THE LONG CON sold "Hedge Berries" as cigarettes
through a
name:key nothing read.
- The default pytest temp directory is unreadable here (WinError 5, environmental).
scripts/run_tests.pymakes a fresh basetemp per run under%TEMP%\claude\clockwork-tests(setCLOCKWORK_TEST_BASETEMP_ROOTto move it), because pytest wipes its basetemp at start and two runs sharing one destroy each other's trees. A barepyteststill needs--basetemp=here (ptmp-<agent>, one per parallel run). - Every test has a limit (pytest-timeout: 300 s, 900 s in a tier,
×
CLOCKWORK_TIMEOUT_SCALE), andrun_tests.pyputs a wall-clock limit on the whole run, stopping the whole process tree (a Job Object here). A barepytestis safe here too: the conftest puts the session in a kill-on-close job, so a timed-out test's children (the thread method'sos._exitskips finalizers) end with the session (v0.21.1 T2 fix round 1). - Windows Defender exclusions (the repo,
Temp\claude, the.venvand base Pythonpython.exe) were added by the owner on 2026-10-07; the import baseline fell sharply. Loopback stall bursts remain, cause unknown. - Loopback TCP connects here sometimes stall: in a plain Python loop of
3000 listen/connect/accept pairs (2026-10-07, nothing else running), 13
were never accepted and many more arrived 1-16 s late, in bursts. Every
hosted test's socket wait is bounded (the bus's wake pair since
v0.21.0), so a stall fails a test rather than hanging the run; a lone
connect timeout in the hosted files is worth one re-run before a hunt.
Under xdist (the parallel phase of
run_tests.py fastandfull) the stalls come oftener and in bursts that fail a whole hosted module at once.run_tests.py fast/fullre-run each serial-phase (loopback, process, mcp_server) failure once, alone, and pass the run when all pass, listing them asFLAKY (passed on solo re-run); a test that fails twice is a bug to hunt, and so is any parallel-phase failure. An idle probe of 1000 pairs here (2026-10-08) had 1-5 over 1 s, worst 3 s. A localllama-serverholding ~22 GB was running during the 2026-10-07 xdist measurements (4-5 GB of 31.5 free); it was gone by the 2026-10-08 hybrid re-measure (17.7 GB free at its end). - Heredocs and
python -cstrings lose backticks to shell command substitution; write commit messages and patch scripts to a file first.