Skip to content

Native EECH campaign: eech_dc on the full engine, Windows build, retail campaigns and regression - #59

Open
fd-shockwave-bot wants to merge 77 commits into
masterfrom
claude/eech-campaign-native-spike-tjkj6g
Open

fd-shockwave-bot wants to merge 77 commits into
masterfrom
claude/eech-campaign-native-spike-tjkj6g

Conversation

@fd-shockwave-bot

Copy link
Copy Markdown

Summary

This PR brings the native EECH campaign work on claude/eech-campaign-native-spike-tjkj6g into master, as the base for the next phase of building. master is an ancestor of the branch, so the merge is clean: 40 commits, head 5f07697b.

The project owner has accepted the branch as it stands. Most of these commits were pushed directly to the integration branch without PR review, and that is recorded here rather than reworked.

What's included (eech-campaign/)

  • eech_dc: all of EECH compiled headless as a Lua 5.1 module, plus the smaller campaign kernel behind a Rust API, with its C-reference corpus.
  • eech-world: a Rust host that embeds Lua, runs the campaign through require ("eech_dc"), and records Tacview.
  • Windows build: eech_dc.dll, eech-world.exe and lua.dll built with MinGW-w64 (tools/build-windows.sh, tools/Dockerfile), with a crash reporter.
  • Retail scenarios: Georgia (map3) and Lebanon (map5), with setup scripts. Retail data is not in the repository.
  • Regression test: tools/regress-windows.ps1, which checks campaign expectations and compares against regression/windows/ baselines.
  • Corrections to original EECH behaviour: S2 and S3 (supply), N1–N3 (NaN terrain and smoke), H1 and U1 (memory errors), and earlier ones, all documented in docs/engine.md.
  • Governance: AGENT.md, PROJECT.md and docs/governance/, merged into the branch through docs(governance): establish project charter and baseline-first agent policy (#51) #53.

Validation (Windows, at 5f07697b)

  • Regression: tools/regress-windows.ps1 -Georgia -Lebanon passes the expectations (Georgia 30/30, Lebanon 39/39), and both campaigns are IDENTICAL to their baselines. Repeat runs are identical.
  • Full war: retail Lebanon played to a conclusion (red wins at 50.4 simulated hours), with no crash and no NaN reports.
  • AddressSanitizer: 24 simulated hours of Lebanon (two seeds) and of Georgia found H1 and U1, both now fixed. Nothing else was reported.

Known limitations

  • Linux baselines: regression/*.json were last recorded at 519e35da, before S2, S3, N1–N3, H1 and U1. Windows is the maintained baseline.
  • CI: .github/workflows/eech-campaign.yml predates the whole-engine build. It uses the runner's default GCC and tests the whole workspace for i686 and under Wine, so its engine jobs may fail until the workflow is reworked.
  • Retail data: the regression needs local game installs (tools/retail-*.sh), so it can't run in hosted CI.
  • Architecture: the whole engine, physical world included, is still inside eech_dc. That is transitional; the world moving into the host is later work.

Review

Needs a formal approval from @fd-starscream-bot before merge (AGENT.md).

🤖 Generated with Claude Code

https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK

claude and others added 30 commits September 23, 2026 23:29
Adds a Cargo workspace (eech-campaign/) that compiles the original EECH
campaign C for the host target and hides it behind a safe Rust API:

- eech-sys: build.rs ports the eech-core-ts C reference extraction spec
  (whole original TUs, verbatim extracts with #line provenance, a reduced
  project.h) and compiles it natively with the cc crate. Two exact-text
  patches with provenance: en_creat.c's va_list reinterpretation (64-bit
  blocker) and ks_updt.c's function-local static. Host layer in csrc/:
  dispatch tables, environment, entries with setjmp aborts and the
  round-toward-zero FP environment, entity identity (index, generation).
- eech-campaign: public API (Campaign, CampaignConfig, World, events,
  snapshot), no unsafe; conformance replay of the eech-core-ts scenario
  language through the same kernel and World boundary.
- corpus: 1401 eech-core-ts scenarios and 3000 float operations exported
  with the output of the 32-bit C reference and the TSTL verdict; the
  x86-64 native module reproduces all of them byte for byte.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…ntory

- Campaign::step drives EECH's update loop: keysites (ks_updt.c) consume
  supplies and stock cargo, the force's low-on-supplies response builds a
  supply mission, the airbase's assignment pass selects a group and stops at
  the slice boundary (assign_primary_task_to_group) with the group and task
  ids. Missing world answers, unported rows and panics fail loudly and
  poison the campaign; ids carry their campaign instance.
- eech-harness: JSON scenario -> HeadlessWorld -> Campaign -> JSON result;
  four recorded semantic scenarios.
- global-state.txt classifies every writable symbol of the kernel (checked
  against nm); resets for the statics found (suitability array via EECH's
  deinitialiser, task.c and ks_int.c load-time values); database digest
  test; B1 probe test; diagnostics name EECH ordinals.
- Build variants: EECH_C_OPT, EECH_FPU_ROUNDING=nearest (investigation),
  EECH_ORIGINAL_STACK_ATTRIBUTES=1 (i686 only).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
- The module builds for x86_64-pc-windows-gnu (LLP64, the DCS platform)
  and passes the whole suite under Wine, the corpus included. highlevl.h's
  WIN32 release ai_log only preprocesses under MSVC: GCC-family Windows
  builds take EECH's own non-WIN32 definition (-UWIN32, finding T1).
- The NULL-dereference outcome of a dedicated replay process is reported on
  Windows through a vectored exception handler.
- cc no longer passes -w (it silenced the -Werror= promotions); our own
  sources build with -Wall -Werror, which found %d with ptrdiff_t indices.
- CI: x86-64, i686, i686 with the original va_list idiom, -O2, Windows
  cross under Wine; rustfmt and clippy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
README and docs/: feasibility report (verdict, milestone evidence,
risks, recommendation), 64-bit classification (B1 blocker, T1-T4
portability, H1-H3 legacy), floating point (RTZ entries, results matrix
across compilers/optimisation/platforms, rounding-mode sensitivity),
global state (singleton, sequential, inventory classes, G1-G6), patches
and compile definitions, ports derived from call sites, source closure
and vendoring, conformance (three-way corpus), DCS constraints.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…affold

- eech-dc: cdylib (mlua 0.10, lua51 + module, modelled on flying-dice/
  dcs-studio's bridge crate): require ("eech_dc") gives eech.new / step /
  snapshot / close, an embedded Lua driver (eech.start: model-time pump,
  world queries, protected callbacks, teardown at mission end) and
  coordinate helpers. No panic paths (restriction lints denied); every
  failure is a Lua error. Tests run the module in a real Lua 5.1 state.
- eech-world: the headless harness host. It links Lua 5.1, creates the
  state a simulator creates (native modules allowed) and loads the DC DLL
  from the build directory with require. Tacview ACMI 2.2 writer using
  EECH tacview.c's geodesy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
The maintained Windows build's 1,225 C files (the aphavoc and modules
Visual Studio projects), compiled on the WIN32 code path against compat/
headers that declare the Windows SDK and DirectX names EECH uses.

- compat/: Win32 types (LLP64 widths), DirectX value types and constants,
  opaque COM interfaces, and prototypes for every Win32/CRT call so no
  pointer result passes through an implicit int declaration.
- csrc/: the headless platform layer. Win32 over POSIX (files, mappings,
  Windows paths resolved case-insensitively, text-mode reads), simulated
  time, debug_fatal unwinding to the entry point, null DirectX COM calls,
  and replacements for the 8 backends that need a real Windows machine
  (WinMain, debug windows, joysticks, GDI fonts, DirectDraw, Direct3D,
  DirectPlay, GUIDs).
- build/patches.rs: 3 source patches (two MSVC-only macros, one union
  with repeated member names) plus the spike's P1 va_list marshaller.

Implicit function declarations, implicit int and int conversions are
build errors. The result links with no undefined symbols.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
eech-engine-sys:
- csrc/eech_engine.c: boot and frame entry points following EECH's own
  dedicated-server path (WinMain, application_main, brief_ and
  full_initialise_game, the game-initialisation phases, flight () up to its
  loop); a frame is one iteration of flight ()'s loop without drawing.
  debug_fatal unwinds to the entry point; FPU round-toward-zero per entry.
- csrc/eech_synth3d.c: writes a 3D object database in EECH's formats
  (bininfo.bin, 3dobjs.*, 3dobjdb.bin, displace.bin, stars.bin) with the
  engine's compiled name tables. The retail database is not in the repo.
- memory-backed Direct3D surfaces and video screen, DirectDraw init.
- 1x1 placeholders for missing UI artwork (.psd, .bmp, .tga), logged.
- patch B2: UI attribute pointers read as int (ui_attrs.c, 9 sites).

eech-map: Luxembourg from OSM + SRTM: terrain (ffp/sec/rgb), road network
(ROADS.dat/.nde/.wp), side map, population placement, campaign script,
popname.dat, bridge.pop, mapinfo.txt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…acview

- eech-engine-sys: airport and FARP scenes in the synthetic 3D database
  (landing, takeoff and holding routes for fixed wing, helicopters and
  vehicles, as routegen.c reads them; hangars and towers as scene links);
  eech_observe.c reports aircraft, vehicles, air defence, ships, infantry,
  weapons in flight and keysites the way EECH's own Tacview logger
  classifies them; clock.
- eech-engine: safe API (boot once per process, frame, objects, clock;
  debug_fatal becomes EngineError::Fatal and poisons the engine).
- eech-dc: the Lua module now contains the whole engine:
  require ('eech_dc') -> boot{...}, engine:frame(ms), engine:objects(),
  engine:clock(), write_3d_database(dir).
- eech-world: the host binary. Embeds Lua 5.1, lets the script require the
  DLL, runs lua/campaign.lua, records what it observes to Tacview ACMI 2.2
  (positions only when changed; destroyed events; keysite captures).
- eech-map: the campaign script now has forces, reserves, task generation,
  division numbers, frontline forces, airbase groups; FARPs and SAM sites
  in the population file.

A one-hour Luxembourg campaign runs headless without faults.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…orted once

- eech_engine_prepare_installation (Lua: dc.prepare_installation): writes
  the generated part of an installation, the 3D database, textures.bin
  (texture names generated at build time from modules/3d/textname.h) and
  brief_en.dat (one briefing per task type). The retail files are not in
  the repository.
- eech-map copies the EECH data the repository carries (setup/common/data:
  formation databases, language, suspension) into the installation.
- EECH_TRACE_FILES=1 logs every file the engine opens or misses.
- keysites are on the update list too; observe reports them once.
- campaign.lua logs the tasks groups are on, per side.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…il for absent fields

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…istic guard

EECH aims through WEAPON_SYSTEM_HEADING/PITCH sub-objects of the launcher's
3D scene; without them WEAPON_AND_TARGET_VECTORS_VALID is never set and no
AI unit fires. The synthetic 3D database now gives every aircraft and
vehicle scene a heading/pitch/muzzle/weapon tree sized from the
weapon_config_database packages its type can carry.

Weapon weights, drag and motor power come only from the GWUT table; the map
tool now installs GWUT1162.CSV (a missile launched without it has a NaN
velocity). Patch W1 guards get_ballistic_pitch_deflection against
|height| > range at point blank (NaN pitch, INT_MIN table index).

Also: object task state in the observation API, more airbase groups and
pads (minimum idle group counts), wrecks and weapons in the Lua summary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Linked hangars carry REGEN_FIXED_WING/HELICOPTER/ROUTED_VEHICLE
sub-objects, so EECH rebuilds losses from the force's hardware reserves.
Without them air tasking stops once losses bring each group type down to
the minimum idle count assign.c holds back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…uncher's side

Without EXPLOS.CSV EECH runs on its compiled explosion database, which
declares components it never initialises (XSMALL_HE: 5 declared, 3 set);
a fresh installation then faults on the first small explosion. eech-map
copies EXPLOS.CSV, METASMOK.CSV and SMOKES.CSV as an installation ships
them. Observed weapons now carry their launcher's side, so Tacview colours
them by coalition.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
A REPAIR or SUPPLY task never starts from the keysite it serves, so a
side whose only airbase is struck out of action is paralysed for the rest
of the campaign. eech-map now gives each side two airbases: the two
largest named aerodromes at least 10 km apart, or one synthesised at the
side's town farthest from its other airbase, placed off water (EECH turns
a keysite on water into an anchorage). Hardware reserves are sized for
days of war.

engine:diagnostics() (eech_engine_diagnostics) logs each force's
reserves and regen queues, and each keysite's state, supplies, regen
sites and air groups; campaign.lua calls it every diagnostics_every
simulated seconds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Troop insertion reads the helicopter's TROOP_TAKEOFF_ROUTE and
TROOP_LANDING_ROUTE (under WAYPOINT_ROUTES) to put infantry down and pick
it up; without them the first troop insertion is fatal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
recordings/luxembourg-6h.zip.acmi: six simulated hours of the dynamic
campaign from the harness. Red's troop insertions capture both blue
airbases; blue's air force dies out without a base to regenerate from.

EECH reseeds its random numbers from the system clock when the session
starts; the simulated clock's epoch now follows the boot seed (seed 1
keeps epoch 0), so the seed selects the run. The host writes object
removals in id order, so the ACMI no longer depends on hash-map order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
eech-map takes a map name (or reads it from the extract's file name):
Georgia covers 40.4-46.7 E, 41.0-43.6 N (256 x 142 sectors; EECH's
campaign map holds at most 128 campaign sectors a side), front at 44.0 E
through Shida Kartli, airbases Kutaisi and Senaki (blue), Vaziani and
Marneuli (red), tertiary roads, English names where the local script is
not Latin, the Black Sea as sea terrain. campaign.lua takes
scenario=luxembourg|georgia.

X1: EECH's GNU convert_float_to_int is 'fistp (%1)', the 16-bit store in
AT&T syntax: values from 32,768 up saturated, so every world coordinate
past 32.7 km fell into the wrong sector. Convert in C (truncation, the
rounding EECH sets). The Luxembourg recording predates this fix and is
being re-recorded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…e resupply fix

Airbases and FARPs only consume ammo and fuel; factories and refineries
make it and SUPPLY missions fly it. eech-map places a factory and a
refinery per side on the largest OSM industrial areas behind the front, as
key templates in a version 2 population file (popread.c creates keysites
from key templates only in version 2 files).

S1: an airbase low on supplies picked itself as its closest supplier
airbase (get_closest_keysite without exclusion), so no airbase was ever
resupplied; exclude the requester.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
The airbases sit 70-160 km behind the front, beyond most helicopter
tasking; FARPs are populated only by FRONTLINE_FORCES (two groups each),
so Georgia gets four a side.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
recordings/georgia-12h.zip.acmi: twelve simulated hours of the Georgia
campaign; recordings/luxembourg-6h.zip.acmi replaces the recording made
before X1 (its sector lookups past 32.7 km were wrong).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…rage tool

- prepare_installation keeps a retail 3D database (textures.pal present)
- tools/retail-map3.sh assembles an install root for the retail Georgia
  campaign (map3, Caspian Black Gold) from retail data kept outside the
  repository; the retail script's 30-minute END_CAMPAIGN FAIL trigger is
  dropped (the original is kept as .retail)
- campaign.lua scenario georgia_retail; the Tacview recorder takes an
  affine map projection: map3 is stretched 1.22 east-west, fitted by
  correlating the retail terrain heights with SRTM (r = 0.978), where
  EECH's own metric origin is off by up to 50 km
- tools/offsite.sh: rclone against one Google Drive folder for install
  data and run outputs shared between sessions

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
The Steam Apache vs Havoc 3dobjdb.bin holds 2,761 of the 3,026 scenes this
source defines; a modern install adds the rest from the community objects
(setup/cohokum/3ddata/objects), which retail-map3.sh now copies in.

- C1: a scene's collision object 0 (the null object) means none on the
  3dobjdb.bin path too (the .EES path already maps it to -1)
- T1, T2: headless, a texture camouflage mismatch or a missing texture
  animation between the community objects and the retail texture set is
  logged instead of fatal (nothing is drawn)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
- recordings/georgia-retail-2h.zip.acmi: 2 simulated hours of the retail
  map3 campaign (FARPs 18, 7 and 3 change hands), with a README section
- tools/acmi-summary.py: losses, weapons, captures and task mix from a
  recording and its run log
- Luxembourg removed: its recording, eech-map spec, campaign.lua scenario
  and the README/engine.md sections; examples now use map16/georgia.chc
- handover: building locally with Docker (rust:bookworm, GCC 12) and the
  local retail data sources

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fd-shockwave-bot and others added 30 commits September 26, 2026 00:51
The M1 review (#62, docs/corrections.md) rejected S2 for the M1 candidate:
its stated need is not reproduced once S3 is present, and S3 explains the
diagnostic that motivated it. Without S2, response_to_force_low_on_supplies
compiles as written: when the chosen supplier holds no crate of the type
asked for, no SUPPLY task is created and the keysite asks again later.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
regression/windows/lebanon_retail.json is re-recorded at 0288ece (S2
removed). Georgia is IDENTICAL to its baseline (retail Georgia has no
factories or refineries, so S2 never acted there); Lebanon keeps all 39
campaign expectations with 23 metrics outside the old baseline's
tolerance: fewer SUPPLY tasks from different suppliers (blue 11 -> 10,
red 52 -> 49), then a diverging war. The new Lebanon metrics are
byte-identical to the #62 review's independent no-S2 run, so the new
baseline was produced twice.

docs/corrections.md records the removal and the before/after,
docs/engine.md drops S2 from the patch table, docs/m1-baseline.md gives
the updated candidate (eech_dc.dll da72f128...) and marks S2 as removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
Review feedback on #63: docs/m1-baseline.md still presented a8661ea as
the candidate and the correction review as an open blocker. It now:
- identifies 0288ece (S2 removed) as the candidate, with eech_dc.dll
  da72f128... and the unchanged host and Lua artifacts;
- keeps a8661ea as provenance only, with a table of which evidence was
  produced where (build identity, lifecycle and regression rerun at
  0288ece; rebuild check, sensitivity and kernel at a8661ea);
- gives the regression figures at 0288ece and the new Lebanon baseline;
- states what is established for the post-S2 candidate, including the
  completed correction review (#62, docs/corrections.md);
- leaves only conditional older-doc reconciliation and the #19
  acceptance review as remaining M1 work.
Documentation only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M1: remove S2, restore the original EECH supplier rule; new Lebanon baseline (#19)
- lua/observations.lua (campaign.lua observe=, observe_every=): structured
  observations from the samples the Tacview recorder receives: appear,
  retype, alive, side, usable and gone events every sample, and full
  snapshots every observe_every seconds and at each metrics checkpoint.
  campaign.lua computes the checkpoint condition once for both; the metrics
  are unchanged.
- tools/input-manifest.py: SHA-256 of every input file of a root, with a
  combined hash; the files the engine writes are listed apart.
- tools/observation-check.py: replays the observations through the
  recorder's rules and requires the Tacview recording's declarations,
  removals and Destroyed events frame by frame, its positions at every
  snapshot within the recorder's change threshold (with its own
  implementation of the map projection), and losses, captures and keysite
  state to agree with the metrics; reports implausible jumps.
- tools/observation-check-selftest.py: the check passes unmodified evidence
  and fails five deliberate discrepancies.
- tools/reference-windows.ps1: the reference scenario, run twice and
  checked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m2-reference.md records the reference scenario and answers whether
EECH World's observations of it can be trusted as physical-world evidence.
The M1 candidate's own binaries (0288ece) with a1941f2's scripts ran
retail Lebanon for 3 simulated hours twice, from two separately assembled
roots with identical inputs (1,179 files, 121a7890...): the Tacview
recording, the structured observations and the metrics are byte-identical.
tools/observation-check.py passes on both: identity and lifecycle exact in
all 10,800 frames, 308,368 snapshot positions within the recorder's
threshold, losses identical three ways, keysite state identical to the
metrics. The metrics equal the M1 Lebanon baseline, so observing does not
perturb the campaign; the M1 regression and lifecycle checks are unchanged.

Fidelity limits found: EECH reuses entity indices, merging some weapon
tracks (23 jumps, 12 launcher changes seen); Tacview attitude can be stale
and its keysites carry no capture or supply state; EECH's session clock
lags host time by 29.6 s in 3 h (float accumulation explains 4.6 s) and
Tacview's reference time is a fixed label.

reference/lebanon-3h/ keeps the recording (30.3 MB zipped), the
observations (8.2 MB), the metrics, the input manifest and the check
report. The checker now breaks plausibility flags down by type and counts
weapon indices held by another launcher; the runner records the baseline
comparison's verdict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M2: a trustworthy EECH World reference scenario and its observation evidence (#23)
- lua/observations.lua, campaign.lua observe_supply=1 (off by default, so
  runs without it write the M2 observations unchanged): task events when
  a unit's group, primary task or operational state changes; supply
  events when a keysite's ammo or fuel moves by 0.5 or more in a sample;
  track events every sample for aircraft whose group is on TASK_SUPPLY.
- campaign.lua loads observations.lua before the engine boots: the boot
  changes the process's working directory (csrc/eech_engine.c chdir to
  <root>/cohokum), so a relative script path no longer resolved after it.
- tools/supply-chain-check.py: every delivery (a keysite level going to
  100) attributed to the SUPPLY group over the receiver at that sample,
  with its assignment, the one-crate pick-up at the supplier with the
  group over it, the continuous 1-second flight between, the end of the
  task, and what the campaign then does with the delivered stock; the
  delivery count must agree with the metrics' resupplied counts.
- tools/supply-chain-check-selftest.py: removing the assignment, the
  pick-up, part of the flight or the aircraft at the drop-off breaks the
  chain.
- tools/reference-windows.ps1 -ObserveSupply: runs the chain check on
  both runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…ained

docs/m3-resupply-loop.md follows one real resupply chain through the
public path (eech-world.exe -> campaign.lua -> require ("eech_dc") ->
engine:objects ()): Il-76 group 88235 "Jester" is assigned TASK_SUPPLY
(1,160.77 s), flies continuously to Halat, where Halat's fuel drops one
crate as it passes 85 m overhead (1,555.67), and on 30 km to Beirut
International, whose fuel goes 10 -> 100 as it passes 43 m overhead
(1,904.58). The campaign then uses the fuel: groups draw on it, another
SUPPLY task picks up from it, and once it is below the request threshold
another SUPPLY task delivers to Beirut again (3,324.68).

Across the run all 32 deliveries are attributed to the SUPPLY group over
the receiver; 31 chains are complete, and the one that is not is a crate
carried over from an earlier task: a drop-off to a group leaves the crate
aboard (mb_msgs.c), recorded as a finding. The deliveries equal the
metrics' resupplied counts; two runs are byte-identical; the metrics and
the recording equal the M1/M2 evidence, and the M2 settings reproduce
M2's files byte for byte.

reference/m3-resupply-loop/ keeps the observations (10.5 MB), the
featured chain's events and the check reports. The runner's title is
generic; the checker's header states the carried-over crate case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M3: one campaign -> world -> campaign resupply loop through the public path (#27)
…next

docs/m3-supply-loss.md follows the naturally occurring loss in the
accepted reference run's retained observations (no new run, no runtime
change): red Mi-17 87568 of SUPPLY group 87567 "Wolfpack" dies at
1,584.66 s while Performing Task, 32 km from Beirut, as two blue MIM-72G
Chaparrals vanish at it. The group falls from two aircraft to one with no
replacement; the survivor turns back at that moment and its task ends at
Beirut with no keysite delivery; the campaign re-tasks the one-aircraft
group at 3,619.24 s and it delivers fuel to Power Station 3 at 4,013.67 s.

No red keysite was requesting fuel when the task was assigned, so its
receiver was a group (by elimination and the code), whose supply is not
reported: the lost cargo and the receiver's shortfall are not observable.
Facts are classified as observed, inferred, not observable and code.

tools/loss-chain-check.py and its self-test; reference/m3-supply-loss/
keeps the report and the records behind each link.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M3: a SUPPLY aircraft destroyed in flight, and the campaign's response (#27)
tools/build-mutant-windows.sh <patch id | control> target/mutants/<dir>
builds the Windows engine from a temporary git worktree of the clean HEAD
(under target/mutants/, removed afterwards), with one patch entry removed
from build/patches.rs in that worktree only (checked to be that entry and
nothing else), in its own cargo target volume, into target/mutants/ only,
stamped MUTANT or CONTROL with the source diff and the patch markers of
the staged sources it compiled. The working tree and the normal build
(tools/build-windows.sh, target/windows, the normal volume) are never
touched.

tools/regress-windows.ps1 -Out: where a run's outputs go (default
unchanged), so a control and a mutant can run side by side.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…tes leading-slash patterns); check the checkout

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
tools/mutation-windows.ps1 builds a control and a test-only mutant with one
patch removed (tools/build-mutant-windows.sh, the same clean commit),
checks that they differ only in that patch (the same commit, image,
toolchain, lua.dll and scripts; a one-entry source diff; the compiled
patch markers differ by that patch alone), gives each mutant run an
identical copy of the inputs, runs tools/regress-windows.ps1 -Exact on
both side by side, and reports DETECTED when the control passes and the
mutant fails, with the metrics that changed (tools/metrics-difference.py).
-RepeatMutant runs the mutant twice and requires identical metrics.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…hers

git worktree prune also removed another session's worktree record, whose
recorded path Windows git took as stale; the builder now removes only the
temporary worktree it created.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m4-s3-mutation.md: tools/mutation-windows.ps1 built a control and a
test-only mutant with S3 removed from the same commit (56b164d), each in a
temporary worktree, and showed they differ in that patch alone (one-entry
source diff; compiled patch markers differ only by S3; the same image,
toolchain, lua.dll and scripts; identical input copies). The accepted
regression (tools/regress-windows.ps1 -Exact) passes the control (30/30,
39/39, IDENTICAL) and fails the mutant in both campaigns: Lebanon 38/39
("airbases resupplied with ammo (0)") and not identical, Georgia not
identical; the tolerance comparison flags both too. Two mutant runs are
byte-identical. The metrics name the behaviour: ammo deliveries vanish
(Lebanon airbases 3 -> 0, military bases 6 -> 0; Georgia airbases 6 -> 1)
while fuel deliveries rise, the signature of ammo tasks loading fuel.

reference/m4-s3-mutation/ keeps the summary, the difference, both builds'
identities, the diff, the input manifests and each run's regression output.
regression/README.md and README.md point to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M4: the accepted regression detects a reintroduced campaign defect (S3 mutant)
regression/pack.json lists every materially accepted M1-M3 behaviour with
where it was accepted, its path, the check that protects it and how
(direct, baseline-sensitive, evidence only, not covered), and the blind
spots; plus the accepted candidate and scripts, the scenarios, the input
manifests, and the hash and revision of every retained artifact, checker
and runner the judgement depends on.

tools/regression-pack.py runs the existing checks and reports each claim:
- retained tier (no retail data, about a minute): artifact, checker and
  manifest hashes; the checkers reproduce the retained reports from the
  retained evidence; their self-tests still detect their discrepancies;
  the S3 mutant's retained metrics are still rejected, and Georgia's
  expectations still miss it (the documented blind spot). Any failure
  stops the pack.
- runtime tier (a build and the retail roots): the roots against their
  manifests, lifecycle-windows.ps1, then regress-windows.ps1 -Exact and
  reference-windows.ps1 -ObserveSupply in parallel, loss-chain-check.py,
  and the fresh evidence against the retained.
EVIDENCE ONLY and NOT COVERED come from the inventory and never pass.

tools/lifecycle-windows.ps1 gains -Out (default unchanged), so that the
pack's runs do not share an output directory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…served run

reference-windows.ps1 judges "observing does not perturb" by comparing the
observed run with the Lebanon baseline. On a changed build that comparison
fails whether or not observation perturbs anything. The pack now asserts
M2.unperturbed by comparing the observed run's metrics with the same
build's unobserved regression run (regress-compare.py --exact). The
runner's own baseline comparison becomes reference.baseline, part of the
baseline-sensitive M2.reference_reproduced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m4-regression-pack.md answers what accepted M1-M3 behaviour the
pack would detect (17 claims, directly protected), what would merely look
different (12, baseline-sensitive) and what could regress undetected
(8 evidence only, 8 not covered), with the blind spots and provenance.

reference/m4-regression-pack/ keeps both runs (pack revision 90acbac):
- accepted: the candidate's binaries (0288ece) with the M3 scripts:
  PASS, all 38 checks; every accepted result reproduced.
- s3-mutant: #68's mutant, not rebuilt: FAIL. M1.lebanon.resupply and
  M1.corrections.S3 fail as directly protected ("airbases resupplied with
  ammo (0)"); every baseline-sensitive claim differs, including S1 and S2,
  which did not change; M3.resupply_loop stays green although all 30
  deliveries were fuel, a blind spot now demonstrated.
- fail-loudly.txt: an altered artifact, a missing one and an edited
  checker each stop the pack before any judgement.

regression/README.md and README.md point to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…#69)

The runner classifies the claims and gives the verdict, so it is now
among the recorded checkers and runners in regression/pack.json, at
90acbac: the version both retained runs used (unchanged since). An edit
to it without a reviewed update of pack.json now fails the retained tier.
13 checkers and runners are recorded, where there were 12.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
reference/m4-regression-pack/retained-only/: --retained-only at 23016c4
passes all 13 retained checks, evidence.checkers over the 13 recorded
checkers and runners, tools/regression-pack.py among them. fail-loudly.txt
gains a fourth case: an edit to tools/regression-pack.py stops the pack
(RETAINED EVIDENCE INCONSISTENT, exit 1). No campaign run was repeated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M4: the accepted M1–M3 regression pack, with coverage, provenance and blind spots
regression/baselines/lineages.json establishes the accepted Windows
baseline as lineage windows-retail, version v1 (the current Georgia and
Lebanon baselines, unchanged, stored under regression/baselines/
windows-retail/v1/; regression/windows/ is its working copy). Versions
are never overwritten or removed; earlier git history is not reconstructed.

tools/baseline-governance.py:
- propose: a red pack report becomes a change record in
  regression/changes/ with every failed claim and check UNCLASSIFIED;
- each delta is classified, with evidence, as one of AGENT.md's four
  classes; no other class is accepted;
- regression: REGRESSION - REJECTED, never advances a baseline;
- approved-defect-correction / intentional-semantic-change: advance only
  with a formal APPROVED GitHub review by fd-starscream-bot of a commit
  holding this exact record (bound by its digest); adopt adds a version,
  keeps the old one and materialises the new current version;
- expected-input-world-variation: never replaces the equivalent-input
  baseline; approved, it founds a separate lineage;
- check: lineages, versions, working copies and records consistent.

regression/changes/2026-09-27-s3-mutant.json classifies every failure of
#69's S3 mutant report as a regression (evidence: #68's attribution, #62's
S3 decision); adopt refuses it and windows-retail v1 stays current.

regression/fixtures/governance/ (test only, not an EECH baseline) and
tools/baseline-governance-selftest.py exercise an approved v1 -> v2
transition with a local test approver, and detect or refuse ten broken
variants (11 of 11 as expected).

tools/regress-windows.ps1 refuses -Update and no longer writes a missing
baseline. tools/regression-pack.py checks the lineage, runs the fixture
self-test, names the current accepted baseline, and gives each failure its
disposition without turning it into a pass. regression/pack.json records
the 15 checkers and runners.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/governance/REVIEW.md gains "Baseline changes": propose, classify every
delta with evidence in AGENT.md's four classes, review, then adopt with a
verified approval reference; versions are never overwritten or removed; a
disposition never makes a failing check pass. BASELINE_RECORD.md points
its "Differences and acceptance" to the checked change records for the
campaign baselines, and regression/README.md replaces -Update with the
governed path.

tools/baseline-governance.py reads a PR's reviews as one UTF-8 JSON line
each (gh api --jq): on Windows the console code page failed on a review
body's UTF-8 text. Exercised read-only on #69's real reviews.
regression/pack.json records its new hash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m4-baseline-governance.md answers: when accepted behaviour changes,
can we tell why, prevent an unapproved re-baseline, preserve what was
accepted, and identify the authoritative baseline? Yes, for the Windows
campaign baselines: windows-retail v1 is current and unchanged.

reference/m4-baseline-governance/, at 61b53bd:
- s3-regression.txt: #69's red S3 report proposed (all UNCLASSIFIED),
  classified as REGRESSION - REJECTED (all 23 failures), adopt refused,
  lineage and working baselines byte-unchanged;
- update-refused.txt: regress-windows.ps1 -Update refused;
- fixture-selftest.txt: the test-only approved v1 -> v2 transition and ten
  broken variants, 11 of 11 as expected;
- s3-mutant/: the full pack on the mutant: behaviour identical to #69,
  every failure with its disposition, verdict FAIL - CLASSIFIED AS
  REGRESSION - REJECTED, exit 1;
- accepted/: the full pack on the accepted build: PASS, 40 checks,
  against windows-retail v1;
- check.txt: lineages and records consistent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M4: accepted baselines change only by a classified, approved, non-destructive record
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants