Repository navigation
Native EECH campaign: eech_dc on the full engine, Windows build, retail campaigns and regression - #59
Open
fd-shockwave-bot wants to merge 77 commits into
Open
fd-shockwave-bot wants to merge 77 commits into
fd-shockwave-bot wants to merge 77 commits into
Conversation
Adds a Cargo workspace (eech-campaign/) that compiles the original EECH campaign C for the host target and hides it behind a safe Rust API: - eech-sys: build.rs ports the eech-core-ts C reference extraction spec (whole original TUs, verbatim extracts with #line provenance, a reduced project.h) and compiles it natively with the cc crate. Two exact-text patches with provenance: en_creat.c's va_list reinterpretation (64-bit blocker) and ks_updt.c's function-local static. Host layer in csrc/: dispatch tables, environment, entries with setjmp aborts and the round-toward-zero FP environment, entity identity (index, generation). - eech-campaign: public API (Campaign, CampaignConfig, World, events, snapshot), no unsafe; conformance replay of the eech-core-ts scenario language through the same kernel and World boundary. - corpus: 1401 eech-core-ts scenarios and 3000 float operations exported with the output of the 32-bit C reference and the TSTL verdict; the x86-64 native module reproduces all of them byte for byte. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…ntory - Campaign::step drives EECH's update loop: keysites (ks_updt.c) consume supplies and stock cargo, the force's low-on-supplies response builds a supply mission, the airbase's assignment pass selects a group and stops at the slice boundary (assign_primary_task_to_group) with the group and task ids. Missing world answers, unported rows and panics fail loudly and poison the campaign; ids carry their campaign instance. - eech-harness: JSON scenario -> HeadlessWorld -> Campaign -> JSON result; four recorded semantic scenarios. - global-state.txt classifies every writable symbol of the kernel (checked against nm); resets for the statics found (suitability array via EECH's deinitialiser, task.c and ks_int.c load-time values); database digest test; B1 probe test; diagnostics name EECH ordinals. - Build variants: EECH_C_OPT, EECH_FPU_ROUNDING=nearest (investigation), EECH_ORIGINAL_STACK_ATTRIBUTES=1 (i686 only). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
- The module builds for x86_64-pc-windows-gnu (LLP64, the DCS platform) and passes the whole suite under Wine, the corpus included. highlevl.h's WIN32 release ai_log only preprocesses under MSVC: GCC-family Windows builds take EECH's own non-WIN32 definition (-UWIN32, finding T1). - The NULL-dereference outcome of a dedicated replay process is reported on Windows through a vectored exception handler. - cc no longer passes -w (it silenced the -Werror= promotions); our own sources build with -Wall -Werror, which found %d with ptrdiff_t indices. - CI: x86-64, i686, i686 with the original va_list idiom, -O2, Windows cross under Wine; rustfmt and clippy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
README and docs/: feasibility report (verdict, milestone evidence, risks, recommendation), 64-bit classification (B1 blocker, T1-T4 portability, H1-H3 legacy), floating point (RTZ entries, results matrix across compilers/optimisation/platforms, rounding-mode sensitivity), global state (singleton, sequential, inventory classes, G1-G6), patches and compile definitions, ports derived from call sites, source closure and vendoring, conformance (three-way corpus), DCS constraints. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…affold
- eech-dc: cdylib (mlua 0.10, lua51 + module, modelled on flying-dice/
dcs-studio's bridge crate): require ("eech_dc") gives eech.new / step /
snapshot / close, an embedded Lua driver (eech.start: model-time pump,
world queries, protected callbacks, teardown at mission end) and
coordinate helpers. No panic paths (restriction lints denied); every
failure is a Lua error. Tests run the module in a real Lua 5.1 state.
- eech-world: the headless harness host. It links Lua 5.1, creates the
state a simulator creates (native modules allowed) and loads the DC DLL
from the build directory with require. Tacview ACMI 2.2 writer using
EECH tacview.c's geodesy.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
The maintained Windows build's 1,225 C files (the aphavoc and modules Visual Studio projects), compiled on the WIN32 code path against compat/ headers that declare the Windows SDK and DirectX names EECH uses. - compat/: Win32 types (LLP64 widths), DirectX value types and constants, opaque COM interfaces, and prototypes for every Win32/CRT call so no pointer result passes through an implicit int declaration. - csrc/: the headless platform layer. Win32 over POSIX (files, mappings, Windows paths resolved case-insensitively, text-mode reads), simulated time, debug_fatal unwinding to the entry point, null DirectX COM calls, and replacements for the 8 backends that need a real Windows machine (WinMain, debug windows, joysticks, GDI fonts, DirectDraw, Direct3D, DirectPlay, GUIDs). - build/patches.rs: 3 source patches (two MSVC-only macros, one union with repeated member names) plus the spike's P1 va_list marshaller. Implicit function declarations, implicit int and int conversions are build errors. The result links with no undefined symbols. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
eech-engine-sys: - csrc/eech_engine.c: boot and frame entry points following EECH's own dedicated-server path (WinMain, application_main, brief_ and full_initialise_game, the game-initialisation phases, flight () up to its loop); a frame is one iteration of flight ()'s loop without drawing. debug_fatal unwinds to the entry point; FPU round-toward-zero per entry. - csrc/eech_synth3d.c: writes a 3D object database in EECH's formats (bininfo.bin, 3dobjs.*, 3dobjdb.bin, displace.bin, stars.bin) with the engine's compiled name tables. The retail database is not in the repo. - memory-backed Direct3D surfaces and video screen, DirectDraw init. - 1x1 placeholders for missing UI artwork (.psd, .bmp, .tga), logged. - patch B2: UI attribute pointers read as int (ui_attrs.c, 9 sites). eech-map: Luxembourg from OSM + SRTM: terrain (ffp/sec/rgb), road network (ROADS.dat/.nde/.wp), side map, population placement, campaign script, popname.dat, bridge.pop, mapinfo.txt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…acview
- eech-engine-sys: airport and FARP scenes in the synthetic 3D database
(landing, takeoff and holding routes for fixed wing, helicopters and
vehicles, as routegen.c reads them; hangars and towers as scene links);
eech_observe.c reports aircraft, vehicles, air defence, ships, infantry,
weapons in flight and keysites the way EECH's own Tacview logger
classifies them; clock.
- eech-engine: safe API (boot once per process, frame, objects, clock;
debug_fatal becomes EngineError::Fatal and poisons the engine).
- eech-dc: the Lua module now contains the whole engine:
require ('eech_dc') -> boot{...}, engine:frame(ms), engine:objects(),
engine:clock(), write_3d_database(dir).
- eech-world: the host binary. Embeds Lua 5.1, lets the script require the
DLL, runs lua/campaign.lua, records what it observes to Tacview ACMI 2.2
(positions only when changed; destroyed events; keysite captures).
- eech-map: the campaign script now has forces, reserves, task generation,
division numbers, frontline forces, airbase groups; FARPs and SAM sites
in the population file.
A one-hour Luxembourg campaign runs headless without faults.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…orted once - eech_engine_prepare_installation (Lua: dc.prepare_installation): writes the generated part of an installation, the 3D database, textures.bin (texture names generated at build time from modules/3d/textname.h) and brief_en.dat (one briefing per task type). The retail files are not in the repository. - eech-map copies the EECH data the repository carries (setup/common/data: formation databases, language, suspension) into the installation. - EECH_TRACE_FILES=1 logs every file the engine opens or misses. - keysites are on the update list too; observe reports them once. - campaign.lua logs the tasks groups are on, per side. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…il for absent fields Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…istic guard EECH aims through WEAPON_SYSTEM_HEADING/PITCH sub-objects of the launcher's 3D scene; without them WEAPON_AND_TARGET_VECTORS_VALID is never set and no AI unit fires. The synthetic 3D database now gives every aircraft and vehicle scene a heading/pitch/muzzle/weapon tree sized from the weapon_config_database packages its type can carry. Weapon weights, drag and motor power come only from the GWUT table; the map tool now installs GWUT1162.CSV (a missile launched without it has a NaN velocity). Patch W1 guards get_ballistic_pitch_deflection against |height| > range at point blank (NaN pitch, INT_MIN table index). Also: object task state in the observation API, more airbase groups and pads (minimum idle group counts), wrecks and weapons in the Lua summary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Linked hangars carry REGEN_FIXED_WING/HELICOPTER/ROUTED_VEHICLE sub-objects, so EECH rebuilds losses from the force's hardware reserves. Without them air tasking stops once losses bring each group type down to the minimum idle count assign.c holds back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…uncher's side Without EXPLOS.CSV EECH runs on its compiled explosion database, which declares components it never initialises (XSMALL_HE: 5 declared, 3 set); a fresh installation then faults on the first small explosion. eech-map copies EXPLOS.CSV, METASMOK.CSV and SMOKES.CSV as an installation ships them. Observed weapons now carry their launcher's side, so Tacview colours them by coalition. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
A REPAIR or SUPPLY task never starts from the keysite it serves, so a side whose only airbase is struck out of action is paralysed for the rest of the campaign. eech-map now gives each side two airbases: the two largest named aerodromes at least 10 km apart, or one synthesised at the side's town farthest from its other airbase, placed off water (EECH turns a keysite on water into an anchorage). Hardware reserves are sized for days of war. engine:diagnostics() (eech_engine_diagnostics) logs each force's reserves and regen queues, and each keysite's state, supplies, regen sites and air groups; campaign.lua calls it every diagnostics_every simulated seconds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Troop insertion reads the helicopter's TROOP_TAKEOFF_ROUTE and TROOP_LANDING_ROUTE (under WAYPOINT_ROUTES) to put infantry down and pick it up; without them the first troop insertion is fatal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
recordings/luxembourg-6h.zip.acmi: six simulated hours of the dynamic campaign from the harness. Red's troop insertions capture both blue airbases; blue's air force dies out without a base to regenerate from. EECH reseeds its random numbers from the system clock when the session starts; the simulated clock's epoch now follows the boot seed (seed 1 keeps epoch 0), so the seed selects the run. The host writes object removals in id order, so the ACMI no longer depends on hash-map order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
eech-map takes a map name (or reads it from the extract's file name): Georgia covers 40.4-46.7 E, 41.0-43.6 N (256 x 142 sectors; EECH's campaign map holds at most 128 campaign sectors a side), front at 44.0 E through Shida Kartli, airbases Kutaisi and Senaki (blue), Vaziani and Marneuli (red), tertiary roads, English names where the local script is not Latin, the Black Sea as sea terrain. campaign.lua takes scenario=luxembourg|georgia. X1: EECH's GNU convert_float_to_int is 'fistp (%1)', the 16-bit store in AT&T syntax: values from 32,768 up saturated, so every world coordinate past 32.7 km fell into the wrong sector. Convert in C (truncation, the rounding EECH sets). The Luxembourg recording predates this fix and is being re-recorded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…e resupply fix Airbases and FARPs only consume ammo and fuel; factories and refineries make it and SUPPLY missions fly it. eech-map places a factory and a refinery per side on the largest OSM industrial areas behind the front, as key templates in a version 2 population file (popread.c creates keysites from key templates only in version 2 files). S1: an airbase low on supplies picked itself as its closest supplier airbase (get_closest_keysite without exclusion), so no airbase was ever resupplied; exclude the requester. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
The airbases sit 70-160 km behind the front, beyond most helicopter tasking; FARPs are populated only by FRONTLINE_FORCES (two groups each), so Georgia gets four a side. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
recordings/georgia-12h.zip.acmi: twelve simulated hours of the Georgia campaign; recordings/luxembourg-6h.zip.acmi replaces the recording made before X1 (its sector lookups past 32.7 km were wrong). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…rage tool - prepare_installation keeps a retail 3D database (textures.pal present) - tools/retail-map3.sh assembles an install root for the retail Georgia campaign (map3, Caspian Black Gold) from retail data kept outside the repository; the retail script's 30-minute END_CAMPAIGN FAIL trigger is dropped (the original is kept as .retail) - campaign.lua scenario georgia_retail; the Tacview recorder takes an affine map projection: map3 is stretched 1.22 east-west, fitted by correlating the retail terrain heights with SRTM (r = 0.978), where EECH's own metric origin is off by up to 50 km - tools/offsite.sh: rclone against one Google Drive folder for install data and run outputs shared between sessions Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
The Steam Apache vs Havoc 3dobjdb.bin holds 2,761 of the 3,026 scenes this source defines; a modern install adds the rest from the community objects (setup/cohokum/3ddata/objects), which retail-map3.sh now copies in. - C1: a scene's collision object 0 (the null object) means none on the 3dobjdb.bin path too (the .EES path already maps it to -1) - T1, T2: headless, a texture camouflage mismatch or a missing texture animation between the community objects and the retail texture set is logged instead of fatal (nothing is drawn) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
…rage Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4ztFDZvVEsA6Wj62yZ7wY
- recordings/georgia-retail-2h.zip.acmi: 2 simulated hours of the retail map3 campaign (FARPs 18, 7 and 3 change hands), with a README section - tools/acmi-summary.py: losses, weapons, captures and task mix from a recording and its run log - Luxembourg removed: its recording, eech-map spec, campaign.lua scenario and the README/engine.md sections; examples now use map16/georgia.chc - handover: building locally with Docker (rust:bookworm, GCC 12) and the local retail data sources Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The M1 review (#62, docs/corrections.md) rejected S2 for the M1 candidate: its stated need is not reproduced once S3 is present, and S3 explains the diagnostic that motivated it. Without S2, response_to_force_low_on_supplies compiles as written: when the chosen supplier holds no crate of the type asked for, no SUPPLY task is created and the keysite asks again later. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
regression/windows/lebanon_retail.json is re-recorded at 0288ece (S2 removed). Georgia is IDENTICAL to its baseline (retail Georgia has no factories or refineries, so S2 never acted there); Lebanon keeps all 39 campaign expectations with 23 metrics outside the old baseline's tolerance: fewer SUPPLY tasks from different suppliers (blue 11 -> 10, red 52 -> 49), then a diverging war. The new Lebanon metrics are byte-identical to the #62 review's independent no-S2 run, so the new baseline was produced twice. docs/corrections.md records the removal and the before/after, docs/engine.md drops S2 from the patch table, docs/m1-baseline.md gives the updated candidate (eech_dc.dll da72f128...) and marks S2 as removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
Review feedback on #63: docs/m1-baseline.md still presented a8661ea as the candidate and the correction review as an open blocker. It now: - identifies 0288ece (S2 removed) as the candidate, with eech_dc.dll da72f128... and the unchanged host and Lua artifacts; - keeps a8661ea as provenance only, with a table of which evidence was produced where (build identity, lifecycle and regression rerun at 0288ece; rebuild check, sensitivity and kernel at a8661ea); - gives the regression figures at 0288ece and the new Lebanon baseline; - states what is established for the post-S2 candidate, including the completed correction review (#62, docs/corrections.md); - leaves only conditional older-doc reconciliation and the #19 acceptance review as remaining M1 work. Documentation only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M1: remove S2, restore the original EECH supplier rule; new Lebanon baseline (#19)
- lua/observations.lua (campaign.lua observe=, observe_every=): structured observations from the samples the Tacview recorder receives: appear, retype, alive, side, usable and gone events every sample, and full snapshots every observe_every seconds and at each metrics checkpoint. campaign.lua computes the checkpoint condition once for both; the metrics are unchanged. - tools/input-manifest.py: SHA-256 of every input file of a root, with a combined hash; the files the engine writes are listed apart. - tools/observation-check.py: replays the observations through the recorder's rules and requires the Tacview recording's declarations, removals and Destroyed events frame by frame, its positions at every snapshot within the recorder's change threshold (with its own implementation of the map projection), and losses, captures and keysite state to agree with the metrics; reports implausible jumps. - tools/observation-check-selftest.py: the check passes unmodified evidence and fails five deliberate discrepancies. - tools/reference-windows.ps1: the reference scenario, run twice and checked. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m2-reference.md records the reference scenario and answers whether EECH World's observations of it can be trusted as physical-world evidence. The M1 candidate's own binaries (0288ece) with a1941f2's scripts ran retail Lebanon for 3 simulated hours twice, from two separately assembled roots with identical inputs (1,179 files, 121a7890...): the Tacview recording, the structured observations and the metrics are byte-identical. tools/observation-check.py passes on both: identity and lifecycle exact in all 10,800 frames, 308,368 snapshot positions within the recorder's threshold, losses identical three ways, keysite state identical to the metrics. The metrics equal the M1 Lebanon baseline, so observing does not perturb the campaign; the M1 regression and lifecycle checks are unchanged. Fidelity limits found: EECH reuses entity indices, merging some weapon tracks (23 jumps, 12 launcher changes seen); Tacview attitude can be stale and its keysites carry no capture or supply state; EECH's session clock lags host time by 29.6 s in 3 h (float accumulation explains 4.6 s) and Tacview's reference time is a fixed label. reference/lebanon-3h/ keeps the recording (30.3 MB zipped), the observations (8.2 MB), the metrics, the input manifest and the check report. The checker now breaks plausibility flags down by type and counts weapon indices held by another launcher; the runner records the baseline comparison's verdict. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M2: a trustworthy EECH World reference scenario and its observation evidence (#23)
- lua/observations.lua, campaign.lua observe_supply=1 (off by default, so runs without it write the M2 observations unchanged): task events when a unit's group, primary task or operational state changes; supply events when a keysite's ammo or fuel moves by 0.5 or more in a sample; track events every sample for aircraft whose group is on TASK_SUPPLY. - campaign.lua loads observations.lua before the engine boots: the boot changes the process's working directory (csrc/eech_engine.c chdir to <root>/cohokum), so a relative script path no longer resolved after it. - tools/supply-chain-check.py: every delivery (a keysite level going to 100) attributed to the SUPPLY group over the receiver at that sample, with its assignment, the one-crate pick-up at the supplier with the group over it, the continuous 1-second flight between, the end of the task, and what the campaign then does with the delivered stock; the delivery count must agree with the metrics' resupplied counts. - tools/supply-chain-check-selftest.py: removing the assignment, the pick-up, part of the flight or the aircraft at the drop-off breaks the chain. - tools/reference-windows.ps1 -ObserveSupply: runs the chain check on both runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…ained
docs/m3-resupply-loop.md follows one real resupply chain through the
public path (eech-world.exe -> campaign.lua -> require ("eech_dc") ->
engine:objects ()): Il-76 group 88235 "Jester" is assigned TASK_SUPPLY
(1,160.77 s), flies continuously to Halat, where Halat's fuel drops one
crate as it passes 85 m overhead (1,555.67), and on 30 km to Beirut
International, whose fuel goes 10 -> 100 as it passes 43 m overhead
(1,904.58). The campaign then uses the fuel: groups draw on it, another
SUPPLY task picks up from it, and once it is below the request threshold
another SUPPLY task delivers to Beirut again (3,324.68).
Across the run all 32 deliveries are attributed to the SUPPLY group over
the receiver; 31 chains are complete, and the one that is not is a crate
carried over from an earlier task: a drop-off to a group leaves the crate
aboard (mb_msgs.c), recorded as a finding. The deliveries equal the
metrics' resupplied counts; two runs are byte-identical; the metrics and
the recording equal the M1/M2 evidence, and the M2 settings reproduce
M2's files byte for byte.
reference/m3-resupply-loop/ keeps the observations (10.5 MB), the
featured chain's events and the check reports. The runner's title is
generic; the checker's header states the carried-over crate case.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M3: one campaign -> world -> campaign resupply loop through the public path (#27)
…next docs/m3-supply-loss.md follows the naturally occurring loss in the accepted reference run's retained observations (no new run, no runtime change): red Mi-17 87568 of SUPPLY group 87567 "Wolfpack" dies at 1,584.66 s while Performing Task, 32 km from Beirut, as two blue MIM-72G Chaparrals vanish at it. The group falls from two aircraft to one with no replacement; the survivor turns back at that moment and its task ends at Beirut with no keysite delivery; the campaign re-tasks the one-aircraft group at 3,619.24 s and it delivers fuel to Power Station 3 at 4,013.67 s. No red keysite was requesting fuel when the task was assigned, so its receiver was a group (by elimination and the code), whose supply is not reported: the lost cargo and the receiver's shortfall are not observable. Facts are classified as observed, inferred, not observable and code. tools/loss-chain-check.py and its self-test; reference/m3-supply-loss/ keeps the report and the records behind each link. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M3: a SUPPLY aircraft destroyed in flight, and the campaign's response (#27)
tools/build-mutant-windows.sh <patch id | control> target/mutants/<dir> builds the Windows engine from a temporary git worktree of the clean HEAD (under target/mutants/, removed afterwards), with one patch entry removed from build/patches.rs in that worktree only (checked to be that entry and nothing else), in its own cargo target volume, into target/mutants/ only, stamped MUTANT or CONTROL with the source diff and the patch markers of the staged sources it compiled. The working tree and the normal build (tools/build-windows.sh, target/windows, the normal volume) are never touched. tools/regress-windows.ps1 -Out: where a run's outputs go (default unchanged), so a control and a mutant can run side by side. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…tes leading-slash patterns); check the checkout Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
tools/mutation-windows.ps1 builds a control and a test-only mutant with one patch removed (tools/build-mutant-windows.sh, the same clean commit), checks that they differ only in that patch (the same commit, image, toolchain, lua.dll and scripts; a one-entry source diff; the compiled patch markers differ by that patch alone), gives each mutant run an identical copy of the inputs, runs tools/regress-windows.ps1 -Exact on both side by side, and reports DETECTED when the control passes and the mutant fails, with the metrics that changed (tools/metrics-difference.py). -RepeatMutant runs the mutant twice and requires identical metrics. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…hers git worktree prune also removed another session's worktree record, whose recorded path Windows git took as stale; the builder now removes only the temporary worktree it created. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m4-s3-mutation.md: tools/mutation-windows.ps1 built a control and a test-only mutant with S3 removed from the same commit (56b164d), each in a temporary worktree, and showed they differ in that patch alone (one-entry source diff; compiled patch markers differ only by S3; the same image, toolchain, lua.dll and scripts; identical input copies). The accepted regression (tools/regress-windows.ps1 -Exact) passes the control (30/30, 39/39, IDENTICAL) and fails the mutant in both campaigns: Lebanon 38/39 ("airbases resupplied with ammo (0)") and not identical, Georgia not identical; the tolerance comparison flags both too. Two mutant runs are byte-identical. The metrics name the behaviour: ammo deliveries vanish (Lebanon airbases 3 -> 0, military bases 6 -> 0; Georgia airbases 6 -> 1) while fuel deliveries rise, the signature of ammo tasks loading fuel. reference/m4-s3-mutation/ keeps the summary, the difference, both builds' identities, the diff, the input manifests and each run's regression output. regression/README.md and README.md point to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M4: the accepted regression detects a reintroduced campaign defect (S3 mutant)
regression/pack.json lists every materially accepted M1-M3 behaviour with where it was accepted, its path, the check that protects it and how (direct, baseline-sensitive, evidence only, not covered), and the blind spots; plus the accepted candidate and scripts, the scenarios, the input manifests, and the hash and revision of every retained artifact, checker and runner the judgement depends on. tools/regression-pack.py runs the existing checks and reports each claim: - retained tier (no retail data, about a minute): artifact, checker and manifest hashes; the checkers reproduce the retained reports from the retained evidence; their self-tests still detect their discrepancies; the S3 mutant's retained metrics are still rejected, and Georgia's expectations still miss it (the documented blind spot). Any failure stops the pack. - runtime tier (a build and the retail roots): the roots against their manifests, lifecycle-windows.ps1, then regress-windows.ps1 -Exact and reference-windows.ps1 -ObserveSupply in parallel, loss-chain-check.py, and the fresh evidence against the retained. EVIDENCE ONLY and NOT COVERED come from the inventory and never pass. tools/lifecycle-windows.ps1 gains -Out (default unchanged), so that the pack's runs do not share an output directory. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…served run reference-windows.ps1 judges "observing does not perturb" by comparing the observed run with the Lebanon baseline. On a changed build that comparison fails whether or not observation perturbs anything. The pack now asserts M2.unperturbed by comparing the observed run's metrics with the same build's unobserved regression run (regress-compare.py --exact). The runner's own baseline comparison becomes reference.baseline, part of the baseline-sensitive M2.reference_reproduced. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m4-regression-pack.md answers what accepted M1-M3 behaviour the pack would detect (17 claims, directly protected), what would merely look different (12, baseline-sensitive) and what could regress undetected (8 evidence only, 8 not covered), with the blind spots and provenance. reference/m4-regression-pack/ keeps both runs (pack revision 90acbac): - accepted: the candidate's binaries (0288ece) with the M3 scripts: PASS, all 38 checks; every accepted result reproduced. - s3-mutant: #68's mutant, not rebuilt: FAIL. M1.lebanon.resupply and M1.corrections.S3 fail as directly protected ("airbases resupplied with ammo (0)"); every baseline-sensitive claim differs, including S1 and S2, which did not change; M3.resupply_loop stays green although all 30 deliveries were fuel, a blind spot now demonstrated. - fail-loudly.txt: an altered artifact, a missing one and an edited checker each stop the pack before any judgement. regression/README.md and README.md point to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
…#69) The runner classifies the claims and gives the verdict, so it is now among the recorded checkers and runners in regression/pack.json, at 90acbac: the version both retained runs used (unchanged since). An edit to it without a reviewed update of pack.json now fails the retained tier. 13 checkers and runners are recorded, where there were 12. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
reference/m4-regression-pack/retained-only/: --retained-only at 23016c4 passes all 13 retained checks, evidence.checkers over the 13 recorded checkers and runners, tools/regression-pack.py among them. fail-loudly.txt gains a fourth case: an edit to tools/regression-pack.py stops the pack (RETAINED EVIDENCE INCONSISTENT, exit 1). No campaign run was repeated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M4: the accepted M1–M3 regression pack, with coverage, provenance and blind spots
regression/baselines/lineages.json establishes the accepted Windows baseline as lineage windows-retail, version v1 (the current Georgia and Lebanon baselines, unchanged, stored under regression/baselines/ windows-retail/v1/; regression/windows/ is its working copy). Versions are never overwritten or removed; earlier git history is not reconstructed. tools/baseline-governance.py: - propose: a red pack report becomes a change record in regression/changes/ with every failed claim and check UNCLASSIFIED; - each delta is classified, with evidence, as one of AGENT.md's four classes; no other class is accepted; - regression: REGRESSION - REJECTED, never advances a baseline; - approved-defect-correction / intentional-semantic-change: advance only with a formal APPROVED GitHub review by fd-starscream-bot of a commit holding this exact record (bound by its digest); adopt adds a version, keeps the old one and materialises the new current version; - expected-input-world-variation: never replaces the equivalent-input baseline; approved, it founds a separate lineage; - check: lineages, versions, working copies and records consistent. regression/changes/2026-09-27-s3-mutant.json classifies every failure of #69's S3 mutant report as a regression (evidence: #68's attribution, #62's S3 decision); adopt refuses it and windows-retail v1 stays current. regression/fixtures/governance/ (test only, not an EECH baseline) and tools/baseline-governance-selftest.py exercise an approved v1 -> v2 transition with a local test approver, and detect or refuse ten broken variants (11 of 11 as expected). tools/regress-windows.ps1 refuses -Update and no longer writes a missing baseline. tools/regression-pack.py checks the lineage, runs the fixture self-test, names the current accepted baseline, and gives each failure its disposition without turning it into a pass. regression/pack.json records the 15 checkers and runners. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/governance/REVIEW.md gains "Baseline changes": propose, classify every delta with evidence in AGENT.md's four classes, review, then adopt with a verified approval reference; versions are never overwritten or removed; a disposition never makes a failing check pass. BASELINE_RECORD.md points its "Differences and acceptance" to the checked change records for the campaign baselines, and regression/README.md replaces -Update with the governed path. tools/baseline-governance.py reads a PR's reviews as one UTF-8 JSON line each (gh api --jq): on Windows the console code page failed on a review body's UTF-8 text. Exercised read-only on #69's real reviews. regression/pack.json records its new hash. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
docs/m4-baseline-governance.md answers: when accepted behaviour changes, can we tell why, prevent an unapproved re-baseline, preserve what was accepted, and identify the authoritative baseline? Yes, for the Windows campaign baselines: windows-retail v1 is current and unchanged. reference/m4-baseline-governance/, at 61b53bd: - s3-regression.txt: #69's red S3 report proposed (all UNCLASSIFIED), classified as REGRESSION - REJECTED (all 23 failures), adopt refused, lineage and working baselines byte-unchanged; - update-refused.txt: regress-windows.ps1 -Update refused; - fixture-selftest.txt: the test-only approved v1 -> v2 transition and ten broken variants, 11 of 11 as expected; - s3-mutant/: the full pack on the mutant: behaviour identical to #69, every failure with its disposition, verdict FAIL - CLASSIFIED AS REGRESSION - REJECTED, exit 1; - accepted/: the full pack on the accepted build: PASS, 40 checks, against windows-retail v1; - check.txt: lineages and records consistent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK
M4: accepted baselines change only by a classified, approved, non-destructive record
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR brings the native EECH campaign work on
claude/eech-campaign-native-spike-tjkj6gintomaster, as the base for the next phase of building.masteris an ancestor of the branch, so the merge is clean: 40 commits, head5f07697b.The project owner has accepted the branch as it stands. Most of these commits were pushed directly to the integration branch without PR review, and that is recorded here rather than reworked.
What's included (
eech-campaign/)eech_dc: all of EECH compiled headless as a Lua 5.1 module, plus the smaller campaign kernel behind a Rust API, with its C-reference corpus.eech-world: a Rust host that embeds Lua, runs the campaign throughrequire ("eech_dc"), and records Tacview.eech_dc.dll,eech-world.exeandlua.dllbuilt with MinGW-w64 (tools/build-windows.sh,tools/Dockerfile), with a crash reporter.tools/regress-windows.ps1, which checks campaign expectations and compares againstregression/windows/baselines.docs/engine.md.AGENT.md,PROJECT.mdanddocs/governance/, merged into the branch through docs(governance): establish project charter and baseline-first agent policy (#51) #53.Validation (Windows, at
5f07697b)tools/regress-windows.ps1 -Georgia -Lebanonpasses the expectations (Georgia 30/30, Lebanon 39/39), and both campaigns are IDENTICAL to their baselines. Repeat runs are identical.Known limitations
regression/*.jsonwere last recorded at519e35da, before S2, S3, N1–N3, H1 and U1. Windows is the maintained baseline..github/workflows/eech-campaign.ymlpredates the whole-engine build. It uses the runner's default GCC and tests the whole workspace for i686 and under Wine, so its engine jobs may fail until the workflow is reworked.tools/retail-*.sh), so it can't run in hosted CI.eech_dc. That is transitional; the world moving into the host is later work.Review
Needs a formal approval from @fd-starscream-bot before merge (AGENT.md).
🤖 Generated with Claude Code
https://claude.ai/code/session_01C64rJFJLY8BTJbC1QimZDK