A development archive

HOW IN THE MOUNTAINS WAS BUILT

A chronological, honest record of the work — every major rebuild, the player complaint that started it, the number that proved it broken, and the number that proved it fixed. Click any thumbnail to open the full illustrated report.

This game was built by an AI agent fleet, working to one standing bar: that a skeptical soldier reading a field manual would recognise the behaviour as real — and ask "an AI built this?". Software is usually shipped with the scaffolding hidden. We do the opposite. This archive ships inside the game, beside the Field Manual, because the way it was made — measure first, fan out specialists, then attack your own answer — is itself part of the thing. Read it as a logbook: the valley pushed back, and each chapter is what pushing back looked like.

9days · Jun 3–26
37chapters
24full reports
26tracked issues
20+test harnesses
1deterministic clock
How to read a chapter
Every entry follows the same arc: the problem (usually a player's own words), the shape of the work (how the fix was approached — that's the craft), the before→after numbers that make "fixed" provable rather than a vibe, and a link to the full report where one exists. Red is the old number; green is the new one.

The five moves

Almost every chapter below is some combination of these. They are the project's working doctrine — the reason the fixes held instead of thrashing.

01

Metricize first

"Feels off" is not actionable until it's a number. Write a headless harness that reproduces the complaint, and baseline it before touching code.

02

Fan out specialists

For open problems, dispatch many independent agents on one axis each, judge their output against a rubric, then synthesize — never one draft iterated.

03

Build a fair oracle

To ask "did it really work?", compute ground truth independently (a BFS flood that obeys the mover's own rules) and act on the gap it exposes.

04

Adversarially verify

Your first answer is a hypothesis. A separate skeptical pass re-reads the code and tries to break the fix — it has caught regressions a green author-run missed.

05

Prove on held-out seeds

Tune on one set of procedural valleys, prove the win on a fresh set you never touched. A win only on the tuned set is a curve-fit until proven.

All the full reports

Each chapter below links its own write-up inline — but here is the complete set in one place. Every one opens the original illustrated report in full, exactly as it was written when the work shipped: the charts, the annotated screenshots (click any image to enlarge it), and — for the soundscape — playable before/after audio.

2026-06-03Squad movement rebuildFour layered bugs ended by one "navigate globally, steer locally" locomotion layer.Open full report → 2026-06-04The map art bible — 160 SVG assetsThe studio art bible: the design system, before/after, and all 160 assets rendered live.Open art bible → 2026-06-05The 30%-reached gapThe router was already good — fatigue, not pathfinding, was eating the march.Open full report → 2026-06-06Call-for-fire realismThe squad stopped dropping mortars on itself: densest-cluster aimpoint + danger-close gate.Open full report → 2026-06-06World-map scale auditThe soldiers were 36× too big — a render lie over an already-authentic sim. Full scorecard.Open full audit → 2026-06-06Civilian atmosphericsThe valley learned a daily rhythm — and to go quiet before an ambush.Open full report → 2026-06-06The deploy loading screenProfiling the 6.5-second frozen button, then staging the work behind real feedback.Open full report → 2026-06-06The 5× campaignFive file-disjoint waves — combat, audio, game-feel, COIN, scale — verified across deployments.Open full report → 2026-06-07Soundscape — 20× immersionFrom mono-and-silent to a living canyon: reverb, ambient bed, occlusion, 5-layer gunfire. Playable A/B audio.Open full report → 2026-06-07World scale, finishedThe audit's deferred half, shipped: COP 170→120 m, platoon 35→41, villages monolith→hamlet, honest gradient — adversarially verified, with the stall hunt the static pass missed.Open full report → 2026-06-08The command surface, re-cut461→1 measured UX defects: keyboard focus rings, a 12px legibility floor, promoted campaign meters, command toasts, an accessible modal + in-game controls. Five expert judges, unanimous.Open full report → 2026-06-08Soldier-scale realismHours-long progressive census + a dwell event-roll, an autonomous squad flank with bounding overwatch (0→100%), and one honest revert: a switchback-pathing fix measured against its own baseline, found to stall a squad, and backed out.Open full report → 2026-06-08The combat outpost, made realA six-specialist survey + judge rebuilt every aspect of the COP: a fire plan sited by a terrain LOS sweep (M2 on the 922 m avenue, Mk19 on 94%-dead ground), the HESCO bastion silhouette, night life-signs, a living garrison, and the watch — a wire assault that now demands a Final Protective Fire decision. copaudit clean ×9.Open full report → 2026-06-08Making the firefight feel realFive fixes across visual + sim + audio. One sub-tick interpolation made bullets travel and resurrected the muzzle flash (0→100% rendered). Suppression finally cuts a pinned man's fire rate (flat → −48%); temperament shapes bursts; ComBloc tracers burn green; the MG hammers (56→8 cracks/burst) — all within ±15% of the old balance, determinism intact.Open full report → 2026-06-09Reading the groundThe map quietly lied: half the walkable valley was stranded behind fake-cliff noise (reachable 48→61%), footpaths were a 5 m smudge, a 56° wall was drawn like a meadow, and cover was a smear decoupled from the rocks you could see. Four measured fixes — a steep-band reconnect (with one cost-floor revert: no win, −1.5× speed, +71% WIA), switchback foot-trails (a trench bug caught in a render), paths as scaled dirt lines + cliffs as sheer rock, and ~11,600 discrete cover objects that are drawn = sim. The combat-cover half was cut on the numbers (a light open-ground stamp dragged firefights, WIA +89%) — the honest deferral to a sub-cell model.Open full report → 2026-06-10Climbing the mountainA squad ordered to a peak OP walked a 3 km arc around the mountain instead of climbing it — the pathfinder's cost was isotropic, so a switchback never paid off, and an 8-direction grid couldn't even draw the traverse angle. The fix is an any-angle (Theta*) tactical planner with a signed-grade cost + a turn penalty, fired only for a real climb and bolted onto the proven router without changing a line of it: held-out detour ×3.58→×3.16, climbable-face OPs ×2.19→×1.31, village routing & terrain byte-identical, and the stall guard the two prior attempts failed now passes. The honest residual: OPs behind a true impassable massif still detour — real terrain, not a planning miss.Open full report → 2026-06-10Behind that rockA boulder on the open slope did nothing for the man behind it — combat read a 5 m averaged cover number, so hugging a rock scored as exposed as standing in the open. The obvious fix (more open-ground cover) had been reverted for grinding every firefight +89% WIA. The missing piece wasn't more cover, it was direction: a discrete object now stops only the round coming from the bearing it faces (a flanker still sees you), and a low rock hides a prone man far more than a standing one. The result is the opposite of a grind — US WIA −22% (held-out −36%), enemy still killed via the flank, 0 stranded — because directional cover resolves fights by maneuver, not attrition.Open full report → 2026-06-10Supplies that biteIssue 021 asked to couple the COP's fortification to combat — a hardened wire should blunt an assault. Metricizing first turned it on its head: the assault doesn't happen. A "complex attack" is a standoff from 340 m; the insurgents never close the wire (no assault behaviour exists), so hesco and claymores have no event to bite on — deferred with the measured reason. The shippable win was the adjacent item: the four supplies that drained daily to zero consequence now cost you — dead batteries blind the patrol at night (2.4× detection lost), dehydration halves fatigue recovery, no medical slows wound recovery — all bounded, all balance-neutral at full stock.Open full report → 2026-06-10Reading the sunThe land classification grew the same vegetation on a north slope and a south slope — no notion of aspect, the most visible fact of a real mountainside. An aspect term now makes forest cling to the shaded draw (59→64%) and scrub own the sunny spur (48→56%). Because vegetation IS cover and concealment, it had to be earned: a 3-strength balance sweep showed the firefight moves chaotically with the terrain, so it ships at 0.05 — the strength that clears the no-stall guard with KIA down on both seed sets, no aspect-caused strandings, and siting byte-identical. On by default; the terrain-dependent wounded count disclosed, not buried.Open full report → 2026-06-10Closing the open issuesOne pass over the whole backlog: four realism wins shipped active (a squad that switchbacks up a face, a soldier who uses that rock, supplies that finally bite, and vegetation that reads the sun) — the fourth earned by a 3-strength balance sweep + held-out A/B that found the one strength (0.05) which improves the sim with KIA down on both seed sets. The two remaining tickets already run their live fix; what's left is a documented terrain floor and a logged do-not. Six issues, six live fixes, zero regressions shipped. The thread: the most valuable artifact in half of these was a probe, not a feature.Open the overview → 2026-06-10Making the people realOne architecture, four places: the enemy cell gets the friendly squad's decide/execute split (ambush volley p90 11.2→0.6 s, a peel that converges while a rout disperses), men hesitate like men, a deterministic callout bus makes the brotherhood audible (man-down 100%, two-run identical) — and the village remembers: the elder physically walks out to the shura, a civilian death writes a named blood debt his household buries at first light, and children trail only the patrols a village trusts. Five new probes caught six real bugs before ship.Open full report → 2026-06-11The mixer and the whistleA per-category sound mixer (combat / ambience / radio / alerts, each with on/off + level, persisted) wired between the buses and master so user trim never fights the contact-duck — plus five audible improvements with oracle numbers: the incoming-shell whistle the valley was missing (2.26 s, deterministic off the fire mission's own countdown), real stereo width for the calm bed (4%→12.4% — the fix was loop offsets, not panning), thunder that crosses the valley, an adhan that finally sounds sung (5.3 Hz vibrato), and three ricochet families. Playable A/B audio throughout.Open full report → 2026-06-11Calibre voicesOne sentence of sim code — stamping which weapon produced each effect — unlocked the sound pass the synth itself had been asking for ("we have no per-weapon id on the Effect"). The .50s now hammer an octave below the 7.62 guns (DShK darker than the PKM by 772 Hz, measured), the RPG announces its launch before the blast, the M320 bloops, bolt guns cycle their bolts, pistols bark, and 60 vs 120 mm mortars grade by calibre (LF share 14.1→19.9%). Plus the sounds that were missing: reload clatter at −51.6 dB, and upgrades that make the IED heave (centroid 5233→3951 Hz), the incoming shell tear instead of sing, and the near miss complete its N-wave. 13/13 oracle assertions; new cues proven on 3 held-out seeds of organic combat. Playable A/B audio throughout.Open full report → 2026-06-13The valley is real — a WebGL terrain rebuildA first attempt was reverted for being "the old flat map with the sun switched on" (blit parity + a colour filter). This rebuild inverts it: the painted albedo is demoted to one input of five, and a WebGL2 fragment recomposes the surface per-pixel from the sim's own arrays — per-landcover materials + detail-normal raking, baked horizon AO, live-sun ridge shadows, a dark-silt flow-advected river, single-scatter aerial perspective, ACES + bloom — under a transparent 2D HUD (the legibility firewall), with a u_detailGain zoom-ramp that collapses to byte-faithful relief at strategic (Δ+0.2 luma). 9 green commits; a 4-critic adversarial pass caught a snow-white river + contrast-crushing haze (1/4 → fixed → fresh-eyes confirmed clears the bar). Terrain = full 10×; assets = grounded (full relight + FX deferred, issue 028).Open full report → 2026-06-26When the tests start defending the bugThis game is built test-first — but a test can quietly stop measuring whether the game is good and start measuring whether it is the same. An audit of all 84 harnesses found exactly one such blocker: a remembered "~8.58 WIA band" — a casualty number that was just the sim's own past output (the day aspect-vegetation shipped), not any real-world rate. It had flagged fewer wounded soldiers as a defect, narrowed a realism win, and reverted features. And it was policed finer than the harness's own noise: four identical-config draws ranged 2.67–9.42 WIA (σ≈2.5). The fix is a charter, not a patch — a gate may never assert the sim's own output, and the COIN win condition now gets the standing gate the firefight wrongly held.Open full report → 2026-07-02When "too squiggly" meant "too straight"Four realism complaints — stuck soldiers, squiggly paths, smooth terrain, a fake outpost — measured before a line of code changed, and two were the opposite of what they looked like. "Squiggly" paths were a near-beeline (ratio 1.12, 66% down the fall line); the wobble was 0.16 px at play zoom. "Getting stuck" wasn't the patrols at all — it was an in-combat freeze (one fighter pinned 678 s, 42% of insurgent contact time blocked → 0). The real wins, front by front: the uphill march re-anchored to FM 3-97.6 (2.45× too fast → doctrine-honest 0.98), foot-trails taught to contour by the USFS half-rule (walker's grade 0.215 → 0.114), the fbm valley walls re-bedded as banded strata (wall reversals 0 → 6–18/km, reachable valley 60.6 → 75.2%), and the KOP dressed in the clutter that is the realism (concertina, parapets, nets, tents, road clipped at the wire). Plus two commits that fixed the test, not the game: a null RNG draw swung the COIN gate ±19 pts, so it now scores the paired best seed. Honest residual: the outpost was dressed, not reshaped — the perfect-circle rebuild is the next campaign (issue 033).Open full report → 2026-07-03The point man waitsThe owner reported two things — the point man runs "unrealistically far forward," and squad members "get stuck on buildings in the villages." Metricized before touching code, and the named suspect was innocent: village qalat walls block 0.6 s of a patrol; the real grind is the outpost's own b-huts (11.6 s) and the terrain/HESCO wire (38.6 s). The actual defect is cause-agnostic — while a follower is genuinely stuck the point man runs at 0.51 m/s and is halted just 1% of the time, so the file strings to 355 m. Two guards meant to catch this were dead code (both keyed on a blockedTimer>6 the watchdog caps at 2 s — one had read 0.0 s on every seed and mis-reassured a prior investigation). Fix: a real halt gated on being blocked, not merely slow, and capped at 45 s per leg — so it fires in villages and on the wire yet stays byte-identical on the tactical-window gate (the issue-031 arbiter). Halted-while-wedged 1%→38%; peaks reined in 12/21→9/21 seeds; held-out clean, 0 new strandings. Honest residual: the ~29 m average lead is nine-man-file geometry, not a defect — and the biggest building grind (COP egress) is deferred to a muster-routing pass.Open full report → 2026-07-16The enemy gets a nameThe insurgency was a weather system — an abstract strength scalar and a memoryless tempo clock; fighters spawned from nothing and exfiltrated into nothing, so the game's own intel feed had nothing to be intelligence about. This wave makes it an organization: 3–5 persistent cells with named leaders who survive between fights, home ground in the draws, physical munitions caches that IED ambushes drain, a decaying patrol-heat memory of your habits — and an intel ladder (unknown → named → located → mapped) climbed above all by won-over villages giving their cell up, so FM 3-24's loop is a mechanic. Same day: relief-of-command stopped being an opening-days dice roll (careful tours censored ~50–60% → 0/3 on the relief tree; battalion now reads an attributed evidence file past a 5-day grace). Honest record kept: the build shipped a conservation bug (exfil +1 printed strength 64→80-cap in one hot day — the acceptance probe asserted the contract's own wrong words) — caught by the skeptical pass, fixed to exact roster conservation, probe 9/9. Combined COIN gate: all 8 PASS, best-pair spread 93 (highest ever) — and the mean's collapse to 25.7 is filed as issue 037, because the adaptive enemy now punishes the harness's fixed-route scripted commander.Open full report → metaHow the CLAUDE.md was built147 MB of transcripts distilled, via a five-workflow tournament, into the doctrine every agent reads first.Open the explainer →
Act I · June 3–5

Making the valley walkable & legible

Before anything else could matter, a squad had to be able to leave the gate, cross the ground, and arrive — and the player had to be able to see it happen.

Traced squad march around the COP, before vs after
The squad's actual traced path around the outpost — the diagram that made the bug obvious.
2026-06-03

The squad declared "on station" 300 m from the village

MovementTerrainUX

A player ordered a squad to the village just right of the gate. It hugged the perimeter wire, bumped it, and "accomplished its mission" without ever reaching the village. On the default seed korengal it stalled 152 m short and called itself on-station 300 m from the objective.

The shape: metricize first — a headless harness reproduced the failure across seven adversarial seeds, four root causes were confirmed against the live code with dumped data, then all four were replaced by one unified "navigate globally, steer locally" locomotion layer.

seeds reaching objective2/75/7 korengal on-station gap300 mon objective 227 m route computed as555 mgoes around access-road bulldozed footprint6213 m²~450 m²

"Navigate globally, steer locally" — one locomotion rebuild to end four layered bugs hiding in plain sight.

Read the full report →
The valley with a properly sited outpost
The valley after the generation pass — the outpost now sits where a real one would.
2026-06-04

Villagers in the wire, gates that faced cliffs — the COP was its own enemy

TerrainUXRealism

The outpost was generating itself into traps: gates opened onto impassable slopes, faced away from every village, and buildings let soldiers walk through walls. Worst of all, calm civilians were strolling into the HESCO wire — 4,467 times per audit run, with no panic anywhere.

The shape: one harness (copaudit.ts) metricized all five open outpost issues plus the wire-pin bug in a single table, with a separate reachability.ts probe as an orthogonal fairness check. The fix was in the generation, not the behaviour.

villager wire-pin ticks44670 gate egress blocked1/90/16 gate faces >90° from villages7/90/16 perimeter ring continuity90%98%

4,467 calm civilians, ambling into the wire at routine waypoints — fixed by fixing the generation, not the behaviour.

Resolved tracked issues 001–005. Engineering record: docs/progress/2026-06-04-terrain/.

Clean route to the far village after the A* rebuild
After the corridor-A* rebuild: a direct route instead of a spiral around the outpost.
2026-06-04

The navigator dutifully walked a 718 m spiral to a 391 m objective

MovementTerrainUX

Squads sent to a far village would orbit the outpost in a spiral, string into a 100 m accordion snake, and spin in place at over 400°/s at the first awkward waypoint. The pathfinder was returning routes that looped 1.5 times around the COP to connect two points 15 m apart.

The shape: two new harnesses turned four bugs into hard numbers, then a corridor-constrained full-resolution A* replaced the broken coarse-plus-repair pathfinder, and a wake-following cohesion model replaced the flickering breadcrumb follower. Verified with before/after trajectory renders.

route length ÷ crow-flies1.77~1.1 loopy routes (>2.5×)14%2% navigator heading jitter443°/s0–2°/s point-man wall-grind ticks2286~16

The navigator dutifully walked every metre of a 718 m spiral to reach a village 391 m away.

Engineering record: docs/progress/2026-06-04-movement/.

Assault overlays during an AI-run firefight
The squad fights itself: base-of-fire and maneuver overlays during an all-AI assault.
2026-06-04

You have the watch — the day individual soldier control died

UXCombatCOIN

The game had a command path that contradicted its own design: players could micro-manage individual men during a firefight. That gutted the central tension — set the conditions before contact, then live with them. Combat had to be 100% AI once rounds crack.

The shape: a ground-up deletion of the old per-soldier order path and a fresh squad-leader coordinator state machine (react → hold / suppress / assault / break), verified by a new SOP-response matrix harness and a cover-discipline probe, then balance-checked across 12 firefights.

balance: 12/12 firefights stable 0.83 KIA · 5.08 WIA · 0 civcas · 0 stranded

You set the conditions beforehand and live with them — doctrine locks the moment rounds crack.

Engineering record: docs/progress/2026-06-04-squad-command/.

The simulation running with newly-activated systems
Finished systems — enemy indirect fire, bleed-out, wind ballistics — finally switched on.
2026-06-04

64 candidates, 10 survivors: waking the dead code

RealismCombatCOIN

The sim was honest in skeleton but full of finished systems nobody had switched on — enemy indirect fire, IED teams, and spotters built with zero callers; civilian casualties tallied but never read; wind stored but touching no ballistics. A 5.56 round wounded the same at 800 m as point-blank, and the insurgency could be ground to zero by attrition alone.

The shape: fan-out discovery — 10 subsystem-expert agents proposed gaps, a judge ranked 64 candidates to 10 (rejecting anything already modelled), each measured with a probe before/after, then hardened by two adversarial review waves that ran their own experiments.

M4 terminal damage @ 500 m30.714.2 insurgent strength, hostile valley (14 d)6680 CIVCAS strike attitude hit−23−35 dense-canopy night detect: eye 0% · thermal 41%

Realism by mechanism, not by fiat — the right casualty ratio fell out of the systems interacting.

Engineering record: docs/progress/2026-06-04-realism/.

Combat overview with new fire-mission reticles and suppression cues
A firefight you can read: fire-mission reticles, suppression arrows, casualty markers.
2026-06-04

Indirect fire was a silent flat disc — now it announces itself

RenderCombatUX

Every piece of combat information — incoming mortars, suppression direction, casualties, IEDs — was invisible or a featureless coloured blob. A player could not tell where fire came from, whether a soldier was pinned or merely suppressed, or that a round was about to land until it silently went off as a flat orange disc.

The shape: design-workflow first — 8 agents (4 design lenses → a 3-judge panel) produced a Combat Visual Language before any code. Reading real sim ticks then revealed two plan-changing facts (indirect rounds detonate instantly; arc progress is exact), and the cues shipped as a render-only layer.

81 in-flight rounds frame cost31.27 ms31.64 ms fx-probe: 7/7 hard checks pass frag arc progress monotonic [0.00 … 0.93]

The hardest part was realising indirect rounds never fly — so the reticle had to read the sim's intent, not its projectiles.

Engineering record: docs/progress/2026-06-04-combat-visual/.

The tactical map after the sprite overhaul
From wireframe boxes to a real outpost. Opens the studio art bible — all 160 assets, live.
2026-06-04

Blue rectangles to bas-relief: 160 hand-authored SVGs remake the map

RenderUX

The tactical map was primitive canvas shapes — blue rectangles for soldiers, wireframe boxes for buildings, coloured dots for everything else. Nothing about the visual layer said this was a defence-grade infantry sim. The outpost looked like labelled wireframes; villages were a house glyph and a name.

The shape: a background multi-agent workflow with one agent per sprite family — each rendered its own SVG and critiqued the rasterised PNG before returning — followed by a four-lens art-director pass driving a fix round. Verified live across every zoom band.

160 SVG assets across 16 families 116 fps at high zoom, sim running units → top-down soldier sprites per role + LOD symbols

Every asset was seen and critiqued by the agent that drew it before a human ever looked at the screen.

Open the art bible →
The redesigned UI overview
One of eight: the command interface redesign — a fuller right column, a used bottom strip.
2026-06-05

Eight patches in one day: the day the valley stopped being broken

TerrainMovementUX

Squads marched out the gate, looped back, and got stuck. Villages spawned inside the wire. Weapons fired off the edge of the world — 7,200 m range on a 3,620 m map. The road network was a series of dead-straight trenches. Eight separate problems, each measured and fixed in one session.

The shape: metricize across eight axes at once with a stack of headless harnesses, then write each fix as a single-writer change to one file so every metric delta is individually attributable. A tiered road network and a route-quality fix were the headline wins.

route-quality mean ratio1.261.01 village/COP footprint overlap2/240/24 road network connectivity36%59% trench "trough" cells3600~0

Weapons were firing 7,200 m on a 3,620 m map — the map was realistic; the numbers just hadn't been checked.

Resolved tracked issue 008. Engineering record: docs/progress/2026-06-05-batch/.

Traced arrival at the far village after the fatigue fix
The far village, finally reached. The fix was the movement economy, not the router.
2026-06-05

The router was fine — fatigue was eating the march

MovementTerrainMethodology

Squads were "setting up short" 273 m from their objective, or stalling halfway across the valley. Only 36% of physically reachable villages were ever reached, and just 13% of those directly opposite the gate. Everyone assumed the A* router was broken.

The shape: a fair oracle — an adversarial harness scored real-sim arrival against an 8-connected BFS ground truth, and per-tick traces apportioned blame to the exact tick. That overturned the hypothesis: fatigue saturation, not the router, was eating the march.

arrival among reachable36%76% opposite-gate REAR bucket13%50% network connectivity59%72% router was already good — mean ratio 1.01

The router was already good — a router rewrite would have fixed nothing.

Read the full report →
Act II · June 6

The deep systems & realism

With the valley walkable, the work turned to the things that make it a simulation — the river, the firefight, the population, the sound, and the soul of counterinsurgency — culminating in one integrated campaign-wide pass.

A squad crossing the river at a ford
A squad fords the river. Before this, in 28% of valleys the far bank was unreachable.
2026-06-06

The river was an impassable chasm splitting the valley

MovementTerrainMethodology

The river bisected the valley as a deeply-incised, cliff-walled chasm with zero crossings — in 28% of seeds the two banks were entirely separate passable components, so patrols simply could not reach the far side. Stranded squads, unreachable villages, a planner that gave up — all of it followed from that one structural failure.

The shape: a fast static structural audit (no sim) baselined the break, a fan-out stuck-case hunt found every stranding class, and the fix targeted the generation root cause — a walkable floodplain crossed at fords and footbridges — plus a river-aware foot planner. An adversarial pass re-measured everything on fresh held-out seeds.

planner returns no route30%2% banks in split components28%0% squads that return home23%81% worst-tick stall473 ms~29 ms

The adversarial pass earned its keep: it caught a civilian-bazaar regression stalling the tick at 473 ms before a shot was fired.

Resolved tracked issue 010. Engineering record: docs/progress/2026-06-06-river-navigation/.

A squad in a 360-degree security halt
A 360° security halt that forms like a real squad — and keeps scanning, not freezing.
2026-06-06

A frozen turret where a security halt should be

MovementUXRealism

Squads setting up 360° security marched men straight through each other to reach raw ring slots — and once halted, soldiers snapped their facing at 700°/s and froze, a dead turret instead of a scanning perimeter. Spacing was double doctrine. The whole thing read as mechanical, not military.

The shape: metricize first with two harnesses, then a three-expert research workflow (FM/ATP 3-21.8, game-AI steering, crowd-sim) converged on the fix — slot assignment by nearest, a settle that scans rather than snaps — re-measured in isolation, determinism confirmed.

max position churn3.71.1 heading jitter110°/s16°/s settle time40.1 s27.4 s halted scan motion0 (frozen)8°/s (live)

"The hardest part of command is watching" — the 360° now forms like a real squad, not a lattice snapping into place.

Engineering record: docs/progress/2026-06-06-movement-realism/.

The outpost interior with guaranteed building clearance
The outpost interior, now navigable — buildings spaced so courtyards never seal.
2026-06-06

The TOC was a dead-end courtyard — squads just kept grinding the wall

TerrainMovementUX

A player reported a squad stuck on buildings inside the outpost — men grinding the TOC wall forever. It looked like a movement bug, but it was a terrain-generation bug the COP audit's own metrics were blind to: buildings placed with zero guaranteed clearance were sealing interior courtyards the pathfinder could not enter.

The shape: metricize with two new harnesses and a render probe that drew the sealed courtyard in red, then a multi-agent workshop diagnosed the blind spot in the audit oracle itself (open space ≠ connected space), and a second adversarial pass verified zero regression.

seeds with unreachable posts~42/600/60 valley-2533 grind/600 s116926 survey-44 grind/600 s1280826 minimum building gap0 m≥10 m

The audit said 80% open — but open space and connected space are not the same thing.

Resolved tracked issue 012. Engineering record: docs/progress/2026-06-06-cop-interior/.

Mortar aimpoint before vs after the densest-cluster fix
Before/after of the mortar aimpoint — averaging all enemies put rounds on the squad.
2026-06-06

The squad stopped dropping mortars on itself

CombatAIRealism

A player report: "the squad called a mortar mission much too close to them, nowhere near the enemy." Both halves were real — the AI averaged all enemy positions (landing rounds on the squad in an L-shaped ambush), and the danger-close gate existed but never actually stopped the call, producing a genuine US fratricide.

The shape: metricize first — a probe intercepted every AI call-for-fire across 40 deployments, scoring the aimpoint against independent ground truth, then let rounds fall to count real fratricide. The fix: a densest-cluster aimpoint, a real danger-close withhold, and an FDC check-fire.

danger-close calls3–4%0% fratricide events10 nearest friendly to impact10 m69 m off-target vs living enemy: 0% (held-out)

It found a one-line averaging mistake and a missing safety the way a soldier would: by measuring where the rounds actually fell.

Read the full report →
The outpost at default zoom showing figure scale
The scale audit. Opens the full scorecard — every dimension of the sim vs the real Korengal.
2026-06-06

The soldiers were 36 metres wide

RenderRealismUX

A player felt the scale "doesn't feel true to life." A cited 13-agent research pass confirmed the worst offender: the renderer painted each soldier roughly 21 m wide at the default zoom — 36× his real 0.6 m footprint — so a 9-man squad at textbook 5.5 m spacing fused into a single blob dominating the 170 m outpost. The sim was authentic; the renderer was lying about it.

The shape: a live-app probe applied the renderer's exact clamp formulas per zoom, cross-referenced against a 13-agent real-world research pass (Korengal geography, FM/ATP 3-21.8, first-hand accounts), producing a verbatim scorecard of the sim against reality.

figure vs true size (default zoom)36×17× figure vs true (tactical zoom)12× squad figures overlaptruefalse

The sim was already authentic; the renderer was lying about it — a 36× exaggeration hiding correct numbers underneath.

Read the full audit →
The COIN strategic HUD
The counterinsurgency HUD — attitudes, cooperation, CERP, directives — finally load-bearing.
2026-06-06

The strategy layer that punished patience and rewarded body-count

COINCampaignSimulation

The game's central claim — "you can win every firefight and still lose the valley" — was mechanically false. Careful counterinsurgency and a pure body-count run scored almost identically. Projects never completed because no element stayed to hold security, so every funded build sabotaged itself. CERP only counted toward zero. Directive deadline-and-penalty code was never read anywhere.

The shape: a campaign-loop harness compared careful-COIN vs body-count across seeds and game-day windows as the authoritative ledger, then a coherent engine wave fixed each root cause in isolation — and a later pass fixed the instrument itself, which had been over-marching careful patrols into ambushes.

careful Δ mean attitude0.1 (flat)+13.5 projects complete / tour02.8 careful vs body-count score spread~036.0 CERP ever rosenoyes

The harness meant to prove COIN worked was itself broken — it over-marched careful patrols into ambushes, doubling KIA and inverting the spread.

Resolved tracked issue 015. Engineering record: docs/progress/2026-06-06-coin-real-game/.

The civilian atmospherics report
The diurnal-rhythm & melt-away report — the calm-before-the-ambush tell, measured.
2026-06-06

The valley never slept — until it learned to go quiet before the shooting

RealismCOINUX

The valley was diurnally flat — civilians wandered 24/7 at constant population, day and night identical. Worse, the flagship counterinsurgency tell the tutorial explicitly teaches — civilians melting away before an ambush — did not exist in the engine. A patient player had nothing to read.

The shape: a probe with an independent oracle measured outdoor occupancy by hour and pre-contact distance-to-home, tuned on one seed and proven on a held-out one, with same-seed determinism asserted by hashing civilian positions across two identical runs.

outdoor occupancy at night100% (flat)~0% midday outdoor≥62% melt cohort closes home-distance0%~50% before any shot

The fields thin over a few seconds — that thinning is the absence an alert player learns to read.

Read the full report →
The in-game audio controls
Every sound synthesized at runtime — no audio files, no audio library, fully deterministic.
2026-06-06

The silence that made you realise: no game sounds here at all

AudioRenderUX

The game had no audio whatsoever — no crack-thump of incoming rounds, no SAW rip against a PKM hammer from the high ground, no mortar shot-and-splash, no radio net under the firefight. The genre's spine was missing, and the player was watching a silent simulation.

The shape: ground-up procedural synthesis with no binary assets and no npm dependency — a strict pure/browser split (a deterministic cue collector in the sim half; Web Audio node graphs in the browser half), verified headless by byte-identical replay and live against a debug ring buffer.

1520 cues scheduled · 0 double-fires two replays of one firefight: byte-identical (5983 cues) lib/sim has ZERO audio imports

Every sound is synthesized procedurally — no binary assets, no npm audio dependency.

Engineering record: docs/progress/2026-06-06-procedural-audio/.

The deploy loading screen
The staged loading screen that replaced a 6.5-second frozen button.
2026-06-06

The frozen button that swallowed six seconds

UXRender

Clicking Deploy froze the browser for roughly 6.5 seconds before the game appeared — no spinner, no progress, nothing — because a 4096×4096 relief bake (16.7 million pixels) ran synchronously inside the click handler before React could ever paint a loading screen.

The shape: profile every deploy phase over the DevTools protocol to find the villain, then stage the work behind a loading screen with double-rAF yields and split the monolithic bake into 40 progressive bands — verified byte-identical across 5 seeds. An honest fix: feedback, not raw speed.

feedback after clickfrozen0.2 ms progress updates during bake042 samples total click→playable ~6.5 s (goal was feedback, not speed)

The deploy is now covered by feedback — it is honestly not faster, and the report says so.

Read the full report →
A night firefight after the integrated campaign wave
A night firefight after the five-wave pass. Opens the full illustrated campaign report.
2026-06-06

Five waves, one campaign: the firefight, the audio, the valley, and the soul of COIN — at once

CombatCOINRender

The game still played like a prototype: US soldiers died at a 5:1 disadvantage, the strategic layer was indistinguishable from a body-count run, soldiers were painted 36× too large, and the valley was silent and diurnally flat. This pass tied the loose threads of the day into one verified campaign-wide release.

The shape: a goal-driven orchestration — a 12-domain recon fan-out produced 63 ranked hypotheses, then five file-disjoint implementation waves (combat soul, audio, game-feel, COIN, scale), each ending in its own adversarial verify gated by the full done-gate suite.

US:enemy exchange ratio~5:1 against~even / favorable COIN tour-score spread0.530.3 suppression pinned fraction0.00027–36% squad figures overlaptruefalse

You can win every firefight and still lose the valley — and now the scoreboard finally proves it.

Read the full report →
Act III · June 7

Cinema & polish

The last pass turned a correct simulation into one that looks and sounds like the place it models — weather, the light a muzzle throws on the hillside, the punch of a near-miss, and a canyon that echoes.

A night detonation lighting the relief
A detonation briefly lights the relief at night — and the figures are now true-scale.
2026-06-07

The valley was a weatherless diagram — now it has weather, dark, and a muzzle's light

RenderCombatUX

The battlefield looked like a CAD schematic: no weather, no lighting from gunfire, no sense of scale, and every soldier painted as a 21-metre blob. A night firefight was indistinguishable from a daytime diagram. This pass added rain, fog, snow and dust, night-light from fires and muzzles, wind-driven smoke, camera punch — and shipped the true-scale figures.

The shape: a render-side cinematic pass that is pure read-out of existing sim state — no writes back into the deterministic engine — verified with the same live-app scale probe, resolving the cheap "representation" half of the scale issue.

figure vs true (default)36×17× soldier as % of outpost0.126~0.010 squad figures overlaptruefalse

A detonation should briefly light the relief — not just add a dot to a diagram.

Resolved tracked issue 014 (render half). Engineering record: docs/progress/2026-06-07-atmospherics/.

Before/after firefight spectrogram
A firefight spectrogram, before vs after. Opens the full report — with playable A/B audio.
2026-06-07

From mono-and-silent to a living canyon: a 20× soundscape rebuild

AudioRealismRender

The earlier audio pass put sound in the game — but it was combat-only, essentially mono, had no reverb, and fell silent between firefights. The valley didn't breathe. This pass rebuilt it into a living place: a canyon that echoes, a wind/river/generator/wildlife bed keyed to the sim's time-of-day and weather, terrain that muffles fire behind a ridge, and 5-layer weapon synthesis with a correct 5.56-vs-7.62 timbre tell.

The shape: a 6-agent research fan-out (gunfire DSP, canyon reverb, ambient design, spatialization, adaptive mixing, valley acoustics) → an additive bus architecture → three parallel module builds → an offline oracle (audio-render.ts) + spectrogram viz tuned to all-green. The adversarial pass caught a real oracle bug (a voice counter that never decremented, so the firefight was measured on only its first 32 cues).

firefight stereo correlation0.981 (mono)0.392 (wide) ambient calm bed−120 dB (silent)−42.6 dB (alive) firefight reverb tail2450 ms7100 ms (canyon) weapon timbre tellM4 ≈ AKM4 6860 > AK 6212 Hz

100% procedural — Web Audio only, zero audio assets, the reverb impulse is generated noise — and deterministic.

Read the full report (with playable A/B audio) →
A village rendered as a ring of distinct walled qalats in orchards
Landigal as a hamlet of distinct qalats. Opens the full completion report.
2026-06-07

World scale, finished: the outpost, the platoon, the villages, the valley

WorldgenRealismRender

The June 6 audit proved the soldiers were painted 36× too big — a render lie over an already-authentic sim — and shipped the loud render fix that day. The riskier sim-scale half was deferred behind the determinism contract. This pass finished it: a platoon-scale outpost, a full-strength roster, a north-draining gradient with honest names, and villages that read as hamlets.

The shape: six atomic commits (one careful writer on the determinism-critical terrain), a 7-agent adversarial verification fan-out (0 blockers) — then the dynamic balance harness caught a stall the static pass couldn't see: a returning patrol frozen in an overlapping-compound wall maze, fixed by spacing the hamlet into discrete qalats.

COP diameter170 m (FOB)120 m (platoon OP) platoon / weapons squad35 / 341 / 9 village footprint1 monolithic box2–5-qalat hamlet stranded patrols / 121 (stall)0

The brief said the gradient was inverted; the code said it already drained north — only the names lied. Code is ground truth; a spec is a hypothesis.

Read the full report →
The command HUD, after vs before
The command HUD, after (top) vs before (bottom). Opens the full UI/UX report with charts and annotated crops.
2026-06-08

The command surface, re-cut: 461 → 1 measured UX defects

UI/UXAccessibilityRender

The HUD a commander stares at for a whole deployment was dense and handsome — and quietly hostile: built on 9–11px text, with zero keyboard-focus visibility, no motion-safety, glyph buttons a screen reader calls "button," and the campaign's five north-star metrics drawn as identical 2.5px slivers. This pass kept every ounce of the milspec soul and fixed the legibility and accessibility under it.

The shape: a CDP audit harness metricized the debt at 461 defects before any edit; an 11-agent survey (5-persona judge panel + 6 single-axis proposers + a synthesis lead) produced a deconflicted backlog; four atomic rounds implemented it; a 5-agent adversarial paired re-judge graded the result and demanded residuals. A focus-floor + a 12px type floor did most of the collapse; an overflow oracle and a keyboard-behaviour harness (6/6 pass) verified nothing broke.

measured UX defects4611 (−99.8%) keyboard focus failures1520 sub-12px text nodes2620 expert quality score /10058.682.4 (5/5 prefer)

The literal 20× lives on the objective axis — a 461× defect reduction, reproduced deterministically — and an independent expert panel unanimously confirmed the direction and magnitude.

Read the full report →
A squad caught mid-flank: a base of fire pins the enemy while a maneuver element bounds to a covered flank
Caught mid-assault: a 4-man base of fire pins the enemy while a 4-man maneuver element bounds to a covered flank — the squad-leader AI's call, no player order. Opens the full report.
2026-06-08

Soldier-scale realism: hours-long dwell, an autonomous flank — and one honest revert

SimCombat AICOINRestraint

At soldier-scale zoom, four things "didn't feel right." Each was turned into a number first. Village dwell was minutes where doctrine is hours, and the census teleported to "done" on arrival; in a firefight the whole squad sat on a base of fire while one man shuffled — it never flanked unless hand-told, and even then it rushed the front. And a squad ordered up a mountain ringed the spur instead of climbing.

The shape: three headless probes baselined on HEAD; population-driven progressive census + a dwell event-roll the player warps through; the squad-leader AI given the maneuver decision (develop-timer SOP lever, covered-flank routing, bounding overwatch) with KIA held flat. The fourth — switchback pathing — was built, measured against a clean HEAD worktree on held-out seeds, found to regress the most load-bearing movement system (a new stall) for a below-target gain, and reverted, then logged as two tracked issues + a reproduction probe.

on-station dwell150–1100 s1–8 h (16–37×) contacts squad flanks unbidden0%22–48% bounding-overwatch discipline0%100% OP detour fix×3.13 readyreverted (kept green)

The wins are nice; the revert is the tell — an agent measured its own plausible fix, found it traded a below-target gain for a movement regression, and backed it out.

Read the full report →
The combat outpost at 02:00 — warm lit windows and floodlights against a dark valley
02:00 at the outpost — the one warm, human place in a cold valley. The night life-signs, the HESCO bastion ring, the lit gate. Opens the full report.
2026-06-08

The combat outpost, made real: a fire plan sited by terrain, a garrison that lives, and the watch

SimRenderCombat AIDoctrineWorkflow

The COP was geometrically correct and dramatically inert — a cluster of b-huts on a tan patch, defended by eight identical nubs, with crew-served guns handed out by array index and no Mk19 at all, a garrison that loitered, and a wire attack you could only watch as a cutscene. copaudit was already 100% green, so this was never a bug hunt: it was taking a system that works and making it convincing.

The shape: a six-specialist design survey (defense, visual, garrison, facilities, gameplay, atmosphere — each pinned to ATP 3-21.8 and the real file:line facts) and a synthesis judge produced a Law-6 unification: one terrain LOS-sweep spine that sites the M2 on the longest avenue and the Mk19 over the worst dead ground, feeds the render silhouette and the mortar pit; one render env seam for the night life-signs; and one reused fire-request loop that turns a wire assault into a Final Protective Fire the commander must clear. Seven atomic commits, each verified.

crew-served sitingblind array indexterrain LOS sweep (M2→922 m, Mk19→94% dead) Mk19 on the wall01, on the dead ground a wire assaultpassive cutscenecommander-cleared FPF fob.hescodead statwritten by details, read on the panel

The art to build a recognisable COP already sat unused in the sprite manifest; the doctrine to defend it sat in the terrain. The pass wired the engine to the ground truth — and copaudit stayed clean (0 sealed pockets ×9).

Read the full report →
A live firefight at the outpost — muzzle flashes at the firing positions, suppression crescents on pinned men, casualty pulses, the troops-in-contact banner
A live firefight, mid-contact — muzzle flashes that finally render, suppression on the pinned men, casualties pulsing. Opens the full report.
2026-06-08

Making the firefight feel real: bullets that travel, a flash that renders, suppression with teeth

RenderSimAudioDoctrineWorkflow

The owner played a firefight and named five things wrong: bullets teleported — appearing halfway between the gun and the target; the feedback was flickery; a man with rounds cracking past his head fired just as much; every combatant felt the same; and the machine guns buzzed. None was a crash — all were feel, which the project's rules forbid touching until it's a number.

The shape: a six-specialist survey + judge ranked the best five fixes spanning visual + sim + audio. Four of the five shared one root cause — the renderer read the 0.1 s sim tick verbatim at 60 fps — so one sub-tick interpolation made bullets travel and resurrected the muzzle flash (which had been aged past its own draw cutoff before any frame could show it). The two simulation changes reshaped the existing random draw without adding one, so a seed still reproduces; balance held within ±15%.

muzzle flashes that render0 / 453453 / 453 per-frame bullet jump123 px / 93% frozen~20 px / 0% fire rate when pinnedflat (~2.2 rds/s)−48% (1.05) synth cracks per MG burst~56 (a roar)8 (a hammer)

The hardest part wasn't the fix — it was the restraint: the personality "panic spray" was tuned down when a first attempt pushed casualties out of tolerance. The win condition was never "more dramatic," it was "more honest, within ±15% of the old balance."

Read the full report →
The valley with its new path network — a scaled dirt MSR following the river, village tracks, and goat-trail threads molding to the slopes
The valley, made legible — the MSR and tracks now read as scaled dirt roads molding to the terrain, the ridges as impassable rock. Opens the full report.
2026-06-09

Reading the ground: traversable steep terrain, footpaths that mold, cliffs you can read, cover you can use

TerrainSimRenderMovement

The map looked alive but quietly lied about itself. A new probe floods the valley the way the mover does and found the tell: 80% of the ground was passable but only 48% was reachable — the hard "no foot traffic above 51°" cutoff salt-and-peppered fake cliffs across climbable faces, shattering the walkable map into islands. The footpaths were a 5 m paint smudge; a 56° wall was drawn the same warm tan as a meadow; and the cover that decides a firefight was one averaged number per cell, decoupled from the rocks you could see.

The shape: four asks, four measured iterations, each baselined and re-measured in isolation. Soften the impassable line to reconnect the steep band (reach 48→61%, real cliffs preserved). Switchback foot-trails up the spurs (a benched first version cut impassable trenches — caught in a trail render, rebuilt to surface-lay). Stroke every captured path centerline as a scaled dirt line, and shade everything above the foot-passable slope as sheer cool rock. And promote the boulders from cosmetic sprites to sim objects the renderer draws from — so the rock you see on rocky ground is where the engine's cover is. (The combat-cover half on open ground was measured and cut: even a light stamp dragged firefights +89% WIA — deferred to a sub-cell model.)

walkable map reachable48%61% footpaths per valley~622–32 cover objects (drawn = sim)011,600 changes reverted on the numbers2

The restraint was the work. A cost-floor tweak that felt right bought no routing win, ran the sim 1.5× slower and spiked WIA 71% — reverted on the numbers. A cover change that looked fine in aggregate turned ridge firefights into grinds — caught only because a same-seed A/B disagreed with the cross-config one. Cross-config comparisons lied twice; same-seed A/Bs told the truth.

Read the full report →
Two trajectory diagrams of the same OP climb: BEFORE a 1281 m route ringing the spur in a wide arc, AFTER a 769 m route climbing nearly straight with one clean switchback
Same OP, same seed. Left: the old isotropic planner rings the spur — a 1281 m arc around the mountain. Right: the any-angle planner climbs it — 769 m, one clean switchback. Opens the full report.
2026-06-10

Climbing the mountain: a squad that switchbacks up a face instead of ringing the spur

MovementPathfindingTerrainRealism

Order a squad to a peak OP and it walked a 3-kilometre arc around the mountain instead of climbing it. The cause was geometric: the foot-movement cost was isotropic — it read a cell's slope magnitude, the same in every heading — so a diagonal traverse bought nothing and a switchback could never be cheaper. And the planner was 8-direction, so the shallow traverse angle a switchback needs fell between the grid headings. Two earlier in-place attempts had been reverted for regressing every route and stalling a squad.

The shape: the scoped rebuild the issue had already named, built to be additive. The slope penalty moves out of the per-cell cost and into the edge — speed now depends on the signed grade along travel, so a long switchback genuinely beats a short scramble (it minimises travel time, not length). A Theta* (any-angle) search lets a leg run at any heading — the geometry the 8-dir grid couldn't draw — with a turn penalty for few clean bends. It fires only for a deliberate squad march to a steep, elevated, tactical objective; world generation and valley-floor routing never trip it, so the terrain and village movement are byte-identical. A same-seed A/B kill-switch proved it: planner off = byte-identical to HEAD.

held-out OP detour×3.58×3.16 climbable-face OP (korengal)×2.19×1.31 switchback jitter (reversals)3–61–2 village routing & terrainbyte-identical

The honesty is in the residual. OPs the probe tucks behind a genuinely impassable massif still detour — and that's real terrain, not a planning miss, so the planner correctly falls through to the proven router rather than paying for a near-whole-map search to shave one adversarial seed. And the balance cost is stated plainly: climbing exposed high ground costs a little more (KIA 0.92→1.08) — the realistic price of the climb, accepted rather than dodged, with the stall guard the prior attempts failed now passing.

Read the full report →
A radial plot of the cover one boulder gives a man tucked behind it — lobes pointing at the rock, ~0 on the flanks, the prone lobe much larger than standing
The cover one boulder gives the man behind it, as the shooter's bearing sweeps around. It points at the rock — and collapses to ~0 on the flanks. The prone lobe dwarfs the standing one. Opens the full report.
2026-06-10

Behind that rock: directional, posture-aware cover a soldier can actually use

CombatCoverRealismBalance

The drawn boulders became real sim objects in June, but the firefight still read a single averaged cover number per 5-metre cell — so a man tucked behind a rock was scored exactly as exposed as one standing in the open (measured: behind 0.08 = open = flank). The rock was decoration to the bullets. And the obvious fix had already been tried and reverted: adding open-ground cover omnidirectionally made both sides survive longer and dragged every firefight into a +89%-WIA grind.

The shape: the missing thing wasn't more cover — it was direction. A boulder only stops the round coming from the bearing it faces. So a discrete object between shooter and target now blocks a posture-scaled fraction of the shot (a low rock hides a prone man far more than a standing one), on the fire-hit path only, and the AI of both sides moves to tuck behind it. Because a flanker's round is never on that line, the squad's already-shipped autonomous flank defeats the cover, and the fight resolves by maneuver instead of a longer attrition clock.

cover behind a boulder0.080.39 cover from the flank0.080.00 US wounded (held-out A/B)6.754.33 enemy still accounted5.505.58

The reverted attempt added cover and casualties rose; this one adds directional cover and casualties fall — US WIA −22% on the tuned seeds, −36% held-out, with the enemy still dying via the flank. Same idea (open-ground cover), opposite result, because of one property: a boulder faces a direction.

Read the full report →
Bar chart: dead batteries cut night detection to 41%, dehydration cuts fatigue recovery to 50%, no medical cuts wound recovery to 54%
The cost of a neglected resupply, measured: dead batteries blind the patrol at night, dehydration and a bare aid bag slow recovery — all bounded so a stocked COP is unaffected. Opens the full report.
2026-06-10

Supplies that bite — and the COP assault that wasn't there

LogisticsCOINCombatLaw 1

Issue 021 wanted the COP's fortification (fob.hesco) to matter in a fight: a hardened wire should blunt a costly assault. Metricizing it first — Law 1 — flipped the whole campaign. A probe staged the complex attack and measured garrison casualties hardened vs neglected: zero either way, across 8 seeds and a 20-minute trace. The trace showed why — the attackers spawn at 340 m and never close the wire. The insurgent AI engages and shoot-and-scoots laterally; there is no assault behaviour at all. A "complex attack" is a standoff from the ridges (the real Korengal), so fortification and claymores have no event to defend against.

The shape: rather than fake a coupling to a non-event, ship the part of 021 that affects the gameplay you actually have — the patrol. Water, food, batteries, and medical drained every day to a floor with zero consequence. Now each bites, as a bounded clamp: dead batteries cost the US their night-vision edge (a 2.4× detection penalty in the dark), dehydration and underfeeding halve fatigue recovery, and a bare aid bag slows wound recovery — while a fully stocked COP plays exactly as before.

night detection (dead batteries)NVG2.4× worse fatigue recovery (dehydrated)100%50% wound recovery (no medical)100%54% balance change at full stock0

The most valuable line of code this campaign was a probe, not a feature. It proved the assault the fortification was meant to stop never happens — saving a coupling to a number that never matters, and pointing at the realism win that does: the resupply you forgot to call.

Read the full report →
Bar chart: forest-on-shaded-slope 59%→64%, scrub-on-sunny-slope 48%→56%
The shipped signal: vegetation now leans the ecologically-correct way — forest to the shaded draw, scrub to the sunny spur. Modest by design, real, and on by default. Opens the full report.
2026-06-10

Reading the sun: vegetation that knows which way a slope faces

TerrainEcologyBalanceRestraint

The valley grew the same vegetation on a north slope and a south slope — no notion of aspect, the single most visible fact of a real mountainside (shaded slopes hold moisture and grow forest; sun-facing ones bake to scrub). An aspect term in classifyLand fixed it — and because vegetation is cover and concealment, it could not be slipped in. A three-strength balance sweep showed the firefight moves chaotically with the terrain (KIA non-monotonic, strandings flickering), so it ships at the one strength that earns its place.

The shape: measure, don't assume. 0.16 ran too hot (+56% KIA, stranded); 0.025 was chaotic (also stranded); 0.05 clears the no-stall guard, gated to the steep faces so village/COP siting stays byte-identical. Held out on fresh seeds: KIA down on both sets (the permanent loss never regresses), the lone stranding pre-existing (aspect adds none), and the wounded count terrain-dependent — disclosed, not buried.

forest on shaded slope59%64% scrub on sunny slope48%56% US KIA (both seed sets)1.171.08 village/COP sitingbyte-identical

It shipped on by default only after the balance earned it — three strengths measured, a held-out check, and the honest caveat (combat intensity varies with the ground, which is itself realistic) stated in the open. The realism is in the vegetation; the rigor is in refusing to ship it until the numbers said it was safe.

Read the full report →
2026-06-10

Closing the open issues — four realism wins shipped, two reasoned closeouts, zero regressions

BacklogMethodRestraint

A single pass over every open issue, run the way this repo demands: metricize before touching code, prove every win on held-out seeds, adversarially verify before "done" — and ship only what the numbers earn. Four realism wins shipped active (019 switchback pathing, 020 directional cover, 021 logistics teeth, 007 aspect-driven vegetation), each with a held-out proof and — for the two combat changes — an adversarial workflow that tried to break it and couldn't. The fourth was the hard one: aspect touches the cover/conceal field a firefight reads, so it took a three-strength balance sweep + a held-out A/B to find the one strength (0.05) that improves the sim — forest-faces-north 59%→64%, scrub-faces-south 48%→56%, KIA down on both seed sets, no aspect-caused strandings, siting byte-identical. It ships on by default. The two remaining tickets (011, 009) already run their active fix in the shipped game — a progressive, cached relief bake behind the loading screen; the reachability-aware spawn-snap + benching guard — so what's left is a logged do-not (resolution-lowering would degrade the relief the realism work added) and a documented terrain floor, not unbuilt work.

The headline is that six issues now carry a live fix and the sim got better without a single shipped regression — including the one (aspect) that took a sweep across strengths to ship honestly instead of guessing. On a project whose success metric is a skeptic asking "did an AI really build this?", the discipline to keep sweeping until the numbers point to the strength that ships — and to hold a perf ticket at its honest deployed mitigation rather than fake a further win — is the part that's hardest to fake.

Read the overview →
The shura, live: the squad ring between the Kandlay compounds, the command log carrying 'Asadullah Saifullah sits down with SSG Toner — the shura begins'
The hollow heart, filled: the elder walks out, sits down 4.1 m from the squad leader — and only then does the shura's business flow. Live, Day 1, +42 s of on-station. Opens the full report.
2026-06-10

Making the people real: group minds decide, individuals execute, everything visible — and the valley remembers

Sim AICOINRenderDoctrineVerification

The owner's standing brief was a dare, not a ticket: make every soldier, fighter, and villager so convincing a veteran forgets he's watching a simulation. Six recon agents swept the code and corrected the brief itself — civilians never respawned (persistent named people all along), enemy cells already existed, near-miss reactions were real physics. The genuine gaps: no group minds, no scenes, nothing visible. The enemy was a row of independent turrets (a five-man trap opened raggedly over 11 s); two same-SOP squads assaulted in lockstep (step-off spread literally 0.00 s); the elder had a name but never appeared; a civilian death was an attitude penalty, not a person.

The shape: one architecture extended, never a second pattern invented — the friendly squad's decide/execute split (squadFight/friendlyBrain) given to the enemy (cell-combat.ts: volley as one, displace by halves, peel to a rally, rout contagion — A/B-proven via an ITM_NOCELL kill-switch), to the village (a staged shura gating the KLE's business on elderMet; named blood-debt grievances buried at first light, settleable by solatia; kids trailing only friendly patrols), plus a deterministic callout bus so internal state becomes watchable. Five new probes caught six real bugs before ship — a 390-emission shout metronome, a half-volley that scooted before its first shot, a dead medic-scene gate (0/1048 ticks), and a medic firing on his own patient, caught by a balance bisect's KIA-up/WIA-down signature. Proven on held-out seeds at n=16.

ambush volley p9011.2 s0.6 s breaking cell → rally (30→90 s)40→61 m rout54→34 m peel man-down callout coverageinvisible100% (80/80), two-run identical step-off spread / commit drift0.00 σ / 0.4 s0.10 σ / 6.0 s kids trail, friendly vs hostile (held-out)9.59 vs 0.00 min held-out balance KIA1.080.94 · civCas 0 · 0 stranded

"Abdul Khan of Kandlay was killed by our fire. His household will remember." — the ledger entry the engine now writes. The residuals are led with, not buried: the pinned-revert mechanism exists but never fired at natural tempo (rvt 0), and the massed-volley spectacle shot was never captured live — three failed attempts documented so the next session doesn't rediscover them.

Read the full report →
Spectrogram of the new incoming-shell whistle: a descending sweep that swells toward impact
The new incoming-shell whistle on the spectrogram: 2.26 s of descending shriek, swelling as the round closes — it ends exactly where the splash begins. Opens the full report (with playable A/B audio).
2026-06-11

The mixer and the whistle: per-category sound control + five audible upgrades

AudioUIVerification

Two asks: let the player turn categories of sound on and off, and make the sound itself better. The mixer rode the existing bus architecture — one user gain per category (combat / ambience / radio / alerts) inserted between each bus and master, because the contact-duck system writes absolute values to the bus gains and a user trim on the same node would be stomped on every duck. The valley reverb return joined the combat category (muting combat must mute its echo), the danger-close klaxon moved to alerts where it semantically lives, and an oracle re-render proved the re-plumbing level-neutral: all 24 scenes within 0.15 dB.

The shape: the five audio upgrades each came from a recorded residual or a missing battlefield signature, and each closed against the offline render oracle. The big one was a synthesis subtlety: the calm bed measured near-mono (corr 0.997) despite panned voices, because every bed loops one shared noise buffer and two in-phase filters of the same source stay correlated no matter how they're panned — the fix was starting each voice at a different offset into the loop, and only then did the panning (river bands straddling the river's real bearing, the generator on the COP's, Haas rain) buy real width. The whistle is pure determinism: the sim already counts down a fire mission's etaS, so the mapper voices one whistle 2.4 s before the first round, once per mission, latched and warp-safe.

calm-bed stereo width4% (corr 0.997)12.4% (corr 0.963) incoming-shell whistledid not exist2.26 s, ends at the splash calm→combat dynamic spread23.0 dB25.8 dB thunder / adhan vibrato / ricochet familiesnone / flat / 1storm roll · 5.3 Hz · 3 mixer plumbing (24 scenes)level-neutral within 0.15 dB

The honest residuals lead the report: the band-split river left the calm bed ~2.9 dB quieter (in band, knob named), and there is still no MEDEVAC helicopter — evacuation is instantaneous in the sim, so a rotor loop has nothing to attach to, and faking one with a timer in the audio layer was the change deliberately not made.

Read the full report (with playable A/B audio) →
Spectrogram comparison: the PKM's energy (top) vs the DShK .50 sitting visibly lower (bottom)
The calibre tell on the spectrogram: the PKM (top) vs the DShK .50 (bottom) — the heavy gun's energy mass sits visibly lower. Opens the full report (with playable A/B audio).
2026-06-11

Calibre voices: every weapon gets its real sound

AudioSimulationVerification

A soldier tells weapons apart by ear before he sees anything — the report's pitch and weight ARE the information. But the audio layer only knew faction and "is it a machine gun", so an M2 .50, an RPG launch, an M9 pistol and an M4 all collapsed into two rifle cracks. The unlock was one sentence of sim code: the engine now stamps which weapon produced each muzzle/blast effect, and the whole calibre ladder fell out of it — measured, solo/near: M9 7701 Hz > M4 6860 > M24 6450 > Enfield 5876 > DShK 5242. The .50s needed one more honesty: the shared 6 kHz attack transient had made every gun equally bright, so the attack itself is now weapon-tinted — the M2's is a low whump.

The shape: new sounds where the battlefield was silent (the RPG's pop-and-whoosh launch — "RPG!" is finally audible as a launch, not just an arrival; the M320's bloop and the Mk19's bolt-thunk; a man two metres away swapping mags at −51.6 dB), per-weapon voice rows refining the class voices (bolt guns cycle their bolts ~0.5 s after the report), and physics upgrades to the loudest moments: the IED now heaves (seismic 26→16 Hz ground wave leads, the airborne crack arrives 8 ms late through the soil cap), the incoming shell tears instead of singing (two incommensurate modulators that never phase-lock), the near miss completes its N-wave. Where the standard metrics went blind — centroid and HF% are bin-count-dominated by broadband noise — the oracle grew a low-band share metric, because a blast's calibre lives below 200 Hz.

weapon voices (distinct)2 faction cracks11 (+5 cue kinds) DShK vs PKM centroidsame voicedarker by 772 Hz IED centroid / LF share5233 Hz / —3951 Hz / 37.9% 60 vs 120 mm mortar LF shareidentical14.1% vs 19.9% oracle assertions813/13 green (39 scenes)

The residuals are named: the RPG launch never fired organically in the three held-out runs (gunners carry one round) — routing proven by direct mapper check, organic occurrence unobserved; and there are still no synthesized screams, deliberately — the recorded negative on synthesized radio voice ("reads cheesy faster than squelch reads sparse") extends to wounded men, with higher stakes.

Read the full report (with playable A/B audio) →
Act · June 13

Making the valley photoreal

The sim was already a defense-grade COIN model and the map was legible — but it still looked like a map. A first WebGL attempt only switched the sun on over the same flat bitmap and was reverted. This is the rebuild that makes the ground a photograph without ever stopping being a chart.

The valley at strategic zoom: terraced cropland quilting the floor, dusty scree slopes, a dark river braiding the bottom, ridge-shadowed spurs — reading as satellite imagery
Strategic zoom, noon — material identity per landcover, ridge shadows, a dark drainage network. The before was a single tinted heightmap. Opens the full report (before/after at every zoom band).
2026-06-13

The valley is real — a WebGL terrain rebuild (attempt #2)

RenderingWebGL2Verification

The first attempt at this overhaul was reverted: it moved the existing flat painted relief onto the GPU byte-for-byte and laid a colour filter and a moving sun over it — the same map with the sun switched on. This rebuild inverts the idea. The painted albedo is demoted from "the image" to one input of five, and a WebGL2 fragment shader recomposes the surface per-pixel from the simulation's own arrays — landcover, slope, cover, the heightfield — against a procedural material library, lit in linear radiance by the live master-clock sun.

The shape: a WebGL2 underlayer beneath a transparent Canvas-2D HUD (the units and contours are a separate, never-graded layer — the legibility firewall). Three passes a frame: terrain → RGBA16F HDR (per-landcover materials + detail-normal raking the live sun, baked horizon AO, multi-km ridge shadows, a dark-silt flow-advected river, aerial perspective + in-fog god-rays), then threshold bloom, then ACES tonemap + time-of-day grade + dither. A u_detailGain zoom-ramp collapses the whole transformation to byte-faithful relief at strategic zoom, so the band the old map got right is provably preserved while close zoom resolves into real ground. Nine commits, each green; reuses the verified Kunar solar model, the bit-faithful heightmap port, and the exponential cast-shadow march.

terrain surfaceflat painted hillshadeper-pixel material recomposition strategic mid-grey luma (firewall)132.2132.4 (byte-faithful kept) the rivernear-white "salt flat"dark silt water adversarial verdict1/4 clears the barfixed, fresh-eyes confirmed standing checkstsc/build/smoke/balance green

Residuals named, partial stated as partial: the terrain got the full 10×, but the assets got grounding (sun-tracked shadows), not a full relight — the 164 sprites are still flat bitmaps. A per-sprite normal/AO relight and FX particles are a scoped follow-on (issue 028), logged rather than shipped half-done. The win the verification earned: a four-critic panel caught a snow-white river and a contrast-crushing haze the build's own author had rationalized away — fixed, then re-confirmed by a fresh reviewer.

Read the full report (before/after at every zoom) →
Act · June 26

Auditing the instruments

A game built test-first lives or dies by its tests. So we turned the lens on the harnesses themselves and asked the uncomfortable question: have any of them stopped measuring good and started measuring same?

The harness charter report — 'When the tests start defending the bug'
The charter that came out of the audit: a gate may never assert the sim's own past output. Opens the full report.
2026-06-26

When the tests start defending the bug

MethodologyVerificationCOIN

An audit of all 84 harness files went looking for tests that had become blockers — and found exactly one, though it was a practice, not a file. A remembered "~8.58 WIA band" had become a target. Its origin: it was simply the wounded-per-deployment that one balance.ts run printed the day aspect-vegetation shipped (terrain.ts:665) — a sim output, crowned as "historical." From then it flagged safer outcomes as defects (issue 026 logged a WIA of 6.92 as "below band, watch it"), set the stopping point for a realism win (issue 027), and reverted features (020, 022).

The shape: measure first. Four independent 12×50 draws of the identical gate config returned 7.08 / 2.67 / 5.00 / 9.42 WIA — the band was being policed at a resolution finer than the harness's own ±2.5 sampling noise. The fix is a rule, not a patch: a charter (docs/wiki/Harnesses.md) with one law — a gate may never assert the sim's own past output — and a re-ordering drawn from the design's soul, "win every firefight, still lose the valley": the COIN win condition now holds the standing gate the firefight wrongly held.

WIA across 4 identical-config draws"~8.58 band"2.67–9.42 (σ≈2.5) win condition (COIN)probe, never gatedstanding GATE (exit 1 if inert) tactical casualtiesa defended targeta diagnostic (σ floor printed) the suite itselfsuspected blockerhealthy — 1 bad habit, now ruled out
Read the full report (the evidence + the charter) →
Act · July 2–3

Making the mountains real

Four things felt fake — the way soldiers moved, the shape of the paths, the smooth walls, the staged outpost. The rule held: measure before you touch code. Two of the four complaints turned out to be the opposite of what they looked like.

The valley at noon after the campaign — banded strata walls and trails that switchback across the contour
The valley after: the walls carry cliff-band/bench striation and the foot-trails traverse the slope instead of charging down the fall line. Opens the full report.
2026-07-02 → 03

When "too squiggly" meant "too straight"

TerrainMovementThe KOPDoctrine

The owner named four realism complaints: soldiers get stuck, the paths are too squiggly, the terrain is too smooth, the outpost feels fake. Metricized first, two inverted. The "squiggly" paths were a near-beeline — planned route ratio 1.12, 66% of it straight down the fall line — and the visible wobble measured 0.16 px at play zoom; real trails needed to traverse, not de-wiggle. And "getting stuck" wasn't the patrols (router within 0.6–5% of optimal) — it was an in-combat freeze where a fighter breaking contact got a beeline into a cliff and re-issued the dead target forever: one unit frozen 678 s.

The shape: three build fronts, each grounded in cited doctrine — the uphill march re-anchored Naismith→FM 3-97.6, the foot-trails to the USFS half-rule, the walls to real Korengal gneiss-and-schist bedding (Junger, Restrepo). Then two commits fixed the test: a single unused RNG draw swung the COIN gate ±19 pts (relief-of-command is an opening-days lottery), so it now scores the paired best seed. Honest residual: the KOP was dressed, not reshaped — the perfect-circle generation rebuild is the next campaign.

insurgent time blocked in contact42%0% uphill march vs FM 3-97.62.45× too fast0.98 (doctrine-honest) trail grade a walker feels0.2150.114 (~USFS 10%) wall reversals / km (bedded rock)0–2.86.2–17.8 reachable valley (mean)60.6%75.2% roads stamped inside the wire90 / 8 seeds
Read the full report (four fronts, the two inversions, and the honest residuals) →
Two snapshots of the same squad at its most-strung march moment: baseline with the point man 52 m ahead, the fix with him reined to 40 m
The same squad, worst march moment: baseline (left) strings the point man 52 m clear of the pack; the fix (right) holds him to 40 m. Opens the full report.
2026-07-03

The point man waits

MovementSquad AIDoctrineHarnesses

Two complaints: the point man runs unrealistically far forward, and men get stuck on buildings in the villages, so the file spreads. The house rule is no fix without a number — so a probe attributed every stalled tick to the cell that actually blocked the man, and the named suspect was innocent. Village qalat walls cause 0.6 s of wedging per patrol; the grind is on the outpost's own b-huts (11.6 s) and terrain/the HESCO wire (38.6 s). I even implemented "make village walls cosmetic" — it moved the number by nothing, and was reverted.

The shape: the real defect is cause-agnostic — whatever a man snags on, the lead barely reacts (halted 1% of the time a follower is stuck). The fix is a genuine halt — the point man takes a knee — gated on a follower being blocked (not merely slow, the case issue 031 proved you must not touch) and capped at 45 s per leg. That budget makes it byte-identical on the tactical-window gate while it reins the file in. Two dead guards surfaced and were fixed: both keyed on a blockedTimer>6 the watchdog caps at 2 s — one a harness column that had read 0.0 s on every seed and told a prior pass "nobody is stuck."

point man halted while a follower is wedged1%38% nav speed while a man is wedged0.51 m/s0.32 m/s seeds with point man >40 m forward12 / 219 / 21 "buildings in villages" — qalat-wall wedgenamed cause0.6 s (innocent) tactical-window gate (issue-031 arbiter)15 / 2715 / 27 (identical) held-out survey-40…55 · new strandings0
Read the full report (the innocent suspect, the two dead metrics, and the budgeted halt) →
Act · July 16

Someone to hunt

The valley could simulate every bullet — but the enemy was a weather system and the campaign's harshest consequence was a dice roll. One session, three fronts: the insurgency becomes a persistent order of battle you learn by winning the population, relief-of-command starts reading an evidence file, and the docs that had begun anchoring fresh sessions to a stale reality were put under the same law as the code.

The Enemy Picture panel live: one cell MAPPED (Gul Wali's group, ~23 fighters, caches 1 found 1 destroyed), one LOCATED, and the mapped home marker on the rain-soaked valley floor
The Enemy Picture, live: one cell MAPPED with its strength, villages and cache ledger, one merely LOCATED, the rest honest ignorance — and the mapped home marker out on the valley floor. Opens the full report.
2026-07-16

The enemy gets a name

Enemy AICOINIntelCampaignHUD

The valley could simulate every bullet, but there was no one to hunt: the insurgency was two scalars and a tempo clock. Fighters materialized, fought, vanished — an escaped fighter cost the enemy nothing — and the game's sourced, reliability-scored intel feed pointed at weather. Now 3–5 persistent cells live in the draws: named leaders who survive between fights (kill one and the cell renames, weakens by a third, and remembers), munitions caches that IED ambushes must drain, and a patrol-heat grid that learns your habits — predictable routes get IED'd. The population is the sensor: a won-over village names its cell, then locates it, then maps it — the Enemy Picture panel, the map markers and the new weekly commander's assessment all read one shared intel gate, so the player sees exactly what the fiction has earned and never ground truth.

The shape: same session, the other rage-quit died — relief-of-command now reads an attributed evidence file (every confidence dock tagged casualties / civcas / directives) and fires only on a nameable pattern past a five-day grace, never a day-3 dice roll. And the honest part: the build shipped a conservation bug — exfil deposited +1 into a roster the fighter never left, printing strength 64→80-cap in a single hot game-day, and the acceptance probe asserted the bug because the contract's wording was wrong. The skeptical pass caught it; exfil is net-zero now, KIA exactly −1, and the probe proves conservation both ways.

careful tours censored by opening-days relief~50–60%0 / 3 persistent enemy state2 scalarscells · leaders · caches · heat Σ cell strengths vs derived scalar (probe, whole run)drift 0.00e+0 pinned-hot 1-day strength (the caught bug)64 → 80-cap64 → 23.6 (real attrition) enemy-network probe (new standing gate)9 / 9 OK combined COIN gate · best-pair spreadall 8 PASS93 (highest measured)
Read the full report (the weather system, the evidence file, the bug we shipped and caught, one live night) →
The method itself

How the machine that built this was built

The doctrine in "The five moves" above is not an afterthought — it was itself mined, tournamented, and verified. This is the most meta chapter, and the most important.

The CLAUDE.md build explainer
The full explainer of how the project's operating doctrine (CLAUDE.md) was built.
2026-06-06

147 MB of transcripts distilled into one file that every agent reads first

MethodologyUX

After 23 real build sessions the project had accumulated 147 MB of transcripts encoding hard-won patterns — verify-as-a-number, fan-out orchestration, adversarial passes — but none of it was reachable to a fresh agent starting cold. Every new session re-discovered the codebase and sometimes re-proposed approaches already measured and rejected.

The shape: a five-workflow pipeline (Mine → Orient → Generate → Tournament → Finalize) extracted the corpus signal, generated five candidate operating philosophies in parallel, ran each through an adversarial critic and a 3-judge panel, then synthesized a champion from the best spine with grafts from the runners-up.

transcript → corpus147 MB464 KB raw → ranked techniques9624 tournament: champion 193/210 vs runner-up 192/210

A CLAUDE.md is not documentation — it is latent-space conditioning: mined, distilled, and won.

Read the explainer →
The development archive's own front page
This page, documenting itself — the practice is now part of the done-gate.
2026-06-07

The archive that documents itself — making the work visible

MethodologyUX

The reports above had been piling up under docs/ — which never ships. A player deploying the game would never see any of it. So the scattered record was pulled into this one chronological archive, published into the part of the tree that does ship, and wired into the Field Manual and the title-screen menu. Transparency was made a feature, not an accident.

The shape: a survey fan-out (one reader per chapter, returning verbatim numbers under a strict schema) fed a hand-written synthesis; then the published layout was rendered through a real static server and every link asserted to resolve before it was called done — and the practice was written into CLAUDE.md so it can't be skipped next time.

reports reachable at deploy09 development story~18 folders1 archive, 21 chapters broken links0 / 30

Software usually hides its scaffolding. Here, showing the work is the work — so the archive ships inside the game.

Engineering record: docs/progress/2026-06-07-dev-archive/.

A note on honesty
Where a fix was partial, the report says so — the deploy got feedback, not raw speed; the world-scale audit fixed the cheap render half and deferred the expensive simulation half. Numbers here are quoted verbatim from the engineering record under docs/progress/ and the issue ledger under docs/issues/. Nothing was rounded up to "fixed."