Five paid runs, rc 0 in all, 1 130 145 tokens (~NOK 7,6 under P20's stated list-price
assumption; no invoice read).
What the anchoring bought, measured off the artefacts rather than inferred from the
flag: cost_baseline_anchored true 6/6, and the falsification arm a4 passed 5 of 5 with
ALL FIVE caught by stage0-baseline -- the one stage that knows what the project buys.
Round 4, re-judged with the same instrument: 4 of 5, every one of them on stage 0b, and
the fifth VALIDATED.
C1 fired live and changed behaviour: distinct documents opened before a declaration went
1/1/1/2/5/13 -> 3/3/5/7/11/12, and THREE declarations were refused mid-run (15 calls, 12
recorded) after which the model read more and declared again. C2: read_dir against a
level the base does not hold fell from 16 of 104 to 8 of 128.
And what it cost, measured just as plainly: 0 of 20 approaches validated (round 4: 4).
All 26 rejections read "unknown cost code" -- because the schedule reaches the VALIDATOR
and never the prompt, and stage 0's refusal names HOW MANY codes the project has and not
WHICH. Step 5 feeds that refusal back verbatim, and "you guessed wrong, there are five
right ones" cannot be corrected. Contrast the magnitude arm, which names the baseline
value and therefore converges.
Four remaining findings, each with a NAMED solution and an estimate -- the first is one
sentence in _reconcile_against_baseline.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>