Commit graph

1 commit

Author SHA1 Message Date
3f04f90261 docs(p22): stress round 6 -- the three findings closed, and what the closing exposed
Five paid runs, rc 0 on all five, 913 320 tokens, no round cap hit, anchored 6 of 6.

DEL A hit hard: naming the project's own cost codes in stage 0's refusal took `priced` from 0 of 20
to 16 of 20 and validated approaches from 0 of 20 to 10 of 20, with invented cost codes down from 26
to 12. The largest single movement any part has produced in six rounds.

DEL B moved a number three rounds had not: declarations pointing at an answer-key concept went
0 of 13 -> 0 of 12 -> 5 of 16, and requirement_hit per approach 0 of 20 -> 3 of 20.

DEL C halved the guessing: read_file against a document the base does not hold went 7 of 52 to
3 of 51, read_dir against a level it does not hold 8 of 128 to 2 of 97. Two of the three remaining
read_file misses got the NEW document clause, so DEL C fired live.

AND THE CLOSING EXPOSED SOMETHING LARGER. must_refuse -- the falsification arm, the one approach per
set the base has no ground for -- went 5 of 5 to 2 of 5. Three were VALIDATED. P21's 5 of 5 was not
sharp: it was achieved because the model invented a cost code, not because the base lacks ground.
The order predicted exactly this. Now the arm discriminates, and it says nothing in the gate speaks
to whether the base supports the direction. That is finding 1 and it stands first, with a named
solution and an operator decision attached.

Also recorded: the order's own cited figures re-measured with denominators before anything was built
on them. "requirement_hit 0 of 12" mixes a per-approach field (0 of 20) with the declaration count
(0 of 12); "0 of 20 validated" and "26 rejections" are two different populations (20 approach rows
vs 20 + 6 own-proposals); and the read-miss figures hold under the context_files definition while
the judge's filesystem one gives 6 of 52, not 7. Rounds 4 and 5 are re-judged with the same
instrument; round 3 is deliberately NOT, because filling in anchored/priced/stage from prose would
be inventing a measurement.

Honesty limits: five paid runs on one question are not a sample; every amount in the price schedules
is invented and the FORM is what is not; the cost figure is an ASSUMPTION at list price with no
invoice read; and must_refuse 2 of 5 is not evidence the system got worse, it is evidence the
previous measurement could not discriminate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 01:50:40 +02:00