feat(validator,generate,run): an identifier a proposal builds on must be in the input, or the verdict falls [skip-docs]

P6 (økt 108) ended in ValidatedProposal (verdict 5fd6272e3725fe68) on two cost
codes -- M-04-01 / M-04-03 -- that appear in NO prompt of that run. Measured
here first, verbatim: validate_proposal(p, baseline=None) validates it; the same
proposal against any non-empty CostBaseline is rejected naming both codes.

So the hole was never "fabrication goes uncaught" -- _reconcile_against_baseline
exists and is right -- but that the falsifier is reached only through
`if baseline is not None`. The input always exists; the baseline does not.

New stage 0b (_ground_against_input), OUTSIDE the baseline branch, after stage 0
so an anchored run's message is byte-identical to before. ONE Rejection, the
validator's own type, naming EVERY ungrounded identifier "; "-joined in the
proposal's own order (økt 94's completeness reason).

The rule has NO pattern -- `code in grounding`, exact substring -- and that is a
measurement: over the delivered corpora (K2 1108 files / 2 005 561 chars, the
three N payloads 8 excerpts each) the identifier forms are heterogeneous, and a
pattern chosen to cover them would be a rule about shapes. Bare numerals are the
one inert class (46 394 occurrences / 2 117 distinct in K2); the rule fails OPEN
there, never closed.

Evidence is three non-model-authored sources: what run_project DELIVERED (the
rendered cut/pointer/chunks plus the base's context_files -- never files, which
would make the type: verdict layer evidence), the project's own cost lines, and
the baseline's codes when anchored. The rendered PROMPT is deliberately NOT
evidence, on two measurements: gen_context IS the debate output on the S2c path,
and from attempt 2 the prompt carries the previous Rejection.reason verbatim --
which for this stage QUOTES the identifier it just refused. Grounding in the
prompt would let the gate's own refusal disarm it on its second round.

Prose scanning was chosen against WITH THE NUMBERS: a typed gate catches 2/2
(P6) and 2/2 (S7c) -- 100% of what reached a verdict. What stays uncaught, said
plainly: an ungrounded identifier that lives only in agent/debate prose and never
becomes an affected_item code (2 of 4 P6, 2 of 4 S7c, 1 of 2 P4).

Iron Law: 9 red / 2 green before the rule existed. Ten mutations all red against
the whole suite, green control 1558 passed / 5 skipped (from 1543/5, superset,
0 removed), golden demo-transcript.stdout BYTE-UNCHANGED (shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Three existing fixtures changed, no gate weakened -- most of all
test_pre_amendment_bundle_runs_unchanged, which sent the SAME FABRICATED code and
asserted it validated: the økt-108 hole written down as an expectation.

No paid run. Order 20260909T113641Z-38938691-from-.claude.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-09 16:04:01 +02:00
commit 277bb95777
12 changed files with 831 additions and 13 deletions

View file

@ -96,6 +96,10 @@ _CLAIM = 200.0
#: NOT ``energy_efficiency`` — that measure would additionally hit the method cap (0.15 * 1000 = 150)
#: and reject a claim of 200 for a reason that has nothing to do with this seam.
_MEASURE = "behovsstyrt_drift"
#: P7: the synthetic context must NAME the line the synthetic proposal restates, exactly as a
#: real run's input does. Without it the grounding stage rejects a sentinel code no input ever
#: mentioned — correctly, and for a reason that has nothing to do with this seam.
_CONTEXT = f"Price schedule line {_CODE} covers demand-driven operation."
def _wire_reply(*, with_band: bool) -> str:
@ -178,7 +182,7 @@ async def test_generation_call_carries_the_strict_schema() -> None:
project = load_reference_projects()[0]
client = _OptionsRecordingChatClient(_wire_reply(with_band=True))
await generate_via_llm(client, project, "", _meter(), max_attempts=1)
await generate_via_llm(client, project, _CONTEXT, _meter(), max_attempts=1)
# Control FIRST: a positive assert over an empty list would pass vacuously.
assert client.seen_options, "no generation call was observed — the assert below proves nothing"
@ -249,14 +253,14 @@ async def test_assumption_bands_keep_the_monte_carlo_falsifier_alive() -> None:
banded = await generate_via_llm(
_OptionsRecordingChatClient(_wire_reply(with_band=True)),
project,
"",
_CONTEXT,
_meter(),
max_attempts=1,
)
bandless = await generate_via_llm(
_OptionsRecordingChatClient(_wire_reply(with_band=False)),
project,
"",
_CONTEXT,
_meter(),
max_attempts=1,
)
@ -311,7 +315,7 @@ async def test_band_round_trip_is_verbatim_and_the_ir_map_form_still_parses() ->
from_array = await generate_via_llm(
_OptionsRecordingChatClient(_wire_reply(with_band=True)),
project,
"",
_CONTEXT,
_meter(),
max_attempts=1,
)
@ -330,7 +334,7 @@ async def test_band_round_trip_is_verbatim_and_the_ir_map_form_still_parses() ->
}
)
from_map = await generate_via_llm(
_OptionsRecordingChatClient(map_form), project, "", _meter(), max_attempts=1
_OptionsRecordingChatClient(map_form), project, _CONTEXT, _meter(), max_attempts=1
)
assert isinstance(from_map.outcome, ValidatedProposal), (
"the IR's own map form stopped parsing — the normalisation replaced rather than extended"