feat(validator,generate,run): an identifier a proposal builds on must be in the input, or the verdict falls [skip-docs]
P6 (økt 108) ended in ValidatedProposal (verdict 5fd6272e3725fe68) on two cost codes -- M-04-01 / M-04-03 -- that appear in NO prompt of that run. Measured here first, verbatim: validate_proposal(p, baseline=None) validates it; the same proposal against any non-empty CostBaseline is rejected naming both codes. So the hole was never "fabrication goes uncaught" -- _reconcile_against_baseline exists and is right -- but that the falsifier is reached only through `if baseline is not None`. The input always exists; the baseline does not. New stage 0b (_ground_against_input), OUTSIDE the baseline branch, after stage 0 so an anchored run's message is byte-identical to before. ONE Rejection, the validator's own type, naming EVERY ungrounded identifier "; "-joined in the proposal's own order (økt 94's completeness reason). The rule has NO pattern -- `code in grounding`, exact substring -- and that is a measurement: over the delivered corpora (K2 1108 files / 2 005 561 chars, the three N payloads 8 excerpts each) the identifier forms are heterogeneous, and a pattern chosen to cover them would be a rule about shapes. Bare numerals are the one inert class (46 394 occurrences / 2 117 distinct in K2); the rule fails OPEN there, never closed. Evidence is three non-model-authored sources: what run_project DELIVERED (the rendered cut/pointer/chunks plus the base's context_files -- never files, which would make the type: verdict layer evidence), the project's own cost lines, and the baseline's codes when anchored. The rendered PROMPT is deliberately NOT evidence, on two measurements: gen_context IS the debate output on the S2c path, and from attempt 2 the prompt carries the previous Rejection.reason verbatim -- which for this stage QUOTES the identifier it just refused. Grounding in the prompt would let the gate's own refusal disarm it on its second round. Prose scanning was chosen against WITH THE NUMBERS: a typed gate catches 2/2 (P6) and 2/2 (S7c) -- 100% of what reached a verdict. What stays uncaught, said plainly: an ungrounded identifier that lives only in agent/debate prose and never becomes an affected_item code (2 of 4 P6, 2 of 4 S7c, 1 of 2 P4). Iron Law: 9 red / 2 green before the rule existed. Ten mutations all red against the whole suite, green control 1558 passed / 5 skipped (from 1543/5, superset, 0 removed), golden demo-transcript.stdout BYTE-UNCHANGED (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f). Three existing fixtures changed, no gate weakened -- most of all test_pre_amendment_bundle_runs_unchanged, which sent the SAME FABRICATED code and asserted it validated: the økt-108 hole written down as an expectation. No paid run. Order 20260909T113641Z-38938691-from-.claude. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
999846a485
commit
277bb95777
12 changed files with 831 additions and 13 deletions
|
|
@ -1053,8 +1053,16 @@ async def run_project(
|
|||
"a pre-pass payload declares a cut of a knowledge base, so it needs the bundle it "
|
||||
"was cut from; this run was given no bundle_dir"
|
||||
)
|
||||
# P7: the delivered base is the run's own evidence for what identifiers EXIST. Built from
|
||||
# ``context_files`` (MAJOR-3/S7a-3's rule), so the ``type: verdict`` layer stays out — a
|
||||
# proposal grounded in a prior verdict would reach the ExpeL fold's material around its gate.
|
||||
bundle_grounding = ""
|
||||
if bundle_dir is not None:
|
||||
bundle = okf.navigate_bundle(bundle_dir)
|
||||
bundle_grounding = "\n".join(
|
||||
"\n".join([f.name, *f.frontmatter.values(), f.body])
|
||||
for f in bundle.context_files
|
||||
)
|
||||
# ONE bundle-id rule (Step 10, slackened S7a-3 pkt. 1): the DECLARED id is the identity and
|
||||
# the mount is carried alongside, so a base delivered under a directory name of its own is
|
||||
# opened rather than refused. What is still refused, before a single model call: a base
|
||||
|
|
@ -1366,6 +1374,11 @@ async def run_project(
|
|||
# an attempt is helped by knowing it; it never enters the record, because the two
|
||||
# falsifiers are never blended.
|
||||
checker_verdict=checker_decision,
|
||||
# P7: what this run was GIVEN, as opposed to what the debate said about it.
|
||||
# ``context`` is the DELIVERED rendering (pre-pass cut / bundle pointer / retrieved
|
||||
# chunks) — never ``gen_context``, which on the debate path is the model's own
|
||||
# summary and would let a code the debate invented ground the proposal repeating it.
|
||||
grounding="\n".join([context, bundle_grounding]),
|
||||
)
|
||||
refinements.extend(generated.refinements)
|
||||
return generated.outcome
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue