docs(p18): stress round 2 -- five paid runs, and the mutation that found a hole
P18 parts D and E (order 20260914T105139Z), plus the two things measuring them turned up. DEL D -- five paid runs (gpt-4-1-mini, azure, PACE_SECONDS=2), same four context sets, SAME parameters on all four (--max-rounds 3 --max-tokens 600000), plus one variance repeat of gate-nordvik. All four free --live-dry-runs first: rc 0, and Grounding-offer numbers IDENTICAL to round 1 (435/272/982/3) -- the control that the Grounding structure did not change what the gate measures. MEASURED, round 1 -> round 2: - runs that died on the token cap: 3 of 7 -> 0 of 5, and two sets now fit a LOWER cap than round 1 had to give them; - guessed read paths: 12 -> 0; - wall time, the two sets whose parameters are directly comparable: n100 140.1 s -> 70 s, n500 73.5 s -> 69 s; - 5 of 31 read_file calls opened documents BEYOND the default window, so the window WAS widened -- the trace does not record which knob (finding 1); - (a) grounded in a fasit concept: 0 of 26, UNCHANGED. That is the mission gap, and DEL A did not close it. The a4 falsification arm fails once in each round, on a different set. r761's a4 is now rejected -- but NOT by B1: the model proposed "Kontraktsum" this time, so the "appears nowhere" arm caught it, and B1's effect on that row is proven offline, not live. NEW failure: tunnel-hauglia a4 VALIDATED on "impulsventilator" (3/270 documents), and fv412 a1 on "bituminost barelag" (4/1133). Both are ordinary Norwegian words from the standard's prose, not cost codes. B1 cannot and should not fell them: this is an ANCHORING defect, not a grounding one, and it is finding 2 with two named remedies and a recommendation. C2 isolated by re-judging round 1 with the new judge: kontrakt-sorasen goes named=3 -> named=1, and the survivor is named_in_measure -- the model's own words. Two of the three were the whole-base snippet artefact. DEL E -- docs/2026-09-14-p18-stressrunde-2.md: round 1 against round 2, what each fix bought (measured, never attributed), the B2 table, variance, and for EACH remaining ugly finding a NAMED solution with an estimate. THE MUTATION THAT FOUND A HOLE. B6 (revert run.py to compose ONE blob instead of one document per concept file) left the WHOLE suite green: 1698 passed / 5 skipped. The composition arm drives _grounding_text with a Grounding it builds ITSELF, so it cannot see what the RUN handed over -- and a blob has exactly one boundary, so the floor can never be reached, the share can never fire, and the measured defect is back intact. The rule is only as good as the boundaries it is given. Arm (h) is the gate that was missing: a crafted base with TWELVE concept files all carrying the same token -- per document 12 of 17 and inert, as one blob 1 of 1 and grounding -- with a control on a code only ONE file carries, which must still validate. Measured RED against exactly that mutation. The mutation was not dropped and the seam was not declared unwitnessed: it got a witness. Also: the debate's own bundle pointer (run.py _bundle_pointer) now explains the window and the filter, alongside the tool description and the navigator instruction updated in9b47e5a-- a description that lies about the body IS the model's instruction (the Fase 3 class). Golden transcript unaffected. Mutations, all against the FULL suite in an isolated worktree, one at a time: DEL A 7 of 7 red (control 1685/5), DEL B+C 9 of 10 red (control 1698/5), the tenth being B6 above. Tables in the report s 9. Verification: uv run pytest -q 1699 passed / 5 skipped (1670 oncfd9079; +29, 0 removed). ruff check + format clean, mypy clean (38 files). Golden demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
7a7c988253
commit
e7ba367f9d
4 changed files with 432 additions and 5 deletions
|
|
@ -219,3 +219,76 @@ def test_the_run_hands_the_gate_one_document_per_concept_file() -> None:
|
|||
assert composed.text == "\n".join(
|
||||
["dokument A", "dokument B", "dokument C", "PRJ-01", "BAS-01"]
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_boundaries_survive_the_real_run_and_not_only_the_composer(
|
||||
tmp_path: Path,
|
||||
) -> None:
|
||||
"""(h) BEHAVIOURAL, over the real ``run_project`` bundle arm — and it exists because a mutation
|
||||
found the gap, not because it was foreseen.
|
||||
|
||||
Reverting ``run.py`` to compose ONE blob instead of one document per concept file left the WHOLE
|
||||
suite green (1698 passed / 5 skipped). Arm (g) above drives ``_grounding_text`` with a
|
||||
``Grounding`` it builds itself, so it can never see what the RUN handed over — exactly the
|
||||
vacuity this repo keeps measuring. The rule is only as good as the boundaries it is given, and
|
||||
a blob has exactly one: with a single document the absolute floor can never be reached, so the
|
||||
share can never fire and the measured defect returns intact.
|
||||
|
||||
The base is crafted so the two implementations must DISAGREE: twelve concept files all carrying
|
||||
the same token, which is past ``_GROUNDING_MIN_INERT_DOCUMENTS``. Per document it is in 12 of 13
|
||||
and inert; as one blob it is in 1 of 1 and grounds. The CONTROL is the same run with a code only
|
||||
ONE file carries, which must still validate — otherwise the arm would also pass on a run that
|
||||
rejects everything.
|
||||
"""
|
||||
import shutil
|
||||
|
||||
from portfolio_optimiser.run import run_project
|
||||
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||
|
||||
base = tmp_path / "base"
|
||||
shutil.copytree(Path("shared/examples/bygg-energi-mikro"), base)
|
||||
links = []
|
||||
for n in range(12):
|
||||
rel = f"seksjon-{n:02d}.md"
|
||||
(base / rel).write_text(
|
||||
f"---\ntype: concept\ntitle: Seksjon {n:02d}\n---\n\n"
|
||||
"KORPUSMERKE-77 gjelder overalt i denne basen.\n"
|
||||
+ ("Kostlinjen EN-ENESTE-01 star bare her.\n" if n == 0 else ""),
|
||||
encoding="utf-8",
|
||||
)
|
||||
links.append(f"- [Seksjon {n:02d}]({rel})")
|
||||
index = base / "index.md"
|
||||
index.write_text(
|
||||
index.read_text(encoding="utf-8") + "\n" + "\n".join(links) + "\n", encoding="utf-8"
|
||||
)
|
||||
|
||||
async def _outcome(code: str) -> object:
|
||||
reply = (
|
||||
f'{{"measure":"LED-retrofit","affected_items":'
|
||||
f'[{{"code":"{code}","quantity":300000,"unit_cost":1.0}}],'
|
||||
f'"claimed_saving_nok":30000}}'
|
||||
)
|
||||
result = await run_project(
|
||||
"BYGG-KONTOR-NORD",
|
||||
"local",
|
||||
docs_dir=str(base),
|
||||
bundle_dir=str(base),
|
||||
client_factory=lambda role: ScriptedChatClient(
|
||||
"Reasoning holds.\nVERDICT: APPROVE" if role == "checker" else reply, role=role
|
||||
),
|
||||
)
|
||||
return result.outcome
|
||||
|
||||
control = await _outcome("EN-ENESTE-01")
|
||||
assert isinstance(control, ValidatedProposal), (
|
||||
f"CONTROL: a code only one concept file carries must still ground — {control}"
|
||||
)
|
||||
|
||||
everywhere = await _outcome("KORPUSMERKE-77")
|
||||
assert isinstance(everywhere, Rejection), (
|
||||
"a token every concept file carries grounded a proposal — the run handed the gate one blob"
|
||||
)
|
||||
# The NUMERATOR is the discriminator: as one blob the token is in 1 of 1, so a refusal naming
|
||||
# 12 can only come from a run that kept the concept files apart.
|
||||
assert "appears in 12 of the " in everywhere.reason, everywhere.reason
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue