feat(run,generate): a run says what its delivered input can ground, before it spends an attempt [skip-docs]
P7 is right and landed, but re-measuring it exposed a consequence no row stated: with the gate live, 29 of 29 cost codes in 13 of 13 delivered proposals fall across the three free recordings (PM's denominator; 14 of 14 in 8 proposals on the PARSEABLE one -- the five blobs that separate the numbers are refused by pydantic's `claimed <= total` and never reach stage 0b). All 29 were invented, so the gate is right; but a gate that always refuses is as useless as one that never does. The cause is that the PROMPT asks for something the input cannot supply. `_build_messages` requires each affected_item to "restate a cost line as the project's price schedule already carries it", while K2's delivered input carries 9 occurrences / 2 distinct code-shaped tokens -- `SHA-01`/`SHA-10`, both document numbers off a page footer -- and `derive_cost_baseline` refuses the base outright. There is no cost line in it to restate. `GroundingOffer(chars, identifiers, cost_lines)` reports it. The PAIR is the diagnosis: "50 identifiers, 0 cost lines" says what neither number says alone. A REPORT, never a gate -- it blocks nothing, because a blocking requirement IS `--require-cost-baseline` (F4/D-3, opt-in, untouched), and `_ground_against_input` is untouched. The callsite is MEASURED, not chosen: `generate.py` composes the grounding per attempt, after `await _fetch_parsed`, so a report there could only speak once an attempt had been paid for; `run.py` binds both halves above the `--live-dry-run` cut and before the first `debate.run`, so the FREE trip says it. `delivered` is bound ONCE and the same variable feeds the report and `_evaluate`; the report composes THROUGH `_grounding_text`, the gate's own composer. A pattern is admissible here and not in the gate, and that is the difference between a report and a falsifier: an unknown form is a token left uncounted -- an under-count, never a false rejection. The forms are transcribed from the measurement; bare numbers are excluded with the number (46 394 / 2 117 in K2). `grounding_offer_notice` is the ONE renderer and is silent when the run CAN anchor -- omission, never an empty row. Load-bearing MEASURED (tests/test_grounding_offer_loadbearing.py, 12 arms), nine mutations all red against the WHOLE suite + green control 1570/5 (from 1558/5, strict superset, 0 removed) and the golden byte-unchanged (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f). Measurement: docs/2026-09-09-p8-forankringstilbudet.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
277bb95777
commit
455d611660
5 changed files with 767 additions and 2 deletions
342
tests/test_grounding_offer_loadbearing.py
Normal file
342
tests/test_grounding_offer_loadbearing.py
Normal file
|
|
@ -0,0 +1,342 @@
|
|||
"""P8 — the run SAYS what its delivered input can ground, before it spends an attempt on a proposal
|
||||
that cannot be grounded.
|
||||
|
||||
P7 (økt 109) is right and landed: an identifier a proposal builds on must appear VERBATIM in the
|
||||
input, or the verdict falls. Re-measuring it exposed a CONSEQUENCE no row stated. Over the three
|
||||
free recordings, with the gate live, **29 of 29 cost codes in 13 of 13 delivered proposals are
|
||||
ungrounded** — every one of them invented. That is not a fault in the gate. It is that the PROMPT
|
||||
asks for something the delivered input cannot supply: ``_build_messages`` says each entry "must
|
||||
restate a cost line as the project's price schedule already carries it", and MEASURED, K2's
|
||||
delivered input carries **two** code-shaped tokens in the prompt (``SHA-01``/``SHA-10``), both of
|
||||
them document numbers off a page footer, and ``derive_cost_baseline`` refuses the base outright —
|
||||
there is no cost line in it to restate. A gate that always refuses is as useless as one that never
|
||||
does, so the run must be able to say which of the two situations it is in.
|
||||
|
||||
**This is a REPORT, never a gate.** It does not block: a run with a null offer still runs, because
|
||||
a blocking requirement is exactly ``--require-cost-baseline``, which F4 settled as opt-in and this
|
||||
order freed. It also does not touch ``_ground_against_input``, which ten mutations hold.
|
||||
|
||||
Arms, each with a named detach point:
|
||||
|
||||
* **(a) the null offer is stated.** A delivered text with no identifier of any measured form and no
|
||||
baseline reports ``identifiers=0, cost_lines=0``. This is the K2 CONTROL — without it a green
|
||||
known positive proves nothing.
|
||||
* **(b) the positive offer is stated.** N100's delivered input carries the BINDING known positive
|
||||
``Krav 3.3.1—13`` (P7 premiss (vii): it stands in 6 of 6 prompts, EM-DASH U+2014), and the offer
|
||||
is positive there. Paired with a control that the search CAN find the string at all.
|
||||
* **(c) the offer is measured on the SAME text the gate will see.** ``grounding_offer`` composes
|
||||
through ``generate._grounding_text`` — the one composer P7 uses — so a report built from a second
|
||||
rendering, free to disagree with the one that was sent, is impossible by construction.
|
||||
* **(d) the offer reaches the outcome the operator reads**, on BOTH surfaces: ``RunResult`` after a
|
||||
full run and ``DryRunReport`` before the first model call. The dry-run arm is what proves the
|
||||
measurement happens BEFORE the three attempts are spent — its client factory RAISES.
|
||||
* **(e) ONE renderer**, silent when the run CAN anchor a cost line (omission, never an empty row —
|
||||
``cost_baseline_notice``'s rule), speaking when it cannot.
|
||||
* **(f) a run that exhausts its attempts on ungrounded identifiers leaves the input-side diagnosis
|
||||
in the outcome**, not only "the model missed".
|
||||
|
||||
A NEW file rather than an extension of ``test_identifier_grounding_loadbearing.py``, stated as the
|
||||
order asks: that file's subject is the GATE and it is held by ten mutations. A report mutated here
|
||||
would go red in a file whose docstring promises a falsifier, and which seam a red test names is the
|
||||
only thing a mutation table is for.
|
||||
|
||||
The fixtures are P7's own tracked generation prompts. Tracked test data on purpose (``scratchpad/``
|
||||
is absent from ``git archive HEAD``) — and REUSED rather than extended with the delivered cuts,
|
||||
because those cuts are 10 kB and 96 kB of a real Norwegian tender and a public road standard, and
|
||||
committing corpus text into a repo published on ``open/`` is a publication decision that belongs to
|
||||
the operator, not to this order. The 8-distinct / 435-distinct figures for the delivered cut and the
|
||||
full grounding are MEASURED and reported in ``docs/2026-09-09-p8-forankringstilbudet.md``; what a
|
||||
test asserts here is the property, on text this repo already ships.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from portfolio_optimiser.generate import GroundingOffer, _grounding_text, grounding_offer
|
||||
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
|
||||
from portfolio_optimiser.reference_domain import CostItem, Project
|
||||
from portfolio_optimiser import run as run_mod
|
||||
from portfolio_optimiser.run import grounding_offer_notice, run_project
|
||||
from portfolio_optimiser.simulation import scripted_factory
|
||||
from portfolio_optimiser.validator import Rejection
|
||||
|
||||
FIXTURES = Path(__file__).parent / "fixtures" / "p7-grounding"
|
||||
SHARED = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
|
||||
#: P7 premiss (vii), the BINDING known positive: the model quoted this requirement number verbatim
|
||||
#: in 6 of 6 replies and it stands in 6 of 6 prompts. EM-DASH (U+2014); the hyphen variant is 0/6.
|
||||
KNOWN_POSITIVE = "Krav 3.3.1—13"
|
||||
|
||||
#: The base with a ``cost-baseline.json`` (one line) and the one without — the pair that makes the
|
||||
#: renderer's omission arm reachable rather than asserted.
|
||||
ANCHORED = SHARED / "veglys-fv-soer"
|
||||
UNANCHORED = SHARED / "bygg-energi-mikro"
|
||||
_UNANCHORED_PID = "BYGG-KONTOR-NORD"
|
||||
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (sim)"}
|
||||
_VALID_REPLY = (
|
||||
'{"measure":"LED-retrofit","affected_items":'
|
||||
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],'
|
||||
'"claimed_saving_nok":30000}'
|
||||
)
|
||||
#: A code that appears NOWHERE in the base — the ungrounded case, which P7's stage 0b refuses on
|
||||
#: every attempt, so the run exhausts its attempts and returns the last ``Rejection``.
|
||||
_UNGROUNDED_REPLY = (
|
||||
'{"measure":"invented","affected_items":'
|
||||
'[{"code":"M-04-01","quantity":300000,"unit_cost":1.0}],'
|
||||
'"claimed_saving_nok":30000}'
|
||||
)
|
||||
_CHECKER_REPLY = "Reasoning holds.\nVERDICT: APPROVE"
|
||||
|
||||
|
||||
def _prompt(name: str) -> str:
|
||||
return (FIXTURES / name).read_text(encoding="utf-8")
|
||||
|
||||
|
||||
def _project(*codes: str) -> Project:
|
||||
"""A project shaped like the bundle path's: ``_project_from_bundle`` builds ``cost_items=()``
|
||||
(MEASURED, P7's own row), so the default carries no codes of its own and the offer measures the
|
||||
delivered text alone."""
|
||||
return Project(
|
||||
id="X",
|
||||
name="X",
|
||||
description="",
|
||||
currency="NOK",
|
||||
cost_items=tuple(
|
||||
CostItem(code=c, description=c, quantity=1.0, unit="stk", unit_cost=1.0)
|
||||
for c in codes
|
||||
),
|
||||
docs_dir="",
|
||||
)
|
||||
|
||||
|
||||
def _baseline(*codes: str) -> CostBaseline:
|
||||
return CostBaseline(
|
||||
project_id="X",
|
||||
items={c: CostBaselineLine(quantity=1.0, unit_cost=1.0) for c in codes},
|
||||
)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ (a) the null offer is stated
|
||||
|
||||
|
||||
def test_a_delivered_input_with_no_cost_line_reports_zero() -> None:
|
||||
"""LOAD-BEARING (a) — the K2 CONTROL. Both K2 generation prompts carry no identifier of any
|
||||
measured form and the run is un-anchored, so the offer is null in BOTH numbers. Without this
|
||||
control a green (b) would be satisfied by a report that always counts positive."""
|
||||
for name in ("p6-k2-generation-prompt.txt", "s7c-k2-generation-prompt.txt"):
|
||||
offer = grounding_offer(_project(), None, _prompt(name))
|
||||
assert offer.identifiers == 0, (name, offer)
|
||||
assert offer.cost_lines == 0, (name, offer)
|
||||
assert offer.chars > 0, "the measurement must have had text to measure"
|
||||
|
||||
|
||||
# -------------------------------------------------------------- (b) the positive offer is stated
|
||||
|
||||
|
||||
def test_the_known_positive_input_reports_a_positive_offer() -> None:
|
||||
"""LOAD-BEARING (b) — the BINDING known positive. Paired with the control that the string is
|
||||
actually there: a report that finds nothing because it searched for nothing would otherwise
|
||||
pass (a) and look measured."""
|
||||
text = _prompt("p4-n100-generation-prompt.txt")
|
||||
assert KNOWN_POSITIVE in text, "the fixture no longer carries the known positive"
|
||||
offer = grounding_offer(_project(), None, text)
|
||||
assert offer.identifiers > 0, offer
|
||||
assert offer.cost_lines == 0, "a road standard carries no cost lines (F4's own finding)"
|
||||
|
||||
|
||||
def test_a_bare_number_is_not_counted_as_an_offer() -> None:
|
||||
"""LOAD-BEARING (b), the other half. K2 carries 46 394 bare-number occurrences over 2 117
|
||||
distinct values (P7 § 2), so counting them would make every report positive and the whole
|
||||
measurement inert — the repo's cardinal class, a gate that can only come out green."""
|
||||
offer = grounding_offer(_project(), None, "1234 5678 90 42.5 1000000")
|
||||
assert offer.identifiers == 0, offer
|
||||
|
||||
|
||||
# ----------------------------------------------- (c) measured on the SAME text the gate will see
|
||||
|
||||
|
||||
def test_the_offer_is_measured_on_the_text_the_gate_will_see() -> None:
|
||||
"""LOAD-BEARING (c) — requirement (2). The offer composes through ``_grounding_text``, the one
|
||||
composer P7's gate uses, so the report and the gate cannot describe different texts. Both of
|
||||
the composer's OTHER two sources are exercised: a project cost line and a baseline code each
|
||||
raise the count, which a report built from ``delivered`` alone cannot do."""
|
||||
delivered = "nothing citable here"
|
||||
project, baseline = _project("PRJ-77"), _baseline("BAS-88")
|
||||
|
||||
offer = grounding_offer(project, baseline, delivered)
|
||||
|
||||
assert offer.chars == len(_grounding_text(project, baseline, delivered))
|
||||
assert offer.identifiers == 2, offer
|
||||
assert offer.cost_lines == 1, offer
|
||||
|
||||
|
||||
# ---------------------------------------------------------- (e) ONE renderer, omission when able
|
||||
|
||||
|
||||
def test_the_notice_speaks_on_a_null_offer_and_is_silent_when_anchored() -> None:
|
||||
"""LOAD-BEARING (e). Silent when the run CAN anchor a cost line (omission, never an empty row);
|
||||
speaking when it cannot, and carrying BOTH numbers — "50 identifiers, 0 cost lines" is the
|
||||
diagnosis, and neither number alone says it."""
|
||||
assert grounding_offer_notice(GroundingOffer(chars=9, identifiers=3, cost_lines=1)) is None
|
||||
assert grounding_offer_notice(None) is None
|
||||
|
||||
line = grounding_offer_notice(GroundingOffer(chars=9, identifiers=3, cost_lines=0))
|
||||
assert line is not None
|
||||
assert "3" in line and "0" in line
|
||||
|
||||
|
||||
# --------------------------------------- (d) the offer reaches BOTH surfaces the operator reads
|
||||
|
||||
|
||||
async def test_the_dry_run_carries_the_offer_before_any_model_call() -> None:
|
||||
"""LOAD-BEARING (d), and the proof of "BEFORE it spends its three attempts".
|
||||
|
||||
Asserted on CALLS, never on client CONSTRUCTION: ``fresh_workflow`` builds the proposer and
|
||||
checker clients EAGERLY, above the dry-run cut (``workflow.py:64``), so a factory that raised
|
||||
would be red against a working implementation. An empty sink is the measurement — session 57's
|
||||
rule that a refusal after the spend is indistinguishable from one before it at the exit code.
|
||||
"""
|
||||
sink: list[str] = []
|
||||
|
||||
report = await run_project(
|
||||
_UNANCHORED_PID,
|
||||
"local",
|
||||
docs_dir=str(UNANCHORED),
|
||||
bundle_dir=str(UNANCHORED),
|
||||
client_factory=scripted_factory({"proposer": "x", "checker": "x"}, sink),
|
||||
live_dry_run=True,
|
||||
)
|
||||
|
||||
assert sink == [], "the offer was measured only after a model call was made"
|
||||
|
||||
offer = getattr(report, "grounding_offer", None)
|
||||
assert offer is not None, "the dry run says nothing about what its input can ground"
|
||||
assert offer.cost_lines == 0, offer
|
||||
assert offer.chars > 0, offer
|
||||
|
||||
|
||||
async def test_the_full_run_carries_the_offer() -> None:
|
||||
"""LOAD-BEARING (d), the other surface — requirement (3). A measurement that never leaves
|
||||
``run_project`` is a log line, not a report."""
|
||||
result = await run_project(
|
||||
_UNANCHORED_PID,
|
||||
"local",
|
||||
docs_dir=str(UNANCHORED),
|
||||
bundle_dir=str(UNANCHORED),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
client_factory=scripted_factory(
|
||||
{"proposer": _VALID_REPLY, "checker": _CHECKER_REPLY}, []
|
||||
),
|
||||
)
|
||||
|
||||
offer = getattr(result, "grounding_offer", None)
|
||||
assert offer is not None, "the run says nothing about what its input can ground"
|
||||
assert offer.cost_lines == 0, offer
|
||||
|
||||
|
||||
async def test_an_anchored_run_reports_its_cost_lines() -> None:
|
||||
"""LOAD-BEARING (d), the CONTROL that ``cost_lines`` is not a constant zero: the anchored base
|
||||
ships exactly one line, and the run reports it."""
|
||||
report = await run_project(
|
||||
"VEGLYS-FV-SOER",
|
||||
"local",
|
||||
docs_dir=str(ANCHORED),
|
||||
bundle_dir=str(ANCHORED),
|
||||
client_factory=scripted_factory({"proposer": "x", "checker": "x"}, []),
|
||||
live_dry_run=True,
|
||||
)
|
||||
|
||||
offer = getattr(report, "grounding_offer", None)
|
||||
assert offer is not None and offer.cost_lines == 1, offer
|
||||
|
||||
|
||||
# ------------------------------------- (f) an exhausted run leaves the INPUT-side diagnosis behind
|
||||
|
||||
|
||||
async def test_a_run_exhausted_on_ungrounded_identifiers_says_the_input_was_the_problem() -> None:
|
||||
"""LOAD-BEARING (f). P7's gate refuses every attempt, the loop returns its last ``Rejection``
|
||||
(``generate.py``: ``last_ruling`` after the attempt loop — it does not crash), and the outcome
|
||||
the operator reads carries the input-side fact as well: this input could anchor NOTHING, so no
|
||||
attempt could ever have succeeded. Without it the record says only that the model missed."""
|
||||
result = await run_project(
|
||||
_UNANCHORED_PID,
|
||||
"local",
|
||||
docs_dir=str(UNANCHORED),
|
||||
bundle_dir=str(UNANCHORED),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
client_factory=scripted_factory(
|
||||
{"proposer": _UNGROUNDED_REPLY, "checker": _CHECKER_REPLY}, []
|
||||
),
|
||||
)
|
||||
|
||||
assert isinstance(result.outcome, Rejection)
|
||||
assert "M-04-01" in result.outcome.reason
|
||||
offer = getattr(result, "grounding_offer", None)
|
||||
assert offer is not None and offer.cost_lines == 0, offer
|
||||
assert grounding_offer_notice(offer) is not None
|
||||
|
||||
|
||||
# ------------------------------------------------ (d) the CLI prints it, on BOTH free and paid
|
||||
|
||||
|
||||
def test_the_cli_dry_run_prints_the_offer(capsys) -> None:
|
||||
"""LOAD-BEARING (d), stdout. A renderer that returns the line while no caller prints it is a
|
||||
measurement the operator never sees — its own seam, so its own arm (the ``cost_baseline_notice``
|
||||
precedent, whose CLI print carries a mutation of its own)."""
|
||||
rc = run_mod.main(
|
||||
[
|
||||
_UNANCHORED_PID,
|
||||
"--docs-dir",
|
||||
str(UNANCHORED),
|
||||
"--bundle-dir",
|
||||
str(UNANCHORED),
|
||||
"--live-dry-run",
|
||||
]
|
||||
)
|
||||
assert rc == 0
|
||||
assert "Grounding offer" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_the_cli_says_nothing_when_the_run_can_anchor(capsys) -> None:
|
||||
"""CONTROL for the omission arm on the SAME surface: the anchored base prints no offer line at
|
||||
all. Without it, a notice that always fired would pass the arm above."""
|
||||
rc = run_mod.main(
|
||||
[
|
||||
"VEGLYS-FV-SOER",
|
||||
"--docs-dir",
|
||||
str(ANCHORED),
|
||||
"--bundle-dir",
|
||||
str(ANCHORED),
|
||||
"--live-dry-run",
|
||||
]
|
||||
)
|
||||
assert rc == 0
|
||||
assert "Grounding offer" not in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_the_cli_full_run_prints_the_offer(tmp_path: Path, capsys) -> None:
|
||||
"""LOAD-BEARING (d), the paid surface's stdout — so the line is a property of a RUN and not of
|
||||
the dry-run branch alone."""
|
||||
replies = tmp_path / "replies.json"
|
||||
replies.write_text(
|
||||
json.dumps({"proposer": _VALID_REPLY, "checker": _CHECKER_REPLY}), encoding="utf-8"
|
||||
)
|
||||
rc = run_mod.main(
|
||||
[
|
||||
_UNANCHORED_PID,
|
||||
"--docs-dir",
|
||||
str(UNANCHORED),
|
||||
"--bundle-dir",
|
||||
str(UNANCHORED),
|
||||
"--scripted-replies",
|
||||
str(replies),
|
||||
"--decision",
|
||||
"approved",
|
||||
"--rationale",
|
||||
"expert reviewed (test)",
|
||||
]
|
||||
)
|
||||
assert rc == 0
|
||||
assert "Grounding offer" in capsys.readouterr().out
|
||||
Loading…
Add table
Add a link
Reference in a new issue