P18 parts B and C (order 20260914T105139Z).
B1 -- stage 0b. P7 made it `item.code in grounding`: plain containment over
ONE concatenated string. P16 ran it against a delivered corpus and measured
what containment cannot tell apart: the falsification arm a4-indeksregulering
put 250 000 NOK on a single cost line coded R761 -- the knowledge base's OWN
NAME, carried by all 2 756 of its concept documents -- and the whole gate
said validated (stage 0 skipped, un-anchored run; checker approve).
The grounding is now carried as the DOCUMENTS it is made of (validator.
Grounding), not as a blob. A structure and not a second argument beside the
text: the boundaries and the text are one fact, and .text is derived, so the
gate and P8's report measure the same characters. run.py composes one
document per concept file where the base is already walked; generate.
_grounding_text folds each cost line in as a one-line document.
N and A are MEASURED, not chosen (14.09, four mounted vegnormal bases):
- every must_cite ref and mandate affected_code in the four context sets --
shortest real identifier is FOUR characters (12.1, 52.1), so N = 3 sits one
below the measurement and cannot refuse anything measured;
- document frequency of every code-shaped token per base -- 1 692 distinct
and NOT ONE reaches 5 %. Highest anywhere 6/446 (1.35 %), highest a fasit
names 3/446 (0.67 %), R761 2 756/2 756 (100 %). A = 0.05 therefore sits
3.7x above the highest real token and 20x below the defect.
Length is NOT what makes the defect inert (R761 is four characters); the
share is. And a share is not a measurement without a denominator big enough
to take one (ansikt 4): one of three is 33 %, so an ABSOLUTE floor of 10
documents gates it. Highest absolute count any real identifier reaches is 6,
and every fixture in the repo is far below 10 -- which is why every pre-P18
gate is UNTOUCHED by this rule rather than exempted from it. Grounding.of
(one document) can never reach the floor by construction.
The refusal NAMES the denominator ("appears in 2756 of the 2756 documents
this run was given"), because Step 5 feeds that reason verbatim into the next
attempt's prompt: a proposer told only "ungrounded" answers with another
token of the same kind.
B2 SPIKE (measured, NOT built) FELLED the order's own alternative: option (b)
"ground in what the run OPENED" was run over P16's 16 code rows -- R761
stands in every OPENED document too, so (b) would NOT have caught the defect,
while B1 makes it inert and still grounds the real process line 65
ASFALTDEKKER (29/2756 = 1.05 %). (b) is not a substitute for B1.
C1 -- --docs-dir is optional once --bundle-dir is given (P16 FUNN 2). On the
bundle path docs_dir is never read: retrieval, the chunk tool and the "no
citable content" check all live in the road branch. Bound ONCE from
--bundle-dir, which is byte-identically what the README already tells an
operator to type by hand. NOT the "--docs-dir omvei": no such path is opened
and the road branch still refuses without a real --docs-dir (own arm).
C2 -- the judge's snippet arm counts only under citation_scope == "narrowed",
as (a) already does (PM decision, P16 s 6.2). P16's reason for (b') being
clean -- snippets are bodies while ref/title live in frontmatter, 0 of 446
n100 bodies -- holds for "Krav 4.1.2-1" but NOT for R761, where a process
number like 12.1 stands in the bodies. Under a whole-base citation list that
mark was "cited" before any model call.
tests: test_inert_identifier_loadbearing.py (7 arms; known positive is P16's
OWN artefact replayed against the base that run was given, known negative is
26 of 26 fasit references still grounding), test_docs_dir_optional_
loadbearing.py (5 arms). test_stress_judge_loadbearing.py's snippet arm split
into narrowed/whole-base -- the pair is the discriminator, same snippet, same
mark, only the scope differs. The grounding tests migrate from str to
Grounding.of (the honest reading of a caller that declared no boundaries).
Verification: uv run pytest -q 1698 passed / 5 skipped (1685 after part A,
strict superset, 0 removed). ruff check + format clean, mypy clean (38
files). Golden demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
339 lines
16 KiB
Python
339 lines
16 KiB
Python
"""P8 — the run SAYS what its delivered input can ground, before it spends an attempt on a proposal
|
|
that cannot be grounded.
|
|
|
|
P7 (økt 109) is right and landed: an identifier a proposal builds on must appear VERBATIM in the
|
|
input, or the verdict falls. Re-measuring it exposed a CONSEQUENCE no row stated. Over the three
|
|
free recordings, with the gate live, **29 of 29 cost codes in 13 of 13 delivered proposals are
|
|
ungrounded** — every one of them invented. That is not a fault in the gate. It is that the PROMPT
|
|
asks for something the delivered input cannot supply: ``_build_messages`` says each entry "must
|
|
restate a cost line as the project's price schedule already carries it", and MEASURED, K2's
|
|
delivered input carries **two** code-shaped tokens in the prompt (``SHA-01``/``SHA-10``), both of
|
|
them document numbers off a page footer, and ``derive_cost_baseline`` refuses the base outright —
|
|
there is no cost line in it to restate. A gate that always refuses is as useless as one that never
|
|
does, so the run must be able to say which of the two situations it is in.
|
|
|
|
**This is a REPORT, never a gate.** It does not block: a run with a null offer still runs, because
|
|
a blocking requirement is exactly ``--require-cost-baseline``, which F4 settled as opt-in and this
|
|
order freed. It also does not touch ``_ground_against_input``, which ten mutations hold.
|
|
|
|
Arms, each with a named detach point:
|
|
|
|
* **(a) the null offer is stated.** A delivered text with no identifier of any measured form and no
|
|
baseline reports ``identifiers=0, cost_lines=0``. This is the K2 CONTROL — without it a green
|
|
known positive proves nothing.
|
|
* **(b) the positive offer is stated.** N100's delivered input carries the BINDING known positive
|
|
``Krav 3.3.1—13`` (P7 premiss (vii): it stands in 6 of 6 prompts, EM-DASH U+2014), and the offer
|
|
is positive there. Paired with a control that the search CAN find the string at all.
|
|
* **(c) the offer is measured on the SAME text the gate will see.** ``grounding_offer`` composes
|
|
through ``generate._grounding_text`` — the one composer P7 uses — so a report built from a second
|
|
rendering, free to disagree with the one that was sent, is impossible by construction.
|
|
* **(d) the offer reaches the outcome the operator reads**, on BOTH surfaces: ``RunResult`` after a
|
|
full run and ``DryRunReport`` before the first model call. The dry-run arm is what proves the
|
|
measurement happens BEFORE the three attempts are spent — its client factory RAISES.
|
|
* **(e) ONE renderer**, silent when the run CAN anchor a cost line (omission, never an empty row —
|
|
``cost_baseline_notice``'s rule), speaking when it cannot.
|
|
* **(f) a run that exhausts its attempts on ungrounded identifiers leaves the input-side diagnosis
|
|
in the outcome**, not only "the model missed".
|
|
|
|
A NEW file rather than an extension of ``test_identifier_grounding_loadbearing.py``, stated as the
|
|
order asks: that file's subject is the GATE and it is held by ten mutations. A report mutated here
|
|
would go red in a file whose docstring promises a falsifier, and which seam a red test names is the
|
|
only thing a mutation table is for.
|
|
|
|
The fixtures are P7's own tracked generation prompts. Tracked test data on purpose (``scratchpad/``
|
|
is absent from ``git archive HEAD``) — and REUSED rather than extended with the delivered cuts,
|
|
because those cuts are 10 kB and 96 kB of a real Norwegian tender and a public road standard, and
|
|
committing corpus text into a repo published on ``open/`` is a publication decision that belongs to
|
|
the operator, not to this order. The 8-distinct / 435-distinct figures for the delivered cut and the
|
|
full grounding are MEASURED and reported in ``docs/2026-09-09-p8-forankringstilbudet.md``; what a
|
|
test asserts here is the property, on text this repo already ships.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
from pathlib import Path
|
|
|
|
from portfolio_optimiser.generate import GroundingOffer, _grounding_text, grounding_offer
|
|
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
|
|
from portfolio_optimiser.reference_domain import CostItem, Project
|
|
from portfolio_optimiser import run as run_mod
|
|
from portfolio_optimiser.run import grounding_offer_notice, run_project
|
|
from portfolio_optimiser.simulation import scripted_factory
|
|
from portfolio_optimiser.validator import Grounding, Rejection
|
|
|
|
FIXTURES = Path(__file__).parent / "fixtures" / "p7-grounding"
|
|
SHARED = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
|
|
|
#: P7 premiss (vii), the BINDING known positive: the model quoted this requirement number verbatim
|
|
#: in 6 of 6 replies and it stands in 6 of 6 prompts. EM-DASH (U+2014); the hyphen variant is 0/6.
|
|
KNOWN_POSITIVE = "Krav 3.3.1—13"
|
|
|
|
#: The base with a ``cost-baseline.json`` (one line) and the one without — the pair that makes the
|
|
#: renderer's omission arm reachable rather than asserted.
|
|
ANCHORED = SHARED / "veglys-fv-soer"
|
|
UNANCHORED = SHARED / "bygg-energi-mikro"
|
|
_UNANCHORED_PID = "BYGG-KONTOR-NORD"
|
|
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (sim)"}
|
|
_VALID_REPLY = (
|
|
'{"measure":"LED-retrofit","affected_items":'
|
|
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],'
|
|
'"claimed_saving_nok":30000}'
|
|
)
|
|
#: A code that appears NOWHERE in the base — the ungrounded case, which P7's stage 0b refuses on
|
|
#: every attempt, so the run exhausts its attempts and returns the last ``Rejection``.
|
|
_UNGROUNDED_REPLY = (
|
|
'{"measure":"invented","affected_items":'
|
|
'[{"code":"M-04-01","quantity":300000,"unit_cost":1.0}],'
|
|
'"claimed_saving_nok":30000}'
|
|
)
|
|
_CHECKER_REPLY = "Reasoning holds.\nVERDICT: APPROVE"
|
|
|
|
|
|
def _prompt(name: str) -> str:
|
|
return (FIXTURES / name).read_text(encoding="utf-8")
|
|
|
|
|
|
def _project(*codes: str) -> Project:
|
|
"""A project shaped like the bundle path's: ``_project_from_bundle`` builds ``cost_items=()``
|
|
(MEASURED, P7's own row), so the default carries no codes of its own and the offer measures the
|
|
delivered text alone."""
|
|
return Project(
|
|
id="X",
|
|
name="X",
|
|
description="",
|
|
currency="NOK",
|
|
cost_items=tuple(
|
|
CostItem(code=c, description=c, quantity=1.0, unit="stk", unit_cost=1.0) for c in codes
|
|
),
|
|
docs_dir="",
|
|
)
|
|
|
|
|
|
def _baseline(*codes: str) -> CostBaseline:
|
|
return CostBaseline(
|
|
project_id="X",
|
|
items={c: CostBaselineLine(quantity=1.0, unit_cost=1.0) for c in codes},
|
|
)
|
|
|
|
|
|
# ------------------------------------------------------------------ (a) the null offer is stated
|
|
|
|
|
|
def test_a_delivered_input_with_no_cost_line_reports_zero() -> None:
|
|
"""LOAD-BEARING (a) — the K2 CONTROL. Both K2 generation prompts carry no identifier of any
|
|
measured form and the run is un-anchored, so the offer is null in BOTH numbers. Without this
|
|
control a green (b) would be satisfied by a report that always counts positive."""
|
|
for name in ("p6-k2-generation-prompt.txt", "s7c-k2-generation-prompt.txt"):
|
|
offer = grounding_offer(_project(), None, Grounding.of(_prompt(name)))
|
|
assert offer.identifiers == 0, (name, offer)
|
|
assert offer.cost_lines == 0, (name, offer)
|
|
assert offer.chars > 0, "the measurement must have had text to measure"
|
|
|
|
|
|
# -------------------------------------------------------------- (b) the positive offer is stated
|
|
|
|
|
|
def test_the_known_positive_input_reports_a_positive_offer() -> None:
|
|
"""LOAD-BEARING (b) — the BINDING known positive. Paired with the control that the string is
|
|
actually there: a report that finds nothing because it searched for nothing would otherwise
|
|
pass (a) and look measured."""
|
|
text = _prompt("p4-n100-generation-prompt.txt")
|
|
assert KNOWN_POSITIVE in text, "the fixture no longer carries the known positive"
|
|
offer = grounding_offer(_project(), None, Grounding.of(text))
|
|
assert offer.identifiers > 0, offer
|
|
assert offer.cost_lines == 0, "a road standard carries no cost lines (F4's own finding)"
|
|
|
|
|
|
def test_a_bare_number_is_not_counted_as_an_offer() -> None:
|
|
"""LOAD-BEARING (b), the other half. K2 carries 46 394 bare-number occurrences over 2 117
|
|
distinct values (P7 § 2), so counting them would make every report positive and the whole
|
|
measurement inert — the repo's cardinal class, a gate that can only come out green."""
|
|
offer = grounding_offer(_project(), None, Grounding.of("1234 5678 90 42.5 1000000"))
|
|
assert offer.identifiers == 0, offer
|
|
|
|
|
|
# ----------------------------------------------- (c) measured on the SAME text the gate will see
|
|
|
|
|
|
def test_the_offer_is_measured_on_the_text_the_gate_will_see() -> None:
|
|
"""LOAD-BEARING (c) — requirement (2). The offer composes through ``_grounding_text``, the one
|
|
composer P7's gate uses, so the report and the gate cannot describe different texts. Both of
|
|
the composer's OTHER two sources are exercised: a project cost line and a baseline code each
|
|
raise the count, which a report built from ``delivered`` alone cannot do."""
|
|
delivered = Grounding.of("nothing citable here")
|
|
project, baseline = _project("PRJ-77"), _baseline("BAS-88")
|
|
|
|
offer = grounding_offer(project, baseline, delivered)
|
|
|
|
assert offer.chars == len(_grounding_text(project, baseline, delivered).text)
|
|
assert offer.identifiers == 2, offer
|
|
assert offer.cost_lines == 1, offer
|
|
|
|
|
|
# ---------------------------------------------------------- (e) ONE renderer, omission when able
|
|
|
|
|
|
def test_the_notice_speaks_on_a_null_offer_and_is_silent_when_anchored() -> None:
|
|
"""LOAD-BEARING (e). Silent when the run CAN anchor a cost line (omission, never an empty row);
|
|
speaking when it cannot, and carrying BOTH numbers — "50 identifiers, 0 cost lines" is the
|
|
diagnosis, and neither number alone says it."""
|
|
assert grounding_offer_notice(GroundingOffer(chars=9, identifiers=3, cost_lines=1)) is None
|
|
assert grounding_offer_notice(None) is None
|
|
|
|
line = grounding_offer_notice(GroundingOffer(chars=9, identifiers=3, cost_lines=0))
|
|
assert line is not None
|
|
assert "3" in line and "0" in line
|
|
|
|
|
|
# --------------------------------------- (d) the offer reaches BOTH surfaces the operator reads
|
|
|
|
|
|
async def test_the_dry_run_carries_the_offer_before_any_model_call() -> None:
|
|
"""LOAD-BEARING (d), and the proof of "BEFORE it spends its three attempts".
|
|
|
|
Asserted on CALLS, never on client CONSTRUCTION: ``fresh_workflow`` builds the proposer and
|
|
checker clients EAGERLY, above the dry-run cut (``workflow.py:64``), so a factory that raised
|
|
would be red against a working implementation. An empty sink is the measurement — session 57's
|
|
rule that a refusal after the spend is indistinguishable from one before it at the exit code.
|
|
"""
|
|
sink: list[str] = []
|
|
|
|
report = await run_project(
|
|
_UNANCHORED_PID,
|
|
"local",
|
|
docs_dir=str(UNANCHORED),
|
|
bundle_dir=str(UNANCHORED),
|
|
client_factory=scripted_factory({"proposer": "x", "checker": "x"}, sink),
|
|
live_dry_run=True,
|
|
)
|
|
|
|
assert sink == [], "the offer was measured only after a model call was made"
|
|
|
|
offer = getattr(report, "grounding_offer", None)
|
|
assert offer is not None, "the dry run says nothing about what its input can ground"
|
|
assert offer.cost_lines == 0, offer
|
|
assert offer.chars > 0, offer
|
|
|
|
|
|
async def test_the_full_run_carries_the_offer() -> None:
|
|
"""LOAD-BEARING (d), the other surface — requirement (3). A measurement that never leaves
|
|
``run_project`` is a log line, not a report."""
|
|
result = await run_project(
|
|
_UNANCHORED_PID,
|
|
"local",
|
|
docs_dir=str(UNANCHORED),
|
|
bundle_dir=str(UNANCHORED),
|
|
verdict_input=_VERDICT_INPUT,
|
|
client_factory=scripted_factory({"proposer": _VALID_REPLY, "checker": _CHECKER_REPLY}, []),
|
|
)
|
|
|
|
offer = getattr(result, "grounding_offer", None)
|
|
assert offer is not None, "the run says nothing about what its input can ground"
|
|
assert offer.cost_lines == 0, offer
|
|
|
|
|
|
async def test_an_anchored_run_reports_its_cost_lines() -> None:
|
|
"""LOAD-BEARING (d), the CONTROL that ``cost_lines`` is not a constant zero: the anchored base
|
|
ships exactly one line, and the run reports it."""
|
|
report = await run_project(
|
|
"VEGLYS-FV-SOER",
|
|
"local",
|
|
docs_dir=str(ANCHORED),
|
|
bundle_dir=str(ANCHORED),
|
|
client_factory=scripted_factory({"proposer": "x", "checker": "x"}, []),
|
|
live_dry_run=True,
|
|
)
|
|
|
|
offer = getattr(report, "grounding_offer", None)
|
|
assert offer is not None and offer.cost_lines == 1, offer
|
|
|
|
|
|
# ------------------------------------- (f) an exhausted run leaves the INPUT-side diagnosis behind
|
|
|
|
|
|
async def test_a_run_exhausted_on_ungrounded_identifiers_says_the_input_was_the_problem() -> None:
|
|
"""LOAD-BEARING (f). P7's gate refuses every attempt, the loop returns its last ``Rejection``
|
|
(``generate.py``: ``last_ruling`` after the attempt loop — it does not crash), and the outcome
|
|
the operator reads carries the input-side fact as well: this input could anchor NOTHING, so no
|
|
attempt could ever have succeeded. Without it the record says only that the model missed."""
|
|
result = await run_project(
|
|
_UNANCHORED_PID,
|
|
"local",
|
|
docs_dir=str(UNANCHORED),
|
|
bundle_dir=str(UNANCHORED),
|
|
verdict_input=_VERDICT_INPUT,
|
|
client_factory=scripted_factory(
|
|
{"proposer": _UNGROUNDED_REPLY, "checker": _CHECKER_REPLY}, []
|
|
),
|
|
)
|
|
|
|
assert isinstance(result.outcome, Rejection)
|
|
assert "M-04-01" in result.outcome.reason
|
|
offer = getattr(result, "grounding_offer", None)
|
|
assert offer is not None and offer.cost_lines == 0, offer
|
|
assert grounding_offer_notice(offer) is not None
|
|
|
|
|
|
# ------------------------------------------------ (d) the CLI prints it, on BOTH free and paid
|
|
|
|
|
|
def test_the_cli_dry_run_prints_the_offer(capsys) -> None:
|
|
"""LOAD-BEARING (d), stdout. A renderer that returns the line while no caller prints it is a
|
|
measurement the operator never sees — its own seam, so its own arm (the ``cost_baseline_notice``
|
|
precedent, whose CLI print carries a mutation of its own)."""
|
|
rc = run_mod.main(
|
|
[
|
|
_UNANCHORED_PID,
|
|
"--docs-dir",
|
|
str(UNANCHORED),
|
|
"--bundle-dir",
|
|
str(UNANCHORED),
|
|
"--live-dry-run",
|
|
]
|
|
)
|
|
assert rc == 0
|
|
assert "Grounding offer" in capsys.readouterr().out
|
|
|
|
|
|
def test_the_cli_says_nothing_when_the_run_can_anchor(capsys) -> None:
|
|
"""CONTROL for the omission arm on the SAME surface: the anchored base prints no offer line at
|
|
all. Without it, a notice that always fired would pass the arm above."""
|
|
rc = run_mod.main(
|
|
[
|
|
"VEGLYS-FV-SOER",
|
|
"--docs-dir",
|
|
str(ANCHORED),
|
|
"--bundle-dir",
|
|
str(ANCHORED),
|
|
"--live-dry-run",
|
|
]
|
|
)
|
|
assert rc == 0
|
|
assert "Grounding offer" not in capsys.readouterr().out
|
|
|
|
|
|
def test_the_cli_full_run_prints_the_offer(tmp_path: Path, capsys) -> None:
|
|
"""LOAD-BEARING (d), the paid surface's stdout — so the line is a property of a RUN and not of
|
|
the dry-run branch alone."""
|
|
replies = tmp_path / "replies.json"
|
|
replies.write_text(
|
|
json.dumps({"proposer": _VALID_REPLY, "checker": _CHECKER_REPLY}), encoding="utf-8"
|
|
)
|
|
rc = run_mod.main(
|
|
[
|
|
_UNANCHORED_PID,
|
|
"--docs-dir",
|
|
str(UNANCHORED),
|
|
"--bundle-dir",
|
|
str(UNANCHORED),
|
|
"--scripted-replies",
|
|
str(replies),
|
|
"--decision",
|
|
"approved",
|
|
"--rationale",
|
|
"expert reviewed (test)",
|
|
]
|
|
)
|
|
assert rc == 0
|
|
assert "Grounding offer" in capsys.readouterr().out
|