feat(p22): stage 0's unknown-code refusal NAMES the codes the project buys

P21 was the first ANCHORED stress round and it bought something real -- must_refuse 5/5, all
five caught on stage0-baseline. It cost something measured just as clearly: 0 of 20 approaches
validated, and 26 of 26 rejections (20 approach rows plus 6 own-proposals -- a LARGER population
than the 20) read `unknown cost code '<invention>': not in project P's cost baseline (5 known
codes)`. The model invented signalregulering_konstruksjon, VENTIL_IMP, RIGG01, baerelag_asfalt
and 22 more, and it could not have done otherwise: the price schedule reaches the VALIDATOR and
never the proposer, and the refusal stated the COUNT of known codes, not one name. Step 5 feeds
that sentence verbatim into the next attempt -- and "you guessed wrong, there are five right
answers" carries nothing to correct towards.

The contrast already lived in the same stage: the MAGNITUDE half NAMES the baseline value, and
that is the half that let the loop converge in session 94. This gives that property to the other
half, in one place, and Step 5 carries it forward for free.

The window is a FIXED COUNT of whole codes (20), never a share, and it counts codes rather than
characters because a character cut can sever a code mid-name and hand the proposer an identifier
that exists nowhere. The cut is announced; a schedule that fits is not marked truncated; the
order is the schedule's own.

Measured with denominators: every cost baseline in this repo or its measured corpora is at most
six codes, and the largest real delivered price schedule measured is K2's prissammenstilling at
14 priced rows. Nothing measured reaches the window; it exists for the R761-style mengdebeskrivelse.

Load-bearing MEASURED (tests/test_named_known_codes_loadbearing.py, 10 arms), eight mutations all
red against the WHOLE suite + green control 1873/5 (from 1863/5, superset, 0 removed) and golden
demo-transcript.stdout BYTE-UNCHANGED (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f):
A1 revert to the bare count (7 red) - A2 a share instead of a fixed window (6) - A3 sorted (1) -
A4 silent cut (2) - A5 character slice (3) - A6 no bound at all (3) - A7 break the rejection_stage
marker (2, one an OLDER independent witness) - A8 grow the magnitude half with a code list
(2, one an OLDER independent witness).

Honesty limits, stated: no paid run yet confirms this changes the outcome live (that is DEL D);
the truncation branch is exercised only synthetically because nothing measured reaches the window;
and the schedule still does not reach the prompt, so the first attempt guesses as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-15 22:52:37 +02:00
commit 0a81de2d76
4 changed files with 302 additions and 3 deletions

View file

@ -2971,6 +2971,46 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
lesninger teller mot gulvet — regelen måler at kjøringen SÅ seg om, ikke at den forsto; ingen
LEVENDE modell har møtt noen av de to nektene (structured-output-grensens klasse); og C2 er
hjelpetekst, ikke en gate — den kan ikke gjøre en gjettet sti riktig, bare billigere å rette.
- **Stadium 0s ukjent-kode-nekt NAVNGIR kodene prosjektet faktisk kjoeper, i et FAST vindu
(P22 DEL A, 15.09):** P21 var den foerste FORANKREDE runden, og den kostet noe maalt like tydelig
som den kjoepte: **0 av 20** tilnaerminger validerte (runde 4: 4 av 20), og **26 av 26**
avvisninger — 20 approach-rader PLUSS 6 egne forslag, altsaa en STOERRE populasjon enn de 20 —
leste `unknown cost code '<paafunn>': not in project P's cost baseline (5 known codes)`. Modellen
fant paa `signalregulering_konstruksjon`, `VENTIL_IMP`, `RIGG01`, `baerelag_asfalt` og 22 til, og
den KUNNE ikke gjort annet: prisskjemaet naar VALIDATOREN og aldri proposeren, og nekten oppga
ANTALLET kjente koder, ikke ett eneste navn. Steg 5 mater setningen ORDRETT inn i neste forsoeks
prompt, saa «du gjettet feil, det finnes fem riktige» baerer ingenting aa korrigere mot.
**Kontrasten bodde allerede i samme stadium:** MAGNITUDE-halvdelen NAVNGIR baseline-verdien, og
DET er halvdelen som lot loekka konvergere i oekt 94. `_known_codes_clause` er den ene
renderingen, og magnitude-halvdelen faar den IKKE — der er setningen alt korrigerbar, og en
kodeliste paa begge ville gjort de to halvdelene uskillbare for en leser (M A8 → 2 roede, hvorav
ett ELDRE uavhengig vitne, `test_stage0_all_violations::test_a_single_violation_is_byte_identical`).
**Vinduet er et FAST ANTALL, aldri en andel** (`_CATALOGUE_EXCERPT_CHARS`-regelen; M A2 → 6 roede,
M A6 → 3), **og det teller KODER, ikke tegn**, av en maalt grunn: et tegn-kutt kan kappe en kode
midt i navnet og gi proposeren en identifikator som finnes ingen steder — `_index_excerpt`-regelen
invertert (M A5 → 3 roede paa hel-kode-armen). **Kuttet ANNONSERES** (`first N:`) mens nevneren
staar i begge grener, og et skjema som PASSER merkes ikke avkortet (`index_truncated`s regel;
M A4 → 2 roede). **Rekkefoelgen er skjemaets egen** — en sortering ville oppfunnet en rangering
prosjektet aldri uttalte (M A3 → 1 roed, den armen ALENE). Vinduet er 20, valgt ved MAALING med
nevner: hver kostbaseline i repoet eller dets maalte korpora er hoeyst SEKS koder (kontekstsettene
5/5/5/5/6, de to `shared/examples` 1 hver, MAJOR-4s avledning av det syntetiske K2-prisskjemaet 3),
og det stoerste EKTE leverte prisskjemaet maalt er K2s `prissammenstilling-sheet-1.md` med 14
prisede rader av 118 linjer — ingenting maalt naar vinduet; det finnes for den umaalte
R761-formede mengdebeskrivelsen, der korpuset erklaerer 2 727 `prosessnr`.
**`rejection_stage` er koblingen P21s «a4 5/5 paa stage0-baseline» hviler paa** og noekler paa
delstrengen `cost baseline (` — hadde den nye klausulen flyttet den, ville hver stadium-0-nekt
blitt omdoept til `other` i stillhet (M A7 → 2 roede, hvorav ett ELDRE uavhengig vitne).
**IKKE GJORT, og det er en beslutning:** skjemaet rendres ALDRI inn i proposer-prompten — en ny
prompt-flate med egen kostnad per kjoering som dessuten ikke fjerner behovet for at nekten er
korrigerbar. Load-bearing MAALT (`tests/test_named_known_codes_loadbearing.py`, 10 armer), **aatte
mutasjoner alle roede mot HELE suiten** + groenn kontroll **1873/5** (fra 1863/5, supersett,
0 fjernet) og golden `demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): A1 tilbake til det bare antallet (7) · A2 andel (6) ·
A3 sortert (1) · A4 stille kutt (2) · A5 tegn-slice (3) · A6 ingen binding (3) · A7 bryt
stage-markoeren (2) · A8 la magnitude-halvdelen vokse en kodeliste (2). **AErlighets-grenser,
uttalt:** ingen betalt kjoering bekrefter ennaa at nekten endrer utfallet levende (DEL D er
maalingen); vinduet er aldri naadd av noe maalt, saa avkortnings-grenen er oevd kun syntetisk; og
skjemaet naar fortsatt ikke prompten, saa foerste forsoek gjetter som foer.
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -757,8 +757,12 @@ when the seam is detached, so the loop cannot silently degrade into theater.
belongs to the project, and this is where you hand it over: FILE is a `cost-baseline.json` — the
same `{"project_id": …, "items": {"<code>": {"quantity": …, "unit_cost": …}}}` shape a bundle may
ship — and it is used INSTEAD of one inside the base. With it, a proposal naming a cost line the
project does not buy is refused as a fabricated line, and one naming a real line with invented
magnitudes is refused with the real ones named, so the next attempt can correct.
project does not buy is refused as a fabricated line — and the refusal NAMES the codes the
project does have, so the next attempt has something to correct towards. A real line with
invented magnitudes is likewise refused with the real figures named. (Measured on a live tender
round before the codes were named: 26 rejections, every one of them an invented code, and not
one run recovered — a refusal that says only "there are five right answers" carries nothing to
aim at. A long schedule is cut to a fixed number of codes and says so.)
In `--across-bundle` mode the same schedule anchors every base: one project, one price schedule.
Mutually exclusive with `--derive-cost-baseline` (two sources for one baseline), and it satisfies

View file

@ -154,6 +154,50 @@ def baseline_from_project(project: Project) -> CostBaseline:
)
#: P22 DEL A - how many of the project's OWN cost codes the stage-0 refusal names.
#:
#: A FIXED WINDOW, never a share of the schedule: a share scales with the corpus again, only with
#: a smaller constant, which is the exact regression the catalogue excerpt's
#: ``_CATALOGUE_EXCERPT_CHARS`` exists for. It counts CODES rather than characters for a measured
#: reason - a character cut can sever a code mid-name and hand the proposer an identifier that
#: exists nowhere, the ``_index_excerpt`` rule inverted (a path that never existed is worse than
#: no path). A count window can only ever emit whole codes.
#:
#: MEASURED with denominators (okt 126): every cost baseline anywhere in this repo or its measured
#: corpora is at most SIX codes - the five context sets carry 5/5/5/5/6, the two shipped
#: ``shared/examples`` baselines carry 1 each, and MAJOR-4's derivation of the synthetic K2 price
#: schedule yields 3 - while the largest REAL delivered price schedule measured is K2's
#: ``prissammenstilling-sheet-1.md``, 14 priced rows of 118 lines. Nothing measured reaches this
#: window. It exists for the unmeasured R761-style mengdebeskrivelse, where a Norwegian road
#: contract is priced BY prosesskode and the corpus declares 2 727 of them.
_KNOWN_CODE_WINDOW: Final = 20
def _known_codes_clause(baseline: CostBaseline) -> str:
"""The parenthetical of stage 0's unknown-code refusal: how many codes the project buys, and
which ones (P22 DEL A).
MEASURED (P21 round 5): all 26 rejections across six paid runs named an invented code and
answered with a COUNT - ``(5 known codes)`` - so the sentence Step 5 feeds verbatim into the
next attempt carried nothing to correct towards, and 0 of 20 approaches validated. The
magnitude half of this same stage NAMES the baseline value, and that is the half that let the
loop converge in okt 94. This is that half's property, given to the other one.
Bound by ``_KNOWN_CODE_WINDOW`` - a fixed count of WHOLE codes, see the constant. The cut is
ANNOUNCED (``first N:``) rather than left for the reader to subtract from the denominator,
which stays in either branch; a schedule that FITS is not marked truncated and gets its whole
list. Omission, never a lie in either direction (``index_truncated``'s rule).
The order is the SCHEDULE's own. A sort would invent a ranking the project never stated, and
the window would then be an alphabetical accident rather than the head of the document the
operator wrote."""
codes = list(baseline.items)
listed = ", ".join(repr(c) for c in codes[:_KNOWN_CODE_WINDOW])
if len(codes) <= _KNOWN_CODE_WINDOW:
return f"{len(codes)} known codes: {listed}"
return f"{len(codes)} known codes, first {_KNOWN_CODE_WINDOW}: {listed}"
def _reconcile_against_baseline(
proposal: SavingsProposal, baseline: CostBaseline, tolerance: float
) -> Rejection | None:
@ -194,7 +238,7 @@ def _reconcile_against_baseline(
if line is None:
violations.append(
f"unknown cost code {item.code!r}: not in project {baseline.project_id}'s "
f"cost baseline ({len(baseline.items)} known codes)"
f"cost baseline ({_known_codes_clause(baseline)})"
)
continue
for field, claimed, actual in (

View file

@ -0,0 +1,211 @@
"""P22 DEL A — stage 0's unknown-code refusal NAMES the codes the project actually buys.
MEASURED (P21 stress round 5, re-measured at the head of økt 126): all 26 rejections across the
six paid runs read ``unknown cost code '<invention>': not in project P's cost baseline (5 known
codes)``, and **0 of 20** approaches validated (round 4: 4 of 20). The model invented
``signalregulering_konstruksjon``, ``VENTIL_IMP``, ``RIGG01``, ``bærelag_asfalt`` and 22 more. It
could not have done otherwise: the price schedule reaches the VALIDATOR and never the proposer,
and the refusal states the COUNT of known codes, not one code's name. Step 5 feeds that sentence
verbatim into the next attempt's prompt — and "you guessed wrong, there are five right answers"
carries nothing to correct towards.
The contrast lives in this repo already. The MAGNITUDE half of the same stage NAMES the baseline
value, and that is the half that let the loop converge in økt 94: the proposer found each correct
figure because the refusal told it what the figure was.
**The window is a fixed COUNT, never a share** (the catalogue-excerpt rule, `_CATALOGUE_EXCERPT_
CHARS`), and it counts CODES rather than characters for a measured reason: a character cut can
sever a code mid-name and hand the proposer an identifier that exists nowhere the
``_index_excerpt`` rule inverted, where a path that never existed is worse than no path. A count
window can only ever emit whole codes.
Measured sizes, with denominators (økt 126): every cost baseline anywhere in this repo or its
measured corpora is at most SIX codes (the five context sets: 5 · 5 · 5 · 5 · 6; the two shipped
``shared/examples`` baselines: 1 each; MAJOR-4's derivation of the synthetic K2 price schedule: 3),
and the largest REAL delivered price schedule measured is K2's ``prissammenstilling-sheet-1.md`` at
14 priced rows of 118 lines. Nothing measured reaches the window; it exists for the unmeasured
R761-style mengdebeskrivelse, where a contract is priced BY prosesskode and the corpus declares
2 727 of them.
"""
from __future__ import annotations
from typing import Final
from portfolio_optimiser.generate import _build_messages
from portfolio_optimiser.ir import (
AffectedItem,
CostBaseline,
CostBaselineLine,
SavingsProposal,
)
from portfolio_optimiser.reference_domain import Project
from portfolio_optimiser.validator import (
_KNOWN_CODE_WINDOW,
Rejection,
ValidatedProposal,
rejection_stage,
validate_proposal,
)
_SMALL: Final = ("RIGG", "ASFALT", "MASSE", "FROST", "SKILT")
def _baseline(codes: tuple[str, ...], project_id: str = "proj") -> CostBaseline:
return CostBaseline(
project_id=project_id,
items={c: CostBaselineLine(code=c, quantity=100.0, unit_cost=1000.0) for c in codes},
)
def _proposal(code: str, *, quantity: float = 100.0, unit_cost: float = 1000.0) -> SavingsProposal:
return SavingsProposal(
project_id="proj",
measure="Reduce scope",
affected_items=[AffectedItem(code=code, quantity=quantity, unit_cost=unit_cost)],
claimed_saving_nok=1000.0,
assumptions={},
)
def _refusal(codes: tuple[str, ...], guess: str = "INVENTED") -> str:
result = validate_proposal(_proposal(guess), baseline=_baseline(codes))
assert isinstance(result, Rejection), "an invented code must never reach validated"
return result.reason
# --- (a) known-POSITIVE: the refusal names them --------------------------------------------------
def test_the_refusal_names_every_code_a_small_schedule_carries() -> None:
"""RED before DEL A: the refusal carried only ``(5 known codes)``. It now names all five, so
the sentence Step 5 feeds forward is CORRECTABLE. The denominator stays a reader must be
able to tell how many exist from the same sentence that lists them."""
reason = _refusal(_SMALL)
for code in _SMALL:
assert code in reason, f"{code!r} must be named: {reason}"
assert "5 known codes" in reason, reason
# --- (b) known-NEGATIVE: the gate can still FELL, and stays silent when it should -----------------
def test_a_code_the_project_really_buys_still_validates() -> None:
"""The control the order demands alongside the positive: a corrected gate can still be untrue
for an independent reason. A proposal on a REAL line validates, so the naming above is caused
by the code being absent not by a stage that now rejects everything."""
result = validate_proposal(_proposal("RIGG"), baseline=_baseline(_SMALL))
assert isinstance(result, ValidatedProposal)
def test_the_magnitude_half_does_not_grow_a_code_list() -> None:
"""A REAL code at an invented magnitude falls on the OTHER half of this stage, which already
names the baseline value. Listing the schedule there would be noise on a sentence that is
already correctable and would make the two halves indistinguishable to a reader."""
result = validate_proposal(_proposal("RIGG", unit_cost=5000.0), baseline=_baseline(_SMALL))
assert isinstance(result, Rejection)
assert "tolerance around the baseline" in result.reason
assert "known codes" not in result.reason, result.reason
# --- (c)/(d) the window is BOUND, and it is a window and not a share ------------------------------
def test_a_large_schedule_is_bound_by_a_fixed_window() -> None:
"""A mengdebeskrivelse priced by prosesskode can carry thousands of lines. The list is bound at
``_KNOWN_CODE_WINDOW`` codes, the denominator still states how many exist, and the cut is
ANNOUNCED rather than left for the reader to subtract (``index_truncated``'s rule)."""
many = tuple(f"P{i:04d}" for i in range(500))
reason = _refusal(many)
named = [c for c in many if c in reason]
assert len(named) == _KNOWN_CODE_WINDOW, f"named {len(named)}, want {_KNOWN_CODE_WINDOW}"
assert "500 known codes" in reason, reason
assert f"first {_KNOWN_CODE_WINDOW}" in reason, reason
def test_the_window_is_a_count_and_not_a_share_of_the_schedule() -> None:
"""A share scales with the schedule again, only with a smaller constant — the exact regression
the catalogue excerpt's fixed window exists for. Two schedules an order of magnitude apart name
the SAME number of codes."""
small_named = sum(
1
for c in tuple(f"P{i:04d}" for i in range(200))
if c in _refusal(tuple(f"P{i:04d}" for i in range(200)))
)
big_named = sum(
1
for c in tuple(f"P{i:04d}" for i in range(2000))
if c in _refusal(tuple(f"P{i:04d}" for i in range(2000)))
)
assert small_named == big_named == _KNOWN_CODE_WINDOW
def test_a_schedule_exactly_at_the_window_is_not_announced_as_cut() -> None:
"""Omission, never a lie in either direction: a schedule that FITS is not marked truncated and
gets its whole list (``index_truncated``'s positive half)."""
exact = tuple(f"P{i:04d}" for i in range(_KNOWN_CODE_WINDOW))
reason = _refusal(exact)
assert all(c in reason for c in exact)
assert "first " not in reason, reason
# --- (e)/(f) the codes it hands back are the project's own, whole and in its own order ------------
def test_the_named_codes_are_the_schedules_own_order_never_sorted() -> None:
"""The schedule's order is the PROJECT's order. A sort would invent a ranking the project never
stated, and the first window would then be an alphabetical accident rather than the head of the
document the operator wrote."""
unsorted = ("ZZ-last", "AA-first", "MM-middle") + tuple(f"P{i:03d}" for i in range(50))
reason = _refusal(unsorted)
named = [c for c in unsorted if c in reason]
assert named[:3] == ["ZZ-last", "AA-first", "MM-middle"], named[:3]
def test_every_named_token_resolves_to_a_real_baseline_line() -> None:
"""The ``_index_excerpt`` property, measured rather than assumed: each code the refusal hands
back is a WHOLE key of the baseline. A character-sliced window could emit ``'P004`` and send
the proposer after an identifier that exists nowhere."""
many = tuple(f"code-{i:03d}-long-enough-to-slice" for i in range(100))
reason = _refusal(many)
listed = reason.split("first ")[1].split(": ", 1)[1].rstrip(")")
tokens = [t.strip().strip("'") for t in listed.split(", ")]
assert len(tokens) == _KNOWN_CODE_WINDOW, tokens
for token in tokens:
assert token in many, f"{token!r} is not a baseline code"
# --- (g) the coupling P21's "a4 5/5 on stage0-baseline" rests on ----------------------------------
def test_the_longer_sentence_is_still_labelled_stage0() -> None:
"""``rejection_stage`` keys on ``"cost baseline ("``. Had the new clause moved that substring,
every stage-0 refusal would have been relabelled ``"other"`` in silence, and P21's headline
measurement (``must_refuse`` 5/5, all on ``stage0-baseline``) would have read as a regression
caused by the fix meant to strengthen it."""
assert rejection_stage(_refusal(_SMALL)) == "stage0-baseline"
# --- (h) Step 5 carries it to the place it has to reach ------------------------------------------
def test_the_named_codes_reach_the_next_attempts_prompt() -> None:
"""The whole point: the refusal is only correctable if it reaches the proposer. Step 5 feeds
the reason VERBATIM into the next attempt, so the codes ride for free this arm measures that
they arrive, rather than trusting that they do."""
reason = _refusal(_SMALL)
messages = _build_messages(
Project(
id="proj",
name="Proj",
description="",
currency="NOK",
cost_items=(),
docs_dir="",
),
"context",
Rejection(proposal=_proposal("INVENTED"), reason=reason),
)
prompt = "\n".join(m.text for m in messages)
for code in _SMALL:
assert code in prompt, f"{code!r} never reached the prompt"