feat(p20): the requirement that is RIGHT, and a clause number that is not a price

Three seams, one commit: A, B and C touch the same four modules (run.py carries
the debate task, the grounding composition and the announcement; okf.py carries
one reference-number vocabulary read by both A and B), so splitting them into
three commits would have meant hunk-level staging of entangled files. Stated
rather than silently restructured.

A — the declaration answers with the DOCUMENT's own words. Measured: 13
declarations over round 3 and P17b, not one naming a fasit concept, while the
tool answered {"declared": true, ...} by echoing the caller's own arguments. It
now returns the document's title and req_number, read off Bundle.context_files
(so the type: verdict layer can never be named back), plus the sentence saying
what the declaration binds. A path the base carries as no concept answers with
empty strings rather than refusing. The commission's success_criteria now reach
the DEBATE task through mandate.criteria_block, the one renderer, empty when
there are none — which is what keeps every un-commissioned prompt, and the
golden, byte-identical.

B — a clause number is not a price. THE ORDER'S OWN RULE WAS FELLED BY
MEASUREMENT: it asks to refuse a code that IS declared req_number/prosessnr,
and neither of its two known positives is. n500 declares seksjon 10.4.1..10.4.4
but never the bare 10.4; r761 declares 2727 prosessnr and 2753 seksjon, none of
them 1.10.4, which occurs once, as prose ("iht. vegnormal N200 kap. 1.10.4").
The COMPLEMENT fires on both and closes the hole _ground_against_input already
admits in writing -- "it fails OPEN on a coincidental match". Unanchored run +
requirement-shaped code + the base declares a vocabulary + the code is not in
it -> refused, naming the denominator. All five of kontrakt-sorasen's real
process codes ARE declared and pass, which is what keeps the one context set
built on real codes measurable. Replayed over all 24 codes of round 3 + P17b:
exactly the two known positives flip validated -> rejected, 22 unchanged.

C — a parse failure no longer burns the round ledger blind. _fetch_parsed takes
a BUILDER instead of a finished message list, so the retry carries the parse
reason; measured, kontrakt-sorasen-04 spent 11 of 12 rounds re-asking the same
question. And announced_subject names the routed bases instead of saying "the
portfolio" for a two-base commission.

Suite 1807/5 (from 1781, +26, 0 removed), golden demo-transcript.stdout
BYTE-UNCHANGED (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f),
ruff and mypy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-15 06:02:46 +02:00
commit c8f0c8f7c4
12 changed files with 1098 additions and 37 deletions

View file

@ -2769,6 +2769,78 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
flaten er BEVISST urørt (feltet er i ingen av hostings tre sett); og `--proposal-review` trås flaten er BEVISST urørt (feltet er i ingen av hostings tre sett); og `--proposal-review` trås
gjennom til ÉN terminal delt av alle basene (motorens egen begrunnelse — dispatchen er gjennom til ÉN terminal delt av alle basene (motorens egen begrunnelse — dispatchen er
sekvensiell). sekvensiell).
- **Et KRAVNUMMER er ikke en PRIS, og basens EGEN nummer-ordliste er dét som sier det — ordrens
regel ble FELT av målingen før noe ble bygget på den (P20 DEL B, 15.09):** tre stressrunder og
ett multi-base-pass bar `validated` forslag hvis kostkode var et kapittelnummer i en vegnormal.
To overlever i utboksene og er kjent-positivene: **`10.4`** (tunnel-hauglia runde 3, n500) og
**`1.10.4`** (lindaas P17b, r761). Begge er GRUNNET i P7s forstand og ingen er INERT i P18/B1s
(`10.4` i 12 av 274 dokumenter, `1.10.4` i **1 av 2 756**); stadium 0 kjørte aldri, fordi ingen
vegnormal-base bærer en kostbaseline. **Ordrens B1 sier: form 2/3 OG «står som
`req_number`/`prosessnr` i toppnivå-frontmatter» → nekt.** MÅLT 15.09: n500 erklærer
`seksjon: 10.4.1``10.4.4` og `req_number: Krav 10.4.3—2`, men **aldri den bare `10.4`** — den er
et seksjons-PREFIKS; og r761 erklærer 2 727 `prosessnr` + 2 753 `seksjon`, hvorav **ingen** er
`1.10.4`, som står ÉN gang, som prosa: «iht. vegnormal N200 Vegbygging kap. 1.10.4». **Den
ordrede regelen fyrer altså på INGEN av sine egne kjent-positive.** KOMPLEMENTET fyrer på BEGGE,
og lukker et hull `_ground_against_input`s egen docstring alt innrømmer skriftlig («it fails OPEN
… on a coincidental match»): for ÉN form — et klausulnummer — gir basen oss ordlista som skiller
en ekte referanse fra et sammentreff. Regelen er derfor: **UFORANKRET + kravformet kode + basen
erklærer en ordliste + koden er IKKE i den → nekt, med nevner.** **Komplementet er også dét som
SPARER det ene kontekstsettet bygget på ekte prosesskoder:** alle fem kodene i
`contexts/kontrakt-sorasen-2027` (`12.1`, `12.12`, `22.1`, `52.11`, `51.1`) ER erklærte
`prosessnr` og passerer — under den ordrede regelen ville hver av dem blitt nektet på en
uforankret r761-kjøring, og settets positive armer blitt umålbare; dét er R761-risikoen ordren
selv navngir, ankommet gjennom døra den ble pekt bort fra. **MÅLT over ALLE 24 koder i runde 3 +
P17b (10 kjøringer):** nøyaktig to er kravformede, de er de to kjent-positive, og offline replay
flipper nøyaktig de to (`validated → rejected`) mens 22 står uendret. **Generalitetsvernet er
`_form_refusal`s mønster:** en input som erklærer INGEN referansenumre kan ikke besvares i en
ordliste den ikke har, så regelen kan ikke fyre der — dét er hva som lar hver pre-P20-fixtur stå
URØRT i stedet for UNNTATT. **`anchored` er et EKSPLISITT flagg, aldri `not anchored_codes`:** en
baseline uten linjer og ingen baseline er ulike fakta (`cost_baseline_anchored`s grunn).
**Ordlista reiser MED teksten** (`Grounding.declared_references`, DEFAULTET —
`skipped_links`-halvdelen: en tom ordliste er et ærlig POSITIVT utsagn), komponert i SAMME vandring
som dokumentene i `run.py` og propagert gjennom `_grounding_text`; to avledninger av én bases
ordliste ville stått fritt til å være uenige (kø-(p)). **`okf.REFERENCE_NUMBER_FIELDS` er ÉN
kilde**, derivert fra `_FILTER_FIELDS` pluss `seksjon` — og BEVISST ikke lagt til i
`_FILTER_FIELDS` selv, fordi hva et `filter`-ord søker i er målt og gatet (P18) og en utvidelse
ville endret `total_matches` for hvert navigatørkall. `code_forms` får en TREDJE verdi,
`requirement`, for en kode basen FAKTISK erklærer (ordrens kjent-negative (c): alle
`must_cite`-referanser klassifiseres slik) — rapport, aldri gaten, som fyrer på komplementet.
Load-bearing MÅLT (`tests/test_requirement_number_gate_loadbearing.py`, 11 armer).
- **Erklæringen svarer med DOKUMENTETS egne ord, og kommisjonens kriterier når leseren som kan
handle på dem (P20 DEL A, 15.09):** P19 DEL A gjorde at en retning MÅ navngi kravet som binder
den, og rungen virker — MÅLT ga runde 3 ni erklæringer over fem betalte kjøringer og P17b fire
over ett pass, og **ikke én av de 13 navnga et fasit-konsept**. Det ingen ba om var at kravet
skulle være det RIKTIGE. To halvdeler manglet: (1) verktøyet svarte `{"declared": true, …}` ved å
ekko kallerens egne tre argumenter, så en modell som hadde erklært feil krav ble fortalt, med de
eneste ordene den fikk, at den hadde lyktes; (2) `Mandate.success_criteria` nådde `announce` og
INGENTING ellers (P19 F2) — skrevet ut for et menneske og holdt tilbake fra den ene leseren som
kunne handle på det. Svaret bærer nå dokumentets EGNE `title`/`req_number`, lest av basen gjennom
`okf.reference_number` (ÉN leser av «hvilket krav er dette», kø-(p)), pluss `binds`-setningen.
**Lest av `Bundle.context_files`, aldri `files`:** en `type: verdict`-fil kan dermed ikke navngis
tilbake ved tittel — det ene laget ingen listing nevner og `read_file` nekter blankt. **En sti
basen ikke bærer som konsept (`index.md`) svarer med TOMME strenger, aldri en nekt:** lese-sporet
har alt godkjent erklæringen, og å gjøre «jeg kan ikke gjengi tittelen din» om til en nekt ville
felt en erklæring kjøringens eget bevis viser ble lest. `mandate.criteria_block` er ENESTE
renderer og TOM uten kriterier — omisjon, aldri en tom overskrift (`announce`-regelen), og dét er
hva som holder hver prompt i hver ukommisjonert kjøring byte-identisk, demoens inkludert.
**A2s plassering er MÅLT:** på utforskningsstien finnes ingenting å bære — `main()` sender
`explore()` ingen `success_criteria` i det hele tatt, så objektivet ER prompten og når oppgaven
alt. Load-bearing MÅLT (`tests/test_right_requirement_loadbearing.py`, 8 armer).
- **En parse-feil brenner ikke lenger rundeboka i stillhet, og annonseringen navngir hva den
handler om (P20 DEL C, 15.09):** `_fetch_parsed` prøvde på nytt med den BYTE-IDENTISKE prompten.
MÅLT (P19 F4): `kontrakt-sorasen-04` etterlot `{run_id}-parse-failures.json` med **elleve** rader,
hver av dem samme feil (`claimed_saving_nok` ≤ 0) — elleve av kjøringens tolv runder, brukt på å
spørre om igjen uten å si hva som var galt. Steg 5s `prior_rejection` bærer VALIDATOR-avvisninger,
og et svar som aldri parset når aldri en validator, så ingen eksisterende blokk kunne bære det.
`_fetch_parsed` tar nå en BYGGER i stedet for en ferdig meldingsliste — kallerens egen
`_build_messages`-binding, så retry-løkka ikke kan komponere en prompt den ytre løkka ikke ville
komponert (kø-(p)) — og grunnen er per RETRY, tom på forsøk 1, så attempt 1 er byte-identisk. Den
verbatime fangsten (Fase 1b funn 1) står URØRT. **C2:** `--across-bundle` tar ingen
`--project-id`, så annonseringen sa «Run mandate for the portfolio» om en kommisjon dispatchet
over to navngitte baser (P17b F5); `announced_subject` navngir de rutede basene ved ERKLÆRT id, og
en base som ikke lar seg løse faller tilbake til katalognavnet — `dimension_label`-presedensen
ordrett: å annonsere skal ALDRI endre hvilken feil en operatør ser. Load-bearing MÅLT
(`tests/test_parse_error_feedback_loadbearing.py`, 7 armer).
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet. - **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase. - Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -481,6 +481,17 @@ when the seam is detached, so the loop cannot silently degrade into theater.
that does not use identifiers cannot be answered in them. A code the project's cost baseline that does not use identifiers cannot be answered in them. A code the project's cost baseline
carries is never refused for its shape, because stage 0 has already ruled it a real line. carries is never refused for its shape, because stage 0 has already ruled it a real line.
**A clause number is not a price.** A standard numbers its own requirements and a process code
numbers its own settlement posts, and both look exactly like a bare decimal: `10.4`, `1.10.4`,
`12.1`. When a run is **unanchored** — the knowledge base ships no cost baseline, so the
validator's stage 0 never runs — a cost code shaped like a clause number is refused unless it is
one of the reference numbers the base itself declares (`req_number`, `prosessnr`, `seksjon` in a
document's own frontmatter). The refusal names the denominator: *"not one of the 2765 this
knowledge base declares"*. A code the base does declare still validates, so a project priced in
real process codes is untouched; an input that declares no reference numbers at all cannot trip
the rule at all. `code_forms` reports the third value, `requirement`, for a code the base does
declare. Anchored runs are unaffected — stage 0 is the stronger falsifier and keeps the ruling.
**Naming the requirement that binds a direction.** Whoever navigates a knowledge base — the **Naming the requirement that binds a direction.** Whoever navigates a knowledge base — the
exploration's hypothesiser and, since the debate started navigating, the proposer — is asked to exploration's hypothesiser and, since the debate started navigating, the proposer — is asked to
name the ONE requirement that binds the direction it commits to, and to declare it with name the ONE requirement that binds the direction it commits to, and to declare it with
@ -491,7 +502,23 @@ when the seam is detached, so the loop cannot silently degrade into theater.
`"why_none"` — a base that holds no requirement for a direction is a finding worth stating, and `"why_none"` — a base that holds no requirement for a direction is a finding worth stating, and
the field is never simply omitted. Where an approach carries one, the proposer's prompt names it the field is never simply omitted. Where an approach carries one, the proposer's prompt names it
and asks for it back verbatim in `measure`, and the declaration is written to and asks for it back verbatim in `measure`, and the declaration is written to
`{run_id}-debate.json` / `{run_id}-exploration.json` under `requirements`. `{run_id}-debate.json` / `{run_id}-exploration.json` under `requirements`. The reply gives back
**the document's own** `title` and `req_number`, read off the base rather than echoed from the
arguments, plus the sentence that says what the declaration binds — so a declaration of the wrong
requirement can be seen to be wrong. The route is named in the instruction as well as the
answer: pass `read_dir` a `filter` word from the approach's own label.
**What the commissioner counts as success reaches the readers.** A mandate's `success_criteria`
used to reach the announcement and nothing else. It is now restated verbatim in the debate's task
message — the prompt where `declare_requirement` is available — through one renderer, and is
omitted entirely when the commission states none.
**A reply that could not be parsed says why, once.** A malformed reply used to be retried with
the byte-identical prompt; measured, one paid run spent eleven of its twelve rounds re-asking a
question it had already answered the same wrong way. The next attempt's prompt now carries the
parse reason — only the reason, never the discarded JSON — as a block of its own, distinct from
a validator rejection (the numbers were refuted) and from expert feedback (a person objected).
The verbatim reply is still captured to `{run_id}-parse-failures.json` as before.
**Answering the plan review (`--plan-review`).** With `enable_plan_review` set, the exploration **Answering the plan review (`--plan-review`).** With `enable_plan_review` set, the exploration
stops before the loop is allowed to run and asks you to sign the plan off. `--plan-review` stops before the loop is allowed to run and asks you to sign the plan off. `--plan-review`

View file

@ -211,9 +211,13 @@ _INSTRUCTIONS: Final = {
HYPOTHESISER_ROLE: ( HYPOTHESISER_ROLE: (
"You shape ONE candidate cost-saving direction at a time from what the navigator found. " "You shape ONE candidate cost-saving direction at a time from what the navigator found. "
"BEFORE you commit to a direction, name the ONE requirement in the knowledge base that " "BEFORE you commit to a direction, name the ONE requirement in the knowledge base that "
"BINDS it: have the navigator find it with read_dir(filter=...) and read it with " "BINDS it: pass read_dir a 'filter' word taken from the approach's own label — "
"read_file, then call declare_requirement with the base id, that path and the " "filter='rundkjoring' finds the level's requirements about roundabouts, and one of them is "
"requirement's own number. A direction with no requirement behind it is a guess. " "the 'Krav 4.1.2-1' you are looking for — read it with read_file, then call "
"declare_requirement with the base id, that path and the requirement's own number. The "
"reply gives back the document's own title and number: if they are not about your measure, "
"you declared the wrong requirement and should filter again. A direction with no "
"requirement behind it is a guess. "
"You may call quick_validate to sanity-check a candidate's numbers; its verdict is " "You may call quick_validate to sanity-check a candidate's numbers; its verdict is "
"ADVISORY and is not the project's decision. When you commit to a direction, end your " "ADVISORY and is not the project's decision. When you commit to a direction, end your "
f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", ' f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", '
@ -1001,6 +1005,26 @@ def _resolve_bundle(index: Mapping[str, str], bundle_id: str) -> str:
return index[bundle_id] return index[bundle_id]
def _declared_document(index: Mapping[str, str], bundle_id: str, path: str) -> tuple[str, str]:
"""``(title, reference_number)`` of the document at ``path``, or ``("", "")`` when the base has
no navigated concept under that name.
Read off the SAME ``Bundle.context_files`` every listing rung is built from (MAJOR-3/S7a-3), so
a declaration can never be answered with the title of a ``type: verdict`` document the one
layer no listing names and ``read_file`` refuses outright.
A path the base does not carry as a concept ``index.md`` is the reachable case answers with
two empty strings rather than raising: the declaration itself has already been accepted by the
read-trace check above, and turning "I cannot restate your title" into a refusal would fail a
declaration the run's own trace proves was read.
"""
bundle = okf.navigate_bundle(_resolve_bundle(index, bundle_id))
for file in bundle.context_files:
if file.name == path:
return okf.unquote_scalar(file.frontmatter.get("title", "")), okf.reference_number(file)
return "", ""
#: Characters of the root index body one catalogue entry may carry. The catalogue's job is to let a #: Characters of the root index body one catalogue entry may carry. The catalogue's job is to let a
#: manager pick a base, not to read one, so the excerpt is a fixed-size window rather than a share #: manager pick a base, not to read one, so the excerpt is a fixed-size window rather than a share
#: of the base: cost then scales with how many bases are configured, which the operator chose, and #: of the base: cost then scales with how many bases are configured, which the operator chose, and
@ -1364,7 +1388,11 @@ def navigator_tools(
"requirement's own number as its frontmatter states it. You must have READ the " "requirement's own number as its frontmatter states it. You must have READ the "
"document with read_file first: a declaration naming a path this run never opened is " "document with read_file first: a declaration naming a path this run never opened is "
"refused, and reading it is the correction. Use read_dir with a 'filter' word to find " "refused, and reading it is the correction. Use read_dir with a 'filter' word to find "
"it, read_file to read it, then declare it." "it (a word from the approach's own label works: filter='rundkjoring' -> "
"'Krav 4.1.2-1'), read_file to read it, then declare it. The reply gives back the "
"document's own "
"title and number, so you can see whether you declared the requirement you meant: a "
"declaration of a requirement that is not about the measure is worth nothing."
), ),
) )
def declare_requirement(bundle_id: str, path: str, ref: str) -> dict[str, Any]: def declare_requirement(bundle_id: str, path: str, ref: str) -> dict[str, Any]:
@ -1386,7 +1414,26 @@ def navigator_tools(
"direction" "direction"
) )
requirements.append(DeclaredRequirement(bundle_id=bundle_id, path=path, ref=ref)) requirements.append(DeclaredRequirement(bundle_id=bundle_id, path=path, ref=ref))
return {"declared": True, "bundle_id": bundle_id, "path": path, "ref": ref} # P20/A1: give back the DOCUMENT's own title and number, read off the base rather than
# echoed from the arguments. MEASURED (P19 round 3, P17b): 13 declarations over 5 runs and
# NOT ONE named a fasit concept — the tool answered ``{"declared": true, ...}`` to every
# declaration, so a model that had declared the wrong requirement was told it had succeeded.
# ``okf.reference_number`` is the ONE reader of "which requirement is this" (kø-(p)), and
# ``binds`` says out loud what the declaration is for: without it the reply is data with no
# instruction, and the instruction is the whole correction.
declared = _declared_document(index, bundle_id, path)
return {
"declared": True,
"bundle_id": bundle_id,
"path": path,
"ref": ref,
"title": declared[0],
"req_number": declared[1],
"binds": (
f"This declaration says {ref} is the requirement the proposal rests on; a "
"declaration of a requirement that is not about the measure is worth nothing."
),
}
tools = [list_bundles, read_bundle, read_dir, read_file] tools = [list_bundles, read_bundle, read_dir, read_file]
if requirements is not None: if requirements is not None:

View file

@ -286,6 +286,7 @@ def _build_messages(
*, *,
approach: Approach | None = None, approach: Approach | None = None,
prior_feedback: str | None = None, prior_feedback: str | None = None,
parse_error: str | None = None,
) -> list[Message]: ) -> list[Message]:
"""Build the hypothesis prompt. When ``prior_rejection`` is set (Step 5, målbilde §5/§7), """Build the hypothesis prompt. When ``prior_rejection`` is set (Step 5, målbilde §5/§7),
append a revision block carrying ONLY the falsification *reason* verbatim never the prior append a revision block carrying ONLY the falsification *reason* verbatim never the prior
@ -308,8 +309,17 @@ def _build_messages(
"your numbers were refuted"; conflating the two would tell the model the machine objected "your numbers were refuted"; conflating the two would tell the model the machine objected
when a person did. ``None`` -> the byte-identical base prompt, like the other two. when a person did. ``None`` -> the byte-identical base prompt, like the other two.
All three are composable, and the ORDER is fixed: base -> approach head -> rejection -> When ``parse_error`` is set (P20/C1) a FOURTH block carries the reason the PREVIOUS reply
feedback. A prompt can legitimately carry a rejection AND a feedback at once that is the could not be parsed the same "only the reason, never the JSON" rule the other two follow.
It is a different instruction from both: a rejection means the numbers were refuted and a
feedback means a person objected, while this one means nothing was ever read. MEASURED (P19
F4): ``_fetch_parsed`` retried with the byte-identical prompt, and one round-3 run
(``kontrakt-sorasen-04``) spent ELEVEN of its twelve rounds on replies that all failed the
same way ``claimed_saving_nok: 0`` because nothing ever told the model what was wrong.
``None`` -> the byte-identical base prompt, like the other three.
All four are composable, and the ORDER is fixed: base -> approach head -> rejection ->
feedback -> parse error. A prompt can legitimately carry a rejection AND a feedback at once that is the
attempt after a revise whose bought attempt the validator then rejected: the human's attempt after a revise whose bought attempt the validator then rejected: the human's
instruction STANDS until the human next answers, while the machine's reason is per-attempt instruction STANDS until the human next answers, while the machine's reason is per-attempt
(only the most recent, as today). (only the most recent, as today).
@ -362,6 +372,14 @@ def _build_messages(
f"Expert feedback: {prior_feedback}\n" f"Expert feedback: {prior_feedback}\n"
"Produce a REVISED SavingsProposal that follows this feedback." "Produce a REVISED SavingsProposal that follows this feedback."
) )
if parse_error is not None:
prompt += (
"\n\nYour previous reply could not be PARSED as a SavingsProposal, so it was "
"discarded before any validator saw it.\n"
f"Reason: {parse_error}\n"
"Reply with a SavingsProposal whose claimed_saving_nok is greater than 0 and whose "
"affected_items each carry code, quantity and unit_cost."
)
return [Message(role="user", contents=[prompt])] return [Message(role="user", contents=[prompt])]
@ -472,7 +490,13 @@ def _grounding_text(
*delivered.documents, *delivered.documents,
*(item.code for item in project.cost_items), *(item.code for item in project.cost_items),
*(() if baseline is None else baseline.items), *(() if baseline is None else baseline.items),
) ),
# P20/B: the base's own vocabulary of clause numbers travels WITH the text it was read
# off. Carried through rather than recomposed: ``run_project`` walks the base once and
# composes both halves there, and a second derivation here would be free to disagree with
# the documents it is supposed to describe (kø-(p)). The two later sources are cost CODES,
# which declare nothing, so they contribute none.
declared_references=delivered.declared_references,
) )
@ -615,10 +639,19 @@ async def generate_via_llm(
from discarding what it already knew. Never a malformed proposal; raises ``BudgetExceeded`` from discarding what it already knew. Never a malformed proposal; raises ``BudgetExceeded``
when the meter cap is crossed.""" when the meter cap is crossed."""
async def _fetch_parsed(messages: list[Message]) -> SavingsProposal: async def _fetch_parsed(build: Callable[[str | None], list[Message]]) -> SavingsProposal:
# Parse-robust: a malformed/text-leaked reply is retried; the meter caps total work. # Parse-robust: a malformed/text-leaked reply is retried; the meter caps total work.
#
# P20/C1: the retry is no longer BLIND. It takes a BUILDER rather than a finished message
# list, because the whole defect was that the same bytes were re-sent: measured, one
# round-3 run burned 11 of its 12 rounds on replies that all failed identically. The
# builder is the caller's own ``_build_messages`` binding, so this loop cannot compose a
# prompt the outer loop would not have composed (kø-(p)); the reason is per-RETRY, like
# ``prior_rejection`` is per-attempt, and starts empty so attempt 1 is byte-identical.
parse_error: str | None = None
while True: while True:
meter.tick_round() # between-attempt bound (BudgetExceeded over cap) meter.tick_round() # between-attempt bound (BudgetExceeded over cap)
messages = build(parse_error)
# Fase 1b, funn 1b: hand the model a GRAMMAR, not a prose request. The prompt's # Fase 1b, funn 1b: hand the model a GRAMMAR, not a prose request. The prompt's
# "Respond with ONLY a JSON object" line stays — a provider that ignores # "Respond with ONLY a JSON object" line stays — a provider that ignores
# ``response_format`` (or a local model that does not implement it) must still be told # ``response_format`` (or a local model that does not implement it) must still be told
@ -633,10 +666,10 @@ async def generate_via_llm(
# Capture BEFORE the retry: this reply was paid for, and once ``continue`` runs the # Capture BEFORE the retry: this reply was paid for, and once ``continue`` runs the
# only record of what the model actually said is gone (Fase 1b, funn 1). Verbatim — # only record of what the model actually said is gone (Fase 1b, funn 1). Verbatim —
# the operator is diagnosing a format failure, so any shortening removes evidence. # the operator is diagnosing a format failure, so any shortening removes evidence.
reason = f"{type(exc).__name__}: {exc}"
if parse_failures is not None: if parse_failures is not None:
parse_failures.append( parse_failures.append(ParseFailure(text=reply.text, error=reason))
ParseFailure(text=reply.text, error=f"{type(exc).__name__}: {exc}") parse_error = reason
)
continue continue
last: Rejection | None = None last: Rejection | None = None
@ -672,14 +705,16 @@ async def generate_via_llm(
# accumulated history (bounded prompt growth). # accumulated history (bounded prompt growth).
if last is not None: if last is not None:
fed_back.append(last) fed_back.append(last)
messages = _build_messages( candidate = await _fetch_parsed(
project, lambda parse_error: _build_messages(
context, project,
prior_rejection=last, context,
approach=approach, prior_rejection=last,
prior_feedback=feedback, approach=approach,
prior_feedback=feedback,
parse_error=parse_error,
)
) )
candidate = await _fetch_parsed(messages)
if pending_revise is not None and reviews is not None: if pending_revise is not None and reviews is not None:
reviews[pending_revise] = replace(reviews[pending_revise], honoured=True) reviews[pending_revise] = replace(reviews[pending_revise], honoured=True)
pending_revise = None pending_revise = None

View file

@ -414,6 +414,29 @@ def announce(
return "\n".join(lines) return "\n".join(lines)
def criteria_block(success_criteria: str) -> str:
"""The ONE rendering of a commission's success criteria INTO a prompt, or ``""``.
MEASURED (P19 F2): ``success_criteria`` reached ``announce`` and nothing else, so the operator's
own statement of what a good answer looks like was printed for a human and withheld from the
only reader who could act on it. The approach's ``description`` has always reached the
generation prompt (``generate._build_messages``); this is its run-level sibling.
ONE composer, for -(p): the debate task and any later prompt that carries the criteria must
say the same thing about them, and two renderings of one commission are free to disagree about
what the operator asked for.
Empty in, empty out omission rather than an empty heading, the ``announce`` rule. That is
what keeps every prompt of every un-commissioned run, the demo's included, byte-identical.
"""
if not success_criteria:
return ""
return (
"\nWhat the commissioner counts as success (restated verbatim from the commission):\n"
f"{success_criteria}\n"
)
def settle( def settle(
coverage: tuple[ApproachOutcome, ...], coverage: tuple[ApproachOutcome, ...],
*, *,

View file

@ -1224,6 +1224,53 @@ _DIRECTORY_PAGE_MAX: Final = 50
#: hence ``unquote_scalar``, this repo's ONE de-quoting rule). #: hence ``unquote_scalar``, this repo's ONE de-quoting rule).
_FILTER_FIELDS: Final = ("req_number", "prosessnr") _FILTER_FIELDS: Final = ("req_number", "prosessnr")
#: Which frontmatter keys DECLARE a reference number — the base's own vocabulary of requirement and
#: process numbers. Derived from ``_FILTER_FIELDS`` rather than restating it (two copies of the two
#: measured keys is the kø-(p) drift), plus ``seksjon``, which is the field that carries the BARE
#: form-3 number on the N corpora and on R761 (``seksjon: '10.4.3'``, ``seksjon: '11.11'``) while
#: ``req_number`` carries the composed ``Krav 10.4.3—1``. MEASURED 15.09 over the four delivered
#: bases: n100 553 distinct declared numbers, n200 1 440, n500 365, r761 2 765.
#:
#: NOT added to ``_FILTER_FIELDS`` itself, and that is deliberate: what a ``filter`` word searches
#: is measured and gated (P18), and widening it would change ``total_matches`` for every navigator
#: call. These two constants answer different questions — "what does a filter look in" and "what
#: does this document declare as its number" — over ONE list of the measured reference keys.
REFERENCE_NUMBER_FIELDS: Final = (*_FILTER_FIELDS, "seksjon")
def reference_number(file: BundleFile) -> str:
"""The document's OWN reference number as its top-level frontmatter states it, or ``""``.
First non-empty of ``_FILTER_FIELDS``, in that order: the N corpora declare ``req_number``
("Krav 4.1.2—1") and R761 declares ``prosessnr`` ("'11.11'", quoted hence ``unquote_scalar``),
and MEASURED no delivered document declares both. ``seksjon`` is deliberately NOT read here:
this answers "which requirement IS this", and a section number names the chapter a requirement
sits in, not the requirement.
Read off ``BundleFile.frontmatter``, so P15's top-level-wins rule applies and a nested
``sources:`` entry can never answer for the concept.
"""
for key in _FILTER_FIELDS:
value = unquote_scalar(file.frontmatter.get(key, ""))
if value:
return value
return ""
def declared_reference_numbers(file: BundleFile) -> tuple[str, ...]:
"""Every reference number this ONE document declares (``REFERENCE_NUMBER_FIELDS``), de-quoted.
The unit of the reference VOCABULARY the validator's stage 0b checks a requirement-shaped code
against (P20/B). Per DOCUMENT rather than per base, because the caller composing the grounding
already walks the base once and the boundaries it composes are the ones the gate must read
(P18/B1's rule: one document is one unit, never a blob).
"""
return tuple(
value
for key in REFERENCE_NUMBER_FIELDS
if (value := unquote_scalar(file.frontmatter.get(key, "")))
)
def _matches_filter(file: BundleFile, needle: str) -> bool: def _matches_filter(file: BundleFile, needle: str) -> bool:
"""Case-insensitive SUBSTRING over the document's title and its reference number. """Case-insensitive SUBSTRING over the document's title and its reference number.

View file

@ -98,6 +98,7 @@ from portfolio_optimiser.mandate import (
MandateRoutingError, MandateRoutingError,
announce, announce,
candidate_from_approach, candidate_from_approach,
criteria_block,
load_mandate, load_mandate,
route_by_bundle, route_by_bundle,
settle, settle,
@ -1131,6 +1132,10 @@ async def run_project(
# gate's share rule needs, and composing them here — where the base is already walked — is what # gate's share rule needs, and composing them here — where the base is already walked — is what
# keeps them from being a second, drifting reconstruction (kø-(p)). # keeps them from being a second, drifting reconstruction (kø-(p)).
bundle_grounding: tuple[str, ...] = () bundle_grounding: tuple[str, ...] = ()
#: P20/B: the reference numbers the base's documents DECLARE, composed in the SAME walk as the
#: documents above, so the gate's vocabulary and the gate's text describe one reading of one
#: base. Empty on the road path, which is what keeps the rule unable to fire there.
bundle_references: tuple[str, ...] = ()
# S2c: a CALLER-OWNED sink for what the debate opens (the ``parse_failures``/``ExplorationTrace`` # S2c: a CALLER-OWNED sink for what the debate opens (the ``parse_failures``/``ExplorationTrace``
# shape). A returned value would be lost on exactly the run that most needs the evidence — a # shape). A returned value would be lost on exactly the run that most needs the evidence — a
# budget stop mid-debate raises out of ``debate.run`` and constructs no ``RunResult`` at all. # budget stop mid-debate raises out of ``debate.run`` and constructs no ``RunResult`` at all.
@ -1151,6 +1156,9 @@ async def run_project(
bundle_grounding = tuple( bundle_grounding = tuple(
"\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files "\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files
) )
bundle_references = tuple(
ref for f in bundle.context_files for ref in okf.declared_reference_numbers(f)
)
# ONE bundle-id rule (Step 10, slackened S7a-3 pkt. 1): the DECLARED id is the identity and # ONE bundle-id rule (Step 10, slackened S7a-3 pkt. 1): the DECLARED id is the identity and
# the mount is carried alongside, so a base delivered under a directory name of its own is # the mount is carried alongside, so a base delivered under a directory name of its own is
# opened rather than refused. What is still refused, before a single model call: a base # opened rather than refused. What is still refused, before a single model call: a base
@ -1272,7 +1280,9 @@ async def run_project(
# ``_fetch_parsed`` has returned, so a report from there could only ever speak once an attempt # ``_fetch_parsed`` has returned, so a report from there could only ever speak once an attempt
# had been paid for. ONE binding feeding both the report and the gate: two compositions of one # had been paid for. ONE binding feeding both the report and the gate: two compositions of one
# text are free to disagree, which is exactly what a report must not be able to do (kø-(p)). # text are free to disagree, which is exactly what a report must not be able to do (kø-(p)).
delivered = Grounding(documents=(context, *bundle_grounding)) delivered = Grounding(
documents=(context, *bundle_grounding), declared_references=bundle_references
)
offer = grounding_offer(project, baseline, delivered) offer = grounding_offer(project, baseline, delivered)
# Trekk B2 (krav 3): configured MCP servers become tools the AGENTS can call during the debate. # Trekk B2 (krav 3): configured MCP servers become tools the AGENTS can call during the debate.
@ -1359,8 +1369,16 @@ async def run_project(
async with AsyncExitStack() as mcp_stack: async with AsyncExitStack() as mcp_stack:
for live_tool in live_mcp_tools: for live_tool in live_mcp_tools:
await mcp_stack.enter_async_context(live_tool) await mcp_stack.enter_async_context(live_tool)
# P20/A2: the commission's success criteria reach the DEBATE — the prompt where
# ``declare_requirement`` is available — and not only the announcement. Composed through
# ``mandate.criteria_block``, the ONE renderer (kø-(p)); empty without a commission, so
# every un-commissioned run's task message is byte-identical, which is what keeps the
# golden transcript unchanged. MEASURED: on the exploration path there is nothing to
# carry — ``main()`` passes ``explore()`` no ``success_criteria`` at all, so its
# objective IS the prompt and already reaches the task.
criteria = criteria_block(mandate.success_criteria) if mandate is not None else ""
result = await debate.run( result = await debate.run(
f"Find a cost-saving measure for {project.id}.\nContext:\n{context}" f"Find a cost-saving measure for {project.id}.{criteria}\nContext:\n{context}"
) )
finally: finally:
# ``finally``, the ``write_parse_failures`` precedent: any exception leaving the debate — # ``finally``, the ``write_parse_failures`` precedent: any exception leaving the debate —
@ -1610,7 +1628,12 @@ async def run_project(
cost_baseline_anchored=baseline is not None, cost_baseline_anchored=baseline is not None,
# P19/B2: what the run made of each code it was handed. Derived from the SAME classifier # P19/B2: what the run made of each code it was handed. Derived from the SAME classifier
# the gate uses (kø-(p)), off the proposal being stamped — never re-read from anywhere. # the gate uses (kø-(p)), off the proposal being stamped — never re-read from anywhere.
code_forms=classify_codes([item.code for item in proposal.affected_items]), code_forms=classify_codes(
[item.code for item in proposal.affected_items],
# P20/B: the THIRD value, ``requirement``, needs the base's own vocabulary. Read off
# the SAME ``delivered`` the gate was handed, never a second composition.
delivered,
),
# WHICH corpus was judged, and whether the base named itself or the mount named it for it. # WHICH corpus was judged, and whether the base named itself or the mount named it for it.
# Read off the SAME resolution the run opened the base with (kø-(p)); ``None`` on the road # Read off the SAME resolution the run opened the base with (kø-(p)); ``None`` on the road
# path, where no knowledge base exists to name. # path, where no knowledge base exists to name.
@ -2290,6 +2313,32 @@ def _write_multibase_summary(
) )
def announced_subject(project_id: str | None, across_bundle: Sequence[str]) -> str:
"""WHO the announcement is about: the project, the routed bases, or the portfolio (P20/C2).
MEASURED (P17b F5): ``--across-bundle`` takes no ``--project-id`` each base's project is read
from that base's own IR projection — so ``args.project_id or "the portfolio"`` announced a
multi-base commission as a portfolio pass, which is a different mode entirely.
The ids are resolved for NAMING ONLY, and a base that cannot be resolved falls back to its
directory name. That is the ``dimension_label`` precedent one line above the call site,
verbatim: a configuration that fails to load is left unnamed here and refused a moment later by
the dispatch, which stays the single owner of that refusal announcing must never change which
error an operator sees.
"""
if project_id:
return project_id
if not across_bundle:
return "the portfolio"
names = []
for raw in across_bundle:
try:
names.append(okf.reconcile_bundle_id(raw).id)
except (okf.BundleIdMismatch, FileNotFoundError, ValueError, OSError):
names.append(Path(raw).name)
return ", ".join(names)
def resolve_bundle_routing( def resolve_bundle_routing(
bundle_dirs: Sequence[str], bundle_dirs: Sequence[str],
) -> tuple[tuple[str, str, str], ...]: ) -> tuple[tuple[str, str, str], ...]:
@ -4045,7 +4094,7 @@ def main(argv: list[str] | None = None) -> int:
print( print(
announce( announce(
mandate, mandate,
project_id=args.project_id or "the portfolio", project_id=announced_subject(args.project_id, args.across_bundle or ()),
# From ARGV, never the constants (P16 B2). The announcement is the one thing # From ARGV, never the constants (P16 B2). The announcement is the one thing
# printed BEFORE the first paid call, and its whole job is to say what this run # printed BEFORE the first paid call, and its whole job is to say what this run
# will do; reading the defaults was correct only while main() could not do # will do; reading the defaults was correct only while main() could not do

View file

@ -250,29 +250,50 @@ _GROUNDING_MIN_INERT_DOCUMENTS: Final = 10
#: Bare numbers are deliberately EXCLUDED, with the number: K2 carries 46 394 occurrences over #: Bare numbers are deliberately EXCLUDED, with the number: K2 carries 46 394 occurrences over
#: 2 117 distinct values (P7 § 2), so counting them would make every report positive and the #: 2 117 distinct values (P7 § 2), so counting them would make every report positive and the
#: measurement inert — the repo's cardinal class, a gate that can only come out green. #: measurement inert — the repo's cardinal class, a gate that can only come out green.
#: The four forms, NAMED rather than reached by index: P20/B needs two of them by themselves
#: (a requirement/process number), and ``IDENTIFIER_FORMS[1:3]`` in a second module would be a
#: positional dependency on a tuple literal — the kø-(p) shape with no compiler to catch it.
_FORM_SEPARATED_UPPER: Final = re.compile(r"\b[A-ZÆØÅ][A-ZÆØÅ0-9]*(?:[-_][A-ZÆØÅ0-9]+)+\b")
_FORM_REQUIREMENT_NUMBER: Final = re.compile(r"Krav\s+\d+(?:\.\d+)*\s*[\u2014-]\s*\d+(?:_\d+)?")
_FORM_PROCESS_NUMBER: Final = re.compile(r"(?<![\d.])[1-9]\d{0,2}(?:\.\d{1,3}){1,4}\b(?!\.\d)")
_FORM_PROCESS_HEADING: Final = re.compile(r"(?<!\d )(?<![\d.])[1-9]\d{0,2} [A-ZÆØÅ]{5,}\b")
IDENTIFIER_FORMS: Final = ( IDENTIFIER_FORMS: Final = (
# ``SHA-01``, ``RIM-02``, ``B-20-00-00``, ``FOR-2011-12-06-1357`` (K2's 50) — and, since P19/B2 # ``SHA-01``, ``RIM-02``, ``B-20-00-00``, ``FOR-2011-12-06-1357`` (K2's 50) — and, since P19/B2
# made the same forms decide ``prose`` vs ``identifier``, an UPPERCASE separated token with no # made the same forms decide ``prose`` vs ``identifier``, an UPPERCASE separated token with no
# digits at all. MEASURED: this repo's own ``ENERGI-TOTAL-EL`` matched neither of the pre-P19 # digits at all. MEASURED: this repo's own ``ENERGI-TOTAL-EL`` matched neither of the pre-P19
# forms, so the classifier called a real cost code prose; a gate is only allowed to be wrong in # forms, so the classifier called a real cost code prose; a gate is only allowed to be wrong in
# the direction that admits too much. # the direction that admits too much.
re.compile(r"\b[A-ZÆØÅ][A-ZÆØÅ0-9]*(?:[-_][A-ZÆØÅ0-9]+)+\b"), _FORM_SEPARATED_UPPER,
# ``Krav 3.3.1—13`` (the N corpora's dominant form). EM-DASH U+2014 AND the hyphen, because the # ``Krav 3.3.1—13`` (the N corpora's dominant form). EM-DASH U+2014 AND the hyphen, because the
# binding known positive is the em-dash spelling and only the em-dash spelling scores 6 of 6. # binding known positive is the em-dash spelling and only the em-dash spelling scores 6 of 6.
# The trailing ``(?:_\d+)?`` is MEASURED, not defensive: one of the 26 fasit references is # The trailing ``(?:_\d+)?`` is MEASURED, not defensive: one of the 26 fasit references is
# ``Krav 3.3.2—1_1``, and without it the classifier called that real reference prose. # ``Krav 3.3.2—1_1``, and without it the classifier called that real reference prose.
re.compile(r"Krav\s+\d+(?:\.\d+)*\s*[\u2014-]\s*\d+(?:_\d+)?"), _FORM_REQUIREMENT_NUMBER,
# R761's process numbers, ``12.1`` / ``52.11`` (P19 B1). MEASURED: all six ``ref`` values in # R761's process numbers, ``12.1`` / ``52.11`` (P19 B1). MEASURED: all six ``ref`` values in
# ``contexts/kontrakt-sorasen-2027/fasit.json`` are of this shape and NEITHER of the first two # ``contexts/kontrakt-sorasen-2027/fasit.json`` are of this shape and NEITHER of the first two
# forms matches one of them, so r761's whole offer was 3 identifiers over 6.5 MB. The trailing # forms matches one of them, so r761's whole offer was 3 identifiers over 6.5 MB. The trailing
# ``(?!\.\d)`` is what keeps a Norwegian date out: ``15.09.2026`` would otherwise contribute # ``(?!\.\d)`` is what keeps a Norwegian date out: ``15.09.2026`` would otherwise contribute
# its ``15.09`` prefix, and a date is not a requirement. # its ``15.09`` prefix, and a date is not a requirement.
re.compile(r"(?<![\d.])[1-9]\d{0,2}(?:\.\d{1,3}){1,4}\b(?!\.\d)"), _FORM_PROCESS_NUMBER,
# ``65 ASFALTDEKKER`` — a process number and its heading, the form a price schedule's section # ``65 ASFALTDEKKER`` — a process number and its heading, the form a price schedule's section
# rows carry (P18 § 2 measured it at 29 of 2 756 documents). # rows carry (P18 § 2 measured it at 29 of 2 756 documents).
re.compile(r"(?<!\d )(?<![\d.])[1-9]\d{0,2} [A-ZÆØÅ]{5,}\b"), _FORM_PROCESS_HEADING,
) )
#: The two forms a REQUIREMENT or PROCESS number takes (P20/B). MEASURED over the four delivered
#: bases: these are the shapes that live in ``req_number`` ("Krav 4.1.2—1"), ``prosessnr``
#: ("'11.11'") and ``seksjon`` ("'10.4.3'") — the fields a base uses to number its own clauses.
#: The other two forms are NOT here: ``SHA-01``-style tokens are what a price schedule's cost lines
#: look like, and ``65 ASFALTDEKKER`` IS a schedule section row.
REQUIREMENT_FORMS: Final = (_FORM_REQUIREMENT_NUMBER, _FORM_PROCESS_NUMBER)
def has_requirement_form(code: str) -> bool:
"""Whether ``code`` is shaped like a requirement or process number. FULL-MATCH, never a search,
for ``has_identifier_form``'s reason: ``impulsventilator 12.1`` is not a clause number."""
return any(form.fullmatch(code) for form in REQUIREMENT_FORMS)
def identifier_tokens(text: str) -> set[str]: def identifier_tokens(text: str) -> set[str]:
"""Every DISTINCT token of any ``IDENTIFIER_FORMS`` shape in ``text``. One reader, two callers. """Every DISTINCT token of any ``IDENTIFIER_FORMS`` shape in ``text``. One reader, two callers.
@ -296,9 +317,27 @@ def has_identifier_form(code: str) -> bool:
return any(form.fullmatch(code) for form in IDENTIFIER_FORMS) return any(form.fullmatch(code) for form in IDENTIFIER_FORMS)
def classify_codes(codes: Sequence[str]) -> dict[str, str]: def classify_codes(codes: Sequence[str], grounding: Grounding | None = None) -> dict[str, str]:
"""``{code: "identifier" | "prose"}`` — P19/B2's report, in ONE place for both consumers.""" """``{code: "identifier" | "prose" | "requirement"}`` — P19/B2's report, widened by P20/B, in
return {code: "identifier" if has_identifier_form(code) else "prose" for code in codes} ONE place for both consumers.
``requirement`` is the third value: a code shaped like a clause number (``REQUIREMENT_FORMS``)
AND declared as one by the input's own documents. It is a REPORT about what the run made of a
code, not the gate the gate is ``_reference_refusal`` below and fires on the COMPLEMENT, a
requirement-shaped code the base's vocabulary does NOT contain.
``grounding=None`` is the pre-P20 answer exactly: without the input there is no vocabulary to
check against, so no code can be called a requirement. ``stress.py`` re-derives with ``None``
for runs that predate the field, and says so.
"""
vocabulary = frozenset() if grounding is None else grounding.reference_vocabulary
out: dict[str, str] = {}
for code in codes:
if has_requirement_form(code) and code in vocabulary:
out[code] = "requirement"
else:
out[code] = "identifier" if has_identifier_form(code) else "prose"
return out
@dataclass(frozen=True) @dataclass(frozen=True)
@ -320,6 +359,16 @@ class Grounding:
""" """
documents: tuple[str, ...] documents: tuple[str, ...]
#: Every reference number the input's documents DECLARE in their own top-level frontmatter
#: (``okf.declared_reference_numbers``), one entry per declaration — the base's own vocabulary
#: of requirement, process and section numbers. P20/B checks a requirement-shaped code against
#: it.
#:
#: DEFAULTED, the ``skipped_links`` half rather than ``cost_baseline_anchored``'s: an empty
#: vocabulary is an honest POSITIVE statement ("this input declares no clause numbers"), and it
#: is what keeps every caller written before today — the road path, every fixture, ``of`` —
#: unchanged by construction, since the gate cannot fire without one.
declared_references: tuple[str, ...] = ()
@classmethod @classmethod
def of(cls, text: str) -> Grounding: def of(cls, text: str) -> Grounding:
@ -335,6 +384,11 @@ class Grounding:
"""How many of the documents contain ``token`` — the numerator, in the unit of the rule.""" """How many of the documents contain ``token`` — the numerator, in the unit of the rule."""
return sum(1 for document in self.documents if token in document) return sum(1 for document in self.documents if token in document)
@cached_property
def reference_vocabulary(self) -> frozenset[str]:
"""The DISTINCT reference numbers this input declares. The denominator P20/B names."""
return frozenset(self.declared_references)
@cached_property @cached_property
def identifiers(self) -> frozenset[str]: def identifiers(self) -> frozenset[str]:
"""Every distinct identifier-shaped token this input OFFERS (P8's count, P19/B3's guard). """Every distinct identifier-shaped token this input OFFERS (P8's count, P19/B3's guard).
@ -401,10 +455,60 @@ def _form_refusal(grounding: Grounding, code: str, anchored_codes: frozenset[str
) )
def _reference_refusal(grounding: Grounding, code: str) -> str | None:
"""Why a requirement-shaped ``code`` cannot be a cost line of an UNANCHORED input (P20/B).
**The order's own rule was FELLED BY MEASUREMENT before anything was built on it.** It reads:
a code is a requirement when it matches form 2 or 3 AND "står som ``req_number``/``prosessnr``
i toppnivå-frontmatter" — refuse that. Measured 15.09 against the two known positives the same
order names:
* ``10.4`` (n500, tunnel-04, ``validated``) is declared NOWHERE in n500's frontmatter. The base
declares ``seksjon: 10.4.1`` ``10.4.4`` and ``req_number: Krav 10.4.32``; the bare ``10.4``
is a section PREFIX that occurs in 12 of 274 documents and is no document's own number;
* ``1.10.4`` (r761, lindaas a4, ``validated``) is not one of r761's 2 727 ``prosessnr`` nor one
of its 2 753 ``seksjon`` values. It occurs in ONE of 2 756 documents, as prose: "iht.
vegnormal N200 Vegbygging kap. 1.10.4".
So the ordered rule fires on NEITHER of its own known positives. The COMPLEMENT does, and it is
the better-grounded rule besides: ``_ground_against_input``'s docstring already admits that this
stage "fails OPEN … on a coincidental match", and for one shape a clause number the base
hands us the vocabulary needed to close exactly that hole. A form-3 token that is NOT one of the
numbers this base declares was matched in prose by accident.
MEASURED over every code of round 3 and P17b (24 codes, 10 runs): exactly two are
requirement-shaped, they are the two known positives, and neither is in its base's vocabulary.
All five of ``contexts/kontrakt-sorasen-2027``'s REAL process codes (``12.1``, ``12.12``,
``22.1``, ``52.11``, ``51.1``) ARE declared ``prosessnr`` and pass which is what keeps the
R761 risk the order names (a process number is both a clause and a settlement post) from
turning into a wholesale refusal of the one context set built on real codes.
**The generality guard, ``_form_refusal``'s pattern:** an input that declares no reference
numbers at all cannot be answered in a vocabulary it does not have, so the rule cannot fire
there. That is what leaves every pre-P20 fixture untouched rather than exempted.
The message NAMES THE DENOMINATOR (ansikt 4, and Step 5 feeds it verbatim into the next
attempt): "not one of the 2 765 it declares" is actionable where "ungrounded" is not.
"""
if not has_requirement_form(code):
return None
vocabulary = grounding.reference_vocabulary
if not vocabulary or code in vocabulary:
return None
sample = ", ".join(sorted(vocabulary)[:3])
return (
f"is shaped like a requirement or process number, but it is not one of the "
f"{len(vocabulary)} this knowledge base declares (for example {sample}) — it was matched "
"in prose by coincidence, and an unanchored base carries no price for a clause number"
)
def _ground_against_input( def _ground_against_input(
proposal: SavingsProposal, proposal: SavingsProposal,
grounding: Grounding, grounding: Grounding,
anchored_codes: frozenset[str] = frozenset(), anchored_codes: frozenset[str] = frozenset(),
*,
anchored: bool = True,
) -> Rejection | None: ) -> Rejection | None:
"""P7: every identifier the proposal builds on must appear VERBATIM in the input it was built """P7: every identifier the proposal builds on must appear VERBATIM in the input it was built
from, or the verdict falls. from, or the verdict falls.
@ -459,6 +563,15 @@ def _ground_against_input(
shapeless = _form_refusal(grounding, item.code, anchored_codes) shapeless = _form_refusal(grounding, item.code, anchored_codes)
if shapeless is not None: if shapeless is not None:
violations.append(f"ungrounded identifier {item.code!r}: it {shapeless}") violations.append(f"ungrounded identifier {item.code!r}: it {shapeless}")
continue
# P20/B: shaped, grounded, not inert — and still a clause number the base never declared.
# UNANCHORED only: with a baseline, stage 0 has already ruled every code that reaches here
# a real line of this project, and the weaker stage must not overrule the stronger one (the
# sentence ``_form_refusal`` and ``_grounding_text`` both carry).
if not anchored:
coincidental = _reference_refusal(grounding, item.code)
if coincidental is not None:
violations.append(f"ungrounded identifier {item.code!r}: it {coincidental}")
if not violations: if not violations:
return None return None
return Rejection(proposal=proposal, reason="; ".join(violations)) return Rejection(proposal=proposal, reason="; ".join(violations))
@ -504,6 +617,11 @@ def validate_proposal(
proposal, proposal,
grounding, grounding,
frozenset() if baseline is None else frozenset(baseline.items), frozenset() if baseline is None else frozenset(baseline.items),
# P20/B: an EXPLICIT flag, never ``not anchored_codes``. A baseline with no items and
# no baseline at all are different facts, and conflating them is the very shape this
# repo refuses elsewhere (``cost_baseline_anchored`` is required without a default for
# the same reason).
anchored=baseline is not None,
) )
if adrift is not None: if adrift is not None:
return adrift return adrift

View file

@ -121,12 +121,17 @@ def test_the_correction_is_to_read_it_and_then_it_is_accepted() -> None:
answer = tools["declare_requirement"].func( answer = tools["declare_requirement"].func(
bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1" bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1"
) )
assert answer == { # P20/A1 widened the reply: the three arguments PLUS the document's own title and number and
"declared": True, # the sentence saying what the declaration binds. Asserted key by key rather than by equality,
"bundle_id": "tunnel-hauglia", # because an exact-dict assert here would fail on every future field while saying nothing about
"path": path, # the one property this arm exists for — that the declaration was ACCEPTED and RECORDED.
"ref": "Krav 12.1", assert answer["declared"] is True
} assert (answer["bundle_id"], answer["path"], answer["ref"]) == (
"tunnel-hauglia",
path,
"Krav 12.1",
)
assert set(answer) == {"declared", "bundle_id", "path", "ref", "title", "req_number", "binds"}
assert declared == [DeclaredRequirement(bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1")] assert declared == [DeclaredRequirement(bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1")]

View file

@ -0,0 +1,129 @@
"""P20 DEL C — a parse failure that does not burn the round ledger, and an announcement that
names what it is about.
**C1, what was measured (P19 F4).** ``_fetch_parsed`` retried a malformed reply with the
BYTE-IDENTICAL prompt. One round-3 run, ``kontrakt-sorasen-04``, left a
``{run_id}-parse-failures.json`` with ELEVEN rows, every one of them the same failure
(``claimed_saving_nok`` 0) eleven of the run's twelve rounds, spent re-asking a question the
model had already answered the same wrong way, because nothing ever told it what was wrong. Step
5's ``prior_rejection`` carries VALIDATOR rejections; a reply that never parsed never reaches a
validator, so no existing block could carry it.
**C2, what was measured (P17b F5).** ``--across-bundle`` takes no ``--project-id``, so the
announcement the one thing printed before the first paid call said "Run mandate for the
portfolio" for a commission dispatched across two named knowledge bases.
What each arm pins:
(a) the reason reaches the NEXT attempt's prompt, verbatim, asserted on what the client received;
(b) attempt 1 is byte-identical: a run whose first reply parses sends the pre-P20 prompt;
(c) the block is per-RETRY once a reply parses, the next attempt does not carry a stale reason;
(d) the evidence artefact is unchanged: the verbatim text is still captured (funn 1 stands);
(e) the announcement names the routed bases by their DECLARED ids;
(f) an unresolvable base falls back to its directory name rather than refusing the
``dimension_label`` precedent, so announcing never changes which error an operator sees;
(g) a single-project run and a portfolio pass announce exactly as before.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from typing import Any
from portfolio_optimiser.budget import Budget, TokenMeter
from portfolio_optimiser.generate import ParseFailure, _build_messages, generate_via_llm
from portfolio_optimiser.reference_domain import Project
from portfolio_optimiser.run import announced_subject
from portfolio_optimiser.simulation import ScriptedChatClient
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
#: The exact failure ``kontrakt-sorasen-04`` produced eleven times: a well-formed JSON object
#: whose ``claimed_saving_nok`` is 0, refused by pydantic before any validator sees it.
_UNPARSEABLE = (
'{"measure":"m","affected_items":[{"code":"C-1","quantity":1000,"unit_cost":100}],'
'"claimed_saving_nok":0}'
)
#: The grounding this loop declares: P7 requires the code to occur in the input verbatim.
_CONTEXT = "The project price schedule carries cost line C-1."
_VALID = (
'{"measure":"m","affected_items":[{"code":"C-1","quantity":1000,"unit_cost":100}],'
'"claimed_saving_nok":5000}'
)
def _project() -> Project:
return Project(id="p", name="n", description="d", currency="NOK", cost_items=(), docs_dir=".")
def _generate(replies: list[str]) -> tuple[Any, list[str], list[ParseFailure]]:
"""Drive the real loop with the repo's ONE scripted client, recording every prompt sent."""
sink: list[str] = []
failures: list[ParseFailure] = []
client = ScriptedChatClient(script=list(replies), sink=sink, role="proposer")
result = asyncio.run(
generate_via_llm(
client,
_project(),
_CONTEXT,
TokenMeter(Budget(max_tokens=100_000, max_rounds=8)),
parse_failures=failures,
)
)
return result, sink, failures
# ------------------------------------------------------------------------------- C1
def test_the_parse_reason_reaches_the_next_attempts_prompt() -> None:
"""(a) The retry is no longer blind — asserted on what the CLIENT received."""
_result, sink, _failures = _generate([_UNPARSEABLE, _VALID])
assert len(sink) >= 2, sink
assert "could not be PARSED" in sink[1]
assert "claimed_saving_nok" in sink[1]
def test_the_first_prompt_is_byte_identical_to_the_pre_p20_one() -> None:
"""(b) Without this, (a) could be satisfied by always appending the block."""
_result, sink, _failures = _generate([_VALID])
assert "could not be PARSED" not in sink[0]
assert sink[0] == _build_messages(_project(), _CONTEXT)[0].text
def test_a_reason_does_not_survive_a_reply_that_parsed() -> None:
"""(c) Per-RETRY, like ``prior_rejection`` is per-attempt: no stale instruction."""
assert "could not be PARSED" not in _build_messages(_project(), "c")[0].text
carried = _build_messages(_project(), "c", parse_error="ValueError: x")[0].text
assert "ValueError: x" in carried
def test_the_verbatim_evidence_is_still_captured() -> None:
"""(d) Fase 1b funn 1 stands: the paid reply's own text survives the retry."""
_result, _sink, failures = _generate([_UNPARSEABLE, _VALID])
assert len(failures) == 1
assert failures[0].text == _UNPARSEABLE
assert failures[0].error.startswith("ValidationError")
# ------------------------------------------------------------------------------- C2
def test_the_announcement_names_the_routed_bases() -> None:
"""(e) Two bases, two declared ids — not "the portfolio"."""
assert announced_subject(None, (str(_TUNNEL),)) == "tunnel-hauglia"
def test_an_unresolvable_base_falls_back_to_its_directory_name(tmp_path: Path) -> None:
"""(f) Naming never changes which error an operator sees (the dimension_label precedent)."""
missing = tmp_path / "not-a-base"
assert announced_subject(None, (str(missing),)) == "not-a-base"
def test_the_two_older_subjects_are_unchanged() -> None:
"""(g) A named project wins; no bases at all is still the portfolio."""
assert announced_subject("proj-1", ()) == "proj-1"
assert announced_subject("proj-1", ("ignored",)) == "proj-1"
assert announced_subject(None, ()) == "the portfolio"

View file

@ -0,0 +1,275 @@
"""P20 DEL B — a clause number is not a price, and the base's own vocabulary is what says so.
**What was measured.** Three paid rounds and one multi-base pass carried FOUR ``validated``
proposals whose cost code was a chapter number of a standard. Two survive in the recorded outboxes
and are this arm's known positives:
* ``10.4`` tunnel-hauglia round 3, base ``vegnormal-n500-2024``, ``validated``;
* ``1.10.4`` lindaas P17b, base ``vegnormal-r761-2025``, ``validated``.
Both are GROUNDED in P7's sense (they occur verbatim in the input) and neither is INERT in
P18/B1's sense (``10.4`` in 12 of 274 documents, ``1.10.4`` in 1 of 2 756). Stage 0 never ran:
no vegnormal base ships a cost baseline. Nothing in the gate could say what they are.
**THE ORDER'S OWN RULE WAS FELLED BY MEASUREMENT BEFORE ANYTHING WAS BUILT ON IT.** B1 reads: a
code is a requirement when it has form 2 or 3 AND "står som ``req_number``/``prosessnr`` i
toppnivå-frontmatter i minst ett av grunnlagets dokumenter" — refuse that. Measured 15.09:
* n500 declares ``seksjon: 10.4.1`` ``10.4.4`` and ``req_number: Krav 10.4.32``. The bare
``10.4`` is declared NOWHERE it is a section PREFIX;
* r761 declares 2 727 ``prosessnr`` and 2 753 ``seksjon``. ``1.10.4`` is NONE of them: it occurs
once, as prose, in "Krav til materialer skal være iht. vegnormal N200 Vegbygging kap. 1.10.4".
The ordered rule therefore fires on NEITHER of its own known positives. The COMPLEMENT fires on
BOTH, and it closes a hole ``_ground_against_input`` already admits in writing "it fails OPEN …
on a coincidental match". For one shape, a clause number, the base hands us the vocabulary needed
to tell a real reference from a coincidence, and that is the rule built here.
The complement is also what SPARES the one context set built on real process codes: all five of
``contexts/kontrakt-sorasen-2027``'s codes are declared ``prosessnr`` and pass. Under the ordered
rule every one of them would have been refused on an unanchored r761 run, and the set's positive
arms would have become unmeasurable the R761 risk the order names, arriving through the door it
was pointed away from.
Measured over EVERY code of round 3 and P17b (24 codes, 10 runs): exactly two are
requirement-shaped, they are the two known positives, and the replay flips exactly those two.
What each arm pins:
(a) known positive the ``10.4`` proposal, replayed against the base it actually ran on, is
``rejected``, and the reason names the denominator;
(b) known positive ditto ``1.10.4`` on r761;
(c) known negative a code the base DOES declare (sorasen's real ``12.1``) still validates;
(d) known negative the gate is OFF when the run is anchored, even for a clause-shaped code;
(e) the generality guard an input that declares no reference numbers at all cannot trip the rule,
which is what leaves every pre-P20 fixture untouched rather than exempted;
(f) K2's identifier forms are untouched: ``SHA-01`` is not requirement-shaped;
(g) ``classify_codes``' third value, and its denominator-free reading (``grounding=None``) that
``stress.py`` re-derives with;
(h) the vocabulary travels WITH the text through ``_grounding_text``, so the gate the generation
loop runs sees what the run composed;
(i) a run composes the vocabulary from the base it opened asserted end-to-end through
``run_project``, not on the composer.
"""
from __future__ import annotations
import json
import os
from pathlib import Path
import pytest
from portfolio_optimiser import okf
from portfolio_optimiser.generate import _grounding_text
from portfolio_optimiser.ir import CostBaseline, SavingsProposal
from portfolio_optimiser.reference_domain import Project
from portfolio_optimiser.validator import (
Grounding,
Rejection,
ValidatedProposal,
classify_codes,
has_requirement_form,
validate_proposal,
)
_DEFAULT_BUNDLE_ROOT = Path.home() / "repos" / "vegnormal-okf" / "build" / "ferdig"
_ROUND3 = Path("scratchpad/p19-stress/tunnel-hauglia-2027")
_P17B = Path("scratchpad/p17b-multibase/lindaas")
def _base(name: str) -> Path:
root = Path(os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", str(_DEFAULT_BUNDLE_ROOT)))
if not (root / name).is_dir():
pytest.skip(f"knowledge base {name!r} is not mounted under {root}")
return root / name
def _grounding_over(name: str) -> Grounding:
"""The delivered base exactly as ``run_project`` composes it — documents AND vocabulary."""
bundle = okf.navigate_bundle(str(_base(name)))
return Grounding(
documents=tuple(
"\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files
),
declared_references=tuple(
ref for f in bundle.context_files for ref in okf.declared_reference_numbers(f)
),
)
def _recorded(path: Path) -> SavingsProposal:
if not path.is_file():
pytest.skip(f"the recorded artefact {path} is not present in this checkout")
return SavingsProposal.model_validate(json.loads(path.read_text(encoding="utf-8"))["proposal"])
def _proposal(code: str, *, saving: float = 1000.0) -> SavingsProposal:
return SavingsProposal(
project_id="p",
measure="m",
affected_items=[{"code": code, "quantity": 10.0, "unit_cost": 1000.0}],
claimed_saving_nok=saving,
)
# ---------------------------------------------------------------------------------- known positives
def test_the_section_number_that_reached_validated_on_n500_is_refused() -> None:
"""(a) tunnel-04's ``10.4``, replayed against the base that run actually opened."""
proposal = _recorded(_ROUND3 / "tunnel-hauglia-2027-04-a4-enhetspris-ventilator-proposal.json")
assert [i.code for i in proposal.affected_items] == ["10.4"]
outcome = validate_proposal(proposal, baseline=None, grounding=_grounding_over("n500-2024"))
assert isinstance(outcome, Rejection)
assert "'10.4'" in outcome.reason
assert "not one of the 365 this knowledge base declares" in outcome.reason
def test_the_process_number_that_reached_validated_on_r761_is_refused() -> None:
"""(b) lindaas a4's ``1.10.4`` — a chapter of ANOTHER standard, quoted in one r761 document."""
proposal = _recorded(_P17B / "lindaas-01-vegnormal-r761-2025-a4-indeksregulering-proposal.json")
assert [i.code for i in proposal.affected_items] == ["1.10.4"]
outcome = validate_proposal(proposal, baseline=None, grounding=_grounding_over("r761-2025"))
assert isinstance(outcome, Rejection)
assert "not one of the 2765 this knowledge base declares" in outcome.reason
# ---------------------------------------------------------------------------------- known negatives
def test_a_process_code_the_base_declares_still_validates() -> None:
"""(c) The arm that keeps this a rule about the corpus and not about shapes.
``12.1`` is ``contexts/kontrakt-sorasen-2027``'s own first code and a REAL declared
``prosessnr`` of R761. Under the ordered rule it would have been refused; it must not be.
"""
grounding = _grounding_over("r761-2025")
assert "12.1" in grounding.reference_vocabulary
outcome = validate_proposal(_proposal("12.1"), baseline=None, grounding=grounding)
assert isinstance(outcome, ValidatedProposal), getattr(outcome, "reason", "")
def test_every_sorasen_code_is_in_the_bases_vocabulary() -> None:
"""(c) The whole context set, not one sample: five real codes, five declared numbers."""
codes = [
code
for approach in json.loads(
Path("contexts/kontrakt-sorasen-2027/mandate.json").read_text(encoding="utf-8")
)["approaches"]
for code in approach.get("affected_codes", [])
]
vocabulary = _grounding_over("r761-2025").reference_vocabulary
shaped = [c for c in codes if has_requirement_form(c)]
assert len(shaped) == 5, shaped
assert [c for c in shaped if c not in vocabulary] == []
def test_the_fasit_references_are_classified_requirement() -> None:
"""(g) known negative (c) of the order: a fasit reference IS a requirement, and says so."""
fasit = json.loads(
Path("contexts/kontrakt-sorasen-2027/fasit.json").read_text(encoding="utf-8")
)
refs = sorted({c["ref"] for entry in fasit["must_cite"] for c in entry["concepts"]})
forms = classify_codes(refs, _grounding_over("r761-2025"))
assert set(forms.values()) == {"requirement"}, forms
def test_an_anchored_run_is_untouched_by_the_rule() -> None:
"""(d) Stage 0 has already ruled; the weaker stage must not overrule the stronger."""
grounding = Grounding(documents=("12.9 is a clause",), declared_references=("12.1", "12.2"))
baseline = CostBaseline(project_id="p", items={"12.9": {"quantity": 10.0, "unit_cost": 1000.0}})
outcome = validate_proposal(_proposal("12.9"), baseline=baseline, grounding=grounding)
assert isinstance(outcome, ValidatedProposal), getattr(outcome, "reason", "")
# The control: the SAME code and the SAME text, unanchored, is refused.
unanchored = validate_proposal(_proposal("12.9"), baseline=None, grounding=grounding)
assert isinstance(unanchored, Rejection)
def test_an_input_that_declares_no_reference_numbers_cannot_trip_the_rule() -> None:
"""(e) The generality guard — and the reason every pre-P20 fixture is untouched."""
grounding = Grounding(documents=("a document mentioning 12.9 once",))
assert grounding.reference_vocabulary == frozenset()
outcome = validate_proposal(_proposal("12.9"), baseline=None, grounding=grounding)
assert isinstance(outcome, ValidatedProposal), getattr(outcome, "reason", "")
def test_a_cost_line_identifier_is_not_requirement_shaped() -> None:
"""(f) K2's 50 identifiers and this repo's own code: shape, measured."""
assert not has_requirement_form("SHA-01")
assert not has_requirement_form("ENERGI-TOTAL-EL")
assert not has_requirement_form("65 ASFALTDEKKER")
assert has_requirement_form("10.4") and has_requirement_form("Krav 4.1.2—1")
def test_classify_codes_without_a_grounding_is_the_pre_p20_answer() -> None:
"""(g) ``stress.py`` re-derives for runs written before the field existed."""
assert classify_codes(["12.1", "SHA-01", "impulsventilator"]) == {
"12.1": "identifier",
"SHA-01": "identifier",
"impulsventilator": "prose",
}
grounding = Grounding(documents=("x",), declared_references=("12.1",))
assert classify_codes(["12.1"], grounding) == {"12.1": "requirement"}
# ---------------------------------------------------------------------------------- the wiring
def test_the_vocabulary_travels_with_the_text_into_the_generation_gate() -> None:
"""(h) ``_grounding_text`` composes the run's three sources; the vocabulary must survive it."""
delivered = Grounding(documents=("d",), declared_references=("12.1",))
project = Project(
id="p", name="n", description="d", currency="NOK", cost_items=(), docs_dir="."
)
composed = _grounding_text(project, None, delivered)
assert composed.reference_vocabulary == frozenset({"12.1"})
def test_a_run_composes_the_vocabulary_from_the_base_it_opened(tmp_path: Path) -> None:
"""(i) End-to-end through ``run_project``: the stamp says ``requirement`` for a declared code.
Asserted on the ARTEFACT a run leaves, never on the composer a vocabulary wired nowhere would
satisfy every arm above and none of this one.
"""
import asyncio
from agent_framework import BaseChatClient
from portfolio_optimiser.run import run_project
from portfolio_optimiser.simulation import ScriptedChatClient
base = tmp_path / "mini"
base.mkdir()
(base / "index.md").write_text(
"---\nbundle_id: mini\n---\n\n- [Krav](krav.md) — one clause.\n", encoding="utf-8"
)
(base / "krav.md").write_text(
"---\ntype: Krav\ntitle: Krav 4.1.2-1\nprosessnr: '12.1'\n---\n\nEn kostlinje 12.1.\n",
encoding="utf-8",
)
(base / "validator-input.json").write_text(
json.dumps({"project_id": "mini-p", "measure": "m", "affected_codes": ["12.1"]}),
encoding="utf-8",
)
reply = (
'{"measure":"m","affected_items":[{"code":"12.1","quantity":10,"unit_cost":1000}],'
'"claimed_saving_nok":1000}'
)
def factory(role: str) -> BaseChatClient:
return ScriptedChatClient(
"Reasoning holds.\nVERDICT: APPROVE" if role == "checker" else reply, role=role
)
result = asyncio.run(
run_project(
"mini-p",
"local",
docs_dir=str(base),
bundle_dir=str(base),
client_factory=factory,
)
)
assert result.provenance.code_forms == {"12.1": "requirement"}
assert result.provenance.validator_decision == "validated"

View file

@ -0,0 +1,234 @@
"""P20 DEL A — the requirement that is RIGHT, and the commission's criteria reaching the reader.
**The measured silence, three rounds and one multi-base pass deep.** P19 DEL A made a direction
NAME the requirement that binds it and made the run refuse a declaration naming a path it never
opened. The declarations then happened and MEASURED (P19 round 3: 9 declarations over 5 paid
runs; P17b: 4 over one two-base pass) **not one of them named a fasit concept**. The rung works;
what nothing asked for was that the requirement be the RIGHT one. Two halves were missing:
* the tool answered ``{"declared": true, ...}`` to every accepted declaration, echoing back the
caller's own three arguments. A model that had declared a requirement about something else was
told, in the only words it got, that it had succeeded;
* the commission's own statement of what a good answer looks like — ``Mandate.success_criteria`` —
reached ``announce`` and NOTHING else (P19 F2). It was printed for a human and withheld from the
only reader who could act on it.
What each arm pins:
(a) A1 the reply carries the DOCUMENT's own ``title`` and ``req_number``, read off the base, plus
the sentence that says what the declaration BINDS. Without the document's own words the reply
cannot contradict a wrong declaration, which is the whole correction;
(b) A1 a path the base carries as no navigated concept (``index.md`` is the reachable case)
answers with empty strings rather than raising: the read trace has already accepted the
declaration, and turning "I cannot restate your title" into a refusal would fail a declaration
the run's own evidence proves was read;
(c) A1 the reply is read off ``Bundle.context_files``, so a ``type: verdict`` document can never
be named back by title. The one layer no listing names stays unnamed even in an answer;
(d) A2 ``criteria_block`` is the ONE renderer, and it is EMPTY without criteria. That omission is
what keeps every prompt of every un-commissioned run byte-identical, the demo's included;
(e) A2 the criteria reach the DEBATE's task message, the prompt where ``declare_requirement`` is
available. Asserted on the text the client actually received, never on the composer;
(f) A2 a run with NO mandate sends the task message byte-identically to before, which is the
half that keeps ``demo-transcript.stdout`` unchanged and is asserted here rather than left to
the golden;
(g) A3 the navigator instruction and the tool description both name the ``read_dir(filter=...)``
route with a worked example. A description that lies about the body IS the model's instruction
(the Fase-3 class), and here the body is new.
"""
from __future__ import annotations
from pathlib import Path
from typing import Any
import pytest
from portfolio_optimiser import explore, okf
from portfolio_optimiser.explore import DeclaredRequirement, ToolCall, navigator_tools
from portfolio_optimiser.mandate import Approach, Mandate, criteria_block
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
_BYGG = _EXAMPLES / "bygg-energi-mikro"
_PID = "BYGG-KONTOR-NORD"
def _wired(bundle_dir: Path) -> tuple[dict[str, Any], list[ToolCall], list[DeclaredRequirement]]:
opened: list[ToolCall] = []
declared: list[DeclaredRequirement] = []
tools = navigator_tools((str(bundle_dir),), opened=opened, requirements=declared)
return {t.name: t for t in tools}, opened, declared
def _declare(tools: dict[str, Any], opened: list[ToolCall], base: str, path: str) -> dict[str, Any]:
"""Read it the way a run does, then declare it — the (b) path of P19 DEL A."""
opened.append(ToolCall(name="read_file", bundle_id=base, path=path))
return tools["declare_requirement"].func(bundle_id=base, path=path, ref="Krav 4.1.2-1")
def _base_with_a_requirement(root: Path) -> Path:
"""A base declaring its own ``req_number`` — the form the delivered N corpora carry.
Crafted rather than taken from ``shared/examples``: MEASURED, no example base declares a
``req_number`` at all, so an arm built on one of them could not tell a reply that reads the
document from one that returns an empty string.
"""
base = root / "n-mini"
base.mkdir()
(base / "index.md").write_text(
"---\nbundle_id: n-mini\n---\n\n- [Krav](krav.md) — one requirement.\n", encoding="utf-8"
)
(base / "krav.md").write_text(
"---\ntype: Krav\ntitle: Krav 4.1.2-1 Rundkjoring\nreq_number: Krav 4.1.2-1\n"
"seksjon: '4.1.2'\n---\n\nEn rundkjoring skal ha …\n",
encoding="utf-8",
)
return base
# ---------------------------------------------------------------------------------------------
# (a)-(c) the reply the model can be contradicted by
# ---------------------------------------------------------------------------------------------
def test_the_reply_carries_the_documents_own_title_and_number(tmp_path: Path) -> None:
"""(a) The base's words, not the caller's — the only thing that can say "wrong one"."""
base = _base_with_a_requirement(tmp_path)
tools, opened, declared = _wired(base)
answer = _declare(tools, opened, "n-mini", "krav.md")
assert answer["title"] == "Krav 4.1.2-1 Rundkjoring"
assert answer["req_number"] == "Krav 4.1.2-1"
assert "is the requirement the proposal rests on" in answer["binds"]
assert declared == [DeclaredRequirement(bundle_id="n-mini", path="krav.md", ref="Krav 4.1.2-1")]
def test_a_path_that_is_no_concept_answers_with_empty_strings(tmp_path: Path) -> None:
"""(b) Accepted-but-unrestatable is not a refusal; the read trace already ruled."""
base = _base_with_a_requirement(tmp_path)
tools, opened, _declared = _wired(base)
answer = _declare(tools, opened, "n-mini", "index.md")
assert answer["declared"] is True
assert (answer["title"], answer["req_number"]) == ("", "")
def test_a_verdict_document_is_never_named_back_by_title(tmp_path: Path) -> None:
"""(c) Read off ``context_files``, the property that drops the ``type: verdict`` layer."""
base = _base_with_a_requirement(tmp_path)
(base / "dom.md").write_text(
"---\ntype: verdict\ntitle: A PRIOR EXPERT JUDGEMENT\nid: v1\n---\n\napproved.\n",
encoding="utf-8",
)
tools, opened, _declared = _wired(base)
answer = _declare(tools, opened, "n-mini", "dom.md")
assert answer["declared"] is True
assert answer["title"] == "", "the verdict layer must not be readable through this answer"
# The control: the SAME base answers the concept's title, so the empty string above is the
# layer being excluded rather than the reader being broken.
assert _declare(tools, opened, "n-mini", "krav.md")["title"] == "Krav 4.1.2-1 Rundkjoring"
def test_the_tunnel_example_answers_its_own_title_and_no_number() -> None:
"""(a) control on a REAL example base: title present, reference number honestly absent."""
tools, opened, _declared = _wired(_TUNNEL)
path = okf.navigate_bundle(str(_TUNNEL)).context_files[0].name
answer = _declare(tools, opened, "tunnel-hauglia", path)
assert answer["title"].startswith("Kilder: tunnelbelysning")
assert answer["req_number"] == "", "this base declares none, and the reply must not invent one"
# ---------------------------------------------------------------------------------------------
# (d)-(f) the criteria reaching the prompt
# ---------------------------------------------------------------------------------------------
def test_the_criteria_renderer_is_empty_without_criteria() -> None:
"""(d) Omission, never an empty heading — the ``announce`` rule, and the golden's guarantee."""
assert criteria_block("") == ""
rendered = criteria_block("a saving the price schedule can carry")
assert "a saving the price schedule can carry" in rendered
assert rendered.startswith("\n")
def test_the_commissions_criteria_reach_the_debate_task(monkeypatch: pytest.MonkeyPatch) -> None:
"""(e) Asserted on what the CLIENT received, never on the composer."""
sent = _run_and_capture_task(
monkeypatch,
mandate=Mandate(
objective="cut cost",
approaches=(Approach(id="a1", label="L", description="D"),),
success_criteria="SENTINEL-CRITERION-ONLY-THE-COMMISSION-CARRIES",
),
)
assert "SENTINEL-CRITERION-ONLY-THE-COMMISSION-CARRIES" in sent[0]
assert "What the commissioner counts as success" in sent[0]
def test_a_run_without_a_mandate_sends_the_task_unchanged(monkeypatch: pytest.MonkeyPatch) -> None:
"""(f) The byte-identity half. Without it (e) could be satisfied by always appending."""
sent = _run_and_capture_task(monkeypatch, mandate=None)
assert sent[0].startswith(f"Find a cost-saving measure for {_PID}.\nContext:\n")
def _run_and_capture_task(monkeypatch: pytest.MonkeyPatch, *, mandate: Mandate | None) -> list[str]:
"""Drive a real ``run_project`` and capture every prompt the debate's clients received.
The same recording-factory seam ``test_debate_navigation_cost_loadbearing`` measures S2c with:
the assertion is on what the CLIENT was handed, never on ``criteria_block``'s return value, so
a renderer wired nowhere cannot satisfy it.
"""
import asyncio
from agent_framework import BaseChatClient
from portfolio_optimiser.run import run_project
from portfolio_optimiser.simulation import ScriptedChatClient
sink: list[str] = []
valid = (
'{"measure":"LED-retrofit","affected_items":'
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],'
'"claimed_saving_nok":30000}'
)
def factory(role: str) -> BaseChatClient:
client = ScriptedChatClient(
"Reasoning holds.\nVERDICT: APPROVE" if role == "checker" else valid, role=role
)
original = client._inner_get_response # type: ignore[attr-defined]
def recording(*, messages, options, stream=False, **kwargs): # type: ignore[no-untyped-def]
sink.append("\n".join(getattr(m, "text", "") or "" for m in messages))
return original(messages=messages, options=options, stream=stream, **kwargs)
client._inner_get_response = recording # type: ignore[attr-defined,method-assign]
return client
asyncio.run(
run_project(
_PID,
"local",
docs_dir=str(_BYGG),
bundle_dir=str(_BYGG),
verdict_input={"decision": "approved", "rationale": "expert reviewed (test)"},
client_factory=factory,
mandate=mandate,
)
)
assert sink, "the debate never reached the client"
return sink
# ---------------------------------------------------------------------------------------------
# (g) the instruction that describes the new body
# ---------------------------------------------------------------------------------------------
def test_both_descriptions_name_the_filter_route() -> None:
"""(g) A description that lies about the body IS the instruction (the Fase-3 class)."""
instruction = explore._INSTRUCTIONS[explore.HYPOTHESISER_ROLE]
assert "read_dir a 'filter' word" in instruction
assert "Krav 4.1.2-1" in instruction, "the worked example is what makes the route concrete"
tools = {t.name: t for t in navigator_tools((str(_TUNNEL),), opened=[], requirements=[])}
description = tools["declare_requirement"].description or ""
assert "filter='rundkjoring'" in description
assert "own title and number" in description