feat(p20): the requirement that is RIGHT, and a clause number that is not a price
Three seams, one commit: A, B and C touch the same four modules (run.py carries
the debate task, the grounding composition and the announcement; okf.py carries
one reference-number vocabulary read by both A and B), so splitting them into
three commits would have meant hunk-level staging of entangled files. Stated
rather than silently restructured.
A — the declaration answers with the DOCUMENT's own words. Measured: 13
declarations over round 3 and P17b, not one naming a fasit concept, while the
tool answered {"declared": true, ...} by echoing the caller's own arguments. It
now returns the document's title and req_number, read off Bundle.context_files
(so the type: verdict layer can never be named back), plus the sentence saying
what the declaration binds. A path the base carries as no concept answers with
empty strings rather than refusing. The commission's success_criteria now reach
the DEBATE task through mandate.criteria_block, the one renderer, empty when
there are none — which is what keeps every un-commissioned prompt, and the
golden, byte-identical.
B — a clause number is not a price. THE ORDER'S OWN RULE WAS FELLED BY
MEASUREMENT: it asks to refuse a code that IS declared req_number/prosessnr,
and neither of its two known positives is. n500 declares seksjon 10.4.1..10.4.4
but never the bare 10.4; r761 declares 2727 prosessnr and 2753 seksjon, none of
them 1.10.4, which occurs once, as prose ("iht. vegnormal N200 kap. 1.10.4").
The COMPLEMENT fires on both and closes the hole _ground_against_input already
admits in writing -- "it fails OPEN on a coincidental match". Unanchored run +
requirement-shaped code + the base declares a vocabulary + the code is not in
it -> refused, naming the denominator. All five of kontrakt-sorasen's real
process codes ARE declared and pass, which is what keeps the one context set
built on real codes measurable. Replayed over all 24 codes of round 3 + P17b:
exactly the two known positives flip validated -> rejected, 22 unchanged.
C — a parse failure no longer burns the round ledger blind. _fetch_parsed takes
a BUILDER instead of a finished message list, so the retry carries the parse
reason; measured, kontrakt-sorasen-04 spent 11 of 12 rounds re-asking the same
question. And announced_subject names the routed bases instead of saying "the
portfolio" for a two-base commission.
Suite 1807/5 (from 1781, +26, 0 removed), golden demo-transcript.stdout
BYTE-UNCHANGED (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f),
ruff and mypy clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
64723c5d89
commit
c8f0c8f7c4
12 changed files with 1098 additions and 37 deletions
72
CLAUDE.md
72
CLAUDE.md
|
|
@ -2769,6 +2769,78 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
|
||||||
flaten er BEVISST urørt (feltet er i ingen av hostings tre sett); og `--proposal-review` trås
|
flaten er BEVISST urørt (feltet er i ingen av hostings tre sett); og `--proposal-review` trås
|
||||||
gjennom til ÉN terminal delt av alle basene (motorens egen begrunnelse — dispatchen er
|
gjennom til ÉN terminal delt av alle basene (motorens egen begrunnelse — dispatchen er
|
||||||
sekvensiell).
|
sekvensiell).
|
||||||
|
- **Et KRAVNUMMER er ikke en PRIS, og basens EGEN nummer-ordliste er dét som sier det — ordrens
|
||||||
|
regel ble FELT av målingen før noe ble bygget på den (P20 DEL B, 15.09):** tre stressrunder og
|
||||||
|
ett multi-base-pass bar `validated` forslag hvis kostkode var et kapittelnummer i en vegnormal.
|
||||||
|
To overlever i utboksene og er kjent-positivene: **`10.4`** (tunnel-hauglia runde 3, n500) og
|
||||||
|
**`1.10.4`** (lindaas P17b, r761). Begge er GRUNNET i P7s forstand og ingen er INERT i P18/B1s
|
||||||
|
(`10.4` i 12 av 274 dokumenter, `1.10.4` i **1 av 2 756**); stadium 0 kjørte aldri, fordi ingen
|
||||||
|
vegnormal-base bærer en kostbaseline. **Ordrens B1 sier: form 2/3 OG «står som
|
||||||
|
`req_number`/`prosessnr` i toppnivå-frontmatter» → nekt.** MÅLT 15.09: n500 erklærer
|
||||||
|
`seksjon: 10.4.1`…`10.4.4` og `req_number: Krav 10.4.3—2`, men **aldri den bare `10.4`** — den er
|
||||||
|
et seksjons-PREFIKS; og r761 erklærer 2 727 `prosessnr` + 2 753 `seksjon`, hvorav **ingen** er
|
||||||
|
`1.10.4`, som står ÉN gang, som prosa: «iht. vegnormal N200 Vegbygging kap. 1.10.4». **Den
|
||||||
|
ordrede regelen fyrer altså på INGEN av sine egne kjent-positive.** KOMPLEMENTET fyrer på BEGGE,
|
||||||
|
og lukker et hull `_ground_against_input`s egen docstring alt innrømmer skriftlig («it fails OPEN
|
||||||
|
… on a coincidental match»): for ÉN form — et klausulnummer — gir basen oss ordlista som skiller
|
||||||
|
en ekte referanse fra et sammentreff. Regelen er derfor: **UFORANKRET + kravformet kode + basen
|
||||||
|
erklærer en ordliste + koden er IKKE i den → nekt, med nevner.** **Komplementet er også dét som
|
||||||
|
SPARER det ene kontekstsettet bygget på ekte prosesskoder:** alle fem kodene i
|
||||||
|
`contexts/kontrakt-sorasen-2027` (`12.1`, `12.12`, `22.1`, `52.11`, `51.1`) ER erklærte
|
||||||
|
`prosessnr` og passerer — under den ordrede regelen ville hver av dem blitt nektet på en
|
||||||
|
uforankret r761-kjøring, og settets positive armer blitt umålbare; dét er R761-risikoen ordren
|
||||||
|
selv navngir, ankommet gjennom døra den ble pekt bort fra. **MÅLT over ALLE 24 koder i runde 3 +
|
||||||
|
P17b (10 kjøringer):** nøyaktig to er kravformede, de er de to kjent-positive, og offline replay
|
||||||
|
flipper nøyaktig de to (`validated → rejected`) mens 22 står uendret. **Generalitetsvernet er
|
||||||
|
`_form_refusal`s mønster:** en input som erklærer INGEN referansenumre kan ikke besvares i en
|
||||||
|
ordliste den ikke har, så regelen kan ikke fyre der — dét er hva som lar hver pre-P20-fixtur stå
|
||||||
|
URØRT i stedet for UNNTATT. **`anchored` er et EKSPLISITT flagg, aldri `not anchored_codes`:** en
|
||||||
|
baseline uten linjer og ingen baseline er ulike fakta (`cost_baseline_anchored`s grunn).
|
||||||
|
**Ordlista reiser MED teksten** (`Grounding.declared_references`, DEFAULTET —
|
||||||
|
`skipped_links`-halvdelen: en tom ordliste er et ærlig POSITIVT utsagn), komponert i SAMME vandring
|
||||||
|
som dokumentene i `run.py` og propagert gjennom `_grounding_text`; to avledninger av én bases
|
||||||
|
ordliste ville stått fritt til å være uenige (kø-(p)). **`okf.REFERENCE_NUMBER_FIELDS` er ÉN
|
||||||
|
kilde**, derivert fra `_FILTER_FIELDS` pluss `seksjon` — og BEVISST ikke lagt til i
|
||||||
|
`_FILTER_FIELDS` selv, fordi hva et `filter`-ord søker i er målt og gatet (P18) og en utvidelse
|
||||||
|
ville endret `total_matches` for hvert navigatørkall. `code_forms` får en TREDJE verdi,
|
||||||
|
`requirement`, for en kode basen FAKTISK erklærer (ordrens kjent-negative (c): alle
|
||||||
|
`must_cite`-referanser klassifiseres slik) — rapport, aldri gaten, som fyrer på komplementet.
|
||||||
|
Load-bearing MÅLT (`tests/test_requirement_number_gate_loadbearing.py`, 11 armer).
|
||||||
|
- **Erklæringen svarer med DOKUMENTETS egne ord, og kommisjonens kriterier når leseren som kan
|
||||||
|
handle på dem (P20 DEL A, 15.09):** P19 DEL A gjorde at en retning MÅ navngi kravet som binder
|
||||||
|
den, og rungen virker — MÅLT ga runde 3 ni erklæringer over fem betalte kjøringer og P17b fire
|
||||||
|
over ett pass, og **ikke én av de 13 navnga et fasit-konsept**. Det ingen ba om var at kravet
|
||||||
|
skulle være det RIKTIGE. To halvdeler manglet: (1) verktøyet svarte `{"declared": true, …}` ved å
|
||||||
|
ekko kallerens egne tre argumenter, så en modell som hadde erklært feil krav ble fortalt, med de
|
||||||
|
eneste ordene den fikk, at den hadde lyktes; (2) `Mandate.success_criteria` nådde `announce` og
|
||||||
|
INGENTING ellers (P19 F2) — skrevet ut for et menneske og holdt tilbake fra den ene leseren som
|
||||||
|
kunne handle på det. Svaret bærer nå dokumentets EGNE `title`/`req_number`, lest av basen gjennom
|
||||||
|
`okf.reference_number` (ÉN leser av «hvilket krav er dette», kø-(p)), pluss `binds`-setningen.
|
||||||
|
**Lest av `Bundle.context_files`, aldri `files`:** en `type: verdict`-fil kan dermed ikke navngis
|
||||||
|
tilbake ved tittel — det ene laget ingen listing nevner og `read_file` nekter blankt. **En sti
|
||||||
|
basen ikke bærer som konsept (`index.md`) svarer med TOMME strenger, aldri en nekt:** lese-sporet
|
||||||
|
har alt godkjent erklæringen, og å gjøre «jeg kan ikke gjengi tittelen din» om til en nekt ville
|
||||||
|
felt en erklæring kjøringens eget bevis viser ble lest. `mandate.criteria_block` er ENESTE
|
||||||
|
renderer og TOM uten kriterier — omisjon, aldri en tom overskrift (`announce`-regelen), og dét er
|
||||||
|
hva som holder hver prompt i hver ukommisjonert kjøring byte-identisk, demoens inkludert.
|
||||||
|
**A2s plassering er MÅLT:** på utforskningsstien finnes ingenting å bære — `main()` sender
|
||||||
|
`explore()` ingen `success_criteria` i det hele tatt, så objektivet ER prompten og når oppgaven
|
||||||
|
alt. Load-bearing MÅLT (`tests/test_right_requirement_loadbearing.py`, 8 armer).
|
||||||
|
- **En parse-feil brenner ikke lenger rundeboka i stillhet, og annonseringen navngir hva den
|
||||||
|
handler om (P20 DEL C, 15.09):** `_fetch_parsed` prøvde på nytt med den BYTE-IDENTISKE prompten.
|
||||||
|
MÅLT (P19 F4): `kontrakt-sorasen-04` etterlot `{run_id}-parse-failures.json` med **elleve** rader,
|
||||||
|
hver av dem samme feil (`claimed_saving_nok` ≤ 0) — elleve av kjøringens tolv runder, brukt på å
|
||||||
|
spørre om igjen uten å si hva som var galt. Steg 5s `prior_rejection` bærer VALIDATOR-avvisninger,
|
||||||
|
og et svar som aldri parset når aldri en validator, så ingen eksisterende blokk kunne bære det.
|
||||||
|
`_fetch_parsed` tar nå en BYGGER i stedet for en ferdig meldingsliste — kallerens egen
|
||||||
|
`_build_messages`-binding, så retry-løkka ikke kan komponere en prompt den ytre løkka ikke ville
|
||||||
|
komponert (kø-(p)) — og grunnen er per RETRY, tom på forsøk 1, så attempt 1 er byte-identisk. Den
|
||||||
|
verbatime fangsten (Fase 1b funn 1) står URØRT. **C2:** `--across-bundle` tar ingen
|
||||||
|
`--project-id`, så annonseringen sa «Run mandate for the portfolio» om en kommisjon dispatchet
|
||||||
|
over to navngitte baser (P17b F5); `announced_subject` navngir de rutede basene ved ERKLÆRT id, og
|
||||||
|
en base som ikke lar seg løse faller tilbake til katalognavnet — `dimension_label`-presedensen
|
||||||
|
ordrett: å annonsere skal ALDRI endre hvilken feil en operatør ser. Load-bearing MÅLT
|
||||||
|
(`tests/test_parse_error_feedback_loadbearing.py`, 7 armer).
|
||||||
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
||||||
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
||||||
|
|
||||||
|
|
|
||||||
29
README.md
29
README.md
|
|
@ -481,6 +481,17 @@ when the seam is detached, so the loop cannot silently degrade into theater.
|
||||||
that does not use identifiers cannot be answered in them. A code the project's cost baseline
|
that does not use identifiers cannot be answered in them. A code the project's cost baseline
|
||||||
carries is never refused for its shape, because stage 0 has already ruled it a real line.
|
carries is never refused for its shape, because stage 0 has already ruled it a real line.
|
||||||
|
|
||||||
|
**A clause number is not a price.** A standard numbers its own requirements and a process code
|
||||||
|
numbers its own settlement posts, and both look exactly like a bare decimal: `10.4`, `1.10.4`,
|
||||||
|
`12.1`. When a run is **unanchored** — the knowledge base ships no cost baseline, so the
|
||||||
|
validator's stage 0 never runs — a cost code shaped like a clause number is refused unless it is
|
||||||
|
one of the reference numbers the base itself declares (`req_number`, `prosessnr`, `seksjon` in a
|
||||||
|
document's own frontmatter). The refusal names the denominator: *"not one of the 2765 this
|
||||||
|
knowledge base declares"*. A code the base does declare still validates, so a project priced in
|
||||||
|
real process codes is untouched; an input that declares no reference numbers at all cannot trip
|
||||||
|
the rule at all. `code_forms` reports the third value, `requirement`, for a code the base does
|
||||||
|
declare. Anchored runs are unaffected — stage 0 is the stronger falsifier and keeps the ruling.
|
||||||
|
|
||||||
**Naming the requirement that binds a direction.** Whoever navigates a knowledge base — the
|
**Naming the requirement that binds a direction.** Whoever navigates a knowledge base — the
|
||||||
exploration's hypothesiser and, since the debate started navigating, the proposer — is asked to
|
exploration's hypothesiser and, since the debate started navigating, the proposer — is asked to
|
||||||
name the ONE requirement that binds the direction it commits to, and to declare it with
|
name the ONE requirement that binds the direction it commits to, and to declare it with
|
||||||
|
|
@ -491,7 +502,23 @@ when the seam is detached, so the loop cannot silently degrade into theater.
|
||||||
`"why_none"` — a base that holds no requirement for a direction is a finding worth stating, and
|
`"why_none"` — a base that holds no requirement for a direction is a finding worth stating, and
|
||||||
the field is never simply omitted. Where an approach carries one, the proposer's prompt names it
|
the field is never simply omitted. Where an approach carries one, the proposer's prompt names it
|
||||||
and asks for it back verbatim in `measure`, and the declaration is written to
|
and asks for it back verbatim in `measure`, and the declaration is written to
|
||||||
`{run_id}-debate.json` / `{run_id}-exploration.json` under `requirements`.
|
`{run_id}-debate.json` / `{run_id}-exploration.json` under `requirements`. The reply gives back
|
||||||
|
**the document's own** `title` and `req_number`, read off the base rather than echoed from the
|
||||||
|
arguments, plus the sentence that says what the declaration binds — so a declaration of the wrong
|
||||||
|
requirement can be seen to be wrong. The route is named in the instruction as well as the
|
||||||
|
answer: pass `read_dir` a `filter` word from the approach's own label.
|
||||||
|
|
||||||
|
**What the commissioner counts as success reaches the readers.** A mandate's `success_criteria`
|
||||||
|
used to reach the announcement and nothing else. It is now restated verbatim in the debate's task
|
||||||
|
message — the prompt where `declare_requirement` is available — through one renderer, and is
|
||||||
|
omitted entirely when the commission states none.
|
||||||
|
|
||||||
|
**A reply that could not be parsed says why, once.** A malformed reply used to be retried with
|
||||||
|
the byte-identical prompt; measured, one paid run spent eleven of its twelve rounds re-asking a
|
||||||
|
question it had already answered the same wrong way. The next attempt's prompt now carries the
|
||||||
|
parse reason — only the reason, never the discarded JSON — as a block of its own, distinct from
|
||||||
|
a validator rejection (the numbers were refuted) and from expert feedback (a person objected).
|
||||||
|
The verbatim reply is still captured to `{run_id}-parse-failures.json` as before.
|
||||||
|
|
||||||
**Answering the plan review (`--plan-review`).** With `enable_plan_review` set, the exploration
|
**Answering the plan review (`--plan-review`).** With `enable_plan_review` set, the exploration
|
||||||
stops before the loop is allowed to run and asks you to sign the plan off. `--plan-review`
|
stops before the loop is allowed to run and asks you to sign the plan off. `--plan-review`
|
||||||
|
|
|
||||||
|
|
@ -211,9 +211,13 @@ _INSTRUCTIONS: Final = {
|
||||||
HYPOTHESISER_ROLE: (
|
HYPOTHESISER_ROLE: (
|
||||||
"You shape ONE candidate cost-saving direction at a time from what the navigator found. "
|
"You shape ONE candidate cost-saving direction at a time from what the navigator found. "
|
||||||
"BEFORE you commit to a direction, name the ONE requirement in the knowledge base that "
|
"BEFORE you commit to a direction, name the ONE requirement in the knowledge base that "
|
||||||
"BINDS it: have the navigator find it with read_dir(filter=...) and read it with "
|
"BINDS it: pass read_dir a 'filter' word taken from the approach's own label — "
|
||||||
"read_file, then call declare_requirement with the base id, that path and the "
|
"filter='rundkjoring' finds the level's requirements about roundabouts, and one of them is "
|
||||||
"requirement's own number. A direction with no requirement behind it is a guess. "
|
"the 'Krav 4.1.2-1' you are looking for — read it with read_file, then call "
|
||||||
|
"declare_requirement with the base id, that path and the requirement's own number. The "
|
||||||
|
"reply gives back the document's own title and number: if they are not about your measure, "
|
||||||
|
"you declared the wrong requirement and should filter again. A direction with no "
|
||||||
|
"requirement behind it is a guess. "
|
||||||
"You may call quick_validate to sanity-check a candidate's numbers; its verdict is "
|
"You may call quick_validate to sanity-check a candidate's numbers; its verdict is "
|
||||||
"ADVISORY and is not the project's decision. When you commit to a direction, end your "
|
"ADVISORY and is not the project's decision. When you commit to a direction, end your "
|
||||||
f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", '
|
f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", '
|
||||||
|
|
@ -1001,6 +1005,26 @@ def _resolve_bundle(index: Mapping[str, str], bundle_id: str) -> str:
|
||||||
return index[bundle_id]
|
return index[bundle_id]
|
||||||
|
|
||||||
|
|
||||||
|
def _declared_document(index: Mapping[str, str], bundle_id: str, path: str) -> tuple[str, str]:
|
||||||
|
"""``(title, reference_number)`` of the document at ``path``, or ``("", "")`` when the base has
|
||||||
|
no navigated concept under that name.
|
||||||
|
|
||||||
|
Read off the SAME ``Bundle.context_files`` every listing rung is built from (MAJOR-3/S7a-3), so
|
||||||
|
a declaration can never be answered with the title of a ``type: verdict`` document — the one
|
||||||
|
layer no listing names and ``read_file`` refuses outright.
|
||||||
|
|
||||||
|
A path the base does not carry as a concept — ``index.md`` is the reachable case — answers with
|
||||||
|
two empty strings rather than raising: the declaration itself has already been accepted by the
|
||||||
|
read-trace check above, and turning "I cannot restate your title" into a refusal would fail a
|
||||||
|
declaration the run's own trace proves was read.
|
||||||
|
"""
|
||||||
|
bundle = okf.navigate_bundle(_resolve_bundle(index, bundle_id))
|
||||||
|
for file in bundle.context_files:
|
||||||
|
if file.name == path:
|
||||||
|
return okf.unquote_scalar(file.frontmatter.get("title", "")), okf.reference_number(file)
|
||||||
|
return "", ""
|
||||||
|
|
||||||
|
|
||||||
#: Characters of the root index body one catalogue entry may carry. The catalogue's job is to let a
|
#: Characters of the root index body one catalogue entry may carry. The catalogue's job is to let a
|
||||||
#: manager pick a base, not to read one, so the excerpt is a fixed-size window rather than a share
|
#: manager pick a base, not to read one, so the excerpt is a fixed-size window rather than a share
|
||||||
#: of the base: cost then scales with how many bases are configured, which the operator chose, and
|
#: of the base: cost then scales with how many bases are configured, which the operator chose, and
|
||||||
|
|
@ -1364,7 +1388,11 @@ def navigator_tools(
|
||||||
"requirement's own number as its frontmatter states it. You must have READ the "
|
"requirement's own number as its frontmatter states it. You must have READ the "
|
||||||
"document with read_file first: a declaration naming a path this run never opened is "
|
"document with read_file first: a declaration naming a path this run never opened is "
|
||||||
"refused, and reading it is the correction. Use read_dir with a 'filter' word to find "
|
"refused, and reading it is the correction. Use read_dir with a 'filter' word to find "
|
||||||
"it, read_file to read it, then declare it."
|
"it (a word from the approach's own label works: filter='rundkjoring' -> "
|
||||||
|
"'Krav 4.1.2-1'), read_file to read it, then declare it. The reply gives back the "
|
||||||
|
"document's own "
|
||||||
|
"title and number, so you can see whether you declared the requirement you meant: a "
|
||||||
|
"declaration of a requirement that is not about the measure is worth nothing."
|
||||||
),
|
),
|
||||||
)
|
)
|
||||||
def declare_requirement(bundle_id: str, path: str, ref: str) -> dict[str, Any]:
|
def declare_requirement(bundle_id: str, path: str, ref: str) -> dict[str, Any]:
|
||||||
|
|
@ -1386,7 +1414,26 @@ def navigator_tools(
|
||||||
"direction"
|
"direction"
|
||||||
)
|
)
|
||||||
requirements.append(DeclaredRequirement(bundle_id=bundle_id, path=path, ref=ref))
|
requirements.append(DeclaredRequirement(bundle_id=bundle_id, path=path, ref=ref))
|
||||||
return {"declared": True, "bundle_id": bundle_id, "path": path, "ref": ref}
|
# P20/A1: give back the DOCUMENT's own title and number, read off the base rather than
|
||||||
|
# echoed from the arguments. MEASURED (P19 round 3, P17b): 13 declarations over 5 runs and
|
||||||
|
# NOT ONE named a fasit concept — the tool answered ``{"declared": true, ...}`` to every
|
||||||
|
# declaration, so a model that had declared the wrong requirement was told it had succeeded.
|
||||||
|
# ``okf.reference_number`` is the ONE reader of "which requirement is this" (kø-(p)), and
|
||||||
|
# ``binds`` says out loud what the declaration is for: without it the reply is data with no
|
||||||
|
# instruction, and the instruction is the whole correction.
|
||||||
|
declared = _declared_document(index, bundle_id, path)
|
||||||
|
return {
|
||||||
|
"declared": True,
|
||||||
|
"bundle_id": bundle_id,
|
||||||
|
"path": path,
|
||||||
|
"ref": ref,
|
||||||
|
"title": declared[0],
|
||||||
|
"req_number": declared[1],
|
||||||
|
"binds": (
|
||||||
|
f"This declaration says {ref} is the requirement the proposal rests on; a "
|
||||||
|
"declaration of a requirement that is not about the measure is worth nothing."
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
tools = [list_bundles, read_bundle, read_dir, read_file]
|
tools = [list_bundles, read_bundle, read_dir, read_file]
|
||||||
if requirements is not None:
|
if requirements is not None:
|
||||||
|
|
|
||||||
|
|
@ -286,6 +286,7 @@ def _build_messages(
|
||||||
*,
|
*,
|
||||||
approach: Approach | None = None,
|
approach: Approach | None = None,
|
||||||
prior_feedback: str | None = None,
|
prior_feedback: str | None = None,
|
||||||
|
parse_error: str | None = None,
|
||||||
) -> list[Message]:
|
) -> list[Message]:
|
||||||
"""Build the hypothesis prompt. When ``prior_rejection`` is set (Step 5, målbilde §5/§7),
|
"""Build the hypothesis prompt. When ``prior_rejection`` is set (Step 5, målbilde §5/§7),
|
||||||
append a revision block carrying ONLY the falsification *reason* verbatim — never the prior
|
append a revision block carrying ONLY the falsification *reason* verbatim — never the prior
|
||||||
|
|
@ -308,8 +309,17 @@ def _build_messages(
|
||||||
"your numbers were refuted"; conflating the two would tell the model the machine objected
|
"your numbers were refuted"; conflating the two would tell the model the machine objected
|
||||||
when a person did. ``None`` -> the byte-identical base prompt, like the other two.
|
when a person did. ``None`` -> the byte-identical base prompt, like the other two.
|
||||||
|
|
||||||
All three are composable, and the ORDER is fixed: base -> approach head -> rejection ->
|
When ``parse_error`` is set (P20/C1) a FOURTH block carries the reason the PREVIOUS reply
|
||||||
feedback. A prompt can legitimately carry a rejection AND a feedback at once — that is the
|
could not be parsed — the same "only the reason, never the JSON" rule the other two follow.
|
||||||
|
It is a different instruction from both: a rejection means the numbers were refuted and a
|
||||||
|
feedback means a person objected, while this one means nothing was ever read. MEASURED (P19
|
||||||
|
F4): ``_fetch_parsed`` retried with the byte-identical prompt, and one round-3 run
|
||||||
|
(``kontrakt-sorasen-04``) spent ELEVEN of its twelve rounds on replies that all failed the
|
||||||
|
same way — ``claimed_saving_nok: 0`` — because nothing ever told the model what was wrong.
|
||||||
|
``None`` -> the byte-identical base prompt, like the other three.
|
||||||
|
|
||||||
|
All four are composable, and the ORDER is fixed: base -> approach head -> rejection ->
|
||||||
|
feedback -> parse error. A prompt can legitimately carry a rejection AND a feedback at once — that is the
|
||||||
attempt after a revise whose bought attempt the validator then rejected: the human's
|
attempt after a revise whose bought attempt the validator then rejected: the human's
|
||||||
instruction STANDS until the human next answers, while the machine's reason is per-attempt
|
instruction STANDS until the human next answers, while the machine's reason is per-attempt
|
||||||
(only the most recent, as today).
|
(only the most recent, as today).
|
||||||
|
|
@ -362,6 +372,14 @@ def _build_messages(
|
||||||
f"Expert feedback: {prior_feedback}\n"
|
f"Expert feedback: {prior_feedback}\n"
|
||||||
"Produce a REVISED SavingsProposal that follows this feedback."
|
"Produce a REVISED SavingsProposal that follows this feedback."
|
||||||
)
|
)
|
||||||
|
if parse_error is not None:
|
||||||
|
prompt += (
|
||||||
|
"\n\nYour previous reply could not be PARSED as a SavingsProposal, so it was "
|
||||||
|
"discarded before any validator saw it.\n"
|
||||||
|
f"Reason: {parse_error}\n"
|
||||||
|
"Reply with a SavingsProposal whose claimed_saving_nok is greater than 0 and whose "
|
||||||
|
"affected_items each carry code, quantity and unit_cost."
|
||||||
|
)
|
||||||
return [Message(role="user", contents=[prompt])]
|
return [Message(role="user", contents=[prompt])]
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -472,7 +490,13 @@ def _grounding_text(
|
||||||
*delivered.documents,
|
*delivered.documents,
|
||||||
*(item.code for item in project.cost_items),
|
*(item.code for item in project.cost_items),
|
||||||
*(() if baseline is None else baseline.items),
|
*(() if baseline is None else baseline.items),
|
||||||
)
|
),
|
||||||
|
# P20/B: the base's own vocabulary of clause numbers travels WITH the text it was read
|
||||||
|
# off. Carried through rather than recomposed: ``run_project`` walks the base once and
|
||||||
|
# composes both halves there, and a second derivation here would be free to disagree with
|
||||||
|
# the documents it is supposed to describe (kø-(p)). The two later sources are cost CODES,
|
||||||
|
# which declare nothing, so they contribute none.
|
||||||
|
declared_references=delivered.declared_references,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -615,10 +639,19 @@ async def generate_via_llm(
|
||||||
from discarding what it already knew. Never a malformed proposal; raises ``BudgetExceeded``
|
from discarding what it already knew. Never a malformed proposal; raises ``BudgetExceeded``
|
||||||
when the meter cap is crossed."""
|
when the meter cap is crossed."""
|
||||||
|
|
||||||
async def _fetch_parsed(messages: list[Message]) -> SavingsProposal:
|
async def _fetch_parsed(build: Callable[[str | None], list[Message]]) -> SavingsProposal:
|
||||||
# Parse-robust: a malformed/text-leaked reply is retried; the meter caps total work.
|
# Parse-robust: a malformed/text-leaked reply is retried; the meter caps total work.
|
||||||
|
#
|
||||||
|
# P20/C1: the retry is no longer BLIND. It takes a BUILDER rather than a finished message
|
||||||
|
# list, because the whole defect was that the same bytes were re-sent: measured, one
|
||||||
|
# round-3 run burned 11 of its 12 rounds on replies that all failed identically. The
|
||||||
|
# builder is the caller's own ``_build_messages`` binding, so this loop cannot compose a
|
||||||
|
# prompt the outer loop would not have composed (kø-(p)); the reason is per-RETRY, like
|
||||||
|
# ``prior_rejection`` is per-attempt, and starts empty so attempt 1 is byte-identical.
|
||||||
|
parse_error: str | None = None
|
||||||
while True:
|
while True:
|
||||||
meter.tick_round() # between-attempt bound (BudgetExceeded over cap)
|
meter.tick_round() # between-attempt bound (BudgetExceeded over cap)
|
||||||
|
messages = build(parse_error)
|
||||||
# Fase 1b, funn 1b: hand the model a GRAMMAR, not a prose request. The prompt's
|
# Fase 1b, funn 1b: hand the model a GRAMMAR, not a prose request. The prompt's
|
||||||
# "Respond with ONLY a JSON object" line stays — a provider that ignores
|
# "Respond with ONLY a JSON object" line stays — a provider that ignores
|
||||||
# ``response_format`` (or a local model that does not implement it) must still be told
|
# ``response_format`` (or a local model that does not implement it) must still be told
|
||||||
|
|
@ -633,10 +666,10 @@ async def generate_via_llm(
|
||||||
# Capture BEFORE the retry: this reply was paid for, and once ``continue`` runs the
|
# Capture BEFORE the retry: this reply was paid for, and once ``continue`` runs the
|
||||||
# only record of what the model actually said is gone (Fase 1b, funn 1). Verbatim —
|
# only record of what the model actually said is gone (Fase 1b, funn 1). Verbatim —
|
||||||
# the operator is diagnosing a format failure, so any shortening removes evidence.
|
# the operator is diagnosing a format failure, so any shortening removes evidence.
|
||||||
|
reason = f"{type(exc).__name__}: {exc}"
|
||||||
if parse_failures is not None:
|
if parse_failures is not None:
|
||||||
parse_failures.append(
|
parse_failures.append(ParseFailure(text=reply.text, error=reason))
|
||||||
ParseFailure(text=reply.text, error=f"{type(exc).__name__}: {exc}")
|
parse_error = reason
|
||||||
)
|
|
||||||
continue
|
continue
|
||||||
|
|
||||||
last: Rejection | None = None
|
last: Rejection | None = None
|
||||||
|
|
@ -672,14 +705,16 @@ async def generate_via_llm(
|
||||||
# accumulated history (bounded prompt growth).
|
# accumulated history (bounded prompt growth).
|
||||||
if last is not None:
|
if last is not None:
|
||||||
fed_back.append(last)
|
fed_back.append(last)
|
||||||
messages = _build_messages(
|
candidate = await _fetch_parsed(
|
||||||
project,
|
lambda parse_error: _build_messages(
|
||||||
context,
|
project,
|
||||||
prior_rejection=last,
|
context,
|
||||||
approach=approach,
|
prior_rejection=last,
|
||||||
prior_feedback=feedback,
|
approach=approach,
|
||||||
|
prior_feedback=feedback,
|
||||||
|
parse_error=parse_error,
|
||||||
|
)
|
||||||
)
|
)
|
||||||
candidate = await _fetch_parsed(messages)
|
|
||||||
if pending_revise is not None and reviews is not None:
|
if pending_revise is not None and reviews is not None:
|
||||||
reviews[pending_revise] = replace(reviews[pending_revise], honoured=True)
|
reviews[pending_revise] = replace(reviews[pending_revise], honoured=True)
|
||||||
pending_revise = None
|
pending_revise = None
|
||||||
|
|
|
||||||
|
|
@ -414,6 +414,29 @@ def announce(
|
||||||
return "\n".join(lines)
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def criteria_block(success_criteria: str) -> str:
|
||||||
|
"""The ONE rendering of a commission's success criteria INTO a prompt, or ``""``.
|
||||||
|
|
||||||
|
MEASURED (P19 F2): ``success_criteria`` reached ``announce`` and nothing else, so the operator's
|
||||||
|
own statement of what a good answer looks like was printed for a human and withheld from the
|
||||||
|
only reader who could act on it. The approach's ``description`` has always reached the
|
||||||
|
generation prompt (``generate._build_messages``); this is its run-level sibling.
|
||||||
|
|
||||||
|
ONE composer, for kø-(p): the debate task and any later prompt that carries the criteria must
|
||||||
|
say the same thing about them, and two renderings of one commission are free to disagree about
|
||||||
|
what the operator asked for.
|
||||||
|
|
||||||
|
Empty in, empty out — omission rather than an empty heading, the ``announce`` rule. That is
|
||||||
|
what keeps every prompt of every un-commissioned run, the demo's included, byte-identical.
|
||||||
|
"""
|
||||||
|
if not success_criteria:
|
||||||
|
return ""
|
||||||
|
return (
|
||||||
|
"\nWhat the commissioner counts as success (restated verbatim from the commission):\n"
|
||||||
|
f"{success_criteria}\n"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def settle(
|
def settle(
|
||||||
coverage: tuple[ApproachOutcome, ...],
|
coverage: tuple[ApproachOutcome, ...],
|
||||||
*,
|
*,
|
||||||
|
|
|
||||||
|
|
@ -1224,6 +1224,53 @@ _DIRECTORY_PAGE_MAX: Final = 50
|
||||||
#: hence ``unquote_scalar``, this repo's ONE de-quoting rule).
|
#: hence ``unquote_scalar``, this repo's ONE de-quoting rule).
|
||||||
_FILTER_FIELDS: Final = ("req_number", "prosessnr")
|
_FILTER_FIELDS: Final = ("req_number", "prosessnr")
|
||||||
|
|
||||||
|
#: Which frontmatter keys DECLARE a reference number — the base's own vocabulary of requirement and
|
||||||
|
#: process numbers. Derived from ``_FILTER_FIELDS`` rather than restating it (two copies of the two
|
||||||
|
#: measured keys is the kø-(p) drift), plus ``seksjon``, which is the field that carries the BARE
|
||||||
|
#: form-3 number on the N corpora and on R761 (``seksjon: '10.4.3'``, ``seksjon: '11.11'``) while
|
||||||
|
#: ``req_number`` carries the composed ``Krav 10.4.3—1``. MEASURED 15.09 over the four delivered
|
||||||
|
#: bases: n100 553 distinct declared numbers, n200 1 440, n500 365, r761 2 765.
|
||||||
|
#:
|
||||||
|
#: NOT added to ``_FILTER_FIELDS`` itself, and that is deliberate: what a ``filter`` word searches
|
||||||
|
#: is measured and gated (P18), and widening it would change ``total_matches`` for every navigator
|
||||||
|
#: call. These two constants answer different questions — "what does a filter look in" and "what
|
||||||
|
#: does this document declare as its number" — over ONE list of the measured reference keys.
|
||||||
|
REFERENCE_NUMBER_FIELDS: Final = (*_FILTER_FIELDS, "seksjon")
|
||||||
|
|
||||||
|
|
||||||
|
def reference_number(file: BundleFile) -> str:
|
||||||
|
"""The document's OWN reference number as its top-level frontmatter states it, or ``""``.
|
||||||
|
|
||||||
|
First non-empty of ``_FILTER_FIELDS``, in that order: the N corpora declare ``req_number``
|
||||||
|
("Krav 4.1.2—1") and R761 declares ``prosessnr`` ("'11.11'", quoted — hence ``unquote_scalar``),
|
||||||
|
and MEASURED no delivered document declares both. ``seksjon`` is deliberately NOT read here:
|
||||||
|
this answers "which requirement IS this", and a section number names the chapter a requirement
|
||||||
|
sits in, not the requirement.
|
||||||
|
|
||||||
|
Read off ``BundleFile.frontmatter``, so P15's top-level-wins rule applies and a nested
|
||||||
|
``sources:`` entry can never answer for the concept.
|
||||||
|
"""
|
||||||
|
for key in _FILTER_FIELDS:
|
||||||
|
value = unquote_scalar(file.frontmatter.get(key, ""))
|
||||||
|
if value:
|
||||||
|
return value
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def declared_reference_numbers(file: BundleFile) -> tuple[str, ...]:
|
||||||
|
"""Every reference number this ONE document declares (``REFERENCE_NUMBER_FIELDS``), de-quoted.
|
||||||
|
|
||||||
|
The unit of the reference VOCABULARY the validator's stage 0b checks a requirement-shaped code
|
||||||
|
against (P20/B). Per DOCUMENT rather than per base, because the caller composing the grounding
|
||||||
|
already walks the base once and the boundaries it composes are the ones the gate must read
|
||||||
|
(P18/B1's rule: one document is one unit, never a blob).
|
||||||
|
"""
|
||||||
|
return tuple(
|
||||||
|
value
|
||||||
|
for key in REFERENCE_NUMBER_FIELDS
|
||||||
|
if (value := unquote_scalar(file.frontmatter.get(key, "")))
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def _matches_filter(file: BundleFile, needle: str) -> bool:
|
def _matches_filter(file: BundleFile, needle: str) -> bool:
|
||||||
"""Case-insensitive SUBSTRING over the document's title and its reference number.
|
"""Case-insensitive SUBSTRING over the document's title and its reference number.
|
||||||
|
|
|
||||||
|
|
@ -98,6 +98,7 @@ from portfolio_optimiser.mandate import (
|
||||||
MandateRoutingError,
|
MandateRoutingError,
|
||||||
announce,
|
announce,
|
||||||
candidate_from_approach,
|
candidate_from_approach,
|
||||||
|
criteria_block,
|
||||||
load_mandate,
|
load_mandate,
|
||||||
route_by_bundle,
|
route_by_bundle,
|
||||||
settle,
|
settle,
|
||||||
|
|
@ -1131,6 +1132,10 @@ async def run_project(
|
||||||
# gate's share rule needs, and composing them here — where the base is already walked — is what
|
# gate's share rule needs, and composing them here — where the base is already walked — is what
|
||||||
# keeps them from being a second, drifting reconstruction (kø-(p)).
|
# keeps them from being a second, drifting reconstruction (kø-(p)).
|
||||||
bundle_grounding: tuple[str, ...] = ()
|
bundle_grounding: tuple[str, ...] = ()
|
||||||
|
#: P20/B: the reference numbers the base's documents DECLARE, composed in the SAME walk as the
|
||||||
|
#: documents above, so the gate's vocabulary and the gate's text describe one reading of one
|
||||||
|
#: base. Empty on the road path, which is what keeps the rule unable to fire there.
|
||||||
|
bundle_references: tuple[str, ...] = ()
|
||||||
# S2c: a CALLER-OWNED sink for what the debate opens (the ``parse_failures``/``ExplorationTrace``
|
# S2c: a CALLER-OWNED sink for what the debate opens (the ``parse_failures``/``ExplorationTrace``
|
||||||
# shape). A returned value would be lost on exactly the run that most needs the evidence — a
|
# shape). A returned value would be lost on exactly the run that most needs the evidence — a
|
||||||
# budget stop mid-debate raises out of ``debate.run`` and constructs no ``RunResult`` at all.
|
# budget stop mid-debate raises out of ``debate.run`` and constructs no ``RunResult`` at all.
|
||||||
|
|
@ -1151,6 +1156,9 @@ async def run_project(
|
||||||
bundle_grounding = tuple(
|
bundle_grounding = tuple(
|
||||||
"\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files
|
"\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files
|
||||||
)
|
)
|
||||||
|
bundle_references = tuple(
|
||||||
|
ref for f in bundle.context_files for ref in okf.declared_reference_numbers(f)
|
||||||
|
)
|
||||||
# ONE bundle-id rule (Step 10, slackened S7a-3 pkt. 1): the DECLARED id is the identity and
|
# ONE bundle-id rule (Step 10, slackened S7a-3 pkt. 1): the DECLARED id is the identity and
|
||||||
# the mount is carried alongside, so a base delivered under a directory name of its own is
|
# the mount is carried alongside, so a base delivered under a directory name of its own is
|
||||||
# opened rather than refused. What is still refused, before a single model call: a base
|
# opened rather than refused. What is still refused, before a single model call: a base
|
||||||
|
|
@ -1272,7 +1280,9 @@ async def run_project(
|
||||||
# ``_fetch_parsed`` has returned, so a report from there could only ever speak once an attempt
|
# ``_fetch_parsed`` has returned, so a report from there could only ever speak once an attempt
|
||||||
# had been paid for. ONE binding feeding both the report and the gate: two compositions of one
|
# had been paid for. ONE binding feeding both the report and the gate: two compositions of one
|
||||||
# text are free to disagree, which is exactly what a report must not be able to do (kø-(p)).
|
# text are free to disagree, which is exactly what a report must not be able to do (kø-(p)).
|
||||||
delivered = Grounding(documents=(context, *bundle_grounding))
|
delivered = Grounding(
|
||||||
|
documents=(context, *bundle_grounding), declared_references=bundle_references
|
||||||
|
)
|
||||||
offer = grounding_offer(project, baseline, delivered)
|
offer = grounding_offer(project, baseline, delivered)
|
||||||
|
|
||||||
# Trekk B2 (krav 3): configured MCP servers become tools the AGENTS can call during the debate.
|
# Trekk B2 (krav 3): configured MCP servers become tools the AGENTS can call during the debate.
|
||||||
|
|
@ -1359,8 +1369,16 @@ async def run_project(
|
||||||
async with AsyncExitStack() as mcp_stack:
|
async with AsyncExitStack() as mcp_stack:
|
||||||
for live_tool in live_mcp_tools:
|
for live_tool in live_mcp_tools:
|
||||||
await mcp_stack.enter_async_context(live_tool)
|
await mcp_stack.enter_async_context(live_tool)
|
||||||
|
# P20/A2: the commission's success criteria reach the DEBATE — the prompt where
|
||||||
|
# ``declare_requirement`` is available — and not only the announcement. Composed through
|
||||||
|
# ``mandate.criteria_block``, the ONE renderer (kø-(p)); empty without a commission, so
|
||||||
|
# every un-commissioned run's task message is byte-identical, which is what keeps the
|
||||||
|
# golden transcript unchanged. MEASURED: on the exploration path there is nothing to
|
||||||
|
# carry — ``main()`` passes ``explore()`` no ``success_criteria`` at all, so its
|
||||||
|
# objective IS the prompt and already reaches the task.
|
||||||
|
criteria = criteria_block(mandate.success_criteria) if mandate is not None else ""
|
||||||
result = await debate.run(
|
result = await debate.run(
|
||||||
f"Find a cost-saving measure for {project.id}.\nContext:\n{context}"
|
f"Find a cost-saving measure for {project.id}.{criteria}\nContext:\n{context}"
|
||||||
)
|
)
|
||||||
finally:
|
finally:
|
||||||
# ``finally``, the ``write_parse_failures`` precedent: any exception leaving the debate —
|
# ``finally``, the ``write_parse_failures`` precedent: any exception leaving the debate —
|
||||||
|
|
@ -1610,7 +1628,12 @@ async def run_project(
|
||||||
cost_baseline_anchored=baseline is not None,
|
cost_baseline_anchored=baseline is not None,
|
||||||
# P19/B2: what the run made of each code it was handed. Derived from the SAME classifier
|
# P19/B2: what the run made of each code it was handed. Derived from the SAME classifier
|
||||||
# the gate uses (kø-(p)), off the proposal being stamped — never re-read from anywhere.
|
# the gate uses (kø-(p)), off the proposal being stamped — never re-read from anywhere.
|
||||||
code_forms=classify_codes([item.code for item in proposal.affected_items]),
|
code_forms=classify_codes(
|
||||||
|
[item.code for item in proposal.affected_items],
|
||||||
|
# P20/B: the THIRD value, ``requirement``, needs the base's own vocabulary. Read off
|
||||||
|
# the SAME ``delivered`` the gate was handed, never a second composition.
|
||||||
|
delivered,
|
||||||
|
),
|
||||||
# WHICH corpus was judged, and whether the base named itself or the mount named it for it.
|
# WHICH corpus was judged, and whether the base named itself or the mount named it for it.
|
||||||
# Read off the SAME resolution the run opened the base with (kø-(p)); ``None`` on the road
|
# Read off the SAME resolution the run opened the base with (kø-(p)); ``None`` on the road
|
||||||
# path, where no knowledge base exists to name.
|
# path, where no knowledge base exists to name.
|
||||||
|
|
@ -2290,6 +2313,32 @@ def _write_multibase_summary(
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def announced_subject(project_id: str | None, across_bundle: Sequence[str]) -> str:
|
||||||
|
"""WHO the announcement is about: the project, the routed bases, or the portfolio (P20/C2).
|
||||||
|
|
||||||
|
MEASURED (P17b F5): ``--across-bundle`` takes no ``--project-id`` — each base's project is read
|
||||||
|
from that base's own IR projection — so ``args.project_id or "the portfolio"`` announced a
|
||||||
|
multi-base commission as a portfolio pass, which is a different mode entirely.
|
||||||
|
|
||||||
|
The ids are resolved for NAMING ONLY, and a base that cannot be resolved falls back to its
|
||||||
|
directory name. That is the ``dimension_label`` precedent one line above the call site,
|
||||||
|
verbatim: a configuration that fails to load is left unnamed here and refused a moment later by
|
||||||
|
the dispatch, which stays the single owner of that refusal — announcing must never change which
|
||||||
|
error an operator sees.
|
||||||
|
"""
|
||||||
|
if project_id:
|
||||||
|
return project_id
|
||||||
|
if not across_bundle:
|
||||||
|
return "the portfolio"
|
||||||
|
names = []
|
||||||
|
for raw in across_bundle:
|
||||||
|
try:
|
||||||
|
names.append(okf.reconcile_bundle_id(raw).id)
|
||||||
|
except (okf.BundleIdMismatch, FileNotFoundError, ValueError, OSError):
|
||||||
|
names.append(Path(raw).name)
|
||||||
|
return ", ".join(names)
|
||||||
|
|
||||||
|
|
||||||
def resolve_bundle_routing(
|
def resolve_bundle_routing(
|
||||||
bundle_dirs: Sequence[str],
|
bundle_dirs: Sequence[str],
|
||||||
) -> tuple[tuple[str, str, str], ...]:
|
) -> tuple[tuple[str, str, str], ...]:
|
||||||
|
|
@ -4045,7 +4094,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
print(
|
print(
|
||||||
announce(
|
announce(
|
||||||
mandate,
|
mandate,
|
||||||
project_id=args.project_id or "the portfolio",
|
project_id=announced_subject(args.project_id, args.across_bundle or ()),
|
||||||
# From ARGV, never the constants (P16 B2). The announcement is the one thing
|
# From ARGV, never the constants (P16 B2). The announcement is the one thing
|
||||||
# printed BEFORE the first paid call, and its whole job is to say what this run
|
# printed BEFORE the first paid call, and its whole job is to say what this run
|
||||||
# will do; reading the defaults was correct only while main() could not do
|
# will do; reading the defaults was correct only while main() could not do
|
||||||
|
|
|
||||||
|
|
@ -250,29 +250,50 @@ _GROUNDING_MIN_INERT_DOCUMENTS: Final = 10
|
||||||
#: Bare numbers are deliberately EXCLUDED, with the number: K2 carries 46 394 occurrences over
|
#: Bare numbers are deliberately EXCLUDED, with the number: K2 carries 46 394 occurrences over
|
||||||
#: 2 117 distinct values (P7 § 2), so counting them would make every report positive and the
|
#: 2 117 distinct values (P7 § 2), so counting them would make every report positive and the
|
||||||
#: measurement inert — the repo's cardinal class, a gate that can only come out green.
|
#: measurement inert — the repo's cardinal class, a gate that can only come out green.
|
||||||
|
#: The four forms, NAMED rather than reached by index: P20/B needs two of them by themselves
|
||||||
|
#: (a requirement/process number), and ``IDENTIFIER_FORMS[1:3]`` in a second module would be a
|
||||||
|
#: positional dependency on a tuple literal — the kø-(p) shape with no compiler to catch it.
|
||||||
|
_FORM_SEPARATED_UPPER: Final = re.compile(r"\b[A-ZÆØÅ][A-ZÆØÅ0-9]*(?:[-_][A-ZÆØÅ0-9]+)+\b")
|
||||||
|
_FORM_REQUIREMENT_NUMBER: Final = re.compile(r"Krav\s+\d+(?:\.\d+)*\s*[\u2014-]\s*\d+(?:_\d+)?")
|
||||||
|
_FORM_PROCESS_NUMBER: Final = re.compile(r"(?<![\d.])[1-9]\d{0,2}(?:\.\d{1,3}){1,4}\b(?!\.\d)")
|
||||||
|
_FORM_PROCESS_HEADING: Final = re.compile(r"(?<!\d )(?<![\d.])[1-9]\d{0,2} [A-ZÆØÅ]{5,}\b")
|
||||||
|
|
||||||
IDENTIFIER_FORMS: Final = (
|
IDENTIFIER_FORMS: Final = (
|
||||||
# ``SHA-01``, ``RIM-02``, ``B-20-00-00``, ``FOR-2011-12-06-1357`` (K2's 50) — and, since P19/B2
|
# ``SHA-01``, ``RIM-02``, ``B-20-00-00``, ``FOR-2011-12-06-1357`` (K2's 50) — and, since P19/B2
|
||||||
# made the same forms decide ``prose`` vs ``identifier``, an UPPERCASE separated token with no
|
# made the same forms decide ``prose`` vs ``identifier``, an UPPERCASE separated token with no
|
||||||
# digits at all. MEASURED: this repo's own ``ENERGI-TOTAL-EL`` matched neither of the pre-P19
|
# digits at all. MEASURED: this repo's own ``ENERGI-TOTAL-EL`` matched neither of the pre-P19
|
||||||
# forms, so the classifier called a real cost code prose; a gate is only allowed to be wrong in
|
# forms, so the classifier called a real cost code prose; a gate is only allowed to be wrong in
|
||||||
# the direction that admits too much.
|
# the direction that admits too much.
|
||||||
re.compile(r"\b[A-ZÆØÅ][A-ZÆØÅ0-9]*(?:[-_][A-ZÆØÅ0-9]+)+\b"),
|
_FORM_SEPARATED_UPPER,
|
||||||
# ``Krav 3.3.1—13`` (the N corpora's dominant form). EM-DASH U+2014 AND the hyphen, because the
|
# ``Krav 3.3.1—13`` (the N corpora's dominant form). EM-DASH U+2014 AND the hyphen, because the
|
||||||
# binding known positive is the em-dash spelling and only the em-dash spelling scores 6 of 6.
|
# binding known positive is the em-dash spelling and only the em-dash spelling scores 6 of 6.
|
||||||
# The trailing ``(?:_\d+)?`` is MEASURED, not defensive: one of the 26 fasit references is
|
# The trailing ``(?:_\d+)?`` is MEASURED, not defensive: one of the 26 fasit references is
|
||||||
# ``Krav 3.3.2—1_1``, and without it the classifier called that real reference prose.
|
# ``Krav 3.3.2—1_1``, and without it the classifier called that real reference prose.
|
||||||
re.compile(r"Krav\s+\d+(?:\.\d+)*\s*[\u2014-]\s*\d+(?:_\d+)?"),
|
_FORM_REQUIREMENT_NUMBER,
|
||||||
# R761's process numbers, ``12.1`` / ``52.11`` (P19 B1). MEASURED: all six ``ref`` values in
|
# R761's process numbers, ``12.1`` / ``52.11`` (P19 B1). MEASURED: all six ``ref`` values in
|
||||||
# ``contexts/kontrakt-sorasen-2027/fasit.json`` are of this shape and NEITHER of the first two
|
# ``contexts/kontrakt-sorasen-2027/fasit.json`` are of this shape and NEITHER of the first two
|
||||||
# forms matches one of them, so r761's whole offer was 3 identifiers over 6.5 MB. The trailing
|
# forms matches one of them, so r761's whole offer was 3 identifiers over 6.5 MB. The trailing
|
||||||
# ``(?!\.\d)`` is what keeps a Norwegian date out: ``15.09.2026`` would otherwise contribute
|
# ``(?!\.\d)`` is what keeps a Norwegian date out: ``15.09.2026`` would otherwise contribute
|
||||||
# its ``15.09`` prefix, and a date is not a requirement.
|
# its ``15.09`` prefix, and a date is not a requirement.
|
||||||
re.compile(r"(?<![\d.])[1-9]\d{0,2}(?:\.\d{1,3}){1,4}\b(?!\.\d)"),
|
_FORM_PROCESS_NUMBER,
|
||||||
# ``65 ASFALTDEKKER`` — a process number and its heading, the form a price schedule's section
|
# ``65 ASFALTDEKKER`` — a process number and its heading, the form a price schedule's section
|
||||||
# rows carry (P18 § 2 measured it at 29 of 2 756 documents).
|
# rows carry (P18 § 2 measured it at 29 of 2 756 documents).
|
||||||
re.compile(r"(?<!\d )(?<![\d.])[1-9]\d{0,2} [A-ZÆØÅ]{5,}\b"),
|
_FORM_PROCESS_HEADING,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
#: The two forms a REQUIREMENT or PROCESS number takes (P20/B). MEASURED over the four delivered
|
||||||
|
#: bases: these are the shapes that live in ``req_number`` ("Krav 4.1.2—1"), ``prosessnr``
|
||||||
|
#: ("'11.11'") and ``seksjon`` ("'10.4.3'") — the fields a base uses to number its own clauses.
|
||||||
|
#: The other two forms are NOT here: ``SHA-01``-style tokens are what a price schedule's cost lines
|
||||||
|
#: look like, and ``65 ASFALTDEKKER`` IS a schedule section row.
|
||||||
|
REQUIREMENT_FORMS: Final = (_FORM_REQUIREMENT_NUMBER, _FORM_PROCESS_NUMBER)
|
||||||
|
|
||||||
|
|
||||||
|
def has_requirement_form(code: str) -> bool:
|
||||||
|
"""Whether ``code`` is shaped like a requirement or process number. FULL-MATCH, never a search,
|
||||||
|
for ``has_identifier_form``'s reason: ``impulsventilator 12.1`` is not a clause number."""
|
||||||
|
return any(form.fullmatch(code) for form in REQUIREMENT_FORMS)
|
||||||
|
|
||||||
|
|
||||||
def identifier_tokens(text: str) -> set[str]:
|
def identifier_tokens(text: str) -> set[str]:
|
||||||
"""Every DISTINCT token of any ``IDENTIFIER_FORMS`` shape in ``text``. One reader, two callers.
|
"""Every DISTINCT token of any ``IDENTIFIER_FORMS`` shape in ``text``. One reader, two callers.
|
||||||
|
|
@ -296,9 +317,27 @@ def has_identifier_form(code: str) -> bool:
|
||||||
return any(form.fullmatch(code) for form in IDENTIFIER_FORMS)
|
return any(form.fullmatch(code) for form in IDENTIFIER_FORMS)
|
||||||
|
|
||||||
|
|
||||||
def classify_codes(codes: Sequence[str]) -> dict[str, str]:
|
def classify_codes(codes: Sequence[str], grounding: Grounding | None = None) -> dict[str, str]:
|
||||||
"""``{code: "identifier" | "prose"}`` — P19/B2's report, in ONE place for both consumers."""
|
"""``{code: "identifier" | "prose" | "requirement"}`` — P19/B2's report, widened by P20/B, in
|
||||||
return {code: "identifier" if has_identifier_form(code) else "prose" for code in codes}
|
ONE place for both consumers.
|
||||||
|
|
||||||
|
``requirement`` is the third value: a code shaped like a clause number (``REQUIREMENT_FORMS``)
|
||||||
|
AND declared as one by the input's own documents. It is a REPORT about what the run made of a
|
||||||
|
code, not the gate — the gate is ``_reference_refusal`` below and fires on the COMPLEMENT, a
|
||||||
|
requirement-shaped code the base's vocabulary does NOT contain.
|
||||||
|
|
||||||
|
``grounding=None`` is the pre-P20 answer exactly: without the input there is no vocabulary to
|
||||||
|
check against, so no code can be called a requirement. ``stress.py`` re-derives with ``None``
|
||||||
|
for runs that predate the field, and says so.
|
||||||
|
"""
|
||||||
|
vocabulary = frozenset() if grounding is None else grounding.reference_vocabulary
|
||||||
|
out: dict[str, str] = {}
|
||||||
|
for code in codes:
|
||||||
|
if has_requirement_form(code) and code in vocabulary:
|
||||||
|
out[code] = "requirement"
|
||||||
|
else:
|
||||||
|
out[code] = "identifier" if has_identifier_form(code) else "prose"
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
|
|
@ -320,6 +359,16 @@ class Grounding:
|
||||||
"""
|
"""
|
||||||
|
|
||||||
documents: tuple[str, ...]
|
documents: tuple[str, ...]
|
||||||
|
#: Every reference number the input's documents DECLARE in their own top-level frontmatter
|
||||||
|
#: (``okf.declared_reference_numbers``), one entry per declaration — the base's own vocabulary
|
||||||
|
#: of requirement, process and section numbers. P20/B checks a requirement-shaped code against
|
||||||
|
#: it.
|
||||||
|
#:
|
||||||
|
#: DEFAULTED, the ``skipped_links`` half rather than ``cost_baseline_anchored``'s: an empty
|
||||||
|
#: vocabulary is an honest POSITIVE statement ("this input declares no clause numbers"), and it
|
||||||
|
#: is what keeps every caller written before today — the road path, every fixture, ``of`` —
|
||||||
|
#: unchanged by construction, since the gate cannot fire without one.
|
||||||
|
declared_references: tuple[str, ...] = ()
|
||||||
|
|
||||||
@classmethod
|
@classmethod
|
||||||
def of(cls, text: str) -> Grounding:
|
def of(cls, text: str) -> Grounding:
|
||||||
|
|
@ -335,6 +384,11 @@ class Grounding:
|
||||||
"""How many of the documents contain ``token`` — the numerator, in the unit of the rule."""
|
"""How many of the documents contain ``token`` — the numerator, in the unit of the rule."""
|
||||||
return sum(1 for document in self.documents if token in document)
|
return sum(1 for document in self.documents if token in document)
|
||||||
|
|
||||||
|
@cached_property
|
||||||
|
def reference_vocabulary(self) -> frozenset[str]:
|
||||||
|
"""The DISTINCT reference numbers this input declares. The denominator P20/B names."""
|
||||||
|
return frozenset(self.declared_references)
|
||||||
|
|
||||||
@cached_property
|
@cached_property
|
||||||
def identifiers(self) -> frozenset[str]:
|
def identifiers(self) -> frozenset[str]:
|
||||||
"""Every distinct identifier-shaped token this input OFFERS (P8's count, P19/B3's guard).
|
"""Every distinct identifier-shaped token this input OFFERS (P8's count, P19/B3's guard).
|
||||||
|
|
@ -401,10 +455,60 @@ def _form_refusal(grounding: Grounding, code: str, anchored_codes: frozenset[str
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _reference_refusal(grounding: Grounding, code: str) -> str | None:
|
||||||
|
"""Why a requirement-shaped ``code`` cannot be a cost line of an UNANCHORED input (P20/B).
|
||||||
|
|
||||||
|
**The order's own rule was FELLED BY MEASUREMENT before anything was built on it.** It reads:
|
||||||
|
a code is a requirement when it matches form 2 or 3 AND "står som ``req_number``/``prosessnr``
|
||||||
|
i toppnivå-frontmatter" — refuse that. Measured 15.09 against the two known positives the same
|
||||||
|
order names:
|
||||||
|
|
||||||
|
* ``10.4`` (n500, tunnel-04, ``validated``) is declared NOWHERE in n500's frontmatter. The base
|
||||||
|
declares ``seksjon: 10.4.1`` … ``10.4.4`` and ``req_number: Krav 10.4.3—2``; the bare ``10.4``
|
||||||
|
is a section PREFIX that occurs in 12 of 274 documents and is no document's own number;
|
||||||
|
* ``1.10.4`` (r761, lindaas a4, ``validated``) is not one of r761's 2 727 ``prosessnr`` nor one
|
||||||
|
of its 2 753 ``seksjon`` values. It occurs in ONE of 2 756 documents, as prose: "iht.
|
||||||
|
vegnormal N200 Vegbygging kap. 1.10.4".
|
||||||
|
|
||||||
|
So the ordered rule fires on NEITHER of its own known positives. The COMPLEMENT does, and it is
|
||||||
|
the better-grounded rule besides: ``_ground_against_input``'s docstring already admits that this
|
||||||
|
stage "fails OPEN … on a coincidental match", and for one shape — a clause number — the base
|
||||||
|
hands us the vocabulary needed to close exactly that hole. A form-3 token that is NOT one of the
|
||||||
|
numbers this base declares was matched in prose by accident.
|
||||||
|
|
||||||
|
MEASURED over every code of round 3 and P17b (24 codes, 10 runs): exactly two are
|
||||||
|
requirement-shaped, they are the two known positives, and neither is in its base's vocabulary.
|
||||||
|
All five of ``contexts/kontrakt-sorasen-2027``'s REAL process codes (``12.1``, ``12.12``,
|
||||||
|
``22.1``, ``52.11``, ``51.1``) ARE declared ``prosessnr`` and pass — which is what keeps the
|
||||||
|
R761 risk the order names (a process number is both a clause and a settlement post) from
|
||||||
|
turning into a wholesale refusal of the one context set built on real codes.
|
||||||
|
|
||||||
|
**The generality guard, ``_form_refusal``'s pattern:** an input that declares no reference
|
||||||
|
numbers at all cannot be answered in a vocabulary it does not have, so the rule cannot fire
|
||||||
|
there. That is what leaves every pre-P20 fixture untouched rather than exempted.
|
||||||
|
|
||||||
|
The message NAMES THE DENOMINATOR (ansikt 4, and Step 5 feeds it verbatim into the next
|
||||||
|
attempt): "not one of the 2 765 it declares" is actionable where "ungrounded" is not.
|
||||||
|
"""
|
||||||
|
if not has_requirement_form(code):
|
||||||
|
return None
|
||||||
|
vocabulary = grounding.reference_vocabulary
|
||||||
|
if not vocabulary or code in vocabulary:
|
||||||
|
return None
|
||||||
|
sample = ", ".join(sorted(vocabulary)[:3])
|
||||||
|
return (
|
||||||
|
f"is shaped like a requirement or process number, but it is not one of the "
|
||||||
|
f"{len(vocabulary)} this knowledge base declares (for example {sample}) — it was matched "
|
||||||
|
"in prose by coincidence, and an unanchored base carries no price for a clause number"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def _ground_against_input(
|
def _ground_against_input(
|
||||||
proposal: SavingsProposal,
|
proposal: SavingsProposal,
|
||||||
grounding: Grounding,
|
grounding: Grounding,
|
||||||
anchored_codes: frozenset[str] = frozenset(),
|
anchored_codes: frozenset[str] = frozenset(),
|
||||||
|
*,
|
||||||
|
anchored: bool = True,
|
||||||
) -> Rejection | None:
|
) -> Rejection | None:
|
||||||
"""P7: every identifier the proposal builds on must appear VERBATIM in the input it was built
|
"""P7: every identifier the proposal builds on must appear VERBATIM in the input it was built
|
||||||
from, or the verdict falls.
|
from, or the verdict falls.
|
||||||
|
|
@ -459,6 +563,15 @@ def _ground_against_input(
|
||||||
shapeless = _form_refusal(grounding, item.code, anchored_codes)
|
shapeless = _form_refusal(grounding, item.code, anchored_codes)
|
||||||
if shapeless is not None:
|
if shapeless is not None:
|
||||||
violations.append(f"ungrounded identifier {item.code!r}: it {shapeless}")
|
violations.append(f"ungrounded identifier {item.code!r}: it {shapeless}")
|
||||||
|
continue
|
||||||
|
# P20/B: shaped, grounded, not inert — and still a clause number the base never declared.
|
||||||
|
# UNANCHORED only: with a baseline, stage 0 has already ruled every code that reaches here
|
||||||
|
# a real line of this project, and the weaker stage must not overrule the stronger one (the
|
||||||
|
# sentence ``_form_refusal`` and ``_grounding_text`` both carry).
|
||||||
|
if not anchored:
|
||||||
|
coincidental = _reference_refusal(grounding, item.code)
|
||||||
|
if coincidental is not None:
|
||||||
|
violations.append(f"ungrounded identifier {item.code!r}: it {coincidental}")
|
||||||
if not violations:
|
if not violations:
|
||||||
return None
|
return None
|
||||||
return Rejection(proposal=proposal, reason="; ".join(violations))
|
return Rejection(proposal=proposal, reason="; ".join(violations))
|
||||||
|
|
@ -504,6 +617,11 @@ def validate_proposal(
|
||||||
proposal,
|
proposal,
|
||||||
grounding,
|
grounding,
|
||||||
frozenset() if baseline is None else frozenset(baseline.items),
|
frozenset() if baseline is None else frozenset(baseline.items),
|
||||||
|
# P20/B: an EXPLICIT flag, never ``not anchored_codes``. A baseline with no items and
|
||||||
|
# no baseline at all are different facts, and conflating them is the very shape this
|
||||||
|
# repo refuses elsewhere (``cost_baseline_anchored`` is required without a default for
|
||||||
|
# the same reason).
|
||||||
|
anchored=baseline is not None,
|
||||||
)
|
)
|
||||||
if adrift is not None:
|
if adrift is not None:
|
||||||
return adrift
|
return adrift
|
||||||
|
|
|
||||||
|
|
@ -121,12 +121,17 @@ def test_the_correction_is_to_read_it_and_then_it_is_accepted() -> None:
|
||||||
answer = tools["declare_requirement"].func(
|
answer = tools["declare_requirement"].func(
|
||||||
bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1"
|
bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1"
|
||||||
)
|
)
|
||||||
assert answer == {
|
# P20/A1 widened the reply: the three arguments PLUS the document's own title and number and
|
||||||
"declared": True,
|
# the sentence saying what the declaration binds. Asserted key by key rather than by equality,
|
||||||
"bundle_id": "tunnel-hauglia",
|
# because an exact-dict assert here would fail on every future field while saying nothing about
|
||||||
"path": path,
|
# the one property this arm exists for — that the declaration was ACCEPTED and RECORDED.
|
||||||
"ref": "Krav 12.1",
|
assert answer["declared"] is True
|
||||||
}
|
assert (answer["bundle_id"], answer["path"], answer["ref"]) == (
|
||||||
|
"tunnel-hauglia",
|
||||||
|
path,
|
||||||
|
"Krav 12.1",
|
||||||
|
)
|
||||||
|
assert set(answer) == {"declared", "bundle_id", "path", "ref", "title", "req_number", "binds"}
|
||||||
assert declared == [DeclaredRequirement(bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1")]
|
assert declared == [DeclaredRequirement(bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1")]
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
129
tests/test_parse_error_feedback_loadbearing.py
Normal file
129
tests/test_parse_error_feedback_loadbearing.py
Normal file
|
|
@ -0,0 +1,129 @@
|
||||||
|
"""P20 DEL C — a parse failure that does not burn the round ledger, and an announcement that
|
||||||
|
names what it is about.
|
||||||
|
|
||||||
|
**C1, what was measured (P19 F4).** ``_fetch_parsed`` retried a malformed reply with the
|
||||||
|
BYTE-IDENTICAL prompt. One round-3 run, ``kontrakt-sorasen-04``, left a
|
||||||
|
``{run_id}-parse-failures.json`` with ELEVEN rows, every one of them the same failure
|
||||||
|
(``claimed_saving_nok`` ≤ 0) — eleven of the run's twelve rounds, spent re-asking a question the
|
||||||
|
model had already answered the same wrong way, because nothing ever told it what was wrong. Step
|
||||||
|
5's ``prior_rejection`` carries VALIDATOR rejections; a reply that never parsed never reaches a
|
||||||
|
validator, so no existing block could carry it.
|
||||||
|
|
||||||
|
**C2, what was measured (P17b F5).** ``--across-bundle`` takes no ``--project-id``, so the
|
||||||
|
announcement — the one thing printed before the first paid call — said "Run mandate for the
|
||||||
|
portfolio" for a commission dispatched across two named knowledge bases.
|
||||||
|
|
||||||
|
What each arm pins:
|
||||||
|
|
||||||
|
(a) the reason reaches the NEXT attempt's prompt, verbatim, asserted on what the client received;
|
||||||
|
(b) attempt 1 is byte-identical: a run whose first reply parses sends the pre-P20 prompt;
|
||||||
|
(c) the block is per-RETRY — once a reply parses, the next attempt does not carry a stale reason;
|
||||||
|
(d) the evidence artefact is unchanged: the verbatim text is still captured (funn 1 stands);
|
||||||
|
(e) the announcement names the routed bases by their DECLARED ids;
|
||||||
|
(f) an unresolvable base falls back to its directory name rather than refusing — the
|
||||||
|
``dimension_label`` precedent, so announcing never changes which error an operator sees;
|
||||||
|
(g) a single-project run and a portfolio pass announce exactly as before.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from portfolio_optimiser.budget import Budget, TokenMeter
|
||||||
|
from portfolio_optimiser.generate import ParseFailure, _build_messages, generate_via_llm
|
||||||
|
from portfolio_optimiser.reference_domain import Project
|
||||||
|
from portfolio_optimiser.run import announced_subject
|
||||||
|
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||||
|
|
||||||
|
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||||
|
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||||
|
|
||||||
|
#: The exact failure ``kontrakt-sorasen-04`` produced eleven times: a well-formed JSON object
|
||||||
|
#: whose ``claimed_saving_nok`` is 0, refused by pydantic before any validator sees it.
|
||||||
|
_UNPARSEABLE = (
|
||||||
|
'{"measure":"m","affected_items":[{"code":"C-1","quantity":1000,"unit_cost":100}],'
|
||||||
|
'"claimed_saving_nok":0}'
|
||||||
|
)
|
||||||
|
#: The grounding this loop declares: P7 requires the code to occur in the input verbatim.
|
||||||
|
_CONTEXT = "The project price schedule carries cost line C-1."
|
||||||
|
_VALID = (
|
||||||
|
'{"measure":"m","affected_items":[{"code":"C-1","quantity":1000,"unit_cost":100}],'
|
||||||
|
'"claimed_saving_nok":5000}'
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _project() -> Project:
|
||||||
|
return Project(id="p", name="n", description="d", currency="NOK", cost_items=(), docs_dir=".")
|
||||||
|
|
||||||
|
|
||||||
|
def _generate(replies: list[str]) -> tuple[Any, list[str], list[ParseFailure]]:
|
||||||
|
"""Drive the real loop with the repo's ONE scripted client, recording every prompt sent."""
|
||||||
|
sink: list[str] = []
|
||||||
|
failures: list[ParseFailure] = []
|
||||||
|
client = ScriptedChatClient(script=list(replies), sink=sink, role="proposer")
|
||||||
|
result = asyncio.run(
|
||||||
|
generate_via_llm(
|
||||||
|
client,
|
||||||
|
_project(),
|
||||||
|
_CONTEXT,
|
||||||
|
TokenMeter(Budget(max_tokens=100_000, max_rounds=8)),
|
||||||
|
parse_failures=failures,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return result, sink, failures
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------------------- C1
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_parse_reason_reaches_the_next_attempts_prompt() -> None:
|
||||||
|
"""(a) The retry is no longer blind — asserted on what the CLIENT received."""
|
||||||
|
_result, sink, _failures = _generate([_UNPARSEABLE, _VALID])
|
||||||
|
assert len(sink) >= 2, sink
|
||||||
|
assert "could not be PARSED" in sink[1]
|
||||||
|
assert "claimed_saving_nok" in sink[1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_first_prompt_is_byte_identical_to_the_pre_p20_one() -> None:
|
||||||
|
"""(b) Without this, (a) could be satisfied by always appending the block."""
|
||||||
|
_result, sink, _failures = _generate([_VALID])
|
||||||
|
assert "could not be PARSED" not in sink[0]
|
||||||
|
assert sink[0] == _build_messages(_project(), _CONTEXT)[0].text
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_reason_does_not_survive_a_reply_that_parsed() -> None:
|
||||||
|
"""(c) Per-RETRY, like ``prior_rejection`` is per-attempt: no stale instruction."""
|
||||||
|
assert "could not be PARSED" not in _build_messages(_project(), "c")[0].text
|
||||||
|
carried = _build_messages(_project(), "c", parse_error="ValueError: x")[0].text
|
||||||
|
assert "ValueError: x" in carried
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_verbatim_evidence_is_still_captured() -> None:
|
||||||
|
"""(d) Fase 1b funn 1 stands: the paid reply's own text survives the retry."""
|
||||||
|
_result, _sink, failures = _generate([_UNPARSEABLE, _VALID])
|
||||||
|
assert len(failures) == 1
|
||||||
|
assert failures[0].text == _UNPARSEABLE
|
||||||
|
assert failures[0].error.startswith("ValidationError")
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------------------- C2
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_announcement_names_the_routed_bases() -> None:
|
||||||
|
"""(e) Two bases, two declared ids — not "the portfolio"."""
|
||||||
|
assert announced_subject(None, (str(_TUNNEL),)) == "tunnel-hauglia"
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_unresolvable_base_falls_back_to_its_directory_name(tmp_path: Path) -> None:
|
||||||
|
"""(f) Naming never changes which error an operator sees (the dimension_label precedent)."""
|
||||||
|
missing = tmp_path / "not-a-base"
|
||||||
|
assert announced_subject(None, (str(missing),)) == "not-a-base"
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_two_older_subjects_are_unchanged() -> None:
|
||||||
|
"""(g) A named project wins; no bases at all is still the portfolio."""
|
||||||
|
assert announced_subject("proj-1", ()) == "proj-1"
|
||||||
|
assert announced_subject("proj-1", ("ignored",)) == "proj-1"
|
||||||
|
assert announced_subject(None, ()) == "the portfolio"
|
||||||
275
tests/test_requirement_number_gate_loadbearing.py
Normal file
275
tests/test_requirement_number_gate_loadbearing.py
Normal file
|
|
@ -0,0 +1,275 @@
|
||||||
|
"""P20 DEL B — a clause number is not a price, and the base's own vocabulary is what says so.
|
||||||
|
|
||||||
|
**What was measured.** Three paid rounds and one multi-base pass carried FOUR ``validated``
|
||||||
|
proposals whose cost code was a chapter number of a standard. Two survive in the recorded outboxes
|
||||||
|
and are this arm's known positives:
|
||||||
|
|
||||||
|
* ``10.4`` — tunnel-hauglia round 3, base ``vegnormal-n500-2024``, ``validated``;
|
||||||
|
* ``1.10.4`` — lindaas P17b, base ``vegnormal-r761-2025``, ``validated``.
|
||||||
|
|
||||||
|
Both are GROUNDED in P7's sense (they occur verbatim in the input) and neither is INERT in
|
||||||
|
P18/B1's sense (``10.4`` in 12 of 274 documents, ``1.10.4`` in 1 of 2 756). Stage 0 never ran:
|
||||||
|
no vegnormal base ships a cost baseline. Nothing in the gate could say what they are.
|
||||||
|
|
||||||
|
**THE ORDER'S OWN RULE WAS FELLED BY MEASUREMENT BEFORE ANYTHING WAS BUILT ON IT.** B1 reads: a
|
||||||
|
code is a requirement when it has form 2 or 3 AND "står som ``req_number``/``prosessnr`` i
|
||||||
|
toppnivå-frontmatter i minst ett av grunnlagets dokumenter" — refuse that. Measured 15.09:
|
||||||
|
|
||||||
|
* n500 declares ``seksjon: 10.4.1`` … ``10.4.4`` and ``req_number: Krav 10.4.3—2``. The bare
|
||||||
|
``10.4`` is declared NOWHERE — it is a section PREFIX;
|
||||||
|
* r761 declares 2 727 ``prosessnr`` and 2 753 ``seksjon``. ``1.10.4`` is NONE of them: it occurs
|
||||||
|
once, as prose, in "Krav til materialer skal være iht. vegnormal N200 Vegbygging kap. 1.10.4".
|
||||||
|
|
||||||
|
The ordered rule therefore fires on NEITHER of its own known positives. The COMPLEMENT fires on
|
||||||
|
BOTH, and it closes a hole ``_ground_against_input`` already admits in writing — "it fails OPEN …
|
||||||
|
on a coincidental match". For one shape, a clause number, the base hands us the vocabulary needed
|
||||||
|
to tell a real reference from a coincidence, and that is the rule built here.
|
||||||
|
|
||||||
|
The complement is also what SPARES the one context set built on real process codes: all five of
|
||||||
|
``contexts/kontrakt-sorasen-2027``'s codes are declared ``prosessnr`` and pass. Under the ordered
|
||||||
|
rule every one of them would have been refused on an unanchored r761 run, and the set's positive
|
||||||
|
arms would have become unmeasurable — the R761 risk the order names, arriving through the door it
|
||||||
|
was pointed away from.
|
||||||
|
|
||||||
|
Measured over EVERY code of round 3 and P17b (24 codes, 10 runs): exactly two are
|
||||||
|
requirement-shaped, they are the two known positives, and the replay flips exactly those two.
|
||||||
|
|
||||||
|
What each arm pins:
|
||||||
|
|
||||||
|
(a) known positive — the ``10.4`` proposal, replayed against the base it actually ran on, is
|
||||||
|
``rejected``, and the reason names the denominator;
|
||||||
|
(b) known positive — ditto ``1.10.4`` on r761;
|
||||||
|
(c) known negative — a code the base DOES declare (sorasen's real ``12.1``) still validates;
|
||||||
|
(d) known negative — the gate is OFF when the run is anchored, even for a clause-shaped code;
|
||||||
|
(e) the generality guard — an input that declares no reference numbers at all cannot trip the rule,
|
||||||
|
which is what leaves every pre-P20 fixture untouched rather than exempted;
|
||||||
|
(f) K2's identifier forms are untouched: ``SHA-01`` is not requirement-shaped;
|
||||||
|
(g) ``classify_codes``' third value, and its denominator-free reading (``grounding=None``) that
|
||||||
|
``stress.py`` re-derives with;
|
||||||
|
(h) the vocabulary travels WITH the text through ``_grounding_text``, so the gate the generation
|
||||||
|
loop runs sees what the run composed;
|
||||||
|
(i) a run composes the vocabulary from the base it opened — asserted end-to-end through
|
||||||
|
``run_project``, not on the composer.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from portfolio_optimiser import okf
|
||||||
|
from portfolio_optimiser.generate import _grounding_text
|
||||||
|
from portfolio_optimiser.ir import CostBaseline, SavingsProposal
|
||||||
|
from portfolio_optimiser.reference_domain import Project
|
||||||
|
from portfolio_optimiser.validator import (
|
||||||
|
Grounding,
|
||||||
|
Rejection,
|
||||||
|
ValidatedProposal,
|
||||||
|
classify_codes,
|
||||||
|
has_requirement_form,
|
||||||
|
validate_proposal,
|
||||||
|
)
|
||||||
|
|
||||||
|
_DEFAULT_BUNDLE_ROOT = Path.home() / "repos" / "vegnormal-okf" / "build" / "ferdig"
|
||||||
|
_ROUND3 = Path("scratchpad/p19-stress/tunnel-hauglia-2027")
|
||||||
|
_P17B = Path("scratchpad/p17b-multibase/lindaas")
|
||||||
|
|
||||||
|
|
||||||
|
def _base(name: str) -> Path:
|
||||||
|
root = Path(os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", str(_DEFAULT_BUNDLE_ROOT)))
|
||||||
|
if not (root / name).is_dir():
|
||||||
|
pytest.skip(f"knowledge base {name!r} is not mounted under {root}")
|
||||||
|
return root / name
|
||||||
|
|
||||||
|
|
||||||
|
def _grounding_over(name: str) -> Grounding:
|
||||||
|
"""The delivered base exactly as ``run_project`` composes it — documents AND vocabulary."""
|
||||||
|
bundle = okf.navigate_bundle(str(_base(name)))
|
||||||
|
return Grounding(
|
||||||
|
documents=tuple(
|
||||||
|
"\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files
|
||||||
|
),
|
||||||
|
declared_references=tuple(
|
||||||
|
ref for f in bundle.context_files for ref in okf.declared_reference_numbers(f)
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _recorded(path: Path) -> SavingsProposal:
|
||||||
|
if not path.is_file():
|
||||||
|
pytest.skip(f"the recorded artefact {path} is not present in this checkout")
|
||||||
|
return SavingsProposal.model_validate(json.loads(path.read_text(encoding="utf-8"))["proposal"])
|
||||||
|
|
||||||
|
|
||||||
|
def _proposal(code: str, *, saving: float = 1000.0) -> SavingsProposal:
|
||||||
|
return SavingsProposal(
|
||||||
|
project_id="p",
|
||||||
|
measure="m",
|
||||||
|
affected_items=[{"code": code, "quantity": 10.0, "unit_cost": 1000.0}],
|
||||||
|
claimed_saving_nok=saving,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------- known positives
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_section_number_that_reached_validated_on_n500_is_refused() -> None:
|
||||||
|
"""(a) tunnel-04's ``10.4``, replayed against the base that run actually opened."""
|
||||||
|
proposal = _recorded(_ROUND3 / "tunnel-hauglia-2027-04-a4-enhetspris-ventilator-proposal.json")
|
||||||
|
assert [i.code for i in proposal.affected_items] == ["10.4"]
|
||||||
|
outcome = validate_proposal(proposal, baseline=None, grounding=_grounding_over("n500-2024"))
|
||||||
|
assert isinstance(outcome, Rejection)
|
||||||
|
assert "'10.4'" in outcome.reason
|
||||||
|
assert "not one of the 365 this knowledge base declares" in outcome.reason
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_process_number_that_reached_validated_on_r761_is_refused() -> None:
|
||||||
|
"""(b) lindaas a4's ``1.10.4`` — a chapter of ANOTHER standard, quoted in one r761 document."""
|
||||||
|
proposal = _recorded(_P17B / "lindaas-01-vegnormal-r761-2025-a4-indeksregulering-proposal.json")
|
||||||
|
assert [i.code for i in proposal.affected_items] == ["1.10.4"]
|
||||||
|
outcome = validate_proposal(proposal, baseline=None, grounding=_grounding_over("r761-2025"))
|
||||||
|
assert isinstance(outcome, Rejection)
|
||||||
|
assert "not one of the 2765 this knowledge base declares" in outcome.reason
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------- known negatives
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_process_code_the_base_declares_still_validates() -> None:
|
||||||
|
"""(c) The arm that keeps this a rule about the corpus and not about shapes.
|
||||||
|
|
||||||
|
``12.1`` is ``contexts/kontrakt-sorasen-2027``'s own first code and a REAL declared
|
||||||
|
``prosessnr`` of R761. Under the ordered rule it would have been refused; it must not be.
|
||||||
|
"""
|
||||||
|
grounding = _grounding_over("r761-2025")
|
||||||
|
assert "12.1" in grounding.reference_vocabulary
|
||||||
|
outcome = validate_proposal(_proposal("12.1"), baseline=None, grounding=grounding)
|
||||||
|
assert isinstance(outcome, ValidatedProposal), getattr(outcome, "reason", "")
|
||||||
|
|
||||||
|
|
||||||
|
def test_every_sorasen_code_is_in_the_bases_vocabulary() -> None:
|
||||||
|
"""(c) The whole context set, not one sample: five real codes, five declared numbers."""
|
||||||
|
codes = [
|
||||||
|
code
|
||||||
|
for approach in json.loads(
|
||||||
|
Path("contexts/kontrakt-sorasen-2027/mandate.json").read_text(encoding="utf-8")
|
||||||
|
)["approaches"]
|
||||||
|
for code in approach.get("affected_codes", [])
|
||||||
|
]
|
||||||
|
vocabulary = _grounding_over("r761-2025").reference_vocabulary
|
||||||
|
shaped = [c for c in codes if has_requirement_form(c)]
|
||||||
|
assert len(shaped) == 5, shaped
|
||||||
|
assert [c for c in shaped if c not in vocabulary] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_fasit_references_are_classified_requirement() -> None:
|
||||||
|
"""(g) known negative (c) of the order: a fasit reference IS a requirement, and says so."""
|
||||||
|
fasit = json.loads(
|
||||||
|
Path("contexts/kontrakt-sorasen-2027/fasit.json").read_text(encoding="utf-8")
|
||||||
|
)
|
||||||
|
refs = sorted({c["ref"] for entry in fasit["must_cite"] for c in entry["concepts"]})
|
||||||
|
forms = classify_codes(refs, _grounding_over("r761-2025"))
|
||||||
|
assert set(forms.values()) == {"requirement"}, forms
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_anchored_run_is_untouched_by_the_rule() -> None:
|
||||||
|
"""(d) Stage 0 has already ruled; the weaker stage must not overrule the stronger."""
|
||||||
|
grounding = Grounding(documents=("12.9 is a clause",), declared_references=("12.1", "12.2"))
|
||||||
|
baseline = CostBaseline(project_id="p", items={"12.9": {"quantity": 10.0, "unit_cost": 1000.0}})
|
||||||
|
outcome = validate_proposal(_proposal("12.9"), baseline=baseline, grounding=grounding)
|
||||||
|
assert isinstance(outcome, ValidatedProposal), getattr(outcome, "reason", "")
|
||||||
|
# The control: the SAME code and the SAME text, unanchored, is refused.
|
||||||
|
unanchored = validate_proposal(_proposal("12.9"), baseline=None, grounding=grounding)
|
||||||
|
assert isinstance(unanchored, Rejection)
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_input_that_declares_no_reference_numbers_cannot_trip_the_rule() -> None:
|
||||||
|
"""(e) The generality guard — and the reason every pre-P20 fixture is untouched."""
|
||||||
|
grounding = Grounding(documents=("a document mentioning 12.9 once",))
|
||||||
|
assert grounding.reference_vocabulary == frozenset()
|
||||||
|
outcome = validate_proposal(_proposal("12.9"), baseline=None, grounding=grounding)
|
||||||
|
assert isinstance(outcome, ValidatedProposal), getattr(outcome, "reason", "")
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_cost_line_identifier_is_not_requirement_shaped() -> None:
|
||||||
|
"""(f) K2's 50 identifiers and this repo's own code: shape, measured."""
|
||||||
|
assert not has_requirement_form("SHA-01")
|
||||||
|
assert not has_requirement_form("ENERGI-TOTAL-EL")
|
||||||
|
assert not has_requirement_form("65 ASFALTDEKKER")
|
||||||
|
assert has_requirement_form("10.4") and has_requirement_form("Krav 4.1.2—1")
|
||||||
|
|
||||||
|
|
||||||
|
def test_classify_codes_without_a_grounding_is_the_pre_p20_answer() -> None:
|
||||||
|
"""(g) ``stress.py`` re-derives for runs written before the field existed."""
|
||||||
|
assert classify_codes(["12.1", "SHA-01", "impulsventilator"]) == {
|
||||||
|
"12.1": "identifier",
|
||||||
|
"SHA-01": "identifier",
|
||||||
|
"impulsventilator": "prose",
|
||||||
|
}
|
||||||
|
grounding = Grounding(documents=("x",), declared_references=("12.1",))
|
||||||
|
assert classify_codes(["12.1"], grounding) == {"12.1": "requirement"}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------- the wiring
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_vocabulary_travels_with_the_text_into_the_generation_gate() -> None:
|
||||||
|
"""(h) ``_grounding_text`` composes the run's three sources; the vocabulary must survive it."""
|
||||||
|
delivered = Grounding(documents=("d",), declared_references=("12.1",))
|
||||||
|
project = Project(
|
||||||
|
id="p", name="n", description="d", currency="NOK", cost_items=(), docs_dir="."
|
||||||
|
)
|
||||||
|
composed = _grounding_text(project, None, delivered)
|
||||||
|
assert composed.reference_vocabulary == frozenset({"12.1"})
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_run_composes_the_vocabulary_from_the_base_it_opened(tmp_path: Path) -> None:
|
||||||
|
"""(i) End-to-end through ``run_project``: the stamp says ``requirement`` for a declared code.
|
||||||
|
|
||||||
|
Asserted on the ARTEFACT a run leaves, never on the composer — a vocabulary wired nowhere would
|
||||||
|
satisfy every arm above and none of this one.
|
||||||
|
"""
|
||||||
|
import asyncio
|
||||||
|
|
||||||
|
from agent_framework import BaseChatClient
|
||||||
|
|
||||||
|
from portfolio_optimiser.run import run_project
|
||||||
|
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||||
|
|
||||||
|
base = tmp_path / "mini"
|
||||||
|
base.mkdir()
|
||||||
|
(base / "index.md").write_text(
|
||||||
|
"---\nbundle_id: mini\n---\n\n- [Krav](krav.md) — one clause.\n", encoding="utf-8"
|
||||||
|
)
|
||||||
|
(base / "krav.md").write_text(
|
||||||
|
"---\ntype: Krav\ntitle: Krav 4.1.2-1\nprosessnr: '12.1'\n---\n\nEn kostlinje 12.1.\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
(base / "validator-input.json").write_text(
|
||||||
|
json.dumps({"project_id": "mini-p", "measure": "m", "affected_codes": ["12.1"]}),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
reply = (
|
||||||
|
'{"measure":"m","affected_items":[{"code":"12.1","quantity":10,"unit_cost":1000}],'
|
||||||
|
'"claimed_saving_nok":1000}'
|
||||||
|
)
|
||||||
|
|
||||||
|
def factory(role: str) -> BaseChatClient:
|
||||||
|
return ScriptedChatClient(
|
||||||
|
"Reasoning holds.\nVERDICT: APPROVE" if role == "checker" else reply, role=role
|
||||||
|
)
|
||||||
|
|
||||||
|
result = asyncio.run(
|
||||||
|
run_project(
|
||||||
|
"mini-p",
|
||||||
|
"local",
|
||||||
|
docs_dir=str(base),
|
||||||
|
bundle_dir=str(base),
|
||||||
|
client_factory=factory,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
assert result.provenance.code_forms == {"12.1": "requirement"}
|
||||||
|
assert result.provenance.validator_decision == "validated"
|
||||||
234
tests/test_right_requirement_loadbearing.py
Normal file
234
tests/test_right_requirement_loadbearing.py
Normal file
|
|
@ -0,0 +1,234 @@
|
||||||
|
"""P20 DEL A — the requirement that is RIGHT, and the commission's criteria reaching the reader.
|
||||||
|
|
||||||
|
**The measured silence, three rounds and one multi-base pass deep.** P19 DEL A made a direction
|
||||||
|
NAME the requirement that binds it and made the run refuse a declaration naming a path it never
|
||||||
|
opened. The declarations then happened — and MEASURED (P19 round 3: 9 declarations over 5 paid
|
||||||
|
runs; P17b: 4 over one two-base pass) **not one of them named a fasit concept**. The rung works;
|
||||||
|
what nothing asked for was that the requirement be the RIGHT one. Two halves were missing:
|
||||||
|
|
||||||
|
* the tool answered ``{"declared": true, ...}`` to every accepted declaration, echoing back the
|
||||||
|
caller's own three arguments. A model that had declared a requirement about something else was
|
||||||
|
told, in the only words it got, that it had succeeded;
|
||||||
|
* the commission's own statement of what a good answer looks like — ``Mandate.success_criteria`` —
|
||||||
|
reached ``announce`` and NOTHING else (P19 F2). It was printed for a human and withheld from the
|
||||||
|
only reader who could act on it.
|
||||||
|
|
||||||
|
What each arm pins:
|
||||||
|
|
||||||
|
(a) A1 — the reply carries the DOCUMENT's own ``title`` and ``req_number``, read off the base, plus
|
||||||
|
the sentence that says what the declaration BINDS. Without the document's own words the reply
|
||||||
|
cannot contradict a wrong declaration, which is the whole correction;
|
||||||
|
(b) A1 — a path the base carries as no navigated concept (``index.md`` is the reachable case)
|
||||||
|
answers with empty strings rather than raising: the read trace has already accepted the
|
||||||
|
declaration, and turning "I cannot restate your title" into a refusal would fail a declaration
|
||||||
|
the run's own evidence proves was read;
|
||||||
|
(c) A1 — the reply is read off ``Bundle.context_files``, so a ``type: verdict`` document can never
|
||||||
|
be named back by title. The one layer no listing names stays unnamed even in an answer;
|
||||||
|
(d) A2 — ``criteria_block`` is the ONE renderer, and it is EMPTY without criteria. That omission is
|
||||||
|
what keeps every prompt of every un-commissioned run byte-identical, the demo's included;
|
||||||
|
(e) A2 — the criteria reach the DEBATE's task message, the prompt where ``declare_requirement`` is
|
||||||
|
available. Asserted on the text the client actually received, never on the composer;
|
||||||
|
(f) A2 — a run with NO mandate sends the task message byte-identically to before, which is the
|
||||||
|
half that keeps ``demo-transcript.stdout`` unchanged and is asserted here rather than left to
|
||||||
|
the golden;
|
||||||
|
(g) A3 — the navigator instruction and the tool description both name the ``read_dir(filter=...)``
|
||||||
|
route with a worked example. A description that lies about the body IS the model's instruction
|
||||||
|
(the Fase-3 class), and here the body is new.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from portfolio_optimiser import explore, okf
|
||||||
|
from portfolio_optimiser.explore import DeclaredRequirement, ToolCall, navigator_tools
|
||||||
|
from portfolio_optimiser.mandate import Approach, Mandate, criteria_block
|
||||||
|
|
||||||
|
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||||
|
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||||
|
_BYGG = _EXAMPLES / "bygg-energi-mikro"
|
||||||
|
_PID = "BYGG-KONTOR-NORD"
|
||||||
|
|
||||||
|
|
||||||
|
def _wired(bundle_dir: Path) -> tuple[dict[str, Any], list[ToolCall], list[DeclaredRequirement]]:
|
||||||
|
opened: list[ToolCall] = []
|
||||||
|
declared: list[DeclaredRequirement] = []
|
||||||
|
tools = navigator_tools((str(bundle_dir),), opened=opened, requirements=declared)
|
||||||
|
return {t.name: t for t in tools}, opened, declared
|
||||||
|
|
||||||
|
|
||||||
|
def _declare(tools: dict[str, Any], opened: list[ToolCall], base: str, path: str) -> dict[str, Any]:
|
||||||
|
"""Read it the way a run does, then declare it — the (b) path of P19 DEL A."""
|
||||||
|
opened.append(ToolCall(name="read_file", bundle_id=base, path=path))
|
||||||
|
return tools["declare_requirement"].func(bundle_id=base, path=path, ref="Krav 4.1.2-1")
|
||||||
|
|
||||||
|
|
||||||
|
def _base_with_a_requirement(root: Path) -> Path:
|
||||||
|
"""A base declaring its own ``req_number`` — the form the delivered N corpora carry.
|
||||||
|
|
||||||
|
Crafted rather than taken from ``shared/examples``: MEASURED, no example base declares a
|
||||||
|
``req_number`` at all, so an arm built on one of them could not tell a reply that reads the
|
||||||
|
document from one that returns an empty string.
|
||||||
|
"""
|
||||||
|
base = root / "n-mini"
|
||||||
|
base.mkdir()
|
||||||
|
(base / "index.md").write_text(
|
||||||
|
"---\nbundle_id: n-mini\n---\n\n- [Krav](krav.md) — one requirement.\n", encoding="utf-8"
|
||||||
|
)
|
||||||
|
(base / "krav.md").write_text(
|
||||||
|
"---\ntype: Krav\ntitle: Krav 4.1.2-1 Rundkjoring\nreq_number: Krav 4.1.2-1\n"
|
||||||
|
"seksjon: '4.1.2'\n---\n\nEn rundkjoring skal ha …\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
return base
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# (a)-(c) the reply the model can be contradicted by
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_reply_carries_the_documents_own_title_and_number(tmp_path: Path) -> None:
|
||||||
|
"""(a) The base's words, not the caller's — the only thing that can say "wrong one"."""
|
||||||
|
base = _base_with_a_requirement(tmp_path)
|
||||||
|
tools, opened, declared = _wired(base)
|
||||||
|
answer = _declare(tools, opened, "n-mini", "krav.md")
|
||||||
|
assert answer["title"] == "Krav 4.1.2-1 Rundkjoring"
|
||||||
|
assert answer["req_number"] == "Krav 4.1.2-1"
|
||||||
|
assert "is the requirement the proposal rests on" in answer["binds"]
|
||||||
|
assert declared == [DeclaredRequirement(bundle_id="n-mini", path="krav.md", ref="Krav 4.1.2-1")]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_path_that_is_no_concept_answers_with_empty_strings(tmp_path: Path) -> None:
|
||||||
|
"""(b) Accepted-but-unrestatable is not a refusal; the read trace already ruled."""
|
||||||
|
base = _base_with_a_requirement(tmp_path)
|
||||||
|
tools, opened, _declared = _wired(base)
|
||||||
|
answer = _declare(tools, opened, "n-mini", "index.md")
|
||||||
|
assert answer["declared"] is True
|
||||||
|
assert (answer["title"], answer["req_number"]) == ("", "")
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_verdict_document_is_never_named_back_by_title(tmp_path: Path) -> None:
|
||||||
|
"""(c) Read off ``context_files``, the property that drops the ``type: verdict`` layer."""
|
||||||
|
base = _base_with_a_requirement(tmp_path)
|
||||||
|
(base / "dom.md").write_text(
|
||||||
|
"---\ntype: verdict\ntitle: A PRIOR EXPERT JUDGEMENT\nid: v1\n---\n\napproved.\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
tools, opened, _declared = _wired(base)
|
||||||
|
answer = _declare(tools, opened, "n-mini", "dom.md")
|
||||||
|
assert answer["declared"] is True
|
||||||
|
assert answer["title"] == "", "the verdict layer must not be readable through this answer"
|
||||||
|
# The control: the SAME base answers the concept's title, so the empty string above is the
|
||||||
|
# layer being excluded rather than the reader being broken.
|
||||||
|
assert _declare(tools, opened, "n-mini", "krav.md")["title"] == "Krav 4.1.2-1 Rundkjoring"
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_tunnel_example_answers_its_own_title_and_no_number() -> None:
|
||||||
|
"""(a) control on a REAL example base: title present, reference number honestly absent."""
|
||||||
|
tools, opened, _declared = _wired(_TUNNEL)
|
||||||
|
path = okf.navigate_bundle(str(_TUNNEL)).context_files[0].name
|
||||||
|
answer = _declare(tools, opened, "tunnel-hauglia", path)
|
||||||
|
assert answer["title"].startswith("Kilder: tunnelbelysning")
|
||||||
|
assert answer["req_number"] == "", "this base declares none, and the reply must not invent one"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# (d)-(f) the criteria reaching the prompt
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_criteria_renderer_is_empty_without_criteria() -> None:
|
||||||
|
"""(d) Omission, never an empty heading — the ``announce`` rule, and the golden's guarantee."""
|
||||||
|
assert criteria_block("") == ""
|
||||||
|
rendered = criteria_block("a saving the price schedule can carry")
|
||||||
|
assert "a saving the price schedule can carry" in rendered
|
||||||
|
assert rendered.startswith("\n")
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_commissions_criteria_reach_the_debate_task(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||||
|
"""(e) Asserted on what the CLIENT received, never on the composer."""
|
||||||
|
sent = _run_and_capture_task(
|
||||||
|
monkeypatch,
|
||||||
|
mandate=Mandate(
|
||||||
|
objective="cut cost",
|
||||||
|
approaches=(Approach(id="a1", label="L", description="D"),),
|
||||||
|
success_criteria="SENTINEL-CRITERION-ONLY-THE-COMMISSION-CARRIES",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
assert "SENTINEL-CRITERION-ONLY-THE-COMMISSION-CARRIES" in sent[0]
|
||||||
|
assert "What the commissioner counts as success" in sent[0]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_run_without_a_mandate_sends_the_task_unchanged(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||||
|
"""(f) The byte-identity half. Without it (e) could be satisfied by always appending."""
|
||||||
|
sent = _run_and_capture_task(monkeypatch, mandate=None)
|
||||||
|
assert sent[0].startswith(f"Find a cost-saving measure for {_PID}.\nContext:\n")
|
||||||
|
|
||||||
|
|
||||||
|
def _run_and_capture_task(monkeypatch: pytest.MonkeyPatch, *, mandate: Mandate | None) -> list[str]:
|
||||||
|
"""Drive a real ``run_project`` and capture every prompt the debate's clients received.
|
||||||
|
|
||||||
|
The same recording-factory seam ``test_debate_navigation_cost_loadbearing`` measures S2c with:
|
||||||
|
the assertion is on what the CLIENT was handed, never on ``criteria_block``'s return value, so
|
||||||
|
a renderer wired nowhere cannot satisfy it.
|
||||||
|
"""
|
||||||
|
import asyncio
|
||||||
|
|
||||||
|
from agent_framework import BaseChatClient
|
||||||
|
|
||||||
|
from portfolio_optimiser.run import run_project
|
||||||
|
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||||
|
|
||||||
|
sink: list[str] = []
|
||||||
|
valid = (
|
||||||
|
'{"measure":"LED-retrofit","affected_items":'
|
||||||
|
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],'
|
||||||
|
'"claimed_saving_nok":30000}'
|
||||||
|
)
|
||||||
|
|
||||||
|
def factory(role: str) -> BaseChatClient:
|
||||||
|
client = ScriptedChatClient(
|
||||||
|
"Reasoning holds.\nVERDICT: APPROVE" if role == "checker" else valid, role=role
|
||||||
|
)
|
||||||
|
original = client._inner_get_response # type: ignore[attr-defined]
|
||||||
|
|
||||||
|
def recording(*, messages, options, stream=False, **kwargs): # type: ignore[no-untyped-def]
|
||||||
|
sink.append("\n".join(getattr(m, "text", "") or "" for m in messages))
|
||||||
|
return original(messages=messages, options=options, stream=stream, **kwargs)
|
||||||
|
|
||||||
|
client._inner_get_response = recording # type: ignore[attr-defined,method-assign]
|
||||||
|
return client
|
||||||
|
|
||||||
|
asyncio.run(
|
||||||
|
run_project(
|
||||||
|
_PID,
|
||||||
|
"local",
|
||||||
|
docs_dir=str(_BYGG),
|
||||||
|
bundle_dir=str(_BYGG),
|
||||||
|
verdict_input={"decision": "approved", "rationale": "expert reviewed (test)"},
|
||||||
|
client_factory=factory,
|
||||||
|
mandate=mandate,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
assert sink, "the debate never reached the client"
|
||||||
|
return sink
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# (g) the instruction that describes the new body
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_both_descriptions_name_the_filter_route() -> None:
|
||||||
|
"""(g) A description that lies about the body IS the instruction (the Fase-3 class)."""
|
||||||
|
instruction = explore._INSTRUCTIONS[explore.HYPOTHESISER_ROLE]
|
||||||
|
assert "read_dir a 'filter' word" in instruction
|
||||||
|
assert "Krav 4.1.2-1" in instruction, "the worked example is what makes the route concrete"
|
||||||
|
tools = {t.name: t for t in navigator_tools((str(_TUNNEL),), opened=[], requirements=[])}
|
||||||
|
description = tools["declare_requirement"].description or ""
|
||||||
|
assert "filter='rundkjoring'" in description
|
||||||
|
assert "own title and number" in description
|
||||||
Loading…
Add table
Add a link
Reference in a new issue