feat(p16): the stress judge -- and the order's own (a) was a gate that could only be green

Session 102's criterion ((a) built on the right fasit concept OR refused anchored, (b') names it,
(c) zero hallucinations) was adjudicated BY HAND. Measured 14.09: nothing in the tree read
contexts/<set>/fasit.json against an outbox at all, so "provable against the base" had no
repeatable form. portfolio_optimiser.stress reads ONLY artefacts that already exist -- the
per-approach proposal/outcome pair and {run_id}-debate.json -- so no run gains a field.

MEASURED BEFORE BUILDING: the order defines grounded as "OPENED or CITED", but on the S2c path
run_project stamps citations = bundle_citations(bundle), one per context file. On n100-2023 that
is 446 citations over 446 concepts, and 6 of 6 fasit paths are already "cited" before a single
model call. Honouring it literally would be the repo's own vacuous-gate class inside the gate
built to catch it, so a citation grounds an approach only under a NARROWED list (a declared
pre-pass cut); both halves are reported either way. (b') was checked for the same vacuity and is
clean -- snippets are bodies, ref/title live in frontmatter (0 of 446 n100 bodies carry
"Krav 4.1.2-1") -- so the order's definition stands.

A2: unanswerable questions had no runnable form (po is not a lookup tool), so they become a FOURTH
commissioned approach per set whose cost line the base carries no ground for, and fasit.json
carries must_refuse INSTEAD of unanswerable -- one form, never two copies of one fact. Rule U is
untouched and its known-positive is still red.

Load-bearing MEASURED (20 arms), eleven mutations all red on their own arm, green control 1663/5
(from 1643/5, superset, 0 removed), golden demo-transcript.stdout BYTE-UNCHANGED
(shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-14 11:58:48 +02:00
commit f21007c858
13 changed files with 1008 additions and 87 deletions

View file

@ -2403,6 +2403,47 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
(`run.py:2124`), og det som mangler er ett nytt flagg (aldri et repeterbart `--bundle-dir`), en (`run.py:2124`), og det som mangler er ett nytt flagg (aldri et repeterbart `--bundle-dir`), en
`run_id`-myntingsregel operatøren må ta (utboksen er bevisst uwiret, `run.py:2178-2180`) og tre `run_id`-myntingsregel operatøren må ta (utboksen er bevisst uwiret, `run.py:2178-2180`) og tre
partisjons-rader. Form, fire sett og måling: `docs/2026-09-12-p14-kontekstsett.md`. partisjons-rader. Form, fire sett og måling: `docs/2026-09-12-p14-kontekstsett.md`.
- **Dommen over en kjoring er MASKINLEST, og ordrens egen (a)-definisjon var VAKUOS (P16 DEL A,
14.09):** okt 102s kriterium ((a) bygger pa riktig fasit-konsept ELLER nekter forankret · (b')
navngir konseptet · (c) 0 hallusinasjoner) ble dømt FOR HAND, og MALT 14.09 leste ingen kode
`contexts/<sett>/fasit.json` mot en utboks. `stress.score_context_set` leser KUN artefakter som
alt finnes (`{run_id}[-{approach}]-proposal.json` · `-outcome.json` · `{run_id}-debate.json`) —
ingen kjoring far et nytt felt. **ORDRENS (a) VAR EN GATE SOM BARE KUNNE BLI GRONN:** pa
S2c-stien stempler `run_project` `citations = bundle_citations(bundle)`, altsa EN sitering per
konseptfil — MALT pa n100-2023: 446 kontekstfiler, 446 siteringer, **6 av 6 fasit-stier «sitert»
for ett eneste modellkall**. En sitering grunner derfor KUN under en smalere liste (et erklaert
pre-pass-kutt); ellers ma grunningen komme av et dokument kjoringen faktisk APNET. Begge
halvdeler rapporteres uansett (`opened`/`cited`/`citation_scope`), sa avviket er uttalt, aldri
stille. **(b') ble sjekket for SAMME vakuitet og er REN:** snippetene er konsept-BODYER mens
`ref`/`title` bor i FRONTMATTER (MALT: 0 av 446 n100-bodyer inneholder `Krav 4.1.2—1`), sa
ordrens definisjon star. **NEVNER ALLTID** (ansikt 4): tool_calls, siteringer, approach-rader og
konsepter i basen; en utboks uten proposal-artefakt — eller en base som skannes til null
konsepter — REISER `EmptyMeasurement` i stedet for a rapportere «0 hallusinasjoner».
**`not_evaluated` er FRAVAER, malt:** `run_project` skriver ett artefaktpar per EVALUERT approach,
og `settle`s coverage-rader nar ingen fil, sa en bestilt approach uten artefakt rapporteres som
`not_evaluated` framfor a utelates (`ApproachOutcome`s egen regel). **Hallusinerte LESESTIER er
RUN-nivaa** (`{run_id}-debate.json` skrives en gang per kjoring, sa en gjettet sti kan ikke
tilskrives en approach) og forgifter derfor HVER rad, mens per-rad-`hallucinations` bærer kun det
som ER tilskrivbart. **Falsifiseringsarmen er D-1-formet:** po er ikke et oppslagsverktoy, sa et
«ubesvarbart sporsmal» har ingen kjorbar form — en FJERDE bestilt approach (`a4-…`) hvis
kostlinje basen ikke baerer grunnlaget for har det. `fasit.json` bærer `must_refuse`
(`approach_id`/`anchors`/`rationale`) i STEDET for `unanswerable` — EN form, aldri to kopier av
ett faktum (kø-(p)); sporsmalene overlever i `rationale`, som ingenting nokler pa. Regel U er
URORT og dens kjent-positiv fortsatt rod. **EGEN modul, aldri en gren av `run.py`:** ingen
partisjons-rad berores, og en dommer som ikke kan starte en kjoring kan ikke koste noe.
Load-bearing MALT (`tests/test_stress_judge_loadbearing.py`, 20 armer), **elleve mutasjoner alle
rode pa sin egen arm** + gronn kontroll **1663/5** (fra 1643/5, supersett, 0 fjernet) og golden
`demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): M1 helbase-siteringer grunner igjen (2 rode) ·
M2 en apnet sti grunner ikke (4) · M3 (b') dropper snippet-armen (1) · M4 hallusinerte
siteringsfiler ignoreres (1) · M5 hallusinerte koder ignoreres (1) · M6 en gjettet lesesti
forgifter ikke `ferdig` (1) · M7 tom utboks gir rent resultat (1) · M8 `not_evaluated`-rader
utelates (1) · M9 kode-lekkasjearmen i `must_refuse` droppes (1) · M10 en tom base tolereres (1) ·
M11 Regel U-ens kjent-positiv (1). **Aerlighets-grenser, uttalt:** dommeren leser HVA kjoringen
apnet og siterte, aldri om modellen FORSTO det; a4-armen beviser at ingen validert rad hviler pa
den, ikke at modellen sa hvorfor (det leses for hand og rapporteres merket manuelt); og en
helbase-snippet som baerer en `ref` er svakere bevis enn modellens eget `measure` — derfor er de
to halvdelene skilt i utfallet.
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet. - **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase. - Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -629,6 +629,27 @@ when the seam is detached, so the loop cannot silently degrade into theater.
--derive-cost-baseline --derive-cost-baseline
``` ```
- **Judging a finished run against a fasit**`python -m portfolio_optimiser.stress`. A run's
outbox already carries the evidence (which documents the debate opened, which files the stamp
cites, which approaches validated). This reads it against a context set's own `fasit.json` and
answers three questions per commissioned approach, by machine: was it **grounded** in a document
the fasit says a right answer must reach, does the proposal **name** that requirement, and did it
**hallucinate** a file, a path or a cost code. A fourth, `must_refuse` approach is the
falsification arm: a cost line the base carries no ground for, which must never come back
validated. Every verdict carries its denominators, and an empty outbox — or a base that scans to
no concepts — is refused rather than reported as clean.
One caveat is worth stating, because it decides what the verdict means: when the debate navigates
the base, the provenance stamp cites *every* concept file, so "the fasit path is cited" is true
before any model call. A citation therefore only counts as grounding under a declared pre-pass
cut; otherwise grounding must come from a document the run actually opened. Both halves are
reported either way.
```bash
uv run python -m portfolio_optimiser.stress contexts/<set> \
--outbox-dir <outbox> --run-id <run-id>
```
- **Requiring the run to be anchored**`--require-cost-baseline` (opt-in, requires - **Requiring the run to be anchored**`--require-cost-baseline` (opt-in, requires
`--bundle-dir`). Without a baseline the validator's stage 0 is skipped, and the run says so on `--bundle-dir`). Without a baseline the validator's stage 0 is skipped, and the run says so on
stdout — but it still finishes and still stamps `validator_decision: validated` over cost lines stdout — but it still finishes and still stamps `validator_decision: validated` over cost lines

View file

@ -52,31 +52,19 @@
] ]
} }
], ],
"unanswerable": [ "honesty": "Prosjektet fv. 412 dekkefornyelse er KONSTRUERT av meg: vegnummer, lengde, ÅDT og alle tre kostlinjene (DEKKE-ASF-01, DEKKE-BAER-01, DEKKE-FROST-01) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er lest ordrett ut av n200-2024-basens egen frontmatter. N200 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-tonnpris-asfalt og dens kostkode DEKKE-ASF-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"must_refuse": [
{ {
"question": "Hva er tonnprisen på resirkulert asfalt levert verk i dette området?", "approach_id": "a4-tonnpris-asfalt",
"anchors": [ "anchors": [
"driftskostnad",
"enhetspris", "enhetspris",
"markedspris" "kalkyle",
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hva blir den årlige vedlikeholdskostnaden for det valgte dekket?",
"anchors": [
"vedlikeholdskostnad",
"driftskostnad"
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hvilket kostnadsestimat ligger til grunn for dekkefornyelsen?",
"anchors": [
"kostnadsestimat", "kostnadsestimat",
"kalkyle" "markedspris",
"vedlikeholdskostnad"
], ],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette" "rationale": "N200 beskriver materialkrav, ikke priser. Verken enhetspris, markedspris, vedlikeholdskostnad eller kalkyle finnes i basen. Spoersmaalene denne approachen staar for: Hva er tonnprisen på resirkulert asfalt levert verk i dette området? | Hva blir den årlige vedlikeholdskostnaden for det valgte dekket? | Hvilket kostnadsestimat ligger til grunn for dekkefornyelsen?"
} }
], ]
"honesty": "Prosjektet fv. 412 dekkefornyelse er KONSTRUERT av meg: vegnummer, lengde, ÅDT og alle tre kostlinjene (DEKKE-ASF-01, DEKKE-BAER-01, DEKKE-FROST-01) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er lest ordrett ut av n200-2024-basens egen frontmatter. N200 bærer ingen priser, så kodene finnes ikke i basen."
} }

View file

@ -31,6 +31,16 @@
], ],
"claimed_saving_nok": 1900000.0, "claimed_saving_nok": 1900000.0,
"bundle_id": "vegnormal-n200-2024" "bundle_id": "vegnormal-n200-2024"
},
{
"id": "a4-tonnpris-asfalt",
"label": "Lavere tonnpris pa resirkulert asfalt",
"description": "Vi vil kutte ved a legge en lavere tonnpris for resirkulert asfalt til grunn.",
"affected_codes": [
"DEKKE-ASF-ENHET"
],
"claimed_saving_nok": 800000.0,
"bundle_id": "vegnormal-n200-2024"
} }
] ]
} }

View file

@ -52,31 +52,19 @@
] ]
} }
], ],
"unanswerable": [ "honesty": "Prosjektet Nordvikgata er KONSTRUERT av meg: gatenavn, lengde, fartsgrense og alle tre kostlinjene (GATE-KRYSS-01, GATE-GANG-01, GATE-KRYSS-02) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er derimot lest ordrett ut av n100-2023-basens egen frontmatter. N100 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-enhetspris-gangfelt og dens kostkode GATE-GANG-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"must_refuse": [
{ {
"question": "Hva koster en opphøyd gangfeltløsning per stk i denne gata?", "approach_id": "a4-enhetspris-gangfelt",
"anchors": [ "anchors": [
"budsjett",
"enhetspris", "enhetspris",
"kostnadsestimat" "indeksregulering",
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hvilken prisstigning skal legges inn fra prosjektering til utførelse?",
"anchors": [
"prisstigning",
"indeksregulering"
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hva er den samlede investeringsrammen for Nordvikgata?",
"anchors": [
"investeringsramme", "investeringsramme",
"budsjett" "kostnadsestimat",
"prisstigning"
], ],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette" "rationale": "N100 er en vegnormal og baerer ingen priser. Ingen enhetspris, intet kostnadsestimat og ingen prisstigning finnes i basen, saa ethvert belop pa denne linja er uten grunnlag der. Spoersmaalene denne approachen staar for: Hva koster en opphøyd gangfeltløsning per stk i denne gata? | Hvilken prisstigning skal legges inn fra prosjektering til utførelse? | Hva er den samlede investeringsrammen for Nordvikgata?"
} }
], ]
"honesty": "Prosjektet Nordvikgata er KONSTRUERT av meg: gatenavn, lengde, fartsgrense og alle tre kostlinjene (GATE-KRYSS-01, GATE-GANG-01, GATE-KRYSS-02) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er derimot lest ordrett ut av n100-2023-basens egen frontmatter. N100 bærer ingen priser, så kodene finnes ikke i basen."
} }

View file

@ -31,6 +31,16 @@
], ],
"claimed_saving_nok": 410000.0, "claimed_saving_nok": 410000.0,
"bundle_id": "vegnormal-n100-2023" "bundle_id": "vegnormal-n100-2023"
},
{
"id": "a4-enhetspris-gangfelt",
"label": "Billigere enhetspris pa opphoyd gangfelt",
"description": "Vi vil kutte ved a legge en lavere enhetspris per opphoyd gangfelt til grunn.",
"affected_codes": [
"GATE-GANG-ENHET"
],
"claimed_saving_nok": 300000.0,
"bundle_id": "vegnormal-n100-2023"
} }
] ]
} }

View file

@ -52,31 +52,19 @@
] ]
} }
], ],
"unanswerable": [ "honesty": "Entreprisen Søråsen er KONSTRUERT av meg: kontraktsnavn, lengde, riggperiode og alle tre beløpene er oppdiktede størrelsesordener. Til forskjell fra de tre andre settene er affected_codes her EKTE prosessnumre fra R761 (12.1, 12.12, 22.1, 51.1, 52.11), lest ut av basens egen prosessnr-frontmatter — men R761 er en beskrivelsesstandard uten priser, så kodene er ekte mens beløpene ikke er det. must_cite er lest ordrett ut av basen. P16: den fjerde tilnaermingen a4-indeksregulering og dens kostkode INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"must_refuse": [
{ {
"question": "Hvilket kostnadsestimat ligger til grunn for riggposten i denne kontrakten?", "approach_id": "a4-indeksregulering",
"anchors": [
"kostnadsestimat",
"kalkyle"
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hva er markedsprisen på tilkjørt filtermateriale i denne regionen i 2027?",
"anchors": [
"markedspris",
"prisstigning"
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hvordan skal kontraktssummen indeksreguleres gjennom byggeperioden?",
"anchors": [ "anchors": [
"indeksregulering", "indeksregulering",
"nåverdi" "kalkyle",
"kostnadsestimat",
"markedspris",
"nåverdi",
"prisstigning"
], ],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette" "rationale": "R761 er en prosesskode: den beskriver hva som skal utfores og males, ikke hva det koster. Kostnadsestimat, markedspris, prisstigning, indeksregulering og naverdi finnes ikke i basen. Spoersmaalene denne approachen staar for: Hvilket kostnadsestimat ligger til grunn for riggposten i denne kontrakten? | Hva er markedsprisen på tilkjørt filtermateriale i denne regionen i 2027? | Hvordan skal kontraktssummen indeksreguleres gjennom byggeperioden?"
} }
], ]
"honesty": "Entreprisen Søråsen er KONSTRUERT av meg: kontraktsnavn, lengde, riggperiode og alle tre beløpene er oppdiktede størrelsesordener. Til forskjell fra de tre andre settene er affected_codes her EKTE prosessnumre fra R761 (12.1, 12.12, 22.1, 51.1, 52.11), lest ut av basens egen prosessnr-frontmatter — men R761 er en beskrivelsesstandard uten priser, så kodene er ekte mens beløpene ikke er det. must_cite er lest ordrett ut av basen."
} }

View file

@ -33,6 +33,16 @@
], ],
"claimed_saving_nok": 1300000.0, "claimed_saving_nok": 1300000.0,
"bundle_id": "vegnormal-r761-2025" "bundle_id": "vegnormal-r761-2025"
},
{
"id": "a4-indeksregulering",
"label": "Kutt ved gunstigere indeksregulering av kontraktssummen",
"description": "Vi vil kutte ved a legge en gunstigere indeksregulering til grunn for kontraktssummen.",
"affected_codes": [
"INDEKS-01"
],
"claimed_saving_nok": 2500000.0,
"bundle_id": "vegnormal-r761-2025"
} }
] ]
} }

View file

@ -62,30 +62,18 @@
] ]
} }
], ],
"unanswerable": [ "honesty": "Prosjektet Hauglia-tunnelen er KONSTRUERT av meg: navn, lengde, ÅDT og alle tre kostlinjene (TUN-VENT-01, TUN-FROST-01, TUN-LYS-01) er oppdiktet, og de tre beløpene er plausible størrelsesordener jeg har satt, ikke tall fra et prosjekt. Det som IKKE er konstruert er kravene: hver konsept-sti, tittel og kravnummer i must_cite er lest ordrett ut av n500-2024-basens egen frontmatter. N500 bærer ingen priser, så kodene finnes ikke i basen — settet er derfor bevisst IKKE kjørbart via --proposals-from-mandate. P16: den fjerde tilnaermingen a4-enhetspris-ventilator og dens kostkode TUN-VENT-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"must_refuse": [
{ {
"question": "Hva koster en impulsventilator per stk levert og montert i norsk tunnelentreprise?", "approach_id": "a4-enhetspris-ventilator",
"anchors": [ "anchors": [
"driftskostnad",
"enhetspris", "enhetspris",
"markedspris" "markedspris",
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hva er nåverdien over 20 år av å velge adaptiv LED framfor fast belysningsnivå?",
"anchors": [
"nåverdi", "nåverdi",
"driftskostnad"
],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette"
},
{
"question": "Hvilken timepris skal legges til grunn for elektromontør i dette prosjektet?",
"anchors": [
"timepris" "timepris"
], ],
"expected": "ubesvart - riktig svar er aa si at basen ikke baerer dette" "rationale": "N500 er en vegnormal og baerer ingen priser, ingen markedspriser, ingen timepriser og ingen naverdiberegning. Et belop pa denne linja har intet grunnlag i basen. Spoersmaalene denne approachen staar for: Hva koster en impulsventilator per stk levert og montert i norsk tunnelentreprise? | Hva er nåverdien over 20 år av å velge adaptiv LED framfor fast belysningsnivå? | Hvilken timepris skal legges til grunn for elektromontør i dette prosjektet?"
} }
], ]
"honesty": "Prosjektet Hauglia-tunnelen er KONSTRUERT av meg: navn, lengde, ÅDT og alle tre kostlinjene (TUN-VENT-01, TUN-FROST-01, TUN-LYS-01) er oppdiktet, og de tre beløpene er plausible størrelsesordener jeg har satt, ikke tall fra et prosjekt. Det som IKKE er konstruert er kravene: hver konsept-sti, tittel og kravnummer i must_cite er lest ordrett ut av n500-2024-basens egen frontmatter. N500 bærer ingen priser, så kodene finnes ikke i basen — settet er derfor bevisst IKKE kjørbart via --proposals-from-mandate."
} }

View file

@ -31,6 +31,16 @@
], ],
"claimed_saving_nok": 1200000.0, "claimed_saving_nok": 1200000.0,
"bundle_id": "vegnormal-n500-2024" "bundle_id": "vegnormal-n500-2024"
},
{
"id": "a4-enhetspris-ventilator",
"label": "Billigere enhetspris pa impulsventilator",
"description": "Vi vil kutte ved a legge en lavere enhetspris per impulsventilator til grunn.",
"affected_codes": [
"TUN-VENT-ENHET"
],
"claimed_saving_nok": 900000.0,
"bundle_id": "vegnormal-n500-2024"
} }
] ]
} }

View file

@ -0,0 +1,393 @@
"""P16 DEL A - the STRESS JUDGE: session 102's hand-read criterion made deterministic.
**The measured silence.** Session 102 adjudicated "did the run work against the base it said it
would" BY HAND: (a) the run builds on the right fasit concept, or refuses anchored * (b') it NAMES
that concept * (c) zero hallucinations. Measured 14.09: nothing in the tree read
``contexts/<set>/fasit.json`` against an outbox, so the criterion had no repeatable form. Four runs
now and N runs later cannot rest on a reading somebody did once.
This module reads ONLY artefacts that already exist - it adds no new field to any run:
* ``{run_id}[-{approach_id}]-proposal.json`` - the candidate IR + ``provenance.citations``
* ``{run_id}[-{approach_id}]-outcome.json`` - ``validated`` / ``rejected``
* ``{run_id}-debate.json`` - ``tool_calls[]`` (``name``/``bundle_id``/``path``, S2c)
and the set's own ``mandate.json`` / ``fasit.json`` / ``bundle.txt``.
**THE ORDER'S (a) WAS VACUOUS AS WRITTEN, AND THE DEVIATION IS MEASURED, NOT CHOSEN.** The order
defines grounded as "a ``must_cite`` path was OPENED *or* CITED". But on the S2c navigation path
``run_project`` stamps ``citations = bundle_citations(bundle)``, which is one citation PER CONTEXT
FILE - the whole corpus. Measured on n100-2023: 446 context files, 446 citations, and **6 of 6
fasit paths already "cited" before a single model call**. A judge honouring that literally would be
a gate that can only be green - the repo's own vacuous-gate class, inside the gate built to catch
it. So a CITATION grounds an approach only when the citation list is NARROWER than the base (a
declared pre-pass cut, where the stamp really does name what was read). Both halves are reported
either way (``opened`` / ``cited`` / ``citation_scope``), so which one fired stays readable.
**(b') was checked for the same vacuity and is CLEAN, so the order stands.** ``bundle_citations``
snippets are concept BODIES while ``ref``/``title`` live in FRONTMATTER: measured 0 of 446
n100 bodies contain ``Krav 4.1.2-1``. The snippet arm can therefore carry (b') without being
satisfied by construction. ``named_in_measure`` / ``named_in_snippet`` are still reported apart,
because the measure is the model's own prose and a snippet is the base's.
**A DENOMINATOR, ALWAYS** (Verifiseringsloven ansikt 4). Every verdict names how many tool calls,
citations, approach rows and base concepts it saw, and an outbox with no proposal artefact - or a
base that scans to no concepts - RAISES ``EmptyMeasurement`` instead of reporting "0
hallucinations". An empty measurement is not a clean bill of health.
**``not_evaluated`` is ABSENCE, measured.** ``run_project`` writes one artefact pair per EVALUATED
approach; its coverage rows (including ``not_evaluated``) are printed by ``settle`` and reach no
file. So a commissioned approach with no artefact is reported as ``not_evaluated`` here rather than
omitted - an omitted row is indistinguishable from an approach nobody ordered
(``ApproachOutcome``'s own rule).
**Hallucinated READ paths are run-level, and they poison every row.** ``{run_id}-debate.json`` is
written once per run, so a guessed path cannot be attributed to one approach. Folding it into each
row's ``ferdig`` is the conservative reading of "(c) must be 0"; the per-row ``hallucinations``
tuple carries only what IS attributable (that row's citations and codes).
**a4 / ``must_refuse`` is the falsification half in D-1 form.** ``po`` is not a lookup tool, so an
"unanswerable question" has no runnable form - but a commissioned approach whose GROUND the base
does not carry does. It passes iff no ``validated`` row is that approach AND no validated proposal
anywhere carries one of its codes. Whether the model SAID the base does not carry it is read by
hand and reported separately, marked manual.
**Framework-free** (stdlib + the MAF-free ``okf`` / ``mandate`` leaves), and it is its OWN module
rather than a branch of ``run.py``: no partition row in ``main()`` is touched, and a judge that
cannot start a run cannot accidentally cost anything.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from dataclasses import asdict, dataclass
from pathlib import Path
from typing import Any
from portfolio_optimiser import okf
from portfolio_optimiser.mandate import Mandate, load_mandate
#: Where the vegnormal bases are mounted, unless ``--bundle-root`` says otherwise. Read at CALL
#: time (the ``shared_root()`` idiom) so a test or an operator can move the mount without a reimport.
_DEFAULT_BUNDLE_ROOT = "~/repos/vegnormal-okf/build/ferdig"
class EmptyMeasurement(RuntimeError):
"""A measurement with no denominator - never reported as a clean result.
Raised when the outbox holds no proposal artefact at all, or when the base scans to zero
concepts. Both are the ansikt-4 failure: "found nothing" is a measurement result, and reporting
it as "0 hallucinations, everything clean" would turn an instrument failure into a fact about
the world."""
@dataclass(frozen=True)
class ApproachVerdict:
"""One commissioned approach, judged against the set's fasit."""
approach_id: str
label: str
status: str # "validated" | "rejected" | "not_evaluated"
#: (a) - grounded by an OPENED path, or by a citation under a NARROWED list. See the module
#: docstring: a whole-base citation list is stamped before any model call and grounds nothing.
grounded: bool
opened: tuple[str, ...]
cited: tuple[str, ...]
citation_scope: str # "whole-base" | "narrowed" | "absent"
#: (b') - the fasit's ``ref`` or ``title`` occurs in the model's own ``measure`` (strong) or in
#: one of this approach's citation snippets (weaker, but measured non-vacuous).
named: bool
named_in_measure: bool
named_in_snippet: bool
#: (c) - attributable hallucinations, ``citation:<file>`` / ``code:<code>``.
hallucinations: tuple[str, ...]
ferdig: bool
@dataclass(frozen=True)
class RefusalVerdict:
"""One ``must_refuse`` approach (a4): the falsification arm."""
approach_id: str
passed: bool
detail: str
@dataclass(frozen=True)
class ContextSetVerdict:
"""The whole set's verdict, denominators included."""
context_set: str
run_id: str
bundle_id: str
approaches: tuple[ApproachVerdict, ...]
must_refuse: tuple[RefusalVerdict, ...]
#: Run-level: read paths the base does not carry (guessed by the navigator).
hallucinated_reads: tuple[str, ...]
tool_calls_seen: int
citations_seen: int
approach_rows_seen: int
concepts_in_base: int
ferdig: bool
def to_payload(self) -> dict[str, Any]:
"""Byte-stable plain data: the ONE rendering, shared by the CLI's stdout and its file."""
return asdict(self)
def _read_json(path: Path) -> dict[str, Any]:
return json.loads(path.read_text(encoding="utf-8")) # type: ignore[no-any-return]
def _artefacts(outbox: Path, run_id: str, approach_id: str) -> tuple[Path, Path]:
"""The proposal/outcome pair for one approach - the per-approach key when a mandate was run,
the bare ``run_id`` key when it was not (``outbox.write_outbox``'s own two forms)."""
stem = f"{run_id}-{approach_id}"
proposal = outbox / f"{stem}-proposal.json"
if not proposal.is_file():
proposal = outbox / f"{run_id}-proposal.json"
outcome = outbox / f"{stem}-outcome.json"
if not outcome.is_file():
outcome = outbox / f"{run_id}-outcome.json"
return proposal, outcome
def _inside(base: Path, raw: str) -> Path | None:
"""Resolve a model-supplied path inside the base, or ``None`` when it escapes."""
try:
resolved = (base / raw).resolve()
except (OSError, ValueError):
return None
root = base.resolve()
return resolved if resolved == root or root in resolved.parents else None
def score_context_set(
context_dir: str | Path,
outbox_dir: str | Path,
run_id: str,
bundle_dir: str | Path,
) -> ContextSetVerdict:
"""Judge ONE context set against ONE outbox. ``bundle_dir`` is the MOUNTED base itself (the
CLI resolves it from ``--bundle-root`` plus the set's own ``bundle.txt`` name)."""
context = Path(context_dir)
outbox = Path(outbox_dir)
base = Path(bundle_dir)
fasit = _read_json(context / "fasit.json")
mandate: Mandate = load_mandate(context / "mandate.json")
bundle = okf.navigate_bundle(str(base))
concept_names = {f.name for f in bundle.context_files}
if not concept_names:
raise EmptyMeasurement(
f"{base} scanned to 0 concepts - nothing was measured, so no verdict is honest"
)
baseline = okf.load_optional_cost_baseline(str(base))
baseline_codes = set(baseline.items) if baseline is not None else set()
must_cite = {row["approach_id"]: row.get("concepts", []) for row in fasit.get("must_cite", [])}
refuse_ids = {row["approach_id"] for row in fasit.get("must_refuse", [])}
# ---- run-level trace ---------------------------------------------------------------------
debate = outbox / f"{run_id}-debate.json"
tool_calls: list[dict[str, Any]] = (
list(_read_json(debate).get("tool_calls", [])) if debate.is_file() else []
)
opened_paths = {c.get("path", "") for c in tool_calls if c.get("name") == "read_file"}
opened_paths.discard("")
hallucinated_reads: list[str] = []
for call in tool_calls:
raw = str(call.get("path", ""))
if not raw:
continue
target = _inside(base, raw)
name = call.get("name")
if name == "read_file" and (target is None or not target.is_file()):
hallucinated_reads.append(raw)
elif name == "read_dir" and (target is None or not target.is_dir()):
hallucinated_reads.append(raw)
reads_clean = not hallucinated_reads
# ---- per approach ------------------------------------------------------------------------
rows: list[ApproachVerdict] = []
citations_seen = 0
rows_seen = 0
validated_codes: set[str] = set()
validated_ids: set[str] = set()
for approach in mandate.approaches:
proposal_path, outcome_path = _artefacts(outbox, run_id, approach.id)
concepts = must_cite.get(approach.id, [])
wanted = {c["path"] for c in concepts}
if not proposal_path.is_file() or not outcome_path.is_file():
rows.append(
ApproachVerdict(
approach_id=approach.id,
label=approach.label,
status="not_evaluated",
grounded=False,
opened=tuple(sorted(wanted & opened_paths)),
cited=(),
citation_scope="absent",
named=False,
named_in_measure=False,
named_in_snippet=False,
hallucinations=(),
ferdig=False,
)
)
continue
rows_seen += 1
payload = _read_json(proposal_path)
proposal = payload.get("proposal", {})
citations = payload.get("provenance", {}).get("citations", [])
citations_seen += len(citations)
status = str(_read_json(outcome_path).get("outcome_type", "rejected"))
if status == "validated":
validated_ids.add(approach.id)
validated_codes.update(
str(item.get("code", "")) for item in proposal.get("affected_items", [])
)
cited_files = {str(c.get("file", "")) for c in citations}
scope = (
"absent"
if not cited_files
else "whole-base"
if concept_names <= cited_files
else "narrowed"
)
opened = tuple(sorted(wanted & opened_paths))
cited = tuple(sorted(wanted & cited_files))
# A whole-base list is stamped before any model work - see the module docstring.
grounded = bool(opened) or (bool(cited) and scope == "narrowed")
measure = str(proposal.get("measure", ""))
snippets = " ".join(str(c.get("snippet", "")) for c in citations)
marks = [m for c in concepts for m in (c.get("ref", ""), c.get("title", "")) if m]
named_in_measure = any(m in measure for m in marks)
named_in_snippet = any(m in snippets for m in marks)
halluc = [f"citation:{f}" for f in sorted(cited_files - concept_names)]
allowed = set(approach.affected_codes) | baseline_codes
halluc += [
f"code:{code}"
for code in sorted(
{str(i.get("code", "")) for i in proposal.get("affected_items", [])} - allowed
)
]
rows.append(
ApproachVerdict(
approach_id=approach.id,
label=approach.label,
status=status,
grounded=grounded,
opened=opened,
cited=cited,
citation_scope=scope,
named=named_in_measure or named_in_snippet,
named_in_measure=named_in_measure,
named_in_snippet=named_in_snippet,
hallucinations=tuple(halluc),
ferdig=(
grounded
and (named_in_measure or named_in_snippet)
and not halluc
and reads_clean
),
)
)
if rows_seen == 0:
raise EmptyMeasurement(
f"{outbox} holds no proposal artefact for run {run_id!r} - "
"an empty outbox is not a clean result"
)
# ---- the falsification arm ---------------------------------------------------------------
refusals: list[RefusalVerdict] = []
for row in fasit.get("must_refuse", []):
rid = row["approach_id"]
commissioned = next((a for a in mandate.approaches if a.id == rid), None)
codes = set(commissioned.affected_codes) if commissioned is not None else set()
leaked = sorted(codes & validated_codes)
if rid in validated_ids:
refusals.append(
RefusalVerdict(
rid, False, f"{rid} was VALIDATED - the base carries no ground for it"
)
)
elif leaked:
refusals.append(
RefusalVerdict(
rid, False, f"{rid}'s codes {leaked} ride inside a validated proposal"
)
)
else:
refusals.append(RefusalVerdict(rid, True, "no validated outcome rests on it"))
judged = [r for r in rows if r.approach_id not in refuse_ids]
return ContextSetVerdict(
context_set=context.name,
run_id=run_id,
bundle_id=str(fasit.get("bundle_id", "")),
approaches=tuple(rows),
must_refuse=tuple(refusals),
hallucinated_reads=tuple(hallucinated_reads),
tool_calls_seen=len(tool_calls),
citations_seen=citations_seen,
approach_rows_seen=rows_seen,
concepts_in_base=len(concept_names),
ferdig=bool(judged) and all(r.ferdig for r in judged) and all(r.passed for r in refusals),
)
def main(argv: list[str] | None = None) -> int:
"""``python -m portfolio_optimiser.stress <context_dir> --outbox-dir D --run-id R``.
Writes the verdict to ``<outbox>/<run_id>-verdict.json`` AND prints it, so a CI reader and a
human reader get the same bytes. Its own module, never a ``run.py`` mode: no partition row is
touched and this entry point cannot start a paid run."""
parser = argparse.ArgumentParser(prog="portfolio-optimiser-stress", description=__doc__)
parser.add_argument("context_dir", help="contexts/<set>")
parser.add_argument("--outbox-dir", required=True)
parser.add_argument("--run-id", required=True)
parser.add_argument(
"--bundle-root",
default=os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", _DEFAULT_BUNDLE_ROOT),
help="directory the set's bundle.txt name is mounted under",
)
args = parser.parse_args(argv)
context = Path(args.context_dir)
declared = dict(
line.split(":", 1) # type: ignore[misc]
for line in (context / "bundle.txt").read_text(encoding="utf-8").splitlines()
if ":" in line
)
base = Path(args.bundle_root).expanduser() / declared["name"].strip()
try:
verdict = score_context_set(context, args.outbox_dir, args.run_id, base)
except EmptyMeasurement as exc:
print(f"stress refused: {exc}", file=sys.stderr)
return 1
payload = json.dumps(verdict.to_payload(), sort_keys=True, indent=2, ensure_ascii=False) + "\n"
out = Path(args.outbox_dir) / f"{args.run_id}-verdict.json"
out.write_text(payload, encoding="utf-8")
print(payload, end="")
return 0
if __name__ == "__main__": # pragma: no cover - exercised by a subprocess test
raise SystemExit(main())

View file

@ -18,13 +18,22 @@ rather than vacuously green: "the anchor was not found" is equally true of a bas
read (Verifiseringsloven, ansikt 4). read (Verifiseringsloven, ansikt 4).
**Rule U** the measurable form of "the base cannot answer this" (documented in **Rule U** the measurable form of "the base cannot answer this" (documented in
``docs/2026-09-12-p14-kontekstsett.md § 2.3``): each unanswerable question declares >= 1 ``anchor``, ``docs/2026-09-12-p14-kontekstsett.md § 2.3``): each ``must_refuse`` row declares >= 1 ``anchor``,
a lowercase word of >= 4 characters, and is admitted **iff every anchor is absent case-insensitive a lowercase word of >= 4 characters, and is admitted **iff every anchor is absent case-insensitive
substring from the WHOLE text (frontmatter + body) of EVERY concept document in the base**. Not substring from the WHOLE text (frontmatter + body) of EVERY concept document in the base**. Not
"shares no keyword with any title": a tunnel question shares "tunnel" with hundreds of titles and "shares no keyword with any title": a tunnel question shares "tunnel" with hundreds of titles and
that proves nothing. What makes a question unanswerable is that the base lacks the SUBJECT, and the that proves nothing. What makes a question unanswerable is that the base lacks the SUBJECT, and the
anchor is that subject. Titles alone would be a proxy the full text costs nothing more to replace anchor is that subject. Titles alone would be a proxy the full text costs nothing more to replace
(measured: 0.77 s for r761-2025, the largest base). (measured: 0.77 s for r761-2025, the largest base).
**P16 A2 moved rule U from ``unanswerable`` to ``must_refuse``, and that is ONE form rather than
two.** ``po`` is not a lookup tool (D-1), so an "unanswerable question" had no runnable form: no
execution path ever consumed those rows, and the falsification half of session 102's criterion was
therefore unprovable by a run. The same fact now rides a FOURTH commissioned approach per set
(``a4-``) whose cost line the base carries no ground for, and ``must_refuse`` names it by
``approach_id`` plus the same anchors. Keeping the questions beside it as a second key would be two
copies of one fact (-(p)) the questions survive inside the row's ``rationale``, which nothing
keys on. Rule U itself is UNCHANGED, and its known-positive is still red.
""" """
from __future__ import annotations from __future__ import annotations
@ -143,7 +152,7 @@ def anchors_are_absent(
if not concepts: if not concepts:
raise ValueError("rule U ran against ZERO concepts: absence here is unmeasured, not false") raise ValueError("rule U ran against ZERO concepts: absence here is unmeasured, not false")
if not anchors: if not anchors:
raise ValueError("an unanswerable question declares no anchors, so nothing was checked") raise ValueError("a must_refuse row declares no anchors, so nothing was checked")
carried = [] carried = []
for anchor in anchors: for anchor in anchors:
if anchor != anchor.lower() or len(anchor) < _MIN_ANCHOR_CHARS: if anchor != anchor.lower() or len(anchor) < _MIN_ANCHOR_CHARS:
@ -217,9 +226,14 @@ def test_the_fasit_names_every_commissioned_approach(set_dir: Path) -> None:
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8")) fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
mandate = load_mandate(set_dir / "mandate.json") mandate = load_mandate(set_dir / "mandate.json")
cited = {row["approach_id"] for row in fasit["must_cite"]} cited = {row["approach_id"] for row in fasit["must_cite"]}
assert cited == {a.id for a in mandate.approaches} refused = {row["approach_id"] for row in fasit["must_refuse"]}
assert not (cited & refused), "an approach is either answerable or the falsification arm"
assert cited | refused == {a.id for a in mandate.approaches}
assert fasit["honesty"].strip(), "DEL 2(iii): what in this set is constructed must be stated" assert fasit["honesty"].strip(), "DEL 2(iii): what in this set is constructed must be stated"
assert len(fasit["unanswerable"]) >= 2, "the order asks for at least two per set" assert refused, "P16 A2: every set carries the falsification arm"
for row in fasit["must_refuse"]:
assert len(row["anchors"]) >= 2, "the order asks for at least two per set"
assert row["rationale"].strip(), "why the base cannot ground it must be stated"
assert (set_dir / "docs").is_dir(), "the form declares a docs/ directory even when it is empty" assert (set_dir / "docs").is_dir(), "the form declares a docs/ directory even when it is empty"
@ -263,11 +277,11 @@ def test_c_rule_u_every_unanswerable_question_is_unanswerable(set_dir: Path) ->
) )
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8")) fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
for row in fasit["unanswerable"]: for row in fasit["must_refuse"]:
carried = anchors_are_absent(row["anchors"], concepts) carried = anchors_are_absent(row["anchors"], concepts)
assert not carried, ( assert not carried, (
f"{set_dir.name}: {declared['name']} DOES carry {carried} over {len(concepts)} " f"{set_dir.name}: {declared['name']} DOES carry {carried} over {len(concepts)} "
f"concepts, so {row['question']!r} is not unanswerable by rule U" f"concepts, so {row['approach_id']!r} is not un-groundable by rule U"
) )

View file

@ -0,0 +1,460 @@
"""P16 DEL A - the stress judge: "beviselig virkning against the base" made DETERMINISTIC.
**The measured silence.** Session 102's criterion ((a) the run builds on the right fasit concept OR
refuses anchored * (b') it NAMES that concept * (c) zero hallucinations) was adjudicated BY HAND.
A hand-read criterion cannot be repeated four times now and N times later, and nothing in the tree
read ``contexts/<set>/fasit.json`` against an outbox at all (measured 14.09: ``grep -rln fasit src
tests`` hit only unrelated files).
``portfolio_optimiser.stress`` reads the artefacts that already carry the evidence -
``{run_id}[-{approach}]-proposal.json`` (proposal + provenance citations),
``-outcome.json`` (validated / rejected), ``{run_id}-debate.json`` (``tool_calls[]`` with
``name``/``bundle_id``/``path``, S2c) - and returns one typed verdict per commissioned approach.
**THE ORDER'S (a) WAS VACUOUS AS WRITTEN, AND THAT IS MEASURED.** The order defines grounded as
"a must_cite path was OPENED *or* CITED". But on the S2c navigation path ``run_project`` stamps
``citations = bundle_citations(bundle)``, which is ONE CITATION PER CONTEXT FILE - the whole corpus.
Measured on n100-2023: 446 context files, 446 citations, and **6 of 6 fasit paths already "cited"
before a single model call**. A judge honouring the order literally would be a gate that can only
be green, which is the repo's own vacuous-gate class, inside the gate built to stop it. So
``grounded`` counts a CITATION only when the citation list is NARROWER than the base (a declared
pre-pass cut); a whole-base list is reported as such and carries nothing. Both halves are reported
either way, so the operator can read which one fired - the deviation is stated, never silent.
**(b') was checked for the same vacuity and is CLEAN.** ``bundle_citations`` snippets are concept
BODIES, and the ``ref``/``title`` live in FRONTMATTER: measured 0 of 446 n100 bodies contain
``Krav 4.1.2-1``. So the snippet arm can carry (b') without being satisfied by construction, and
the order's definition is kept. Which half fired is still reported.
**A denominator, always** (Verifiseringsloven ansikt 4): every verdict names how many tool calls,
citations and approach rows it saw. An outbox with no artefacts RAISES rather than reporting
"0 hallucinations" - an empty measurement is not a clean bill of health.
**a4 / must_refuse is the falsification half in D-1 form.** ``po`` is not a lookup tool, so an
"unanswerable question" has no runnable form; a commissioned approach whose GROUND the base does
not carry does. It passes iff no ``validated`` row is that approach and no validated proposal
carries its code.
Arms: (a) grounded by an opened path * (b) a whole-base citation list cannot ground * (c) a
narrowed list can * (d) "refuses anchored" counts * (e) (b') by measure and by snippet *
(f) hallucinated citation files, read paths and codes * (g) must_refuse passes and fails *
(h) an empty outbox raises * (i) an empty base raises * (j) not_evaluated is absence of an
artefact * (k) the CLI writes the verdict file.
"""
from __future__ import annotations
import json
import subprocess
import sys
from pathlib import Path
import pytest
from portfolio_optimiser import stress
# --------------------------------------------------------------------------------------------
# A synthetic minibase + a synthetic outbox. Nothing here touches a real bundle or a model.
# --------------------------------------------------------------------------------------------
_GOOD = "krav/N1/id-good.md"
_OTHER = "krav/N1/id-other.md"
_REF = "Krav 1.2.3-4"
_TITLE = "Krav 1.2.3-4 Rundkjoringer"
def _minibase(root: Path) -> Path:
base = root / "minibase"
(base / "krav" / "N1").mkdir(parents=True)
(base / "index.md").write_text(
"---\nbundle_id: minibase\n---\n\n- [good](krav/N1/id-good.md)\n"
"- [other](krav/N1/id-other.md)\n",
encoding="utf-8",
)
(base / _GOOD).write_text(
f'---\ntype: concept\ntitle: "{_TITLE}"\nreq_number: "{_REF}"\n---\n\nBody of the good one.\n',
encoding="utf-8",
)
(base / _OTHER).write_text(
'---\ntype: concept\ntitle: "Other"\nreq_number: "Krav 9.9.9-9"\n---\n\nAnother body.\n',
encoding="utf-8",
)
return base
def _context(root: Path, *, must_refuse: bool = True) -> Path:
ctx = root / "ctx"
(ctx / "docs").mkdir(parents=True)
(ctx / "bundle.txt").write_text("name: minibase\nbundle_id: minibase\n", encoding="utf-8")
approaches = [
{
"id": "a1",
"label": "First approach",
"description": "",
"affected_codes": ["CODE-1"],
"claimed_saving_nok": 1000.0,
"bundle_id": "minibase",
}
]
fasit: dict[str, object] = {
"project_id": "proj",
"bundle": "minibase",
"bundle_id": "minibase",
"must_cite": [
{
"approach_id": "a1",
"rationale": "why",
"concepts": [{"path": _GOOD, "title": _TITLE, "ref": _REF}],
}
],
"must_refuse": [],
"honesty": "synthetic",
}
if must_refuse:
approaches.append(
{
"id": "a4",
"label": "Unit price cut",
"description": "",
"affected_codes": ["CODE-4"],
"claimed_saving_nok": 500.0,
"bundle_id": "minibase",
}
)
fasit["must_refuse"] = [
{"approach_id": "a4", "anchors": ["enhetspris"], "rationale": "base carries no prices"}
]
(ctx / "mandate.json").write_text(
json.dumps(
{"objective": "o", "success_criteria": "s", "approaches": approaches},
ensure_ascii=False,
indent=2,
),
encoding="utf-8",
)
(ctx / "fasit.json").write_text(
json.dumps(fasit, ensure_ascii=False, indent=2), encoding="utf-8"
)
return ctx
def _write_outbox(
outbox: Path,
run_id: str,
*,
approach_id: str,
measure: str = f"Simplify per {_REF}",
codes: list[str] | None = None,
citation_files: list[str] | None = None,
citation_snippet: str = "Body of the good one.",
decision: str = "validated",
tool_calls: list[dict[str, str]] | None = None,
) -> None:
outbox.mkdir(parents=True, exist_ok=True)
codes = ["CODE-1"] if codes is None else codes
citation_files = [_GOOD, _OTHER] if citation_files is None else citation_files
stem = f"{run_id}-{approach_id}"
(outbox / f"{stem}-proposal.json").write_text(
json.dumps(
{
"run_id": run_id,
"approach_id": approach_id,
"proposal": {
"project_id": "proj",
"measure": measure,
"affected_items": [
{"code": c, "quantity": 1.0, "unit_cost": 2000.0} for c in codes
],
"claimed_saving_nok": 1000.0,
"assumptions": {},
},
"provenance": {
"citations": [
{
"file": f,
"locator": {"start_index": 0, "end_index": 1},
"snippet": citation_snippet,
}
for f in citation_files
],
"model": "m",
"role": "proposer",
"validator_decision": decision,
"token_usage": 10,
"cost_baseline_anchored": False,
"bundle_id_source": None,
"external_calls": [],
},
},
indent=2,
),
encoding="utf-8",
)
(outbox / f"{stem}-outcome.json").write_text(
json.dumps(
{
"run_id": run_id,
"approach_id": approach_id,
"outcome_type": "validated" if decision == "validated" else "rejected",
**({"reason": "no"} if decision != "validated" else {}),
"checker_verdict": None,
"verdict_id": "vid",
},
indent=2,
),
encoding="utf-8",
)
if tool_calls is not None:
(outbox / f"{run_id}-debate.json").write_text(
json.dumps({"run_id": run_id, "tool_calls": tool_calls}, indent=2), encoding="utf-8"
)
def _opened(path: str) -> list[dict[str, str]]:
return [{"name": "read_file", "bundle_id": "minibase", "path": path}]
def _judge(tmp_path: Path, **kw: object) -> stress.ContextSetVerdict:
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=bool(kw.pop("must_refuse", False)))
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", **kw) # type: ignore[arg-type]
return stress.score_context_set(ctx, outbox, "r1", base)
# --------------------------------------------------------------------------------------------
# (a) grounded by an OPENED path - the non-vacuous half.
# --------------------------------------------------------------------------------------------
def test_a_a_must_cite_path_that_was_opened_grounds_the_approach(tmp_path: Path) -> None:
verdict = _judge(tmp_path, tool_calls=_opened(_GOOD))
row = verdict.approaches[0]
assert row.grounded is True
assert row.opened == (_GOOD,)
def test_a_opening_some_other_document_does_not_ground_it(tmp_path: Path) -> None:
verdict = _judge(tmp_path, tool_calls=_opened(_OTHER))
row = verdict.approaches[0]
assert row.grounded is False
assert row.opened == ()
# --------------------------------------------------------------------------------------------
# (b)+(c) the MEASURED vacuity: a whole-base citation list cannot ground; a narrowed one can.
# --------------------------------------------------------------------------------------------
def test_b_a_whole_base_citation_list_cannot_ground_an_approach(tmp_path: Path) -> None:
"""Measured on n100-2023: 446 context files, 446 citations, 6/6 fasit paths 'cited' before
any model call. Honouring the order literally would make (a) green by construction."""
verdict = _judge(tmp_path, tool_calls=[])
row = verdict.approaches[0]
assert row.cited == (_GOOD,), "the path IS in the citation list"
assert row.citation_scope == "whole-base"
assert row.grounded is False, "a whole-base list is stamped before any model work"
def test_c_a_narrowed_citation_list_does_ground_an_approach(tmp_path: Path) -> None:
verdict = _judge(tmp_path, tool_calls=[], citation_files=[_GOOD])
row = verdict.approaches[0]
assert row.citation_scope == "narrowed"
assert row.grounded is True
# --------------------------------------------------------------------------------------------
# (d) "refuses anchored" counts as (a) - session 102's own second limb.
# --------------------------------------------------------------------------------------------
def test_d_a_rejected_approach_that_opened_the_requirement_is_grounded(tmp_path: Path) -> None:
verdict = _judge(tmp_path, decision="rejected", tool_calls=_opened(_GOOD))
row = verdict.approaches[0]
assert row.status == "rejected"
assert row.grounded is True
# --------------------------------------------------------------------------------------------
# (e) (b') named - by the model's own prose, or by a citation snippet.
# --------------------------------------------------------------------------------------------
def test_e_the_ref_in_the_measure_names_the_concept(tmp_path: Path) -> None:
row = _judge(tmp_path, tool_calls=_opened(_GOOD)).approaches[0]
assert row.named is True
assert row.named_in_measure is True
def test_e_a_measure_that_names_nothing_is_carried_only_by_a_snippet(tmp_path: Path) -> None:
row = _judge(
tmp_path, measure="Do it cheaper", citation_snippet=f"see {_REF}", tool_calls=_opened(_GOOD)
).approaches[0]
assert row.named_in_measure is False
assert row.named_in_snippet is True
assert row.named is True
def test_e_naming_neither_way_fails_b_prime(tmp_path: Path) -> None:
row = _judge(tmp_path, measure="Do it cheaper", tool_calls=_opened(_GOOD)).approaches[0]
assert row.named is False
assert row.ferdig is False
# --------------------------------------------------------------------------------------------
# (f) (c) hallucinations: citation files, read paths, and codes.
# --------------------------------------------------------------------------------------------
def test_f_a_citation_to_a_file_the_base_does_not_carry_is_a_hallucination(
tmp_path: Path,
) -> None:
row = _judge(
tmp_path, citation_files=[_GOOD, "krav/N1/id-invented.md"], tool_calls=_opened(_GOOD)
).approaches[0]
assert row.hallucinations == ("citation:krav/N1/id-invented.md",)
assert row.ferdig is False
def test_f_a_read_path_the_base_does_not_carry_is_a_hallucination(tmp_path: Path) -> None:
verdict = _judge(tmp_path, tool_calls=_opened(_GOOD) + _opened("krav/N1/id-guessed.md"))
assert verdict.hallucinated_reads == ("krav/N1/id-guessed.md",)
assert verdict.approaches[0].ferdig is False
def test_f_a_code_in_neither_the_mandate_nor_the_baseline_is_a_hallucination(
tmp_path: Path,
) -> None:
row = _judge(tmp_path, codes=["CODE-1", "INVENTED-9"], tool_calls=_opened(_GOOD)).approaches[0]
assert row.hallucinations == ("code:INVENTED-9",)
def test_f_a_clean_run_is_ferdig(tmp_path: Path) -> None:
verdict = _judge(tmp_path, tool_calls=_opened(_GOOD))
row = verdict.approaches[0]
assert (row.grounded, row.named, row.hallucinations) == (True, True, ())
assert row.ferdig is True
# --------------------------------------------------------------------------------------------
# (g) must_refuse / a4 - the falsification arm.
# --------------------------------------------------------------------------------------------
def test_g_an_a4_nobody_validated_passes(tmp_path: Path) -> None:
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=True)
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
verdict = stress.score_context_set(ctx, outbox, "r1", base)
assert verdict.must_refuse[0].passed is True
assert verdict.must_refuse[0].approach_id == "a4"
def test_g_a_validated_a4_is_a_falsification_failure(tmp_path: Path) -> None:
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=True)
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
_write_outbox(outbox, "r1", approach_id="a4", codes=["CODE-4"], decision="validated")
verdict = stress.score_context_set(ctx, outbox, "r1", base)
assert verdict.must_refuse[0].passed is False
assert "VALIDATED" in verdict.must_refuse[0].detail
def test_g_a4s_code_inside_another_validated_proposal_also_fails(tmp_path: Path) -> None:
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=True)
outbox = tmp_path / "out"
_write_outbox(
outbox, "r1", approach_id="a1", codes=["CODE-1", "CODE-4"], tool_calls=_opened(_GOOD)
)
verdict = stress.score_context_set(ctx, outbox, "r1", base)
assert verdict.must_refuse[0].passed is False
# --------------------------------------------------------------------------------------------
# (h)+(i) denominators: an empty measurement is never a clean bill of health.
# --------------------------------------------------------------------------------------------
def test_h_an_outbox_with_no_artefacts_raises_never_reports_zero_hallucinations(
tmp_path: Path,
) -> None:
base = _minibase(tmp_path)
ctx = _context(tmp_path)
empty = tmp_path / "out"
empty.mkdir()
with pytest.raises(stress.EmptyMeasurement, match="no proposal artefact"):
stress.score_context_set(ctx, empty, "r1", base)
def test_i_a_base_that_scans_to_nothing_raises(tmp_path: Path) -> None:
ctx = _context(tmp_path)
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
hollow = tmp_path / "hollow"
hollow.mkdir()
(hollow / "index.md").write_text("---\nbundle_id: minibase\n---\n\nnothing\n", encoding="utf-8")
with pytest.raises(stress.EmptyMeasurement, match="concepts"):
stress.score_context_set(ctx, outbox, "r1", hollow)
def test_h_the_denominators_are_always_reported(tmp_path: Path) -> None:
verdict = _judge(tmp_path, tool_calls=_opened(_GOOD))
assert verdict.tool_calls_seen == 1
assert verdict.citations_seen == 2
assert verdict.approach_rows_seen == 1
assert verdict.concepts_in_base == 2
# --------------------------------------------------------------------------------------------
# (j) a commissioned approach with NO artefact is not_evaluated, never omitted.
# --------------------------------------------------------------------------------------------
def test_j_a_commissioned_approach_with_no_artefact_is_not_evaluated(tmp_path: Path) -> None:
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=True)
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
verdict = stress.score_context_set(ctx, outbox, "r1", base)
rows = {r.approach_id: r for r in verdict.approaches}
assert rows["a4"].status == "not_evaluated"
assert rows["a4"].ferdig is False
assert len(verdict.approaches) == 2, "an omitted row reads as an approach nobody ordered"
# --------------------------------------------------------------------------------------------
# (k) the CLI writes the verdict beside the artefacts it judged.
# --------------------------------------------------------------------------------------------
def test_k_the_cli_writes_the_verdict_file_and_prints_it(tmp_path: Path) -> None:
base = _minibase(tmp_path)
ctx = _context(tmp_path)
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
proc = subprocess.run(
[
sys.executable,
"-m",
"portfolio_optimiser.stress",
str(ctx),
"--outbox-dir",
str(outbox),
"--run-id",
"r1",
"--bundle-root",
str(base.parent),
],
capture_output=True,
text=True,
cwd=Path(__file__).resolve().parents[1],
)
assert proc.returncode == 0, proc.stderr
written = outbox / "r1-verdict.json"
assert written.is_file()
payload = json.loads(written.read_text(encoding="utf-8"))
assert payload["ferdig"] is True
assert json.loads(proc.stdout)["ferdig"] is True