feat(p21): the PROJECT carries the price, so a run against a road normal can be anchored

Four paid stress rounds ran entirely UN-ANCHORED, all of them, because the one file
loader reads cost-baseline.json out of the BUNDLE and no vegnormal ships one: N100,
N200, N500 and R761 are knowledge, and knowledge carries requirements, never amounts.
The validator's stage 0 -- the one stage that tells an invented cost line from a line
this project actually buys -- was skipped in every single run, so "validated" could not
mean what it says. P20 G1/G2 measured real R761 process numbers (12.11 three times on
Soraasen, 1.1.1 on Lindaas) validating with amounts nobody had anywhere.

--cost-baseline FILE is PM decision (e), taken over the three alternatives P20 wrote
down. A LOADED object, never a path (prepass_payload's rule): the CLI owns the file and
loads it ONCE, so the notice, the stamp and every base of an --across-bundle pass all
descend from one read. ONE parse, two doors -- load_cost_baseline delegates to
load_cost_baseline_file -- while safe_resolve stays on the bundle door alone, because a
project's own schedule is legitimately outside every base. No tolerant twin: this path
exists only because an operator NAMED a file.

DEL B: five anchored context sets, a1-a3 with their line and a4 with none, so stage 0 is
what catches the falsification arm. THE ORDER'S OWN ARM (h) WAS FELLED BY MEASUREMENT:
"no baseline code is a requirement number the base declares" is measured 0 of 4 on the
project-coded sets and 5 of 5 on kontrakt-sorasen -- which is what R761 Prosesskoden IS,
a bill of quantities priced BY process code. The complement keeps both, and the order's
own mutation still bites.

DEL B3: the judge reports anchored (off the run's own stamp), priced per row, and WHICH
falsifier caught the falsification arm.

Load-bearing MEASURED, five mutations all red against the WHOLE suite, green control
1850/5 (from 1809/5, superset, 0 removed), golden byte-unchanged:
A3(i) the flag is read but the baseline is unused (3 red) . A3(ii) only the first base
gets it (1) . A3(iii) report_forbidden drops it (1) . B2(i) a4 gets a line (1, arm (g)
alone) . B2(ii) a code swapped to 12.11 (2, arms (f) and (h)).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-15 10:49:10 +02:00
commit 7b4f85d77c
20 changed files with 1259 additions and 15 deletions

View file

@ -2848,6 +2848,76 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
en base som ikke lar seg løse faller tilbake til katalognavnet — `dimension_label`-presedensen
ordrett: å annonsere skal ALDRI endre hvilken feil en operatør ser. Load-bearing MÅLT
(`tests/test_parse_error_feedback_loadbearing.py`, 7 armer).
- **PRISEN HØRER TIL PROSJEKTET, ikke til kunnskapsbasen — og ordrens egen B2-regel ble FELT av
måling (P21 DEL A+B, 15.09):** fire betalte runder (P16/P18/P19/P17b/P20) kjørte **UFORANKRET,
alle sammen**, fordi den ene fil-lasteren leser `cost-baseline.json` ut av BUNDLE-katalogen og
ingen vegnormal bærer et prisskjema: N100/N200/N500/R761 er KUNNSKAP, og kunnskap bærer krav,
aldri beløp. Validatorens **stadium 0** — det ENE stadiet som skiller en oppdiktet kostlinje fra
en linje dette prosjektet faktisk kjøper — ble derfor hoppet over i hver eneste av dem, og
`validated` kunne ikke bety det det sier: P20 G1/G2 målte EKTE R761-prosessnumre (`12.11` ×3 på
sorasen, `1.1.1` på lindaas) som validerte med beløp ingen hadde noe sted.
`--cost-baseline FILE` er PM-beslutning **(e)**, valgt over tre alternativer P20 skrev ned:
(a) nekt enhver kravformet kode uforankret ville gjort det ENE realistiske kontekstsettet
umålbart, (b) `--require-cost-baseline` som default ville etterlatt ingen stresstest, og
(c) K2s prisskjema er nektet av MAJOR-4s egen uttalte grense. **Et LASTET objekt, aldri en sti**
(`prepass_payload`-regelen, og `mandate=`/`dimension=` før den): CLI-en eier fila, biblioteks-
sømmen tar det validerte artefaktet — lastet ÉN gang, så notisen, stempelet og hver base i et
`--across-bundle`-pass stammer fra én lesing (kø-(p)). **ÉN parse, TO dører:**
`okf.load_cost_baseline_file` er hvor bytene tolkes og `load_cost_baseline` delegerer til den;
det som skiller er OPPLØSNINGEN — `safe_resolve` blir værende på bundle-døra ALENE, fordi et
prosjekts eget prisskjema legitimt ligger utenfor hver base. **Ingen tolerant tvilling**, og det
er motsatt av `load_optional_cost_baseline`: bundle-fila er fraværende som default (en base
skrevet før amendmentet er legitimt uforankret), mens denne stien finnes KUN fordi en operatør
NAVNGA en fil — å tolerere dens fravær ville besvart en eksplisitt ordre med en stille uforankret
kjøring (`load_mandate`-regelen). **Gjensidig utelukkende med `--derive-cost-baseline`, håndhevet
BEGGE steder:** CLI-en nekter VED NAVN (så operatøren hører hvilke to flagg som kolliderer) og
`run_project` reiser `ValueError` (så en bibliotekskaller ikke kan nå en tilstand CLI-en nekter).
**`cost_baseline_source_notice` er en ANDRE renderer ved siden av `cost_baseline_notice`, aldri
en utvidelse av den:** de sier ULIKE fakta og kan ikke være uenige (å oppgi flagget INNEBÆRER
forankret, så nøyaktig én av de to kan rendres), og den POSITIVE linja er et BEVISST avvik fra
omisjons-regelen — `proposal_review_notice`-avviket, av samme grunn: stillhet her er TVETYDIG,
for en operatør som ga et prisskjema kan ikke skille «fila di forankret kjøringen» fra «basen
hadde sin egen» eller fra «flagget ble droppet». Uten fil returnerer den `None`, så omisjonen
beholdes nøyaktig der den er entydig. **`--across-bundle` får SAMME skjema per base** (ett
prosjekt, ett prisskjema) — det er den ene ankringsparameteren som IKKE er en bundle-sak, og en
base med og en uten ville forankret halve kommisjonen mens stempelet rapporterte ankring for den
halvdelen som tilfeldigvis kjørte først. Tre partisjons-rader (`--portfolio` og
`report_forbidden` nekter VED NAVN med rc-0-kontroll; live-dry-run er en WIRING, så den frie
turen sier det samme som den betalte). **DEL B: fem forankrede kontekstsett** — hvert
`contexts/<sett>/cost-baseline.json` har 48 linjer, a1a3 har SIN linje og a4/`must_refuse` har
INGEN, så stadium 0 er det som fanger falsifiseringsarmen. **Beskrivelser er BEVISST UTELATT fra
JSON-en:** `CostBaselineLine` har ingen slik nøkkel, pydantic ignorerer ekstra felt i stillhet,
og en fixtur hvis innhold droppes taust er en løgn — teksten bor i `mandate.json`s
label/description, som er nøklet på samme kode (kø-(p)). **ORDRENS ARM (h) BLE FELT AV MÅLING
FØR NOE BLE BYGGET PÅ DEN:** regelen «ingen baseline-kode er et kravnummer basen erklærer» er
MÅLT mot `okf.declared_reference_numbers` over de fire monterte basene — de fire prosjektkodede
settene bærer **0**, og `kontrakt-sorasen-2027` bærer **5 av 5** (`12.1`, `12.12`, `22.1`,
`52.11`, `51.1` er ekte R761-`prosessnr`). Det er ikke et uhell i settet; det er hva R761
Prosesskoden ER — en norsk vegkontrakts mengdebeskrivelse prises BY prosesskode — så ordrens
regel ville tvunget fram en omskriving av nettopp det settet beslutning (e) ble valgt for å
bevare. **KOMPLEMENTET beholder begge:** et prisskjema kan prise det kommisjonen NAVNGIR, og kan
ikke INNFØRE en korpus-identifikator som en kostlinje ingen bestilte. Ordrens egen mutasjon biter
fortsatt (bytt en kode til `12.11`, en erklært `prosessnr` ingen approach bestiller → arm (h)
rød). **DEL B3:** dommeren rapporterer `anchored` (lest av kjøringens EGET stempel, aldri
re-avledet), `priced` per rad (mot settets eget skjema, rapportert enten kjøringen var forankret
eller ei, så runde 14 kan dømmes med samme instrument) og `stage` per `must_refuse`-rad fra
`validator.rejection_stage`, som bor ved siden av setningene den nøkler på (kø-(p)) og er en
RAPPORT, aldri en gate — derfor er `"other"` et ærlig svar der og ville ikke vært det inne i
pipelinen. Load-bearing MÅLT (`tests/test_cost_baseline_flag_loadbearing.py` 15 armer +
`tests/test_context_sets_loadbearing.py` +11 + `tests/test_stress_judge_loadbearing.py` +6 +
`tests/test_prose_code_form_loadbearing.py` +2), **fem mutasjoner alle røde mot HELE suiten** +
grønn kontroll **1850/5** (fra 1809/5, supersett, 0 fjernet) og golden
`demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): A3(i) flagget leses men baselinen brukes ikke (3
røde) · A3(ii) bare første base får skjemaet (1) · A3(iii) `report_forbidden` slipper flagget (1)
· B2(i) a4 får en linje (1, arm (g) alene) · B2(ii) en kode byttet til `12.11` (2, armene (f) og
(h)). **Ærlighets-grenser, uttalt:** skjemaet når IKKE prompten — modellen må fortsatt oppgi
mengde og enhetspris selv, og stadium 0s avvisning navngir BASELINE-verdien, så Steg 5s
tilbakemating er hva som lar løkka konvergere (økt 94s måling); den hostede flaten er BEVISST
urørt (feltet er i ingen av hostings tre sett, så den generiske 400-en svarer og Fase 4es to
halvdeler står); portefølje-armen er ikke wiret (flagget er nektet der ved navn, så en utskrift
ville vært død kode); og alle beløp i de fem skjemaene er OPPDIKTEDE størrelsesordener, som hvert
setts eget `honesty`-felt sier.
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -750,6 +750,25 @@ when the seam is detached, so the loop cannot silently degrade into theater.
--derive-cost-baseline --require-cost-baseline
```
- **The project's own price schedule**`--cost-baseline FILE` (requires `--bundle-dir` or
`--across-bundle`). A knowledge base carries what is REQUIRED, not what things cost: a road
normal, a standard or a regulation has requirements and no amounts, so a run against one has
nothing for the validator's stage 0 to reconcile against and that stage is skipped. The price
belongs to the project, and this is where you hand it over: FILE is a `cost-baseline.json` — the
same `{"project_id": …, "items": {"<code>": {"quantity": …, "unit_cost": …}}}` shape a bundle may
ship — and it is used INSTEAD of one inside the base. With it, a proposal naming a cost line the
project does not buy is refused as a fabricated line, and one naming a real line with invented
magnitudes is refused with the real ones named, so the next attempt can correct.
In `--across-bundle` mode the same schedule anchors every base: one project, one price schedule.
Mutually exclusive with `--derive-cost-baseline` (two sources for one baseline), and it satisfies
`--require-cost-baseline`. A missing or malformed file refuses the run before anything starts.
```bash
uv run python -m portfolio_optimiser.run PROSJEKT-1 --bundle-dir <bundle> \
--cost-baseline prisskjema.json --require-cost-baseline
```
`--explore` is refused together with `--mandate` — they are two sources of one mandate, and
merging would silently overwrite what you wrote. To seed an exploration with a domain expert's

View file

@ -0,0 +1,25 @@
{
"project_id": "dekke-og-kontrakt-lindaas-2027",
"items": {
"LIND-FORST-01": {
"quantity": 24600.0,
"unit_cost": 285.0
},
"LIND-FILT-01": {
"quantity": 18400.0,
"unit_cost": 210.0
},
"LIND-RIGG-01": {
"quantity": 1.0,
"unit_cost": 5900000.0
},
"LIND-ASF-01": {
"quantity": 4100.0,
"unit_cost": 640.0
},
"LIND-GRFT-01": {
"quantity": 2050.0,
"unit_cost": 1380.0
}
}
}

View file

@ -50,7 +50,7 @@
]
}
],
"honesty": "Prosjektet fv. 218 Lindaas er KONSTRUERT av meg: vegnummer, lengde, AADT og alle fire kostlinjene (LIND-FORST-01, LIND-FILT-01, LIND-RIGG-01, LIND-INDEKS-01) er oppdiktet, og beloepene er satte stoerrelsesordener. Kravene og prosessene i must_cite er lest ordrett ut av basenes egen frontmatter (n200-2024 for a1/a2, r761-2025 for a3). Verken N200 eller R761 baerer priser, saa kodene finnes ikke i noen av basene. Dette er det FOERSTE settet som spenner TO baser: a1/a2 rutes mot n200-2024 og a3/a4 mot r761-2025, og det er hele grunnen til at settet finnes -- P17b maaler at EN kommisjon kan kjoeres over flere kunnskapsbaser. Den fjerde tilnaermingen a4-indeksregulering og dens kostkode LIND-INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje INGEN av de to basene baerer grunnlaget for. MAALT 15.09: 'enhetspris' ble FORKASTET som anker fordi r761-2025 baerer ordet i 70 av 2 756 konsepter -- et anker som holder for ett sett med EN base holder ikke noedvendigvis for et sett med to.",
"honesty": "Prosjektet fv. 218 Lindaas er KONSTRUERT av meg: vegnummer, lengde, AADT og alle fire kostlinjene (LIND-FORST-01, LIND-FILT-01, LIND-RIGG-01, LIND-INDEKS-01) er oppdiktet, og beloepene er satte stoerrelsesordener. Kravene og prosessene i must_cite er lest ordrett ut av basenes egen frontmatter (n200-2024 for a1/a2, r761-2025 for a3). Verken N200 eller R761 baerer priser, saa kodene finnes ikke i noen av basene. Dette er det FOERSTE settet som spenner TO baser: a1/a2 rutes mot n200-2024 og a3/a4 mot r761-2025, og det er hele grunnen til at settet finnes -- P17b maaler at EN kommisjon kan kjoeres over flere kunnskapsbaser. Den fjerde tilnaermingen a4-indeksregulering og dens kostkode LIND-INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje INGEN av de to basene baerer grunnlaget for. MAALT 15.09: 'enhetspris' ble FORKASTET som anker fordi r761-2025 baerer ordet i 70 av 2 756 konsepter -- et anker som holder for ett sett med EN base holder ikke noedvendigvis for et sett med to. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema, og SAMME fil gjelder begge basene i multi-base-passet (ett prosjekt, ett prisskjema). Mengdene og enhetsprisene er OPPDIKTEDE stoerrelsesordener; verken N200 eller R761 baerer priser. a4s LIND-INDEKS-01 har ingen linje.",
"must_refuse": [
{
"approach_id": "a4-indeksregulering",

View file

@ -0,0 +1,25 @@
{
"project_id": "fv412-dekkefornyelse-2027",
"items": {
"DEKKE-ASF-01": {
"quantity": 10800.0,
"unit_cost": 1150.0
},
"DEKKE-BAER-01": {
"quantity": 6300.0,
"unit_cost": 980.0
},
"DEKKE-FROST-01": {
"quantity": 37800.0,
"unit_cost": 215.0
},
"DEKKE-GRV-01": {
"quantity": 12600.0,
"unit_cost": 145.0
},
"DEKKE-SKILT-01": {
"quantity": 1.0,
"unit_cost": 1450000.0
}
}
}

View file

@ -52,7 +52,7 @@
]
}
],
"honesty": "Prosjektet fv. 412 dekkefornyelse er KONSTRUERT av meg: vegnummer, lengde, ÅDT og alle tre kostlinjene (DEKKE-ASF-01, DEKKE-BAER-01, DEKKE-FROST-01) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er lest ordrett ut av n200-2024-basens egen frontmatter. N200 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-tonnpris-asfalt og dens kostkode DEKKE-ASF-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"honesty": "Prosjektet fv. 412 dekkefornyelse er KONSTRUERT av meg: vegnummer, lengde, ÅDT og alle tre kostlinjene (DEKKE-ASF-01, DEKKE-BAER-01, DEKKE-FROST-01) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er lest ordrett ut av n200-2024-basens egen frontmatter. N200 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-tonnpris-asfalt og dens kostkode DEKKE-ASF-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema. Mengdene og enhetsprisene der er OPPDIKTEDE stoerrelsesordener, satt slik at hver av a1-a3 har sin kostlinje og a4 IKKE har en; N200 baerer ingen priser, saa ingen av tallene er lest noe sted.",
"must_refuse": [
{
"approach_id": "a4-tonnpris-asfalt",

View file

@ -0,0 +1,25 @@
{
"project_id": "gate-nordvik-2027",
"items": {
"GATE-KRYSS-01": {
"quantity": 1.0,
"unit_cost": 9400000.0
},
"GATE-GANG-01": {
"quantity": 6.0,
"unit_cost": 310000.0
},
"GATE-KRYSS-02": {
"quantity": 8.0,
"unit_cost": 155000.0
},
"GATE-DEKKE-01": {
"quantity": 5400.0,
"unit_cost": 1250.0
},
"GATE-VA-01": {
"quantity": 900.0,
"unit_cost": 4800.0
}
}
}

View file

@ -52,7 +52,7 @@
]
}
],
"honesty": "Prosjektet Nordvikgata er KONSTRUERT av meg: gatenavn, lengde, fartsgrense og alle tre kostlinjene (GATE-KRYSS-01, GATE-GANG-01, GATE-KRYSS-02) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er derimot lest ordrett ut av n100-2023-basens egen frontmatter. N100 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-enhetspris-gangfelt og dens kostkode GATE-GANG-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"honesty": "Prosjektet Nordvikgata er KONSTRUERT av meg: gatenavn, lengde, fartsgrense og alle tre kostlinjene (GATE-KRYSS-01, GATE-GANG-01, GATE-KRYSS-02) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er derimot lest ordrett ut av n100-2023-basens egen frontmatter. N100 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-enhetspris-gangfelt og dens kostkode GATE-GANG-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema. Mengdene og enhetsprisene der er OPPDIKTEDE stoerrelsesordener, satt slik at hver av a1-a3 har sin kostlinje og a4 IKKE har en; N100 baerer ingen priser, saa ingen av tallene er lest noe sted.",
"must_refuse": [
{
"approach_id": "a4-enhetspris-gangfelt",

View file

@ -0,0 +1,25 @@
{
"project_id": "kontrakt-sorasen-2027",
"items": {
"12.1": {
"quantity": 1.0,
"unit_cost": 6400000.0
},
"12.12": {
"quantity": 22.0,
"unit_cost": 340000.0
},
"22.1": {
"quantity": 94000.0,
"unit_cost": 118.0
},
"52.11": {
"quantity": 41000.0,
"unit_cost": 265.0
},
"51.1": {
"quantity": 31000.0,
"unit_cost": 96.0
}
}
}

View file

@ -52,7 +52,7 @@
]
}
],
"honesty": "Entreprisen Søråsen er KONSTRUERT av meg: kontraktsnavn, lengde, riggperiode og alle tre beløpene er oppdiktede størrelsesordener. Til forskjell fra de tre andre settene er affected_codes her EKTE prosessnumre fra R761 (12.1, 12.12, 22.1, 51.1, 52.11), lest ut av basens egen prosessnr-frontmatter — men R761 er en beskrivelsesstandard uten priser, så kodene er ekte mens beløpene ikke er det. must_cite er lest ordrett ut av basen. P16: den fjerde tilnaermingen a4-indeksregulering og dens kostkode INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"honesty": "Entreprisen Søråsen er KONSTRUERT av meg: kontraktsnavn, lengde, riggperiode og alle tre beløpene er oppdiktede størrelsesordener. Til forskjell fra de tre andre settene er affected_codes her EKTE prosessnumre fra R761 (12.1, 12.12, 22.1, 51.1, 52.11), lest ut av basens egen prosessnr-frontmatter — men R761 er en beskrivelsesstandard uten priser, så kodene er ekte mens beløpene ikke er det. must_cite er lest ordrett ut av basen. P16: den fjerde tilnaermingen a4-indeksregulering og dens kostkode INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema, kodet med de samme EKTE prosessnumrene, som er slik en norsk vegkontrakt faktisk prises. Mengdene og enhetsprisene er OPPDIKTEDE stoerrelsesordener; R761 baerer ingen priser, saa ingen av tallene er lest noe sted. a4s INDEKS-01 har ingen linje.",
"must_refuse": [
{
"approach_id": "a4-indeksregulering",

View file

@ -0,0 +1,29 @@
{
"project_id": "tunnel-hauglia-2027",
"items": {
"TUN-VENT-01": {
"quantity": 14.0,
"unit_cost": 465000.0
},
"TUN-FROST-01": {
"quantity": 360.0,
"unit_cost": 21500.0
},
"TUN-LYS-01": {
"quantity": 2400.0,
"unit_cost": 1850.0
},
"TUN-SPRENG-01": {
"quantity": 168000.0,
"unit_cost": 410.0
},
"TUN-SIKRING-01": {
"quantity": 2400.0,
"unit_cost": 6900.0
},
"TUN-PORTAL-01": {
"quantity": 2.0,
"unit_cost": 3850000.0
}
}
}

View file

@ -62,7 +62,7 @@
]
}
],
"honesty": "Prosjektet Hauglia-tunnelen er KONSTRUERT av meg: navn, lengde, ÅDT og alle tre kostlinjene (TUN-VENT-01, TUN-FROST-01, TUN-LYS-01) er oppdiktet, og de tre beløpene er plausible størrelsesordener jeg har satt, ikke tall fra et prosjekt. Det som IKKE er konstruert er kravene: hver konsept-sti, tittel og kravnummer i must_cite er lest ordrett ut av n500-2024-basens egen frontmatter. N500 bærer ingen priser, så kodene finnes ikke i basen — settet er derfor bevisst IKKE kjørbart via --proposals-from-mandate. P16: den fjerde tilnaermingen a4-enhetspris-ventilator og dens kostkode TUN-VENT-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
"honesty": "Prosjektet Hauglia-tunnelen er KONSTRUERT av meg: navn, lengde, ÅDT og alle tre kostlinjene (TUN-VENT-01, TUN-FROST-01, TUN-LYS-01) er oppdiktet, og de tre beløpene er plausible størrelsesordener jeg har satt, ikke tall fra et prosjekt. Det som IKKE er konstruert er kravene: hver konsept-sti, tittel og kravnummer i must_cite er lest ordrett ut av n500-2024-basens egen frontmatter. N500 bærer ingen priser, så kodene finnes ikke i basen — settet er derfor bevisst IKKE kjørbart via --proposals-from-mandate. P16: den fjerde tilnaermingen a4-enhetspris-ventilator og dens kostkode TUN-VENT-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (6 linjer) — prosjektets prisskjema. Mengdene og enhetsprisene der er OPPDIKTEDE stoerrelsesordener, satt slik at hver av a1-a3 har sin kostlinje og a4 IKKE har en; N500 baerer ingen priser, saa ingen av tallene er lest noe sted.",
"must_refuse": [
{
"approach_id": "a4-enhetspris-ventilator",

View file

@ -1601,6 +1601,35 @@ def load_cost_baseline(bundle_dir: str, name: str = _COST_BASELINE) -> CostBasel
resolved = Path(safe_resolve(bundle_dir, name))
if not resolved.is_file():
raise FileNotFoundError(f"cost baseline not found in bundle: {name!r}")
return load_cost_baseline_file(str(resolved))
def load_cost_baseline_file(path: str) -> CostBaseline:
"""Load a cost baseline from a file that is NOT inside a knowledge base: the PROJECT's own price
schedule (P21, ``--cost-baseline``).
The measured reason it exists. Four paid rounds (P16/P18/P19/P17b/P20) ran entirely UN-ANCHORED,
because the only file loader reads ``cost-baseline.json`` out of the bundle directory and no road
normal carries a price schedule: a vegnormal is KNOWLEDGE, and the price belongs to the PROJECT.
Stage 0 was therefore skipped in every one of them, and "validated" could not mean anything
P20 G1/G2 measured real process numbers validating with invented amounts. This is the third door
into ``CostBaseline`` alongside the bundle file and ``derive_cost_baseline``, and the only one
whose input is the project rather than the corpus.
**The SAME parse, never a second one** (-(p)): ``load_cost_baseline`` resolves inside the
bundle and then delegates here, so the two doors cannot disagree about what a baseline file is.
What differs is the resolution ``safe_resolve`` is the ONE in-/out-of-bundle test and stays on
the bundle door alone, because a project's own schedule is legitimately outside every base.
Fail-fast, the error CLASSES of ``load_cost_baseline``: a missing file raises
``FileNotFoundError`` and malformed content raises ``pydantic.ValidationError``. There is no
optional twin, and that is deliberate: the bundle file is absent by default (a base authored
before the amendment is legitimately un-anchored), whereas this path exists only because an
operator NAMED a file tolerating its absence would answer an explicit order with a silently
un-anchored run (``load_mandate``'s rule)."""
resolved = Path(path)
if not resolved.is_file():
raise FileNotFoundError(f"cost baseline not found: {path!r}")
return CostBaseline.model_validate_json(resolved.read_text(encoding="utf-8"))

View file

@ -88,7 +88,7 @@ from portfolio_optimiser.generate import (
generate_via_llm,
grounding_offer,
)
from portfolio_optimiser.ir import SavingsProposal
from portfolio_optimiser.ir import CostBaseline, SavingsProposal
from portfolio_optimiser.mandate import (
OWN_PROPOSAL_ID,
Approach,
@ -822,6 +822,36 @@ def cost_baseline_notice(anchored: bool) -> str | None:
return None if anchored else _UNANCHORED_NOTICE
def cost_baseline_source_notice(path: str | None, lines: int) -> str | None:
"""Render where this run's cost baseline came from, or ``None`` when nobody named a file (P21).
A SECOND renderer beside ``cost_baseline_notice``, never a widening of it, because the two say
DIFFERENT facts and cannot disagree: that one warns that stage 0 is SKIPPED, this one names the
file an operator chose and how many lines it carries. Supplying ``--cost-baseline`` implies
anchored, so exactly one of the two can ever render.
**A POSITIVE line, and that is a deliberate departure from the omission rule** its neighbours
follow (``cost_baseline_notice``, ``skipped_links_notice``, ``unkeyed_verdicts_notice``) the
same departure ``proposal_review_notice`` makes, for the same reason. Silence here is
AMBIGUOUS: an operator who passed a project price schedule cannot tell "your file anchored this
run" from "the bundle happened to ship its own" or from "the flag was dropped somewhere", and
the whole point of the flag is that stage 0 now judges. Without a file the renderer returns
``None``, so the omission is kept exactly where it is unambiguous.
It takes the ALREADY-RESOLVED path and count rather than re-reading the file: a renderer that
opened it again would be a second resolution of the same fact, free to drift from the baseline
the run was actually given (``cost_baseline_notice``'s rule).
English, like every other line this CLI prints."""
if path is None:
return None
plural = "" if lines == 1 else "s"
return (
f" Cost baseline: {lines} line{plural} from {path} — the validator's stage 0 reconciles "
"every proposed cost line against this project's own schedule"
)
def grounding_offer_notice(offer: GroundingOffer | None) -> str | None:
"""Render the one line that says what this run's delivered input can ground, or ``None`` when
there is nothing to warn about (P8).
@ -1018,6 +1048,13 @@ async def run_project(
#: file loader, byte-identically.
derive_cost_baseline: bool = False,
require_cost_baseline: bool = False,
#: P21: the PROJECT's own price schedule, supplied by the caller instead of read out of the
#: knowledge base. A LOADED object, never a path — ``prepass_payload``'s rule, and
#: ``mandate=``/``dimension=``' before it: the CLI owns the file, the library seam takes the
#: validated artefact. ``None`` (the default) leaves every existing run on the bundle loader,
#: byte-identically. Mutually exclusive with ``derive_cost_baseline``: two sources for one
#: baseline would have to silently pick one, and the picked one would be a policy nobody wrote.
cost_baseline: CostBaseline | None = None,
dimension: Dimension | None = None,
store: VerdictStore | None = None,
verdict_dir: str | None = None,
@ -1081,6 +1118,17 @@ async def run_project(
semantics, so over a structural tie the resulting order is deterministic but arbitrary;
retrieval *quality* arrives only with an embedder injected via ``embedder=`` or
``--embedder-config``. Default false keeps the structural ranking exactly as before."""
# 0. Fail-fast: TWO sources for ONE baseline, refused rather than merged (P21). A run holding
# both would have to pick silently, and the picked one would be a policy nobody wrote down —
# ``--prepass-payload``/``--prepass-seed``'s rule, and ``portfolio_meter``/``meter_factory``'s
# before it. Checked HERE and not only at the CLI, because the library seam has the same two
# arguments and a library caller must not be able to reach a state the CLI refuses by name.
if cost_baseline is not None and derive_cost_baseline:
raise ValueError(
"cost_baseline and derive_cost_baseline are two sources for one baseline (a file the "
"caller supplies, and a schedule derived from the knowledge base); pass exactly one"
)
# 0. Fail-fast: an outbox write is byte-deterministic and keyed on run_id — no wall-clock default.
if outbox_dir is not None and run_id is None:
raise ValueError(
@ -1173,8 +1221,15 @@ async def run_project(
# silently downgraded order, which is what ``load_mandate`` fail-fasts against. This one
# resolution serves BOTH the full run and the ``live_dry_run`` report below, so the dry-run
# arm cannot drift away from what a real run would anchor on.
# P21 takes precedence over BOTH bundle-side sources, and it is the only one whose input
# is the PROJECT rather than the corpus: a vegnormal is knowledge and carries no prices, so
# before this every paid run measured (P16/P18/P19/P17b/P20) was un-anchored and stage 0
# never ran. Mutual exclusion with ``derive_cost_baseline`` is enforced at the top of this
# function, so the ``if`` below is an ordering and not a silent pick.
baseline = (
okf.derive_cost_baseline(bundle, project_id=project_id)
cost_baseline
if cost_baseline is not None
else okf.derive_cost_baseline(bundle, project_id=project_id)
if derive_cost_baseline
else okf.load_optional_cost_baseline(bundle_dir)
)
@ -2429,6 +2484,12 @@ async def run_mandate_across_bundles(
#: silent drop on the paid path while the free drill honoured them, which is the F4 class.
derive_cost_baseline: bool = False,
require_cost_baseline: bool = False,
#: P21, and threaded UNCHANGED into every base: ONE project has ONE price schedule, so the same
#: baseline anchors each base's run. That is the one anchoring argument which is NOT a bundle
#: concern — the two above are read out of the base being handed over, this one is the project's
#: own, and giving base k a baseline and base k+1 none would anchor half a commission while the
#: stamp reported anchoring for the half that happened to run first.
cost_baseline: CostBaseline | None = None,
) -> MultiBaseResult:
"""Evaluate ONE commission across SEVERAL knowledge bases — the multi-base dispatch (§ C.7).
@ -2531,6 +2592,7 @@ async def run_mandate_across_bundles(
run_id=base_run_id or None,
derive_cost_baseline=derive_cost_baseline,
require_cost_baseline=require_cost_baseline,
cost_baseline=cost_baseline,
notify=lambda verdict: minted_here.append(verdict.id),
verdict_input=verdict_input,
verdict_dir=verdict_dir,
@ -2967,6 +3029,20 @@ def main(argv: list[str] | None = None) -> int:
"guesses: an unpriced or ambiguous schedule stops the run"
),
)
parser.add_argument(
"--cost-baseline",
default=None,
metavar="FILE",
help=(
"anchor the validator's stage 0 to THIS PROJECT's own price schedule (P21): FILE is a "
"cost-baseline.json (the same {project_id, items:{code:{quantity,unit_cost}}} shape a "
"bundle may ship) and it is used INSTEAD of one inside --bundle-dir. The price belongs "
"to the project, not to the knowledge base — a road normal carries requirements, never "
"amounts — so without this a run against one is un-anchored and stage 0 is skipped. "
"Applies to every base in --across-bundle mode: one project, one schedule. Mutually "
"exclusive with --derive-cost-baseline; satisfies --require-cost-baseline"
),
)
parser.add_argument(
"--require-cost-baseline",
action="store_true",
@ -3072,6 +3148,10 @@ def main(argv: list[str] | None = None) -> int:
# Same reason, same rung: report mode returns above every run dispatch, so leaving it
# out is a SILENT DROP of a guarantee the operator asked for by name.
"--require-cost-baseline": args.require_cost_baseline,
# P21, same rung and same reason: report mode returns ABOVE every dispatch that could
# honour a project price schedule, so an omission here would accept the file, anchor
# nothing, and exit 0 — a silent drop rather than a refusal (the F4 class).
"--cost-baseline": args.cost_baseline is not None,
# Same reason, one flag later: report mode returns above the S7b dispatch too.
"--proposals-from-mandate": args.proposals_from_mandate,
"PROJECT_ID": args.project_id is not None,
@ -3183,6 +3263,11 @@ def main(argv: list[str] | None = None) -> int:
# anchored by construction — so here the flag could only ever pass. BY NAME rather
# than falling through to the --bundle-dir requirement, its neighbours' reason.
"--require-cost-baseline": args.require_cost_baseline,
# P21: ONE project's price schedule, and a portfolio pass keys on PROJECTS — each of
# which already carries its own ``cost_items`` and is anchored by construction. A
# single file could only ever be right for one row out of N. BY NAME rather than
# falling through to the --bundle-dir requirement, its neighbours' reason.
"--cost-baseline": args.cost_baseline,
# It reads ONE base's schedule and settles ONE commission against it, so it sits on the
# same side of the partition as the flag it requires. BY NAME rather than falling
# through to "requires --derive-cost-baseline": an operator who wrote --portfolio
@ -3364,6 +3449,37 @@ def main(argv: list[str] | None = None) -> int:
)
return 1
# P21, and its neighbours' reason exactly: on the road path the baseline IS the project's own
# ``cost_items``, so a file here would be a SECOND source for one fact with nothing to break
# the tie — and a flag that can only ever be shadowed is a claim the surface makes about
# itself. Refused BY NAME rather than left to surface as a project lookup that ignored it.
if (
not args.portfolio
and not args.across_bundle
and args.cost_baseline is not None
and args.bundle_dir is None
):
print(
"run refused: --cost-baseline requires --bundle-dir or --across-bundle (the road path "
"is already anchored by the reference project's own cost_items, so a second schedule "
"there would have nothing to anchor that those do not)",
file=sys.stderr,
)
return 1
# TWO sources for ONE baseline, refused rather than merged — the same decision
# ``run_project`` enforces at its own seam, said HERE by name so an operator who typed both
# hears which two flags conflict instead of getting a library ValueError's traceback. Not
# nested under either flag's branch, for the F4 reason its neighbours are not.
if args.cost_baseline is not None and args.derive_cost_baseline:
print(
"run refused: --cost-baseline and --derive-cost-baseline are two sources for one "
"baseline (a file you supply, and a schedule derived from a table in the knowledge "
"base); pass exactly one",
file=sys.stderr,
)
return 1
# The SEEDING arm's three refusals, placed ABOVE the replacing arm's block on purpose: given
# both flags, the block below would answer with "--prepass-payload and --explore cannot be
# combined", which names neither of the two flags the operator actually put in conflict. At
@ -3815,6 +3931,20 @@ def main(argv: list[str] | None = None) -> int:
print(f"run refused: {exc}", file=sys.stderr)
return 1
# The PROJECT's price schedule, loaded fail-fast alongside the commission and for the same
# reason: a baseline that cannot be read is not a run to start UN-anchored instead. Degrading
# it to "no baseline" would answer an operator who asked for stage 0 by name with a run in
# which stage 0 is skipped — ``load_mandate``'s rule, and the exact silence four paid rounds
# were measured inside. Loaded ONCE and passed as an object, so the notice below, the stamp
# and every base of an ``--across-bundle`` pass all descend from one read (kø-(p)).
cost_baseline: CostBaseline | None = None
if args.cost_baseline is not None:
try:
cost_baseline = okf.load_cost_baseline_file(args.cost_baseline)
except (FileNotFoundError, ValidationError, ValueError) as exc:
print(f"run refused: {exc}", file=sys.stderr)
return 1
# The declared cut, loaded fail-fast alongside the commission and for the same reason: a
# payload that cannot be read is not a run to start with a navigating debate instead. Missing,
# not JSON, or not the shape the models require — all three land on the refusal surface with
@ -4191,6 +4321,15 @@ def main(argv: list[str] | None = None) -> int:
print(f"run refused: {exc}", file=sys.stderr)
return 1
# P21, printed ONCE for the pass rather than per base, and that is the opposite placement
# from its neighbours below FOR A REASON: they read each base's OWN stamp, whereas ONE
# project has ONE price schedule and the same baseline anchors every base here. Per base it
# would read as N schedules, which is the claim this flag exists to deny.
source_notice = cost_baseline_source_notice(
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
)
if source_notice is not None:
print(source_notice)
if args.live_dry_run:
for bundle_id, bundle_dir, project_id in resolved:
try:
@ -4210,6 +4349,7 @@ def main(argv: list[str] | None = None) -> int:
max_tokens=args.max_tokens,
derive_cost_baseline=args.derive_cost_baseline,
require_cost_baseline=args.require_cost_baseline,
cost_baseline=cost_baseline,
live_dry_run=True,
)
)
@ -4269,6 +4409,7 @@ def main(argv: list[str] | None = None) -> int:
),
derive_cost_baseline=args.derive_cost_baseline,
require_cost_baseline=args.require_cost_baseline,
cost_baseline=cost_baseline,
outbox_for=outbox_for,
)
)
@ -4302,6 +4443,15 @@ def main(argv: list[str] | None = None) -> int:
multi=multi,
)
# P21, printed ONCE for the pass rather than per base, and that is the opposite placement
# from its neighbours below FOR A REASON: they read each base's OWN stamp, whereas ONE
# project has ONE price schedule and the same baseline anchors every base here. Per base it
# would read as N schedules, which is the claim this flag exists to deny.
source_notice = cost_baseline_source_notice(
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
)
if source_notice is not None:
print(source_notice)
for bundle_run in multi.runs:
print(f"--- {bundle_run.bundle_id} ({bundle_run.project_id}) ---")
print(
@ -4440,6 +4590,7 @@ def main(argv: list[str] | None = None) -> int:
verdict_input=_verdict_input_from_args(args),
derive_cost_baseline=args.derive_cost_baseline,
require_cost_baseline=args.require_cost_baseline,
cost_baseline=cost_baseline,
mcp_servers=mcp_servers,
live_dry_run=True,
)
@ -4470,6 +4621,14 @@ def main(argv: list[str] | None = None) -> int:
notice = cost_baseline_notice(report.cost_baseline_anchored)
if notice is not None:
print(notice)
# P21's positive half, on the FREE trip: the operator who named a project price schedule
# learns on the drill that it was read and how many lines it carries, rather than paying
# for a run to find out. Silence without the flag (omission, never an empty row).
source_notice = cost_baseline_source_notice(
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
)
if source_notice is not None:
print(source_notice)
# P8, printed next to the line it qualifies: "stage 0 is skipped" says the gate lost a
# falsifier; this says what the input could have offered it instead. On the FREE trip, so
# an operator learns a run cannot be grounded without paying three attempts to find out.
@ -4519,6 +4678,7 @@ def main(argv: list[str] | None = None) -> int:
semantic_retrieval=args.semantic_retrieval,
derive_cost_baseline=args.derive_cost_baseline,
require_cost_baseline=args.require_cost_baseline,
cost_baseline=cost_baseline,
client_factory=scripted_client_factory,
mandate=mandate,
mcp_servers=mcp_servers,
@ -4567,6 +4727,12 @@ def main(argv: list[str] | None = None) -> int:
notice = cost_baseline_notice(result.provenance.cost_baseline_anchored)
if notice is not None:
print(notice)
# Same renderer on the paid run, so the drill and the run it rehearses say the same thing.
source_notice = cost_baseline_source_notice(
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
)
if source_notice is not None:
print(source_notice)
# Same renderer on the full run, read off the run's OWN measurement: a run that spent every
# attempt being refused as ungrounded is exactly where the input-side fact costs the most.
offer_notice = grounding_offer_notice(result.grounding_offer)

View file

@ -76,7 +76,7 @@ from typing import Any
from portfolio_optimiser import okf
from portfolio_optimiser.mandate import Mandate, load_mandate
from portfolio_optimiser.validator import classify_codes
from portfolio_optimiser.validator import classify_codes, rejection_stage
#: Where the vegnormal bases are mounted, unless ``--bundle-root`` says otherwise. Read at CALL
#: time (the ``shared_root()`` idiom) so a test or an operator can move the mount without a reimport.
@ -125,6 +125,12 @@ class ApproachVerdict:
#: wrote one, and RE-DERIVED with the same classifier when it did not, so rounds 1 and 2 -
#: written before the field existed - can be re-judged with the same instrument.
prose_codes: tuple[str, ...]
#: P21 B3 - this row's ``affected_item`` codes ARE lines of the project's own price schedule.
#: Measured against ``contexts/<set>/cost-baseline.json`` (the schedule the run is given with
#: ``--cost-baseline``) and reported whether or not the run was anchored, so rounds written
#: before the schedule existed can be re-judged with the same instrument. ``False`` for a row
#: with no proposal, and for one whose codes the project does not buy.
priced: bool
#: P19 D2 - WHY this row was not evaluated: ``rounds`` / ``tokens`` when a cap cut the run
#: short, ``absent`` when the artefact is simply missing and no coverage file says otherwise,
#: and ``""`` for a row that WAS evaluated. Before this, "no artefact" could not be told from
@ -141,6 +147,14 @@ class RefusalVerdict:
approach_id: str
passed: bool
detail: str
#: P21 B3 - WHICH falsifier refused it, from ``validator.rejection_stage`` over the artefact's
#: own reason. ``stage0-baseline`` is the answer this whole order exists to make reachable: it
#: is the only stage that knows what the PROJECT buys, and before a project price schedule it
#: was skipped in every paid run. ``""`` when the arm produced no rejection to classify (it was
#: validated, or never evaluated) - an honest absence rather than a stage nobody reached.
#: REQUIRED without a default (``cost_baseline_anchored``'s rule): every construction site has
#: to say which falsifier spoke, and a default would let one of the three forget.
stage: str
@dataclass(frozen=True)
@ -172,6 +186,12 @@ class ContextSetVerdict:
#: P19 D2 - ``BudgetExceeded.kind`` when a cap cut the run short, ``""`` when nothing did, and
#: ``"absent"`` when the run wrote no coverage file at all (every run before today).
stop_reason: str
#: P21 B3 - whether the run's OWN stamp says stage 0 had a baseline to reconcile against, read
#: off ``provenance.cost_baseline_anchored`` rather than re-derived from a file: the judge
#: reports what the run DID, and a second resolution here would be free to disagree with it.
#: ``False`` when no artefact carried one (every round before P21). REQUIRED without a
#: default, for the reason ``ProvenanceStamp.cost_baseline_anchored`` is.
anchored: bool
tool_calls_seen: int
citations_seen: int
approach_rows_seen: int
@ -299,6 +319,19 @@ def score_context_set(
baseline = okf.load_optional_cost_baseline(str(base))
baseline_codes = set(baseline.items) if baseline is not None else set()
# P21 B3: the PROJECT's own price schedule, which is where the prices live — a road normal
# carries requirements and no amounts, so the ``load_optional_cost_baseline`` above finds
# nothing on every one of the four bases (measured). Read from the SET, which is the same file
# the run is given with ``--cost-baseline``, and read whether or not the run was anchored: that
# is what lets rounds written before the schedule existed be re-judged with this instrument.
# It never overwrites ``baseline_codes`` above — the hallucination arm's allowance is about
# what the BASE could ground, and merging the two would let a priced code launder a
# hallucinated one.
priced_codes: set[str] = set()
project_schedule = okf.load_optional_cost_baseline(str(context))
if project_schedule is not None:
priced_codes = set(project_schedule.items)
must_cite = {row["approach_id"]: row.get("concepts", []) for row in fasit.get("must_cite", [])}
refuse_ids = {row["approach_id"] for row in fasit.get("must_refuse", [])}
@ -353,6 +386,9 @@ def score_context_set(
str(_read_json(coverage_path).get("stop_reason", "")) if coverage_path.is_file() else ""
)
coverage_seen = coverage_path.is_file()
# P21 B3: read off the run's OWN stamp, accumulated over the artefacts below. A run is anchored
# or it is not, so ANY artefact saying so is the run saying so.
anchored = False
# ---- per approach ------------------------------------------------------------------------
rows: list[ApproachVerdict] = []
@ -385,6 +421,7 @@ def score_context_set(
requirement_source=_attributable(approach, declared_paths)[1],
requirement_hit=bool(set(_attributable(approach, declared_paths)[0]) & wanted),
prose_codes=(),
priced=False,
not_evaluated_reason=stop_reason or "absent",
ferdig=False,
)
@ -394,6 +431,7 @@ def score_context_set(
rows_seen += 1
payload = _read_json(proposal_path)
token_usage = max(token_usage, int(payload.get("provenance", {}).get("token_usage", 0)))
anchored = anchored or bool(payload.get("provenance", {}).get("cost_baseline_anchored"))
proposal = payload.get("proposal", {})
citations = payload.get("provenance", {}).get("citations", [])
citations_seen += len(citations)
@ -463,6 +501,7 @@ def score_context_set(
requirement_source=requirement_source,
requirement_hit=requirement_hit,
prose_codes=prose_codes,
priced=bool(codes) and all(c in priced_codes for c in codes),
not_evaluated_reason="",
ferdig=(
grounded
@ -491,20 +530,29 @@ def score_context_set(
commissioned = next((a for a in judged_approaches if a.id == rid), None)
refuse_codes = set(commissioned.affected_codes) if commissioned is not None else set()
leaked = sorted(refuse_codes & validated_codes)
# P21 B3: WHICH falsifier spoke, from the arm's own artefact. ``validator.rejection_stage``
# owns the classification because it owns the sentences (kø-(p)); this only reads the
# reason the run wrote down.
stage = ""
_, arm_outcome = _artefacts(outbox, run_id, rid)
if arm_outcome.is_file():
reason = str(_read_json(arm_outcome).get("reason", ""))
if reason:
stage = rejection_stage(reason)
if rid in validated_ids:
refusals.append(
RefusalVerdict(
rid, False, f"{rid} was VALIDATED - the base carries no ground for it"
rid, False, f"{rid} was VALIDATED - the base carries no ground for it", stage
)
)
elif leaked:
refusals.append(
RefusalVerdict(
rid, False, f"{rid}'s codes {leaked} ride inside a validated proposal"
rid, False, f"{rid}'s codes {leaked} ride inside a validated proposal", stage
)
)
else:
refusals.append(RefusalVerdict(rid, True, "no validated outcome rests on it"))
refusals.append(RefusalVerdict(rid, True, "no validated outcome rests on it", stage))
judged = [r for r in rows if r.approach_id not in refuse_ids]
return ContextSetVerdict(
@ -517,6 +565,7 @@ def score_context_set(
requirements_declared=tuple(declared_paths),
token_usage=token_usage,
stop_reason=stop_reason if coverage_seen else "absent",
anchored=anchored,
filter_calls=sum(1 for c in tool_calls if str(c.get("filter", ""))),
paged_calls=sum(
1 for c in tool_calls if int(c.get("offset", 0) or 0) or int(c.get("limit", 0) or 0)

View file

@ -670,6 +670,42 @@ def validate_proposal(
return ValidatedProposal(proposal=proposal, p10=p10, p50=p50, p90=p90, nominal_feasible=nominal)
#: P21 B3 — which falsifier a ``Rejection`` came from, keyed on the SENTENCES this module writes.
#:
#: It lives HERE, next to the wordings, and never in the judge: a second copy of "what a stage 0
#: refusal looks like" would be free to drift from the sentence the validator actually emits, and
#: the reader most likely to be misled is the one re-judging a paid run months later (kø-(p)).
#:
#: In PIPELINE order, which is also the only order that can be right: ``validate_proposal`` returns
#: at the FIRST failing stage, so one reason carries violations from exactly one of them.
_REJECTION_STAGES: Final = (
("stage0-baseline", ("cost baseline (", "tolerance around the baseline ")),
("stage0b-grounding", ("ungrounded identifier ",)),
("stage4-p90", ("exceeds P90 feasible",)),
("stage4b-nominal", ("exceeds the nominal feasible",)),
("stage5-method-cap", ("method cap",)),
)
def rejection_stage(reason: str) -> str:
"""Which stage of the deterministic gate wrote ``reason`` — ``"other"`` when none did.
A REPORT, never a gate: nothing branches on the answer, so an unrecognised sentence costs a
label and not a verdict. That is why ``"other"`` is an honest answer here and would not be one
inside the pipeline.
The measured reason it exists (P21 B3). Before a project price schedule, the ``must_refuse``
arm of every context set fell when it fell at all on stage 0b, P7's grounding check, which
can only say "this identifier is not in the delivered text". Stage 0 is the one stage that
knows what the PROJECT buys, and it was skipped in every paid run measured, because no road
normal ships a ``cost-baseline.json``. Saying which stage caught the falsification arm is how a
reader can tell an anchored refusal from an un-anchored one that happened to land."""
for stage, markers in _REJECTION_STAGES:
if any(marker in reason for marker in markers):
return stage
return "other"
def self_repair(
generate: Callable[[int], SavingsProposal],
*,

View file

@ -46,6 +46,7 @@ import pytest
from pydantic import ValidationError
from portfolio_optimiser import okf
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
from portfolio_optimiser.mandate import load_mandate
from portfolio_optimiser.stress import read_bundle_declarations
@ -492,3 +493,169 @@ def test_the_fasit_titles_are_distinct_not_the_collapsed_sources_title(set_dir:
"okf.parse_frontmatter collapsed these titles onto the sources block again — the P15 fix "
"in okf._frontmatter_from_text has regressed"
)
# --------------------------------------------------------------------------------------------
# (f) + (g) + (h): the set is ANCHORED (P21 B2).
#
# The measured reason these exist. Four paid rounds ran entirely un-anchored, because the only
# file loader reads ``cost-baseline.json`` out of the BUNDLE and no road normal carries prices — a
# vegnormal is knowledge, the price belongs to the PROJECT. With ``--cost-baseline`` the project
# supplies its own schedule, so the validator's stage 0 judges again: (f) every answerable approach
# has a line to reconcile against, and (g) the falsification arm has NONE, so the code it proposes
# is refused as "not in the project's cost baseline" — by stage 0, the one stage that can tell an
# invented line from a real one, instead of by the weaker downstream gates.
#
# (f) and (g) are SEPARATE arms rather than one loop over all approaches, because they are opposite
# claims about opposite rows: a single arm asserting "exactly the non-refuse codes are present"
# would go red for either defect and name neither.
# --------------------------------------------------------------------------------------------
def _set_baseline(set_dir: Path) -> CostBaseline:
return okf.load_cost_baseline_file(str(set_dir / "cost-baseline.json"))
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
def test_f_every_answerable_approach_has_a_cost_line(set_dir: Path) -> None:
"""Unconditional — the schedule is the PROJECT's and needs no knowledge base to read.
The total is asserted against the approach's own estimate as well as the code's presence:
``SavingsProposal`` refuses ``claimed_saving_nok > sum(affected_items.total)``, so a line that
exists but is smaller than the saving commissioned against it would make the approach
unbuildable a set that looks anchored and cannot be run.
"""
baseline = _set_baseline(set_dir)
assert 4 <= len(baseline.items) <= 8, (
f"{set_dir.name}: {len(baseline.items)} cost lines — the order asks for 4-8"
)
assert baseline.project_id == set_dir.name, (
f"{set_dir.name}: the schedule names project {baseline.project_id!r}"
)
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
refused = {row["approach_id"] for row in fasit["must_refuse"]}
for approach in load_mandate(set_dir / "mandate.json").approaches:
if approach.id in refused:
continue
missing = [c for c in approach.affected_codes if c not in baseline.items]
assert not missing, (
f"{set_dir.name}: {approach.id} is answerable but {missing} carry no line in the "
f"project's schedule ({sorted(baseline.items)})"
)
total = sum(
baseline.items[c].quantity * baseline.items[c].unit_cost
for c in approach.affected_codes
)
assert approach.claimed_saving_nok is not None
assert total >= approach.claimed_saving_nok, (
f"{set_dir.name}: {approach.id} claims {approach.claimed_saving_nok:g} against lines "
f"totalling {total:g} — no proposal on it can satisfy claimed <= total"
)
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
def test_g_the_falsification_arm_has_no_cost_line(set_dir: Path) -> None:
"""The ``must_refuse`` approach's code is ABSENT, so stage 0 is what catches it.
This is the half that makes the anchoring worth measuring rather than just present: rule U
already says the base carries no GROUND for that line, and P7's stage 0b says the identifier is
ungrounded in the delivered input but neither of those is the stage that knows what this
project actually buys. Stage 0 is, and it can only speak when the schedule exists.
"""
baseline = _set_baseline(set_dir)
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
by_id = {a.id: a for a in load_mandate(set_dir / "mandate.json").approaches}
assert fasit["must_refuse"], f"{set_dir.name} declares no falsification arm"
for row in fasit["must_refuse"]:
approach = by_id[row["approach_id"]]
carried = [c for c in approach.affected_codes if c in baseline.items]
assert not carried, (
f"{set_dir.name}: the falsification arm {approach.id} carries {carried} in the "
"project's schedule, so stage 0 would ACCEPT the line it exists to refuse"
)
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
def test_h_no_cost_line_smuggles_in_an_uncommissioned_requirement_number(set_dir: Path) -> None:
"""No line of the schedule is a reference number the base declares AND nobody commissions.
**The order words this arm as "no baseline code is a requirement number the base declares", and
that rule was FELLED BY MEASUREMENT before anything was built on it.** Measured 15.09 against
``okf.declared_reference_numbers`` over the four mounted bases: the four project-coded sets
carry 0 such codes, and ``kontrakt-sorasen-2027`` carries FIVE of five ``12.1``, ``12.12``,
``22.1``, ``52.11``, ``51.1`` are real R761 ``prosessnr``. That is not an accident in the set;
it is what R761 Prosesskoden IS. A Norwegian road contract's bill of quantities is priced BY
process code, so the project's schedule and the corpus's vocabulary share an identifier
namespace by design and the order's rule would have forced a rewrite of the ONE set P20's
decision (e) was chosen to preserve.
The COMPLEMENT keeps both: a schedule may price what the commission names, and may not
INTRODUCE a corpus identifier as a cost line nobody ordered. The order's own mutation still
bites swapping a code for ``12.11`` (a declared ``prosessnr`` no approach commissions) goes
red here while the five real process codes pass because an approach names each of them.
Needs the base (the vocabulary is the base's), so it SKIPS with the root named.
"""
baseline = _set_baseline(set_dir)
commissioned = {
code
for approach in load_mandate(set_dir / "mandate.json").approaches
for code in approach.affected_codes
}
declared: set[str] = set()
concepts = 0
for block in read_bundle_txt(set_dir / "bundle.txt"):
bundle = okf.navigate_bundle(str(_require_base(block)))
for f in bundle.context_files:
concepts += 1
declared |= set(okf.declared_reference_numbers(f))
assert concepts >= 100, (
f"{set_dir.name}: scanned {concepts} concepts — too few to be the base(s)"
)
assert declared, f"{set_dir.name}: the base(s) declare NO reference numbers — nothing to test"
smuggled = sorted(c for c in baseline.items if c in declared and c not in commissioned)
assert not smuggled, (
f"{set_dir.name}: cost line(s) {smuggled} are reference numbers the knowledge base "
"declares and no approach commissions — the schedule would be introducing the corpus's "
"own identifiers as prices nobody ordered"
)
def test_known_positive_f_a_missing_cost_line_is_caught(tmp_path: Path) -> None:
baseline = _set_baseline(_CONTEXT_ROOT / "gate-nordvik-2027")
assert "GATE-KRYSS-01" in baseline.items
stripped = CostBaseline(
project_id=baseline.project_id,
items={k: v for k, v in baseline.items.items() if k != "GATE-KRYSS-01"},
)
assert "GATE-KRYSS-01" not in stripped.items
def test_known_positive_g_a_line_for_the_falsification_arm_is_caught() -> None:
"""The order's mutation (i): give a4 a line, and (g)'s assertion must fail on this set."""
baseline = _set_baseline(_CONTEXT_ROOT / "gate-nordvik-2027")
priced = dict(baseline.items)
priced["GATE-GANG-ENHET"] = CostBaselineLine(quantity=6, unit_cost=50_000.0)
fasit = json.loads((_CONTEXT_ROOT / "gate-nordvik-2027" / "fasit.json").read_text("utf-8"))
by_id = {
a.id: a
for a in load_mandate(_CONTEXT_ROOT / "gate-nordvik-2027" / "mandate.json").approaches
}
for row in fasit["must_refuse"]:
carried = [c for c in by_id[row["approach_id"]].affected_codes if c in priced]
assert carried == ["GATE-GANG-ENHET"]
def test_known_positive_h_an_uncommissioned_requirement_number_is_caught() -> None:
"""The order's mutation (ii): swap a code for ``12.11``.
Driven against a KNOWN vocabulary rather than the mounted base, so this known-positive runs
unconditionally a control that skipped with the base would leave the arm's discriminator
unproven on exactly the machines that cannot run the arm.
"""
declared = {"12.1", "12.11", "12.12"}
commissioned = {"12.1", "12.12"}
assert sorted(c for c in {"12.1", "12.12"} if c in declared and c not in commissioned) == []
assert sorted(c for c in {"12.1", "12.11"} if c in declared and c not in commissioned) == [
"12.11"
]

View file

@ -0,0 +1,396 @@
"""P21 DEL A - the PROJECT carries the price, so a run against a road normal can be anchored.
**The measurement this closes.** Four paid stress rounds (P16/P18/P19/P17b/P20) ran ENTIRELY
un-anchored. The cause is one line: the only file loader reads ``cost-baseline.json`` out of the
BUNDLE directory (``okf.load_optional_cost_baseline``), and no vegnormal ships one N100, N200,
N500 and R761 are knowledge, and knowledge carries requirements, never amounts. The validator's
stage 0 the one stage that can tell an invented cost line from a line this project actually buys
was therefore skipped in every single one, and ``validated`` could not mean what it says: P20
G1/G2 measured real R761 process numbers (``12.11`` three times on Søråsen, ``1.1.1`` on Lindås)
validating with amounts nobody had anywhere.
``--cost-baseline FILE`` is PM decision (e), taken over three alternatives P20 wrote down: (a)
refusing every requirement-shaped code un-anchored would make the one realistic context set
unmeasurable, (b) ``--require-cost-baseline`` as a default would leave no stress test at all, and
(c) K2's priced schedule is refused by MAJOR-4's own stated limit. (e) puts the price where it
belongs with the project and stage 0 judges again.
Arms: (a) the library seam anchors * (b) control: without it the same base is un-anchored *
(c) it is the SAME baseline the validator is handed, so stage 0 really judges * (d) two sources for
one baseline are refused at the library seam * (e) the free trip anchors too * (f) a missing file
and (g) a malformed one are refused with ``load_cost_baseline``'s own error classes * (h)(i)(j)(k)
four CLI refusals BY NAME, each with an rc-0 control on an argv that would otherwise be accepted *
(l) the CLI wiring, measured on the stamp * (m) the notice * (n) EVERY base of an ``--across-bundle``
pass gets the SAME schedule.
"""
from __future__ import annotations
import json
import shutil
from pathlib import Path
from typing import Any
import pytest
from pydantic import ValidationError
from portfolio_optimiser import okf, run
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
from portfolio_optimiser.simulation import ScriptedChatClient
from portfolio_optimiser.validator import Rejection, validate_proposal
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
#: A base that ships NO ``cost-baseline.json`` — the whole class this flag exists for.
_UNPRICED_SOURCE = _EXAMPLES / "bygg-energi-mikro"
_IR_PROJECTION = {
"project_id": "P-KNOWLEDGE",
"measure": "PLACEHOLDER - authored by this test to satisfy the bundle contract",
"affected_items": [{"code": "KNOW-1", "quantity": 10, "unit_cost": 100.0}],
"claimed_saving_nok": 500.0,
}
def _runnable(tmp_path: Path, *, name: str = "base") -> str:
root = tmp_path / name
shutil.copytree(_UNPRICED_SOURCE, root)
(root / "validator-input.json").write_text(json.dumps(_IR_PROJECTION), encoding="utf-8")
assert not (root / "cost-baseline.json").exists(), "the fixture must be UNPRICED"
return str(root)
def _schedule(tmp_path: Path, *, name: str = "cost-baseline.json", **codes: float) -> str:
"""The PROJECT's own price schedule, written OUTSIDE every knowledge base — which is the whole
point: ``safe_resolve`` guards the bundle door, and a project's schedule is legitimately not in
a bundle."""
path = tmp_path / name
path.write_text(
json.dumps(
{
"project_id": "P-KNOWLEDGE",
"items": {
code: {"quantity": 10.0, "unit_cost": unit} for code, unit in codes.items()
},
}
),
encoding="utf-8",
)
return str(path)
def _scripted(sink: list[str] | None = None) -> Any:
return lambda role: ScriptedChatClient(sink=sink, role=role, default_reply="ok")
# --- (a)/(b)/(c)/(d)/(e) the library seam ---------------------------------------------------------
async def test_a_project_schedule_anchors_a_base_that_ships_none(tmp_path: Path) -> None:
bundle_dir = _runnable(tmp_path)
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
report = await run.run_project(
"P-KNOWLEDGE",
"local",
docs_dir=bundle_dir,
bundle_dir=bundle_dir,
cost_baseline=baseline,
live_dry_run=True,
)
assert isinstance(report, run.DryRunReport)
assert report.cost_baseline_anchored is True
async def test_control_without_the_schedule_the_same_base_is_unanchored(tmp_path: Path) -> None:
"""The discriminator for the arm above: the SAME base, the flag removed. Without this, "it is
anchored" could be a property of the fixture rather than of the file."""
bundle_dir = _runnable(tmp_path)
report = await run.run_project(
"P-KNOWLEDGE", "local", docs_dir=bundle_dir, bundle_dir=bundle_dir, live_dry_run=True
)
assert isinstance(report, run.DryRunReport)
assert report.cost_baseline_anchored is False
def test_the_supplied_schedule_is_what_stage_0_judges_against() -> None:
"""The stamp says anchored; this says the anchoring DOES something.
An implementation that read the file, stamped ``True`` and handed the validator ``None`` would
pass every other arm here the mutation the order names as (i). Stage 0 either refuses a code
the project does not buy or it does not, and that is the only thing worth having.
"""
baseline = CostBaseline(
project_id="P-KNOWLEDGE",
items={"RIGG": CostBaselineLine(quantity=10.0, unit_cost=1000.0)},
)
from portfolio_optimiser.ir import AffectedItem, SavingsProposal
invented = SavingsProposal(
project_id="P-KNOWLEDGE",
measure="m",
affected_items=[AffectedItem(code="INDEKS-01", quantity=10.0, unit_cost=1000.0)],
claimed_saving_nok=100.0,
)
outcome = validate_proposal(invented, baseline=baseline)
assert isinstance(outcome, Rejection)
assert "cost baseline" in outcome.reason
# The control: the SAME magnitudes on a code the project DOES buy are not refused by stage 0.
real = SavingsProposal(
project_id="P-KNOWLEDGE",
measure="m",
affected_items=[AffectedItem(code="RIGG", quantity=10.0, unit_cost=1000.0)],
claimed_saving_nok=100.0,
)
second = validate_proposal(real, baseline=baseline)
assert not (isinstance(second, Rejection) and "cost baseline" in second.reason)
async def test_two_sources_for_one_baseline_are_refused_at_the_library_seam(
tmp_path: Path,
) -> None:
"""(d) Checked in ``run_project`` and not only in the CLI: the library takes the same two
arguments, and a library caller must not reach a state the CLI refuses by name."""
bundle_dir = _runnable(tmp_path)
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
with pytest.raises(ValueError, match="two sources for one baseline"):
await run.run_project(
"P-KNOWLEDGE",
"local",
docs_dir=bundle_dir,
bundle_dir=bundle_dir,
cost_baseline=baseline,
derive_cost_baseline=True,
live_dry_run=True,
)
async def test_it_satisfies_the_anchoring_requirement(tmp_path: Path) -> None:
"""(e) ``--require-cost-baseline`` is the guarantee; this is one of the two ways to meet it."""
bundle_dir = _runnable(tmp_path)
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
report = await run.run_project(
"P-KNOWLEDGE",
"local",
docs_dir=bundle_dir,
bundle_dir=bundle_dir,
cost_baseline=baseline,
require_cost_baseline=True,
live_dry_run=True,
)
assert isinstance(report, run.DryRunReport)
assert report.cost_baseline_anchored is True
# --- (f)/(g) the loader's error classes -----------------------------------------------------------
def test_a_missing_schedule_is_refused(tmp_path: Path) -> None:
with pytest.raises(FileNotFoundError):
okf.load_cost_baseline_file(str(tmp_path / "nope.json"))
def test_a_malformed_schedule_is_refused(tmp_path: Path) -> None:
"""``load_cost_baseline``'s own classes, and there is no tolerant twin: this path exists only
because an operator NAMED a file, so degrading its absence would answer an explicit order with
a silently un-anchored run."""
bad = tmp_path / "bad.json"
bad.write_text('{"project_id": "P", "items": {"X": {"quantity": -1, "unit_cost": 0}}}', "utf-8")
with pytest.raises(ValidationError):
okf.load_cost_baseline_file(str(bad))
def test_the_bundle_loader_still_parses_through_the_same_seam(tmp_path: Path) -> None:
"""ONE parse, two doors (kø-(p)). What differs is the RESOLUTION: ``safe_resolve`` stays on the
bundle door alone, because a project's own schedule is legitimately outside every base."""
base = tmp_path / "b"
base.mkdir()
(base / "cost-baseline.json").write_text(
json.dumps({"project_id": "P", "items": {"X": {"quantity": 1, "unit_cost": 2}}}), "utf-8"
)
assert okf.load_cost_baseline(str(base)).items["X"].unit_cost == 2.0
# --- (h)/(i)/(j)/(k) four CLI refusals, each with an rc-0 control ----------------------------------
def test_cli_requires_a_knowledge_base(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
"""(h) On the road path the baseline IS ``Project.cost_items``, so a file there is a second
source for one fact with nothing to break the tie (``--require-cost-baseline``'s reason)."""
rc = run.main(["P1", "--docs-dir", "docs", "--cost-baseline", _schedule(tmp_path, RIGG=1000.0)])
assert rc == 1
assert "--cost-baseline" in capsys.readouterr().err
def test_cli_refuses_two_sources_by_name(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""(i) BY NAME so the operator hears WHICH two flags conflict, rather than a traceback."""
bundle_dir = _runnable(tmp_path)
schedule = _schedule(tmp_path, RIGG=1000.0)
assert (
run.main(
[
"P-KNOWLEDGE",
"--bundle-dir",
bundle_dir,
"--cost-baseline",
schedule,
"--live-dry-run",
]
)
== 0
), "the control argv must be ACCEPTED"
capsys.readouterr()
rc = run.main(
[
"P-KNOWLEDGE",
"--bundle-dir",
bundle_dir,
"--cost-baseline",
schedule,
"--derive-cost-baseline",
"--live-dry-run",
]
)
assert rc == 1
err = capsys.readouterr().err
assert "--cost-baseline" in err and "--derive-cost-baseline" in err
def test_cli_is_refused_in_portfolio_mode(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""(j) A portfolio pass keys on PROJECTS, each already anchored by its own ``cost_items``, so
ONE file could be right for at most one row out of N. BY NAME, its neighbours' reason."""
rc = run.main(["--portfolio", "--cost-baseline", _schedule(tmp_path, RIGG=1000.0)])
assert rc == 1
assert "--portfolio" in capsys.readouterr().err
def test_cli_is_refused_in_report_mode(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
"""(k) Report mode returns ABOVE every dispatch, so an omission from the allowlist is a SILENT
DROP the file would be accepted, nothing anchored, exit 0 (the F4 gap)."""
ledger = tmp_path / "ledger.json"
ledger.write_text("[]", encoding="utf-8")
assert run.main(["--report", "--ledger", str(ledger)]) == 0, "the control argv must be ACCEPTED"
capsys.readouterr()
rc = run.main(
["--report", "--ledger", str(ledger), "--cost-baseline", _schedule(tmp_path, RIGG=1000.0)]
)
assert rc == 1
assert "mode-exclusive" in capsys.readouterr().err
# --- (l)/(m) the CLI wiring and the notice --------------------------------------------------------
def test_cli_wiring_anchors_the_dry_run_and_says_so(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]
) -> None:
"""(l)+(m) The flag must REACH ``run_project``, and the operator must be able to see that it
did on the FREE trip. rc 0 alone proves neither, so the discriminators are the two lines: the
un-anchored warning is GONE and the positive line names the file and the count."""
monkeypatch.delenv("PORTFOLIO_MODEL_MAP", raising=False)
bundle_dir = _runnable(tmp_path)
argv = ["P-KNOWLEDGE", "--bundle-dir", bundle_dir, "--live-dry-run"]
assert run.main(argv) == 0
before = capsys.readouterr().out
assert "Cost baseline: NONE in the bundle" in before
assert "Cost baseline: 2 lines from" not in before
schedule = _schedule(tmp_path, RIGG=1000.0, ASFALT=250.0)
assert run.main([*argv, "--cost-baseline", schedule]) == 0
after = capsys.readouterr().out
assert "Cost baseline: NONE in the bundle" not in after
assert f"Cost baseline: 2 lines from {schedule}" in after
def test_the_notice_is_omitted_when_nobody_named_a_file() -> None:
"""Omission where it is unambiguous — there is exactly one way to supply a schedule, so
silence means nobody did (``cost_baseline_notice``'s rule, kept)."""
assert run.cost_baseline_source_notice(None, 0) is None
assert run.cost_baseline_source_notice("x.json", 1) == (
" Cost baseline: 1 line from x.json — the validator's stage 0 reconciles every proposed "
"cost line against this project's own schedule"
)
# --- (n) every base of a multi-base pass ----------------------------------------------------------
async def test_every_base_of_an_across_bundle_pass_gets_the_same_schedule(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""(n) ONE project has ONE price schedule, so it anchors EVERY base.
The order's mutation (ii) is "only the first base gets it", and the assert is therefore per
BASE: a recorder that stopped at the first call would pass on exactly that mutation the
vacuous-gate class this repo keeps measuring. Both calls are recorded and both must carry the
SAME object, because two reads of one file is already one resolution too many (-(p)).
"""
from portfolio_optimiser.mandate import Approach, Mandate
first = _runnable(tmp_path, name="one")
second = _runnable(tmp_path, name="two")
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
calls: list[dict[str, Any]] = []
async def _recorder(project_id: str, profile: Any = "local", **kwargs: Any) -> Any:
calls.append({"project_id": project_id, **kwargs})
class _Stub:
coverage: tuple[Any, ...] = ()
provenance = None
return _Stub()
monkeypatch.setattr(run, "run_project", _recorder)
mandate = Mandate(
objective="o",
success_criteria="s",
approaches=[
Approach(
id="a1",
label="one",
affected_codes=["RIGG"],
claimed_saving_nok=1.0,
bundle_id="one",
),
Approach(
id="a2",
label="two",
affected_codes=["RIGG"],
claimed_saving_nok=1.0,
bundle_id="two",
),
],
)
await run.run_mandate_across_bundles(mandate, [first, second], "local", cost_baseline=baseline)
assert len(calls) == 2, f"the dispatch ran {len(calls)} base(s), not two"
assert [c["bundle_dir"] for c in calls] == [first, second]
assert [c.get("cost_baseline") for c in calls] == [baseline, baseline], (
"a base was dispatched without the project's own schedule"
)
# The control: without the flag, no base is handed one — so the arm above measures the flag
# rather than a default.
calls.clear()
await run.run_mandate_across_bundles(mandate, [first, second], "local")
assert [c.get("cost_baseline") for c in calls] == [None, None]

View file

@ -209,3 +209,64 @@ def test_the_classification_is_reported_and_is_the_gates_own() -> None:
verdict = validate_proposal(_proposal(code), grounding=Grounding(_OFFERING + (code,)))
refused = isinstance(verdict, Rejection) and "has no identifier form" in verdict.reason
assert refused == (kind == "prose"), (code, kind, verdict)
# --------------------------------------------------------------------------------------------
# P21 B3: ``rejection_stage`` — which falsifier wrote a reason. A REPORT, never a gate: nothing
# branches on it, so an unrecognised sentence costs a label rather than a verdict. It lives beside
# the sentences it keys on, so the classifier and the wordings cannot drift apart (kø-(p)).
# --------------------------------------------------------------------------------------------
def test_p21_rejection_stage_names_each_stage_from_its_own_sentence() -> None:
from portfolio_optimiser.validator import rejection_stage
assert (
rejection_stage("unknown cost code 'X': not in project p's cost baseline (3 known codes)")
== "stage0-baseline"
)
assert (
rejection_stage(
"quantity 5 for cost code 'X' is outside the 5.0% tolerance around the baseline "
"quantity 9"
)
== "stage0-baseline"
)
assert (
rejection_stage("ungrounded identifier 'X': it appears nowhere in the input (9 chars)")
== "stage0b-grounding"
)
assert rejection_stage("claimed saving 9 exceeds P90 feasible 4") == "stage4-p90"
assert (
rejection_stage("claimed saving 9 exceeds the nominal feasible 4 at the items' stated")
== "stage4b-nominal"
)
assert (
rejection_stage("claimed 9 exceeds the energy_efficiency method cap 4 (stricter)")
== "stage5-method-cap"
)
# The honest answer for a sentence this module did not write — a label, never a verdict.
assert rejection_stage("something else entirely") == "other"
def test_p21_rejection_stage_is_keyed_on_the_sentences_the_validator_emits() -> None:
"""The control: the markers are not a private paraphrase but the text stage 0 really writes.
Without it the classifier could key on wording nothing emits and every arm above would still
be green the vacuous-gate class, on a reporter.
"""
from portfolio_optimiser.ir import AffectedItem, CostBaseline, CostBaselineLine, SavingsProposal
from portfolio_optimiser.validator import Rejection, rejection_stage, validate_proposal
baseline = CostBaseline(
project_id="proj", items={"REAL-1": CostBaselineLine(quantity=10.0, unit_cost=100.0)}
)
proposal = SavingsProposal(
project_id="proj",
measure="m",
affected_items=[AffectedItem(code="FAKE-1", quantity=10.0, unit_cost=100.0)],
claimed_saving_nok=100.0,
)
outcome = validate_proposal(proposal, baseline=baseline)
assert isinstance(outcome, Rejection)
assert rejection_stage(outcome.reason) == "stage0-baseline"

View file

@ -82,9 +82,27 @@ def _minibase(root: Path) -> Path:
return base
def _context(root: Path, *, must_refuse: bool = True) -> Path:
def _context(
root: Path, *, must_refuse: bool = True, schedule: dict[str, float] | None = None
) -> Path:
ctx = root / "ctx"
(ctx / "docs").mkdir(parents=True)
if schedule is not None:
# P21 B3: the PROJECT's own price schedule, which is the file a run is handed with
# ``--cost-baseline`` and the one the judge measures ``priced`` against.
(ctx / "cost-baseline.json").write_text(
json.dumps(
{
"project_id": "proj",
"items": {
code: {"quantity": 1.0, "unit_cost": unit}
for code, unit in schedule.items()
},
},
indent=2,
),
encoding="utf-8",
)
(ctx / "bundle.txt").write_text("name: minibase\nbundle_id: minibase\n", encoding="utf-8")
approaches = [
{
@ -149,6 +167,8 @@ def _write_outbox(
citation_snippet: str = "Body of the good one.",
decision: str = "validated",
tool_calls: list[dict[str, str]] | None = None,
anchored: bool = False,
reason: str = "no",
) -> None:
outbox.mkdir(parents=True, exist_ok=True)
codes = ["CODE-1"] if codes is None else codes
@ -181,7 +201,7 @@ def _write_outbox(
"role": "proposer",
"validator_decision": decision,
"token_usage": 10,
"cost_baseline_anchored": False,
"cost_baseline_anchored": anchored,
"bundle_id_source": None,
"external_calls": [],
},
@ -196,7 +216,7 @@ def _write_outbox(
"run_id": run_id,
"approach_id": approach_id,
"outcome_type": "validated" if decision == "validated" else "rejected",
**({"reason": "no"} if decision != "validated" else {}),
**({"reason": reason} if decision != "validated" else {}),
"checker_verdict": None,
"verdict_id": "vid",
},
@ -216,7 +236,12 @@ def _opened(path: str) -> list[dict[str, str]]:
def _judge(tmp_path: Path, **kw: object) -> stress.ContextSetVerdict:
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=bool(kw.pop("must_refuse", False)))
schedule = kw.pop("schedule", None)
ctx = _context(
tmp_path,
must_refuse=bool(kw.pop("must_refuse", False)),
schedule=schedule, # type: ignore[arg-type]
)
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", **kw) # type: ignore[arg-type]
return stress.score_context_set(ctx, outbox, "r1", base)
@ -629,3 +654,100 @@ def test_the_cli_refuses_to_guess_which_base_a_multi_base_outbox_is_for(
assert stress.main([*argv, "--bundle", "nowhere"]) == 1
assert "nowhere" in capsys.readouterr().err
# --------------------------------------------------------------------------------------------
# P21 B3: the judge reports what the run was ANCHORED on, which codes the project actually PRICES,
# and WHICH falsifier caught the falsification arm.
#
# The measured reason. Rounds 1-4 all ran un-anchored — a vegnormal ships no ``cost-baseline.json``
# and the only file loader read one out of the bundle — so stage 0 never spoke and the ``a4`` arm
# fell, when it fell, on P7's grounding check. "It was refused" and "the stage that knows what this
# project buys refused it" are different facts, and only the second is what anchoring bought.
# --------------------------------------------------------------------------------------------
def test_p21_priced_is_true_when_the_project_schedule_carries_the_code(tmp_path: Path) -> None:
verdict = _judge(tmp_path, schedule={"CODE-1": 2000.0})
assert verdict.approaches[0].priced is True
def test_p21_priced_is_false_for_a_code_the_project_does_not_buy(tmp_path: Path) -> None:
"""The discriminator: the SAME schedule, a proposal on a code it does not carry."""
verdict = _judge(tmp_path, schedule={"CODE-1": 2000.0}, codes=["CODE-9"])
assert verdict.approaches[0].priced is False
# ... and the control, so the arm cannot be green by measuring nothing.
assert (
_judge(tmp_path / "b", schedule={"CODE-9": 2000.0}, codes=["CODE-9"]).approaches[0].priced
is True
)
def test_p21_priced_is_false_without_a_project_schedule(tmp_path: Path) -> None:
"""Every round before P21: no schedule, so nothing is priced — reported, never guessed."""
assert _judge(tmp_path).approaches[0].priced is False
def test_p21_anchored_follows_the_runs_own_stamp(tmp_path: Path) -> None:
"""Read off ``provenance.cost_baseline_anchored``, not re-derived from the set's files.
Both arms, because a field that is constant is not a measurement: an artefact stamped
un-anchored must report ``False`` EVEN WHEN the set ships a schedule the judge says what the
run did, and a run that was never given the file is not anchored by the file existing.
"""
assert _judge(tmp_path, anchored=True, schedule={"CODE-1": 2000.0}).anchored is True
assert _judge(tmp_path / "b", anchored=False, schedule={"CODE-1": 2000.0}).anchored is False
def test_p21_the_falsification_arm_reports_which_stage_caught_it(tmp_path: Path) -> None:
"""``stage0-baseline`` is the answer anchoring buys; ``stage0b-grounding`` is what round 4 got.
Driven through the whole judge rather than through ``rejection_stage`` alone, because the
seam being gated is that the judge READS the arm's own outcome artefact — a classifier that
was never called would leave every arm reporting ``""`` and the arms below still green.
"""
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=True, schedule={"CODE-1": 2000.0})
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
_write_outbox(
outbox,
"r1",
approach_id="a4",
codes=["CODE-4"],
decision="rejected",
reason=("unknown cost code 'CODE-4': not in project proj's cost baseline (1 known codes)"),
)
verdict = stress.score_context_set(ctx, outbox, "r1", base)
assert [(r.approach_id, r.passed, r.stage) for r in verdict.must_refuse] == [
("a4", True, "stage0-baseline")
]
other = tmp_path / "b"
base2 = _minibase(other)
ctx2 = _context(other, must_refuse=True)
outbox2 = other / "out"
_write_outbox(outbox2, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
_write_outbox(
outbox2,
"r1",
approach_id="a4",
codes=["CODE-4"],
decision="rejected",
reason="ungrounded identifier 'CODE-4': it appears nowhere in the input (10 chars)",
)
assert stress.score_context_set(ctx2, outbox2, "r1", base2).must_refuse[0].stage == (
"stage0b-grounding"
)
def test_p21_a_validated_falsification_arm_reports_no_stage(tmp_path: Path) -> None:
"""``""`` when nothing refused it — an honest absence, never a stage nobody reached."""
base = _minibase(tmp_path)
ctx = _context(tmp_path, must_refuse=True)
outbox = tmp_path / "out"
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
_write_outbox(outbox, "r1", approach_id="a4", codes=["CODE-4"])
verdict = stress.score_context_set(ctx, outbox, "r1", base)
assert verdict.must_refuse[0].passed is False
assert verdict.must_refuse[0].stage == ""