feat(p21): the PROJECT carries the price, so a run against a road normal can be anchored
Four paid stress rounds ran entirely UN-ANCHORED, all of them, because the one file loader reads cost-baseline.json out of the BUNDLE and no vegnormal ships one: N100, N200, N500 and R761 are knowledge, and knowledge carries requirements, never amounts. The validator's stage 0 -- the one stage that tells an invented cost line from a line this project actually buys -- was skipped in every single run, so "validated" could not mean what it says. P20 G1/G2 measured real R761 process numbers (12.11 three times on Soraasen, 1.1.1 on Lindaas) validating with amounts nobody had anywhere. --cost-baseline FILE is PM decision (e), taken over the three alternatives P20 wrote down. A LOADED object, never a path (prepass_payload's rule): the CLI owns the file and loads it ONCE, so the notice, the stamp and every base of an --across-bundle pass all descend from one read. ONE parse, two doors -- load_cost_baseline delegates to load_cost_baseline_file -- while safe_resolve stays on the bundle door alone, because a project's own schedule is legitimately outside every base. No tolerant twin: this path exists only because an operator NAMED a file. DEL B: five anchored context sets, a1-a3 with their line and a4 with none, so stage 0 is what catches the falsification arm. THE ORDER'S OWN ARM (h) WAS FELLED BY MEASUREMENT: "no baseline code is a requirement number the base declares" is measured 0 of 4 on the project-coded sets and 5 of 5 on kontrakt-sorasen -- which is what R761 Prosesskoden IS, a bill of quantities priced BY process code. The complement keeps both, and the order's own mutation still bites. DEL B3: the judge reports anchored (off the run's own stamp), priced per row, and WHICH falsifier caught the falsification arm. Load-bearing MEASURED, five mutations all red against the WHOLE suite, green control 1850/5 (from 1809/5, superset, 0 removed), golden byte-unchanged: A3(i) the flag is read but the baseline is unused (3 red) . A3(ii) only the first base gets it (1) . A3(iii) report_forbidden drops it (1) . B2(i) a4 gets a line (1, arm (g) alone) . B2(ii) a code swapped to 12.11 (2, arms (f) and (h)). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
587480f050
commit
7b4f85d77c
20 changed files with 1259 additions and 15 deletions
70
CLAUDE.md
70
CLAUDE.md
|
|
@ -2848,6 +2848,76 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
|
||||||
en base som ikke lar seg løse faller tilbake til katalognavnet — `dimension_label`-presedensen
|
en base som ikke lar seg løse faller tilbake til katalognavnet — `dimension_label`-presedensen
|
||||||
ordrett: å annonsere skal ALDRI endre hvilken feil en operatør ser. Load-bearing MÅLT
|
ordrett: å annonsere skal ALDRI endre hvilken feil en operatør ser. Load-bearing MÅLT
|
||||||
(`tests/test_parse_error_feedback_loadbearing.py`, 7 armer).
|
(`tests/test_parse_error_feedback_loadbearing.py`, 7 armer).
|
||||||
|
- **PRISEN HØRER TIL PROSJEKTET, ikke til kunnskapsbasen — og ordrens egen B2-regel ble FELT av
|
||||||
|
måling (P21 DEL A+B, 15.09):** fire betalte runder (P16/P18/P19/P17b/P20) kjørte **UFORANKRET,
|
||||||
|
alle sammen**, fordi den ene fil-lasteren leser `cost-baseline.json` ut av BUNDLE-katalogen og
|
||||||
|
ingen vegnormal bærer et prisskjema: N100/N200/N500/R761 er KUNNSKAP, og kunnskap bærer krav,
|
||||||
|
aldri beløp. Validatorens **stadium 0** — det ENE stadiet som skiller en oppdiktet kostlinje fra
|
||||||
|
en linje dette prosjektet faktisk kjøper — ble derfor hoppet over i hver eneste av dem, og
|
||||||
|
`validated` kunne ikke bety det det sier: P20 G1/G2 målte EKTE R761-prosessnumre (`12.11` ×3 på
|
||||||
|
sorasen, `1.1.1` på lindaas) som validerte med beløp ingen hadde noe sted.
|
||||||
|
`--cost-baseline FILE` er PM-beslutning **(e)**, valgt over tre alternativer P20 skrev ned:
|
||||||
|
(a) nekt enhver kravformet kode uforankret ville gjort det ENE realistiske kontekstsettet
|
||||||
|
umålbart, (b) `--require-cost-baseline` som default ville etterlatt ingen stresstest, og
|
||||||
|
(c) K2s prisskjema er nektet av MAJOR-4s egen uttalte grense. **Et LASTET objekt, aldri en sti**
|
||||||
|
(`prepass_payload`-regelen, og `mandate=`/`dimension=` før den): CLI-en eier fila, biblioteks-
|
||||||
|
sømmen tar det validerte artefaktet — lastet ÉN gang, så notisen, stempelet og hver base i et
|
||||||
|
`--across-bundle`-pass stammer fra én lesing (kø-(p)). **ÉN parse, TO dører:**
|
||||||
|
`okf.load_cost_baseline_file` er hvor bytene tolkes og `load_cost_baseline` delegerer til den;
|
||||||
|
det som skiller er OPPLØSNINGEN — `safe_resolve` blir værende på bundle-døra ALENE, fordi et
|
||||||
|
prosjekts eget prisskjema legitimt ligger utenfor hver base. **Ingen tolerant tvilling**, og det
|
||||||
|
er motsatt av `load_optional_cost_baseline`: bundle-fila er fraværende som default (en base
|
||||||
|
skrevet før amendmentet er legitimt uforankret), mens denne stien finnes KUN fordi en operatør
|
||||||
|
NAVNGA en fil — å tolerere dens fravær ville besvart en eksplisitt ordre med en stille uforankret
|
||||||
|
kjøring (`load_mandate`-regelen). **Gjensidig utelukkende med `--derive-cost-baseline`, håndhevet
|
||||||
|
BEGGE steder:** CLI-en nekter VED NAVN (så operatøren hører hvilke to flagg som kolliderer) og
|
||||||
|
`run_project` reiser `ValueError` (så en bibliotekskaller ikke kan nå en tilstand CLI-en nekter).
|
||||||
|
**`cost_baseline_source_notice` er en ANDRE renderer ved siden av `cost_baseline_notice`, aldri
|
||||||
|
en utvidelse av den:** de sier ULIKE fakta og kan ikke være uenige (å oppgi flagget INNEBÆRER
|
||||||
|
forankret, så nøyaktig én av de to kan rendres), og den POSITIVE linja er et BEVISST avvik fra
|
||||||
|
omisjons-regelen — `proposal_review_notice`-avviket, av samme grunn: stillhet her er TVETYDIG,
|
||||||
|
for en operatør som ga et prisskjema kan ikke skille «fila di forankret kjøringen» fra «basen
|
||||||
|
hadde sin egen» eller fra «flagget ble droppet». Uten fil returnerer den `None`, så omisjonen
|
||||||
|
beholdes nøyaktig der den er entydig. **`--across-bundle` får SAMME skjema per base** (ett
|
||||||
|
prosjekt, ett prisskjema) — det er den ene ankringsparameteren som IKKE er en bundle-sak, og en
|
||||||
|
base med og en uten ville forankret halve kommisjonen mens stempelet rapporterte ankring for den
|
||||||
|
halvdelen som tilfeldigvis kjørte først. Tre partisjons-rader (`--portfolio` og
|
||||||
|
`report_forbidden` nekter VED NAVN med rc-0-kontroll; live-dry-run er en WIRING, så den frie
|
||||||
|
turen sier det samme som den betalte). **DEL B: fem forankrede kontekstsett** — hvert
|
||||||
|
`contexts/<sett>/cost-baseline.json` har 4–8 linjer, a1–a3 har SIN linje og a4/`must_refuse` har
|
||||||
|
INGEN, så stadium 0 er det som fanger falsifiseringsarmen. **Beskrivelser er BEVISST UTELATT fra
|
||||||
|
JSON-en:** `CostBaselineLine` har ingen slik nøkkel, pydantic ignorerer ekstra felt i stillhet,
|
||||||
|
og en fixtur hvis innhold droppes taust er en løgn — teksten bor i `mandate.json`s
|
||||||
|
label/description, som er nøklet på samme kode (kø-(p)). **ORDRENS ARM (h) BLE FELT AV MÅLING
|
||||||
|
FØR NOE BLE BYGGET PÅ DEN:** regelen «ingen baseline-kode er et kravnummer basen erklærer» er
|
||||||
|
MÅLT mot `okf.declared_reference_numbers` over de fire monterte basene — de fire prosjektkodede
|
||||||
|
settene bærer **0**, og `kontrakt-sorasen-2027` bærer **5 av 5** (`12.1`, `12.12`, `22.1`,
|
||||||
|
`52.11`, `51.1` er ekte R761-`prosessnr`). Det er ikke et uhell i settet; det er hva R761
|
||||||
|
Prosesskoden ER — en norsk vegkontrakts mengdebeskrivelse prises BY prosesskode — så ordrens
|
||||||
|
regel ville tvunget fram en omskriving av nettopp det settet beslutning (e) ble valgt for å
|
||||||
|
bevare. **KOMPLEMENTET beholder begge:** et prisskjema kan prise det kommisjonen NAVNGIR, og kan
|
||||||
|
ikke INNFØRE en korpus-identifikator som en kostlinje ingen bestilte. Ordrens egen mutasjon biter
|
||||||
|
fortsatt (bytt en kode til `12.11`, en erklært `prosessnr` ingen approach bestiller → arm (h)
|
||||||
|
rød). **DEL B3:** dommeren rapporterer `anchored` (lest av kjøringens EGET stempel, aldri
|
||||||
|
re-avledet), `priced` per rad (mot settets eget skjema, rapportert enten kjøringen var forankret
|
||||||
|
eller ei, så runde 1–4 kan dømmes med samme instrument) og `stage` per `must_refuse`-rad fra
|
||||||
|
`validator.rejection_stage`, som bor ved siden av setningene den nøkler på (kø-(p)) og er en
|
||||||
|
RAPPORT, aldri en gate — derfor er `"other"` et ærlig svar der og ville ikke vært det inne i
|
||||||
|
pipelinen. Load-bearing MÅLT (`tests/test_cost_baseline_flag_loadbearing.py` 15 armer +
|
||||||
|
`tests/test_context_sets_loadbearing.py` +11 + `tests/test_stress_judge_loadbearing.py` +6 +
|
||||||
|
`tests/test_prose_code_form_loadbearing.py` +2), **fem mutasjoner alle røde mot HELE suiten** +
|
||||||
|
grønn kontroll **1850/5** (fra 1809/5, supersett, 0 fjernet) og golden
|
||||||
|
`demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
|
||||||
|
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): A3(i) flagget leses men baselinen brukes ikke (3
|
||||||
|
røde) · A3(ii) bare første base får skjemaet (1) · A3(iii) `report_forbidden` slipper flagget (1)
|
||||||
|
· B2(i) a4 får en linje (1, arm (g) alene) · B2(ii) en kode byttet til `12.11` (2, armene (f) og
|
||||||
|
(h)). **Ærlighets-grenser, uttalt:** skjemaet når IKKE prompten — modellen må fortsatt oppgi
|
||||||
|
mengde og enhetspris selv, og stadium 0s avvisning navngir BASELINE-verdien, så Steg 5s
|
||||||
|
tilbakemating er hva som lar løkka konvergere (økt 94s måling); den hostede flaten er BEVISST
|
||||||
|
urørt (feltet er i ingen av hostings tre sett, så den generiske 400-en svarer og Fase 4es to
|
||||||
|
halvdeler står); portefølje-armen er ikke wiret (flagget er nektet der ved navn, så en utskrift
|
||||||
|
ville vært død kode); og alle beløp i de fem skjemaene er OPPDIKTEDE størrelsesordener, som hvert
|
||||||
|
setts eget `honesty`-felt sier.
|
||||||
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
||||||
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
||||||
|
|
||||||
|
|
|
||||||
19
README.md
19
README.md
|
|
@ -750,6 +750,25 @@ when the seam is detached, so the loop cannot silently degrade into theater.
|
||||||
--derive-cost-baseline --require-cost-baseline
|
--derive-cost-baseline --require-cost-baseline
|
||||||
```
|
```
|
||||||
|
|
||||||
|
- **The project's own price schedule** — `--cost-baseline FILE` (requires `--bundle-dir` or
|
||||||
|
`--across-bundle`). A knowledge base carries what is REQUIRED, not what things cost: a road
|
||||||
|
normal, a standard or a regulation has requirements and no amounts, so a run against one has
|
||||||
|
nothing for the validator's stage 0 to reconcile against and that stage is skipped. The price
|
||||||
|
belongs to the project, and this is where you hand it over: FILE is a `cost-baseline.json` — the
|
||||||
|
same `{"project_id": …, "items": {"<code>": {"quantity": …, "unit_cost": …}}}` shape a bundle may
|
||||||
|
ship — and it is used INSTEAD of one inside the base. With it, a proposal naming a cost line the
|
||||||
|
project does not buy is refused as a fabricated line, and one naming a real line with invented
|
||||||
|
magnitudes is refused with the real ones named, so the next attempt can correct.
|
||||||
|
|
||||||
|
In `--across-bundle` mode the same schedule anchors every base: one project, one price schedule.
|
||||||
|
Mutually exclusive with `--derive-cost-baseline` (two sources for one baseline), and it satisfies
|
||||||
|
`--require-cost-baseline`. A missing or malformed file refuses the run before anything starts.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
uv run python -m portfolio_optimiser.run PROSJEKT-1 --bundle-dir <bundle> \
|
||||||
|
--cost-baseline prisskjema.json --require-cost-baseline
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
`--explore` is refused together with `--mandate` — they are two sources of one mandate, and
|
`--explore` is refused together with `--mandate` — they are two sources of one mandate, and
|
||||||
merging would silently overwrite what you wrote. To seed an exploration with a domain expert's
|
merging would silently overwrite what you wrote. To seed an exploration with a domain expert's
|
||||||
|
|
|
||||||
25
contexts/dekke-og-kontrakt-lindaas-2027/cost-baseline.json
Normal file
25
contexts/dekke-og-kontrakt-lindaas-2027/cost-baseline.json
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
{
|
||||||
|
"project_id": "dekke-og-kontrakt-lindaas-2027",
|
||||||
|
"items": {
|
||||||
|
"LIND-FORST-01": {
|
||||||
|
"quantity": 24600.0,
|
||||||
|
"unit_cost": 285.0
|
||||||
|
},
|
||||||
|
"LIND-FILT-01": {
|
||||||
|
"quantity": 18400.0,
|
||||||
|
"unit_cost": 210.0
|
||||||
|
},
|
||||||
|
"LIND-RIGG-01": {
|
||||||
|
"quantity": 1.0,
|
||||||
|
"unit_cost": 5900000.0
|
||||||
|
},
|
||||||
|
"LIND-ASF-01": {
|
||||||
|
"quantity": 4100.0,
|
||||||
|
"unit_cost": 640.0
|
||||||
|
},
|
||||||
|
"LIND-GRFT-01": {
|
||||||
|
"quantity": 2050.0,
|
||||||
|
"unit_cost": 1380.0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
@ -50,7 +50,7 @@
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"honesty": "Prosjektet fv. 218 Lindaas er KONSTRUERT av meg: vegnummer, lengde, AADT og alle fire kostlinjene (LIND-FORST-01, LIND-FILT-01, LIND-RIGG-01, LIND-INDEKS-01) er oppdiktet, og beloepene er satte stoerrelsesordener. Kravene og prosessene i must_cite er lest ordrett ut av basenes egen frontmatter (n200-2024 for a1/a2, r761-2025 for a3). Verken N200 eller R761 baerer priser, saa kodene finnes ikke i noen av basene. Dette er det FOERSTE settet som spenner TO baser: a1/a2 rutes mot n200-2024 og a3/a4 mot r761-2025, og det er hele grunnen til at settet finnes -- P17b maaler at EN kommisjon kan kjoeres over flere kunnskapsbaser. Den fjerde tilnaermingen a4-indeksregulering og dens kostkode LIND-INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje INGEN av de to basene baerer grunnlaget for. MAALT 15.09: 'enhetspris' ble FORKASTET som anker fordi r761-2025 baerer ordet i 70 av 2 756 konsepter -- et anker som holder for ett sett med EN base holder ikke noedvendigvis for et sett med to.",
|
"honesty": "Prosjektet fv. 218 Lindaas er KONSTRUERT av meg: vegnummer, lengde, AADT og alle fire kostlinjene (LIND-FORST-01, LIND-FILT-01, LIND-RIGG-01, LIND-INDEKS-01) er oppdiktet, og beloepene er satte stoerrelsesordener. Kravene og prosessene i must_cite er lest ordrett ut av basenes egen frontmatter (n200-2024 for a1/a2, r761-2025 for a3). Verken N200 eller R761 baerer priser, saa kodene finnes ikke i noen av basene. Dette er det FOERSTE settet som spenner TO baser: a1/a2 rutes mot n200-2024 og a3/a4 mot r761-2025, og det er hele grunnen til at settet finnes -- P17b maaler at EN kommisjon kan kjoeres over flere kunnskapsbaser. Den fjerde tilnaermingen a4-indeksregulering og dens kostkode LIND-INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje INGEN av de to basene baerer grunnlaget for. MAALT 15.09: 'enhetspris' ble FORKASTET som anker fordi r761-2025 baerer ordet i 70 av 2 756 konsepter -- et anker som holder for ett sett med EN base holder ikke noedvendigvis for et sett med to. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema, og SAMME fil gjelder begge basene i multi-base-passet (ett prosjekt, ett prisskjema). Mengdene og enhetsprisene er OPPDIKTEDE stoerrelsesordener; verken N200 eller R761 baerer priser. a4s LIND-INDEKS-01 har ingen linje.",
|
||||||
"must_refuse": [
|
"must_refuse": [
|
||||||
{
|
{
|
||||||
"approach_id": "a4-indeksregulering",
|
"approach_id": "a4-indeksregulering",
|
||||||
|
|
|
||||||
25
contexts/fv412-dekkefornyelse-2027/cost-baseline.json
Normal file
25
contexts/fv412-dekkefornyelse-2027/cost-baseline.json
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
{
|
||||||
|
"project_id": "fv412-dekkefornyelse-2027",
|
||||||
|
"items": {
|
||||||
|
"DEKKE-ASF-01": {
|
||||||
|
"quantity": 10800.0,
|
||||||
|
"unit_cost": 1150.0
|
||||||
|
},
|
||||||
|
"DEKKE-BAER-01": {
|
||||||
|
"quantity": 6300.0,
|
||||||
|
"unit_cost": 980.0
|
||||||
|
},
|
||||||
|
"DEKKE-FROST-01": {
|
||||||
|
"quantity": 37800.0,
|
||||||
|
"unit_cost": 215.0
|
||||||
|
},
|
||||||
|
"DEKKE-GRV-01": {
|
||||||
|
"quantity": 12600.0,
|
||||||
|
"unit_cost": 145.0
|
||||||
|
},
|
||||||
|
"DEKKE-SKILT-01": {
|
||||||
|
"quantity": 1.0,
|
||||||
|
"unit_cost": 1450000.0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
@ -52,7 +52,7 @@
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"honesty": "Prosjektet fv. 412 dekkefornyelse er KONSTRUERT av meg: vegnummer, lengde, ÅDT og alle tre kostlinjene (DEKKE-ASF-01, DEKKE-BAER-01, DEKKE-FROST-01) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er lest ordrett ut av n200-2024-basens egen frontmatter. N200 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-tonnpris-asfalt og dens kostkode DEKKE-ASF-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
|
"honesty": "Prosjektet fv. 412 dekkefornyelse er KONSTRUERT av meg: vegnummer, lengde, ÅDT og alle tre kostlinjene (DEKKE-ASF-01, DEKKE-BAER-01, DEKKE-FROST-01) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er lest ordrett ut av n200-2024-basens egen frontmatter. N200 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-tonnpris-asfalt og dens kostkode DEKKE-ASF-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema. Mengdene og enhetsprisene der er OPPDIKTEDE stoerrelsesordener, satt slik at hver av a1-a3 har sin kostlinje og a4 IKKE har en; N200 baerer ingen priser, saa ingen av tallene er lest noe sted.",
|
||||||
"must_refuse": [
|
"must_refuse": [
|
||||||
{
|
{
|
||||||
"approach_id": "a4-tonnpris-asfalt",
|
"approach_id": "a4-tonnpris-asfalt",
|
||||||
|
|
|
||||||
25
contexts/gate-nordvik-2027/cost-baseline.json
Normal file
25
contexts/gate-nordvik-2027/cost-baseline.json
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
{
|
||||||
|
"project_id": "gate-nordvik-2027",
|
||||||
|
"items": {
|
||||||
|
"GATE-KRYSS-01": {
|
||||||
|
"quantity": 1.0,
|
||||||
|
"unit_cost": 9400000.0
|
||||||
|
},
|
||||||
|
"GATE-GANG-01": {
|
||||||
|
"quantity": 6.0,
|
||||||
|
"unit_cost": 310000.0
|
||||||
|
},
|
||||||
|
"GATE-KRYSS-02": {
|
||||||
|
"quantity": 8.0,
|
||||||
|
"unit_cost": 155000.0
|
||||||
|
},
|
||||||
|
"GATE-DEKKE-01": {
|
||||||
|
"quantity": 5400.0,
|
||||||
|
"unit_cost": 1250.0
|
||||||
|
},
|
||||||
|
"GATE-VA-01": {
|
||||||
|
"quantity": 900.0,
|
||||||
|
"unit_cost": 4800.0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
@ -52,7 +52,7 @@
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"honesty": "Prosjektet Nordvikgata er KONSTRUERT av meg: gatenavn, lengde, fartsgrense og alle tre kostlinjene (GATE-KRYSS-01, GATE-GANG-01, GATE-KRYSS-02) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er derimot lest ordrett ut av n100-2023-basens egen frontmatter. N100 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-enhetspris-gangfelt og dens kostkode GATE-GANG-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
|
"honesty": "Prosjektet Nordvikgata er KONSTRUERT av meg: gatenavn, lengde, fartsgrense og alle tre kostlinjene (GATE-KRYSS-01, GATE-GANG-01, GATE-KRYSS-02) er oppdiktet, og beløpene er satte størrelsesordener. Kravene i must_cite er derimot lest ordrett ut av n100-2023-basens egen frontmatter. N100 bærer ingen priser, så kodene finnes ikke i basen. P16: den fjerde tilnaermingen a4-enhetspris-gangfelt og dens kostkode GATE-GANG-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema. Mengdene og enhetsprisene der er OPPDIKTEDE stoerrelsesordener, satt slik at hver av a1-a3 har sin kostlinje og a4 IKKE har en; N100 baerer ingen priser, saa ingen av tallene er lest noe sted.",
|
||||||
"must_refuse": [
|
"must_refuse": [
|
||||||
{
|
{
|
||||||
"approach_id": "a4-enhetspris-gangfelt",
|
"approach_id": "a4-enhetspris-gangfelt",
|
||||||
|
|
|
||||||
25
contexts/kontrakt-sorasen-2027/cost-baseline.json
Normal file
25
contexts/kontrakt-sorasen-2027/cost-baseline.json
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
{
|
||||||
|
"project_id": "kontrakt-sorasen-2027",
|
||||||
|
"items": {
|
||||||
|
"12.1": {
|
||||||
|
"quantity": 1.0,
|
||||||
|
"unit_cost": 6400000.0
|
||||||
|
},
|
||||||
|
"12.12": {
|
||||||
|
"quantity": 22.0,
|
||||||
|
"unit_cost": 340000.0
|
||||||
|
},
|
||||||
|
"22.1": {
|
||||||
|
"quantity": 94000.0,
|
||||||
|
"unit_cost": 118.0
|
||||||
|
},
|
||||||
|
"52.11": {
|
||||||
|
"quantity": 41000.0,
|
||||||
|
"unit_cost": 265.0
|
||||||
|
},
|
||||||
|
"51.1": {
|
||||||
|
"quantity": 31000.0,
|
||||||
|
"unit_cost": 96.0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
@ -52,7 +52,7 @@
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"honesty": "Entreprisen Søråsen er KONSTRUERT av meg: kontraktsnavn, lengde, riggperiode og alle tre beløpene er oppdiktede størrelsesordener. Til forskjell fra de tre andre settene er affected_codes her EKTE prosessnumre fra R761 (12.1, 12.12, 22.1, 51.1, 52.11), lest ut av basens egen prosessnr-frontmatter — men R761 er en beskrivelsesstandard uten priser, så kodene er ekte mens beløpene ikke er det. must_cite er lest ordrett ut av basen. P16: den fjerde tilnaermingen a4-indeksregulering og dens kostkode INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
|
"honesty": "Entreprisen Søråsen er KONSTRUERT av meg: kontraktsnavn, lengde, riggperiode og alle tre beløpene er oppdiktede størrelsesordener. Til forskjell fra de tre andre settene er affected_codes her EKTE prosessnumre fra R761 (12.1, 12.12, 22.1, 51.1, 52.11), lest ut av basens egen prosessnr-frontmatter — men R761 er en beskrivelsesstandard uten priser, så kodene er ekte mens beløpene ikke er det. must_cite er lest ordrett ut av basen. P16: den fjerde tilnaermingen a4-indeksregulering og dens kostkode INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (5 linjer) — prosjektets prisskjema, kodet med de samme EKTE prosessnumrene, som er slik en norsk vegkontrakt faktisk prises. Mengdene og enhetsprisene er OPPDIKTEDE stoerrelsesordener; R761 baerer ingen priser, saa ingen av tallene er lest noe sted. a4s INDEKS-01 har ingen linje.",
|
||||||
"must_refuse": [
|
"must_refuse": [
|
||||||
{
|
{
|
||||||
"approach_id": "a4-indeksregulering",
|
"approach_id": "a4-indeksregulering",
|
||||||
|
|
|
||||||
29
contexts/tunnel-hauglia-2027/cost-baseline.json
Normal file
29
contexts/tunnel-hauglia-2027/cost-baseline.json
Normal file
|
|
@ -0,0 +1,29 @@
|
||||||
|
{
|
||||||
|
"project_id": "tunnel-hauglia-2027",
|
||||||
|
"items": {
|
||||||
|
"TUN-VENT-01": {
|
||||||
|
"quantity": 14.0,
|
||||||
|
"unit_cost": 465000.0
|
||||||
|
},
|
||||||
|
"TUN-FROST-01": {
|
||||||
|
"quantity": 360.0,
|
||||||
|
"unit_cost": 21500.0
|
||||||
|
},
|
||||||
|
"TUN-LYS-01": {
|
||||||
|
"quantity": 2400.0,
|
||||||
|
"unit_cost": 1850.0
|
||||||
|
},
|
||||||
|
"TUN-SPRENG-01": {
|
||||||
|
"quantity": 168000.0,
|
||||||
|
"unit_cost": 410.0
|
||||||
|
},
|
||||||
|
"TUN-SIKRING-01": {
|
||||||
|
"quantity": 2400.0,
|
||||||
|
"unit_cost": 6900.0
|
||||||
|
},
|
||||||
|
"TUN-PORTAL-01": {
|
||||||
|
"quantity": 2.0,
|
||||||
|
"unit_cost": 3850000.0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
@ -62,7 +62,7 @@
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"honesty": "Prosjektet Hauglia-tunnelen er KONSTRUERT av meg: navn, lengde, ÅDT og alle tre kostlinjene (TUN-VENT-01, TUN-FROST-01, TUN-LYS-01) er oppdiktet, og de tre beløpene er plausible størrelsesordener jeg har satt, ikke tall fra et prosjekt. Det som IKKE er konstruert er kravene: hver konsept-sti, tittel og kravnummer i must_cite er lest ordrett ut av n500-2024-basens egen frontmatter. N500 bærer ingen priser, så kodene finnes ikke i basen — settet er derfor bevisst IKKE kjørbart via --proposals-from-mandate. P16: den fjerde tilnaermingen a4-enhetspris-ventilator og dens kostkode TUN-VENT-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden.",
|
"honesty": "Prosjektet Hauglia-tunnelen er KONSTRUERT av meg: navn, lengde, ÅDT og alle tre kostlinjene (TUN-VENT-01, TUN-FROST-01, TUN-LYS-01) er oppdiktet, og de tre beløpene er plausible størrelsesordener jeg har satt, ikke tall fra et prosjekt. Det som IKKE er konstruert er kravene: hver konsept-sti, tittel og kravnummer i must_cite er lest ordrett ut av n500-2024-basens egen frontmatter. N500 bærer ingen priser, så kodene finnes ikke i basen — settet er derfor bevisst IKKE kjørbart via --proposals-from-mandate. P16: den fjerde tilnaermingen a4-enhetspris-ventilator og dens kostkode TUN-VENT-ENHET er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje basen ikke baerer grunnlaget for. Beloepet er en satt stoerrelsesorden. P21: settet har naa sin egen cost-baseline.json (6 linjer) — prosjektets prisskjema. Mengdene og enhetsprisene der er OPPDIKTEDE stoerrelsesordener, satt slik at hver av a1-a3 har sin kostlinje og a4 IKKE har en; N500 baerer ingen priser, saa ingen av tallene er lest noe sted.",
|
||||||
"must_refuse": [
|
"must_refuse": [
|
||||||
{
|
{
|
||||||
"approach_id": "a4-enhetspris-ventilator",
|
"approach_id": "a4-enhetspris-ventilator",
|
||||||
|
|
|
||||||
|
|
@ -1601,6 +1601,35 @@ def load_cost_baseline(bundle_dir: str, name: str = _COST_BASELINE) -> CostBasel
|
||||||
resolved = Path(safe_resolve(bundle_dir, name))
|
resolved = Path(safe_resolve(bundle_dir, name))
|
||||||
if not resolved.is_file():
|
if not resolved.is_file():
|
||||||
raise FileNotFoundError(f"cost baseline not found in bundle: {name!r}")
|
raise FileNotFoundError(f"cost baseline not found in bundle: {name!r}")
|
||||||
|
return load_cost_baseline_file(str(resolved))
|
||||||
|
|
||||||
|
|
||||||
|
def load_cost_baseline_file(path: str) -> CostBaseline:
|
||||||
|
"""Load a cost baseline from a file that is NOT inside a knowledge base: the PROJECT's own price
|
||||||
|
schedule (P21, ``--cost-baseline``).
|
||||||
|
|
||||||
|
The measured reason it exists. Four paid rounds (P16/P18/P19/P17b/P20) ran entirely UN-ANCHORED,
|
||||||
|
because the only file loader reads ``cost-baseline.json`` out of the bundle directory and no road
|
||||||
|
normal carries a price schedule: a vegnormal is KNOWLEDGE, and the price belongs to the PROJECT.
|
||||||
|
Stage 0 was therefore skipped in every one of them, and "validated" could not mean anything —
|
||||||
|
P20 G1/G2 measured real process numbers validating with invented amounts. This is the third door
|
||||||
|
into ``CostBaseline`` alongside the bundle file and ``derive_cost_baseline``, and the only one
|
||||||
|
whose input is the project rather than the corpus.
|
||||||
|
|
||||||
|
**The SAME parse, never a second one** (kø-(p)): ``load_cost_baseline`` resolves inside the
|
||||||
|
bundle and then delegates here, so the two doors cannot disagree about what a baseline file is.
|
||||||
|
What differs is the resolution — ``safe_resolve`` is the ONE in-/out-of-bundle test and stays on
|
||||||
|
the bundle door alone, because a project's own schedule is legitimately outside every base.
|
||||||
|
|
||||||
|
Fail-fast, the error CLASSES of ``load_cost_baseline``: a missing file raises
|
||||||
|
``FileNotFoundError`` and malformed content raises ``pydantic.ValidationError``. There is no
|
||||||
|
optional twin, and that is deliberate: the bundle file is absent by default (a base authored
|
||||||
|
before the amendment is legitimately un-anchored), whereas this path exists only because an
|
||||||
|
operator NAMED a file — tolerating its absence would answer an explicit order with a silently
|
||||||
|
un-anchored run (``load_mandate``'s rule)."""
|
||||||
|
resolved = Path(path)
|
||||||
|
if not resolved.is_file():
|
||||||
|
raise FileNotFoundError(f"cost baseline not found: {path!r}")
|
||||||
return CostBaseline.model_validate_json(resolved.read_text(encoding="utf-8"))
|
return CostBaseline.model_validate_json(resolved.read_text(encoding="utf-8"))
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -88,7 +88,7 @@ from portfolio_optimiser.generate import (
|
||||||
generate_via_llm,
|
generate_via_llm,
|
||||||
grounding_offer,
|
grounding_offer,
|
||||||
)
|
)
|
||||||
from portfolio_optimiser.ir import SavingsProposal
|
from portfolio_optimiser.ir import CostBaseline, SavingsProposal
|
||||||
from portfolio_optimiser.mandate import (
|
from portfolio_optimiser.mandate import (
|
||||||
OWN_PROPOSAL_ID,
|
OWN_PROPOSAL_ID,
|
||||||
Approach,
|
Approach,
|
||||||
|
|
@ -822,6 +822,36 @@ def cost_baseline_notice(anchored: bool) -> str | None:
|
||||||
return None if anchored else _UNANCHORED_NOTICE
|
return None if anchored else _UNANCHORED_NOTICE
|
||||||
|
|
||||||
|
|
||||||
|
def cost_baseline_source_notice(path: str | None, lines: int) -> str | None:
|
||||||
|
"""Render where this run's cost baseline came from, or ``None`` when nobody named a file (P21).
|
||||||
|
|
||||||
|
A SECOND renderer beside ``cost_baseline_notice``, never a widening of it, because the two say
|
||||||
|
DIFFERENT facts and cannot disagree: that one warns that stage 0 is SKIPPED, this one names the
|
||||||
|
file an operator chose and how many lines it carries. Supplying ``--cost-baseline`` implies
|
||||||
|
anchored, so exactly one of the two can ever render.
|
||||||
|
|
||||||
|
**A POSITIVE line, and that is a deliberate departure from the omission rule** its neighbours
|
||||||
|
follow (``cost_baseline_notice``, ``skipped_links_notice``, ``unkeyed_verdicts_notice``) — the
|
||||||
|
same departure ``proposal_review_notice`` makes, for the same reason. Silence here is
|
||||||
|
AMBIGUOUS: an operator who passed a project price schedule cannot tell "your file anchored this
|
||||||
|
run" from "the bundle happened to ship its own" or from "the flag was dropped somewhere", and
|
||||||
|
the whole point of the flag is that stage 0 now judges. Without a file the renderer returns
|
||||||
|
``None``, so the omission is kept exactly where it is unambiguous.
|
||||||
|
|
||||||
|
It takes the ALREADY-RESOLVED path and count rather than re-reading the file: a renderer that
|
||||||
|
opened it again would be a second resolution of the same fact, free to drift from the baseline
|
||||||
|
the run was actually given (``cost_baseline_notice``'s rule).
|
||||||
|
|
||||||
|
English, like every other line this CLI prints."""
|
||||||
|
if path is None:
|
||||||
|
return None
|
||||||
|
plural = "" if lines == 1 else "s"
|
||||||
|
return (
|
||||||
|
f" Cost baseline: {lines} line{plural} from {path} — the validator's stage 0 reconciles "
|
||||||
|
"every proposed cost line against this project's own schedule"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def grounding_offer_notice(offer: GroundingOffer | None) -> str | None:
|
def grounding_offer_notice(offer: GroundingOffer | None) -> str | None:
|
||||||
"""Render the one line that says what this run's delivered input can ground, or ``None`` when
|
"""Render the one line that says what this run's delivered input can ground, or ``None`` when
|
||||||
there is nothing to warn about (P8).
|
there is nothing to warn about (P8).
|
||||||
|
|
@ -1018,6 +1048,13 @@ async def run_project(
|
||||||
#: file loader, byte-identically.
|
#: file loader, byte-identically.
|
||||||
derive_cost_baseline: bool = False,
|
derive_cost_baseline: bool = False,
|
||||||
require_cost_baseline: bool = False,
|
require_cost_baseline: bool = False,
|
||||||
|
#: P21: the PROJECT's own price schedule, supplied by the caller instead of read out of the
|
||||||
|
#: knowledge base. A LOADED object, never a path — ``prepass_payload``'s rule, and
|
||||||
|
#: ``mandate=``/``dimension=``' before it: the CLI owns the file, the library seam takes the
|
||||||
|
#: validated artefact. ``None`` (the default) leaves every existing run on the bundle loader,
|
||||||
|
#: byte-identically. Mutually exclusive with ``derive_cost_baseline``: two sources for one
|
||||||
|
#: baseline would have to silently pick one, and the picked one would be a policy nobody wrote.
|
||||||
|
cost_baseline: CostBaseline | None = None,
|
||||||
dimension: Dimension | None = None,
|
dimension: Dimension | None = None,
|
||||||
store: VerdictStore | None = None,
|
store: VerdictStore | None = None,
|
||||||
verdict_dir: str | None = None,
|
verdict_dir: str | None = None,
|
||||||
|
|
@ -1081,6 +1118,17 @@ async def run_project(
|
||||||
semantics, so over a structural tie the resulting order is deterministic but arbitrary;
|
semantics, so over a structural tie the resulting order is deterministic but arbitrary;
|
||||||
retrieval *quality* arrives only with an embedder injected via ``embedder=`` or
|
retrieval *quality* arrives only with an embedder injected via ``embedder=`` or
|
||||||
``--embedder-config``. Default false keeps the structural ranking exactly as before."""
|
``--embedder-config``. Default false keeps the structural ranking exactly as before."""
|
||||||
|
# 0. Fail-fast: TWO sources for ONE baseline, refused rather than merged (P21). A run holding
|
||||||
|
# both would have to pick silently, and the picked one would be a policy nobody wrote down —
|
||||||
|
# ``--prepass-payload``/``--prepass-seed``'s rule, and ``portfolio_meter``/``meter_factory``'s
|
||||||
|
# before it. Checked HERE and not only at the CLI, because the library seam has the same two
|
||||||
|
# arguments and a library caller must not be able to reach a state the CLI refuses by name.
|
||||||
|
if cost_baseline is not None and derive_cost_baseline:
|
||||||
|
raise ValueError(
|
||||||
|
"cost_baseline and derive_cost_baseline are two sources for one baseline (a file the "
|
||||||
|
"caller supplies, and a schedule derived from the knowledge base); pass exactly one"
|
||||||
|
)
|
||||||
|
|
||||||
# 0. Fail-fast: an outbox write is byte-deterministic and keyed on run_id — no wall-clock default.
|
# 0. Fail-fast: an outbox write is byte-deterministic and keyed on run_id — no wall-clock default.
|
||||||
if outbox_dir is not None and run_id is None:
|
if outbox_dir is not None and run_id is None:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
|
|
@ -1173,8 +1221,15 @@ async def run_project(
|
||||||
# silently downgraded order, which is what ``load_mandate`` fail-fasts against. This one
|
# silently downgraded order, which is what ``load_mandate`` fail-fasts against. This one
|
||||||
# resolution serves BOTH the full run and the ``live_dry_run`` report below, so the dry-run
|
# resolution serves BOTH the full run and the ``live_dry_run`` report below, so the dry-run
|
||||||
# arm cannot drift away from what a real run would anchor on.
|
# arm cannot drift away from what a real run would anchor on.
|
||||||
|
# P21 takes precedence over BOTH bundle-side sources, and it is the only one whose input
|
||||||
|
# is the PROJECT rather than the corpus: a vegnormal is knowledge and carries no prices, so
|
||||||
|
# before this every paid run measured (P16/P18/P19/P17b/P20) was un-anchored and stage 0
|
||||||
|
# never ran. Mutual exclusion with ``derive_cost_baseline`` is enforced at the top of this
|
||||||
|
# function, so the ``if`` below is an ordering and not a silent pick.
|
||||||
baseline = (
|
baseline = (
|
||||||
okf.derive_cost_baseline(bundle, project_id=project_id)
|
cost_baseline
|
||||||
|
if cost_baseline is not None
|
||||||
|
else okf.derive_cost_baseline(bundle, project_id=project_id)
|
||||||
if derive_cost_baseline
|
if derive_cost_baseline
|
||||||
else okf.load_optional_cost_baseline(bundle_dir)
|
else okf.load_optional_cost_baseline(bundle_dir)
|
||||||
)
|
)
|
||||||
|
|
@ -2429,6 +2484,12 @@ async def run_mandate_across_bundles(
|
||||||
#: silent drop on the paid path while the free drill honoured them, which is the F4 class.
|
#: silent drop on the paid path while the free drill honoured them, which is the F4 class.
|
||||||
derive_cost_baseline: bool = False,
|
derive_cost_baseline: bool = False,
|
||||||
require_cost_baseline: bool = False,
|
require_cost_baseline: bool = False,
|
||||||
|
#: P21, and threaded UNCHANGED into every base: ONE project has ONE price schedule, so the same
|
||||||
|
#: baseline anchors each base's run. That is the one anchoring argument which is NOT a bundle
|
||||||
|
#: concern — the two above are read out of the base being handed over, this one is the project's
|
||||||
|
#: own, and giving base k a baseline and base k+1 none would anchor half a commission while the
|
||||||
|
#: stamp reported anchoring for the half that happened to run first.
|
||||||
|
cost_baseline: CostBaseline | None = None,
|
||||||
) -> MultiBaseResult:
|
) -> MultiBaseResult:
|
||||||
"""Evaluate ONE commission across SEVERAL knowledge bases — the multi-base dispatch (§ C.7).
|
"""Evaluate ONE commission across SEVERAL knowledge bases — the multi-base dispatch (§ C.7).
|
||||||
|
|
||||||
|
|
@ -2531,6 +2592,7 @@ async def run_mandate_across_bundles(
|
||||||
run_id=base_run_id or None,
|
run_id=base_run_id or None,
|
||||||
derive_cost_baseline=derive_cost_baseline,
|
derive_cost_baseline=derive_cost_baseline,
|
||||||
require_cost_baseline=require_cost_baseline,
|
require_cost_baseline=require_cost_baseline,
|
||||||
|
cost_baseline=cost_baseline,
|
||||||
notify=lambda verdict: minted_here.append(verdict.id),
|
notify=lambda verdict: minted_here.append(verdict.id),
|
||||||
verdict_input=verdict_input,
|
verdict_input=verdict_input,
|
||||||
verdict_dir=verdict_dir,
|
verdict_dir=verdict_dir,
|
||||||
|
|
@ -2967,6 +3029,20 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
"guesses: an unpriced or ambiguous schedule stops the run"
|
"guesses: an unpriced or ambiguous schedule stops the run"
|
||||||
),
|
),
|
||||||
)
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--cost-baseline",
|
||||||
|
default=None,
|
||||||
|
metavar="FILE",
|
||||||
|
help=(
|
||||||
|
"anchor the validator's stage 0 to THIS PROJECT's own price schedule (P21): FILE is a "
|
||||||
|
"cost-baseline.json (the same {project_id, items:{code:{quantity,unit_cost}}} shape a "
|
||||||
|
"bundle may ship) and it is used INSTEAD of one inside --bundle-dir. The price belongs "
|
||||||
|
"to the project, not to the knowledge base — a road normal carries requirements, never "
|
||||||
|
"amounts — so without this a run against one is un-anchored and stage 0 is skipped. "
|
||||||
|
"Applies to every base in --across-bundle mode: one project, one schedule. Mutually "
|
||||||
|
"exclusive with --derive-cost-baseline; satisfies --require-cost-baseline"
|
||||||
|
),
|
||||||
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--require-cost-baseline",
|
"--require-cost-baseline",
|
||||||
action="store_true",
|
action="store_true",
|
||||||
|
|
@ -3072,6 +3148,10 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
# Same reason, same rung: report mode returns above every run dispatch, so leaving it
|
# Same reason, same rung: report mode returns above every run dispatch, so leaving it
|
||||||
# out is a SILENT DROP of a guarantee the operator asked for by name.
|
# out is a SILENT DROP of a guarantee the operator asked for by name.
|
||||||
"--require-cost-baseline": args.require_cost_baseline,
|
"--require-cost-baseline": args.require_cost_baseline,
|
||||||
|
# P21, same rung and same reason: report mode returns ABOVE every dispatch that could
|
||||||
|
# honour a project price schedule, so an omission here would accept the file, anchor
|
||||||
|
# nothing, and exit 0 — a silent drop rather than a refusal (the F4 class).
|
||||||
|
"--cost-baseline": args.cost_baseline is not None,
|
||||||
# Same reason, one flag later: report mode returns above the S7b dispatch too.
|
# Same reason, one flag later: report mode returns above the S7b dispatch too.
|
||||||
"--proposals-from-mandate": args.proposals_from_mandate,
|
"--proposals-from-mandate": args.proposals_from_mandate,
|
||||||
"PROJECT_ID": args.project_id is not None,
|
"PROJECT_ID": args.project_id is not None,
|
||||||
|
|
@ -3183,6 +3263,11 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
# anchored by construction — so here the flag could only ever pass. BY NAME rather
|
# anchored by construction — so here the flag could only ever pass. BY NAME rather
|
||||||
# than falling through to the --bundle-dir requirement, its neighbours' reason.
|
# than falling through to the --bundle-dir requirement, its neighbours' reason.
|
||||||
"--require-cost-baseline": args.require_cost_baseline,
|
"--require-cost-baseline": args.require_cost_baseline,
|
||||||
|
# P21: ONE project's price schedule, and a portfolio pass keys on PROJECTS — each of
|
||||||
|
# which already carries its own ``cost_items`` and is anchored by construction. A
|
||||||
|
# single file could only ever be right for one row out of N. BY NAME rather than
|
||||||
|
# falling through to the --bundle-dir requirement, its neighbours' reason.
|
||||||
|
"--cost-baseline": args.cost_baseline,
|
||||||
# It reads ONE base's schedule and settles ONE commission against it, so it sits on the
|
# It reads ONE base's schedule and settles ONE commission against it, so it sits on the
|
||||||
# same side of the partition as the flag it requires. BY NAME rather than falling
|
# same side of the partition as the flag it requires. BY NAME rather than falling
|
||||||
# through to "requires --derive-cost-baseline": an operator who wrote --portfolio
|
# through to "requires --derive-cost-baseline": an operator who wrote --portfolio
|
||||||
|
|
@ -3364,6 +3449,37 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
)
|
)
|
||||||
return 1
|
return 1
|
||||||
|
|
||||||
|
# P21, and its neighbours' reason exactly: on the road path the baseline IS the project's own
|
||||||
|
# ``cost_items``, so a file here would be a SECOND source for one fact with nothing to break
|
||||||
|
# the tie — and a flag that can only ever be shadowed is a claim the surface makes about
|
||||||
|
# itself. Refused BY NAME rather than left to surface as a project lookup that ignored it.
|
||||||
|
if (
|
||||||
|
not args.portfolio
|
||||||
|
and not args.across_bundle
|
||||||
|
and args.cost_baseline is not None
|
||||||
|
and args.bundle_dir is None
|
||||||
|
):
|
||||||
|
print(
|
||||||
|
"run refused: --cost-baseline requires --bundle-dir or --across-bundle (the road path "
|
||||||
|
"is already anchored by the reference project's own cost_items, so a second schedule "
|
||||||
|
"there would have nothing to anchor that those do not)",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
# TWO sources for ONE baseline, refused rather than merged — the same decision
|
||||||
|
# ``run_project`` enforces at its own seam, said HERE by name so an operator who typed both
|
||||||
|
# hears which two flags conflict instead of getting a library ValueError's traceback. Not
|
||||||
|
# nested under either flag's branch, for the F4 reason its neighbours are not.
|
||||||
|
if args.cost_baseline is not None and args.derive_cost_baseline:
|
||||||
|
print(
|
||||||
|
"run refused: --cost-baseline and --derive-cost-baseline are two sources for one "
|
||||||
|
"baseline (a file you supply, and a schedule derived from a table in the knowledge "
|
||||||
|
"base); pass exactly one",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1
|
||||||
|
|
||||||
# The SEEDING arm's three refusals, placed ABOVE the replacing arm's block on purpose: given
|
# The SEEDING arm's three refusals, placed ABOVE the replacing arm's block on purpose: given
|
||||||
# both flags, the block below would answer with "--prepass-payload and --explore cannot be
|
# both flags, the block below would answer with "--prepass-payload and --explore cannot be
|
||||||
# combined", which names neither of the two flags the operator actually put in conflict. At
|
# combined", which names neither of the two flags the operator actually put in conflict. At
|
||||||
|
|
@ -3815,6 +3931,20 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
print(f"run refused: {exc}", file=sys.stderr)
|
print(f"run refused: {exc}", file=sys.stderr)
|
||||||
return 1
|
return 1
|
||||||
|
|
||||||
|
# The PROJECT's price schedule, loaded fail-fast alongside the commission and for the same
|
||||||
|
# reason: a baseline that cannot be read is not a run to start UN-anchored instead. Degrading
|
||||||
|
# it to "no baseline" would answer an operator who asked for stage 0 by name with a run in
|
||||||
|
# which stage 0 is skipped — ``load_mandate``'s rule, and the exact silence four paid rounds
|
||||||
|
# were measured inside. Loaded ONCE and passed as an object, so the notice below, the stamp
|
||||||
|
# and every base of an ``--across-bundle`` pass all descend from one read (kø-(p)).
|
||||||
|
cost_baseline: CostBaseline | None = None
|
||||||
|
if args.cost_baseline is not None:
|
||||||
|
try:
|
||||||
|
cost_baseline = okf.load_cost_baseline_file(args.cost_baseline)
|
||||||
|
except (FileNotFoundError, ValidationError, ValueError) as exc:
|
||||||
|
print(f"run refused: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
# The declared cut, loaded fail-fast alongside the commission and for the same reason: a
|
# The declared cut, loaded fail-fast alongside the commission and for the same reason: a
|
||||||
# payload that cannot be read is not a run to start with a navigating debate instead. Missing,
|
# payload that cannot be read is not a run to start with a navigating debate instead. Missing,
|
||||||
# not JSON, or not the shape the models require — all three land on the refusal surface with
|
# not JSON, or not the shape the models require — all three land on the refusal surface with
|
||||||
|
|
@ -4191,6 +4321,15 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
print(f"run refused: {exc}", file=sys.stderr)
|
print(f"run refused: {exc}", file=sys.stderr)
|
||||||
return 1
|
return 1
|
||||||
|
|
||||||
|
# P21, printed ONCE for the pass rather than per base, and that is the opposite placement
|
||||||
|
# from its neighbours below FOR A REASON: they read each base's OWN stamp, whereas ONE
|
||||||
|
# project has ONE price schedule and the same baseline anchors every base here. Per base it
|
||||||
|
# would read as N schedules, which is the claim this flag exists to deny.
|
||||||
|
source_notice = cost_baseline_source_notice(
|
||||||
|
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
|
||||||
|
)
|
||||||
|
if source_notice is not None:
|
||||||
|
print(source_notice)
|
||||||
if args.live_dry_run:
|
if args.live_dry_run:
|
||||||
for bundle_id, bundle_dir, project_id in resolved:
|
for bundle_id, bundle_dir, project_id in resolved:
|
||||||
try:
|
try:
|
||||||
|
|
@ -4210,6 +4349,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
max_tokens=args.max_tokens,
|
max_tokens=args.max_tokens,
|
||||||
derive_cost_baseline=args.derive_cost_baseline,
|
derive_cost_baseline=args.derive_cost_baseline,
|
||||||
require_cost_baseline=args.require_cost_baseline,
|
require_cost_baseline=args.require_cost_baseline,
|
||||||
|
cost_baseline=cost_baseline,
|
||||||
live_dry_run=True,
|
live_dry_run=True,
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|
@ -4269,6 +4409,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
),
|
),
|
||||||
derive_cost_baseline=args.derive_cost_baseline,
|
derive_cost_baseline=args.derive_cost_baseline,
|
||||||
require_cost_baseline=args.require_cost_baseline,
|
require_cost_baseline=args.require_cost_baseline,
|
||||||
|
cost_baseline=cost_baseline,
|
||||||
outbox_for=outbox_for,
|
outbox_for=outbox_for,
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|
@ -4302,6 +4443,15 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
multi=multi,
|
multi=multi,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# P21, printed ONCE for the pass rather than per base, and that is the opposite placement
|
||||||
|
# from its neighbours below FOR A REASON: they read each base's OWN stamp, whereas ONE
|
||||||
|
# project has ONE price schedule and the same baseline anchors every base here. Per base it
|
||||||
|
# would read as N schedules, which is the claim this flag exists to deny.
|
||||||
|
source_notice = cost_baseline_source_notice(
|
||||||
|
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
|
||||||
|
)
|
||||||
|
if source_notice is not None:
|
||||||
|
print(source_notice)
|
||||||
for bundle_run in multi.runs:
|
for bundle_run in multi.runs:
|
||||||
print(f"--- {bundle_run.bundle_id} ({bundle_run.project_id}) ---")
|
print(f"--- {bundle_run.bundle_id} ({bundle_run.project_id}) ---")
|
||||||
print(
|
print(
|
||||||
|
|
@ -4440,6 +4590,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
verdict_input=_verdict_input_from_args(args),
|
verdict_input=_verdict_input_from_args(args),
|
||||||
derive_cost_baseline=args.derive_cost_baseline,
|
derive_cost_baseline=args.derive_cost_baseline,
|
||||||
require_cost_baseline=args.require_cost_baseline,
|
require_cost_baseline=args.require_cost_baseline,
|
||||||
|
cost_baseline=cost_baseline,
|
||||||
mcp_servers=mcp_servers,
|
mcp_servers=mcp_servers,
|
||||||
live_dry_run=True,
|
live_dry_run=True,
|
||||||
)
|
)
|
||||||
|
|
@ -4470,6 +4621,14 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
notice = cost_baseline_notice(report.cost_baseline_anchored)
|
notice = cost_baseline_notice(report.cost_baseline_anchored)
|
||||||
if notice is not None:
|
if notice is not None:
|
||||||
print(notice)
|
print(notice)
|
||||||
|
# P21's positive half, on the FREE trip: the operator who named a project price schedule
|
||||||
|
# learns on the drill that it was read and how many lines it carries, rather than paying
|
||||||
|
# for a run to find out. Silence without the flag (omission, never an empty row).
|
||||||
|
source_notice = cost_baseline_source_notice(
|
||||||
|
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
|
||||||
|
)
|
||||||
|
if source_notice is not None:
|
||||||
|
print(source_notice)
|
||||||
# P8, printed next to the line it qualifies: "stage 0 is skipped" says the gate lost a
|
# P8, printed next to the line it qualifies: "stage 0 is skipped" says the gate lost a
|
||||||
# falsifier; this says what the input could have offered it instead. On the FREE trip, so
|
# falsifier; this says what the input could have offered it instead. On the FREE trip, so
|
||||||
# an operator learns a run cannot be grounded without paying three attempts to find out.
|
# an operator learns a run cannot be grounded without paying three attempts to find out.
|
||||||
|
|
@ -4519,6 +4678,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
semantic_retrieval=args.semantic_retrieval,
|
semantic_retrieval=args.semantic_retrieval,
|
||||||
derive_cost_baseline=args.derive_cost_baseline,
|
derive_cost_baseline=args.derive_cost_baseline,
|
||||||
require_cost_baseline=args.require_cost_baseline,
|
require_cost_baseline=args.require_cost_baseline,
|
||||||
|
cost_baseline=cost_baseline,
|
||||||
client_factory=scripted_client_factory,
|
client_factory=scripted_client_factory,
|
||||||
mandate=mandate,
|
mandate=mandate,
|
||||||
mcp_servers=mcp_servers,
|
mcp_servers=mcp_servers,
|
||||||
|
|
@ -4567,6 +4727,12 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
notice = cost_baseline_notice(result.provenance.cost_baseline_anchored)
|
notice = cost_baseline_notice(result.provenance.cost_baseline_anchored)
|
||||||
if notice is not None:
|
if notice is not None:
|
||||||
print(notice)
|
print(notice)
|
||||||
|
# Same renderer on the paid run, so the drill and the run it rehearses say the same thing.
|
||||||
|
source_notice = cost_baseline_source_notice(
|
||||||
|
args.cost_baseline, 0 if cost_baseline is None else len(cost_baseline.items)
|
||||||
|
)
|
||||||
|
if source_notice is not None:
|
||||||
|
print(source_notice)
|
||||||
# Same renderer on the full run, read off the run's OWN measurement: a run that spent every
|
# Same renderer on the full run, read off the run's OWN measurement: a run that spent every
|
||||||
# attempt being refused as ungrounded is exactly where the input-side fact costs the most.
|
# attempt being refused as ungrounded is exactly where the input-side fact costs the most.
|
||||||
offer_notice = grounding_offer_notice(result.grounding_offer)
|
offer_notice = grounding_offer_notice(result.grounding_offer)
|
||||||
|
|
|
||||||
|
|
@ -76,7 +76,7 @@ from typing import Any
|
||||||
|
|
||||||
from portfolio_optimiser import okf
|
from portfolio_optimiser import okf
|
||||||
from portfolio_optimiser.mandate import Mandate, load_mandate
|
from portfolio_optimiser.mandate import Mandate, load_mandate
|
||||||
from portfolio_optimiser.validator import classify_codes
|
from portfolio_optimiser.validator import classify_codes, rejection_stage
|
||||||
|
|
||||||
#: Where the vegnormal bases are mounted, unless ``--bundle-root`` says otherwise. Read at CALL
|
#: Where the vegnormal bases are mounted, unless ``--bundle-root`` says otherwise. Read at CALL
|
||||||
#: time (the ``shared_root()`` idiom) so a test or an operator can move the mount without a reimport.
|
#: time (the ``shared_root()`` idiom) so a test or an operator can move the mount without a reimport.
|
||||||
|
|
@ -125,6 +125,12 @@ class ApproachVerdict:
|
||||||
#: wrote one, and RE-DERIVED with the same classifier when it did not, so rounds 1 and 2 -
|
#: wrote one, and RE-DERIVED with the same classifier when it did not, so rounds 1 and 2 -
|
||||||
#: written before the field existed - can be re-judged with the same instrument.
|
#: written before the field existed - can be re-judged with the same instrument.
|
||||||
prose_codes: tuple[str, ...]
|
prose_codes: tuple[str, ...]
|
||||||
|
#: P21 B3 - this row's ``affected_item`` codes ARE lines of the project's own price schedule.
|
||||||
|
#: Measured against ``contexts/<set>/cost-baseline.json`` (the schedule the run is given with
|
||||||
|
#: ``--cost-baseline``) and reported whether or not the run was anchored, so rounds written
|
||||||
|
#: before the schedule existed can be re-judged with the same instrument. ``False`` for a row
|
||||||
|
#: with no proposal, and for one whose codes the project does not buy.
|
||||||
|
priced: bool
|
||||||
#: P19 D2 - WHY this row was not evaluated: ``rounds`` / ``tokens`` when a cap cut the run
|
#: P19 D2 - WHY this row was not evaluated: ``rounds`` / ``tokens`` when a cap cut the run
|
||||||
#: short, ``absent`` when the artefact is simply missing and no coverage file says otherwise,
|
#: short, ``absent`` when the artefact is simply missing and no coverage file says otherwise,
|
||||||
#: and ``""`` for a row that WAS evaluated. Before this, "no artefact" could not be told from
|
#: and ``""`` for a row that WAS evaluated. Before this, "no artefact" could not be told from
|
||||||
|
|
@ -141,6 +147,14 @@ class RefusalVerdict:
|
||||||
approach_id: str
|
approach_id: str
|
||||||
passed: bool
|
passed: bool
|
||||||
detail: str
|
detail: str
|
||||||
|
#: P21 B3 - WHICH falsifier refused it, from ``validator.rejection_stage`` over the artefact's
|
||||||
|
#: own reason. ``stage0-baseline`` is the answer this whole order exists to make reachable: it
|
||||||
|
#: is the only stage that knows what the PROJECT buys, and before a project price schedule it
|
||||||
|
#: was skipped in every paid run. ``""`` when the arm produced no rejection to classify (it was
|
||||||
|
#: validated, or never evaluated) - an honest absence rather than a stage nobody reached.
|
||||||
|
#: REQUIRED without a default (``cost_baseline_anchored``'s rule): every construction site has
|
||||||
|
#: to say which falsifier spoke, and a default would let one of the three forget.
|
||||||
|
stage: str
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
|
|
@ -172,6 +186,12 @@ class ContextSetVerdict:
|
||||||
#: P19 D2 - ``BudgetExceeded.kind`` when a cap cut the run short, ``""`` when nothing did, and
|
#: P19 D2 - ``BudgetExceeded.kind`` when a cap cut the run short, ``""`` when nothing did, and
|
||||||
#: ``"absent"`` when the run wrote no coverage file at all (every run before today).
|
#: ``"absent"`` when the run wrote no coverage file at all (every run before today).
|
||||||
stop_reason: str
|
stop_reason: str
|
||||||
|
#: P21 B3 - whether the run's OWN stamp says stage 0 had a baseline to reconcile against, read
|
||||||
|
#: off ``provenance.cost_baseline_anchored`` rather than re-derived from a file: the judge
|
||||||
|
#: reports what the run DID, and a second resolution here would be free to disagree with it.
|
||||||
|
#: ``False`` when no artefact carried one (every round before P21). REQUIRED without a
|
||||||
|
#: default, for the reason ``ProvenanceStamp.cost_baseline_anchored`` is.
|
||||||
|
anchored: bool
|
||||||
tool_calls_seen: int
|
tool_calls_seen: int
|
||||||
citations_seen: int
|
citations_seen: int
|
||||||
approach_rows_seen: int
|
approach_rows_seen: int
|
||||||
|
|
@ -299,6 +319,19 @@ def score_context_set(
|
||||||
baseline = okf.load_optional_cost_baseline(str(base))
|
baseline = okf.load_optional_cost_baseline(str(base))
|
||||||
baseline_codes = set(baseline.items) if baseline is not None else set()
|
baseline_codes = set(baseline.items) if baseline is not None else set()
|
||||||
|
|
||||||
|
# P21 B3: the PROJECT's own price schedule, which is where the prices live — a road normal
|
||||||
|
# carries requirements and no amounts, so the ``load_optional_cost_baseline`` above finds
|
||||||
|
# nothing on every one of the four bases (measured). Read from the SET, which is the same file
|
||||||
|
# the run is given with ``--cost-baseline``, and read whether or not the run was anchored: that
|
||||||
|
# is what lets rounds written before the schedule existed be re-judged with this instrument.
|
||||||
|
# It never overwrites ``baseline_codes`` above — the hallucination arm's allowance is about
|
||||||
|
# what the BASE could ground, and merging the two would let a priced code launder a
|
||||||
|
# hallucinated one.
|
||||||
|
priced_codes: set[str] = set()
|
||||||
|
project_schedule = okf.load_optional_cost_baseline(str(context))
|
||||||
|
if project_schedule is not None:
|
||||||
|
priced_codes = set(project_schedule.items)
|
||||||
|
|
||||||
must_cite = {row["approach_id"]: row.get("concepts", []) for row in fasit.get("must_cite", [])}
|
must_cite = {row["approach_id"]: row.get("concepts", []) for row in fasit.get("must_cite", [])}
|
||||||
refuse_ids = {row["approach_id"] for row in fasit.get("must_refuse", [])}
|
refuse_ids = {row["approach_id"] for row in fasit.get("must_refuse", [])}
|
||||||
|
|
||||||
|
|
@ -353,6 +386,9 @@ def score_context_set(
|
||||||
str(_read_json(coverage_path).get("stop_reason", "")) if coverage_path.is_file() else ""
|
str(_read_json(coverage_path).get("stop_reason", "")) if coverage_path.is_file() else ""
|
||||||
)
|
)
|
||||||
coverage_seen = coverage_path.is_file()
|
coverage_seen = coverage_path.is_file()
|
||||||
|
# P21 B3: read off the run's OWN stamp, accumulated over the artefacts below. A run is anchored
|
||||||
|
# or it is not, so ANY artefact saying so is the run saying so.
|
||||||
|
anchored = False
|
||||||
|
|
||||||
# ---- per approach ------------------------------------------------------------------------
|
# ---- per approach ------------------------------------------------------------------------
|
||||||
rows: list[ApproachVerdict] = []
|
rows: list[ApproachVerdict] = []
|
||||||
|
|
@ -385,6 +421,7 @@ def score_context_set(
|
||||||
requirement_source=_attributable(approach, declared_paths)[1],
|
requirement_source=_attributable(approach, declared_paths)[1],
|
||||||
requirement_hit=bool(set(_attributable(approach, declared_paths)[0]) & wanted),
|
requirement_hit=bool(set(_attributable(approach, declared_paths)[0]) & wanted),
|
||||||
prose_codes=(),
|
prose_codes=(),
|
||||||
|
priced=False,
|
||||||
not_evaluated_reason=stop_reason or "absent",
|
not_evaluated_reason=stop_reason or "absent",
|
||||||
ferdig=False,
|
ferdig=False,
|
||||||
)
|
)
|
||||||
|
|
@ -394,6 +431,7 @@ def score_context_set(
|
||||||
rows_seen += 1
|
rows_seen += 1
|
||||||
payload = _read_json(proposal_path)
|
payload = _read_json(proposal_path)
|
||||||
token_usage = max(token_usage, int(payload.get("provenance", {}).get("token_usage", 0)))
|
token_usage = max(token_usage, int(payload.get("provenance", {}).get("token_usage", 0)))
|
||||||
|
anchored = anchored or bool(payload.get("provenance", {}).get("cost_baseline_anchored"))
|
||||||
proposal = payload.get("proposal", {})
|
proposal = payload.get("proposal", {})
|
||||||
citations = payload.get("provenance", {}).get("citations", [])
|
citations = payload.get("provenance", {}).get("citations", [])
|
||||||
citations_seen += len(citations)
|
citations_seen += len(citations)
|
||||||
|
|
@ -463,6 +501,7 @@ def score_context_set(
|
||||||
requirement_source=requirement_source,
|
requirement_source=requirement_source,
|
||||||
requirement_hit=requirement_hit,
|
requirement_hit=requirement_hit,
|
||||||
prose_codes=prose_codes,
|
prose_codes=prose_codes,
|
||||||
|
priced=bool(codes) and all(c in priced_codes for c in codes),
|
||||||
not_evaluated_reason="",
|
not_evaluated_reason="",
|
||||||
ferdig=(
|
ferdig=(
|
||||||
grounded
|
grounded
|
||||||
|
|
@ -491,20 +530,29 @@ def score_context_set(
|
||||||
commissioned = next((a for a in judged_approaches if a.id == rid), None)
|
commissioned = next((a for a in judged_approaches if a.id == rid), None)
|
||||||
refuse_codes = set(commissioned.affected_codes) if commissioned is not None else set()
|
refuse_codes = set(commissioned.affected_codes) if commissioned is not None else set()
|
||||||
leaked = sorted(refuse_codes & validated_codes)
|
leaked = sorted(refuse_codes & validated_codes)
|
||||||
|
# P21 B3: WHICH falsifier spoke, from the arm's own artefact. ``validator.rejection_stage``
|
||||||
|
# owns the classification because it owns the sentences (kø-(p)); this only reads the
|
||||||
|
# reason the run wrote down.
|
||||||
|
stage = ""
|
||||||
|
_, arm_outcome = _artefacts(outbox, run_id, rid)
|
||||||
|
if arm_outcome.is_file():
|
||||||
|
reason = str(_read_json(arm_outcome).get("reason", ""))
|
||||||
|
if reason:
|
||||||
|
stage = rejection_stage(reason)
|
||||||
if rid in validated_ids:
|
if rid in validated_ids:
|
||||||
refusals.append(
|
refusals.append(
|
||||||
RefusalVerdict(
|
RefusalVerdict(
|
||||||
rid, False, f"{rid} was VALIDATED - the base carries no ground for it"
|
rid, False, f"{rid} was VALIDATED - the base carries no ground for it", stage
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
elif leaked:
|
elif leaked:
|
||||||
refusals.append(
|
refusals.append(
|
||||||
RefusalVerdict(
|
RefusalVerdict(
|
||||||
rid, False, f"{rid}'s codes {leaked} ride inside a validated proposal"
|
rid, False, f"{rid}'s codes {leaked} ride inside a validated proposal", stage
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
else:
|
else:
|
||||||
refusals.append(RefusalVerdict(rid, True, "no validated outcome rests on it"))
|
refusals.append(RefusalVerdict(rid, True, "no validated outcome rests on it", stage))
|
||||||
|
|
||||||
judged = [r for r in rows if r.approach_id not in refuse_ids]
|
judged = [r for r in rows if r.approach_id not in refuse_ids]
|
||||||
return ContextSetVerdict(
|
return ContextSetVerdict(
|
||||||
|
|
@ -517,6 +565,7 @@ def score_context_set(
|
||||||
requirements_declared=tuple(declared_paths),
|
requirements_declared=tuple(declared_paths),
|
||||||
token_usage=token_usage,
|
token_usage=token_usage,
|
||||||
stop_reason=stop_reason if coverage_seen else "absent",
|
stop_reason=stop_reason if coverage_seen else "absent",
|
||||||
|
anchored=anchored,
|
||||||
filter_calls=sum(1 for c in tool_calls if str(c.get("filter", ""))),
|
filter_calls=sum(1 for c in tool_calls if str(c.get("filter", ""))),
|
||||||
paged_calls=sum(
|
paged_calls=sum(
|
||||||
1 for c in tool_calls if int(c.get("offset", 0) or 0) or int(c.get("limit", 0) or 0)
|
1 for c in tool_calls if int(c.get("offset", 0) or 0) or int(c.get("limit", 0) or 0)
|
||||||
|
|
|
||||||
|
|
@ -670,6 +670,42 @@ def validate_proposal(
|
||||||
return ValidatedProposal(proposal=proposal, p10=p10, p50=p50, p90=p90, nominal_feasible=nominal)
|
return ValidatedProposal(proposal=proposal, p10=p10, p50=p50, p90=p90, nominal_feasible=nominal)
|
||||||
|
|
||||||
|
|
||||||
|
#: P21 B3 — which falsifier a ``Rejection`` came from, keyed on the SENTENCES this module writes.
|
||||||
|
#:
|
||||||
|
#: It lives HERE, next to the wordings, and never in the judge: a second copy of "what a stage 0
|
||||||
|
#: refusal looks like" would be free to drift from the sentence the validator actually emits, and
|
||||||
|
#: the reader most likely to be misled is the one re-judging a paid run months later (kø-(p)).
|
||||||
|
#:
|
||||||
|
#: In PIPELINE order, which is also the only order that can be right: ``validate_proposal`` returns
|
||||||
|
#: at the FIRST failing stage, so one reason carries violations from exactly one of them.
|
||||||
|
_REJECTION_STAGES: Final = (
|
||||||
|
("stage0-baseline", ("cost baseline (", "tolerance around the baseline ")),
|
||||||
|
("stage0b-grounding", ("ungrounded identifier ",)),
|
||||||
|
("stage4-p90", ("exceeds P90 feasible",)),
|
||||||
|
("stage4b-nominal", ("exceeds the nominal feasible",)),
|
||||||
|
("stage5-method-cap", ("method cap",)),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def rejection_stage(reason: str) -> str:
|
||||||
|
"""Which stage of the deterministic gate wrote ``reason`` — ``"other"`` when none did.
|
||||||
|
|
||||||
|
A REPORT, never a gate: nothing branches on the answer, so an unrecognised sentence costs a
|
||||||
|
label and not a verdict. That is why ``"other"`` is an honest answer here and would not be one
|
||||||
|
inside the pipeline.
|
||||||
|
|
||||||
|
The measured reason it exists (P21 B3). Before a project price schedule, the ``must_refuse``
|
||||||
|
arm of every context set fell — when it fell at all — on stage 0b, P7's grounding check, which
|
||||||
|
can only say "this identifier is not in the delivered text". Stage 0 is the one stage that
|
||||||
|
knows what the PROJECT buys, and it was skipped in every paid run measured, because no road
|
||||||
|
normal ships a ``cost-baseline.json``. Saying which stage caught the falsification arm is how a
|
||||||
|
reader can tell an anchored refusal from an un-anchored one that happened to land."""
|
||||||
|
for stage, markers in _REJECTION_STAGES:
|
||||||
|
if any(marker in reason for marker in markers):
|
||||||
|
return stage
|
||||||
|
return "other"
|
||||||
|
|
||||||
|
|
||||||
def self_repair(
|
def self_repair(
|
||||||
generate: Callable[[int], SavingsProposal],
|
generate: Callable[[int], SavingsProposal],
|
||||||
*,
|
*,
|
||||||
|
|
|
||||||
|
|
@ -46,6 +46,7 @@ import pytest
|
||||||
from pydantic import ValidationError
|
from pydantic import ValidationError
|
||||||
|
|
||||||
from portfolio_optimiser import okf
|
from portfolio_optimiser import okf
|
||||||
|
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
|
||||||
from portfolio_optimiser.mandate import load_mandate
|
from portfolio_optimiser.mandate import load_mandate
|
||||||
from portfolio_optimiser.stress import read_bundle_declarations
|
from portfolio_optimiser.stress import read_bundle_declarations
|
||||||
|
|
||||||
|
|
@ -492,3 +493,169 @@ def test_the_fasit_titles_are_distinct_not_the_collapsed_sources_title(set_dir:
|
||||||
"okf.parse_frontmatter collapsed these titles onto the sources block again — the P15 fix "
|
"okf.parse_frontmatter collapsed these titles onto the sources block again — the P15 fix "
|
||||||
"in okf._frontmatter_from_text has regressed"
|
"in okf._frontmatter_from_text has regressed"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
# (f) + (g) + (h): the set is ANCHORED (P21 B2).
|
||||||
|
#
|
||||||
|
# The measured reason these exist. Four paid rounds ran entirely un-anchored, because the only
|
||||||
|
# file loader reads ``cost-baseline.json`` out of the BUNDLE and no road normal carries prices — a
|
||||||
|
# vegnormal is knowledge, the price belongs to the PROJECT. With ``--cost-baseline`` the project
|
||||||
|
# supplies its own schedule, so the validator's stage 0 judges again: (f) every answerable approach
|
||||||
|
# has a line to reconcile against, and (g) the falsification arm has NONE, so the code it proposes
|
||||||
|
# is refused as "not in the project's cost baseline" — by stage 0, the one stage that can tell an
|
||||||
|
# invented line from a real one, instead of by the weaker downstream gates.
|
||||||
|
#
|
||||||
|
# (f) and (g) are SEPARATE arms rather than one loop over all approaches, because they are opposite
|
||||||
|
# claims about opposite rows: a single arm asserting "exactly the non-refuse codes are present"
|
||||||
|
# would go red for either defect and name neither.
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _set_baseline(set_dir: Path) -> CostBaseline:
|
||||||
|
return okf.load_cost_baseline_file(str(set_dir / "cost-baseline.json"))
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
|
def test_f_every_answerable_approach_has_a_cost_line(set_dir: Path) -> None:
|
||||||
|
"""Unconditional — the schedule is the PROJECT's and needs no knowledge base to read.
|
||||||
|
|
||||||
|
The total is asserted against the approach's own estimate as well as the code's presence:
|
||||||
|
``SavingsProposal`` refuses ``claimed_saving_nok > sum(affected_items.total)``, so a line that
|
||||||
|
exists but is smaller than the saving commissioned against it would make the approach
|
||||||
|
unbuildable — a set that looks anchored and cannot be run.
|
||||||
|
"""
|
||||||
|
baseline = _set_baseline(set_dir)
|
||||||
|
assert 4 <= len(baseline.items) <= 8, (
|
||||||
|
f"{set_dir.name}: {len(baseline.items)} cost lines — the order asks for 4-8"
|
||||||
|
)
|
||||||
|
assert baseline.project_id == set_dir.name, (
|
||||||
|
f"{set_dir.name}: the schedule names project {baseline.project_id!r}"
|
||||||
|
)
|
||||||
|
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
||||||
|
refused = {row["approach_id"] for row in fasit["must_refuse"]}
|
||||||
|
for approach in load_mandate(set_dir / "mandate.json").approaches:
|
||||||
|
if approach.id in refused:
|
||||||
|
continue
|
||||||
|
missing = [c for c in approach.affected_codes if c not in baseline.items]
|
||||||
|
assert not missing, (
|
||||||
|
f"{set_dir.name}: {approach.id} is answerable but {missing} carry no line in the "
|
||||||
|
f"project's schedule ({sorted(baseline.items)})"
|
||||||
|
)
|
||||||
|
total = sum(
|
||||||
|
baseline.items[c].quantity * baseline.items[c].unit_cost
|
||||||
|
for c in approach.affected_codes
|
||||||
|
)
|
||||||
|
assert approach.claimed_saving_nok is not None
|
||||||
|
assert total >= approach.claimed_saving_nok, (
|
||||||
|
f"{set_dir.name}: {approach.id} claims {approach.claimed_saving_nok:g} against lines "
|
||||||
|
f"totalling {total:g} — no proposal on it can satisfy claimed <= total"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
|
def test_g_the_falsification_arm_has_no_cost_line(set_dir: Path) -> None:
|
||||||
|
"""The ``must_refuse`` approach's code is ABSENT, so stage 0 is what catches it.
|
||||||
|
|
||||||
|
This is the half that makes the anchoring worth measuring rather than just present: rule U
|
||||||
|
already says the base carries no GROUND for that line, and P7's stage 0b says the identifier is
|
||||||
|
ungrounded in the delivered input — but neither of those is the stage that knows what this
|
||||||
|
project actually buys. Stage 0 is, and it can only speak when the schedule exists.
|
||||||
|
"""
|
||||||
|
baseline = _set_baseline(set_dir)
|
||||||
|
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
||||||
|
by_id = {a.id: a for a in load_mandate(set_dir / "mandate.json").approaches}
|
||||||
|
assert fasit["must_refuse"], f"{set_dir.name} declares no falsification arm"
|
||||||
|
for row in fasit["must_refuse"]:
|
||||||
|
approach = by_id[row["approach_id"]]
|
||||||
|
carried = [c for c in approach.affected_codes if c in baseline.items]
|
||||||
|
assert not carried, (
|
||||||
|
f"{set_dir.name}: the falsification arm {approach.id} carries {carried} in the "
|
||||||
|
"project's schedule, so stage 0 would ACCEPT the line it exists to refuse"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
|
def test_h_no_cost_line_smuggles_in_an_uncommissioned_requirement_number(set_dir: Path) -> None:
|
||||||
|
"""No line of the schedule is a reference number the base declares AND nobody commissions.
|
||||||
|
|
||||||
|
**The order words this arm as "no baseline code is a requirement number the base declares", and
|
||||||
|
that rule was FELLED BY MEASUREMENT before anything was built on it.** Measured 15.09 against
|
||||||
|
``okf.declared_reference_numbers`` over the four mounted bases: the four project-coded sets
|
||||||
|
carry 0 such codes, and ``kontrakt-sorasen-2027`` carries FIVE of five — ``12.1``, ``12.12``,
|
||||||
|
``22.1``, ``52.11``, ``51.1`` are real R761 ``prosessnr``. That is not an accident in the set;
|
||||||
|
it is what R761 Prosesskoden IS. A Norwegian road contract's bill of quantities is priced BY
|
||||||
|
process code, so the project's schedule and the corpus's vocabulary share an identifier
|
||||||
|
namespace by design — and the order's rule would have forced a rewrite of the ONE set P20's
|
||||||
|
decision (e) was chosen to preserve.
|
||||||
|
|
||||||
|
The COMPLEMENT keeps both: a schedule may price what the commission names, and may not
|
||||||
|
INTRODUCE a corpus identifier as a cost line nobody ordered. The order's own mutation still
|
||||||
|
bites — swapping a code for ``12.11`` (a declared ``prosessnr`` no approach commissions) goes
|
||||||
|
red here — while the five real process codes pass because an approach names each of them.
|
||||||
|
|
||||||
|
Needs the base (the vocabulary is the base's), so it SKIPS with the root named.
|
||||||
|
"""
|
||||||
|
baseline = _set_baseline(set_dir)
|
||||||
|
commissioned = {
|
||||||
|
code
|
||||||
|
for approach in load_mandate(set_dir / "mandate.json").approaches
|
||||||
|
for code in approach.affected_codes
|
||||||
|
}
|
||||||
|
declared: set[str] = set()
|
||||||
|
concepts = 0
|
||||||
|
for block in read_bundle_txt(set_dir / "bundle.txt"):
|
||||||
|
bundle = okf.navigate_bundle(str(_require_base(block)))
|
||||||
|
for f in bundle.context_files:
|
||||||
|
concepts += 1
|
||||||
|
declared |= set(okf.declared_reference_numbers(f))
|
||||||
|
assert concepts >= 100, (
|
||||||
|
f"{set_dir.name}: scanned {concepts} concepts — too few to be the base(s)"
|
||||||
|
)
|
||||||
|
assert declared, f"{set_dir.name}: the base(s) declare NO reference numbers — nothing to test"
|
||||||
|
smuggled = sorted(c for c in baseline.items if c in declared and c not in commissioned)
|
||||||
|
assert not smuggled, (
|
||||||
|
f"{set_dir.name}: cost line(s) {smuggled} are reference numbers the knowledge base "
|
||||||
|
"declares and no approach commissions — the schedule would be introducing the corpus's "
|
||||||
|
"own identifiers as prices nobody ordered"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_known_positive_f_a_missing_cost_line_is_caught(tmp_path: Path) -> None:
|
||||||
|
baseline = _set_baseline(_CONTEXT_ROOT / "gate-nordvik-2027")
|
||||||
|
assert "GATE-KRYSS-01" in baseline.items
|
||||||
|
stripped = CostBaseline(
|
||||||
|
project_id=baseline.project_id,
|
||||||
|
items={k: v for k, v in baseline.items.items() if k != "GATE-KRYSS-01"},
|
||||||
|
)
|
||||||
|
assert "GATE-KRYSS-01" not in stripped.items
|
||||||
|
|
||||||
|
|
||||||
|
def test_known_positive_g_a_line_for_the_falsification_arm_is_caught() -> None:
|
||||||
|
"""The order's mutation (i): give a4 a line, and (g)'s assertion must fail on this set."""
|
||||||
|
baseline = _set_baseline(_CONTEXT_ROOT / "gate-nordvik-2027")
|
||||||
|
priced = dict(baseline.items)
|
||||||
|
priced["GATE-GANG-ENHET"] = CostBaselineLine(quantity=6, unit_cost=50_000.0)
|
||||||
|
fasit = json.loads((_CONTEXT_ROOT / "gate-nordvik-2027" / "fasit.json").read_text("utf-8"))
|
||||||
|
by_id = {
|
||||||
|
a.id: a
|
||||||
|
for a in load_mandate(_CONTEXT_ROOT / "gate-nordvik-2027" / "mandate.json").approaches
|
||||||
|
}
|
||||||
|
for row in fasit["must_refuse"]:
|
||||||
|
carried = [c for c in by_id[row["approach_id"]].affected_codes if c in priced]
|
||||||
|
assert carried == ["GATE-GANG-ENHET"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_known_positive_h_an_uncommissioned_requirement_number_is_caught() -> None:
|
||||||
|
"""The order's mutation (ii): swap a code for ``12.11``.
|
||||||
|
|
||||||
|
Driven against a KNOWN vocabulary rather than the mounted base, so this known-positive runs
|
||||||
|
unconditionally — a control that skipped with the base would leave the arm's discriminator
|
||||||
|
unproven on exactly the machines that cannot run the arm.
|
||||||
|
"""
|
||||||
|
declared = {"12.1", "12.11", "12.12"}
|
||||||
|
commissioned = {"12.1", "12.12"}
|
||||||
|
assert sorted(c for c in {"12.1", "12.12"} if c in declared and c not in commissioned) == []
|
||||||
|
assert sorted(c for c in {"12.1", "12.11"} if c in declared and c not in commissioned) == [
|
||||||
|
"12.11"
|
||||||
|
]
|
||||||
|
|
|
||||||
396
tests/test_cost_baseline_flag_loadbearing.py
Normal file
396
tests/test_cost_baseline_flag_loadbearing.py
Normal file
|
|
@ -0,0 +1,396 @@
|
||||||
|
"""P21 DEL A - the PROJECT carries the price, so a run against a road normal can be anchored.
|
||||||
|
|
||||||
|
**The measurement this closes.** Four paid stress rounds (P16/P18/P19/P17b/P20) ran ENTIRELY
|
||||||
|
un-anchored. The cause is one line: the only file loader reads ``cost-baseline.json`` out of the
|
||||||
|
BUNDLE directory (``okf.load_optional_cost_baseline``), and no vegnormal ships one — N100, N200,
|
||||||
|
N500 and R761 are knowledge, and knowledge carries requirements, never amounts. The validator's
|
||||||
|
stage 0 — the one stage that can tell an invented cost line from a line this project actually buys
|
||||||
|
— was therefore skipped in every single one, and ``validated`` could not mean what it says: P20
|
||||||
|
G1/G2 measured real R761 process numbers (``12.11`` three times on Søråsen, ``1.1.1`` on Lindås)
|
||||||
|
validating with amounts nobody had anywhere.
|
||||||
|
|
||||||
|
``--cost-baseline FILE`` is PM decision (e), taken over three alternatives P20 wrote down: (a)
|
||||||
|
refusing every requirement-shaped code un-anchored would make the one realistic context set
|
||||||
|
unmeasurable, (b) ``--require-cost-baseline`` as a default would leave no stress test at all, and
|
||||||
|
(c) K2's priced schedule is refused by MAJOR-4's own stated limit. (e) puts the price where it
|
||||||
|
belongs — with the project — and stage 0 judges again.
|
||||||
|
|
||||||
|
Arms: (a) the library seam anchors * (b) control: without it the same base is un-anchored *
|
||||||
|
(c) it is the SAME baseline the validator is handed, so stage 0 really judges * (d) two sources for
|
||||||
|
one baseline are refused at the library seam * (e) the free trip anchors too * (f) a missing file
|
||||||
|
and (g) a malformed one are refused with ``load_cost_baseline``'s own error classes * (h)(i)(j)(k)
|
||||||
|
four CLI refusals BY NAME, each with an rc-0 control on an argv that would otherwise be accepted *
|
||||||
|
(l) the CLI wiring, measured on the stamp * (m) the notice * (n) EVERY base of an ``--across-bundle``
|
||||||
|
pass gets the SAME schedule.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import shutil
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
from pydantic import ValidationError
|
||||||
|
|
||||||
|
from portfolio_optimiser import okf, run
|
||||||
|
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
|
||||||
|
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||||
|
from portfolio_optimiser.validator import Rejection, validate_proposal
|
||||||
|
|
||||||
|
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||||
|
#: A base that ships NO ``cost-baseline.json`` — the whole class this flag exists for.
|
||||||
|
_UNPRICED_SOURCE = _EXAMPLES / "bygg-energi-mikro"
|
||||||
|
|
||||||
|
_IR_PROJECTION = {
|
||||||
|
"project_id": "P-KNOWLEDGE",
|
||||||
|
"measure": "PLACEHOLDER - authored by this test to satisfy the bundle contract",
|
||||||
|
"affected_items": [{"code": "KNOW-1", "quantity": 10, "unit_cost": 100.0}],
|
||||||
|
"claimed_saving_nok": 500.0,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _runnable(tmp_path: Path, *, name: str = "base") -> str:
|
||||||
|
root = tmp_path / name
|
||||||
|
shutil.copytree(_UNPRICED_SOURCE, root)
|
||||||
|
(root / "validator-input.json").write_text(json.dumps(_IR_PROJECTION), encoding="utf-8")
|
||||||
|
assert not (root / "cost-baseline.json").exists(), "the fixture must be UNPRICED"
|
||||||
|
return str(root)
|
||||||
|
|
||||||
|
|
||||||
|
def _schedule(tmp_path: Path, *, name: str = "cost-baseline.json", **codes: float) -> str:
|
||||||
|
"""The PROJECT's own price schedule, written OUTSIDE every knowledge base — which is the whole
|
||||||
|
point: ``safe_resolve`` guards the bundle door, and a project's schedule is legitimately not in
|
||||||
|
a bundle."""
|
||||||
|
path = tmp_path / name
|
||||||
|
path.write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"project_id": "P-KNOWLEDGE",
|
||||||
|
"items": {
|
||||||
|
code: {"quantity": 10.0, "unit_cost": unit} for code, unit in codes.items()
|
||||||
|
},
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
return str(path)
|
||||||
|
|
||||||
|
|
||||||
|
def _scripted(sink: list[str] | None = None) -> Any:
|
||||||
|
return lambda role: ScriptedChatClient(sink=sink, role=role, default_reply="ok")
|
||||||
|
|
||||||
|
|
||||||
|
# --- (a)/(b)/(c)/(d)/(e) the library seam ---------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
async def test_a_project_schedule_anchors_a_base_that_ships_none(tmp_path: Path) -> None:
|
||||||
|
bundle_dir = _runnable(tmp_path)
|
||||||
|
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
|
||||||
|
|
||||||
|
report = await run.run_project(
|
||||||
|
"P-KNOWLEDGE",
|
||||||
|
"local",
|
||||||
|
docs_dir=bundle_dir,
|
||||||
|
bundle_dir=bundle_dir,
|
||||||
|
cost_baseline=baseline,
|
||||||
|
live_dry_run=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert isinstance(report, run.DryRunReport)
|
||||||
|
assert report.cost_baseline_anchored is True
|
||||||
|
|
||||||
|
|
||||||
|
async def test_control_without_the_schedule_the_same_base_is_unanchored(tmp_path: Path) -> None:
|
||||||
|
"""The discriminator for the arm above: the SAME base, the flag removed. Without this, "it is
|
||||||
|
anchored" could be a property of the fixture rather than of the file."""
|
||||||
|
bundle_dir = _runnable(tmp_path)
|
||||||
|
|
||||||
|
report = await run.run_project(
|
||||||
|
"P-KNOWLEDGE", "local", docs_dir=bundle_dir, bundle_dir=bundle_dir, live_dry_run=True
|
||||||
|
)
|
||||||
|
|
||||||
|
assert isinstance(report, run.DryRunReport)
|
||||||
|
assert report.cost_baseline_anchored is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_supplied_schedule_is_what_stage_0_judges_against() -> None:
|
||||||
|
"""The stamp says anchored; this says the anchoring DOES something.
|
||||||
|
|
||||||
|
An implementation that read the file, stamped ``True`` and handed the validator ``None`` would
|
||||||
|
pass every other arm here — the mutation the order names as (i). Stage 0 either refuses a code
|
||||||
|
the project does not buy or it does not, and that is the only thing worth having.
|
||||||
|
"""
|
||||||
|
baseline = CostBaseline(
|
||||||
|
project_id="P-KNOWLEDGE",
|
||||||
|
items={"RIGG": CostBaselineLine(quantity=10.0, unit_cost=1000.0)},
|
||||||
|
)
|
||||||
|
from portfolio_optimiser.ir import AffectedItem, SavingsProposal
|
||||||
|
|
||||||
|
invented = SavingsProposal(
|
||||||
|
project_id="P-KNOWLEDGE",
|
||||||
|
measure="m",
|
||||||
|
affected_items=[AffectedItem(code="INDEKS-01", quantity=10.0, unit_cost=1000.0)],
|
||||||
|
claimed_saving_nok=100.0,
|
||||||
|
)
|
||||||
|
outcome = validate_proposal(invented, baseline=baseline)
|
||||||
|
assert isinstance(outcome, Rejection)
|
||||||
|
assert "cost baseline" in outcome.reason
|
||||||
|
# The control: the SAME magnitudes on a code the project DOES buy are not refused by stage 0.
|
||||||
|
real = SavingsProposal(
|
||||||
|
project_id="P-KNOWLEDGE",
|
||||||
|
measure="m",
|
||||||
|
affected_items=[AffectedItem(code="RIGG", quantity=10.0, unit_cost=1000.0)],
|
||||||
|
claimed_saving_nok=100.0,
|
||||||
|
)
|
||||||
|
second = validate_proposal(real, baseline=baseline)
|
||||||
|
assert not (isinstance(second, Rejection) and "cost baseline" in second.reason)
|
||||||
|
|
||||||
|
|
||||||
|
async def test_two_sources_for_one_baseline_are_refused_at_the_library_seam(
|
||||||
|
tmp_path: Path,
|
||||||
|
) -> None:
|
||||||
|
"""(d) Checked in ``run_project`` and not only in the CLI: the library takes the same two
|
||||||
|
arguments, and a library caller must not reach a state the CLI refuses by name."""
|
||||||
|
bundle_dir = _runnable(tmp_path)
|
||||||
|
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
|
||||||
|
|
||||||
|
with pytest.raises(ValueError, match="two sources for one baseline"):
|
||||||
|
await run.run_project(
|
||||||
|
"P-KNOWLEDGE",
|
||||||
|
"local",
|
||||||
|
docs_dir=bundle_dir,
|
||||||
|
bundle_dir=bundle_dir,
|
||||||
|
cost_baseline=baseline,
|
||||||
|
derive_cost_baseline=True,
|
||||||
|
live_dry_run=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def test_it_satisfies_the_anchoring_requirement(tmp_path: Path) -> None:
|
||||||
|
"""(e) ``--require-cost-baseline`` is the guarantee; this is one of the two ways to meet it."""
|
||||||
|
bundle_dir = _runnable(tmp_path)
|
||||||
|
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
|
||||||
|
|
||||||
|
report = await run.run_project(
|
||||||
|
"P-KNOWLEDGE",
|
||||||
|
"local",
|
||||||
|
docs_dir=bundle_dir,
|
||||||
|
bundle_dir=bundle_dir,
|
||||||
|
cost_baseline=baseline,
|
||||||
|
require_cost_baseline=True,
|
||||||
|
live_dry_run=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert isinstance(report, run.DryRunReport)
|
||||||
|
assert report.cost_baseline_anchored is True
|
||||||
|
|
||||||
|
|
||||||
|
# --- (f)/(g) the loader's error classes -----------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_missing_schedule_is_refused(tmp_path: Path) -> None:
|
||||||
|
with pytest.raises(FileNotFoundError):
|
||||||
|
okf.load_cost_baseline_file(str(tmp_path / "nope.json"))
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_malformed_schedule_is_refused(tmp_path: Path) -> None:
|
||||||
|
"""``load_cost_baseline``'s own classes, and there is no tolerant twin: this path exists only
|
||||||
|
because an operator NAMED a file, so degrading its absence would answer an explicit order with
|
||||||
|
a silently un-anchored run."""
|
||||||
|
bad = tmp_path / "bad.json"
|
||||||
|
bad.write_text('{"project_id": "P", "items": {"X": {"quantity": -1, "unit_cost": 0}}}', "utf-8")
|
||||||
|
with pytest.raises(ValidationError):
|
||||||
|
okf.load_cost_baseline_file(str(bad))
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_bundle_loader_still_parses_through_the_same_seam(tmp_path: Path) -> None:
|
||||||
|
"""ONE parse, two doors (kø-(p)). What differs is the RESOLUTION: ``safe_resolve`` stays on the
|
||||||
|
bundle door alone, because a project's own schedule is legitimately outside every base."""
|
||||||
|
base = tmp_path / "b"
|
||||||
|
base.mkdir()
|
||||||
|
(base / "cost-baseline.json").write_text(
|
||||||
|
json.dumps({"project_id": "P", "items": {"X": {"quantity": 1, "unit_cost": 2}}}), "utf-8"
|
||||||
|
)
|
||||||
|
assert okf.load_cost_baseline(str(base)).items["X"].unit_cost == 2.0
|
||||||
|
|
||||||
|
|
||||||
|
# --- (h)/(i)/(j)/(k) four CLI refusals, each with an rc-0 control ----------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_requires_a_knowledge_base(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
|
||||||
|
"""(h) On the road path the baseline IS ``Project.cost_items``, so a file there is a second
|
||||||
|
source for one fact with nothing to break the tie (``--require-cost-baseline``'s reason)."""
|
||||||
|
rc = run.main(["P1", "--docs-dir", "docs", "--cost-baseline", _schedule(tmp_path, RIGG=1000.0)])
|
||||||
|
|
||||||
|
assert rc == 1
|
||||||
|
assert "--cost-baseline" in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_refuses_two_sources_by_name(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
"""(i) BY NAME so the operator hears WHICH two flags conflict, rather than a traceback."""
|
||||||
|
bundle_dir = _runnable(tmp_path)
|
||||||
|
schedule = _schedule(tmp_path, RIGG=1000.0)
|
||||||
|
|
||||||
|
assert (
|
||||||
|
run.main(
|
||||||
|
[
|
||||||
|
"P-KNOWLEDGE",
|
||||||
|
"--bundle-dir",
|
||||||
|
bundle_dir,
|
||||||
|
"--cost-baseline",
|
||||||
|
schedule,
|
||||||
|
"--live-dry-run",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
== 0
|
||||||
|
), "the control argv must be ACCEPTED"
|
||||||
|
capsys.readouterr()
|
||||||
|
|
||||||
|
rc = run.main(
|
||||||
|
[
|
||||||
|
"P-KNOWLEDGE",
|
||||||
|
"--bundle-dir",
|
||||||
|
bundle_dir,
|
||||||
|
"--cost-baseline",
|
||||||
|
schedule,
|
||||||
|
"--derive-cost-baseline",
|
||||||
|
"--live-dry-run",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 1
|
||||||
|
err = capsys.readouterr().err
|
||||||
|
assert "--cost-baseline" in err and "--derive-cost-baseline" in err
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_is_refused_in_portfolio_mode(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
"""(j) A portfolio pass keys on PROJECTS, each already anchored by its own ``cost_items``, so
|
||||||
|
ONE file could be right for at most one row out of N. BY NAME, its neighbours' reason."""
|
||||||
|
rc = run.main(["--portfolio", "--cost-baseline", _schedule(tmp_path, RIGG=1000.0)])
|
||||||
|
|
||||||
|
assert rc == 1
|
||||||
|
assert "--portfolio" in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_is_refused_in_report_mode(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
|
||||||
|
"""(k) Report mode returns ABOVE every dispatch, so an omission from the allowlist is a SILENT
|
||||||
|
DROP — the file would be accepted, nothing anchored, exit 0 (the F4 gap)."""
|
||||||
|
ledger = tmp_path / "ledger.json"
|
||||||
|
ledger.write_text("[]", encoding="utf-8")
|
||||||
|
|
||||||
|
assert run.main(["--report", "--ledger", str(ledger)]) == 0, "the control argv must be ACCEPTED"
|
||||||
|
capsys.readouterr()
|
||||||
|
|
||||||
|
rc = run.main(
|
||||||
|
["--report", "--ledger", str(ledger), "--cost-baseline", _schedule(tmp_path, RIGG=1000.0)]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 1
|
||||||
|
assert "mode-exclusive" in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
# --- (l)/(m) the CLI wiring and the notice --------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_wiring_anchors_the_dry_run_and_says_so(
|
||||||
|
tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
"""(l)+(m) The flag must REACH ``run_project``, and the operator must be able to see that it
|
||||||
|
did on the FREE trip. rc 0 alone proves neither, so the discriminators are the two lines: the
|
||||||
|
un-anchored warning is GONE and the positive line names the file and the count."""
|
||||||
|
monkeypatch.delenv("PORTFOLIO_MODEL_MAP", raising=False)
|
||||||
|
bundle_dir = _runnable(tmp_path)
|
||||||
|
argv = ["P-KNOWLEDGE", "--bundle-dir", bundle_dir, "--live-dry-run"]
|
||||||
|
|
||||||
|
assert run.main(argv) == 0
|
||||||
|
before = capsys.readouterr().out
|
||||||
|
assert "Cost baseline: NONE in the bundle" in before
|
||||||
|
assert "Cost baseline: 2 lines from" not in before
|
||||||
|
|
||||||
|
schedule = _schedule(tmp_path, RIGG=1000.0, ASFALT=250.0)
|
||||||
|
assert run.main([*argv, "--cost-baseline", schedule]) == 0
|
||||||
|
after = capsys.readouterr().out
|
||||||
|
|
||||||
|
assert "Cost baseline: NONE in the bundle" not in after
|
||||||
|
assert f"Cost baseline: 2 lines from {schedule}" in after
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_notice_is_omitted_when_nobody_named_a_file() -> None:
|
||||||
|
"""Omission where it is unambiguous — there is exactly one way to supply a schedule, so
|
||||||
|
silence means nobody did (``cost_baseline_notice``'s rule, kept)."""
|
||||||
|
assert run.cost_baseline_source_notice(None, 0) is None
|
||||||
|
assert run.cost_baseline_source_notice("x.json", 1) == (
|
||||||
|
" Cost baseline: 1 line from x.json — the validator's stage 0 reconciles every proposed "
|
||||||
|
"cost line against this project's own schedule"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# --- (n) every base of a multi-base pass ----------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
async def test_every_base_of_an_across_bundle_pass_gets_the_same_schedule(
|
||||||
|
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||||
|
) -> None:
|
||||||
|
"""(n) ONE project has ONE price schedule, so it anchors EVERY base.
|
||||||
|
|
||||||
|
The order's mutation (ii) is "only the first base gets it", and the assert is therefore per
|
||||||
|
BASE: a recorder that stopped at the first call would pass on exactly that mutation — the
|
||||||
|
vacuous-gate class this repo keeps measuring. Both calls are recorded and both must carry the
|
||||||
|
SAME object, because two reads of one file is already one resolution too many (kø-(p)).
|
||||||
|
"""
|
||||||
|
from portfolio_optimiser.mandate import Approach, Mandate
|
||||||
|
|
||||||
|
first = _runnable(tmp_path, name="one")
|
||||||
|
second = _runnable(tmp_path, name="two")
|
||||||
|
baseline = okf.load_cost_baseline_file(_schedule(tmp_path, RIGG=1000.0))
|
||||||
|
calls: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
async def _recorder(project_id: str, profile: Any = "local", **kwargs: Any) -> Any:
|
||||||
|
calls.append({"project_id": project_id, **kwargs})
|
||||||
|
|
||||||
|
class _Stub:
|
||||||
|
coverage: tuple[Any, ...] = ()
|
||||||
|
provenance = None
|
||||||
|
|
||||||
|
return _Stub()
|
||||||
|
|
||||||
|
monkeypatch.setattr(run, "run_project", _recorder)
|
||||||
|
|
||||||
|
mandate = Mandate(
|
||||||
|
objective="o",
|
||||||
|
success_criteria="s",
|
||||||
|
approaches=[
|
||||||
|
Approach(
|
||||||
|
id="a1",
|
||||||
|
label="one",
|
||||||
|
affected_codes=["RIGG"],
|
||||||
|
claimed_saving_nok=1.0,
|
||||||
|
bundle_id="one",
|
||||||
|
),
|
||||||
|
Approach(
|
||||||
|
id="a2",
|
||||||
|
label="two",
|
||||||
|
affected_codes=["RIGG"],
|
||||||
|
claimed_saving_nok=1.0,
|
||||||
|
bundle_id="two",
|
||||||
|
),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
await run.run_mandate_across_bundles(mandate, [first, second], "local", cost_baseline=baseline)
|
||||||
|
|
||||||
|
assert len(calls) == 2, f"the dispatch ran {len(calls)} base(s), not two"
|
||||||
|
assert [c["bundle_dir"] for c in calls] == [first, second]
|
||||||
|
assert [c.get("cost_baseline") for c in calls] == [baseline, baseline], (
|
||||||
|
"a base was dispatched without the project's own schedule"
|
||||||
|
)
|
||||||
|
# The control: without the flag, no base is handed one — so the arm above measures the flag
|
||||||
|
# rather than a default.
|
||||||
|
calls.clear()
|
||||||
|
await run.run_mandate_across_bundles(mandate, [first, second], "local")
|
||||||
|
assert [c.get("cost_baseline") for c in calls] == [None, None]
|
||||||
|
|
@ -209,3 +209,64 @@ def test_the_classification_is_reported_and_is_the_gates_own() -> None:
|
||||||
verdict = validate_proposal(_proposal(code), grounding=Grounding(_OFFERING + (code,)))
|
verdict = validate_proposal(_proposal(code), grounding=Grounding(_OFFERING + (code,)))
|
||||||
refused = isinstance(verdict, Rejection) and "has no identifier form" in verdict.reason
|
refused = isinstance(verdict, Rejection) and "has no identifier form" in verdict.reason
|
||||||
assert refused == (kind == "prose"), (code, kind, verdict)
|
assert refused == (kind == "prose"), (code, kind, verdict)
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
# P21 B3: ``rejection_stage`` — which falsifier wrote a reason. A REPORT, never a gate: nothing
|
||||||
|
# branches on it, so an unrecognised sentence costs a label rather than a verdict. It lives beside
|
||||||
|
# the sentences it keys on, so the classifier and the wordings cannot drift apart (kø-(p)).
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_rejection_stage_names_each_stage_from_its_own_sentence() -> None:
|
||||||
|
from portfolio_optimiser.validator import rejection_stage
|
||||||
|
|
||||||
|
assert (
|
||||||
|
rejection_stage("unknown cost code 'X': not in project p's cost baseline (3 known codes)")
|
||||||
|
== "stage0-baseline"
|
||||||
|
)
|
||||||
|
assert (
|
||||||
|
rejection_stage(
|
||||||
|
"quantity 5 for cost code 'X' is outside the 5.0% tolerance around the baseline "
|
||||||
|
"quantity 9"
|
||||||
|
)
|
||||||
|
== "stage0-baseline"
|
||||||
|
)
|
||||||
|
assert (
|
||||||
|
rejection_stage("ungrounded identifier 'X': it appears nowhere in the input (9 chars)")
|
||||||
|
== "stage0b-grounding"
|
||||||
|
)
|
||||||
|
assert rejection_stage("claimed saving 9 exceeds P90 feasible 4") == "stage4-p90"
|
||||||
|
assert (
|
||||||
|
rejection_stage("claimed saving 9 exceeds the nominal feasible 4 at the items' stated")
|
||||||
|
== "stage4b-nominal"
|
||||||
|
)
|
||||||
|
assert (
|
||||||
|
rejection_stage("claimed 9 exceeds the energy_efficiency method cap 4 (stricter)")
|
||||||
|
== "stage5-method-cap"
|
||||||
|
)
|
||||||
|
# The honest answer for a sentence this module did not write — a label, never a verdict.
|
||||||
|
assert rejection_stage("something else entirely") == "other"
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_rejection_stage_is_keyed_on_the_sentences_the_validator_emits() -> None:
|
||||||
|
"""The control: the markers are not a private paraphrase but the text stage 0 really writes.
|
||||||
|
|
||||||
|
Without it the classifier could key on wording nothing emits and every arm above would still
|
||||||
|
be green — the vacuous-gate class, on a reporter.
|
||||||
|
"""
|
||||||
|
from portfolio_optimiser.ir import AffectedItem, CostBaseline, CostBaselineLine, SavingsProposal
|
||||||
|
from portfolio_optimiser.validator import Rejection, rejection_stage, validate_proposal
|
||||||
|
|
||||||
|
baseline = CostBaseline(
|
||||||
|
project_id="proj", items={"REAL-1": CostBaselineLine(quantity=10.0, unit_cost=100.0)}
|
||||||
|
)
|
||||||
|
proposal = SavingsProposal(
|
||||||
|
project_id="proj",
|
||||||
|
measure="m",
|
||||||
|
affected_items=[AffectedItem(code="FAKE-1", quantity=10.0, unit_cost=100.0)],
|
||||||
|
claimed_saving_nok=100.0,
|
||||||
|
)
|
||||||
|
outcome = validate_proposal(proposal, baseline=baseline)
|
||||||
|
assert isinstance(outcome, Rejection)
|
||||||
|
assert rejection_stage(outcome.reason) == "stage0-baseline"
|
||||||
|
|
|
||||||
|
|
@ -82,9 +82,27 @@ def _minibase(root: Path) -> Path:
|
||||||
return base
|
return base
|
||||||
|
|
||||||
|
|
||||||
def _context(root: Path, *, must_refuse: bool = True) -> Path:
|
def _context(
|
||||||
|
root: Path, *, must_refuse: bool = True, schedule: dict[str, float] | None = None
|
||||||
|
) -> Path:
|
||||||
ctx = root / "ctx"
|
ctx = root / "ctx"
|
||||||
(ctx / "docs").mkdir(parents=True)
|
(ctx / "docs").mkdir(parents=True)
|
||||||
|
if schedule is not None:
|
||||||
|
# P21 B3: the PROJECT's own price schedule, which is the file a run is handed with
|
||||||
|
# ``--cost-baseline`` and the one the judge measures ``priced`` against.
|
||||||
|
(ctx / "cost-baseline.json").write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"project_id": "proj",
|
||||||
|
"items": {
|
||||||
|
code: {"quantity": 1.0, "unit_cost": unit}
|
||||||
|
for code, unit in schedule.items()
|
||||||
|
},
|
||||||
|
},
|
||||||
|
indent=2,
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
(ctx / "bundle.txt").write_text("name: minibase\nbundle_id: minibase\n", encoding="utf-8")
|
(ctx / "bundle.txt").write_text("name: minibase\nbundle_id: minibase\n", encoding="utf-8")
|
||||||
approaches = [
|
approaches = [
|
||||||
{
|
{
|
||||||
|
|
@ -149,6 +167,8 @@ def _write_outbox(
|
||||||
citation_snippet: str = "Body of the good one.",
|
citation_snippet: str = "Body of the good one.",
|
||||||
decision: str = "validated",
|
decision: str = "validated",
|
||||||
tool_calls: list[dict[str, str]] | None = None,
|
tool_calls: list[dict[str, str]] | None = None,
|
||||||
|
anchored: bool = False,
|
||||||
|
reason: str = "no",
|
||||||
) -> None:
|
) -> None:
|
||||||
outbox.mkdir(parents=True, exist_ok=True)
|
outbox.mkdir(parents=True, exist_ok=True)
|
||||||
codes = ["CODE-1"] if codes is None else codes
|
codes = ["CODE-1"] if codes is None else codes
|
||||||
|
|
@ -181,7 +201,7 @@ def _write_outbox(
|
||||||
"role": "proposer",
|
"role": "proposer",
|
||||||
"validator_decision": decision,
|
"validator_decision": decision,
|
||||||
"token_usage": 10,
|
"token_usage": 10,
|
||||||
"cost_baseline_anchored": False,
|
"cost_baseline_anchored": anchored,
|
||||||
"bundle_id_source": None,
|
"bundle_id_source": None,
|
||||||
"external_calls": [],
|
"external_calls": [],
|
||||||
},
|
},
|
||||||
|
|
@ -196,7 +216,7 @@ def _write_outbox(
|
||||||
"run_id": run_id,
|
"run_id": run_id,
|
||||||
"approach_id": approach_id,
|
"approach_id": approach_id,
|
||||||
"outcome_type": "validated" if decision == "validated" else "rejected",
|
"outcome_type": "validated" if decision == "validated" else "rejected",
|
||||||
**({"reason": "no"} if decision != "validated" else {}),
|
**({"reason": reason} if decision != "validated" else {}),
|
||||||
"checker_verdict": None,
|
"checker_verdict": None,
|
||||||
"verdict_id": "vid",
|
"verdict_id": "vid",
|
||||||
},
|
},
|
||||||
|
|
@ -216,7 +236,12 @@ def _opened(path: str) -> list[dict[str, str]]:
|
||||||
|
|
||||||
def _judge(tmp_path: Path, **kw: object) -> stress.ContextSetVerdict:
|
def _judge(tmp_path: Path, **kw: object) -> stress.ContextSetVerdict:
|
||||||
base = _minibase(tmp_path)
|
base = _minibase(tmp_path)
|
||||||
ctx = _context(tmp_path, must_refuse=bool(kw.pop("must_refuse", False)))
|
schedule = kw.pop("schedule", None)
|
||||||
|
ctx = _context(
|
||||||
|
tmp_path,
|
||||||
|
must_refuse=bool(kw.pop("must_refuse", False)),
|
||||||
|
schedule=schedule, # type: ignore[arg-type]
|
||||||
|
)
|
||||||
outbox = tmp_path / "out"
|
outbox = tmp_path / "out"
|
||||||
_write_outbox(outbox, "r1", approach_id="a1", **kw) # type: ignore[arg-type]
|
_write_outbox(outbox, "r1", approach_id="a1", **kw) # type: ignore[arg-type]
|
||||||
return stress.score_context_set(ctx, outbox, "r1", base)
|
return stress.score_context_set(ctx, outbox, "r1", base)
|
||||||
|
|
@ -629,3 +654,100 @@ def test_the_cli_refuses_to_guess_which_base_a_multi_base_outbox_is_for(
|
||||||
|
|
||||||
assert stress.main([*argv, "--bundle", "nowhere"]) == 1
|
assert stress.main([*argv, "--bundle", "nowhere"]) == 1
|
||||||
assert "nowhere" in capsys.readouterr().err
|
assert "nowhere" in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
# P21 B3: the judge reports what the run was ANCHORED on, which codes the project actually PRICES,
|
||||||
|
# and WHICH falsifier caught the falsification arm.
|
||||||
|
#
|
||||||
|
# The measured reason. Rounds 1-4 all ran un-anchored — a vegnormal ships no ``cost-baseline.json``
|
||||||
|
# and the only file loader read one out of the bundle — so stage 0 never spoke and the ``a4`` arm
|
||||||
|
# fell, when it fell, on P7's grounding check. "It was refused" and "the stage that knows what this
|
||||||
|
# project buys refused it" are different facts, and only the second is what anchoring bought.
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_priced_is_true_when_the_project_schedule_carries_the_code(tmp_path: Path) -> None:
|
||||||
|
verdict = _judge(tmp_path, schedule={"CODE-1": 2000.0})
|
||||||
|
assert verdict.approaches[0].priced is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_priced_is_false_for_a_code_the_project_does_not_buy(tmp_path: Path) -> None:
|
||||||
|
"""The discriminator: the SAME schedule, a proposal on a code it does not carry."""
|
||||||
|
verdict = _judge(tmp_path, schedule={"CODE-1": 2000.0}, codes=["CODE-9"])
|
||||||
|
assert verdict.approaches[0].priced is False
|
||||||
|
# ... and the control, so the arm cannot be green by measuring nothing.
|
||||||
|
assert (
|
||||||
|
_judge(tmp_path / "b", schedule={"CODE-9": 2000.0}, codes=["CODE-9"]).approaches[0].priced
|
||||||
|
is True
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_priced_is_false_without_a_project_schedule(tmp_path: Path) -> None:
|
||||||
|
"""Every round before P21: no schedule, so nothing is priced — reported, never guessed."""
|
||||||
|
assert _judge(tmp_path).approaches[0].priced is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_anchored_follows_the_runs_own_stamp(tmp_path: Path) -> None:
|
||||||
|
"""Read off ``provenance.cost_baseline_anchored``, not re-derived from the set's files.
|
||||||
|
|
||||||
|
Both arms, because a field that is constant is not a measurement: an artefact stamped
|
||||||
|
un-anchored must report ``False`` EVEN WHEN the set ships a schedule — the judge says what the
|
||||||
|
run did, and a run that was never given the file is not anchored by the file existing.
|
||||||
|
"""
|
||||||
|
assert _judge(tmp_path, anchored=True, schedule={"CODE-1": 2000.0}).anchored is True
|
||||||
|
assert _judge(tmp_path / "b", anchored=False, schedule={"CODE-1": 2000.0}).anchored is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_the_falsification_arm_reports_which_stage_caught_it(tmp_path: Path) -> None:
|
||||||
|
"""``stage0-baseline`` is the answer anchoring buys; ``stage0b-grounding`` is what round 4 got.
|
||||||
|
|
||||||
|
Driven through the whole judge rather than through ``rejection_stage`` alone, because the
|
||||||
|
seam being gated is that the judge READS the arm's own outcome artefact — a classifier that
|
||||||
|
was never called would leave every arm reporting ``""`` and the arms below still green.
|
||||||
|
"""
|
||||||
|
base = _minibase(tmp_path)
|
||||||
|
ctx = _context(tmp_path, must_refuse=True, schedule={"CODE-1": 2000.0})
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
|
||||||
|
_write_outbox(
|
||||||
|
outbox,
|
||||||
|
"r1",
|
||||||
|
approach_id="a4",
|
||||||
|
codes=["CODE-4"],
|
||||||
|
decision="rejected",
|
||||||
|
reason=("unknown cost code 'CODE-4': not in project proj's cost baseline (1 known codes)"),
|
||||||
|
)
|
||||||
|
verdict = stress.score_context_set(ctx, outbox, "r1", base)
|
||||||
|
assert [(r.approach_id, r.passed, r.stage) for r in verdict.must_refuse] == [
|
||||||
|
("a4", True, "stage0-baseline")
|
||||||
|
]
|
||||||
|
|
||||||
|
other = tmp_path / "b"
|
||||||
|
base2 = _minibase(other)
|
||||||
|
ctx2 = _context(other, must_refuse=True)
|
||||||
|
outbox2 = other / "out"
|
||||||
|
_write_outbox(outbox2, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
|
||||||
|
_write_outbox(
|
||||||
|
outbox2,
|
||||||
|
"r1",
|
||||||
|
approach_id="a4",
|
||||||
|
codes=["CODE-4"],
|
||||||
|
decision="rejected",
|
||||||
|
reason="ungrounded identifier 'CODE-4': it appears nowhere in the input (10 chars)",
|
||||||
|
)
|
||||||
|
assert stress.score_context_set(ctx2, outbox2, "r1", base2).must_refuse[0].stage == (
|
||||||
|
"stage0b-grounding"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_p21_a_validated_falsification_arm_reports_no_stage(tmp_path: Path) -> None:
|
||||||
|
"""``""`` when nothing refused it — an honest absence, never a stage nobody reached."""
|
||||||
|
base = _minibase(tmp_path)
|
||||||
|
ctx = _context(tmp_path, must_refuse=True)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
_write_outbox(outbox, "r1", approach_id="a1", tool_calls=_opened(_GOOD))
|
||||||
|
_write_outbox(outbox, "r1", approach_id="a4", codes=["CODE-4"])
|
||||||
|
verdict = stress.score_context_set(ctx, outbox, "r1", base)
|
||||||
|
assert verdict.must_refuse[0].passed is False
|
||||||
|
assert verdict.must_refuse[0].stage == ""
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue