feat(round-builder): one command turns a run's outbox into a round the gate can read [skip-docs]
python -m portfolio_optimiser.evals.round_builder --outbox <dir> --round <n> --ran-at <ISO> writes <rounds-dir>/<n>/ with the run's artefacts COPIED in, outcome.json derived from that copy, and report.md -- the one artefact in a round a domain expert reads and corrects. Round 0 of the v1 criterion can now be made; it counted 0 of 3 because it could not be, which is a different failure from a round nobody had held. What it derives it derives with the gate's own functions rather than a second copy: verify_run decides whether the run stands up to itself (an artefact contradicting its coverage row, a half-missing family and a stray artefact are all refused AT THE SOURCE, before a byte is written), stage_of gives column (c), row_changed gives the report's "changed since the previous round", parse_time refuses a stamp without a zone, safe_rounds_dir refuses a round directory the repo would commit. The validated total is ledger.to_ore per amount, summed as integers. Two things it never does, and both are the point. It never writes the operator's attestation -- the gate stops at FORM OK without one, and that is correct, because no arrangement of files can witness that a run happened. And it never invents: --ran-at is required because no outbox artefact carries a clock, and feedback_ids stays empty because no run records which feedback item produced which row. The report says "ingen tilbakemelding forklarer dette" on every changed row rather than hiding that model noise and an answered objection look alike. Chosen and why: --ran-at as a required argument rather than the coverage file's mtime, because an mtime is a filesystem attribute one call sets and reading it as evidence made row 2 green on a tree nothing had run in (18.09). The report carries no raw stage identifier -- every stage sentence is "<short name>: <explanation>" so the one-line diff of what changed has words a reader can act on. A citation shows its COUNT, because a run that cited 446 places and one that cited one must not look the same. [skip-docs]: the ledger row is in docs/invarianter.md, which is where this repo's rules live. README is the product's front door and this is an operator tool behind `python -m`, the same class as costsim/hitl/preflight, which README deliberately does not carry; v1-rounds/ is gitignored internal machinery and the gate itself is not in README either. CLAUDE.md was emptied of exactly this kind of row in session 130 and is not the place to put one back. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
0fa612a22f
commit
db38cfc1a4
4 changed files with 605 additions and 27 deletions
|
|
@ -3110,6 +3110,30 @@
|
||||||
tall og bevisregisteret er pinnet mot kilden i testen. Småfunn: rundekatalog inne i repoet uten
|
tall og bevisregisteret er pinnet mot kilden i testen. Småfunn: rundekatalog inne i repoet uten
|
||||||
gitignore og manglende `--stress-root`/`--bundle-root` gir exit 2; en typeannotasjon teller ikke
|
gitignore og manglende `--stress-root`/`--bundle-root` gir exit 2; en typeannotasjon teller ikke
|
||||||
som kallsted. Reviewens 20 mutanter kjørt på nytt: **20 av 20 røde**.
|
som kallsted. Reviewens 20 mutanter kjørt på nytt: **20 av 20 røde**.
|
||||||
|
- **Rundekatalogen BYGGES av en kommando, og den kommandoen skriver aldri attesteringen (19.09):**
|
||||||
|
målt read-only @ `6807946` bandt ingenting en kjørings utboks til `v1-rounds/<n>/` (40 treff på
|
||||||
|
`v1-rounds|rounds_dir`, alle i gaten, dens tester, `.gitignore`, hovedboken og et deck; 0 i
|
||||||
|
`run.py`/`outbox.py`/`scripts`) og ingen steder i `src/` skrev markdown — runde 0 sto på 0 av 3
|
||||||
|
fordi den ikke KUNNE lages, ikke fordi ingen hadde holdt den. `python -m
|
||||||
|
portfolio_optimiser.evals.round_builder --outbox <dir> --round <n> --ran-at <ISO>` skriver
|
||||||
|
`<n>/outbox/` som KOPI, `<n>/outcome.json` UTLEDET av den kopien, og `<n>/report.md`. Binderen
|
||||||
|
dømmer ikke med egne regler: `verify_run` avgjør om kjøringen står inne for seg selv (motsagt
|
||||||
|
artefakt, halv familie, streifende fil — alle avvist VED KILDEN, før en byte skrives),
|
||||||
|
`stage_of` gir kolonne (c), `row_changed` gir «Endret siden forrige runde», `parse_time` nekter
|
||||||
|
et tidsstempel uten sone, `safe_rounds_dir` nekter en rundekatalog repoet ville committet;
|
||||||
|
summen er `ledger.to_ore` per beløp. **To ting gjør den ALDRI:** skriver operatørens
|
||||||
|
`attestering.txt` (gaten stopper på FORM OK uten den — ingen filsamling kan vitne om at en
|
||||||
|
kjøring skjedde), og finner på. `--ran-at` er PÅKREVD fordi ingen utboksartefakt bærer en
|
||||||
|
klokke, og `feedback_ids` står tomt fordi ingen kjøring sporer hvilken tilbakemelding som ga
|
||||||
|
hvilken rad; rapporten skriver «Ingen tilbakemelding forklarer dette» på hver endret rad i
|
||||||
|
stedet for å skjule at modellstøy og et besvart innspill ser like ut. Rapportens stadie-navn er
|
||||||
|
prosa, aldri `stage4-p90`, og et sitat bærer ANTALLET siterte steder. Load-bearing MÅLT
|
||||||
|
(`tests/test_round_builder_loadbearing.py`, 28 armer, hvert tall talt en gang til fra
|
||||||
|
fixturens egen tabell); MÅLT på fire EKTE arkiverte utbokser (`tunnel-hauglia-2027` -04/-06/-07
|
||||||
|
/-08): gaten leser rundene, rad 1 = **FORM OK, IKKE BEVIST**. **Ærlighets-grense, uttalt:** rad
|
||||||
|
2 blir RØD og ikke FORM OK på de samme rundene — radene endret seg (2, 5, 5), men ingen endring
|
||||||
|
er sporet til en feedback-id, fordi sporingen ikke finnes ennå. Den hører i oversettelsen
|
||||||
|
`feedback.json` → kjøringens input, som er en egen ordre.
|
||||||
- **Et forslag uten tilnærmingens EGEN erklæring kan ikke bære `validated` (rad 6, 17.09):** målt
|
- **Et forslag uten tilnærmingens EGEN erklæring kan ikke bære `validated` (rad 6, 17.09):** målt
|
||||||
på stressrunde 6 hadde alle 10 validerte tilnærmingene bare kjørings-erklæringer, som ingen kan
|
på stressrunde 6 hadde alle 10 validerte tilnærmingene bare kjørings-erklæringer, som ingen kan
|
||||||
knytte til én tilnærming — og tre falsifiseringsarmer validerte. `declare_requirement` tar derfor
|
knytte til én tilnærming — og tre falsifiseringsarmer validerte. `declare_requirement` tar derfor
|
||||||
|
|
|
||||||
|
|
@ -5,26 +5,72 @@ and a round number and writes ``<rounds-dir>/<n>/`` in the shape ``v1_gate --hel
|
||||||
``outbox/`` copied from the run, ``outcome.json`` DERIVED from that copy, and ``report.md``, the
|
``outbox/`` copied from the run, ``outcome.json`` DERIVED from that copy, and ``report.md``, the
|
||||||
one artefact in the round a domain expert is meant to read and correct.
|
one artefact in the round a domain expert is meant to read and correct.
|
||||||
|
|
||||||
Two things it never does, and both are the point. It never writes ``attestering.txt``: that file
|
Measured 19.09: nothing bound a run to a round directory (every mention of ``v1-rounds`` in the
|
||||||
is the operator's statement that a round was actually held, and a builder that could produce it
|
repository was in the gate, its tests, ``.gitignore``, the ledger and a deck), and nothing under
|
||||||
would turn rows 1-2 back into something a directory can fake. And it never invents — ``ran_at`` is
|
``src/`` wrote markdown at all. Round 0 therefore counted 0 of 3 because it could not be MADE,
|
||||||
an argument because no outbox artefact carries a clock (they are byte-deterministic by contract),
|
which is a different failure from a round nobody had held.
|
||||||
and ``feedback_ids`` stays empty because no run records which feedback item produced which row.
|
|
||||||
An empty list is the honest reading of a run that tracked nothing.
|
What it verifies, it verifies with the gate's own functions rather than with a second copy:
|
||||||
|
``verify_run`` for whether the run stands up to itself, ``stage_of`` for column (c),
|
||||||
|
``parse_time`` for the run's stamp, ``safe_rounds_dir`` for where a round may be written. A
|
||||||
|
builder that judged its output by its own rules could write rounds the reader refuses.
|
||||||
|
|
||||||
|
Two things it never does, and both are the point. It never writes the operator's attestation
|
||||||
|
file: that file is the statement that a round was actually held, and a builder that could produce
|
||||||
|
it would put the gate's one un-computable step back inside the machine. And it never invents —
|
||||||
|
``ran_at`` is an argument because no outbox artefact carries a clock (they are byte-deterministic
|
||||||
|
by contract), and ``feedback_ids`` stays empty because no run records which feedback item produced
|
||||||
|
which row. An empty list is the honest reading of a run that tracked nothing.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import shutil
|
||||||
|
import sys
|
||||||
from collections.abc import Mapping, Sequence
|
from collections.abc import Mapping, Sequence
|
||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
|
from portfolio_optimiser.evals.v1_gate import (
|
||||||
|
DEFAULT_ROUNDS_DIR,
|
||||||
|
RUN_OUTBOX,
|
||||||
|
parse_time,
|
||||||
|
row_changed,
|
||||||
|
safe_rounds_dir,
|
||||||
|
stage_of,
|
||||||
|
verify_run,
|
||||||
|
)
|
||||||
|
from portfolio_optimiser.ledger import to_ore
|
||||||
|
|
||||||
|
_REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||||
|
_COVERAGE_SUFFIX = "-coverage.json"
|
||||||
|
#: The last round the gate reads (``ROUNDS_CONTRACT``): 0 is the baseline, 1-3 carry feedback.
|
||||||
|
_LAST_ROUND = 3
|
||||||
|
#: A citation is shown to place the claim, not to reproduce the source. Longer than this and the
|
||||||
|
#: report stops being something a person reads in one pass.
|
||||||
|
_SNIPPET_MAX = 220
|
||||||
|
|
||||||
|
|
||||||
class RoundBuildError(Exception):
|
class RoundBuildError(Exception):
|
||||||
"""A round that cannot be built from what is on disk, with the reason a reader can act on."""
|
"""A round that cannot be built from what is on disk, with the reason a reader can act on."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class Run:
|
||||||
|
"""One finished run, as its outbox presents it — already checked against itself."""
|
||||||
|
|
||||||
|
outbox: Path
|
||||||
|
run_id: str
|
||||||
|
rows: tuple[dict[str, Any], ...]
|
||||||
|
stop_reason: str
|
||||||
|
artefacts: tuple[str, ...]
|
||||||
|
ignored: tuple[str, ...]
|
||||||
|
proposals: dict[str, dict[str, Any]]
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
class Built:
|
class Built:
|
||||||
"""What one build produced, in numbers the caller can check against the source directory."""
|
"""What one build produced, in numbers the caller can check against the source directory."""
|
||||||
|
|
@ -39,29 +85,416 @@ class Built:
|
||||||
|
|
||||||
|
|
||||||
#: Every rejection stage ``validator.rejection_stage`` can name, in the words a domain expert
|
#: Every rejection stage ``validator.rejection_stage`` can name, in the words a domain expert
|
||||||
#: reads — plus the two non-rejection labels a coverage row can carry.
|
#: reads — plus the two labels a coverage row carries that are not rejections. A stage added to
|
||||||
STAGE_PROSE: dict[str, str] = {}
|
#: the validator without a sentence here reaches the expert as a bare identifier, so the test
|
||||||
|
#: suite pins this table against the validator's own list rather than against itself.
|
||||||
|
#:
|
||||||
|
#: Every sentence is "<kort navn>: <forklaring>", because both halves are needed in different
|
||||||
|
#: places — the full sentence heads a refusal, the short name fits inside a one-line diff of what
|
||||||
|
#: changed since the previous round. A raw ``stage4-p90`` in either is an internal identifier
|
||||||
|
#: reaching a reader who has no way to look it up.
|
||||||
|
STAGE_PROSE: dict[str, str] = {
|
||||||
|
"stage0-baseline": (
|
||||||
|
"prosjektets egen kostnadsbasis: en kostnadslinje tiltaket bygger på stemmer ikke med "
|
||||||
|
"det prosjektet faktisk har budsjettert"
|
||||||
|
),
|
||||||
|
"stage0b-grounding": (
|
||||||
|
"forankringen i kunnskapsbasen: en kode tiltaket bygger på står ikke i teksten som ble "
|
||||||
|
"levert til kjøringen"
|
||||||
|
),
|
||||||
|
"stage4-p90": (
|
||||||
|
"usikkerhetsberegningen: den påståtte besparelsen er større enn det de gunstigste ti "
|
||||||
|
"prosentene av utfallene gir"
|
||||||
|
),
|
||||||
|
"stage4b-nominal": (
|
||||||
|
"det nominelt mulige: den påståtte besparelsen er større enn tiltaket kan gi selv i "
|
||||||
|
"beste fall"
|
||||||
|
),
|
||||||
|
"stage5-method-cap": (
|
||||||
|
"metodetaket: metoden gir erfaringsmessig ikke en så stor andel av kostnaden tilbake"
|
||||||
|
),
|
||||||
|
"unsupported": (
|
||||||
|
"manglende krav: tallene holdt, men ingen krav i kunnskapsbasen ble erklært å binde "
|
||||||
|
"tiltaket, så retningen står uten hjemmel"
|
||||||
|
),
|
||||||
|
"other": ("en kontroll rapporten ikke kjenner navnet på: begrunnelsen står ordrett under"),
|
||||||
|
"not_evaluated": ("ingenting: kjøringen stoppet før tilnærmingen ble vurdert i det hele tatt"),
|
||||||
|
}
|
||||||
|
|
||||||
ROUND_BUILD_CONTRACT = ""
|
ROUND_BUILD_CONTRACT = """\
|
||||||
|
Skriver EN runde i formen v1-gaten leser (se `v1_gate --help` for hele kontrakten):
|
||||||
|
|
||||||
|
<rounds-dir>/<n>/outbox/ kjøringens egne artefakter, KOPIERT hit (aldri lenket)
|
||||||
|
<rounds-dir>/<n>/outcome.json utledet av coverage + artefaktene, aldri håndskrevet
|
||||||
|
<rounds-dir>/<n>/report.md rapporten fagpersonen leser og retter
|
||||||
|
<rounds-dir>/<n>/feedback.json tilbakemeldingen runden svarer på (n = 1, 2, 3)
|
||||||
|
|
||||||
|
To filer skriver denne kommandoen ALDRI, og begge er med vilje:
|
||||||
|
|
||||||
|
<rounds-dir>/<n>/attestering.txt operatørens bekreftelse på at runden faktisk ble holdt.
|
||||||
|
Uten den stopper gaten på FORM OK, IKKE BEVIST — og det er riktig: ingen filer kan vise
|
||||||
|
at en kjøring skjedde. Formen står i `v1_gate --help`.
|
||||||
|
<rounds-dir>/3/report.kept.md runde 3-rapporten slik fagpersonen beholdt den.
|
||||||
|
|
||||||
|
--ran-at er påkrevd fordi ingen artefakt i utboksen bærer en klokke: filene er
|
||||||
|
byte-deterministiske med vilje. Tidspunktet er operatørens opplysning, ikke en måling.
|
||||||
|
|
||||||
|
feedback_ids står tomt på hver rad. Ingen kjøring sporer i dag hvilken tilbakemelding som
|
||||||
|
førte til hvilken rad, og binderen finner ikke på en sporing som ikke finnes.
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
def read_run(outbox_dir: Path) -> Any:
|
# ---------------------------------------------------------------------------------------------
|
||||||
"""The run an outbox directory holds, or ``RoundBuildError`` saying why it holds none."""
|
# Reading the run
|
||||||
return None
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _read_json(path: Path) -> Any:
|
||||||
|
try:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8"))
|
||||||
|
except FileNotFoundError as exc:
|
||||||
|
raise RoundBuildError(f"{path} mangler") from exc
|
||||||
|
except (ValueError, OSError) as exc:
|
||||||
|
raise RoundBuildError(f"{path} kan ikke leses ({exc!r})") from exc
|
||||||
|
|
||||||
|
|
||||||
|
def read_run(outbox_dir: Path) -> Run:
|
||||||
|
"""The run an outbox directory holds, or ``RoundBuildError`` saying why it holds none.
|
||||||
|
|
||||||
|
The run is checked against ITSELF here, before anything is written, with the gate's own
|
||||||
|
``verify_run``: an artefact that contradicts its coverage row, a family with a missing half,
|
||||||
|
and a stray artefact naming an approach nobody commissioned are all refused at the source. A
|
||||||
|
builder that copied first and let the reader find those would hand the operator a round that
|
||||||
|
is red for something this function already knew.
|
||||||
|
"""
|
||||||
|
if not outbox_dir.is_dir():
|
||||||
|
raise RoundBuildError(f"{outbox_dir} er ingen katalog")
|
||||||
|
covers = sorted(p.name for p in outbox_dir.iterdir() if p.name.endswith(_COVERAGE_SUFFIX))
|
||||||
|
if not covers:
|
||||||
|
raise RoundBuildError(
|
||||||
|
f"{outbox_dir} har ingen <kjøring>-coverage.json. Den fila skrives bare når "
|
||||||
|
"kjøringen hadde BÅDE --mandate og --outbox-dir, så denne utboksen kan ikke si "
|
||||||
|
"hvilke tilnærminger som ble bestilt eller hvorfor hver av dem endte som den gjorde. "
|
||||||
|
"En runde kan ikke bygges av en kjøring uten mandat."
|
||||||
|
)
|
||||||
|
if len(covers) > 1:
|
||||||
|
raise RoundBuildError(
|
||||||
|
f"{outbox_dir} holder {len(covers)} kjøringer ({', '.join(covers)}) — en runde "
|
||||||
|
"presenterer ÉN kjøring, og hvilken kan ikke avgjøres av sorteringsrekkefølge"
|
||||||
|
)
|
||||||
|
run_id = covers[0][: -len(_COVERAGE_SUFFIX)]
|
||||||
|
payload = _read_json(outbox_dir / covers[0])
|
||||||
|
try:
|
||||||
|
rows = [dict(row) for row in payload["rows"]]
|
||||||
|
stamped = str(payload["run_id"])
|
||||||
|
stop_reason = str(payload.get("stop_reason", ""))
|
||||||
|
ids = [str(row["id"]) for row in rows]
|
||||||
|
except (KeyError, TypeError) as exc:
|
||||||
|
raise RoundBuildError(f"{covers[0]} mangler et felt ({exc!r})") from exc
|
||||||
|
if stamped != run_id:
|
||||||
|
raise RoundBuildError(f"{covers[0]} er stemplet kjøring {stamped!r}, ikke {run_id!r}")
|
||||||
|
if not rows:
|
||||||
|
raise RoundBuildError(f"{covers[0]}: kjøringen evaluerte ingen tilnærming")
|
||||||
|
if len(set(ids)) != len(ids):
|
||||||
|
raise RoundBuildError(f"{covers[0]}: samme tilnærming står flere ganger")
|
||||||
|
|
||||||
|
unheld = verify_run(outbox_dir, run_id, rows)
|
||||||
|
if unheld:
|
||||||
|
raise RoundBuildError(f"kjøringen {run_id!r} står ikke inne for seg selv — {unheld}")
|
||||||
|
|
||||||
|
artefacts, ignored = [], []
|
||||||
|
for path in sorted(outbox_dir.iterdir()):
|
||||||
|
if path.is_file() and path.name.startswith(f"{run_id}-") and path.suffix == ".json":
|
||||||
|
artefacts.append(path.name)
|
||||||
|
else:
|
||||||
|
ignored.append(path.name)
|
||||||
|
proposals = {
|
||||||
|
aid: _read_json(outbox_dir / f"{run_id}-{aid}-proposal.json")
|
||||||
|
for aid, row in zip(ids, rows, strict=True)
|
||||||
|
if row["status"] != "not_evaluated"
|
||||||
|
}
|
||||||
|
return Run(
|
||||||
|
outbox=outbox_dir,
|
||||||
|
run_id=run_id,
|
||||||
|
rows=tuple(rows),
|
||||||
|
stop_reason=stop_reason,
|
||||||
|
artefacts=tuple(artefacts),
|
||||||
|
ignored=tuple(ignored),
|
||||||
|
proposals=proposals,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# The outcome file
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
def derive_outcome(
|
def derive_outcome(
|
||||||
run: Any, *, ran_at: str, previous: Mapping[str, Any] | None = None
|
run: Run, *, ran_at: str, previous: Mapping[str, Any] | None = None
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
"""``outcome.json`` computed from the run's own coverage and per-approach artefacts."""
|
"""``outcome.json`` computed from the run's own coverage and per-approach artefacts.
|
||||||
return {}
|
|
||||||
|
Every column is derived: (b) and (d) from the coverage row the artefacts were just checked
|
||||||
|
against, (c) from the gate's ``stage_of``. ``removed`` is the set difference against the
|
||||||
|
PREVIOUS round's outcome — derivation, not declaration. The one thing no file knows is when
|
||||||
|
the run happened, and that is why ``ran_at`` is an argument.
|
||||||
|
"""
|
||||||
|
if parse_time(ran_at) is None:
|
||||||
|
raise RoundBuildError(
|
||||||
|
f"--ran-at {ran_at!r} er ikke et ISO-8601-tidspunkt med tidssone "
|
||||||
|
"(f.eks. 2026-09-18T11:20:00+02:00). Gaten leser dette som kjøringens tidspunkt."
|
||||||
|
)
|
||||||
|
approaches = [
|
||||||
|
{
|
||||||
|
"id": str(row["id"]),
|
||||||
|
"validated": row["status"] == "validated",
|
||||||
|
"stage": stage_of(str(row["status"]), str(row.get("detail", ""))),
|
||||||
|
"validated_nok": row.get("saving_nok") if row["status"] == "validated" else None,
|
||||||
|
"feedback_ids": [],
|
||||||
|
}
|
||||||
|
for row in run.rows
|
||||||
|
]
|
||||||
|
here = {row["id"] for row in approaches}
|
||||||
|
gone = (
|
||||||
|
[
|
||||||
|
{"id": str(row["id"]), "feedback_ids": []}
|
||||||
|
for row in previous.get("approaches", ())
|
||||||
|
if str(row["id"]) not in here
|
||||||
|
]
|
||||||
|
if previous is not None
|
||||||
|
else []
|
||||||
|
)
|
||||||
|
return {"run_id": run.run_id, "ran_at": ran_at, "approaches": approaches, "removed": gone}
|
||||||
|
|
||||||
|
|
||||||
|
def validated_ore(outcome: Mapping[str, Any]) -> int:
|
||||||
|
"""The round's validated saving in øre — quantized PER AMOUNT and summed as integers (kø-(p)).
|
||||||
|
|
||||||
|
``ledger.to_ore``'s rule, not a second one: three 60000.005 NOK lines are 18000003 øre this
|
||||||
|
way and 18000001 if the floats are summed first, and each row is a real amount. Only
|
||||||
|
``validated`` rows count; an ``unsupported`` row's numbers held but its direction has no
|
||||||
|
hjemmel, and a rejected row's figure is a claim the validator refused.
|
||||||
|
"""
|
||||||
|
return sum(
|
||||||
|
to_ore(float(row["validated_nok"]))
|
||||||
|
for row in outcome["approaches"]
|
||||||
|
if row["validated"] and row["validated_nok"] is not None
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# The report — the one artefact a person reads
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _kroner(ore: int) -> str:
|
||||||
|
sign = "-" if ore < 0 else ""
|
||||||
|
whole, rest = divmod(abs(ore), 100)
|
||||||
|
return f"{sign}{whole:,}".replace(",", " ") + f",{rest:02d}"
|
||||||
|
|
||||||
|
|
||||||
|
def _antall(value: float) -> str:
|
||||||
|
text = f"{value:,.2f}".replace(",", " ").replace(".", ",")
|
||||||
|
return text[:-3] if text.endswith(",00") else text
|
||||||
|
|
||||||
|
|
||||||
|
def _stage_short(stage: str) -> str:
|
||||||
|
"""The stage's short name: the clause before the colon of its sentence, never the id."""
|
||||||
|
return STAGE_PROSE[stage].split(":", 1)[0]
|
||||||
|
|
||||||
|
|
||||||
|
def _status_word(row: Mapping[str, Any]) -> str:
|
||||||
|
return {
|
||||||
|
"validated": "validert",
|
||||||
|
"rejected": "avvist",
|
||||||
|
"unsupported": "validert, men uten erklært krav",
|
||||||
|
"not_evaluated": "ikke vurdert",
|
||||||
|
}.get(str(row["status"]), str(row["status"]))
|
||||||
|
|
||||||
|
|
||||||
|
def _lines_about(run: Run, aid: str) -> list[str]:
|
||||||
|
"""The proposal's own words: what it proposed, which cost lines it touched, and one source."""
|
||||||
|
payload = run.proposals.get(aid)
|
||||||
|
if payload is None:
|
||||||
|
return []
|
||||||
|
ir = dict(payload.get("proposal", {}))
|
||||||
|
out: list[str] = []
|
||||||
|
if ir.get("measure"):
|
||||||
|
out += ["", str(ir["measure"])]
|
||||||
|
items = [
|
||||||
|
f"{item['code']} ({_antall(float(item['quantity']))} × "
|
||||||
|
f"{_kroner(to_ore(float(item['unit_cost'])))} kroner)"
|
||||||
|
for item in ir.get("affected_items", ())
|
||||||
|
]
|
||||||
|
if items:
|
||||||
|
out += ["", "Berørte kostnadslinjer: " + "; ".join(items) + "."]
|
||||||
|
citations = list(payload.get("provenance", {}).get("citations", ()))
|
||||||
|
if citations:
|
||||||
|
snippet = " ".join(str(citations[0].get("snippet", "")).split())
|
||||||
|
if len(snippet) > _SNIPPET_MAX:
|
||||||
|
snippet = snippet[:_SNIPPET_MAX].rstrip() + " …"
|
||||||
|
if snippet:
|
||||||
|
# The COUNT is part of the citation: a run that cited 446 places and one that cited
|
||||||
|
# one both show a single quote here, and a reader who cannot tell them apart cannot
|
||||||
|
# tell a grounded proposal from a decorated one.
|
||||||
|
out += [
|
||||||
|
"",
|
||||||
|
f"Kilde (1 av {len(citations)} siterte steder): «{snippet}» — "
|
||||||
|
f"{citations[0].get('file', 'ukjent fil')}",
|
||||||
|
]
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def _change_prose(before: Mapping[str, Any], after: Mapping[str, Any]) -> str:
|
||||||
|
def side(row: Mapping[str, Any]) -> str:
|
||||||
|
if row["validated"]:
|
||||||
|
return f"validert til {_kroner(to_ore(float(row['validated_nok'])))} kroner"
|
||||||
|
stage = str(row.get("stage") or "")
|
||||||
|
return "ikke vurdert" if stage == "not_evaluated" else f"avvist ({_stage_short(stage)})"
|
||||||
|
|
||||||
|
return f"{side(before)} → {side(after)}"
|
||||||
|
|
||||||
|
|
||||||
def build_report(
|
def build_report(
|
||||||
run: Any, n: int, outcome: Mapping[str, Any], *, previous: Mapping[str, Any] | None = None
|
run: Run, n: int, outcome: Mapping[str, Any], *, previous: Mapping[str, Any] | None = None
|
||||||
) -> str:
|
) -> str:
|
||||||
"""``report.md`` — what a domain expert reads and corrects, in Norwegian prose."""
|
"""``report.md`` — what a domain expert reads and corrects, in Norwegian prose.
|
||||||
return ""
|
|
||||||
|
Deterministic for a given (outbox, round, previous): the sections are a fixed sequence and
|
||||||
|
every list follows the coverage's OWN order. Row 4 of the gate counts content lines in order,
|
||||||
|
so a report that reshuffled between builds would read as an expert's edit.
|
||||||
|
"""
|
||||||
|
rows = {str(row["id"]): row for row in outcome["approaches"]}
|
||||||
|
fell = [row for row in run.rows if str(row["status"]) in ("rejected", "unsupported")]
|
||||||
|
held = [row for row in run.rows if str(row["status"]) == "validated"]
|
||||||
|
missed = [row for row in run.rows if str(row["status"]) == "not_evaluated"]
|
||||||
|
|
||||||
|
out: list[str] = [f"# Rapport fra runde {n} — kjøring {run.run_id}", ""]
|
||||||
|
out += [
|
||||||
|
"Denne rapporten er bygget maskinelt fra kjøringens egen utboks; ingen modell har",
|
||||||
|
"skrevet den. Den sier hvilke kostnadstiltak systemet foreslo, hvilke som holdt den",
|
||||||
|
"deterministiske kontrollen, og hvorfor de øvrige falt. Rett den fritt — det du beholder",
|
||||||
|
"og det du skriver om, er selve målingen.",
|
||||||
|
"",
|
||||||
|
]
|
||||||
|
|
||||||
|
out += ["## Hva ble vurdert", ""]
|
||||||
|
stopped = (
|
||||||
|
f" Kjøringen stoppet før den var ferdig ({run.stop_reason})." if run.stop_reason else ""
|
||||||
|
)
|
||||||
|
out += [
|
||||||
|
f"Kommisjonen ba om {len(run.rows)} tilnærminger, og kjøringen rakk "
|
||||||
|
f"{len(run.rows) - len(missed)} av dem.{stopped}",
|
||||||
|
"",
|
||||||
|
]
|
||||||
|
out += [f"- **{row['label']}** — {_status_word(row)}" for row in run.rows]
|
||||||
|
out += [""]
|
||||||
|
|
||||||
|
out += ["## Hva holdt, og hva det er verdt", ""]
|
||||||
|
if held:
|
||||||
|
out += [
|
||||||
|
f"{len(held)} av {len(run.rows)} tilnærminger holdt kontrollen. Samlet validert "
|
||||||
|
f"besparelse: {_kroner(validated_ore(outcome))} kroner.",
|
||||||
|
"",
|
||||||
|
]
|
||||||
|
for row in held:
|
||||||
|
aid = str(row["id"])
|
||||||
|
amount = _kroner(to_ore(float(rows[aid]["validated_nok"])))
|
||||||
|
out += [f"### {row['label']} — {amount} kroner"]
|
||||||
|
out += _lines_about(run, aid)
|
||||||
|
out += [""]
|
||||||
|
else:
|
||||||
|
out += ["Ingen tilnærming holdt kontrollen i denne kjøringen.", ""]
|
||||||
|
|
||||||
|
out += ["## Hva falt, og hvorfor", ""]
|
||||||
|
if fell:
|
||||||
|
out += [f"{len(fell)} tilnærminger ble ikke godtatt.", ""]
|
||||||
|
for row in fell:
|
||||||
|
aid = str(row["id"])
|
||||||
|
out += [f"### {row['label']} — falt på {STAGE_PROSE[str(rows[aid]['stage'])]}"]
|
||||||
|
out += _lines_about(run, aid)
|
||||||
|
out += ["", f"Kontrollens egen begrunnelse, ordrett: «{row.get('detail', '')}»", ""]
|
||||||
|
else:
|
||||||
|
out += ["Ingen tilnærming ble avvist i denne kjøringen.", ""]
|
||||||
|
|
||||||
|
out += ["## Hva kjøringen aldri rakk", ""]
|
||||||
|
if missed:
|
||||||
|
out += [
|
||||||
|
f"{len(missed)} tilnærminger ble aldri vurdert. De er verken godtatt eller avvist.",
|
||||||
|
"",
|
||||||
|
]
|
||||||
|
out += [
|
||||||
|
f"- **{row['label']}** — {row.get('detail') or 'ingen grunn oppgitt'}" for row in missed
|
||||||
|
]
|
||||||
|
out += [""]
|
||||||
|
else:
|
||||||
|
out += ["Kjøringen rakk alle tilnærmingene kommisjonen ba om.", ""]
|
||||||
|
|
||||||
|
if n >= 1:
|
||||||
|
out += ["## Endret siden forrige runde", ""]
|
||||||
|
out += _changed_section(run, n, outcome, previous or {})
|
||||||
|
|
||||||
|
out += ["## Slik leser du tallene", ""]
|
||||||
|
out += [
|
||||||
|
"- «Validert» betyr at forslagets egne tall holdt en deterministisk kontroll mot",
|
||||||
|
" prosjektets kostnadsbasis og en usikkerhetsberegning. Det er ikke en beslutning om å",
|
||||||
|
" gjennomføre tiltaket, og ingen har vurdert om tiltaket er faglig forsvarlig.",
|
||||||
|
"- «Validert, men uten erklært krav» betyr at tallene holdt, men at ingen krav i",
|
||||||
|
" kunnskapsbasen ble erklært å binde tiltaket. Beløpet telles ikke med i summen.",
|
||||||
|
"- Beløpene er kvantisert til hele øre per beløp før de summeres, aldri etter.",
|
||||||
|
"- Om kjøringen faktisk ble gjort, og når, står i ingen fil her. Det bekrefter du selv.",
|
||||||
|
"",
|
||||||
|
]
|
||||||
|
return "\n".join(out).rstrip("\n") + "\n"
|
||||||
|
|
||||||
|
|
||||||
|
def _changed_section(
|
||||||
|
run: Run, n: int, outcome: Mapping[str, Any], previous: Mapping[str, Any]
|
||||||
|
) -> list[str]:
|
||||||
|
"""Every row that moved against round n-1, and for each one whether a feedback id explains it.
|
||||||
|
|
||||||
|
Model noise between two runs looks exactly like an answered objection, so a row no feedback id
|
||||||
|
accounts for has to SAY so. Today that is every row: no run records the tracking.
|
||||||
|
"""
|
||||||
|
before = {str(row["id"]): row for row in previous.get("approaches", ())}
|
||||||
|
label_of = {str(row["id"]): str(row["label"]) for row in run.rows}
|
||||||
|
|
||||||
|
def why(row: Mapping[str, Any]) -> str:
|
||||||
|
ids = list(row.get("feedback_ids", ()))
|
||||||
|
return f"Utløst av {', '.join(ids)}" if ids else "Ingen tilbakemelding forklarer dette"
|
||||||
|
|
||||||
|
out = [f"Sammenlignet med runde {n - 1}:", ""]
|
||||||
|
moved = 0
|
||||||
|
for row in outcome["approaches"]:
|
||||||
|
aid = str(row["id"])
|
||||||
|
was = before.get(aid)
|
||||||
|
name = label_of.get(aid, aid)
|
||||||
|
if was is None:
|
||||||
|
moved += 1
|
||||||
|
out.append(f"- **{name}** — ny i denne runden. {why(row)}.")
|
||||||
|
elif row_changed(was, row):
|
||||||
|
moved += 1
|
||||||
|
out.append(f"- **{name}** — {_change_prose(was, row)}. {why(row)}.")
|
||||||
|
for row in outcome["removed"]:
|
||||||
|
moved += 1
|
||||||
|
out.append(
|
||||||
|
f"- **{row['id']}** — tilnærmingen er borte: den står ikke i denne kjøringens "
|
||||||
|
f"liste i det hele tatt. {why(row)}."
|
||||||
|
)
|
||||||
|
if not moved:
|
||||||
|
out.append("- Ingenting endret seg i (a)-(d) over støygrensen.")
|
||||||
|
out.append("")
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# Building the round
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _dump(payload: Mapping[str, Any]) -> str:
|
||||||
|
"""Byte-deterministic on-disk form, mirroring ``outbox._dump`` with Norwegian text kept as is."""
|
||||||
|
return json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
|
||||||
|
|
||||||
|
|
||||||
def build_round(
|
def build_round(
|
||||||
|
|
@ -72,12 +505,119 @@ def build_round(
|
||||||
ran_at: str,
|
ran_at: str,
|
||||||
feedback: Path | None = None,
|
feedback: Path | None = None,
|
||||||
) -> Built:
|
) -> Built:
|
||||||
"""Write ``<rounds-dir>/<n>/`` from ``outbox_dir``, or refuse without touching the tree."""
|
"""Write ``<rounds-dir>/<n>/`` from ``outbox_dir``, or refuse without touching the tree.
|
||||||
return Built(rounds_dir / str(n), "", (), (), (), (), 0)
|
|
||||||
|
Every check that can fail runs BEFORE the first byte is written, so a refusal never leaves a
|
||||||
|
half-built round for the next command to read as a whole one.
|
||||||
|
"""
|
||||||
|
if not 0 <= n <= _LAST_ROUND:
|
||||||
|
raise RoundBuildError(f"runde {n} finnes ikke i kontrakten — gaten leser 0 til 3")
|
||||||
|
if not safe_rounds_dir(rounds_dir):
|
||||||
|
raise RoundBuildError(
|
||||||
|
f"{rounds_dir} ligger i repoet uten å være ignorert av gitignore — fagpersonens "
|
||||||
|
"tilbakemelding kunne da bli committet til den offentlige remoten"
|
||||||
|
)
|
||||||
|
round_dir = rounds_dir / str(n)
|
||||||
|
if round_dir.exists():
|
||||||
|
raise RoundBuildError(
|
||||||
|
f"{round_dir} finnes allerede. En runde holder fagpersonens egen tilbakemelding og "
|
||||||
|
"rettinger; den overskrives aldri. Flytt eller slett den selv om den skal bygges om."
|
||||||
|
)
|
||||||
|
if n == 0 and feedback is not None:
|
||||||
|
raise RoundBuildError(
|
||||||
|
"runde 0 er grunnkjøringen og svarer på ingenting — det finnes ingen rapport noen "
|
||||||
|
"kan ha kommentert ennå, så den tar ingen feedback.json"
|
||||||
|
)
|
||||||
|
previous: dict[str, Any] | None = None
|
||||||
|
if n >= 1:
|
||||||
|
if feedback is None:
|
||||||
|
raise RoundBuildError(
|
||||||
|
f"runde {n} er tilbakemelding på rapport {n - 1} og så en kjøring: den trenger "
|
||||||
|
"fagpersonens feedback.json (--feedback)"
|
||||||
|
)
|
||||||
|
if not feedback.is_file():
|
||||||
|
raise RoundBuildError(f"--feedback {feedback} finnes ikke")
|
||||||
|
earlier = rounds_dir / str(n - 1) / "outcome.json"
|
||||||
|
if not earlier.is_file():
|
||||||
|
raise RoundBuildError(
|
||||||
|
f"runde {n} måles mot runde {n - 1}, og {earlier} finnes ikke — bygg den runden "
|
||||||
|
"først"
|
||||||
|
)
|
||||||
|
previous = _read_json(earlier)
|
||||||
|
|
||||||
|
run = read_run(outbox_dir)
|
||||||
|
outcome = derive_outcome(run, ran_at=ran_at, previous=previous)
|
||||||
|
report = build_report(run, n, outcome, previous=previous)
|
||||||
|
|
||||||
|
outbox = round_dir / RUN_OUTBOX
|
||||||
|
outbox.mkdir(parents=True)
|
||||||
|
for name in run.artefacts:
|
||||||
|
shutil.copyfile(run.outbox / name, outbox / name)
|
||||||
|
(round_dir / "outcome.json").write_text(_dump(outcome), encoding="utf-8")
|
||||||
|
(round_dir / "report.md").write_text(report, encoding="utf-8")
|
||||||
|
if feedback is not None:
|
||||||
|
shutil.copyfile(feedback, round_dir / "feedback.json")
|
||||||
|
|
||||||
|
return Built(
|
||||||
|
round_dir=round_dir,
|
||||||
|
run_id=run.run_id,
|
||||||
|
copied=run.artefacts,
|
||||||
|
ignored=run.ignored,
|
||||||
|
evaluated=tuple(str(r["id"]) for r in run.rows if str(r["status"]) != "not_evaluated"),
|
||||||
|
not_evaluated=tuple(str(r["id"]) for r in run.rows if str(r["status"]) == "not_evaluated"),
|
||||||
|
validated_ore=validated_ore(outcome),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def main(argv: Sequence[str] | None = None) -> int:
|
def main(argv: Sequence[str] | None = None) -> int:
|
||||||
"""Command-line front door; 0 on a built round, 1 on a refusal, 2 on wrong usage."""
|
"""Command-line front door; 0 on a built round, 1 on a refusal, 2 on wrong usage."""
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
prog="python -m portfolio_optimiser.evals.round_builder",
|
||||||
|
description="Bygg én rundekatalog av en ferdig kjørings utboks, i formen v1-gaten leser. "
|
||||||
|
"Deterministisk og offline: ingen modellkall, intet nett, ingen klokke.",
|
||||||
|
epilog=ROUND_BUILD_CONTRACT,
|
||||||
|
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||||
|
)
|
||||||
|
parser.add_argument("--outbox", required=True, help="kjøringens utboks-katalog")
|
||||||
|
parser.add_argument("--round", type=int, required=True, help="rundenummeret (0-3)")
|
||||||
|
parser.add_argument(
|
||||||
|
"--rounds-dir",
|
||||||
|
default=None,
|
||||||
|
help=f"rundekatalogen (default {DEFAULT_ROUNDS_DIR}/, gitignored)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--ran-at",
|
||||||
|
required=True,
|
||||||
|
help="da kjøringen ble gjort, ISO-8601 MED tidssone. Påkrevd fordi ingen artefakt i "
|
||||||
|
"utboksen bærer en klokke — dette er din opplysning, ikke en måling",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--feedback", default=None, help="fagpersonens feedback.json (påkrevd for runde 1-3)"
|
||||||
|
)
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
|
rounds_dir = Path(args.rounds_dir) if args.rounds_dir else _REPO_ROOT / DEFAULT_ROUNDS_DIR
|
||||||
|
try:
|
||||||
|
built = build_round(
|
||||||
|
Path(args.outbox),
|
||||||
|
rounds_dir,
|
||||||
|
args.round,
|
||||||
|
ran_at=args.ran_at,
|
||||||
|
feedback=Path(args.feedback) if args.feedback else None,
|
||||||
|
)
|
||||||
|
except RoundBuildError as refused:
|
||||||
|
print(f"runden ble ikke bygget: {refused}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
print(f"{built.round_dir} bygget av kjøring {built.run_id}")
|
||||||
|
print(
|
||||||
|
f" {len(built.copied)} artefakt(er) kopiert, {len(built.ignored)} fil(er) utenfor "
|
||||||
|
f"kjøringen ble stående igjen: {', '.join(built.ignored) or 'ingen'}"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f" {len(built.evaluated)} tilnærming(er) vurdert, {len(built.not_evaluated)} ikke; "
|
||||||
|
f"validert besparelse {_kroner(built.validated_ore)} kroner"
|
||||||
|
)
|
||||||
|
print(" attesteringen skriver denne kommandoen aldri — gaten stopper på FORM OK uten den")
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -208,7 +208,9 @@ def _is_ai_text(text: str, ai: tuple[str, str]) -> bool:
|
||||||
return any(line in item for line in lines.splitlines() if line)
|
return any(line in item for line in lines.splitlines() if line)
|
||||||
|
|
||||||
|
|
||||||
def _parse_time(value: Any) -> datetime | None:
|
def parse_time(value: Any) -> datetime | None:
|
||||||
|
"""An ISO-8601 instant WITH a zone, or ``None``. Public so ``round_builder`` refuses a stamp
|
||||||
|
at the moment the operator types it rather than writing a round that reads unparseable."""
|
||||||
try:
|
try:
|
||||||
stamp = datetime.fromisoformat(str(value))
|
stamp = datetime.fromisoformat(str(value))
|
||||||
except ValueError:
|
except ValueError:
|
||||||
|
|
@ -256,7 +258,7 @@ def _declared_run(round_dir: Path) -> tuple[str, datetime | None]:
|
||||||
data, _ = _read_json(round_dir / "outcome.json")
|
data, _ = _read_json(round_dir / "outcome.json")
|
||||||
if not isinstance(data, Mapping):
|
if not isinstance(data, Mapping):
|
||||||
return "", None
|
return "", None
|
||||||
return str(data.get("run_id", "")).strip(), _parse_time(data.get("ran_at"))
|
return str(data.get("run_id", "")).strip(), parse_time(data.get("ran_at"))
|
||||||
|
|
||||||
|
|
||||||
def read_attestation(round_dir: Path, now: datetime | None = None) -> Attestation:
|
def read_attestation(round_dir: Path, now: datetime | None = None) -> Attestation:
|
||||||
|
|
@ -382,7 +384,7 @@ def _read_feedback_file(round_dir: Path, ai: tuple[str, str] | None) -> tuple[Fe
|
||||||
return None, f"{name}: feedback.json uleselig ({exc!r})"
|
return None, f"{name}: feedback.json uleselig ({exc!r})"
|
||||||
if not author:
|
if not author:
|
||||||
return None, f"{name}: feedback.json navngir ingen fagperson"
|
return None, f"{name}: feedback.json navngir ingen fagperson"
|
||||||
given_at = _parse_time(given_raw)
|
given_at = parse_time(given_raw)
|
||||||
if given_at is None:
|
if given_at is None:
|
||||||
return None, f"{name}: given_at er ikke et ISO-tidsstempel med tidssone"
|
return None, f"{name}: given_at er ikke et ISO-tidsstempel med tidssone"
|
||||||
if ai is None:
|
if ai is None:
|
||||||
|
|
@ -500,7 +502,12 @@ class Outcome:
|
||||||
ran_at: datetime
|
ran_at: datetime
|
||||||
|
|
||||||
|
|
||||||
def _stage_of(status: str, detail: str) -> str:
|
def stage_of(status: str, detail: str) -> str:
|
||||||
|
"""Column (c) of an outcome row: which stage of the deterministic gate ended this approach.
|
||||||
|
|
||||||
|
Public because ``round_builder`` DERIVES the outcome file this gate then re-checks against the
|
||||||
|
same coverage. Two copies of this mapping would make a round that the builder wrote fail the
|
||||||
|
reader that asked for it."""
|
||||||
if status == "validated":
|
if status == "validated":
|
||||||
return ""
|
return ""
|
||||||
if status == "not_evaluated":
|
if status == "not_evaluated":
|
||||||
|
|
@ -670,7 +677,7 @@ def read_outcome(round_dir: Path) -> tuple[Outcome | None, str]:
|
||||||
str(r["id"]): set(map(str, r.get("feedback_ids", ()))) for r in data.get("removed", ())
|
str(r["id"]): set(map(str, r.get("feedback_ids", ()))) for r in data.get("removed", ())
|
||||||
}
|
}
|
||||||
run_id = str(data["run_id"]).strip()
|
run_id = str(data["run_id"]).strip()
|
||||||
ran_at = _parse_time(data["ran_at"])
|
ran_at = parse_time(data["ran_at"])
|
||||||
except FileNotFoundError:
|
except FileNotFoundError:
|
||||||
return None, f"{path} mangler"
|
return None, f"{path} mangler"
|
||||||
except (ValueError, KeyError, TypeError) as exc:
|
except (ValueError, KeyError, TypeError) as exc:
|
||||||
|
|
@ -698,7 +705,7 @@ def read_outcome(round_dir: Path) -> tuple[Outcome | None, str]:
|
||||||
truth = {
|
truth = {
|
||||||
str(r["id"]): (
|
str(r["id"]): (
|
||||||
r["status"] == "validated",
|
r["status"] == "validated",
|
||||||
_stage_of(str(r["status"]), str(r.get("detail", ""))),
|
stage_of(str(r["status"]), str(r.get("detail", ""))),
|
||||||
r.get("saving_nok") if r["status"] == "validated" else None,
|
r.get("saving_nok") if r["status"] == "validated" else None,
|
||||||
)
|
)
|
||||||
for r in coverage
|
for r in coverage
|
||||||
|
|
|
||||||
|
|
@ -430,6 +430,9 @@ def test_the_stage_vocabulary_covers_every_stage_the_validator_can_name(tmp_path
|
||||||
assert stages | {"other", "not_evaluated"} == set(rb.STAGE_PROSE)
|
assert stages | {"other", "not_evaluated"} == set(rb.STAGE_PROSE)
|
||||||
assert len(stages) == 6
|
assert len(stages) == 6
|
||||||
assert all(len(text) > 20 for text in rb.STAGE_PROSE.values())
|
assert all(len(text) > 20 for text in rb.STAGE_PROSE.values())
|
||||||
|
# Both halves are load-bearing: the short name before the colon is what a one-line diff of
|
||||||
|
# what changed can carry, and without it the diff falls back on the raw stage id.
|
||||||
|
assert all(": " in text for text in rb.STAGE_PROSE.values())
|
||||||
|
|
||||||
|
|
||||||
def test_the_report_carries_no_json_dump_and_no_unexplained_identifier(tmp_path: Path) -> None:
|
def test_the_report_carries_no_json_dump_and_no_unexplained_identifier(tmp_path: Path) -> None:
|
||||||
|
|
@ -438,6 +441,7 @@ def test_the_report_carries_no_json_dump_and_no_unexplained_identifier(tmp_path:
|
||||||
text = _report(tmp_path)
|
text = _report(tmp_path)
|
||||||
assert "verdict_id" not in text
|
assert "verdict_id" not in text
|
||||||
assert "claimed_saving_nok" not in text
|
assert "claimed_saving_nok" not in text
|
||||||
|
assert not any(stage in text for stage in rb.STAGE_PROSE if stage.startswith("stage"))
|
||||||
assert '{"' not in text and "```json" not in text
|
assert '{"' not in text and "```json" not in text
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -475,7 +479,8 @@ def test_the_report_follows_the_runs_own_order_not_a_sorted_one(tmp_path: Path)
|
||||||
considered = considered.split("\n## ", 1)[0]
|
considered = considered.split("\n## ", 1)[0]
|
||||||
at = [considered.index(label) for _a, label, *_ in reversed_spec]
|
at = [considered.index(label) for _a, label, *_ in reversed_spec]
|
||||||
assert at == sorted(at), "the report does not follow the coverage's own order"
|
assert at == sorted(at), "the report does not follow the coverage's own order"
|
||||||
assert at != sorted(considered.index(label) for _a, label, *_ in _SPEC), "unreversed"
|
forward = [considered.index(label) for _a, label, *_ in _SPEC]
|
||||||
|
assert forward == sorted(forward, reverse=True), "the report sorted instead of following it"
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
@ -547,6 +552,8 @@ def test_round_one_reports_what_changed_and_admits_when_nothing_explains_it(
|
||||||
section = text.split("## Endret siden forrige runde", 1)[1].split("\n## ", 1)[0]
|
section = text.split("## Endret siden forrige runde", 1)[1].split("\n## ", 1)[0]
|
||||||
assert flipped[1] in section
|
assert flipped[1] in section
|
||||||
assert "ingen tilbakemelding forklarer dette" in section.lower()
|
assert "ingen tilbakemelding forklarer dette" in section.lower()
|
||||||
|
assert "stage4-p90" not in section, "a raw stage id reached the expert"
|
||||||
|
assert rb.STAGE_PROSE["stage4-p90"].split(":")[0] in section
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue