The four arms that were red on assert, in the order the reader meets them.
A DANGLING SYMLINK as the round directory now refuses instead of tracebacking.
exists() FOLLOWS a link, so a dangling one answers False and slipped straight past
"a round is never overwritten"; the build then died on the filesystem's own
FileExistsError. is_symlink() is checked FIRST, the message says what a link would cost
(the round's content would sit somewhere the gate does not measure), and the link is
left exactly as it was found — nothing is written, exit code 1 like every other refusal.
ONE CITATION LIST SHARED BY EVERY PROPOSAL is now stated once. The cause was measured
before anything was written, because "the builder reads the wrong field" and "the outbox
says the same thing five times" want opposite fixes: in all four archived runs every
proposal carries a byte-identical 270-citation list — the run's whole retrieved context,
stamped once per proposal. No report can make that quote say something about the
individual measure. So when every proposal carries the same list, the report says so
once, says what the list actually is ("hva kjøringen leste, ikke hva det enkelte tiltaket
bygger på"), and drops the five copies. When the lists differ, nothing changes: the quote
and its COUNT stay under each proposal, which is where they mean something.
THE SAME COST LINE ON BOTH SIDES OF THE VERDICT is named where it happens. The 19.09
report refused TUN-LYS-01 under one label and validated the same line under another and
said nothing, so a reader met two figures for one budget line with no way to see they
collided. Both sides now carry the sentence, in the run's own row order.
A REMOVED APPROACH is shown by the label the expert saw, with the id in parentheses. It
is the one row whose human name is not in this run's coverage, so the label is read from
the coverage inside the PREVIOUS round's own outbox — derived from what the round already
carries, not a new column in outcome.json. An unreadable coverage falls back on the bare
id; a missing label is not a reason to refuse a round.
39 of 39 arms green. Nothing here touches what the gate reads: outcome.json keeps its
four columns and the builder still never writes the attestation.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
python -m portfolio_optimiser.evals.round_builder --outbox <dir> --round <n> --ran-at <ISO>
writes <rounds-dir>/<n>/ with the run's artefacts COPIED in, outcome.json derived from that
copy, and report.md -- the one artefact in a round a domain expert reads and corrects. Round 0
of the v1 criterion can now be made; it counted 0 of 3 because it could not be, which is a
different failure from a round nobody had held.
What it derives it derives with the gate's own functions rather than a second copy: verify_run
decides whether the run stands up to itself (an artefact contradicting its coverage row, a
half-missing family and a stray artefact are all refused AT THE SOURCE, before a byte is
written), stage_of gives column (c), row_changed gives the report's "changed since the previous
round", parse_time refuses a stamp without a zone, safe_rounds_dir refuses a round directory the
repo would commit. The validated total is ledger.to_ore per amount, summed as integers.
Two things it never does, and both are the point. It never writes the operator's attestation --
the gate stops at FORM OK without one, and that is correct, because no arrangement of files can
witness that a run happened. And it never invents: --ran-at is required because no outbox
artefact carries a clock, and feedback_ids stays empty because no run records which feedback
item produced which row. The report says "ingen tilbakemelding forklarer dette" on every changed
row rather than hiding that model noise and an answered objection look alike.
Chosen and why: --ran-at as a required argument rather than the coverage file's mtime, because
an mtime is a filesystem attribute one call sets and reading it as evidence made row 2 green on
a tree nothing had run in (18.09). The report carries no raw stage identifier -- every stage
sentence is "<short name>: <explanation>" so the one-line diff of what changed has words a
reader can act on. A citation shows its COUNT, because a run that cited 446 places and one that
cited one must not look the same.
[skip-docs]: the ledger row is in docs/invarianter.md, which is where this repo's rules live.
README is the product's front door and this is an operator tool behind `python -m`, the same
class as costsim/hitl/preflight, which README deliberately does not carry; v1-rounds/ is
gitignored internal machinery and the gate itself is not in README either. CLAUDE.md was emptied
of exactly this kind of row in session 130 and is not the place to put one back.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Round 0 of the v1 criterion cannot be made today: nothing binds a finished run's outbox to
<rounds-dir>/<n>/, and nothing in src writes markdown a domain expert could read. These tests
state what a builder has to do before one exists, and every number they assert is counted a
second time from the fixture's own table rather than read back from the builder.
Red on assertions, not on import: round_builder.py lands as a contract -- dataclass, signatures,
neutral returns -- so each test fails in its own body.
Two gate helpers become public rather than being copied: row_changed (the report's "changed since
the previous round" section must not disagree with the gate about what changed) and
safe_rounds_dir (the builder CREATES the directory the gate only reads, and the writer is where a
leak of the expert's feedback has to be stopped).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>