The gap, found by the mutation sweep of 2026-07-25: verdict_dir=args.verdict_dir
→ None in run.py's execute_portfolio call left the suite 603/603 GREEN. The
flag was wired but not guarded — the inner merge (test_portfolio_learning_
loadbearing.py), the argparse refusal (--verdict-dir without --portfolio) and
the README↔--help sync all stay green under that mutation, so none of them
covered the forwarding itself.
One load-bearing test, no production code. It authors an expert verdict into a
tmp portfolio inbox — keyed on the bundle's own codes + measure type so it ranks
into the fold, with a distinct saving so its id cannot collide with the bundle's
seed — drives main(["--portfolio", …, "--verdict-dir", X]) with the scripted
client, and asserts the verdict's id AND a marker token (present nowhere in the
bundle) reach the proposer prompt.
Detach proof (mutation restored from a COPY, never git checkout): the wire
mutated to None → 1 failed, 603 passed, and the failure is this test alone.
Restored → 604 passed, ruff format left 71 files unchanged, ruff check + mypy
--strict clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQu2xxwedckjU56byu1aUG
The last ungated build session: the operator now drives the whole build from the
command line, and the documents claim exactly what the code does (§1).
run.py becomes the collecting entrance. Exactly one of --bundle (one project) or
--portfolio (N projects from a schema-validated reference config, with
--verdict-dir as the portfolio-level expert inbox) is required; both and neither
are refused. --goals loads a goal contract and checks it against --ledger's
realized sum BEFORE the first model call: the §8 caps bound spend, the goal bounds
achievement, so a hard target the book already meets stops the run at exit 4
without constructing a client. A soft target reached is a flag and the run
continues; an absent ledger is an empty book, so the goal is still evaluated,
never skipped. The one declared goal also drives --value-report's goal progress —
one contract, never two figures that can disagree.
The portfolio path persists nothing (K3 returns typed results; the outbox names
pairs by run_id, which a portfolio pass has none of). Rather than accept
--out/--outbox/--run-id/--value-report/--inbox/--live-dry-run there and silently
ignore them, the entrance refuses them and says why. run_portfolio is imported
lazily — portfolio.py imports this module, so a module-level import is circular.
Three seams, each detach-proven RED:
- unwire the goal check → the run proceeds and spends → red
- unwire the portfolio branch → the configured projects never run → red
- document a flag no CLI offers → the README honesty grep goes red
That last one is the doc-sync made load-bearing: the test reads README.md,
collects every --flag it documents (excluding third-party dev-tooling lines) and
asserts each exists in the --help of a CLI the README names. The drift it exists
to close was real — README claimed 562 tests, CHANGELOG claimed 265, actual 597.
Docs synced to the code: README gains an operator-CLI section and honest goal/
portfolio descriptions, CHANGELOG is rewritten to what actually shipped, and
docs/oppskrift-kunnskapsbase.md delivers D-H point 1 — the documented team
process for building a knowledge base, with the honest 1–2 week expectation and
every factory-dependent step (verdict translation, demo path) marked NOT BUILT.
597 passed · ruff clean · mypy strict clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQu2xxwedckjU56byu1aUG