The Step-7 trace line said "lang fil-løkke" while the verdict arrived as a
function argument (`verdict_input`) — the short, in-run capture. The long loop
was tested but never exercised by the thing on stage.
An expert now drops a real verdict FILE (`write_verdict`) into an inbox between
the runs, and Run B is given `verdict_dir=`, so `run_project` merges it into the
store before the Step-1 fold.
Not done as the plan point was worded, and the difference is load-bearing:
routing the PERSONA verdict through the inbox would have put ONE marker on two
paths — Step 7 (inbox) and Step 8 (promotion) both end in Run B's prompt, so
either could carry it alone and `test_simulation_loadbearing.py`'s promotion
assertion would have stayed green with promotion detached. A second verdict with
its own marker keeps both seams independently red-able; `simulate_learning_loop`
raises when the two markers are equal. The inbox sits beside the bundle copy,
never inside it, and the id is an explicit sentinel (a minted id would collide
with the promoted verdict's, and `VerdictStore.add` is first-write-wins).
766 -> 769 passed (773 collected). Criterion 6 re-measured: stdout byte-identical
across two runs; stderr unchanged at 6 lines. Mutations measured against the full
suite, four red + a green control: detach `verdict_dir=` · point Run B at an empty
folder while the file is still written · marker set to `realization_rate: 0.82`
(measured present in the verdict seed) · marker set to `energy performance gap`
(measured present in a navigated concept file) · benign rename of the inbox dir.
Honesty limit found while measuring: the last two mutations fell on the causality
assertion, not the Run A control — generation prompts carry the debate output, not
the bundle context. The pair holds, but each assert defends a different property.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVYDeJ9evZicgU5r3roZVW
I1: Funn 1 was measured one directory wide; the repo ships a working S4.0
baseline fixture and run.py:516 reads it. The anchored dry-run + the 10%%
deviation test move from Tuesday to the weekend (P4 pt 0); Tuesday becomes a
re-measurement with an explicit abort path (I4: pre-pull hash, reset rule,
18:00 NO-GO). I2: stderr damping decided YES, in the weekend BEFORE pinning —
measured today stderr is six lines, one deliberately non-deterministic. I3:
[project.scripts] moves off freeze day to before the fresh-clone measurement.
I5: a demo runbook post (P4.5) at the freeze. I6: every §4 claim re-measured
today on HEAD bb3df79; the 08-09 datings were commits from 2026-08-06.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UDHSsyMuBASJcRapddciHL
Operator directive: full week available, weekend included, new quota,
high priority. The calendar now starts Friday with P1, pulls the whole
P4 advance (fresh-clone criterion, golden transcript, stderr muting,
both honesty sentences) into the weekend against the micro reserve, and
makes Monday dress rehearsal #0 — the NO-GO outcome is fully verified
BEFORE Tuesday's GO gate, leaving Tuesday/Wednesday thin: pull+measure,
re-measure, freeze, release cut, tag.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019xhQpH4oQBaf8dxCkXuB8Z
Operator decision 2026-08-09. Track 1 (complete v1, incl. other repos):
S1.a = P1 step-7 inbox, S1.b = P2 content gate, S1.c = release cut
(1.0.0 synced in four places, CHANGELOG, [project.scripts] moved in from
P9, tag only AFTER a green dress rehearsal). Other-repo accounting is
measured: commons already ordered with the Tuesday deadline and a
reserve, okf/guard/po-claude need nothing — no new coord message. Each
post carries a named degradation so v1 stays honestly complete at every
level. Track 2 (convincing demo): P3 + P4 + rehearsal + one spoken
mandate sentence, optional stderr-noise muting before the freeze. The
O4-vs-tag conflict is flagged for the operator, not decided.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019xhQpH4oQBaf8dxCkXuB8Z
Fable-review 2026-08-09 made durable: P1-P10 in plain language with the
commands behind every number (evidence table §4). Pre-demo: step-7 inbox
wired into the walkthrough (P1), Spor B sharpening (P2), the stage-0
first-contact check on Tuesday's GO (P3), fresh-clone/stderr/golden
criteria plus two honesty sentences on Wednesday (P4). Post-demo: CLI
portfolio cap (P6), one consolidated commons amendment (P7), method
skill [Voyage] (P8), and an explicit NULL for orchestration swaps (P10).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019xhQpH4oQBaf8dxCkXuB8Z
A prompt that lives only in a conversation dies at /clear, so it goes in the repo.
Three tracks, in the order the operator weighted them: the MAF feature set, the demo, and -- as
the actual deliverable rather than an appendix -- a ranked list of what the week's quota should
buy. Each item carries a mechanism, a hard [FØR TORSDAG]/[ETTER DEMOEN] tag, a cost in SESSIONS
rather than hours, and what would go red if the item were done. An item nothing can falsify is an
opinion, not a finding.
The measured starting points are embedded so the session does not re-derive them wrongly: two
debate agents rather than three, one orchestration in use out of the installed surface, a
hand-rolled portfolio fan-out, and a capability map organised by NEED that therefore never
compares TOPOLOGIES. The map is not stale on version -- 1.9.0/1.0.0 is what is installed -- which
matters, because "the map is old" would be the easy wrong conclusion.
The prompt carries its own discipline because Fable runs without an advisor: every figure must be
produced by a command shown beside it, and premises in STATE and in plan documents are named as
premises. This repo has measured at least three of them wrong, most recently today.
It is also forbidden from smuggling feature work in front of the demo. The demo is a hard date;
the feature set is not. Arguing otherwise is allowed -- but only out loud, with the consequence
spelled out.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoHJCKBTjFKcjsfEQyGbzh
The demo shows "download and run". Implying you can point this at your own sources and build a
knowledge base claims three things the code does not carry -- and A5 (the code may not claim more
than it does) binds the presenter too, not just the source.
Measured first, and one measurement changed the plan: the guard is NOT v0.2 alpha. That figure came
from our own 2026-07-16 inclusion plan, which is a premise rather than a fact. It is v0.3.4, seven
published tags, `dependencies = []` -- stdlib only. Our okf pin (v0.3.2) declares no dependencies
either, so the guard is not coupled to it, and the 0.3.5-vs-0.4.0 release argument concerns the
release AFTER v0.3.4. Adoption moved from risky to tractable on that one reading.
The three claims, made precise: the demo bundle was hand-curated (honesty), the ingest path writes
unscanned (buildable), and the generic bundle factory does not exist (deferred at O1, not buildable
in four days). Two close with code, one with a sentence.
The two tracks are separated on a measured fact: `simulation.py` does not import `ingest`, so Door A
work cannot disturb what Wednesday freezes. Criterion 5 is the one that proves it -- the walkthrough
must stay byte-identical.
Four decisions are named as decisions rather than settled silently: which policy preset, fail-closed
versus flag-and-write, where the guard's report lands in provenance, and keeping `--strict`
meaningful across a seam that ships no py.typed. The honesty paragraph is written in BOTH variants
up front, so Wednesday is an observation and not a judgement call on stage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoHJCKBTjFKcjsfEQyGbzh
The scripted proposer answered one hard-coded pair of proposals. A second project meant a second
hand-written selector, written under demo-week time pressure -- the risk the week plan names
explicitly (§4, risk 2). It is now a registry: `ScriptedCandidate` entries selected by
`scripted_proposer`, plus `project_id` as an argument to `simulate_learning_loop`.
The open decision was WHAT identifies the candidate in the prompt blob; the plan flagged it as
unverified, so it was measured. Two prompt shapes reach the selector: the debate prompt carries the
whole bundle context, the generation prompt carries `Project: {id} - {name}` plus -- as its context
-- the debate output, which is the selector's own earlier reply. So the cost code and the measure
name are present in the generation prompt only because the script put them there; keying on them
would key the script on its own output. The project id is the one identifier both shapes carry and
the framework stamps.
Validation, never repair: no match, or more than one, raises `ScriptedCandidateError`. A default
reply would answer an unregistered project with another project's numbers, which on screen is
indistinguishable from a correct run; an ambiguous blob is a data problem that must surface at the
rehearsal rather than be decided by registry order.
Load-bearing MEASURED against the whole suite, five mutations all red plus a green control: detach
the project keying - one global flip key - fall back on an unknown project - first-match on an
ambiguous prompt - detach the `project_id` argument. The flip-key test was rewritten mid-measurement
because its first form asserted on the FIRST registry entry, where "the matched candidate's key" and
"candidates[0]'s key" coincide -- it could not separate the two implementations, and proved nothing.
766 passed / 4 skipped. Simulation still exits 0, still prints eight labelled steps, still
byte-identical across two runs.
[skip-docs] README is deliberately untouched: O4 defers the README rewrite to 14-15 August, after
the demo has produced the evidence for the level-2 claim. CLAUDE.md carries the invariant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoHJCKBTjFKcjsfEQyGbzh
The simulation proved the loop but printed only four of its eight steps, so a
listener could not follow what they were looking at without narration. This is
presentation only: every value printed is read off the RunResult the run already
returned -- nothing is recomputed against the bundle, nothing is inferred, and
the run path is untouched. The week plan's assumption ("seven of eight steps are
pure presentation") therefore held; no new logic was needed.
Run A walks steps 1-7, the promotion between the runs IS step 8, and Run B is
not re-numbered -- it shows only what changed, which is the marker reaching the
hypothesis prompt. Two honesty limits are visible in what is printed rather than
papered over: `retrieved` is the post-hoc proposal-keyed retrieval, not the
Step-1 fold (the marker line is what evidences the fold reaching the prompt),
and the run carries the checker's DECISION, not its prose -- the decision is
what gates, so it is what is shown.
The working-copy path moves to stderr: mkdtemp is the one non-deterministic
value in the output, and stdout must be byte-identical across runs for the dress
rehearsal's diff check. Status tokens stay VALIDATED/REJECTED in English on
purpose -- the same vocabulary as provenance.validator_decision, which the
Step-6 line prints verbatim.
Verified: `... | grep -cE "^ *Steg [1-8]"` -> 8; two runs byte-identical on
stdout; both a REJECTED and a VALIDATED line for the same candidate; suite
759 passed / 4 skipped; ruff + mypy clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoHJCKBTjFKcjsfEQyGbzh
generate_via_llm consumed each validator Rejection internally (`last`), fed it into the
next attempt's prompt, and dropped it. So Step 5 was real but unobservable: a caller could
see THAT a proposal validated, never that it validated on attempt 2 after the deterministic
validator falsified attempt 1. It was the one step of the eight with no output to show.
The seam is a typed return value -- GenerationResult(outcome, refinements) -- rather than an
out-parameter or a callback: a returned value cannot be silently lost by a caller that forgets
to pass a collector, and mypy forces every call site to acknowledge it.
refinements carries ONLY rejections that were actually fed back. When the attempt budget runs
out the final rejection IS outcome; counting it here would be double-counting, and the bounded
control test goes red on the collect-everything implementation that gets this wrong.
The loop's bound is untouched: max_attempts and meter.tick_round stand, and `last` still drives
the prompt alone, so prompt growth is unchanged. run.py accumulates across _evaluate calls, so
_evaluate_mandate is untouched; RunResult.refinements defaults (the coverage precedent) and is
concatenated across approaches rather than keyed per approach -- stated as an honesty limit.
The simulation now shows it: the scripted proposer overclaims 250000, which the validator
falsifies against P90 = 90000, and the corrected 30000 validates. Only the overclaim is
scripted -- the rejection is computed. scripted_factory takes a per-role reply selector so this
needs no second scripted client body.
README records the two accuracy changes only (Step 5 is now inspectable; the simulation trace
shows the correction). The level-2 publishing claim stays deferred until after the demo (O4).
Load-bearing MEASURED against the full suite with a control, four mutations all red:
detach the returned history (4 tests) - collect-everything (control only) - detach the run
wiring (2 tests) - revert the simulation's proposer to a constant (the demo-protection test).
Control: 759 passed / 4 skipped; ruff, format and mypy clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CcWFcREUi6YPjEpN3ACDP
Seven of the eight steps already have their data in RunResult and need a print;
one does not exist at all. Putting that distinction in a table is the point of
this plan -- it turns "show all eight steps" from an unbounded week into one
build on Friday and presentation work over the weekend.
The go/no-go on Tuesday is deliberate. The content is being built in another
repo on a deadline nobody here controls, so the week is designed to survive it
not arriving rather than to hope it does. The content-keyed reply selector
lands Monday, before the content, for the same reason: a new project should
then be a data entry rather than a hand-written script under time pressure.
Honesty framing is section 1 rather than a footnote, because the demo's own
subject is a system that refuses to claim more than it proves.
The 20 claims went un-corrected, so they stand as confirmed. What the operator
actually decided were the four choices the QA exposed: hand-built example with
the factory path explicitly deferred, step 5 built and shown live, commons
ordered with a fallback, README after the demo rather than before.
Two measurements are recorded because they bound the order, not because they
are interesting: bundle_context renders every navigated file's full body, and
summary-first reading is not built -- so the full 15-30 measure library would
put 40-90k characters into every hypothesis prompt. The order is size-capped
for that reason and says so.
Also recorded: simulate_learning_loop already takes the bundle directory as a
parameter, so new content plugs into an existing seam. The cost is the scripted
replies, which are written against the LED case.
The demo-week brief was written by a session that read its way to the
intention through documents other sessions had written. Two of its frames
were overturned by the primary sources inside one conversation, so the
operator stopped planning and commissioned this: read the primary sources
directly, state the understanding back as numbered claims, and capture the
corrections where they survive.
Six gaps in the picture the brief rests on, all measured rather than argued:
- The intention has a SECOND axis that STATE's list of five primary sources
never named. review-2026-07 (F1-F14) and sesjonsplan-fase2-6 (S2.0-S5.4,
D-A-D-I, M1-M3) are where most of the repo's 31 modules come from: 20
S-numbers, 18 with hits in src/+tests/. A plan written from the five named
sources alone would describe a repo with eight steps and miss two thirds
of what is there.
- D-H's DECIDED demo path ("clone -> unzip -> factory builds -> loop runs")
is factory-dependent, and the factory (D-G/T0, `okf-toolkit`) does not
exist -- measured, not assumed. The brief's "anyone who downloads the repo
can run exactly the same" IS that path.
- The realistic example's content model is already decided (D-F): knowledge
types with required source citation, strict separation from the verdicts.
The commission to commons must reference it, not invent one.
- The demo is the programme's level-2 publishing proof (D-I), with an
honesty ceiling agreed in advance and a README update as its consequence.
- The shared spec covers the loop + ingest and NONE of the surplus: mandate,
notify, ledger, value report, cost simulation, dimension, portfolio
budget, concurrency, preflight all measure 0 mentions. The comparison is
therefore of the SPEC'd core, not of this repo.
- Step 5 is not a presentation-layer concern: `generate_via_llm` consumes
the intermediate rejection internally, and today's demo validates on the
first attempt, so the refinement never triggers. Steps 2 and 6 ARE
printable from data RunResult already carries.
Two inventory numbers spot-checked independently (759 collected; the offline
simulation re-run, output identical). Nothing here is sourced from STATE.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TvjgY5NBg16D7kgQf14s6B
STATE said the release text was already written and only needed dating. Measured:
CHANGELOG.md was last touched at 9e149c6, 53 commits back, so [Unreleased]
described the repo as of S3.1 — and claimed 512 passing tests where the suite
measures 755.
For a FIRST tag there is no predecessor, so the section is not "what changed" but
"what this version is". A description that stops at S3.1 does not under-report a
delta; it misrepresents the artefact being tagged, on a public mirror.
Audited the existing bullets for claims that had gone FALSE rather than merely
stale — the class a test-count catch is only a sample of. Four checked, all still
true at HEAD: numpy confined to semretrieval.py; value_report deliberately not
wired into costsim; the simulation module and knowledge-base recipe present; the
two run modes still two.
Added, grouped by seam rather than by commit: S4.0 cost-baseline anchoring, S2.7,
S3.2 per-candidate verdict keying, S3.3 wave concurrency, S3.4 global token cap,
S2.2/S2.4 ingest transports, --scripted-replies, --mandate, --mcp-config, the
per-approach outbox, and external-call provenance. New Fixed section for the
seven correctness fixes, with a lead-in saying plainly that none of them ever
shipped. Notes gained the deployer-owns-DPIA scope boundary and the pull-only
subtree rule.
The test-suite bullet now records TWO cases where "every seam is load-bearing"
did not hold until measured, not one: the S3.1 CLI-level gap, and four of five
BudgetExceeded raise sites free to report any `observed`.
No code change. Suite measured green at 755 passed / 4 skipped after the edit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HVekyKcvT4ah6jnCT8R8fe
The egress declaration (Trekk B3) says what a run MAY contact. It cannot say what
it DID: after the run, nothing distinguished "the agents queried the price
register" from "the agents ignored it", and a proposal resting on an external
service should be traceable to it.
ToolCallRecorder(FunctionMiddleware) mirrors BudgetMiddleware(ChatMiddleware) one
layer down — that one observes the debate's chat calls, this one its tool calls.
It observes only: call_next is always awaited, so a trace can never alter the run
it traces. The record lands on ProvenanceStamp.external_calls, read AFTER the
debate so it is a record rather than an intention.
MEASURED, not assumed, before any of it was written: FunctionMiddleware fires for
a tool served over a REAL MCP stdio subprocess, and context.function.name carries
the BARE tool name with no server prefix. That measurement decided the design —
MAF cannot tell us which server a tool came from, so attribution comes from our own
config, and a name allowed by two servers is recorded UNATTRIBUTED (server="")
rather than credited to the first match. Naming a service that may never have been
contacted is the one place a guess must not go.
Only CONFIGURED tools are recorded. The middleware fires for every function the
agents invoke, including the in-process retrieve_cost_docs on the road path;
logging those would turn the record into a false egress claim. An empty list is a
positive statement — nothing outside this process was contacted — which is why it
is always serialized rather than omitted.
Honesty limit, written on ExternalCall itself: this is the call and its source. It
is NOT evidence that the service's answer reached the proposal, nor a verified
rendering of that answer.
One finding, and it is the reason for measuring rather than trusting green: the
road-path negative test was VACUOUS. Its scripted tool call named an argument the
tool does not declare (code vs query), MAF rejected the call before invocation, and
the test asserted an empty record against a run where no tool ran at all — green
under the exact mutation it existed to catch. It now spies on the recorder and
asserts the invocation genuinely reached it before asserting it was not recorded.
This is last session's lesson again: a scenario that cannot distinguish two
implementations proves nothing.
The tool-call double is registered in test_scripted_client_consolidation.py's
_DELEGATING_OVERRIDES — it cannot live in the reply_selector seam, which returns a
reply STRING, and a response that is not text is its whole subject.
Load-bearing MEASURED (tests/test_b4_mcp_call_trace_loadbearing.py) against the
whole 755-test suite, four mutations all red: detach the recorder from the debate
middleware · record every function invocation · attribute an ambiguous name to the
first server · stop reading the recorder into provenance. Control: a run with no
configured servers records nothing, so the empty record is a real answer and not
the only one the seam can produce.
Ran it, not just tested it: the real recorder against a real MCP server subprocess
returns ExternalCall(server='prisregister', tool='lookup_unit_price'), and a
scripted CLI run's outbox artefact carries the empty list.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VtRd8y1PDPGwkrRXFhubqr
A run commissioned to evaluate three approaches wrote ONE proposal artefact, so
only the approach it selected could ever receive a verdict. The other two were
evaluated, reported in the settlement, and then taught the learning loop nothing.
The defect class is a key collapse, and it had two halves — fixing either alone
leaves it intact:
* the WRITER wrote one pair per run, so the non-selected approaches never existed
on disk;
* the READER (hitl._read_outbox_proposals) joins proposal to outcome on the
run_id FIELD read from file CONTENT, never the filename. Three files sharing
one run_id collapse onto one dict key, last write wins — so widening only the
filename would have produced three artefacts and still one pending row. This is
the S3.2 collision class: two rows under one key silently become one.
Artefacts are now keyed {run_id}-{approach_id}-*.json AND carry approach_id in the
payload; the join key is (run_id, approach_id). Two properties make them genuinely
judgeable rather than merely present:
* verdict_id is minted per approach (verdicts.verdict_key, the S3.2 content hash)
— reusing the run's single id would let one delivered verdict clear all three
from the queue;
* provenance.validator_decision follows ITS OWN approach — the run's stamp would
report a rejected candidate as validated, and nothing downstream could correct it.
verdicts.verdict_key is public so a run can stamp the key a verdict WILL arrive
under without capturing a decision nobody has made; it delegates to _mint_id
rather than restating the hash (the (p) rule: one keying rule, one copy).
The per-approach set REPLACES the run-level pair rather than joining it — the
selected approach is already among them, and writing both would count it twice in
hitl pending. The selected one carries the run's final outcome, so the outbox can
never disagree with the RunResult; the others carry the validator's verdict, the
only falsifier that ran on them.
mandate.py is deliberately untouched: hanging a ValidatedProposal off a coverage
row would drag validator — and pulp — into a module kept to pydantic+stdlib for
D7 portability, so _evaluate_mandate returns the evaluated outcomes alongside.
Ran it, not just tested it: a real CLI run wrote six artefacts and hitl pending
listed three rows. It also showed the honest edge — three approaches that produce
an identical candidate share one content-hash key, so one verdict settles all
three. That is correct (they were one candidate), and it is now documented.
Load-bearing MEASURED (tests/test_a5_per_approach_artifacts_loadbearing.py) against
the whole 750-test suite, five mutations all red: detach the per-approach writer ·
drop approach_id from the join key · reuse the run's verdict id · reuse the run's
provenance stamp · widen the filename but not the payload. Control: on a full
detach exactly the 5 new tests fail and 745 pre-existing ones stay green — the
no-mandate path is inert, and writes neither the filename segment nor the field.
Docs: bestille-en-kjoring.md (what the commissioner gets) + ekspert-svar.md (what
the expert's queue looks like, and that "rejected" is the validator's verdict on
the numbers, never a professional judgement of the idea).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VtRd8y1PDPGwkrRXFhubqr
Krav 3, and the operator chose the run path explicitly: the external service must
be reachable WHILE the run works, not only when documents are ingested. Until now
the run path had one in-process tool against a local folder — and on the bundle
path the agents had no tools at all.
MAF already ships the client (MCPStdioTool / MCPStreamableHTTPTool, verified in
the pinned 1.9.0 with allowed_tools and request_timeout), so `mcp_tools.py` owns
only what MAF cannot decide for us: which servers a run may contact, which of
their tools it may call, how long it waits, and where the credential comes from.
This is a DIFFERENT seam from ingest_mcp.py on purpose — that one pulls source
documents before a run and speaks to null-argument tools. Same protocol, different
job.
Every refusal is a live hazard, not tidiness. An empty allowlist would let the far
end decide what the agents may call, so naming the tools is mandatory. A
non-positive timeout is an unbounded wait against a third party. An unknown field
is refused rather than ignored, which is also what keeps a literal secret from
being parked in the config — there is no field for one, only the NAME of an env
var. A named-but-unset credential refuses instead of calling anonymously, because
an anonymous call can succeed with the wrong scope.
Egress is declared, always. Every server and permitted tool is named in the run
announcement before the first call — including when no --mandate is given, which
was a real hole: the announcement only printed with a commission, so configuring
servers without one would have contacted third parties with nothing printed at
all. --live-dry-run still opens nothing, because the tools are entered after the
dry-run cut: the promise to stop before the first call now covers egress too.
Threaded through BOTH modes. A flag accepted in one mode and silently dropped in
the other is the defect class this CLI refuses by name.
Load-bearing MEASURED against the whole 744-test suite, four mutations all red:
build the tools but never hand them to the agents (2) · never enter the
AsyncExitStack, so they are constructed and useless (1) · never declare the egress
(2) · drop the allowlist on the built client (1).
Two live docs claimed MCP was unwired in the run path; both corrected rather than
left to rot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULCqjLF61rehj5cZmdUoR3
`docs/bestille-en-kjoring.md` is the commissioning half of the expert-facing pair
(`ekspert-svar.md` is the judging half): the mandate file field by field, how to
run it, and — separated deliberately — what a commission does NOT do. It directs
what is evaluated, never what is approved.
Registered in _LIVE_DOCS, so it cannot silently fall behind the code.
The example output in it is COPIED FROM A REAL RUN, not composed, and running
that run is what found the defect fixed here: three approaches against the same
cost line each validated at 30000 NOK, and the settlement printed
"Validated total: 90000 NOK". Commissioned approaches are ALTERNATIVES — they
usually attack the same line — so summing them reports money the project cannot
realise. A domain expert reading that total would reasonably believe the run
found 90k.
The settlement now reports how many approaches held and which one the run
carries: a selection, not an arithmetic claim. That also removes the last money
addition from this module, which is the right place for it not to be.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULCqjLF61rehj5cZmdUoR3
Krav 2: a run must be clear about what it shall do and achieve. `--mandate`
makes both ends explicit. BEFORE the first (paid) model call the run prints what
it was commissioned to do — objective, every named approach, whether its own
proposals are allowed, the scope, the caps, and which external services it will
contact. AFTER the run it settles: one row per approach, validated with the
figure, rejected with the validator's reason, or not evaluated with why.
The announcement is placed with the other refusals and above the scripted
banner, for the reason the required-args guard was hoisted there: a refused run
must not first print a banner about work it never did. `--live-dry-run`
announces without contacting anything, so a commission can be inspected before
it costs money.
`settle` renders the goal verdict it is GIVEN and never decides it. `ledger.to_ore`
is the framework's one NOK->øre conversion and the goal comparison already runs on
quantised integers, but `mandate.py` cannot import it without dragging `verdicts`
— and therefore agent_framework — into a deliberately framework-neutral module,
while a private copy of a money conversion is exactly the (p) defect. So the
caller decides and this renders; a goal figure without a decided verdict makes no
claim at all.
The caps the announcement prints come from named constants shared with
`run_project`/`run_portfolio`'s defaults — a second copy could drift and make the
announcement describe a run that never happened.
Load-bearing MEASURED against the whole 717-test suite, four mutations all red:
detach the announcement (4) · detach the settlement print (2) · make the CLI
loader tolerant (2) · announce the commission but never hand it to the run (2).
The last one is the one that matters: without it, a run could print a commission
it had no intention of executing.
DEVIATION from the approved plan, stated rather than quietly dropped: --goals is
still refused outside portfolio mode. Accepting it in single-project mode would
have admitted a flag whose documented function (the goal-stop against the ledger)
still does nothing there — the same accepted-but-inert defect --embedder-config
was just fixed for. The mandate's success_criteria carries "what shall this run
achieve" in the expert's own words instead; the numeric target stays portfolio-level.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULCqjLF61rehj5cZmdUoR3
Carrying an approach into the prompt only half-answers krav 1. The expert asked
for their approaches to be CONCRETELY EVALUATED, which means each must reach a
verdict and each verdict must be visible. `run_project(mandate=...)` evaluates
every commissioned approach in turn — the run's own proposal last, when allowed —
each under the SAME meter. No new loop: the caps already in force are the bound.
`RunResult.coverage` is the settlement, one row per approach: validated (with the
figure), rejected (with the validator's reason verbatim), or not_evaluated (with
why). `not_evaluated` is the row that earns the type its keep — an approach the
run never reached must be reported as unreached, because an omitted row is
indistinguishable from an approach nobody ordered. That silence is the defect
class krav 1 is asking us to remove.
Budget exhaustion mid-list is reported, not swallowed. But if the FIRST approach
exhausts it there is nothing honest to return, so BudgetExceeded propagates
exactly as before — a run that produced nothing must still fail loudly.
RunResult stays single-outcome (portfolio aggregation, outbox artefacts and HITL
keying all rest on that). The choice is deterministic: highest validated saving,
ties by mandate order — never whichever ran last.
Load-bearing MEASURED against the whole 702-test suite. FIVE mutations red:
evaluate only the first approach (5 red) · drop the rejected rows (3) · ignore
allow_own_proposals (1) · select produced[-1] (1) · select produced[0] (1).
The sixth measurement is why this commit exists in this shape: the ordering
mutation FIRST STAYED GREEN. The test had placed the bigger approach last, where
"highest saving" and "whichever ran last" give the same answer, so an
order-dependent implementation passed it. A scenario that cannot separate two
implementations proves nothing about either — the test now pins BOTH orderings,
and each mutation direction fails one of them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULCqjLF61rehj5cZmdUoR3
Operator feedback: fagpersoner must be able to name the approaches a run shall
evaluate for a project, and/or ask the system for its own. Today the hypothesis
prompt is hardcoded ("Propose ONE concrete cost-saving measure") and the only
expert-facing lever, --dimension-config, FILTERS what may pass the scoping gate
rather than DIRECTING what is spent attempts on. This is the input that was
missing.
`mandate.py` is the typed commission + a fail-fast loader (mirrors
`load_dimension`/`load_goal_config`): missing or malformed refuses, because a run
must never proceed on a silently degraded commission — the coverage report would
then describe work nobody ordered. Stdlib + pydantic only, so it joins
`_MAF_FREE_MODULES` and can be mirrored to the D7 sibling.
Two refusals carry real defect classes: an EMPTY commission (no approaches and no
own proposals) is a caller error, not a result; and a duplicate approach id — or
one claiming the reserved OWN_PROPOSAL_ID — would collapse two coverage rows onto
one key (the S3.2 key-collision class), which is exactly the silence the coverage
report exists to prevent.
The numeric target is deliberately NOT duplicated here: it already lives in
GoalContract, and two copies of one number drift apart ((p) precedent). The
mandate carries intent; `announce` merely restates the figure.
`_build_messages(approach=...)` switches the opening instruction from *find one*
to *quantify THIS one*, carrying the expert's label and description VERBATIM —
the description is the reason the approach is worth trying, the one part the model
cannot infer from cost data. `approach=None` is byte-identical to the previous
prompt, so every existing run and golden is untouched.
The gate is unmoved: `validate_proposal` is called exactly as before. A
commissioned approach gets no discount — the expert directs what is EVALUATED,
never what is APPROVED.
Load-bearing MEASURED against the whole 695-test suite, four mutations all red:
detach the approach injection (2 red, control stayed green) · let a commissioned
approach bypass the validator · make the mandate loader tolerant · drop the
empty-commission refusal.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULCqjLF61rehj5cZmdUoR3
CORRECTION TO b33ea00. That commit message called the v0.4.0 behaviour a
regression that "breaks our §6 removal path". That claim is WRONG, and this
commit is the record of it — the history is not rewritten.
What v0.4.0 actually introduced is §10.2 per-manifest ownership.
`_is_ingest_owned` (materialize.py:133-160 at v0.5.0a2) compares the
`ingest_manifest` reference's STEM against the running manifest's stem, and
its own comment says why the stem rather than the whole stamp: an EDITED
manifest (new sha -> new stamp) must still reclaim the files its previous run
wrote, while a DIFFERENT manifest sharing the bundle keeps its own.
Our fixture named manifests `manifest-{len(extractions)}.json`, so "the same
manifest, edited" silently became `manifest-2.json` -> `manifest-1.json` —
two stems, i.e. two rival manifests. Refusing to overwrite was CORRECT
behaviour. The fixture was written when ownership was manifest-agnostic and
the naming was pure convenience; v0.4.0 made that convenience load-bearing.
`_write_project` now takes an optional `filename`, and the re-ingest case
pins one stem across both runs, so it means "an edited manifest" on every
version rather than by accident.
MEASURED both ways: green at the pinned v0.3.2 (the distinction is invisible
there), and the WHOLE suite green at v0.5.0a2 — 668 passed. So there is no
technical blocker to the version move at all. What remains is not technical:
1. v0.4.0+ makes `llm-ingestion-guard>=0.2,<0.3` a hard runtime dependency
(v0.3.1/v0.3.2: `dependencies = []`), which moves two invariants this
repo publishes. Operator's call.
2. v0.5.0a2 is an alpha its own CHANGELOG scopes to a named pilot set this
repo is not in. llm-ingestion-okf's call; asked via coord.
The false report was corrected upstream the same hour it was sent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GyAbxJoyypnLLUDcMvnKh8
The pin had sat at v0.3.1 with STATE calling the hold "deliberate" and
recording no reason. Measured: no coord message ever announced v0.4.0 or
v0.5.0a* to this repo, so the hold was drift wearing a decision's clothes.
v0.3.2 is a pure fix (frontmatter and index labels emit verbatim; only
source_query is whitespace-collapsed, per ingest-spec §5), keeps
`dependencies = []`, and is green here: 668 passed.
WHY NOT FURTHER, both measured rather than assumed:
1. v0.4.0 introduces a REGRESSION that breaks our §6 removal path.
Bisected v0.3.2 OK / v0.4.0 RED with a minimal repro: materialize a
bundle, then re-materialize it with a CHANGED manifest, and the library
no longer recognises its own stamp —
MaterializationError: generated filename 'ingest-costs.md' collides
with an existing file that does not carry the ingest stamp
The stamp carries the manifest's name+hash (`ingest_manifest: m2@…`), so
editing a manifest makes every file it previously wrote look curated.
Re-ingesting the SAME manifest is fine, which is why fixtures miss it.
It is `tests/test_ingest_loadbearing.py::test_reingest_with_active_
removal_preserves_promoted_and_curated` that catches it. Reported
upstream; not ours to fix.
2. Everything past v0.3.1 adds `llm-ingestion-guard>=0.2,<0.3` as a HARD
runtime dependency (v0.3.1/v0.3.2: `dependencies = []`). That flips two
documented invariants here — pyproject's "zero runtime deps" comment and
the STATE marker line the guard repo reads machine-readably ("not a
runtime dependency today"). An operator decision, not a version bump.
3. v0.5.0a2 is an alpha whose own CHANGELOG scopes it to a named pilot set
— portfolio-optimiser-claude, the marketplace catalog, claude-code-llm-wiki
— and says "do not pin this tag outside the pilot set", with the v0.2
surface free to change without a deprecation cycle. This repo is not a
pilot. Joining is llm-ingestion-okf's call, requested via coord.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GyAbxJoyypnLLUDcMvnKh8
Walked Door A from a fresh clone: materializing the file-family golden
manifest writes index.md plus one concept file per extraction, and pointing
the run at the result is refused —
run refused: IR projection not found in bundle: 'validator-input.json'
A clean fail-fast, but nothing adopter-facing said it was coming, while the
README actively invites it ("swap --bundle-dir for your own bundle"). The
run path needs the bundle's IR projection, which ingest does not and cannot
produce: ingest materializes source documents, the projection states the
candidate measure. Both docs now say so, with the shape reference named.
Also corrects a live-doc claim that was wrong in both halves: the MCP
timeout is `anyio.fail_after` nested inside both task groups, not
`asyncio.wait_for`, and `tests/test_ingest_golden_mcp.py` covers it
(verified — 2 passing timeout tests). And no bundled example ships a
`cost-baseline.json`, so the text no longer implies one does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GyAbxJoyypnLLUDcMvnKh8
Walked from a fresh clone: `--embedder-config` was accepted in every mode
without `--semantic-retrieval` and then had no effect whatsoever. MEASURED,
not inferred — an injected embedder is consulted ZERO times with the flag
off and once with it on, because the only consumer is the HybridRanker that
flag builds; the default StructuralRetriever takes no embedder at all.
That is the silent-ignore this CLI's flag contract exists to prevent, and
the same ground on which `--semantic-retrieval` itself is already refused
when it cannot take effect.
REFUSED, not wired — the opposite call from `--scripted-replies` in
portfolio mode, and for a stated reason: there the seam already existed, so
refusing would have left a whole mode without an offline door. Here there
is nothing to wire to.
Mode-independent (both modes gate the embedder on the same flag) and placed
ABOVE the scripted door, mirroring the required-args hoist: a refused run
must not first print a banner claiming a scripted loop closed.
Five mutations against the WHOLE suite, all red, each isolating one seam:
detach the refusal (3 red) · scope it to single-project mode (portfolio arm
red) · move it below the banner (banner arm red, rc intact) · build the
ranker unconditionally (the zero-consultation measurement red) · ignore the
injected embedder (its control red).
663 -> 668 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GyAbxJoyypnLLUDcMvnKh8
An independent reviewer found it and the claim was verified by measurement
before being accepted, not taken on trust.
test_portfolio_scripted_pass_makes_no_real_client raised from the patched
_default_factory to prove the scripted portfolio pass never builds a production
client. It cannot: the factory is called inside run_project's coroutine, and
run_portfolio gathers with return_exceptions=True, so the AssertionError was
collected into a RunFailure and never escaped. Measured directly -- with BOTH
the client_factory wiring and the rc rule detached, the blade stayed GREEN. It
was going red on rc alone, which means it was testing defect B while claiming
to test defect A. The original seven-mutation sweep did not catch this because
each mutation was applied singly, and dropping the wiring alone still flips rc.
Replaced with a call sentinel: a list appended inside the factory and asserted
in the test body, which the wave handler cannot swallow. Re-measured -- red on
the wiring detach alone, and red on both detaches together.
The docstring now also states what the sweep could not: blades 1 and 8 are
environment-conditional. The local profile points at loopback, so on a machine
running a local model server the wiring detach would make real calls and could
complete the pass. Their red was real on the machine it was measured on and is
not portable; blade 2's is.
Separately, the budget-stop print is marked as defensive and currently
unreachable from main(), because main() never constructs a PortfolioMeter and
every write to budget_stop is gated on one -- the strict=True precedent
directly above says untested future-proofing must be labelled as such. The
README claim that a portfolio pass reports a cap stop is corrected to say the
cap has no CLI flag yet. Noted for whoever wires that door: BudgetRefused is a
RuntimeError and the existing except clause would not catch it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0118noV9rCfrdREH26XqZB5z
The operator is not a domain expert, so the domain content is mine to own --
and the one thing the loop asks a human for is exactly the thing no example
existed for. docs/ekspert-svar.md is written for whoever has to deliver the
verdict: the two forms a judgement can take (a --rationale string during the
run, a JSON file in the inbox for later runs), where each field comes from, and
four complete paste-ready answers.
Every command and every verdict in it was RUN from a fresh clone before it was
written. The `hitl pending` line quoted is verbatim output. The rejection
answers close the gap STATE has carried since the demo shipped: the README
shows the VALIDATOR refusing a number, but nothing showed an EXPERT refusing a
proposal whose numbers are fine -- the only judgement in the whole loop that a
machine cannot make. Two rejection shapes are given, because "not feasible
here" and "right measure, wrong cost base" teach the system different things.
Everything is marked AI-authored and not verified professional judgement.
Also corrects the --outbox-dir help text, which claimed sharing a folder with
--verdict-dir "re-ingests raw agent output past the Step-8 promotion gate".
Measured, by pointing both at one folder and running twice: it does not. The
outbox artefacts are named {run_id}-*.json and carry none of the verdict keys,
so the tolerant inbox loader skips them and the run is unaffected. The hazard is
real but latent -- a future verdict-shaped artefact in the outbox -- so the
warning stays and says what is actually true. This also answers STATE's open
question about enforcing the distinction in the CLI: no. There is no reachable
contamination to refuse, and a guard for an unreachable case is the kind of
error handling this repo declines to write.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0118noV9rCfrdREH26XqZB5z
Steps 6 and 7, each run verbatim from a fresh clone before being written.
Step 6 is the portfolio pass, which the walkthrough could not reach until the
scripted door was wired into portfolio mode. It is also the clearest single
demonstration the framework has: all four reference projects carry a cost line
01.1 at four different amounts, so one unchanged proposal yields one
ValidatedProposal and three Rejection -- the gate is anchored to each project's
own baseline, not to the proposal's internal arithmetic. The shared verdict id
is explained rather than hidden: a verdict is keyed on the candidate, not the
project, and that key is how a later run finds the earlier judgement.
Step 7 documents the value report and, more importantly, the gap a downloader
hits first: nothing in the shipped code writes a savings ledger. Measured, not
assumed -- no .save call on a ledger exists outside the library API. That is by
design and is now said out loud: the ledger records savings actually realized in
the world, which is not a conclusion the system may draw from its own proposals.
A validated proposal is a claim; a ledger entry is a result. The refusal a
downloader sees before creating one is quoted, and the library snippet that
creates one is shown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0118noV9rCfrdREH26XqZB5z
Walking --portfolio end-to-end as a downloader would -- which no session had
done -- surfaced two defects of a class this repo already legislates against.
(A) --scripted-replies was silently DROPPED in portfolio mode. main()'s
portfolio dispatch returned before the block that builds the scripted client
factory, and the flag was absent from the single_only refusal set: neither
honoured nor refused. Measured against the shipped reference portfolio: no
banner, four real model calls attempted, four APIConnectionError. The previous
session joined this flag to the --report allowlist and missed the portfolio
partition. Resolved by WIRING rather than refusing -- run_portfolio already
exposes the same client_factory seam, and refusing would have left portfolio
mode with no offline door at all for an adopter without a model budget. The
scripted block is hoisted above the dispatch; the single-project required-arg
and semantic-retrieval refusals are hoisted with it so an incomplete argv is
still refused BEFORE the honesty banner could claim a scripted run happened,
and the refusal order within single-project mode is unchanged.
(B) A portfolio pass reported one of its four outcome channels. failures
(S3.3 collect-and-continue) and budget_stop (S3.4 global cap) never reached the
operator and rc was unconditionally 0, so the four-failure pass above printed
NOTHING and exited 0 -- silence read as success. BudgetStop is a separate field
precisely so exhaustion can be told from success; the CLI showed neither.
Failures now print to stderr with project id, error type and message; the
budget stop prints its four numbers; rc is 1 iff something raised. A budget
stop alone stays rc 0: exhaustion is a structured stop the operator asked for
by setting a cap, not a crash. Completed runs still print, so the non-zero rc
does not undo collect-and-continue.
Load-bearing MEASURED against the whole 662-test suite, seven mutations all
red, including both controls: detach the client_factory wiring - make the
banner a single-project courtesy again - detach the failure print - revert rc
to 0 - detach the budget-stop print - print the failure line unconditionally
(control) - print the budget-stop line unconditionally (control).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0118noV9rCfrdREH26XqZB5z
Operator ruling 2026-08-05, which settles decision (g): planning documents are
generally never public, and what OUR OWN sessions generate does not go out on
the forge at all. The example itself stays public so others can run the
process.
`.claude/projects/` is the Voyage session workbench -- 25 briefs/plans/reviews
this project's own sessions produced. Untracked and gitignored, exactly as
STATE.md already is, and for the same stated reason: this repo has a public
mirror, so that class of material is local-only rather than tracked.
The line is drawn at who wrote the document, and it is drawn deliberately:
`docs/plan/`, `docs/research/` and `docs/rapport/` stay tracked. Those are
curated, dated documents written for the repo's readers, three of them linked
from the README as the decision record. Move that line if it was meant wider.
Two files were NOT process artifacts and are not deleted. Both
`build_fixture.py` scripts are cited by tracked tests
(`test_ingest_golden_sql.py`, `test_ingest_golden_http.py`) as the documented
rebuild path for byte-exact goldens -- reproduction code that had landed in the
wrong directory. Moved next to the goldens they build; both docstrings updated,
so no tracked file is left pointing into an untracked tree (verified: the only
remaining `.claude/projects` string in a tracked file is the .gitignore rule
itself). One prose reference in the dated Foundry auth recipe was dropped for
the same reason.
652 tests still pass.
Does NOT address the 27 of these already readable on open/ since the S12
release -- untracking stops future publication only. That retraction is a
separate operator decision and is deliberately not taken here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GWsexbQjPo9rsV3aUE54ZS
Measured before writing: the commands existed but were scattered across the
mode partition, and there was no path a newcomer could walk end to end. The
capability gap that made a complete offline walk impossible is closed in
3abc61b; this is the door onto it in the README.
Five steps, each RUN FROM A FRESH CLONE before being written down (git clone +
uv sync + uv run pytest -> 652 passed): read the knowledge base, watch the
learning loop close, run the loop with your own scripted answers, watch the
validator say NO, and price a real run before spending anything. Step 4 is the
one that was missing entirely -- the demo only ever showed a yes, and a
refusal carries far more weight than another approval.
Also documents that `Rejection (..., decision=approved)` is not a
contradiction: the first is the validator's outcome, the second echoes the
human's recorded verdict. Documented rather than changed -- altering a public
output format is the operator's call, not a side effect of writing docs.
`costsim` gets its first mention in the README at all; it was finished, tested
and completely invisible from the surface.
Also points the commons link at open/ (published 2026-08-04). NB: the
repo-standard v0.3.0 register still lists 19 repos and does not know that repo,
so pointing at the correct live URL now trips a false LINK-DEAD. Reported to
org-ops; the URL answers HTTP 200.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GWsexbQjPo9rsV3aUE54ZS
An adopter without an API budget had two half-doors and no whole one.
`--live-dry-run` takes their own bundle but stops before the first model call
(`run_project` returns a DryRunReport), while `portfolio_optimiser.simulation`
runs the complete loop but only over ITS bundle with ITS scripted answers.
The seam for the missing third case -- the whole loop over your OWN data,
offline -- already existed as `run_project(client_factory=...)` and had zero
CLI exposure. This is the door onto that one seam, not a second implementation
of it (`scripted_factory` is imported lazily; `simulation` imports `run`, so a
module-level import would be circular).
The honesty banner is part of the feature, not decoration (maalbilde §1): a
scripted run that reads like a model run is worse than having no offline mode,
so every scripted invocation prints what is real (context navigation, debate
plumbing, deterministic validator, verdict) and what is not (the answers).
The two offline modes are mutually exclusive rather than one silently winning,
`--report` mode refuses the new flag by allowlist, and a replies file that
cannot serve the run is refused at the door rather than surfacing as a KeyError
mid-run.
Load-bearing MEASURED against the whole suite (645 -> 652), six mutations all
red: detach the wiring · detach the banner · detach the dry-run exclusivity ·
drop the flag from the --report allowlist · make the loader tolerant · control
(print the banner unconditionally).
The --report blade was measured GREEN first: with a non-existent ledger path
the load failure refused before the gate and masked it entirely. Rewritten
against a valid saved ledger, so rc 1 can only come from mode-exclusivity.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GWsexbQjPo9rsV3aUE54ZS
CONTRIBUTING.md told strangers to fork and "open a Pull Request". Measured
against the forge itself, that door is shut: the API reports
has_pull_requests=false / has_issues=true for open/portfolio-optimiser, and
the org's published CONVENTIONS.md states the position -- "Issues velkommen
som signaler. PRs ikke akseptert. Fork-and-own er anbefalt adopsjonsmodell."
Replaces the PR workflow with the two routes that are actually open (issues
as signals, fork-and-own), and keeps the engineering standards the section
carried -- load-bearing tests, blocking validator, ruff/mypy, Conventional
Commits -- reframed as what the project holds itself to, which is what a
forker needs and what an issue is weighed against.
Also points CLAUDE.md's commons link at open/portfolio-optimiser-commons
(published 2026-08-04). The `commons` git remote still resolves via ktg/ and
is deliberately left alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GWsexbQjPo9rsV3aUE54ZS
ddaae5d chore(release): publiseringsklar for open/ — README for standalone rot + MIT + policy-filer
f98b287 docs(plan): V1 RATIFISERT — og :275 er en andre tabellrad, ikke prosa
d6bced7 docs(plan): SS11 ankrer ikke SS8 — funnet var reelt, men ikke raden som ble bestilt
3174475 docs(plan): §7.2 — feilanker-failuremoden var ikke hypotetisk, den inntraff
git-subtree-dir: shared
git-subtree-split: ddaae5d637ba4ee8425291d99cfe5c8f7b632001
Two rules existed. verdicts._unquote took quotes off correctly; the
bundle_context title renderer stripped only `"`. Measured before the fix:
`title: " Spaced "` rendered as `## concept: Spaced ` (the whitespace
half kø-(a) named), and `title: 'Single'` rendered its quotes verbatim —
both ordinary YAML a hand-authoring curator writes, and both reach the
agent's read-context.
The (p) defect class: a duplicated conversion drifts, and the drifted copy
decides something. Same fix shape — ONE source, owned by the module that
owns parse_frontmatter. okf.unquote_scalar is now the rule; verdicts
delegates by identity, so the structural key that _mint_id hashes is
unchanged.
The commons-owned nav-goldens could never have caught this: every golden
title is double-quoted with no inner whitespace, so both rules render them
byte-identically. That is asserted as a control, and it goes RED if a
future golden gains a discriminating title.
Load-bearing MEASURED against the whole suite, three mutations:
- weaken the renderer back to .strip('"') -> ONLY the 2 new rows red,
643 others green (incl. nav-goldens) = it covers ground nothing did
- reintroduce a private _unquote copy in verdicts -> only the identity
test red, 644 green (the copy is behaviourally identical TODAY, which
is exactly why identity is the only thing that catches the class)
- (control, kø-(n)) add `import numpy` to a second src module -> the
existing test_semretrieval_is_the_sole_numpy_importer goes red alone,
confirming that gate is live rather than green-but-dead
638 -> 645 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DVMsih6tXx39VMyq6wJ5H7
docs/extending.md claimed SEMANTIC_WEIGHT_DEFAULT = 0.5 long after the code
lowered it to 0.25. It was found by accident while editing the neighbouring
line; nothing in the suite would ever have caught it. This is that gate.
Both bounds were MEASURED, not assumed:
- SCREAMING_CASE discriminates exactly. Across every numeric `name = value`
code span in docs/ it selects the 2 real module constants and rejects all
10 kwargs/locals (max_attempts=3, concurrency=3, realiseringsgrad=0.79).
Ordinary prose stays freely editable, so the gate has no reason to be
switched off.
- Dated documents are observations, not contract. A spike finding or a July
review records what was true when measured; rewriting it to track the code
would falsify the record.
Measuring also rewrote the ambiguity rule: SEMANTIC_WEIGHT_DEFAULT is bound
in both semretrieval and run (a re-export), so a "same name in two modules"
check would have been RED on today's code. Only DIVERGENT values are refused.
Fail-closed throughout, per write_concept_file / read_spend: an unknown
constant name is an error rather than a skip, and a document that is neither
listed live nor recognisably archived goes RED asking to be classified —
otherwise a new guide would be silently unguarded.
Load-bearing MEASURED against the whole 638-test suite:
- drift the DOC (the original defect) -> only this gate goes red; the other
637 stay green, so it covers ground nothing else did
- drift the CODE -> this gate and the semretrieval weight gate both go red
- make an unknown name tolerant -> red
- drop the only citing doc from the live list -> red (twice: coverage and
classification)
- add a new unclassified guide -> red
- (control) remove this gate entirely, code still drifted -> the adjacent
semretrieval gate still goes red, so nothing is masked in either direction
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D9AAyWtMqr4HjftKaegTtS
Ran `repo-standard` (v0.1.1, class `standalone`) and fixed everything it
flagged as ERROR, plus the WARN links that were genuinely dead.
README first screen:
- opening line is now byte-identical to the forge description, so
description == catalog == README is machine-checkable (badges moved below).
- `## Install` (required for class `standalone`): clone + `uv sync`, stated as
clone-only because the shared spec, persona skill and example bundles under
`shared/` are read from the working tree at run time. `uv run pytest` named as
the verification, with the fact that no CI runner exists said out loud rather
than implied by a badge.
- `## Non-goals` (required): the five limits already binding in CLAUDE.md —
not a compliance product, not a portfolio-level reallocator, not autonomous
decision-making, not turnkey, not a model benchmark.
Dead relative links (measured, not guessed):
- `docs/plan/2026-07-10-sesjonsplan-fase2-6.md` pointed at
`../2026-07-14-revisjonspakke-DF-DI.md` six times; the file sits in
`docs/plan/`, not `docs/`. (The sibling `../review-2026-07.md` links are
correct and untouched.)
- the Fase-1 spike brief linked repo-root-relative from
`.claude/projects/…/`; re-anchored with `../../../`.
The one remaining README ERROR was a gate false positive: `checkInternalLinks`
resolves targets against `git ls-files`, which lists files only, so a link to a
directory can never resolve. `[shared/](shared/)` now points at
`shared/README.md` — a better target anyway, since that file carries the
pull-only subtree rule. Not fixed here: the classifier lives in another repo.
Remaining WARNs are all inside `shared/`, deliberately untouched: it is a
pull-only commons subtree, and the nav-golden files are byte-level fixtures
that gate `test_nav_golden_*` — four of them are OKF bundle-internal links,
and the `/etc/passwd` ones are the negative escape fixture doing its job.
Suite green: 630 passed, 4 skipped (markdown-only diff; no test touched).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ri3aVJPfynCZtHRhesCzUH
`cosine`'s docstring claimed its guard was load-bearing because "a NaN reaching the
ranking sort key would corrupt ordering silently rather than failing loudly" — but the
guard tested `norm == 0.0` only, which a NaN or inf norm passes straight through. The
claim was prose, not behaviour.
Measured, not assumed: `cosine(unit, nan_vector)` AND `cosine(unit, inf_vector)` both
returned `nan`, and a NaN sort key made ranking INPUT-ORDER-DEPENDENT — six permutations
of the same three candidates produced four distinct orderings. That defeats the total
order `HybridRanker` documents ("`id` makes the result independent of input order").
Refuse rather than coerce, and deliberately NOT symmetric with the zero-norm branch: a
zero vector is a legitimate handled state (`FakeEmbedder` returns `np.zeros` by design),
whereas a non-finite component only ever means the INJECTED embedder is broken. Scoring
it `0.0` would launder that into "no semantic similarity" while ranking proceeded on a
forged signal — validation, never repair, mirroring `read_spend`.
Reachable via the documented `Embedder` extension point, not the shipped fake; scoped to
the norms (90% principle — a finite-normed dot-product overflow is not chased).
Also corrects `docs/extending.md`, which stated `SEMANTIC_WEIGHT_DEFAULT = 0.5` while the
code has said `0.25` since the weight was lowered.
625 -> 630 tests. Load-bearing MEASURED against the WHOLE suite, five mutations all red:
detach the guard entirely · coerce to 0.0 instead of raising · check only the first norm ·
drop "non-finite" from the message · (control) detach the zero-norm branch, which fails
ONLY the zero-norm test — the new guard does not mask the existing one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018V9vNBmxAmgJ2JMoHByiHS
Documents the anyio.fail_after-vs-asyncio.wait_for finding and the
cancelled_caught ownership gate, alongside the existing kø-x task-group
invariant it extends.
Advisor review of the prior commit (5269b7d) found the except TimeoutError
branch was wider than the brief asked for: builtin TimeoutError is also
socket.timeout (3.10+) and asyncio.TimeoutError (3.11+), so any TimeoutError
reaching that clause got relabeled mcp_timeout regardless of source. Gate on
anyio.CancelScope.cancelled_caught instead, mirroring _unwrap_ingest_error's
ownership rule (own it, wrap it; otherwise, untouched).
Measured: no live trigger exists today (MCP's own internal read-timeout
converts to McpError before reaching us; a server-side TimeoutError becomes
an ordinary isError result) -- pinned with a synthetic test raising from
StdioServerParameters construction, inside our fail_after scope but before
either nested task group, so it arrives ungrouped. Four mutations red against
the full 625-test suite: drop the translation, revert to asyncio.wait_for,
relabel the code, and drop the cancelled_caught gate.
Also promotes anyio to a declared direct dependency (was transitive via mcp
only) -- ingest_mcp.py now imports it directly.
asyncio.wait_for cancelling stdio_call_tool's run() from outside the anyio
task groups it awaits (stdio_client, ClientSession) never surfaced a
TimeoutError: measured against a real hanging server, the mismatch produced
an anyio.BrokenResourceError wrapped in a BaseExceptionGroup instead. Moving
the deadline to anyio.fail_after, nested inside both task groups, lets
anyio tear down its own structure cleanly and raise a plain TimeoutError,
which is now translated into IngestError(code="mcp_timeout") alongside the
mcp_tool_error/mcp_non_text_content family.
BudgetExceeded carries kind/limit/observed as ONE structured stop event, but only
`observed` was undefended. Measured against the whole suite before writing anything:
four of five raise sites (TokenMeter.charge, tick_round, and BOTH arms of exhausted())
could report any value at all without a single one of 621 tests noticing. Only
PortfolioMeter.check was covered.
What hid it: spikes/_harness.py carries its OWN copy of BudgetExceeded/TokenMeter, so
the spike suite's `observed` assert never touched the shipped module — the production
tick_round had no direct test whatsoever.
exhausted() is the only site that CHOOSES a ledger (the S3.4 pre-call guard), so a
refusal naming portfolio_tokens while reporting the run's own spend would misdirect
every reader of it. Both arms are pinned with observed != limit on purpose: at
exactly-exhausted the two coincide, and a test written there would pass on an
implementation that echoed the cap back as the spend.
No defect in the values themselves (unlike kø-x and kø-p) — the triple was coherent at
all five sites; the gap was purely coverage.
Load-bearing MEASURED: nine mutations, all red — five observed mutations (including the
control) and four echo mutations. 621 -> 623 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LW749xcXQmVEgdipB6KNm4
Two quantization orders existed and met at exactly one comparison.
SavingsLedger quantizes every realized candidate to integer øre and sums the
ints; run.py's goal baselines summed Project.total_cost FLOATS across items and
projects and quantized the total once. _goal_limit_if_reached compared the
former against a threshold derived from the latter — so whether a portfolio pass
stops early was decided by two differently-computed sides.
Measured divergence: three 60000.005 NOK lines are 18000003 øre quantized first
but 18000001 summed first (the float sum drifts to 180000.01499999998).
Decision: quantize per cost line, then sum integers. Each CostItem IS a money
amount — S4.0 made per-line quantity/unit_cost the validator's ground truth — and
integer addition is associative, keeping totals order-independent under the D-D
wave model, which the float fold is not.
ledger.to_ore is now the framework's one NOK->øre conversion; run.py imports it
rather than keeping a private copy (the S4.0 REPLIES precedent).
Measuring the mutations found two further gaps, both now closed: the per-project
baseline is a SECOND call site whose mutation survived the whole suite, and
realize bypassing to_ore with a raw float*100 was caught by nothing.
Load-bearing MEASURED (tests/test_money_quantization_loadbearing.py), five
mutations all red: detach the portfolio baseline · detach the per-project
baseline · reintroduce a private copy in run.py · change the rounding mode · let
realize bypass to_ore. 615 -> 621 tests.
Honesty boundary: sum_claimed_saving_nok (run.py:_aggregate) is deliberately
untouched — a float NOK reporting field that is never quantized and never
compared against the ledger, hence outside the ordering defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WiY53sm8JFqk7NN75g5wRS
`stdio_call_tool` shipped never having been executed end to end — docs said so
explicitly. Running it found a real defect: `stdio_client` and `ClientSession` are
each an anyio task group, and anyio re-packages anything leaving one in a
`BaseExceptionGroup`. Both errors the transport raises from inside the session
(`mcp_tool_error`, `mcp_non_text_content`) therefore reached callers as exception
groups, never as the `IngestError` the whole Door A path catches and switches on by
`code`. No canned-tool test could see this: they never enter a task group.
`_unwrap_ingest_error` recovers the owned error and re-raises it; anything unowned is
re-raised untouched, so this narrows an exception group rather than blanket-catching.
Duck-typed on `.exceptions` because `except*`/`ExceptionGroup` are 3.11+ and this
project supports >=3.10.
Verified against a REAL server subprocess (a local process costs no model tokens, so
the repo's cost discipline is untouched; the contract tests still spawn nothing):
`examples/ingest-golden-mcp/` + `tests/test_ingest_golden_mcp.py` — byte-identical
golden extraction mirroring the http/sql goldens, plus the tool-error and
missing-`server_ref` branches.
Also recorded: a server on the ingest path must expose a NULL-ARGUMENT tool, so
`datasource.build_mcp_server` cannot serve it (`retrieve_cost_docs(query)` has a
required parameter, verified to return an error result). The two are separate seams
by design.
Load-bearing MEASURED, five mutations all RED: detach the unwrap · detach
`initialize()` · make the error code generic · detach the `isError` branch · change
one byte of the served body.
612 -> 615 tests. ruff + format + mypy clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WiY53sm8JFqk7NN75g5wRS
Every stage of validate_proposal reasoned only about numbers the proposal itself
supplied, so an internally-consistent hallucination cleared the whole gate (F3).
A new stage 0 reconciles each affected_item against the project's CostBaseline
before the CBC solve: an unknown cost code is rejected, and a real code carrying
a quantity/unit_cost outside the configured tolerance (5% default, relative to
the baseline value) is rejected. Validation, never repair.
The baseline argument is OPTIONAL (None = pre-S4.0 behaviour), but both run
paths set it: the road path projects project.cost_items, the bundle path loads
cost-baseline.json when the bundle ships one. Bundles written before the
amendment stay un-anchored, so the commons-owned goldens run byte-identically;
a baseline that exists but is malformed still raises on both loaders.
F8: the method-specific cap now comes from the METHOD_CAPS registry (measure
type -> fraction, injectable) instead of an energy_efficiency string comparison.
The baseline format and tolerance semantics were decided locally — the commons
amendment (D-A pt. 2) never arrived, exactly as in S3.2. D7 mirroring stays open.
Three portfolio fixtures quoted cost codes belonging to OTHER projects; the new
gate caught them. They now quote each project's own lines, and the two copied
REPLIES tables import the single source instead of drifting from it.
Load-bearing measured (tests/test_s40_cost_baseline_loadbearing.py), six
mutations all red: detach the reconciliation stage; detach the magnitude
tolerance; detach the road wiring; detach the bundle wiring; ignore the injected
cap registry; make the optional loader tolerant of malformed content. Control:
with the road wiring detached the repaired portfolio fixtures still pass, so
they are not masking the seam. 597 -> 612 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JdwK7bQ4BZkWH4t8MRDKb4
seed_store_from_bundle keyed EVERY `type: verdict` file on bundle_candidate_features — the single
candidate the bundle's validator-input.json describes. A bundle carrying verdicts about several
candidates collapsed them onto one key, so a verdict about candidate B scored a perfect structural
match against candidate A's query and could be folded into A's hypothesis prompt. The ExpeL
substrate was single-candidate by construction.
A verdict file may now carry its own structural key in frontmatter (affected_codes / measure_type /
claimed_saving_nok); absent, keying falls back to the bundle candidate, so every pre-S3.2 seed keeps
working unchanged. promote_verdict writes the three fields, so a promoted verdict — frequently about
a different candidate than the target bundle's projection — does not impersonate that candidate.
Semantics decided HERE, not pulled: commons' seeding rule (method-spec §3 Steg 1 + bundle example)
has not arrived; we said we would build locally first. D7 mirroring stays open.
- ALL THREE fields or none. A partial declaration raises VerdictFrontmatterError rather than merging
with the bundle candidate, which would mint a key belonging to NEITHER candidate. Validation,
never repair (mirrors write_concept_file); the tolerant-skip rule belongs to the RAW inbox layer.
- claimed_saving_nok parses via json.loads — the SAME literal rule the IR projection went through —
and is written back with str() of the raw value. _mint_id hashes that value, so 30000 and 30000.0
are different keys; a normalising writer would split one candidate's signal across two ids.
- The structural key is signal-free, so it does not weaken the Step-8 no-leak property (Test C green).
Load-bearing MEASURED, five mutations all red: detach per-verdict keying · detach the fields
promote_verdict writes · make a partial/unparseable key tolerant · normalise the magnitude on write ·
remove the fallback (control — breaks the step1 suite at collection, proving the fallback bears load).
589 -> 597 tests. Full gate green (pytest, ruff, mypy).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QkjvTTxrg9LTrmghebfiij
Two tightenings, each measured by a detached-mutation run:
(1) The validator now blocks a claim above the CBC nominal feasible, in ADDITION
to the P90 stage. Neither dominates the other: an upward-skewed assumption band
lifts P90 ABOVE nominal -- so P90 alone passed review counterexample #1 (claim
100k, nominal 90k, band [0.70, 1.40], measured P90 121057) -- while a
downward-skewed band pushes P90 below it. Independent gate, same Rejection type,
existing rejections keep their existing reason.
(2) An assumption band must enclose its item's unit_cost (low <= unit_cost <=
high, inclusive). A band that misses it states a different price rather than an
uncertainty, and every Monte Carlo draw would then sample away from the item's
stated cost. Checked exactly where the Monte Carlo looks bands up -- per affected
item, by code; a band keyed to no affected item is never sampled and so has no
unit_cost to enclose.
The premise was re-verified against ground truth before building on it, not
taken from STATE: 05.2 unit_cost 215 in (200,230), 03.1 310 in (290,330),
ENERGI-TOTAL-EL 1.0 in [0.70,1.40] and (0.8,1.2). No fixture violates it.
The LLM path already catches ValidationError as a meter-bounded retry
(generate.py:138), so the new invariant cannot crash a run.
Mutations, all RED: detach the nominal block; drop the model_validator
decorator; make the enclosure strict. tests/test_bygg_energi_mikro.py and the
commons golden are UNCHANGED and green -- the regression proof.
586 -> 589 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017DDXwUqHVAQQeYE7X1TXy5
Three items on one seam — what a FAILED project does to the wave loop — plus the
snapshot copy they sit next to.
(v) The catch is BaseException, not Exception, and that width was ungated. The
existing collect-and-continue test raises RuntimeError, so it stays green when
the handler is narrowed: measured, the whole of test_portfolio_concurrent_
loadbearing.py (13 tests) passes under the narrowing. asyncio.CancelledError is
the one realistic vector that separates the two — probed first, gather(
return_exceptions=True) COLLECTS it, while KeyboardInterrupt propagates
regardless and could never be helped by a wider catch. Narrowed, a cancelled
member is cast into runs as a fake RunResult and the pass dies in _aggregate,
pointing away from its cause. RED measured.
(t) sum_token_usage excludes a failed project's spend, and that is the honest
answer, not a bug: a run that died before producing a stamp has no provenance,
and inventing one is the fabrication RunFailure exists to avoid. What needed
gating is that those tokens still reach the ledger the global cap is enforced
against — otherwise a repeatedly-failing project burns budget while the meter
reads clean. Pins meter.spent as the pass's real cost, sum_token_usage as the
completed-run subtotal, and their difference as exactly the failed spend. RED
measured against the likely "fix" (sourcing sum_token_usage from the meter),
which is wrong because a seeded meter also carries EARLIER passes' spend; 21
existing budget/portfolio tests stay green under it.
(s) _wave_snapshot uses dataclasses.replace, so a field added later is carried
without touching the function. Not cosmetic: measured, dropping retriever by
hand-enumerating left all 585 tests green — the Step-2 coverage its docstring
credited no longer existed, so the S3.1 retriever seam could be downgraded
mid-pass in silence. Now gated by a property test derived from
dataclasses.fields (not a field count, the shape rejected earlier). The explicit
verdicts copy is retained and separately gated: replace(store) alone shares the
caller's list and takes the byte-identical determinism test RED.
strict=True on the zip is documented as deliberately untested — measured green
when dropped, since gather is built from exactly snapshots, so a test could only
go red by manufacturing a mismatch and would exercise zip rather than this pass.
The new double is registered in the S2.5 consolidation guard's delegating-
overrides list rather than the guard being weakened; it already delegates via
super()._inner_get_response, which test_delegating_overrides_call_super now
enforces on it.
583 -> 586 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MbgTCEZma764i1rHTrzceU