The scripted proposer answered one hard-coded pair of proposals. A second project meant a second
hand-written selector, written under demo-week time pressure -- the risk the week plan names
explicitly (§4, risk 2). It is now a registry: `ScriptedCandidate` entries selected by
`scripted_proposer`, plus `project_id` as an argument to `simulate_learning_loop`.
The open decision was WHAT identifies the candidate in the prompt blob; the plan flagged it as
unverified, so it was measured. Two prompt shapes reach the selector: the debate prompt carries the
whole bundle context, the generation prompt carries `Project: {id} - {name}` plus -- as its context
-- the debate output, which is the selector's own earlier reply. So the cost code and the measure
name are present in the generation prompt only because the script put them there; keying on them
would key the script on its own output. The project id is the one identifier both shapes carry and
the framework stamps.
Validation, never repair: no match, or more than one, raises `ScriptedCandidateError`. A default
reply would answer an unregistered project with another project's numbers, which on screen is
indistinguishable from a correct run; an ambiguous blob is a data problem that must surface at the
rehearsal rather than be decided by registry order.
Load-bearing MEASURED against the whole suite, five mutations all red plus a green control: detach
the project keying - one global flip key - fall back on an unknown project - first-match on an
ambiguous prompt - detach the `project_id` argument. The flip-key test was rewritten mid-measurement
because its first form asserted on the FIRST registry entry, where "the matched candidate's key" and
"candidates[0]'s key" coincide -- it could not separate the two implementations, and proved nothing.
766 passed / 4 skipped. Simulation still exits 0, still prints eight labelled steps, still
byte-identical across two runs.
[skip-docs] README is deliberately untouched: O4 defers the README rewrite to 14-15 August, after
the demo has produced the evidence for the level-2 claim. CLAUDE.md carries the invariant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoHJCKBTjFKcjsfEQyGbzh
219 lines
18 KiB
Markdown
219 lines
18 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to this project will be documented in this file.
|
||
|
||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||
|
||
## [Unreleased]
|
||
|
||
### Added
|
||
- Step 5 is now observable: `generate_via_llm` returns a `GenerationResult` carrying the validator
|
||
falsifications that informed a later attempt, surfaced on `RunResult.refinements`. The offline
|
||
simulation exercises it — the scripted proposer overclaims, the deterministic validator falsifies
|
||
the number, and the refined proposal validates.
|
||
|
||
### Changed
|
||
- **Breaking (library API):** `generate_via_llm` returns `GenerationResult` instead of
|
||
`ValidatedProposal | Rejection`; read `.outcome` for the previous value. The refinement loop's
|
||
bound is unchanged (`max_attempts` + token meter).
|
||
- `simulation.scripted_factory` accepts a per-role reply *selector* over `(prompt, role)` as well as
|
||
a constant reply, so a scripted role can answer differently on a later attempt.
|
||
- The offline simulation's scripted proposer is now a candidate **registry** rather than a
|
||
hand-written reply: `simulation.scripted_proposer(candidates)` builds the selector from
|
||
`ScriptedCandidate` entries keyed on the project id the prompt names, and
|
||
`simulate_learning_loop` takes `project_id` alongside `bundle_dir`. Adding a project to the
|
||
walkthrough is a data entry. A prompt matching no entry — or more than one — raises
|
||
`ScriptedCandidateError` rather than answering with another project's numbers.
|
||
|
||
## [0.1.0] - 2026-08-06
|
||
|
||
First tagged release. There is no prior release, so the entries below describe what this version
|
||
**is**, not what changed since a predecessor.
|
||
|
||
### Added
|
||
- Deterministic backbone: mandatory blocking validator (solver + Monte Carlo against the shared
|
||
golden suite), budget meter with hard fail-fast caps, provenance stamping.
|
||
- Agentic learning loop wired end to end, one load-bearing seam at a time (target-picture steps
|
||
1, 3/4, 5, 7, 8): OKF-navigated bundle context with the gated ExpeL fold, maker-checker debate
|
||
where the checker gates the reasoning, informed refinement (previous rejection reason fed into
|
||
the next bounded attempt), async verdict file inbox, and gated wiki promotion (fail-closed).
|
||
- Offline end-to-end simulation proving the learning loop closes with a scripted client
|
||
(`uv run python -m portfolio_optimiser.simulation`) — plumbing proof, not live-model proof.
|
||
- Framework-neutral shared core in `shared/`: OKF concept + example bundle, golden validator
|
||
suite, and the expert-reviewer persona as an Agent Skill.
|
||
- CLI parity (S5.3): `run.py` drives the whole method from the command line in **two modes** —
|
||
single-project (`--dimension-config`, `--outbox-dir`/`--run-id`, plus the already-wired
|
||
`--bundle-dir`/`--verdict-dir`) and portfolio (`--portfolio` with `--goals`/`--ledger`), the
|
||
latter printing an observable `goal reached: …` line when a savings goal is met. Adds a
|
||
fail-fast `load_dimension` loader and structured refusals (rc 1, no traceback) for misuse and
|
||
mode-exclusivity violations. The prior-verdict fold is on the `--bundle-dir` path only; a
|
||
`--docs-dir`-only run is single-shot.
|
||
- Value report (S5.4): a read-only `--report [--json] --ledger <file>` surface on `run.py` that
|
||
rolls up the accumulated `SavingsLedger` — per-project + portfolio totals (dimension-free-deduped
|
||
integer øre), flagged cross-dimension overlaps (each counted once), and per-entry provenance — as
|
||
a human table or deterministic JSON. Makes no model calls; mode-exclusive (only `--ledger`/`--json`
|
||
permitted with `--report`, which requires `--ledger`). Honest scope boundary: the report core
|
||
(`value_report.py`) now **exists** but is deliberately **not** wired into `costsim`'s
|
||
`kost_mot_verdi` placeholder — that cost-vs-value integration is a separate, deferred step. The
|
||
`costsim` seam note was reworded from the stale "fylles av S5.4 verdirapport" to a truthful
|
||
forward reference so `costsim`'s own output no longer claims the wiring is done.
|
||
- Semantic retrieval seam (S3.1): a new MAF-free `semretrieval.py` adds an `Embedder`/`Retriever`
|
||
pair and a `HybridRanker` blending a numpy cosine term over the embedded feature triple (sorted
|
||
cost codes, measure type, magnitude bucket) with the existing structural score, exposed as
|
||
`--semantic-retrieval`. **What ships is the seam, not better retrieval quality**: the bundled
|
||
`FakeEmbedder` is a deterministic sha256 projection with no semantics, so over a structural tie
|
||
the order is deterministic but arbitrary. A real embedder is selected from a CLOSED registry via
|
||
`--embedder-config` / `build_embedder` — deliberately never an import path, so a config file can
|
||
never name arbitrary code to load.
|
||
**Off by default and additive**: with no retriever passed the store delegates to
|
||
`StructuralRetriever`, which reproduces the pre-seam ranking exactly, so the text-excluded
|
||
default and every existing test are unchanged. The embedding excludes `description`, matching
|
||
`similarity` ("text is ignored by design") and the verdict-id hash — so a flag-on run reads no
|
||
surface text either, and a genuine expert verdict can no longer be outranked by the framework's
|
||
own echo of the query. The ranker is passed PER CALL, never assigned to the caller's store, so
|
||
the opt-in cannot outlive the run that asked for it. Accepted in both run modes; in
|
||
single-project mode it requires `--bundle-dir` and `--verdict-dir` and is refused — never
|
||
silently ignored — without them.
|
||
Determinism, precisely: ranking order rests on the total order `(-round(score, 9), id)`; the
|
||
BLAS thread pins set before numpy is imported (`VECLIB_MAXIMUM_THREADS` for Accelerate,
|
||
`OPENBLAS`/`MKL`/`OMP` for other backends) defend the narrower claim that vector artifacts are
|
||
byte-identical across environments. Ships an optional, rebuildable `vectors.npy` +
|
||
`vectors.jsonl` store (byte-identical regardless of insertion order; fail-fast on a row/line
|
||
mismatch; missing loads as `None`) — an authoring primitive with no caller in `src/`, offered to
|
||
extenders like `write_verdict` and `promote_verdict`. numpy is confined to `semretrieval.py` and
|
||
never enters `okf.py`, `retrieval.py` or `shared/`.
|
||
- Cost-baseline anchoring of the deterministic gate (S4.0): `validate_proposal(..., baseline=...)`
|
||
reconciles every `affected_item` against the project's real `CostBaseline` in a **stage 0, before
|
||
the solver** — two independent rejections (unknown cost code; a real code whose
|
||
`quantity`/`unit_cost` falls outside a configurable tolerance, default 5 %, of the baseline value).
|
||
Before this, every stage reasoned only about numbers the proposal itself supplied, so an internally
|
||
consistent hallucination cleared the whole gate. Validation, never repair: the proposal is
|
||
rejected, never silently corrected to the baseline. The argument is **optional** (`None` reproduces
|
||
the earlier behaviour) but both run paths set it — the road path always, the bundle path only when
|
||
the bundle ships a `cost-baseline.json`, since a pre-amendment bundle is legitimately unanchored.
|
||
A baseline that exists but is malformed raises on both loaders rather than reading as "no
|
||
baseline". Method caps are looked up in an injectable `METHOD_CAPS` registry, so a second measure
|
||
type is data rather than an edit to the validator.
|
||
- Deterministic stage-2 bound and band enclosure in the validator (S2.7).
|
||
- Per-candidate verdict keying (S3.2): `seed_store_from_bundle` reads each `type: verdict` file's own
|
||
structural frontmatter (`affected_codes` / `measure_type` / `claimed_saving_nok`) instead of
|
||
collapsing a multi-candidate bundle onto the single IR projection, where a verdict about candidate
|
||
B scored a perfect structural match against candidate A's query. All three fields or none — a
|
||
partial declaration raises rather than being merged with the bundle candidate, which would mint a
|
||
key belonging to neither. Bundles that declare none fall back to the previous keying, so every
|
||
pre-S3.2 seed is unchanged.
|
||
- Portfolio concurrency (S3.3): wave-partitioned execution with a per-project snapshot and a
|
||
deterministic merge barrier, pinned by a concurrent-equals-sequential contract test; `k < 1` is
|
||
fail-fast. Failures are collected and the pass continues, surfaced as `RunFailure` slots rather
|
||
than aborting the portfolio. Goal-stop is evaluated at wave boundaries, with the intra-wave
|
||
semantics documented and pinned.
|
||
- Global portfolio token cap (S3.4/F10): `PortfolioBudget` + `PortfolioMeter` form one ledger across
|
||
a whole portfolio pass — and, seeded from a spend file, across passes — while the per-run
|
||
`Budget`/`TokenMeter` is unchanged. Three enforcement points: a **startup refusal** when the
|
||
remainder cannot fund a single run; **wave admission**, where an unfundable project is never
|
||
started (a project that is merely aborted has already cost calls) and the pass stops structurally
|
||
with the completed runs preserved; and a **pre-call guard** in `BudgetMiddleware`, so a call the
|
||
remainder cannot pay for is refused rather than made. Because a whole wave is checked against the
|
||
same pre-wave remainder, admission reserves each member's requirement. Resource exhaustion is
|
||
reported in its own `budget_stop` field, never merged into `stop_reason` — a goal stop is success.
|
||
The spend file is our own accounting state, so corrupt content raises instead of reading as zero.
|
||
- Ingest transports (S2.2, S2.4): the MCP connector as a transport inside the http family, and a
|
||
time-bounded default http transport in front of the pinned ingest library. A server on the ingest
|
||
path must expose a **zero-argument tool** (the URL carries both coordinates), which is a
|
||
deliberately separate seam from the in-run data-source demo server.
|
||
- Azure/Foundry offline preflight config gate (`preflight.py`, S4.1).
|
||
- Offline live-dry-run drill (`--live-dry-run`, S4.2): walks the whole path up to the eager
|
||
client build and stops before the first model call — zero chat calls.
|
||
- Out-of-band HITL verdict routing CLI (`hitl.py`, S5.1): code-prefix routing on the candidate id.
|
||
- Expert-notification contract (`notify.py`, S5.2): the declared `Notifier` with `ConsoleNotifier`,
|
||
`FileNotifier` (byte-deterministic JSONL), and `WebhookNotifier`, plus a fail-fast
|
||
`build_notifier`; the webhook is the only egress point, fail-closed behind an explicit per-run
|
||
`allow_egress` opt-in. Not auto-wired into `run.py`.
|
||
- Offline whole-loop CLI door (`--scripted-replies`): drives the entire method over the caller's own
|
||
data with no model budget, in **both** run modes. Resolved before the portfolio dispatch, because
|
||
placing it after made `--portfolio --scripted-replies` silently drop the flag and attempt four real
|
||
model calls (measured). Mutually exclusive with `--live-dry-run`: both are offline and they
|
||
contradict, so the combination is refused rather than letting one quietly win.
|
||
- Run mandate (`mandate.py`, `--mandate`): a domain expert commissions **which approaches** a run
|
||
evaluates, the run announces what it will do before doing it, and it settles for that — every
|
||
commissioned approach is evaluated and every one appears in a coverage report. A mandate is not a
|
||
dimension (`dimension.admits` filters what may pass the scoping gate; a mandate directs what the
|
||
run spends its attempts on) and not a goal (the numeric target keeps its single home in
|
||
`GoalContract`/`--goals` and is merely restated in the announcement). Two construction-time
|
||
refusals: an empty commission, and a duplicate approach id — including one claiming the reserved
|
||
`own-proposal` — since the id keys each coverage row and two rows under one key collapse silently.
|
||
Pure module: pydantic + stdlib only, D7-portable, guarded MAF-free.
|
||
- MCP servers as in-run tools (`mcp_tools.py`, `--mcp-config`): concrete external servers become
|
||
tools the agents can call **during** the debate. Opt-in and additive — with no config there are
|
||
zero network calls and the tool list is unchanged. Three rules are load-bearing: an **allowlist is
|
||
required** (an empty one would let the far end decide what the agents may call), **every server and
|
||
every permitted tool is named in the announcement before the first call** (also without
|
||
`--mandate` — no undeclared egress), and `--live-dry-run` opens nothing. A malformed config is
|
||
refused rather than degraded to "no external services", which would make the announcement describe
|
||
a run nobody configured. This is a separate seam from the ingest-path MCP transport above.
|
||
- Per-approach outbox artefacts: a run commissioned to evaluate three approaches previously wrote one
|
||
proposal artefact, so only the selected approach could ever receive a verdict and the others taught
|
||
the learning loop nothing. Artefacts are now keyed `{run_id}-{approach_id}-*.json` and carry
|
||
`approach_id` in the payload, with the join key `(run_id, approach_id)` — widening only the
|
||
filename would not have worked, because the reader joins on the `run_id` field read from file
|
||
content. Two properties make them genuinely judgeable: `verdict_id` is minted per approach (so one
|
||
delivered verdict cannot clear all three from the queue), and `provenance.validator_decision`
|
||
follows its own approach rather than the run's. `verdicts.verdict_key` is the single public
|
||
verdict-key rule.
|
||
- External-call provenance: the egress declaration says what a run **may** contact;
|
||
`ToolCallRecorder` records what it **did**, landing on `ProvenanceStamp.external_calls` read after
|
||
the debate. It only observes — `call_next` is always awaited, so a trace can never alter the run it
|
||
traces. Only configured tools are recorded; logging in-process functions would turn the record into
|
||
a false egress claim, and an empty list is therefore a positive statement that nothing outside the
|
||
process was contacted. Attribution comes from our own config because MAF reports the bare tool name
|
||
with no server prefix (measured, not assumed), so a name allowed by two servers is recorded
|
||
unattributed rather than credited to the first match. Stated honesty limit: this is the call and
|
||
its source — not evidence that the service's answer reached the proposal.
|
||
- `docs/knowledge-base-recipe.md` (S5.3, D-H item 1): the documented team process (technical +
|
||
domain expert) for building a knowledge base, with the honest 1–2 week expectation.
|
||
- Test suite: 755 passing tests (4 skips are live-provider-only). Every wired seam is covered by a
|
||
load-bearing test that goes red when the seam is detached, and each carries a recorded detach point
|
||
verified by mutation. That property is maintained by measurement, not by assumption, and this
|
||
release contains two recorded cases where it did not hold until it was checked: the S3.1
|
||
`--semantic-retrieval` flag was covered only below the CLI, so hardcoding it off at `main()` level
|
||
left the suite green; and four of the five `BudgetExceeded` raise sites could report any value in
|
||
the `observed` field without a single test noticing. Both are now pinned.
|
||
|
||
### Fixed
|
||
Defects found and closed before this first tag. Nothing here was ever released; they are recorded
|
||
because each one changed an invariant an adopter depends on.
|
||
|
||
- Money quantisation happens in one order, from one source: `ledger.to_ore` is the framework's single
|
||
NOK→øre conversion, applied **per amount**, after which integers are summed. Previously the goal
|
||
baseline summed floats and quantised once while the ledger summed per-candidate integers, and the
|
||
two orders met in exactly one place — the check that decides whether a portfolio pass stops early
|
||
on a percentage goal. Measured divergence: three lines of `60000.005` NOK are `18000003` øre
|
||
quantised first but `18000001` summed first, enough to flip a goal. Both call sites were fixed; a
|
||
fix to only one survived the whole suite (measured).
|
||
- The MCP stdio ingest transport now runs against a real server, and its error contract is repaired:
|
||
`stdio_client` and `ClientSession` are each a task group, and anyio wraps everything leaving one in
|
||
an exception group, so the module's own `IngestError` reached callers as an exception group and
|
||
never as the type the ingest door switches on. The owned error is unwrapped and re-raised; anything
|
||
else is re-raised untouched. No canned-tool test could have caught this — the defect only appears
|
||
once the code is actually run.
|
||
- The MCP timeout composes with anyio's own cancel scope (`anyio.fail_after` nested inside both task
|
||
groups) instead of `asyncio.wait_for` from outside, which cancelled the structure anyio owns and
|
||
produced a `BrokenResourceError` inside an exception group rather than a timeout. The translation
|
||
is narrowed by `CancelScope.cancelled_caught`, so only the scope that hit its own deadline earns
|
||
the timeout code — an unconditional `except TimeoutError` would mislabel any `TimeoutError`, since
|
||
the builtin is also `socket.timeout` and `asyncio.TimeoutError`. `anyio` is now a declared direct
|
||
dependency rather than a transitive one.
|
||
- Frontmatter scalars are unquoted by exactly one rule (`okf.unquote_scalar`).
|
||
- A non-finite embedding is refused rather than scored.
|
||
- `--embedder-config` is refused when `--semantic-retrieval` is absent, rather than silently dropped.
|
||
- The portfolio CLI no longer swallows its offline door or its failures.
|
||
- Constants cited in the live documentation are gated against the code, so a doc and the value it
|
||
quotes cannot drift apart unnoticed.
|
||
|
||
### Notes
|
||
- Licensed under the MIT License (see `LICENSE`).
|
||
- Scope boundary: this is a technical framework. The deployer owns DPIA, risk assessment and
|
||
processing purpose; the framework ships only the technical preconditions (local-only operation,
|
||
provenance, no silent egress).
|
||
- The `shared/` directory is a git subtree of `portfolio-optimiser-commons` and is pull-only.
|