STATE said the release text was already written and only needed dating. Measured:
CHANGELOG.md was last touched at 9e149c6, 53 commits back, so [Unreleased]
described the repo as of S3.1 — and claimed 512 passing tests where the suite
measures 755.
For a FIRST tag there is no predecessor, so the section is not "what changed" but
"what this version is". A description that stops at S3.1 does not under-report a
delta; it misrepresents the artefact being tagged, on a public mirror.
Audited the existing bullets for claims that had gone FALSE rather than merely
stale — the class a test-count catch is only a sample of. Four checked, all still
true at HEAD: numpy confined to semretrieval.py; value_report deliberately not
wired into costsim; the simulation module and knowledge-base recipe present; the
two run modes still two.
Added, grouped by seam rather than by commit: S4.0 cost-baseline anchoring, S2.7,
S3.2 per-candidate verdict keying, S3.3 wave concurrency, S3.4 global token cap,
S2.2/S2.4 ingest transports, --scripted-replies, --mandate, --mcp-config, the
per-approach outbox, and external-call provenance. New Fixed section for the
seven correctness fixes, with a lead-in saying plainly that none of them ever
shipped. Notes gained the deployer-owns-DPIA scope boundary and the pull-only
subtree rule.
The test-suite bullet now records TWO cases where "every seam is load-bearing"
did not hold until measured, not one: the S3.1 CLI-level gap, and four of five
BudgetExceeded raise sites free to report any `observed`.
No code change. Suite measured green at 755 passed / 4 skipped after the edit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HVekyKcvT4ah6jnCT8R8fe
198 lines
16 KiB
Markdown
198 lines
16 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to this project will be documented in this file.
|
||
|
||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||
|
||
## [0.1.0] - 2026-08-06
|
||
|
||
First tagged release. There is no prior release, so the entries below describe what this version
|
||
**is**, not what changed since a predecessor.
|
||
|
||
### Added
|
||
- Deterministic backbone: mandatory blocking validator (solver + Monte Carlo against the shared
|
||
golden suite), budget meter with hard fail-fast caps, provenance stamping.
|
||
- Agentic learning loop wired end to end, one load-bearing seam at a time (target-picture steps
|
||
1, 3/4, 5, 7, 8): OKF-navigated bundle context with the gated ExpeL fold, maker-checker debate
|
||
where the checker gates the reasoning, informed refinement (previous rejection reason fed into
|
||
the next bounded attempt), async verdict file inbox, and gated wiki promotion (fail-closed).
|
||
- Offline end-to-end simulation proving the learning loop closes with a scripted client
|
||
(`uv run python -m portfolio_optimiser.simulation`) — plumbing proof, not live-model proof.
|
||
- Framework-neutral shared core in `shared/`: OKF concept + example bundle, golden validator
|
||
suite, and the expert-reviewer persona as an Agent Skill.
|
||
- CLI parity (S5.3): `run.py` drives the whole method from the command line in **two modes** —
|
||
single-project (`--dimension-config`, `--outbox-dir`/`--run-id`, plus the already-wired
|
||
`--bundle-dir`/`--verdict-dir`) and portfolio (`--portfolio` with `--goals`/`--ledger`), the
|
||
latter printing an observable `goal reached: …` line when a savings goal is met. Adds a
|
||
fail-fast `load_dimension` loader and structured refusals (rc 1, no traceback) for misuse and
|
||
mode-exclusivity violations. The prior-verdict fold is on the `--bundle-dir` path only; a
|
||
`--docs-dir`-only run is single-shot.
|
||
- Value report (S5.4): a read-only `--report [--json] --ledger <file>` surface on `run.py` that
|
||
rolls up the accumulated `SavingsLedger` — per-project + portfolio totals (dimension-free-deduped
|
||
integer øre), flagged cross-dimension overlaps (each counted once), and per-entry provenance — as
|
||
a human table or deterministic JSON. Makes no model calls; mode-exclusive (only `--ledger`/`--json`
|
||
permitted with `--report`, which requires `--ledger`). Honest scope boundary: the report core
|
||
(`value_report.py`) now **exists** but is deliberately **not** wired into `costsim`'s
|
||
`kost_mot_verdi` placeholder — that cost-vs-value integration is a separate, deferred step. The
|
||
`costsim` seam note was reworded from the stale "fylles av S5.4 verdirapport" to a truthful
|
||
forward reference so `costsim`'s own output no longer claims the wiring is done.
|
||
- Semantic retrieval seam (S3.1): a new MAF-free `semretrieval.py` adds an `Embedder`/`Retriever`
|
||
pair and a `HybridRanker` blending a numpy cosine term over the embedded feature triple (sorted
|
||
cost codes, measure type, magnitude bucket) with the existing structural score, exposed as
|
||
`--semantic-retrieval`. **What ships is the seam, not better retrieval quality**: the bundled
|
||
`FakeEmbedder` is a deterministic sha256 projection with no semantics, so over a structural tie
|
||
the order is deterministic but arbitrary. A real embedder is selected from a CLOSED registry via
|
||
`--embedder-config` / `build_embedder` — deliberately never an import path, so a config file can
|
||
never name arbitrary code to load.
|
||
**Off by default and additive**: with no retriever passed the store delegates to
|
||
`StructuralRetriever`, which reproduces the pre-seam ranking exactly, so the text-excluded
|
||
default and every existing test are unchanged. The embedding excludes `description`, matching
|
||
`similarity` ("text is ignored by design") and the verdict-id hash — so a flag-on run reads no
|
||
surface text either, and a genuine expert verdict can no longer be outranked by the framework's
|
||
own echo of the query. The ranker is passed PER CALL, never assigned to the caller's store, so
|
||
the opt-in cannot outlive the run that asked for it. Accepted in both run modes; in
|
||
single-project mode it requires `--bundle-dir` and `--verdict-dir` and is refused — never
|
||
silently ignored — without them.
|
||
Determinism, precisely: ranking order rests on the total order `(-round(score, 9), id)`; the
|
||
BLAS thread pins set before numpy is imported (`VECLIB_MAXIMUM_THREADS` for Accelerate,
|
||
`OPENBLAS`/`MKL`/`OMP` for other backends) defend the narrower claim that vector artifacts are
|
||
byte-identical across environments. Ships an optional, rebuildable `vectors.npy` +
|
||
`vectors.jsonl` store (byte-identical regardless of insertion order; fail-fast on a row/line
|
||
mismatch; missing loads as `None`) — an authoring primitive with no caller in `src/`, offered to
|
||
extenders like `write_verdict` and `promote_verdict`. numpy is confined to `semretrieval.py` and
|
||
never enters `okf.py`, `retrieval.py` or `shared/`.
|
||
- Cost-baseline anchoring of the deterministic gate (S4.0): `validate_proposal(..., baseline=...)`
|
||
reconciles every `affected_item` against the project's real `CostBaseline` in a **stage 0, before
|
||
the solver** — two independent rejections (unknown cost code; a real code whose
|
||
`quantity`/`unit_cost` falls outside a configurable tolerance, default 5 %, of the baseline value).
|
||
Before this, every stage reasoned only about numbers the proposal itself supplied, so an internally
|
||
consistent hallucination cleared the whole gate. Validation, never repair: the proposal is
|
||
rejected, never silently corrected to the baseline. The argument is **optional** (`None` reproduces
|
||
the earlier behaviour) but both run paths set it — the road path always, the bundle path only when
|
||
the bundle ships a `cost-baseline.json`, since a pre-amendment bundle is legitimately unanchored.
|
||
A baseline that exists but is malformed raises on both loaders rather than reading as "no
|
||
baseline". Method caps are looked up in an injectable `METHOD_CAPS` registry, so a second measure
|
||
type is data rather than an edit to the validator.
|
||
- Deterministic stage-2 bound and band enclosure in the validator (S2.7).
|
||
- Per-candidate verdict keying (S3.2): `seed_store_from_bundle` reads each `type: verdict` file's own
|
||
structural frontmatter (`affected_codes` / `measure_type` / `claimed_saving_nok`) instead of
|
||
collapsing a multi-candidate bundle onto the single IR projection, where a verdict about candidate
|
||
B scored a perfect structural match against candidate A's query. All three fields or none — a
|
||
partial declaration raises rather than being merged with the bundle candidate, which would mint a
|
||
key belonging to neither. Bundles that declare none fall back to the previous keying, so every
|
||
pre-S3.2 seed is unchanged.
|
||
- Portfolio concurrency (S3.3): wave-partitioned execution with a per-project snapshot and a
|
||
deterministic merge barrier, pinned by a concurrent-equals-sequential contract test; `k < 1` is
|
||
fail-fast. Failures are collected and the pass continues, surfaced as `RunFailure` slots rather
|
||
than aborting the portfolio. Goal-stop is evaluated at wave boundaries, with the intra-wave
|
||
semantics documented and pinned.
|
||
- Global portfolio token cap (S3.4/F10): `PortfolioBudget` + `PortfolioMeter` form one ledger across
|
||
a whole portfolio pass — and, seeded from a spend file, across passes — while the per-run
|
||
`Budget`/`TokenMeter` is unchanged. Three enforcement points: a **startup refusal** when the
|
||
remainder cannot fund a single run; **wave admission**, where an unfundable project is never
|
||
started (a project that is merely aborted has already cost calls) and the pass stops structurally
|
||
with the completed runs preserved; and a **pre-call guard** in `BudgetMiddleware`, so a call the
|
||
remainder cannot pay for is refused rather than made. Because a whole wave is checked against the
|
||
same pre-wave remainder, admission reserves each member's requirement. Resource exhaustion is
|
||
reported in its own `budget_stop` field, never merged into `stop_reason` — a goal stop is success.
|
||
The spend file is our own accounting state, so corrupt content raises instead of reading as zero.
|
||
- Ingest transports (S2.2, S2.4): the MCP connector as a transport inside the http family, and a
|
||
time-bounded default http transport in front of the pinned ingest library. A server on the ingest
|
||
path must expose a **zero-argument tool** (the URL carries both coordinates), which is a
|
||
deliberately separate seam from the in-run data-source demo server.
|
||
- Azure/Foundry offline preflight config gate (`preflight.py`, S4.1).
|
||
- Offline live-dry-run drill (`--live-dry-run`, S4.2): walks the whole path up to the eager
|
||
client build and stops before the first model call — zero chat calls.
|
||
- Out-of-band HITL verdict routing CLI (`hitl.py`, S5.1): code-prefix routing on the candidate id.
|
||
- Expert-notification contract (`notify.py`, S5.2): the declared `Notifier` with `ConsoleNotifier`,
|
||
`FileNotifier` (byte-deterministic JSONL), and `WebhookNotifier`, plus a fail-fast
|
||
`build_notifier`; the webhook is the only egress point, fail-closed behind an explicit per-run
|
||
`allow_egress` opt-in. Not auto-wired into `run.py`.
|
||
- Offline whole-loop CLI door (`--scripted-replies`): drives the entire method over the caller's own
|
||
data with no model budget, in **both** run modes. Resolved before the portfolio dispatch, because
|
||
placing it after made `--portfolio --scripted-replies` silently drop the flag and attempt four real
|
||
model calls (measured). Mutually exclusive with `--live-dry-run`: both are offline and they
|
||
contradict, so the combination is refused rather than letting one quietly win.
|
||
- Run mandate (`mandate.py`, `--mandate`): a domain expert commissions **which approaches** a run
|
||
evaluates, the run announces what it will do before doing it, and it settles for that — every
|
||
commissioned approach is evaluated and every one appears in a coverage report. A mandate is not a
|
||
dimension (`dimension.admits` filters what may pass the scoping gate; a mandate directs what the
|
||
run spends its attempts on) and not a goal (the numeric target keeps its single home in
|
||
`GoalContract`/`--goals` and is merely restated in the announcement). Two construction-time
|
||
refusals: an empty commission, and a duplicate approach id — including one claiming the reserved
|
||
`own-proposal` — since the id keys each coverage row and two rows under one key collapse silently.
|
||
Pure module: pydantic + stdlib only, D7-portable, guarded MAF-free.
|
||
- MCP servers as in-run tools (`mcp_tools.py`, `--mcp-config`): concrete external servers become
|
||
tools the agents can call **during** the debate. Opt-in and additive — with no config there are
|
||
zero network calls and the tool list is unchanged. Three rules are load-bearing: an **allowlist is
|
||
required** (an empty one would let the far end decide what the agents may call), **every server and
|
||
every permitted tool is named in the announcement before the first call** (also without
|
||
`--mandate` — no undeclared egress), and `--live-dry-run` opens nothing. A malformed config is
|
||
refused rather than degraded to "no external services", which would make the announcement describe
|
||
a run nobody configured. This is a separate seam from the ingest-path MCP transport above.
|
||
- Per-approach outbox artefacts: a run commissioned to evaluate three approaches previously wrote one
|
||
proposal artefact, so only the selected approach could ever receive a verdict and the others taught
|
||
the learning loop nothing. Artefacts are now keyed `{run_id}-{approach_id}-*.json` and carry
|
||
`approach_id` in the payload, with the join key `(run_id, approach_id)` — widening only the
|
||
filename would not have worked, because the reader joins on the `run_id` field read from file
|
||
content. Two properties make them genuinely judgeable: `verdict_id` is minted per approach (so one
|
||
delivered verdict cannot clear all three from the queue), and `provenance.validator_decision`
|
||
follows its own approach rather than the run's. `verdicts.verdict_key` is the single public
|
||
verdict-key rule.
|
||
- External-call provenance: the egress declaration says what a run **may** contact;
|
||
`ToolCallRecorder` records what it **did**, landing on `ProvenanceStamp.external_calls` read after
|
||
the debate. It only observes — `call_next` is always awaited, so a trace can never alter the run it
|
||
traces. Only configured tools are recorded; logging in-process functions would turn the record into
|
||
a false egress claim, and an empty list is therefore a positive statement that nothing outside the
|
||
process was contacted. Attribution comes from our own config because MAF reports the bare tool name
|
||
with no server prefix (measured, not assumed), so a name allowed by two servers is recorded
|
||
unattributed rather than credited to the first match. Stated honesty limit: this is the call and
|
||
its source — not evidence that the service's answer reached the proposal.
|
||
- `docs/knowledge-base-recipe.md` (S5.3, D-H item 1): the documented team process (technical +
|
||
domain expert) for building a knowledge base, with the honest 1–2 week expectation.
|
||
- Test suite: 755 passing tests (4 skips are live-provider-only). Every wired seam is covered by a
|
||
load-bearing test that goes red when the seam is detached, and each carries a recorded detach point
|
||
verified by mutation. That property is maintained by measurement, not by assumption, and this
|
||
release contains two recorded cases where it did not hold until it was checked: the S3.1
|
||
`--semantic-retrieval` flag was covered only below the CLI, so hardcoding it off at `main()` level
|
||
left the suite green; and four of the five `BudgetExceeded` raise sites could report any value in
|
||
the `observed` field without a single test noticing. Both are now pinned.
|
||
|
||
### Fixed
|
||
Defects found and closed before this first tag. Nothing here was ever released; they are recorded
|
||
because each one changed an invariant an adopter depends on.
|
||
|
||
- Money quantisation happens in one order, from one source: `ledger.to_ore` is the framework's single
|
||
NOK→øre conversion, applied **per amount**, after which integers are summed. Previously the goal
|
||
baseline summed floats and quantised once while the ledger summed per-candidate integers, and the
|
||
two orders met in exactly one place — the check that decides whether a portfolio pass stops early
|
||
on a percentage goal. Measured divergence: three lines of `60000.005` NOK are `18000003` øre
|
||
quantised first but `18000001` summed first, enough to flip a goal. Both call sites were fixed; a
|
||
fix to only one survived the whole suite (measured).
|
||
- The MCP stdio ingest transport now runs against a real server, and its error contract is repaired:
|
||
`stdio_client` and `ClientSession` are each a task group, and anyio wraps everything leaving one in
|
||
an exception group, so the module's own `IngestError` reached callers as an exception group and
|
||
never as the type the ingest door switches on. The owned error is unwrapped and re-raised; anything
|
||
else is re-raised untouched. No canned-tool test could have caught this — the defect only appears
|
||
once the code is actually run.
|
||
- The MCP timeout composes with anyio's own cancel scope (`anyio.fail_after` nested inside both task
|
||
groups) instead of `asyncio.wait_for` from outside, which cancelled the structure anyio owns and
|
||
produced a `BrokenResourceError` inside an exception group rather than a timeout. The translation
|
||
is narrowed by `CancelScope.cancelled_caught`, so only the scope that hit its own deadline earns
|
||
the timeout code — an unconditional `except TimeoutError` would mislabel any `TimeoutError`, since
|
||
the builtin is also `socket.timeout` and `asyncio.TimeoutError`. `anyio` is now a declared direct
|
||
dependency rather than a transitive one.
|
||
- Frontmatter scalars are unquoted by exactly one rule (`okf.unquote_scalar`).
|
||
- A non-finite embedding is refused rather than scored.
|
||
- `--embedder-config` is refused when `--semantic-retrieval` is absent, rather than silently dropped.
|
||
- The portfolio CLI no longer swallows its offline door or its failures.
|
||
- Constants cited in the live documentation are gated against the code, so a doc and the value it
|
||
quotes cannot drift apart unnoticed.
|
||
|
||
### Notes
|
||
- Licensed under the MIT License (see `LICENSE`).
|
||
- Scope boundary: this is a technical framework. The deployer owns DPIA, risk assessment and
|
||
processing purpose; the framework ships only the technical preconditions (local-only operation,
|
||
provenance, no silent egress).
|
||
- The `shared/` directory is a git subtree of `portfolio-optimiser-commons` and is pull-only.
|