Datoen er LEST med `git log -1 --format=%cs` på Y, ikke skrevet på forhånd og ikke tatt fra veggklokka. Re-leses på denne commiten før taggen settes (runbookens §5 punkt 8): faller midnatt mellom Y og Z, står gårsdagens dato i commiten som faktisk tagges. Denne commiten er Z. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DWeYxduEnFQeynXbrtEA6o
293 lines
24 KiB
Markdown
293 lines
24 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to this project will be documented in this file.
|
||
|
||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||
|
||
## [1.0.0] - 2026-08-12
|
||
|
||
### Added
|
||
- Step 5 is now observable: `generate_via_llm` returns a `GenerationResult` carrying the validator
|
||
falsifications that informed a later attempt, surfaced on `RunResult.refinements`. The offline
|
||
simulation exercises it — the scripted proposer overclaims, the deterministic validator falsifies
|
||
the number, and the refined proposal validates.
|
||
- Two console entry points ship with `uv sync`: `portfolio-optimiser` (the CLI) and
|
||
`portfolio-optimiser-demo` (the offline walkthrough). Deliberately two of the package's five
|
||
`main()` functions — `costsim`, `hitl` and `preflight` stay operator tools invoked as modules, and
|
||
every name here is one a release has to carry. Both invocation forms write byte-identical stdout.
|
||
The entry points are tested against the INSTALLED distribution's metadata rather than the TOML: a
|
||
`[project.scripts]` line that has never been synced is a claim, not a command.
|
||
- The offline walkthrough's transcript is checked in as a golden fixture
|
||
(`tests/golden/demo-transcript.stdout` and `.stderr`). Self-identity across two runs cannot detect
|
||
a regression — two runs of a regressed walkthrough agree exactly as well as two runs of a correct
|
||
one — so the fixture leaves the process. stdout is pinned verbatim; stderr is normalised on exactly
|
||
two measured environment spans, the `site-packages` prefix and the temporary directory, leaving the
|
||
`po-sim-` prefix visible because that belongs to the program rather than the environment. A
|
||
companion control forbids the mask from widening: a normaliser that dropped whole lines, with the
|
||
fixture regenerated beneath it, would keep both equality tests green.
|
||
|
||
### Changed
|
||
- **Breaking (library API):** `generate_via_llm` returns `GenerationResult` instead of
|
||
`ValidatedProposal | Rejection`; read `.outcome` for the previous value. The refinement loop's
|
||
bound is unchanged (`max_attempts` + token meter).
|
||
- `simulation.scripted_factory` accepts a per-role reply *selector* over `(prompt, role)` as well as
|
||
a constant reply, so a scripted role can answer differently on a later attempt.
|
||
- The offline simulation's scripted proposer is now a candidate **registry** rather than a
|
||
hand-written reply: `simulation.scripted_proposer(candidates)` builds the selector from
|
||
`ScriptedCandidate` entries keyed on the project id the prompt names, and
|
||
`simulate_learning_loop` takes `project_id` alongside `bundle_dir`. Adding a project to the
|
||
walkthrough is a data entry. A prompt matching no entry — or more than one — raises
|
||
`ScriptedCandidateError` rather than answering with another project's numbers.
|
||
- The offline simulation now EXERCISES the Step-7 file inbox it narrates. The verdict previously
|
||
arrived as a function argument — the short, in-run capture — while the trace line described the
|
||
long file loop. An expert now writes a real verdict file into an inbox between the two runs, and
|
||
the second run is given `verdict_dir=`, so `run_project` merges it before the Step-1 fold. The
|
||
inbox sits beside the bundle copy and never inside it: a verdict file within the bundle would reach
|
||
the next run as navigable context, which is a different mechanism wearing this one's clothes. The
|
||
two time-scales carry SEPARATE markers by construction, since one marker on both paths would let
|
||
either seam alone satisfy the assertion and leave the other free to rot; `simulate_learning_loop`
|
||
refuses equal markers.
|
||
- The offline walkthrough runs ANCHORED. Its deterministic gate reconciles each proposal against the
|
||
project's real cost lines, which activates only when the knowledge base ships a `cost-baseline.json`;
|
||
without one the gate reasoned solely about numbers the proposal supplied itself. The walkthrough now
|
||
reads that file through exactly the seam a delivered knowledge base would use. For the synthetic
|
||
fallback bundle, which cannot receive the file inside the pull-only `shared/` subtree, the baseline
|
||
is DERIVED IN CODE from the scripted register rather than typed beside it — two sources of the same
|
||
numbers drift, and drift is precisely what the walkthrough's own 10 % probe models. The declared
|
||
baseline is printed, because an anchoring nobody can see is one nobody can check.
|
||
- The walkthrough's stderr is quieter. The expected round-cap notice is dropped by a filter on the
|
||
emitting logger, keyed on the message and installed by `main()` — never at import, so a library
|
||
consumer keeps its own logging configuration. The two `ExperimentalWarning` lines are deliberately
|
||
NOT damped: they fire while the package `__init__` imports the agent framework, always before the
|
||
simulation's own imports and under both invocation forms, so silencing them would mean filtering
|
||
warnings inside the library on every consumer's behalf. They are pinned in the golden fixture
|
||
instead. stderr went from six lines to four.
|
||
- The walkthrough DERIVES its provenance sentence for prior verdicts instead of stating it. The line
|
||
above already computes the count, so a hand-written split would be a second copy of the same fact,
|
||
free to drift the moment a knowledge base ships another seeded verdict.
|
||
- The shared expert-reviewer persona's canonical example verdict is worded domain-neutrally
|
||
("i tilsvarende anlegg" rather than "i kontorbygg"), pulled from the upstream commons repository.
|
||
The walkthrough prints that `rationale` verbatim, so the wording was a building-type justification
|
||
read out over a road-lighting project; it could not be fixed downstream, because overriding the
|
||
text locally would re-stub the very artifact the shared skill exists to make load-bearing. The
|
||
`marker` value is byte-unchanged, and the pinned transcript fixture was regenerated against a
|
||
prediction written before the pull — the printed line is clipped at a fixed width, so the swap
|
||
moves the tail as well, and a regeneration without a written prediction could not tell that
|
||
expected shift apart from drift.
|
||
|
||
### Security
|
||
- Door A — the ingest path that materialises externally sourced documents into a knowledge base — can
|
||
now scan generated content before it is published, through `ingest.materialize_gated`. The gate is
|
||
**opt-in and requested by name**: `materialize` itself stays ungated by design, because golden
|
||
suites pin its bytes and a caller that wants the gate asks for it.
|
||
The seam sits *around* materialisation rather than inside it. The pinned upstream stages in memory
|
||
and then performs its own disk phase, with no callback between the two, so a gate placed "at the
|
||
write point" could only have run after the bytes had landed — a cleanup, not a gate. Instead the
|
||
bundle is COPIED, materialised into the copy, scanned, and then published or discarded as a whole.
|
||
The copy is load-bearing rather than convenient: the upstream's ownership scan, its collision gate
|
||
against curated content, and its index merge all read the EXISTING bundle, so staging into an empty
|
||
directory would publish a bundle stripped of its curated neighbours and their index links — data
|
||
loss dressed as a security fix.
|
||
Trust follows ORIGIN, never channel. The outcome is per BUNDLE, since partial publication would
|
||
leave a bundle and index answering to no manifest, while diagnostics are per DOCUMENT so a single
|
||
run reports every finding rather than only the first. Findings are written to the bundle's `log.md`
|
||
and never to concept frontmatter, where four golden suites pin the bytes.
|
||
|
||
### Notes
|
||
- The `1.0.0` version signals a stable public surface, not a finished research programme. Two
|
||
boundaries are open and named rather than implied: the ingest stamp predicate has diverged from the
|
||
upstream specification (a value literal here, a structured field upstream) and does not touch the
|
||
run path, and the mirroring of several seams to the sibling implementation is outstanding.
|
||
|
||
## [0.1.0] - 2026-08-06
|
||
|
||
First tagged release. There is no prior release, so the entries below describe what this version
|
||
**is**, not what changed since a predecessor.
|
||
|
||
### Added
|
||
- Deterministic backbone: mandatory blocking validator (solver + Monte Carlo against the shared
|
||
golden suite), budget meter with hard fail-fast caps, provenance stamping.
|
||
- Agentic learning loop wired end to end, one load-bearing seam at a time (target-picture steps
|
||
1, 3/4, 5, 7, 8): OKF-navigated bundle context with the gated ExpeL fold, maker-checker debate
|
||
where the checker gates the reasoning, informed refinement (previous rejection reason fed into
|
||
the next bounded attempt), async verdict file inbox, and gated wiki promotion (fail-closed).
|
||
- Offline end-to-end simulation proving the learning loop closes with a scripted client
|
||
(`uv run python -m portfolio_optimiser.simulation`) — plumbing proof, not live-model proof.
|
||
- Framework-neutral shared core in `shared/`: OKF concept + example bundle, golden validator
|
||
suite, and the expert-reviewer persona as an Agent Skill.
|
||
- CLI parity (S5.3): `run.py` drives the whole method from the command line in **two modes** —
|
||
single-project (`--dimension-config`, `--outbox-dir`/`--run-id`, plus the already-wired
|
||
`--bundle-dir`/`--verdict-dir`) and portfolio (`--portfolio` with `--goals`/`--ledger`), the
|
||
latter printing an observable `goal reached: …` line when a savings goal is met. Adds a
|
||
fail-fast `load_dimension` loader and structured refusals (rc 1, no traceback) for misuse and
|
||
mode-exclusivity violations. The prior-verdict fold is on the `--bundle-dir` path only; a
|
||
`--docs-dir`-only run is single-shot.
|
||
- Value report (S5.4): a read-only `--report [--json] --ledger <file>` surface on `run.py` that
|
||
rolls up the accumulated `SavingsLedger` — per-project + portfolio totals (dimension-free-deduped
|
||
integer øre), flagged cross-dimension overlaps (each counted once), and per-entry provenance — as
|
||
a human table or deterministic JSON. Makes no model calls; mode-exclusive (only `--ledger`/`--json`
|
||
permitted with `--report`, which requires `--ledger`). Honest scope boundary: the report core
|
||
(`value_report.py`) now **exists** but is deliberately **not** wired into `costsim`'s
|
||
`kost_mot_verdi` placeholder — that cost-vs-value integration is a separate, deferred step. The
|
||
`costsim` seam note was reworded from the stale "fylles av S5.4 verdirapport" to a truthful
|
||
forward reference so `costsim`'s own output no longer claims the wiring is done.
|
||
- Semantic retrieval seam (S3.1): a new MAF-free `semretrieval.py` adds an `Embedder`/`Retriever`
|
||
pair and a `HybridRanker` blending a numpy cosine term over the embedded feature triple (sorted
|
||
cost codes, measure type, magnitude bucket) with the existing structural score, exposed as
|
||
`--semantic-retrieval`. **What ships is the seam, not better retrieval quality**: the bundled
|
||
`FakeEmbedder` is a deterministic sha256 projection with no semantics, so over a structural tie
|
||
the order is deterministic but arbitrary. A real embedder is selected from a CLOSED registry via
|
||
`--embedder-config` / `build_embedder` — deliberately never an import path, so a config file can
|
||
never name arbitrary code to load.
|
||
**Off by default and additive**: with no retriever passed the store delegates to
|
||
`StructuralRetriever`, which reproduces the pre-seam ranking exactly, so the text-excluded
|
||
default and every existing test are unchanged. The embedding excludes `description`, matching
|
||
`similarity` ("text is ignored by design") and the verdict-id hash — so a flag-on run reads no
|
||
surface text either, and a genuine expert verdict can no longer be outranked by the framework's
|
||
own echo of the query. The ranker is passed PER CALL, never assigned to the caller's store, so
|
||
the opt-in cannot outlive the run that asked for it. Accepted in both run modes; in
|
||
single-project mode it requires `--bundle-dir` and `--verdict-dir` and is refused — never
|
||
silently ignored — without them.
|
||
Determinism, precisely: ranking order rests on the total order `(-round(score, 9), id)`; the
|
||
BLAS thread pins set before numpy is imported (`VECLIB_MAXIMUM_THREADS` for Accelerate,
|
||
`OPENBLAS`/`MKL`/`OMP` for other backends) defend the narrower claim that vector artifacts are
|
||
byte-identical across environments. Ships an optional, rebuildable `vectors.npy` +
|
||
`vectors.jsonl` store (byte-identical regardless of insertion order; fail-fast on a row/line
|
||
mismatch; missing loads as `None`) — an authoring primitive with no caller in `src/`, offered to
|
||
extenders like `write_verdict` and `promote_verdict`. numpy is confined to `semretrieval.py` and
|
||
never enters `okf.py`, `retrieval.py` or `shared/`.
|
||
- Cost-baseline anchoring of the deterministic gate (S4.0): `validate_proposal(..., baseline=...)`
|
||
reconciles every `affected_item` against the project's real `CostBaseline` in a **stage 0, before
|
||
the solver** — two independent rejections (unknown cost code; a real code whose
|
||
`quantity`/`unit_cost` falls outside a configurable tolerance, default 5 %, of the baseline value).
|
||
Before this, every stage reasoned only about numbers the proposal itself supplied, so an internally
|
||
consistent hallucination cleared the whole gate. Validation, never repair: the proposal is
|
||
rejected, never silently corrected to the baseline. The argument is **optional** (`None` reproduces
|
||
the earlier behaviour) but both run paths set it — the road path always, the bundle path only when
|
||
the bundle ships a `cost-baseline.json`, since a pre-amendment bundle is legitimately unanchored.
|
||
A baseline that exists but is malformed raises on both loaders rather than reading as "no
|
||
baseline". Method caps are looked up in an injectable `METHOD_CAPS` registry, so a second measure
|
||
type is data rather than an edit to the validator.
|
||
- Deterministic stage-2 bound and band enclosure in the validator (S2.7).
|
||
- Per-candidate verdict keying (S3.2): `seed_store_from_bundle` reads each `type: verdict` file's own
|
||
structural frontmatter (`affected_codes` / `measure_type` / `claimed_saving_nok`) instead of
|
||
collapsing a multi-candidate bundle onto the single IR projection, where a verdict about candidate
|
||
B scored a perfect structural match against candidate A's query. All three fields or none — a
|
||
partial declaration raises rather than being merged with the bundle candidate, which would mint a
|
||
key belonging to neither. Bundles that declare none fall back to the previous keying, so every
|
||
pre-S3.2 seed is unchanged.
|
||
- Portfolio concurrency (S3.3): wave-partitioned execution with a per-project snapshot and a
|
||
deterministic merge barrier, pinned by a concurrent-equals-sequential contract test; `k < 1` is
|
||
fail-fast. Failures are collected and the pass continues, surfaced as `RunFailure` slots rather
|
||
than aborting the portfolio. Goal-stop is evaluated at wave boundaries, with the intra-wave
|
||
semantics documented and pinned.
|
||
- Global portfolio token cap (S3.4/F10): `PortfolioBudget` + `PortfolioMeter` form one ledger across
|
||
a whole portfolio pass — and, seeded from a spend file, across passes — while the per-run
|
||
`Budget`/`TokenMeter` is unchanged. Three enforcement points: a **startup refusal** when the
|
||
remainder cannot fund a single run; **wave admission**, where an unfundable project is never
|
||
started (a project that is merely aborted has already cost calls) and the pass stops structurally
|
||
with the completed runs preserved; and a **pre-call guard** in `BudgetMiddleware`, so a call the
|
||
remainder cannot pay for is refused rather than made. Because a whole wave is checked against the
|
||
same pre-wave remainder, admission reserves each member's requirement. Resource exhaustion is
|
||
reported in its own `budget_stop` field, never merged into `stop_reason` — a goal stop is success.
|
||
The spend file is our own accounting state, so corrupt content raises instead of reading as zero.
|
||
- Ingest transports (S2.2, S2.4): the MCP connector as a transport inside the http family, and a
|
||
time-bounded default http transport in front of the pinned ingest library. A server on the ingest
|
||
path must expose a **zero-argument tool** (the URL carries both coordinates), which is a
|
||
deliberately separate seam from the in-run data-source demo server.
|
||
- Azure/Foundry offline preflight config gate (`preflight.py`, S4.1).
|
||
- Offline live-dry-run drill (`--live-dry-run`, S4.2): walks the whole path up to the eager
|
||
client build and stops before the first model call — zero chat calls.
|
||
- Out-of-band HITL verdict routing CLI (`hitl.py`, S5.1): code-prefix routing on the candidate id.
|
||
- Expert-notification contract (`notify.py`, S5.2): the declared `Notifier` with `ConsoleNotifier`,
|
||
`FileNotifier` (byte-deterministic JSONL), and `WebhookNotifier`, plus a fail-fast
|
||
`build_notifier`; the webhook is the only egress point, fail-closed behind an explicit per-run
|
||
`allow_egress` opt-in. Not auto-wired into `run.py`.
|
||
- Offline whole-loop CLI door (`--scripted-replies`): drives the entire method over the caller's own
|
||
data with no model budget, in **both** run modes. Resolved before the portfolio dispatch, because
|
||
placing it after made `--portfolio --scripted-replies` silently drop the flag and attempt four real
|
||
model calls (measured). Mutually exclusive with `--live-dry-run`: both are offline and they
|
||
contradict, so the combination is refused rather than letting one quietly win.
|
||
- Run mandate (`mandate.py`, `--mandate`): a domain expert commissions **which approaches** a run
|
||
evaluates, the run announces what it will do before doing it, and it settles for that — every
|
||
commissioned approach is evaluated and every one appears in a coverage report. A mandate is not a
|
||
dimension (`dimension.admits` filters what may pass the scoping gate; a mandate directs what the
|
||
run spends its attempts on) and not a goal (the numeric target keeps its single home in
|
||
`GoalContract`/`--goals` and is merely restated in the announcement). Two construction-time
|
||
refusals: an empty commission, and a duplicate approach id — including one claiming the reserved
|
||
`own-proposal` — since the id keys each coverage row and two rows under one key collapse silently.
|
||
Pure module: pydantic + stdlib only, D7-portable, guarded MAF-free.
|
||
- MCP servers as in-run tools (`mcp_tools.py`, `--mcp-config`): concrete external servers become
|
||
tools the agents can call **during** the debate. Opt-in and additive — with no config there are
|
||
zero network calls and the tool list is unchanged. Three rules are load-bearing: an **allowlist is
|
||
required** (an empty one would let the far end decide what the agents may call), **every server and
|
||
every permitted tool is named in the announcement before the first call** (also without
|
||
`--mandate` — no undeclared egress), and `--live-dry-run` opens nothing. A malformed config is
|
||
refused rather than degraded to "no external services", which would make the announcement describe
|
||
a run nobody configured. This is a separate seam from the ingest-path MCP transport above.
|
||
- Per-approach outbox artefacts: a run commissioned to evaluate three approaches previously wrote one
|
||
proposal artefact, so only the selected approach could ever receive a verdict and the others taught
|
||
the learning loop nothing. Artefacts are now keyed `{run_id}-{approach_id}-*.json` and carry
|
||
`approach_id` in the payload, with the join key `(run_id, approach_id)` — widening only the
|
||
filename would not have worked, because the reader joins on the `run_id` field read from file
|
||
content. Two properties make them genuinely judgeable: `verdict_id` is minted per approach (so one
|
||
delivered verdict cannot clear all three from the queue), and `provenance.validator_decision`
|
||
follows its own approach rather than the run's. `verdicts.verdict_key` is the single public
|
||
verdict-key rule.
|
||
- External-call provenance: the egress declaration says what a run **may** contact;
|
||
`ToolCallRecorder` records what it **did**, landing on `ProvenanceStamp.external_calls` read after
|
||
the debate. It only observes — `call_next` is always awaited, so a trace can never alter the run it
|
||
traces. Only configured tools are recorded; logging in-process functions would turn the record into
|
||
a false egress claim, and an empty list is therefore a positive statement that nothing outside the
|
||
process was contacted. Attribution comes from our own config because MAF reports the bare tool name
|
||
with no server prefix (measured, not assumed), so a name allowed by two servers is recorded
|
||
unattributed rather than credited to the first match. Stated honesty limit: this is the call and
|
||
its source — not evidence that the service's answer reached the proposal.
|
||
- `docs/knowledge-base-recipe.md` (S5.3, D-H item 1): the documented team process (technical +
|
||
domain expert) for building a knowledge base, with the honest 1–2 week expectation.
|
||
- Test suite: 755 passing tests (4 skips are live-provider-only). Every wired seam is covered by a
|
||
load-bearing test that goes red when the seam is detached, and each carries a recorded detach point
|
||
verified by mutation. That property is maintained by measurement, not by assumption, and this
|
||
release contains two recorded cases where it did not hold until it was checked: the S3.1
|
||
`--semantic-retrieval` flag was covered only below the CLI, so hardcoding it off at `main()` level
|
||
left the suite green; and four of the five `BudgetExceeded` raise sites could report any value in
|
||
the `observed` field without a single test noticing. Both are now pinned.
|
||
|
||
### Fixed
|
||
Defects found and closed before this first tag. Nothing here was ever released; they are recorded
|
||
because each one changed an invariant an adopter depends on.
|
||
|
||
- Money quantisation happens in one order, from one source: `ledger.to_ore` is the framework's single
|
||
NOK→øre conversion, applied **per amount**, after which integers are summed. Previously the goal
|
||
baseline summed floats and quantised once while the ledger summed per-candidate integers, and the
|
||
two orders met in exactly one place — the check that decides whether a portfolio pass stops early
|
||
on a percentage goal. Measured divergence: three lines of `60000.005` NOK are `18000003` øre
|
||
quantised first but `18000001` summed first, enough to flip a goal. Both call sites were fixed; a
|
||
fix to only one survived the whole suite (measured).
|
||
- The MCP stdio ingest transport now runs against a real server, and its error contract is repaired:
|
||
`stdio_client` and `ClientSession` are each a task group, and anyio wraps everything leaving one in
|
||
an exception group, so the module's own `IngestError` reached callers as an exception group and
|
||
never as the type the ingest door switches on. The owned error is unwrapped and re-raised; anything
|
||
else is re-raised untouched. No canned-tool test could have caught this — the defect only appears
|
||
once the code is actually run.
|
||
- The MCP timeout composes with anyio's own cancel scope (`anyio.fail_after` nested inside both task
|
||
groups) instead of `asyncio.wait_for` from outside, which cancelled the structure anyio owns and
|
||
produced a `BrokenResourceError` inside an exception group rather than a timeout. The translation
|
||
is narrowed by `CancelScope.cancelled_caught`, so only the scope that hit its own deadline earns
|
||
the timeout code — an unconditional `except TimeoutError` would mislabel any `TimeoutError`, since
|
||
the builtin is also `socket.timeout` and `asyncio.TimeoutError`. `anyio` is now a declared direct
|
||
dependency rather than a transitive one.
|
||
- Frontmatter scalars are unquoted by exactly one rule (`okf.unquote_scalar`).
|
||
- A non-finite embedding is refused rather than scored.
|
||
- `--embedder-config` is refused when `--semantic-retrieval` is absent, rather than silently dropped.
|
||
- The portfolio CLI no longer swallows its offline door or its failures.
|
||
- Constants cited in the live documentation are gated against the code, so a doc and the value it
|
||
quotes cannot drift apart unnoticed.
|
||
|
||
### Notes
|
||
- Licensed under the MIT License (see `LICENSE`).
|
||
- Scope boundary: this is a technical framework. The deployer owns DPIA, risk assessment and
|
||
processing purpose; the framework ships only the technical preconditions (local-only operation,
|
||
provenance, no silent egress).
|
||
- The `shared/` directory is a git subtree of `portfolio-optimiser-commons` and is pull-only.
|