portfolio-optimiser/CHANGELOG.md
Kjell Tore Guttormsen 28a420ab97 feat(3): flaten sier sant om seg selv, og to gater måler at den fortsetter å gjøre det
Fase 3 (AAA+ på publisert flate). Tre av planens premisser falt på måling og er
rettet FØR handling, ikke etterpå:

* GOVERNANCE-raden hadde feil tiltak. Planen sa «skriv den»; org-ops D11 sier én
  kanonisk fil som hvert repo LENKER, og filen er nå publisert (målt: HTTP 200 på
  open/repo-standard). Å skrive vår egen ville gjort oss til kopi nr. 12 av en
  fil D11-bølgen holder på å rydde vekk. README lenker den, i samme form som
  repo-mailbox bruker, og bus-faktor 1 står uttalt i den kanoniske teksten.
* Release-objektet for v1.0.0 FINNES allerede på open/ (id 155, CHANGELOG-kropp,
  siden rendrer) — det som mangler er vedlegg, ikke objektet.
* WARN RELEASE-STALE fyrer ikke, og kan ikke: regelen sammenligner utgivelse mot
  tagg og er strukturelt blind for repo med null utgivelser (org-ops hovedbok
  #18). Gaten var OK/20 sjekker FØR arbeidet startet, så den kan ikke tjene som
  verifikasjon for denne fasen. Bevisene er Forgejo-APIet, filinnholdet og
  ren-klon-kjøringen.

A5-defekten rettet: env.template:21 sa at credential resolves via
DefaultAzureCredential. Den har aldri gjort det — backends.py:149 konstruerer
ManagedIdentityCredential eller AzureCliCredential, og Learns MAF-veiledning
navngir den spesifikke credentialen NETTOPP for å unngå probing. En operatør som
kopierte templaten ble fortalt at feil identitet ville bli brukt.

To load-bearing gater (Iron Law: begge røde før fiksen, 2 failed / 7 passed):

1. env.template navngir de credentials backends.py faktisk konstruerer, og ingen
   linje utgir DefaultAzureCredential for å være mekanismen. LINJEFORANKRET, ikke
   delstreng: backends.py NAVNGIR klassen fire ganger i kommentarene som
   begrunner hvorfor den ikke brukes, så en fil-bred substring-gate ville vært
   rød på nøyaktig den prosaen den beskytter (repoets 08-09-klasse, fjerde gang).
2. README-ens wheel-filnavn bærer versjonen bygget stempler på fila. Uten den
   ville en versjonsbump stille etterlatt en publisert install-kommando som peker
   på en fil som ikke finnes.

Hver positiv assert er paret med en KONTROLL på at det søkes etter noe som
finnes — en ekstraktor som stille finner null lager en gate som bare kan bli
grønn.

MUTASJONER MÅLT MOT HELE SUITEN, begge røde på riktig test og på INGEN annen:
gjeninnfør den usanne credential-påstanden (2 røde, 844 grønne) · la
wheel-filnavnet drifte til 1.0.0 (1 rød, 845 grønne). Restaurert fra scratchpad
+ shasum -c mellom hver. Bumpen selv var den andre mutasjonen: pyproject 1.0.0 →
1.1.0 gjorde README-gaten rød alene, før README ble rettet.

SECURITY.md: varslingsfrist (minst én minor-release og aldri under 30 dager
mellom kunngjøring og fjerning, med sikkerhetskritisk fjerning som uttalt
unntak). Støttetabellen er bevisst VERSJONSFRI — et release-nummer skrevet der
ville drevet ved neste tagg, altså samme defektklasse som gate 2 fanger.

CLAUDE.md beholdt på flaten med en engelsk innramming øverst (operatørvalg): den
sier hva fila er for en fremmed. Innholdet er repoets sterkeste bevis på at hver
beslutning er målt; å fjerne det ville fjernet bevis, ikke friksjon.

Versjon 1.1.0 — synket i pyproject, __init__, test_smoke og README-kommandoen.
1.0.0-treet kan ikke produsere en kjørbar wheel (force-include kom etter taggen,
målt: git show v1.0.0:pyproject.toml har den ikke), så en wheel hengt på den
utgivelsen ville vært nøyaktig den usanne påstanden denne fasen finnes for å
fjerne. Operatøren valgte bumpen framfor et vedlegg som ikke virker.

846 passed / 4 skipped (fra 837). ruff + format + mypy rene.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011ckyg3Pc6k7FRuR6fDGQLJ
2026-08-14 06:57:25 +02:00

343 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [1.1.0] - 2026-08-14
The release that makes the distribution stand on its own. `1.0.0` shipped a framework that could
only run from a checkout; this one runs from an installed wheel and inside a Foundry-hosted
container, and it is the first release whose artefacts are published alongside it.
### Added
- `shared/` now travels **as packaged data**. The wheel carries a byte-identical mirror of the tree
under `portfolio_optimiser/_shared/`, and `shared_root()` resolves at call time in a fixed order:
`PORTFOLIO_SHARED_ROOT`, then the working tree's `shared/` when one exists, then the packaged
copy. The working tree stays authoritative in a checkout — that is what keeps the pull-only
subtree contract and the byte-exact goldens untouched. Measured before and after: the `1.0.0`
wheel carried 58 files and none under `shared/`; this one carries 122, of which 64 are the mirror.
- A **hosted entry point**: `main.py` wraps `run_project` on a single asyncio loop and serves the
Foundry hosting contract (`GET /readiness`, `POST /invocations`, SIGTERM → exit 0), with a
`Dockerfile` and `azure.yaml` beside it. The invocation payload is whitelisted against
`run_project`'s own signature — an unknown field is refused by name with a 400 rather than
silently dropped — and a rejected proposal is a successful run (200), because the negative outcome
belongs to the payload and never to the transport.
- Wheel-install instructions in the README. A wheel is not installable on its own: two dependencies
are pinned to git tags, and `[tool.uv.sources]` does not travel with wheel metadata, so both must
be supplied as requirements alongside the wheel. The published command is the one that was
measured (65 packages, exit 0), not one composed afterwards.
- Two gates on claims the **published surface** makes about itself: that `env.template` names the
credentials `backends.py` actually constructs, and that the README's wheel-install command spells
the version the build stamps on the file. Both read raw text and are line-anchored, because prose
is the only place these claims live.
### Changed
- The AZURE profile now reads **its own environment** rather than the operator's laptop. The
endpoint resolves to the first non-empty of `PORTFOLIO_FOUNDRY_PROJECT_ENDPOINT` and Foundry's
injected `FOUNDRY_PROJECT_ENDPOINT`; precedence applies to *values*, so an exported-but-empty name
falls through instead of masking a real one. The credential follows the same environment:
`AzureCliCredential` on a developer host, `ManagedIdentityCredential` when
`FOUNDRY_HOSTING_ENVIRONMENT` holds a non-empty value — never `DefaultAzureCredential`, whose
probing would walk a credential chain that cannot succeed in a container and turn a config error
into a slow one.
- `SECURITY.md` states a **deprecation notice period**: at least one minor release and no fewer than
30 days between announcement and removal, with security-critical removals named as the explicit
exception. The supported-versions table is deliberately version-free, since a release number
written there would drift at the next tag.
- The README links the organisation's single canonical `GOVERNANCE.md` instead of vendoring a copy,
and states the maintenance model (solo-maintained, no SLA, fork-and-own) on the first screen.
- `CLAUDE.md` opens with an English note explaining what the file is for a visitor: the working
agreement with the AI agent that builds this repository, doubling as its invariant ledger.
### Fixed
- `env.template` claimed the AZURE profile resolved its credential through `DefaultAzureCredential`.
It never has. An operator copying the template was told the wrong identity would be used.
## [1.0.0] - 2026-08-12
### Added
- Step 5 is now observable: `generate_via_llm` returns a `GenerationResult` carrying the validator
falsifications that informed a later attempt, surfaced on `RunResult.refinements`. The offline
simulation exercises it — the scripted proposer overclaims, the deterministic validator falsifies
the number, and the refined proposal validates.
- Two console entry points ship with `uv sync`: `portfolio-optimiser` (the CLI) and
`portfolio-optimiser-demo` (the offline walkthrough). Deliberately two of the package's five
`main()` functions — `costsim`, `hitl` and `preflight` stay operator tools invoked as modules, and
every name here is one a release has to carry. Both invocation forms write byte-identical stdout.
The entry points are tested against the INSTALLED distribution's metadata rather than the TOML: a
`[project.scripts]` line that has never been synced is a claim, not a command.
- The offline walkthrough's transcript is checked in as a golden fixture
(`tests/golden/demo-transcript.stdout` and `.stderr`). Self-identity across two runs cannot detect
a regression — two runs of a regressed walkthrough agree exactly as well as two runs of a correct
one — so the fixture leaves the process. stdout is pinned verbatim; stderr is normalised on exactly
two measured environment spans, the `site-packages` prefix and the temporary directory, leaving the
`po-sim-` prefix visible because that belongs to the program rather than the environment. A
companion control forbids the mask from widening: a normaliser that dropped whole lines, with the
fixture regenerated beneath it, would keep both equality tests green.
### Changed
- **Breaking (library API):** `generate_via_llm` returns `GenerationResult` instead of
`ValidatedProposal | Rejection`; read `.outcome` for the previous value. The refinement loop's
bound is unchanged (`max_attempts` + token meter).
- `simulation.scripted_factory` accepts a per-role reply *selector* over `(prompt, role)` as well as
a constant reply, so a scripted role can answer differently on a later attempt.
- The offline simulation's scripted proposer is now a candidate **registry** rather than a
hand-written reply: `simulation.scripted_proposer(candidates)` builds the selector from
`ScriptedCandidate` entries keyed on the project id the prompt names, and
`simulate_learning_loop` takes `project_id` alongside `bundle_dir`. Adding a project to the
walkthrough is a data entry. A prompt matching no entry — or more than one — raises
`ScriptedCandidateError` rather than answering with another project's numbers.
- The offline simulation now EXERCISES the Step-7 file inbox it narrates. The verdict previously
arrived as a function argument — the short, in-run capture — while the trace line described the
long file loop. An expert now writes a real verdict file into an inbox between the two runs, and
the second run is given `verdict_dir=`, so `run_project` merges it before the Step-1 fold. The
inbox sits beside the bundle copy and never inside it: a verdict file within the bundle would reach
the next run as navigable context, which is a different mechanism wearing this one's clothes. The
two time-scales carry SEPARATE markers by construction, since one marker on both paths would let
either seam alone satisfy the assertion and leave the other free to rot; `simulate_learning_loop`
refuses equal markers.
- The offline walkthrough runs ANCHORED. Its deterministic gate reconciles each proposal against the
project's real cost lines, which activates only when the knowledge base ships a `cost-baseline.json`;
without one the gate reasoned solely about numbers the proposal supplied itself. The walkthrough now
reads that file through exactly the seam a delivered knowledge base would use. For the synthetic
fallback bundle, which cannot receive the file inside the pull-only `shared/` subtree, the baseline
is DERIVED IN CODE from the scripted register rather than typed beside it — two sources of the same
numbers drift, and drift is precisely what the walkthrough's own 10 % probe models. The declared
baseline is printed, because an anchoring nobody can see is one nobody can check.
- The walkthrough's stderr is quieter. The expected round-cap notice is dropped by a filter on the
emitting logger, keyed on the message and installed by `main()` — never at import, so a library
consumer keeps its own logging configuration. The two `ExperimentalWarning` lines are deliberately
NOT damped: they fire while the package `__init__` imports the agent framework, always before the
simulation's own imports and under both invocation forms, so silencing them would mean filtering
warnings inside the library on every consumer's behalf. They are pinned in the golden fixture
instead. stderr went from six lines to four.
- The walkthrough DERIVES its provenance sentence for prior verdicts instead of stating it. The line
above already computes the count, so a hand-written split would be a second copy of the same fact,
free to drift the moment a knowledge base ships another seeded verdict.
- The shared expert-reviewer persona's canonical example verdict is worded domain-neutrally
("i tilsvarende anlegg" rather than "i kontorbygg"), pulled from the upstream commons repository.
The walkthrough prints that `rationale` verbatim, so the wording was a building-type justification
read out over a road-lighting project; it could not be fixed downstream, because overriding the
text locally would re-stub the very artifact the shared skill exists to make load-bearing. The
`marker` value is byte-unchanged, and the pinned transcript fixture was regenerated against a
prediction written before the pull — the printed line is clipped at a fixed width, so the swap
moves the tail as well, and a regeneration without a written prediction could not tell that
expected shift apart from drift.
### Security
- Door A — the ingest path that materialises externally sourced documents into a knowledge base — can
now scan generated content before it is published, through `ingest.materialize_gated`. The gate is
**opt-in and requested by name**: `materialize` itself stays ungated by design, because golden
suites pin its bytes and a caller that wants the gate asks for it.
The seam sits *around* materialisation rather than inside it. The pinned upstream stages in memory
and then performs its own disk phase, with no callback between the two, so a gate placed "at the
write point" could only have run after the bytes had landed — a cleanup, not a gate. Instead the
bundle is COPIED, materialised into the copy, scanned, and then published or discarded as a whole.
The copy is load-bearing rather than convenient: the upstream's ownership scan, its collision gate
against curated content, and its index merge all read the EXISTING bundle, so staging into an empty
directory would publish a bundle stripped of its curated neighbours and their index links — data
loss dressed as a security fix.
Trust follows ORIGIN, never channel. The outcome is per BUNDLE, since partial publication would
leave a bundle and index answering to no manifest, while diagnostics are per DOCUMENT so a single
run reports every finding rather than only the first. Findings are written to the bundle's `log.md`
and never to concept frontmatter, where four golden suites pin the bytes.
### Notes
- The `1.0.0` version signals a stable public surface, not a finished research programme. Two
boundaries are open and named rather than implied: the ingest stamp predicate has diverged from the
upstream specification (a value literal here, a structured field upstream) and does not touch the
run path, and the mirroring of several seams to the sibling implementation is outstanding.
## [0.1.0] - 2026-08-06
First tagged release. There is no prior release, so the entries below describe what this version
**is**, not what changed since a predecessor.
### Added
- Deterministic backbone: mandatory blocking validator (solver + Monte Carlo against the shared
golden suite), budget meter with hard fail-fast caps, provenance stamping.
- Agentic learning loop wired end to end, one load-bearing seam at a time (target-picture steps
1, 3/4, 5, 7, 8): OKF-navigated bundle context with the gated ExpeL fold, maker-checker debate
where the checker gates the reasoning, informed refinement (previous rejection reason fed into
the next bounded attempt), async verdict file inbox, and gated wiki promotion (fail-closed).
- Offline end-to-end simulation proving the learning loop closes with a scripted client
(`uv run python -m portfolio_optimiser.simulation`) — plumbing proof, not live-model proof.
- Framework-neutral shared core in `shared/`: OKF concept + example bundle, golden validator
suite, and the expert-reviewer persona as an Agent Skill.
- CLI parity (S5.3): `run.py` drives the whole method from the command line in **two modes**
single-project (`--dimension-config`, `--outbox-dir`/`--run-id`, plus the already-wired
`--bundle-dir`/`--verdict-dir`) and portfolio (`--portfolio` with `--goals`/`--ledger`), the
latter printing an observable `goal reached: …` line when a savings goal is met. Adds a
fail-fast `load_dimension` loader and structured refusals (rc 1, no traceback) for misuse and
mode-exclusivity violations. The prior-verdict fold is on the `--bundle-dir` path only; a
`--docs-dir`-only run is single-shot.
- Value report (S5.4): a read-only `--report [--json] --ledger <file>` surface on `run.py` that
rolls up the accumulated `SavingsLedger` — per-project + portfolio totals (dimension-free-deduped
integer øre), flagged cross-dimension overlaps (each counted once), and per-entry provenance — as
a human table or deterministic JSON. Makes no model calls; mode-exclusive (only `--ledger`/`--json`
permitted with `--report`, which requires `--ledger`). Honest scope boundary: the report core
(`value_report.py`) now **exists** but is deliberately **not** wired into `costsim`'s
`kost_mot_verdi` placeholder — that cost-vs-value integration is a separate, deferred step. The
`costsim` seam note was reworded from the stale "fylles av S5.4 verdirapport" to a truthful
forward reference so `costsim`'s own output no longer claims the wiring is done.
- Semantic retrieval seam (S3.1): a new MAF-free `semretrieval.py` adds an `Embedder`/`Retriever`
pair and a `HybridRanker` blending a numpy cosine term over the embedded feature triple (sorted
cost codes, measure type, magnitude bucket) with the existing structural score, exposed as
`--semantic-retrieval`. **What ships is the seam, not better retrieval quality**: the bundled
`FakeEmbedder` is a deterministic sha256 projection with no semantics, so over a structural tie
the order is deterministic but arbitrary. A real embedder is selected from a CLOSED registry via
`--embedder-config` / `build_embedder` — deliberately never an import path, so a config file can
never name arbitrary code to load.
**Off by default and additive**: with no retriever passed the store delegates to
`StructuralRetriever`, which reproduces the pre-seam ranking exactly, so the text-excluded
default and every existing test are unchanged. The embedding excludes `description`, matching
`similarity` ("text is ignored by design") and the verdict-id hash — so a flag-on run reads no
surface text either, and a genuine expert verdict can no longer be outranked by the framework's
own echo of the query. The ranker is passed PER CALL, never assigned to the caller's store, so
the opt-in cannot outlive the run that asked for it. Accepted in both run modes; in
single-project mode it requires `--bundle-dir` and `--verdict-dir` and is refused — never
silently ignored — without them.
Determinism, precisely: ranking order rests on the total order `(-round(score, 9), id)`; the
BLAS thread pins set before numpy is imported (`VECLIB_MAXIMUM_THREADS` for Accelerate,
`OPENBLAS`/`MKL`/`OMP` for other backends) defend the narrower claim that vector artifacts are
byte-identical across environments. Ships an optional, rebuildable `vectors.npy` +
`vectors.jsonl` store (byte-identical regardless of insertion order; fail-fast on a row/line
mismatch; missing loads as `None`) — an authoring primitive with no caller in `src/`, offered to
extenders like `write_verdict` and `promote_verdict`. numpy is confined to `semretrieval.py` and
never enters `okf.py`, `retrieval.py` or `shared/`.
- Cost-baseline anchoring of the deterministic gate (S4.0): `validate_proposal(..., baseline=...)`
reconciles every `affected_item` against the project's real `CostBaseline` in a **stage 0, before
the solver** — two independent rejections (unknown cost code; a real code whose
`quantity`/`unit_cost` falls outside a configurable tolerance, default 5 %, of the baseline value).
Before this, every stage reasoned only about numbers the proposal itself supplied, so an internally
consistent hallucination cleared the whole gate. Validation, never repair: the proposal is
rejected, never silently corrected to the baseline. The argument is **optional** (`None` reproduces
the earlier behaviour) but both run paths set it — the road path always, the bundle path only when
the bundle ships a `cost-baseline.json`, since a pre-amendment bundle is legitimately unanchored.
A baseline that exists but is malformed raises on both loaders rather than reading as "no
baseline". Method caps are looked up in an injectable `METHOD_CAPS` registry, so a second measure
type is data rather than an edit to the validator.
- Deterministic stage-2 bound and band enclosure in the validator (S2.7).
- Per-candidate verdict keying (S3.2): `seed_store_from_bundle` reads each `type: verdict` file's own
structural frontmatter (`affected_codes` / `measure_type` / `claimed_saving_nok`) instead of
collapsing a multi-candidate bundle onto the single IR projection, where a verdict about candidate
B scored a perfect structural match against candidate A's query. All three fields or none — a
partial declaration raises rather than being merged with the bundle candidate, which would mint a
key belonging to neither. Bundles that declare none fall back to the previous keying, so every
pre-S3.2 seed is unchanged.
- Portfolio concurrency (S3.3): wave-partitioned execution with a per-project snapshot and a
deterministic merge barrier, pinned by a concurrent-equals-sequential contract test; `k < 1` is
fail-fast. Failures are collected and the pass continues, surfaced as `RunFailure` slots rather
than aborting the portfolio. Goal-stop is evaluated at wave boundaries, with the intra-wave
semantics documented and pinned.
- Global portfolio token cap (S3.4/F10): `PortfolioBudget` + `PortfolioMeter` form one ledger across
a whole portfolio pass — and, seeded from a spend file, across passes — while the per-run
`Budget`/`TokenMeter` is unchanged. Three enforcement points: a **startup refusal** when the
remainder cannot fund a single run; **wave admission**, where an unfundable project is never
started (a project that is merely aborted has already cost calls) and the pass stops structurally
with the completed runs preserved; and a **pre-call guard** in `BudgetMiddleware`, so a call the
remainder cannot pay for is refused rather than made. Because a whole wave is checked against the
same pre-wave remainder, admission reserves each member's requirement. Resource exhaustion is
reported in its own `budget_stop` field, never merged into `stop_reason` — a goal stop is success.
The spend file is our own accounting state, so corrupt content raises instead of reading as zero.
- Ingest transports (S2.2, S2.4): the MCP connector as a transport inside the http family, and a
time-bounded default http transport in front of the pinned ingest library. A server on the ingest
path must expose a **zero-argument tool** (the URL carries both coordinates), which is a
deliberately separate seam from the in-run data-source demo server.
- Azure/Foundry offline preflight config gate (`preflight.py`, S4.1).
- Offline live-dry-run drill (`--live-dry-run`, S4.2): walks the whole path up to the eager
client build and stops before the first model call — zero chat calls.
- Out-of-band HITL verdict routing CLI (`hitl.py`, S5.1): code-prefix routing on the candidate id.
- Expert-notification contract (`notify.py`, S5.2): the declared `Notifier` with `ConsoleNotifier`,
`FileNotifier` (byte-deterministic JSONL), and `WebhookNotifier`, plus a fail-fast
`build_notifier`; the webhook is the only egress point, fail-closed behind an explicit per-run
`allow_egress` opt-in. Not auto-wired into `run.py`.
- Offline whole-loop CLI door (`--scripted-replies`): drives the entire method over the caller's own
data with no model budget, in **both** run modes. Resolved before the portfolio dispatch, because
placing it after made `--portfolio --scripted-replies` silently drop the flag and attempt four real
model calls (measured). Mutually exclusive with `--live-dry-run`: both are offline and they
contradict, so the combination is refused rather than letting one quietly win.
- Run mandate (`mandate.py`, `--mandate`): a domain expert commissions **which approaches** a run
evaluates, the run announces what it will do before doing it, and it settles for that — every
commissioned approach is evaluated and every one appears in a coverage report. A mandate is not a
dimension (`dimension.admits` filters what may pass the scoping gate; a mandate directs what the
run spends its attempts on) and not a goal (the numeric target keeps its single home in
`GoalContract`/`--goals` and is merely restated in the announcement). Two construction-time
refusals: an empty commission, and a duplicate approach id — including one claiming the reserved
`own-proposal` — since the id keys each coverage row and two rows under one key collapse silently.
Pure module: pydantic + stdlib only, D7-portable, guarded MAF-free.
- MCP servers as in-run tools (`mcp_tools.py`, `--mcp-config`): concrete external servers become
tools the agents can call **during** the debate. Opt-in and additive — with no config there are
zero network calls and the tool list is unchanged. Three rules are load-bearing: an **allowlist is
required** (an empty one would let the far end decide what the agents may call), **every server and
every permitted tool is named in the announcement before the first call** (also without
`--mandate` — no undeclared egress), and `--live-dry-run` opens nothing. A malformed config is
refused rather than degraded to "no external services", which would make the announcement describe
a run nobody configured. This is a separate seam from the ingest-path MCP transport above.
- Per-approach outbox artefacts: a run commissioned to evaluate three approaches previously wrote one
proposal artefact, so only the selected approach could ever receive a verdict and the others taught
the learning loop nothing. Artefacts are now keyed `{run_id}-{approach_id}-*.json` and carry
`approach_id` in the payload, with the join key `(run_id, approach_id)` — widening only the
filename would not have worked, because the reader joins on the `run_id` field read from file
content. Two properties make them genuinely judgeable: `verdict_id` is minted per approach (so one
delivered verdict cannot clear all three from the queue), and `provenance.validator_decision`
follows its own approach rather than the run's. `verdicts.verdict_key` is the single public
verdict-key rule.
- External-call provenance: the egress declaration says what a run **may** contact;
`ToolCallRecorder` records what it **did**, landing on `ProvenanceStamp.external_calls` read after
the debate. It only observes — `call_next` is always awaited, so a trace can never alter the run it
traces. Only configured tools are recorded; logging in-process functions would turn the record into
a false egress claim, and an empty list is therefore a positive statement that nothing outside the
process was contacted. Attribution comes from our own config because MAF reports the bare tool name
with no server prefix (measured, not assumed), so a name allowed by two servers is recorded
unattributed rather than credited to the first match. Stated honesty limit: this is the call and
its source — not evidence that the service's answer reached the proposal.
- `docs/knowledge-base-recipe.md` (S5.3, D-H item 1): the documented team process (technical +
domain expert) for building a knowledge base, with the honest 12 week expectation.
- Test suite: 755 passing tests (4 skips are live-provider-only). Every wired seam is covered by a
load-bearing test that goes red when the seam is detached, and each carries a recorded detach point
verified by mutation. That property is maintained by measurement, not by assumption, and this
release contains two recorded cases where it did not hold until it was checked: the S3.1
`--semantic-retrieval` flag was covered only below the CLI, so hardcoding it off at `main()` level
left the suite green; and four of the five `BudgetExceeded` raise sites could report any value in
the `observed` field without a single test noticing. Both are now pinned.
### Fixed
Defects found and closed before this first tag. Nothing here was ever released; they are recorded
because each one changed an invariant an adopter depends on.
- Money quantisation happens in one order, from one source: `ledger.to_ore` is the framework's single
NOK→øre conversion, applied **per amount**, after which integers are summed. Previously the goal
baseline summed floats and quantised once while the ledger summed per-candidate integers, and the
two orders met in exactly one place — the check that decides whether a portfolio pass stops early
on a percentage goal. Measured divergence: three lines of `60000.005` NOK are `18000003` øre
quantised first but `18000001` summed first, enough to flip a goal. Both call sites were fixed; a
fix to only one survived the whole suite (measured).
- The MCP stdio ingest transport now runs against a real server, and its error contract is repaired:
`stdio_client` and `ClientSession` are each a task group, and anyio wraps everything leaving one in
an exception group, so the module's own `IngestError` reached callers as an exception group and
never as the type the ingest door switches on. The owned error is unwrapped and re-raised; anything
else is re-raised untouched. No canned-tool test could have caught this — the defect only appears
once the code is actually run.
- The MCP timeout composes with anyio's own cancel scope (`anyio.fail_after` nested inside both task
groups) instead of `asyncio.wait_for` from outside, which cancelled the structure anyio owns and
produced a `BrokenResourceError` inside an exception group rather than a timeout. The translation
is narrowed by `CancelScope.cancelled_caught`, so only the scope that hit its own deadline earns
the timeout code — an unconditional `except TimeoutError` would mislabel any `TimeoutError`, since
the builtin is also `socket.timeout` and `asyncio.TimeoutError`. `anyio` is now a declared direct
dependency rather than a transitive one.
- Frontmatter scalars are unquoted by exactly one rule (`okf.unquote_scalar`).
- A non-finite embedding is refused rather than scored.
- `--embedder-config` is refused when `--semantic-retrieval` is absent, rather than silently dropped.
- The portfolio CLI no longer swallows its offline door or its failures.
- Constants cited in the live documentation are gated against the code, so a doc and the value it
quotes cannot drift apart unnoticed.
### Notes
- Licensed under the MIT License (see `LICENSE`).
- Scope boundary: this is a technical framework. The deployer owns DPIA, risk assessment and
processing purpose; the framework ships only the technical preconditions (local-only operation,
provenance, no silent egress).
- The `shared/` directory is a git subtree of `portfolio-optimiser-commons` and is pull-only.