Version 1.0.0 across the four sites that carry it — pyproject.toml, __init__.py, uv.lock, test_smoke.py. Measured that these are the only four: README carries no version badge, and CHANGELOG's `## [0.1.0]` is history rather than a bump site. The heading stays `[Unreleased]`. STATE authorises the CHANGELOG CONTENT now and holds the TAG until after Wednesday's freeze, so stamping `## [1.0.0] - <date>` today would be a future-dated claim about an event that has not happened — and one to rewrite if the dress rehearsal fails or the freeze slips. `pyproject` at 1.0.0 with `[Unreleased]` populated is the release-prep state, not an inconsistency; nothing machine-reads the CHANGELOG (measured). The global versjonssync rule is read as CONTENT, not heading. Tag day is then one atomic move: rename the heading, stamp the date, tag. The re-lock was the hazard, and it was gated rather than assumed. Bumping the version stales `uv.lock`, and the next `uv run` would have re-locked it invisibly against a RANGE dependency (`agent-framework-core>=1.9.0,<2`) — while the two ExperimentalWarning lines are pinned byte-for-byte in the stderr golden, and STATE's own okf note records that a bare sync is enough for a guard to stop guarding with no local diff. So: bump, then `uv lock` EXPLICITLY, then diff before any test ran. The diff is the single `portfolio-optimiser` version line; agent-framework-core, llm-ingestion-okf (v0.3.2) and llm-ingestion-guard (v0.3.4) are untouched, and uv.lock was re-checked AFTER the suite to confirm no silent re-lock. CHANGELOG prose for the six feat commits `[Unreleased]` did not cover — it carried only Step 5 and the scripted registry. Console entry points and the golden transcript are Added; the Step-7 inbox, the anchored walkthrough, the stderr damping and the derived provenance sentence are Changed, scoped as the OFFLINE SIMULATION rather than framework runtime, since they change what the walkthrough exercises and not the library's behaviour. The content gate is Security, and carries its opt-in qualifier: `materialize` stays ungated by design and `materialize_gated` is asked for by name — an entry claiming "ingest now scans content before writing" without that clause would overclaim, and it sits next to the sentence read on stage Thursday. A Notes line names the two open boundaries (ingest stamp spec divergence, D7 mirroring) so 1.0.0 reads as a stable surface rather than a finished programme. Measured, not asserted: 810 passed / 4 skipped unchanged · ruff + mypy clean (31 source files) · no `0.1.0` remaining outside .venv/shared · and the demo RUN, not just tested — stdout byte-identical to tests/golden/demo-transcript.stdout, exit 0, 61 stdout / 4 stderr lines, matching dress rehearsal nr. 0. The version string appears nowhere in either golden (0 hits), so the bump could not move the fasit. Two STATE premises corrected by measurement: 24 commits since v0.1.0, not 23; and eight feat commits exist since the tag, of which six were undocumented. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ue1AnPZYsC9Tk7e5Tyv8Fv
23 KiB
23 KiB
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Added
- Step 5 is now observable:
generate_via_llmreturns aGenerationResultcarrying the validator falsifications that informed a later attempt, surfaced onRunResult.refinements. The offline simulation exercises it — the scripted proposer overclaims, the deterministic validator falsifies the number, and the refined proposal validates. - Two console entry points ship with
uv sync:portfolio-optimiser(the CLI) andportfolio-optimiser-demo(the offline walkthrough). Deliberately two of the package's fivemain()functions —costsim,hitlandpreflightstay operator tools invoked as modules, and every name here is one a release has to carry. Both invocation forms write byte-identical stdout. The entry points are tested against the INSTALLED distribution's metadata rather than the TOML: a[project.scripts]line that has never been synced is a claim, not a command. - The offline walkthrough's transcript is checked in as a golden fixture
(
tests/golden/demo-transcript.stdoutand.stderr). Self-identity across two runs cannot detect a regression — two runs of a regressed walkthrough agree exactly as well as two runs of a correct one — so the fixture leaves the process. stdout is pinned verbatim; stderr is normalised on exactly two measured environment spans, thesite-packagesprefix and the temporary directory, leaving thepo-sim-prefix visible because that belongs to the program rather than the environment. A companion control forbids the mask from widening: a normaliser that dropped whole lines, with the fixture regenerated beneath it, would keep both equality tests green.
Changed
- Breaking (library API):
generate_via_llmreturnsGenerationResultinstead ofValidatedProposal | Rejection; read.outcomefor the previous value. The refinement loop's bound is unchanged (max_attempts+ token meter). simulation.scripted_factoryaccepts a per-role reply selector over(prompt, role)as well as a constant reply, so a scripted role can answer differently on a later attempt.- The offline simulation's scripted proposer is now a candidate registry rather than a
hand-written reply:
simulation.scripted_proposer(candidates)builds the selector fromScriptedCandidateentries keyed on the project id the prompt names, andsimulate_learning_looptakesproject_idalongsidebundle_dir. Adding a project to the walkthrough is a data entry. A prompt matching no entry — or more than one — raisesScriptedCandidateErrorrather than answering with another project's numbers. - The offline simulation now EXERCISES the Step-7 file inbox it narrates. The verdict previously
arrived as a function argument — the short, in-run capture — while the trace line described the
long file loop. An expert now writes a real verdict file into an inbox between the two runs, and
the second run is given
verdict_dir=, sorun_projectmerges it before the Step-1 fold. The inbox sits beside the bundle copy and never inside it: a verdict file within the bundle would reach the next run as navigable context, which is a different mechanism wearing this one's clothes. The two time-scales carry SEPARATE markers by construction, since one marker on both paths would let either seam alone satisfy the assertion and leave the other free to rot;simulate_learning_looprefuses equal markers. - The offline walkthrough runs ANCHORED. Its deterministic gate reconciles each proposal against the
project's real cost lines, which activates only when the knowledge base ships a
cost-baseline.json; without one the gate reasoned solely about numbers the proposal supplied itself. The walkthrough now reads that file through exactly the seam a delivered knowledge base would use. For the synthetic fallback bundle, which cannot receive the file inside the pull-onlyshared/subtree, the baseline is DERIVED IN CODE from the scripted register rather than typed beside it — two sources of the same numbers drift, and drift is precisely what the walkthrough's own 10 % probe models. The declared baseline is printed, because an anchoring nobody can see is one nobody can check. - The walkthrough's stderr is quieter. The expected round-cap notice is dropped by a filter on the
emitting logger, keyed on the message and installed by
main()— never at import, so a library consumer keeps its own logging configuration. The twoExperimentalWarninglines are deliberately NOT damped: they fire while the package__init__imports the agent framework, always before the simulation's own imports and under both invocation forms, so silencing them would mean filtering warnings inside the library on every consumer's behalf. They are pinned in the golden fixture instead. stderr went from six lines to four. - The walkthrough DERIVES its provenance sentence for prior verdicts instead of stating it. The line above already computes the count, so a hand-written split would be a second copy of the same fact, free to drift the moment a knowledge base ships another seeded verdict.
Security
- Door A — the ingest path that materialises externally sourced documents into a knowledge base — can
now scan generated content before it is published, through
ingest.materialize_gated. The gate is opt-in and requested by name:materializeitself stays ungated by design, because golden suites pin its bytes and a caller that wants the gate asks for it. The seam sits around materialisation rather than inside it. The pinned upstream stages in memory and then performs its own disk phase, with no callback between the two, so a gate placed "at the write point" could only have run after the bytes had landed — a cleanup, not a gate. Instead the bundle is COPIED, materialised into the copy, scanned, and then published or discarded as a whole. The copy is load-bearing rather than convenient: the upstream's ownership scan, its collision gate against curated content, and its index merge all read the EXISTING bundle, so staging into an empty directory would publish a bundle stripped of its curated neighbours and their index links — data loss dressed as a security fix. Trust follows ORIGIN, never channel. The outcome is per BUNDLE, since partial publication would leave a bundle and index answering to no manifest, while diagnostics are per DOCUMENT so a single run reports every finding rather than only the first. Findings are written to the bundle'slog.mdand never to concept frontmatter, where four golden suites pin the bytes.
Notes
- The
1.0.0version signals a stable public surface, not a finished research programme. Two boundaries are open and named rather than implied: the ingest stamp predicate has diverged from the upstream specification (a value literal here, a structured field upstream) and does not touch the run path, and the mirroring of several seams to the sibling implementation is outstanding.
[0.1.0] - 2026-08-06
First tagged release. There is no prior release, so the entries below describe what this version is, not what changed since a predecessor.
Added
- Deterministic backbone: mandatory blocking validator (solver + Monte Carlo against the shared golden suite), budget meter with hard fail-fast caps, provenance stamping.
- Agentic learning loop wired end to end, one load-bearing seam at a time (target-picture steps 1, 3/4, 5, 7, 8): OKF-navigated bundle context with the gated ExpeL fold, maker-checker debate where the checker gates the reasoning, informed refinement (previous rejection reason fed into the next bounded attempt), async verdict file inbox, and gated wiki promotion (fail-closed).
- Offline end-to-end simulation proving the learning loop closes with a scripted client
(
uv run python -m portfolio_optimiser.simulation) — plumbing proof, not live-model proof. - Framework-neutral shared core in
shared/: OKF concept + example bundle, golden validator suite, and the expert-reviewer persona as an Agent Skill. - CLI parity (S5.3):
run.pydrives the whole method from the command line in two modes — single-project (--dimension-config,--outbox-dir/--run-id, plus the already-wired--bundle-dir/--verdict-dir) and portfolio (--portfoliowith--goals/--ledger), the latter printing an observablegoal reached: …line when a savings goal is met. Adds a fail-fastload_dimensionloader and structured refusals (rc 1, no traceback) for misuse and mode-exclusivity violations. The prior-verdict fold is on the--bundle-dirpath only; a--docs-dir-only run is single-shot. - Value report (S5.4): a read-only
--report [--json] --ledger <file>surface onrun.pythat rolls up the accumulatedSavingsLedger— per-project + portfolio totals (dimension-free-deduped integer øre), flagged cross-dimension overlaps (each counted once), and per-entry provenance — as a human table or deterministic JSON. Makes no model calls; mode-exclusive (only--ledger/--jsonpermitted with--report, which requires--ledger). Honest scope boundary: the report core (value_report.py) now exists but is deliberately not wired intocostsim'skost_mot_verdiplaceholder — that cost-vs-value integration is a separate, deferred step. Thecostsimseam note was reworded from the stale "fylles av S5.4 verdirapport" to a truthful forward reference socostsim's own output no longer claims the wiring is done. - Semantic retrieval seam (S3.1): a new MAF-free
semretrieval.pyadds anEmbedder/Retrieverpair and aHybridRankerblending a numpy cosine term over the embedded feature triple (sorted cost codes, measure type, magnitude bucket) with the existing structural score, exposed as--semantic-retrieval. What ships is the seam, not better retrieval quality: the bundledFakeEmbedderis a deterministic sha256 projection with no semantics, so over a structural tie the order is deterministic but arbitrary. A real embedder is selected from a CLOSED registry via--embedder-config/build_embedder— deliberately never an import path, so a config file can never name arbitrary code to load. Off by default and additive: with no retriever passed the store delegates toStructuralRetriever, which reproduces the pre-seam ranking exactly, so the text-excluded default and every existing test are unchanged. The embedding excludesdescription, matchingsimilarity("text is ignored by design") and the verdict-id hash — so a flag-on run reads no surface text either, and a genuine expert verdict can no longer be outranked by the framework's own echo of the query. The ranker is passed PER CALL, never assigned to the caller's store, so the opt-in cannot outlive the run that asked for it. Accepted in both run modes; in single-project mode it requires--bundle-dirand--verdict-dirand is refused — never silently ignored — without them. Determinism, precisely: ranking order rests on the total order(-round(score, 9), id); the BLAS thread pins set before numpy is imported (VECLIB_MAXIMUM_THREADSfor Accelerate,OPENBLAS/MKL/OMPfor other backends) defend the narrower claim that vector artifacts are byte-identical across environments. Ships an optional, rebuildablevectors.npy+vectors.jsonlstore (byte-identical regardless of insertion order; fail-fast on a row/line mismatch; missing loads asNone) — an authoring primitive with no caller insrc/, offered to extenders likewrite_verdictandpromote_verdict. numpy is confined tosemretrieval.pyand never entersokf.py,retrieval.pyorshared/. - Cost-baseline anchoring of the deterministic gate (S4.0):
validate_proposal(..., baseline=...)reconciles everyaffected_itemagainst the project's realCostBaselinein a stage 0, before the solver — two independent rejections (unknown cost code; a real code whosequantity/unit_costfalls outside a configurable tolerance, default 5 %, of the baseline value). Before this, every stage reasoned only about numbers the proposal itself supplied, so an internally consistent hallucination cleared the whole gate. Validation, never repair: the proposal is rejected, never silently corrected to the baseline. The argument is optional (Nonereproduces the earlier behaviour) but both run paths set it — the road path always, the bundle path only when the bundle ships acost-baseline.json, since a pre-amendment bundle is legitimately unanchored. A baseline that exists but is malformed raises on both loaders rather than reading as "no baseline". Method caps are looked up in an injectableMETHOD_CAPSregistry, so a second measure type is data rather than an edit to the validator. - Deterministic stage-2 bound and band enclosure in the validator (S2.7).
- Per-candidate verdict keying (S3.2):
seed_store_from_bundlereads eachtype: verdictfile's own structural frontmatter (affected_codes/measure_type/claimed_saving_nok) instead of collapsing a multi-candidate bundle onto the single IR projection, where a verdict about candidate B scored a perfect structural match against candidate A's query. All three fields or none — a partial declaration raises rather than being merged with the bundle candidate, which would mint a key belonging to neither. Bundles that declare none fall back to the previous keying, so every pre-S3.2 seed is unchanged. - Portfolio concurrency (S3.3): wave-partitioned execution with a per-project snapshot and a
deterministic merge barrier, pinned by a concurrent-equals-sequential contract test;
k < 1is fail-fast. Failures are collected and the pass continues, surfaced asRunFailureslots rather than aborting the portfolio. Goal-stop is evaluated at wave boundaries, with the intra-wave semantics documented and pinned. - Global portfolio token cap (S3.4/F10):
PortfolioBudget+PortfolioMeterform one ledger across a whole portfolio pass — and, seeded from a spend file, across passes — while the per-runBudget/TokenMeteris unchanged. Three enforcement points: a startup refusal when the remainder cannot fund a single run; wave admission, where an unfundable project is never started (a project that is merely aborted has already cost calls) and the pass stops structurally with the completed runs preserved; and a pre-call guard inBudgetMiddleware, so a call the remainder cannot pay for is refused rather than made. Because a whole wave is checked against the same pre-wave remainder, admission reserves each member's requirement. Resource exhaustion is reported in its ownbudget_stopfield, never merged intostop_reason— a goal stop is success. The spend file is our own accounting state, so corrupt content raises instead of reading as zero. - Ingest transports (S2.2, S2.4): the MCP connector as a transport inside the http family, and a time-bounded default http transport in front of the pinned ingest library. A server on the ingest path must expose a zero-argument tool (the URL carries both coordinates), which is a deliberately separate seam from the in-run data-source demo server.
- Azure/Foundry offline preflight config gate (
preflight.py, S4.1). - Offline live-dry-run drill (
--live-dry-run, S4.2): walks the whole path up to the eager client build and stops before the first model call — zero chat calls. - Out-of-band HITL verdict routing CLI (
hitl.py, S5.1): code-prefix routing on the candidate id. - Expert-notification contract (
notify.py, S5.2): the declaredNotifierwithConsoleNotifier,FileNotifier(byte-deterministic JSONL), andWebhookNotifier, plus a fail-fastbuild_notifier; the webhook is the only egress point, fail-closed behind an explicit per-runallow_egressopt-in. Not auto-wired intorun.py. - Offline whole-loop CLI door (
--scripted-replies): drives the entire method over the caller's own data with no model budget, in both run modes. Resolved before the portfolio dispatch, because placing it after made--portfolio --scripted-repliessilently drop the flag and attempt four real model calls (measured). Mutually exclusive with--live-dry-run: both are offline and they contradict, so the combination is refused rather than letting one quietly win. - Run mandate (
mandate.py,--mandate): a domain expert commissions which approaches a run evaluates, the run announces what it will do before doing it, and it settles for that — every commissioned approach is evaluated and every one appears in a coverage report. A mandate is not a dimension (dimension.admitsfilters what may pass the scoping gate; a mandate directs what the run spends its attempts on) and not a goal (the numeric target keeps its single home inGoalContract/--goalsand is merely restated in the announcement). Two construction-time refusals: an empty commission, and a duplicate approach id — including one claiming the reservedown-proposal— since the id keys each coverage row and two rows under one key collapse silently. Pure module: pydantic + stdlib only, D7-portable, guarded MAF-free. - MCP servers as in-run tools (
mcp_tools.py,--mcp-config): concrete external servers become tools the agents can call during the debate. Opt-in and additive — with no config there are zero network calls and the tool list is unchanged. Three rules are load-bearing: an allowlist is required (an empty one would let the far end decide what the agents may call), every server and every permitted tool is named in the announcement before the first call (also without--mandate— no undeclared egress), and--live-dry-runopens nothing. A malformed config is refused rather than degraded to "no external services", which would make the announcement describe a run nobody configured. This is a separate seam from the ingest-path MCP transport above. - Per-approach outbox artefacts: a run commissioned to evaluate three approaches previously wrote one
proposal artefact, so only the selected approach could ever receive a verdict and the others taught
the learning loop nothing. Artefacts are now keyed
{run_id}-{approach_id}-*.jsonand carryapproach_idin the payload, with the join key(run_id, approach_id)— widening only the filename would not have worked, because the reader joins on therun_idfield read from file content. Two properties make them genuinely judgeable:verdict_idis minted per approach (so one delivered verdict cannot clear all three from the queue), andprovenance.validator_decisionfollows its own approach rather than the run's.verdicts.verdict_keyis the single public verdict-key rule. - External-call provenance: the egress declaration says what a run may contact;
ToolCallRecorderrecords what it did, landing onProvenanceStamp.external_callsread after the debate. It only observes —call_nextis always awaited, so a trace can never alter the run it traces. Only configured tools are recorded; logging in-process functions would turn the record into a false egress claim, and an empty list is therefore a positive statement that nothing outside the process was contacted. Attribution comes from our own config because MAF reports the bare tool name with no server prefix (measured, not assumed), so a name allowed by two servers is recorded unattributed rather than credited to the first match. Stated honesty limit: this is the call and its source — not evidence that the service's answer reached the proposal. docs/knowledge-base-recipe.md(S5.3, D-H item 1): the documented team process (technical + domain expert) for building a knowledge base, with the honest 1–2 week expectation.- Test suite: 755 passing tests (4 skips are live-provider-only). Every wired seam is covered by a
load-bearing test that goes red when the seam is detached, and each carries a recorded detach point
verified by mutation. That property is maintained by measurement, not by assumption, and this
release contains two recorded cases where it did not hold until it was checked: the S3.1
--semantic-retrievalflag was covered only below the CLI, so hardcoding it off atmain()level left the suite green; and four of the fiveBudgetExceededraise sites could report any value in theobservedfield without a single test noticing. Both are now pinned.
Fixed
Defects found and closed before this first tag. Nothing here was ever released; they are recorded because each one changed an invariant an adopter depends on.
- Money quantisation happens in one order, from one source:
ledger.to_oreis the framework's single NOK→øre conversion, applied per amount, after which integers are summed. Previously the goal baseline summed floats and quantised once while the ledger summed per-candidate integers, and the two orders met in exactly one place — the check that decides whether a portfolio pass stops early on a percentage goal. Measured divergence: three lines of60000.005NOK are18000003øre quantised first but18000001summed first, enough to flip a goal. Both call sites were fixed; a fix to only one survived the whole suite (measured). - The MCP stdio ingest transport now runs against a real server, and its error contract is repaired:
stdio_clientandClientSessionare each a task group, and anyio wraps everything leaving one in an exception group, so the module's ownIngestErrorreached callers as an exception group and never as the type the ingest door switches on. The owned error is unwrapped and re-raised; anything else is re-raised untouched. No canned-tool test could have caught this — the defect only appears once the code is actually run. - The MCP timeout composes with anyio's own cancel scope (
anyio.fail_afternested inside both task groups) instead ofasyncio.wait_forfrom outside, which cancelled the structure anyio owns and produced aBrokenResourceErrorinside an exception group rather than a timeout. The translation is narrowed byCancelScope.cancelled_caught, so only the scope that hit its own deadline earns the timeout code — an unconditionalexcept TimeoutErrorwould mislabel anyTimeoutError, since the builtin is alsosocket.timeoutandasyncio.TimeoutError.anyiois now a declared direct dependency rather than a transitive one. - Frontmatter scalars are unquoted by exactly one rule (
okf.unquote_scalar). - A non-finite embedding is refused rather than scored.
--embedder-configis refused when--semantic-retrievalis absent, rather than silently dropped.- The portfolio CLI no longer swallows its offline door or its failures.
- Constants cited in the live documentation are gated against the code, so a doc and the value it quotes cannot drift apart unnoticed.
Notes
- Licensed under the MIT License (see
LICENSE). - Scope boundary: this is a technical framework. The deployer owns DPIA, risk assessment and processing purpose; the framework ships only the technical preconditions (local-only operation, provenance, no silent egress).
- The
shared/directory is a git subtree ofportfolio-optimiser-commonsand is pull-only.