portfolio-optimiser-claude/CHANGELOG.md
Kjell Tore Guttormsen 0b69354c74 docs(changelog): cut 0.1.0 — the version the two open gaps allow
The repo has never had a tag, so this is version one. The number is chosen
from maturity, not from habit: the shared spec's V1 `generated` shape is
landed upstream but not adopted here (the golden is the library's emission,
so adoption is gated on the ingest pin swap), and the SDK pin reaches
further than the range whose premises are source-verified. Both gaps are
held open by tests on purpose. 1.0.0 would claim a settled surface this
implementation does not have.

pyproject.toml already declares 0.1.0 and needs no change; README carries
three badges, none of them versioned, and src/ declares no __version__ —
enumerated, not assumed, so nothing else can drift from the tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014eYPddPVYPMxc5L7a4nxvA
2026-08-17 13:33:19 +02:00

68 lines
4.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
## [0.1.0] - 2026-08-17
### Added
**The method, implemented (D7).** Sibling implementation of the portfolio-optimiser method
on the Claude Agent SDK, built from the shared frozen spec + golden suite alone — never by
reverse-engineering the MAF sibling.
- **Deterministic backbone** — the typed cost-IR, the mandatory blocking validator (frozen by
the shared golden suite, its only oracle), first-class provenance, and fail-fast startup
contracts including the role → model map.
- **Agentic loop** — bounded generation, makerchecker debate, the validator gate, and
informed refinement; the budget meter admits no unbounded loop anywhere, and carries an
optional pre-call run-total USD belt on top of the post-charge token/round caps.
- **Learning loop** — the OKF context seam (navigation, never chunk-stuffing), the ExpeL-style
experience fold, the async expert-verdict inbox, and the fail-closed promotion gate.
- **Ingest layer** — deterministic CSV (`file`) and SQLite (`sql`) connectors in front of the
loop, materializing OKF bundles the unchanged loop consumes; frozen by byte-identical golden
extractions. `http`/MCP is an extension point this repo does **not** build, and a manifest
naming it is rejected fail-fast.
- **Value layer** — the fail-closed savings ledger (dimension-free sum, no double-counting),
the hard/soft goal contract, the outbox output layer, the HITL pending/routing view, the
pre-run cost simulation over schema-validated pricing config, the SDK/API preflight, opt-in
notification sinks (webhook egress only behind an explicit per-run flag), and the per-run
value report (modelled → expert-corrected → realized, goal progress, quantified learning
effect, cost against value).
- **Operator CLI** — one collecting entrance (`run.py`) for a single project (`--bundle`) or a
portfolio (`--portfolio`, `--verdict-dir`), with `--goals` + `--ledger` stopping a run before
any model call when the target is already met, and `--value-report` projecting what the run
delivered. Standalone entrances for `valuereport`, `hitl`, `costsim` and `preflight`. Every
flag the README documents is checked against the actual `--help` output by a load-bearing
test.
- **Knowledge-base recipe** — the documented team process for building the OKF bundles the
framework reads (`docs/oppskrift-kunnskapsbase.md`), with an honest 12 week expectation.
- **Traceable run cost** — the provenance stamp records which SDK build produced the run,
read from the producing client rather than the environment, so the SDK's cost estimate can
be traced to the price table that computed it. A run not produced by the SDK reports `null`
instead of borrowing the installed version.
- **The programme's one live model run** (S10) — executed and validated at a documented
$0.127514, its four artifacts committed as fixed reference output under `runs/s10/`. That
record is never edited after the fact: it predates the `sdk_version` field and is left
without one rather than back-filled with a guess.
- **Load-bearing tests** — every seam is proven by a test that goes red when the seam is
detached; the whole suite runs offline, with no API key and no network.
### Notes
- **Why `0.1.0` and not `1.0.0`.** Two gaps are documented and deliberately held open rather
than papered over: the frozen spec's §7 `generated` shape (V1) has landed in the shared spec
but is *not* yet adopted in the goldens — the golden is the library's own emission, so
adoption is gated on the ingest-library pin swap, and
`test_ingest_stamp_conformance_loadbearing.py` is the ratchet that goes red the moment a
`materialize()` run reaches the new shape. And the SDK pin (`>=0.2.111,<0.3`) reaches further
than the range whose premises are source-verified (through 0.2.110). A `1.0.0` would claim a
settled surface this implementation does not yet have.
- Honesty rule (method spec §1): no artifact in this repo claims more than the implementation
does. Scripted stand-ins are labelled as such, unbuilt extension points are named as
unbuilt, and a figure the data does not carry is reported unmarked rather than back-filled.
- Licensed under the MIT License (see `LICENSE`).