The guard checked whether the installed SDK satisfied the pin. Nobody had ever checked whether anyone had READ it. Those are different questions, and the gap between them was a whole version range: pinned >=0.2.111,<0.3, premises source-verified through 0.2.110, installed 0.2.120. Every build in between was admissible and unexamined — `uv sync --upgrade` would have kept 806 tests green on an SDK no one had opened. Written red first: a guard handed 0.2.140 returned it without complaint. _VERIFIED_THROUGH is the ratchet. It records the newest build actually read at source, and a newer one fails naming the five premises to re-check. The pin is untouched and was never the defect — measurement dissolved the premise that it needed lifting. It was not too narrow but too wide, and a wider permission is not repaired by widening it further. The premises themselves were prose the failure message recited. Nothing tested them, so one that stopped being true would have surfaced on the one live paid run (S10, D6). They are now a table introspected against the installed package, with the printed prose derived from that same table so a checked attribute cannot go unreported or a reported one unchecked. The premise introspection structurally cannot see — that query() yields an AssistantMessage then a closing ResultMessage — is named apart, and is the honest reason the human reading still has to happen. Value-proved, not merely named: disabling the ratchet reds 1 test, stubbing the inventory to "no gaps" reds 3, re-hardcoding the prose reds 1, and lowering _VERIFIED_THROUGH below the installed build reds the real installed-version test rather than only a monkeypatched one. 0.2.139 read at source (0.2.120 -> 0.2.139, latest on PyPI today; STATE said 0.2.134, measured 08-09 and stale). The public query.py is byte-identical, every premise field keeps its type and default, and the parser changes are additive. One needed a look: 0.2.139 added a skills path defaulting setting_sources to ["user", "project"], which would have undone the S10 isolation fix — it fires only on None, so the explicit [] is out of reach. Prose carrying stale version claims moved with the reading, never ahead of it: each was re-verified at 0.2.139 before being restated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014dKDjVG7qrBh9NkAAxutqN
6.2 KiB
6.2 KiB
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Changed
- The SDK pin now has a proof behind it, not only a permission.
claude-agent-sdkmoved 0.2.120 → 0.2.139, and the version guard gained_VERIFIED_THROUGH: the newest build whose source was actually read. The pin (>=0.2.111,<0.3, unchanged) says whatuvmay resolve; it never said anyone had looked. Those are different questions, and conflating them left 0.2.111–0.2.999 admissible while the premises were source-verified only through 0.2.110 —uv sync --upgradewould have kept the suite green on an unexamined build. A build newer than the last one read now fails RED, naming the premises to re-check. A third question — "is something newer upstream?" — is deliberately NOT asked: it needs the network, and this suite stays offline. - The SDK premises are checked, not merely printed. They had existed only as a
sentence the failure message recited; nothing verified them, so a premise that stopped
being true would have surfaced on the one live paid run. They are now a table
introspected against the installed package, with the operator-facing prose derived from
that same table so the two cannot drift. The one premise introspection structurally
cannot see — that
query()yields anAssistantMessagethen a closingResultMessage— is named separately, and is the honest reason a human reading is still required. - Re-verified at 0.2.139 source level: the public
query.pyis byte-identical to 0.2.120, every premise field keeps its type and default, and the parser changes are additive (a neworiginpassthrough, a newConversationResetMessagethis client ignores). 0.2.139 also added a skills path that can defaultsetting_sourcesto["user", "project"]; it fires only onNone, so the explicit[]that carries the S10 isolation fix is out of its reach.
[0.1.0] - 2026-08-17
Added
The method, implemented (D7). Sibling implementation of the portfolio-optimiser method on the Claude Agent SDK, built from the shared frozen spec + golden suite alone — never by reverse-engineering the MAF sibling.
- Deterministic backbone — the typed cost-IR, the mandatory blocking validator (frozen by the shared golden suite, its only oracle), first-class provenance, and fail-fast startup contracts including the role → model map.
- Agentic loop — bounded generation, maker–checker debate, the validator gate, and informed refinement; the budget meter admits no unbounded loop anywhere, and carries an optional pre-call run-total USD belt on top of the post-charge token/round caps.
- Learning loop — the OKF context seam (navigation, never chunk-stuffing), the ExpeL-style experience fold, the async expert-verdict inbox, and the fail-closed promotion gate.
- Ingest layer — deterministic CSV (
file) and SQLite (sql) connectors in front of the loop, materializing OKF bundles the unchanged loop consumes; frozen by byte-identical golden extractions.http/MCP is an extension point this repo does not build, and a manifest naming it is rejected fail-fast. - Value layer — the fail-closed savings ledger (dimension-free sum, no double-counting), the hard/soft goal contract, the outbox output layer, the HITL pending/routing view, the pre-run cost simulation over schema-validated pricing config, the SDK/API preflight, opt-in notification sinks (webhook egress only behind an explicit per-run flag), and the per-run value report (modelled → expert-corrected → realized, goal progress, quantified learning effect, cost against value).
- Operator CLI — one collecting entrance (
run.py) for a single project (--bundle) or a portfolio (--portfolio,--verdict-dir), with--goals+--ledgerstopping a run before any model call when the target is already met, and--value-reportprojecting what the run delivered. Standalone entrances forvaluereport,hitl,costsimandpreflight. Every flag the README documents is checked against the actual--helpoutput by a load-bearing test. - Knowledge-base recipe — the documented team process for building the OKF bundles the
framework reads (
docs/oppskrift-kunnskapsbase.md), with an honest 1–2 week expectation. - Traceable run cost — the provenance stamp records which SDK build produced the run,
read from the producing client rather than the environment, so the SDK's cost estimate can
be traced to the price table that computed it. A run not produced by the SDK reports
nullinstead of borrowing the installed version. - The programme's one live model run (S10) — executed and validated at a documented
$0.127514, its four artifacts committed as fixed reference output under
runs/s10/. That record is never edited after the fact: it predates thesdk_versionfield and is left without one rather than back-filled with a guess. - Load-bearing tests — every seam is proven by a test that goes red when the seam is detached; the whole suite runs offline, with no API key and no network.
Notes
- Why
0.1.0and not1.0.0. Two gaps are documented and deliberately held open rather than papered over: the frozen spec's §7generatedshape (V1) has landed in the shared spec but is not yet adopted in the goldens — the golden is the library's own emission, so adoption is gated on the ingest-library pin swap, andtest_ingest_stamp_conformance_loadbearing.pyis the ratchet that goes red the moment amaterialize()run reaches the new shape. And the SDK pin (>=0.2.111,<0.3) reaches further than the range whose premises are source-verified (through 0.2.110). A1.0.0would claim a settled surface this implementation does not yet have. - Honesty rule (method spec §1): no artifact in this repo claims more than the implementation does. Scripted stand-ins are labelled as such, unbuilt extension points are named as unbuilt, and a figure the data does not carry is reported unmarked rather than back-filled.
- Licensed under the MIT License (see
LICENSE).