Ran `repo-standard` (v0.1.1, class `standalone`) and fixed everything it flagged as ERROR, plus the WARN links that were genuinely dead. README first screen: - opening line is now byte-identical to the forge description, so description == catalog == README is machine-checkable (badges moved below). - `## Install` (required for class `standalone`): clone + `uv sync`, stated as clone-only because the shared spec, persona skill and example bundles under `shared/` are read from the working tree at run time. `uv run pytest` named as the verification, with the fact that no CI runner exists said out loud rather than implied by a badge. - `## Non-goals` (required): the five limits already binding in CLAUDE.md — not a compliance product, not a portfolio-level reallocator, not autonomous decision-making, not turnkey, not a model benchmark. Dead relative links (measured, not guessed): - `docs/plan/2026-07-10-sesjonsplan-fase2-6.md` pointed at `../2026-07-14-revisjonspakke-DF-DI.md` six times; the file sits in `docs/plan/`, not `docs/`. (The sibling `../review-2026-07.md` links are correct and untouched.) - the Fase-1 spike brief linked repo-root-relative from `.claude/projects/…/`; re-anchored with `../../../`. The one remaining README ERROR was a gate false positive: `checkInternalLinks` resolves targets against `git ls-files`, which lists files only, so a link to a directory can never resolve. `[shared/](shared/)` now points at `shared/README.md` — a better target anyway, since that file carries the pull-only subtree rule. Not fixed here: the classifier lives in another repo. Remaining WARNs are all inside `shared/`, deliberately untouched: it is a pull-only commons subtree, and the nav-golden files are byte-level fixtures that gate `test_nav_golden_*` — four of them are OKF bundle-internal links, and the `/etc/passwd` ones are the negative escape fixture doing its job. Suite green: 630 passed, 4 skipped (markdown-only diff; no test touched). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ri3aVJPfynCZtHRhesCzUH
16 KiB
portfolio-optimiser
Generic, open framework on Microsoft Agent Framework (MAF): multi-agent cost-saving proposals gated by a mandatory deterministic validator, with HITL learning.
A generic, open framework — built on Microsoft Agent Framework (MAF) — that finds cost savings inside each project of a portfolio of independent projects. A swarm of agents generates candidate measures; a mandatory deterministic validator (solver + Monte Carlo) decides the numbers; domain experts judge the outcomes (human-in-the-loop); and the system learns from their verdicts across runs.
Install
Python ≥3.10, with uv. The package is not published to a package
index — install it from source:
git clone https://git.fromaitochitta.com/open/portfolio-optimiser.git
cd portfolio-optimiser
uv sync
Clone rather than install into an existing environment: the shared spec, the persona skill and the
example bundles under shared/ are read from the working tree at run time.
Verify the install by running the whole suite from the clean clone:
uv run pytest
There is no CI runner in this organization, so nothing runs that suite automatically — the command above is the verification.
Non-goals
- Not a compliance product. It ships the technical prerequisites — local-only operation, provenance on every proposal, no silent data egress — and stops there. Processing purpose, DPIA and risk assessment stay with the deploying organization.
- Not a portfolio-level reallocator. It finds savings inside each project. Moving budget between projects, ranking projects against one another and portfolio governance sit above the method and are out of scope.
- Not autonomous decision-making. The deterministic validator can only block; approving a measure is a domain expert's call (human-in-the-loop), and the framework implements nothing on the agents' say-so.
- Not a turnkey vertical solution. The aim is a generic core with explicit extension points (data sources, cost models, personas) — not the last 10% of any one domain.
- Not a model benchmark. The end-to-end proof runs offline against a scripted stand-in client: it shows that the loop closes, not how well a given LLM proposes or judges.
Status: the full 8-step agentic loop is wired and proven with load-bearing tests, and the end-to-end proof is an offline simulation with a scripted stand-in client — no live-model run yet. The ingest layer (real data sources) is implemented — file/CSV and SQL on both stacks with bit-identical golden extractions from the shared spec, plus HTTP as a MAF-only demonstrated extension point against a local mock — but exercised only against committed fixtures: no bundle has yet been materialized from a live source. A sibling implementation of the same method on the Claude Agents SDK is built in parallel from the same shared spec.
Disclaimer — technical framework only. Deploying organizations own their processing purposes and assessments (DPIA, risk/ROS, security review). The framework ships the technical prerequisites — local-only mode, provenance, no silent data egress — but makes no compliance guarantees.
Built on an LLM wiki: Karpathy's idea, Google's format
The knowledge architecture is the heart of the project, and it is deliberately not ours:
- The idea is Andrej Karpathy's "LLM wiki": instead of pointing a model at documents written for people, you curate a small, versioned body of knowledge written for the model to read — concept files, explicit structure, explicit links.
- The format is Google Cloud's Open Knowledge Format (OKF)
(open spec, v0.1), which formalizes that pattern: a knowledge bundle is a directory of
markdown files with YAML frontmatter (one required field,
type), a reservedindex.mdentry point, and intra-bundle cross-links forming an emergent graph. Custom frontmatter fields are allowed and must be preserved — which is exactly where this project's own layers (expert verdicts, ingest provenance) live.
Because OKF is open and vendor-neutral, the same bundles are consumed unchanged by both reference implementations (MAF and the Claude Agents SDK sibling) — the knowledge outlives any particular agent stack.
Not RAG. Agents read a bundle by navigating it — index.md first, then its
cross-links, with progressive disclosure — never by keyword retrieval or stuffing the whole
bundle into a prompt. Query-time retrieval against the bundle is explicitly forbidden by the
method spec: it would leak the verdict layer around the learning gate.
AI-first, humans on top
A traditional wiki is built for people — optimized for humans finding and reading information, with machine access bolted on afterwards. This project inverts that order, and is a concrete example of what that looks like:
- The wiki (the OKF bundle) is written for the model: it is the agent's working memory and the substrate the learning loop reads from and promotes into.
- The human affordances are layers on top: experts judge outcomes by dropping a plain JSON verdict file in an inbox folder; an explicit, fail-closed promotion gate is the only path by which an approved verdict becomes wiki knowledge; reports and reviews are rendered from the machine-readable layers.
Humans stay decisive — nothing enters the wiki without an approval — but the primary reader of every file is the model, not a person browsing.
How it works
One run, one project, eight steps — with the learning loop closing across runs:
- Understand — navigate the project's OKF bundle; fold the candidate's prior expert verdicts into the hypothesis prompt (ExpeL-style, retrieved structurally, never by text).
- Hypothesise — one typed candidate measure (strict IR, fail-fast schema).
- Debate — a maker-checker pair argues the reasoning (round-capped).
- Validate — two falsifiers on the same candidate: the deterministic validator gates the numbers (blocking, never optional) and the checker gates the reasoning. The validator is anchored to the project's declared cost baseline, so a proposal cannot invent the cost lines it claims to save against.
- Refine — a rejected attempt retries informed by the rejection reason, under hard attempt and token caps. Unbounded loops are forbidden everywhere.
- Propose or discard — a validated proposal with risk percentiles, or a typed rejection.
- Expert feedback — days later, an expert drops a verdict file in an inbox folder; a later run picks it up. Fully resumable; no live session assumed.
- Promote — an approved verdict is lifted into the wiki as a
type: verdictconcept file, navigable by the next run. The gate is fail-closed: raw agent output never self-promotes.
Every proposal carries provenance (citations into the bundle, model, validator decision, token usage). Every seam above is protected by a load-bearing test — a test designed to fail when the seam is detached, so the loop cannot silently degrade into theater.
How it is set up
-
One shared, framework-neutral core (
shared/, a git subtree ofportfolio-optimiser-commons): the business concept, the normative method spec and ingest spec, the expert-reviewer persona as an Agent Skill, and an example bundle with a golden suite as the only ground truth. Both stacks implement from the spec alone. -
Per project: one OKF bundle — the bundled examples are hand-curated; the ingest layer that materializes a bundle from a source (file catalogues/CSV + SQL, HTTP as a MAF-only demonstrated extension point) via a deterministic, schema-validated manifest that runs before the loop is implemented and exercised against committed fixtures — no bundle has yet been materialized from a live source.
-
Run: the
run.pyCLI has three modes — a documented partition, since one invocation cannot exercise every flag:- Single-project —
PROJECT_ID --docs-dir <dir>, plus optional--bundle-dir,--verdict-dir,--outbox-dir(which requires--run-id),--dimension-config,--semantic-retrieval,--decision/--rationale, and--live-dry-run. - Portfolio —
--portfolio, plus optional--goals,--ledger,--dimension-config,--semantic-retrieval; it stops early and prints agoal reached: …line when the accumulated ledger meets a goal. - Value report (S5.4, read-only) —
--report --ledger <file>rolls up the ledger's realized savings to stdout: per-project totals, the portfolio total, flagged cross-dimension overlaps (each counted once), and per-entry provenance. Add--jsonfor deterministic JSON instead of the human table. It makes no model calls and is mode-exclusive — only--ledger/--jsonare permitted alongside--report;--reportrequires--ledger, and a stray--jsonwithout--reportis refused (rc 1, never silently ignored).
# Single-project, offline drill (builds contracts + clients, stops before the first model call): uv run python -m portfolio_optimiser.run FV42-GSV-E1 --docs-dir <docs> --bundle-dir <bundle> --live-dry-run # Portfolio run with a savings goal checked against an accumulated ledger: uv run python -m portfolio_optimiser.run --portfolio --goals goals.json --ledger ledger.json # Read-only value report over an accumulated ledger (human table; add --json for JSON): uv run python -m portfolio_optimiser.run --report --ledger ledger.json--semantic-retrieval(S3.1) is an opt-in ranking change, off by default. Off, prior verdicts are ranked exactly as before: a structural score over the affected cost-code set, measure type and magnitude bucket, with surface text deliberately excluded. On, that score is blended with a cosine term over the same structural triple, which lets a prior verdict on a different cost-code set outrank one that ties structurally.What this ships is the seam, not better retrieval. The bundled
FakeEmbedderis a deterministic sha256 projection carrying no semantics, so over a structural tie the resulting order is deterministic but arbitrary. Retrieval quality depends entirely on injecting a real embedder —--embedder-configselects one from a closed registry (never an import path; a config file can never name arbitrary code to load), anddocs/extending.mddocuments theEmbedderprotocol. The embedding excludesdescription, matching the structural score and the verdict-id hash, so a flag-on run reads no surface text either.The flag is accepted in both run modes, but in single-project mode it requires
--bundle-dirand--verdict-dir: without them it cannot take effect, and the run is refused rather than silently ignoring the flag. Nothing about a flag-off run changes, and no savings claim depends on it.A bundle may hold verdicts about several candidates, while its
validator-input.jsondescribes only one. Atype: verdictfile therefore may declare its own retrieval key in frontmatter —affected_codes,measure_type,claimed_saving_nok— and is keyed on that; omit them and it falls back to the bundle's candidate, exactly as before. The three are all or nothing: a partial declaration is refused rather than merged with the bundle candidate, since the merge would produce a key belonging to neither.promote_verdictwrites all three, so a promoted verdict about one candidate never surfaces for another.A bundle may also ship a
cost-baseline.json— the project's actual cost lines,{code: {quantity, unit_cost}}— and when it does, the deterministic validator reconciles every affected item of a proposal against it before anything else runs. A cost code the project does not have is rejected, and so is a real code carrying a quantity or unit cost outside the configured tolerance (5% by default, relative to the baseline value). Without it, every stage of the gate reasons only about numbers the proposal supplied itself, so an internally consistent hallucination passes. The reconciliation validates; it never repairs a proposal into the baseline. A bundle that ships no baseline is simply un-anchored and runs exactly as before, while a baseline that is present but malformed is an error rather than a silent fall-back to un-anchored. On the reference-domain (non-bundle) path the project's own cost items are the baseline, so those runs are always anchored.The prior-verdict fold — the learning step — happens only on the
--bundle-dirpath; a plain--docs-dir-only run is single-shot (no fold).--decision/--rationaleapply to the single-project path only and are inert in portfolio mode.--outbox-dirmust differ from--verdict-dir: writing the raw outbox into a folder later read as an inbox would re-ingest raw agent output past the promotion gate (self-contamination) — documented here, deliberately not CLI-enforced. Stop criteria and budget caps are required at startup. Try the offline end-to-end proof (no model, no network):uv run python -m portfolio_optimiser.simulation. - Single-project —
-
A global token cap across the whole portfolio, enforced before the call. Per-run caps alone let N projects cost N times that with no ceiling over the pass. Pass a
PortfolioMeter(PortfolioBudget(max_total_tokens=…, max_tokens_per_run=…))torun_portfolioand one ledger bounds the entire pass — and, seeded frombudget.read_spend, a series of passes. It bites in three places: a remainder that cannot fund one run refuses the pass at startup (BudgetRefused); a project that cannot be funded is never started, stopping the pass structurally (budget_stop, completed runs preserved); and a chat call the remainder cannot pay for is refused rather than made (the post-charge check remains, since real usage is only knowable after the response). Spend persists viabudget.write_spend, which takes an explicit stamp and no wall-clock default, so the file is byte-deterministic. Python API only — not yet exposed on the CLI.
What this enables
The reference case is portfolio cost review (the example bundle is a building-energy measure), but the architecture is designed to generalize to any setting with the same shape — candidate measures inside independent projects, numbers a deterministic tool can check, and judgement only an expert has:
- Portfolio reviews — cost savings, energy efficiency, maintenance and procurement measures, proposed per project and validated against the project's own data.
- Compounding organizational memory — approved expert verdicts become navigable knowledge; the next run's hypotheses start from what experts actually decided, including realization gaps no solver can compute.
- Auditable AI — an unbroken provenance chain from expert decision back through proposal, bundle file and text span, and (with ingest) to the source system, query, and timestamp.
- Vendor-neutral knowledge — the same bundles drive two different agent stacks; switching frameworks does not orphan the organization's curated knowledge.
Docs
- Building a knowledge base — the team recipe (technical + domain expert) for curating a bundle, with the honest expectation that a good base takes 1–2 weeks of dedicated work.
- Target picture — the agentic loop + OKF knowledge architecture (north star).
- Prior-art & platform research (incl. implementation register §15).
- Ingest target picture — connectors and the ingest layer (frozen 2026-07-03).
Stack & develop
Python ≥3.10 · MAF via the split GA packages (see pyproject.toml) · uv. Backend profiles:
Azure/Foundry (full) + local (fallback).
uv sync
uv run pytest
uv run ruff check .