Generic, open framework on Microsoft Agent Framework (MAF): multi-agent cost-saving proposals gated by a mandatory deterministic validator, with HITL learning.
Find a file
Kjell Tore Guttormsen 93608d008d docs(readme): one coherent start-to-finish walkthrough a downloader can follow
Measured before writing: the commands existed but were scattered across the
mode partition, and there was no path a newcomer could walk end to end. The
capability gap that made a complete offline walk impossible is closed in
3abc61b; this is the door onto it in the README.

Five steps, each RUN FROM A FRESH CLONE before being written down (git clone +
uv sync + uv run pytest -> 652 passed): read the knowledge base, watch the
learning loop close, run the loop with your own scripted answers, watch the
validator say NO, and price a real run before spending anything. Step 4 is the
one that was missing entirely -- the demo only ever showed a yes, and a
refusal carries far more weight than another approval.

Also documents that `Rejection (..., decision=approved)` is not a
contradiction: the first is the validator's outcome, the second echoes the
human's recorded verdict. Documented rather than changed -- altering a public
output format is the operator's call, not a side effect of writing docs.

`costsim` gets its first mention in the README at all; it was finished, tested
and completely invisible from the surface.

Also points the commons link at open/ (published 2026-08-04). NB: the
repo-standard v0.3.0 register still lists 19 repos and does not know that repo,
so pointing at the correct live URL now trips a false LINK-DEAD. Reported to
org-ops; the URL answers HTTP 200.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GWsexbQjPo9rsV3aUE54ZS
2026-08-05 08:59:37 +02:00
.claude/projects docs(repo): meet the org repo-standard gate — 0 ERROR 2026-08-03 21:56:19 +02:00
docs Squashed 'shared/' changes from a2b57d2..ddaae5d 2026-08-05 08:24:14 +02:00
examples fix(ingest): run the MCP stdio transport against a real server, and repair its error contract (kø-x) 2026-08-03 17:56:19 +02:00
shared Merge commit 'e0fa223591' 2026-08-05 08:24:14 +02:00
spikes fix(fase1): spike B fan-out measures real conversation bleed, not a counter 2026-06-24 11:09:55 +02:00
src/portfolio_optimiser feat(run): a CLI door onto the offline whole-loop run (--scripted-replies) 2026-08-05 08:55:37 +02:00
tests feat(run): a CLI door onto the offline whole-loop run (--scripted-replies) 2026-08-05 08:55:37 +02:00
.gitignore build(s31): add numpy dependency for hybrid verdict retriever 2026-07-25 06:07:07 +02:00
.python-version feat: initial scaffold (Python framework on Microsoft Agent Framework) 2026-06-23 22:01:22 +02:00
CHANGELOG.md docs(s31): close the review's honesty gap — narrow semantic claims to the shipped mechanism 2026-07-25 13:00:13 +02:00
CLAUDE.md docs(repo): align the public contribution text with the forge's actual doors 2026-08-05 08:29:43 +02:00
CODE_OF_CONDUCT.md Squashed 'shared/' changes from a2b57d2..ddaae5d 2026-08-05 08:24:14 +02:00
CONTRIBUTING.md docs(repo): align the public contribution text with the forge's actual doors 2026-08-05 08:29:43 +02:00
env.template docs(fase2): add env.template documenting both profiles + no-egress notes 2026-06-26 00:43:26 +02:00
LICENSE Squashed 'shared/' changes from a2b57d2..ddaae5d 2026-08-05 08:24:14 +02:00
pyproject.toml fix(ingest): narrow mcp_timeout to our OWN deadline, not the exception type (kø-z follow-up) 2026-08-03 21:22:00 +02:00
README.md docs(readme): one coherent start-to-finish walkthrough a downloader can follow 2026-08-05 08:59:37 +02:00
SECURITY.md Squashed 'shared/' changes from a2b57d2..ddaae5d 2026-08-05 08:24:14 +02:00
uv.lock fix(ingest): narrow mcp_timeout to our OWN deadline, not the exception type (kø-z follow-up) 2026-08-03 21:22:00 +02:00

portfolio-optimiser

Generic, open framework on Microsoft Agent Framework (MAF): multi-agent cost-saving proposals gated by a mandatory deterministic validator, with HITL learning.

License: MIT Python Built on Microsoft Agent Framework

A generic, open framework — built on Microsoft Agent Framework (MAF) — that finds cost savings inside each project of a portfolio of independent projects. A swarm of agents generates candidate measures; a mandatory deterministic validator (solver + Monte Carlo) decides the numbers; domain experts judge the outcomes (human-in-the-loop); and the system learns from their verdicts across runs.

Install

Python ≥3.10, with uv. The package is not published to a package index — install it from source:

git clone https://git.fromaitochitta.com/open/portfolio-optimiser.git
cd portfolio-optimiser
uv sync

Clone rather than install into an existing environment: the shared spec, the persona skill and the example bundles under shared/ are read from the working tree at run time.

Verify the install by running the whole suite from the clean clone:

uv run pytest

There is no CI runner in this organization, so nothing runs that suite automatically — the command above is the verification.

Walk the whole chain offline

Five commands, no API key, no network, no cost. They exercise the real loop — context navigation over the knowledge base, the maker/checker debate, the deterministic validator, the verdict — with scripted stand-ins for the agents' answers. Every scripted invocation prints a banner saying so, because a scripted run that reads like a model run would be worse than having no offline mode at all. What this shows is that the loop closes and the gate bites; it does not show how well a given model would propose or judge.

1 — Look at the knowledge base. It is curated markdown, not a black box:

ls shared/examples/bygg-energi-mikro/

2 — Watch the learning loop close. Two runs separated by an expert approval, with the second demonstrably informed by the first:

uv run python -m portfolio_optimiser.simulation

The trace ends with the approved verdict's marker present in Run B's prompt and absent from Run A's — knowledge crossing runs purely through the file-backed wiki (promote → re-seed → fold).

3 — Run the loop over a knowledge base, with answers you supply. Write the stand-in replies, then point the CLI at the bundle:

cat > replies.json <<'JSON'
{
  "proposer": "{\"measure\":\"LED-retrofit\",\"affected_items\":[{\"code\":\"ENERGI-TOTAL-EL\",\"quantity\":300000,\"unit_cost\":1.0}],\"claimed_saving_nok\":30000}",
  "checker": "The numbers are within a feasible range. VERDICT: APPROVE"
}
JSON

uv run python -m portfolio_optimiser.run BYGG-KONTOR-NORD \
  --docs-dir shared/examples/bygg-energi-mikro \
  --bundle-dir shared/examples/bygg-energi-mikro \
  --scripted-replies replies.json

Ends in ValidatedProposal. Swap --bundle-dir/--docs-dir for your own bundle to run it over your own data — that is the point of this door, and the reason it is not the same thing as step 2.

4 — Watch it say no. Raise claimed_saving_nok to 250000 in replies.json and run the same command again. The outcome becomes Rejection: the deterministic validator refuses a saving the project's own numbers cannot support, no matter how confidently the proposer asserted it. This is the part of the method that carries the weight — the agents propose, and something that cannot be argued with decides.

Read that summary line carefully: Rejection (verdict id=…, decision=approved) is not a contradiction. Rejection is the validator's outcome, while decision= echoes the human's recorded verdict — here the --decision default, since nobody reviewed this run. The two are deliberately separate: a machine gate that blocks, and a human judgement that approves, are different questions and are never collapsed into one field.

5 — See what it would cost with a real model, before spending anything:

uv run python -m portfolio_optimiser.costsim --projects 4 --profile local

Modelled upper bounds per role and model, with the source of each price quoted. --profile local prices the free local backend; the estimate is a ceiling, not a bill.

--live-dry-run is a different, narrower drill: it builds contracts, clients and budget against your own configuration and stops before the first model call. It verifies the setup; it does not run the loop. --scripted-replies runs the whole loop. The two are mutually exclusive and passing both is refused rather than one silently winning.

Non-goals

  • Not a compliance product. It ships the technical prerequisites — local-only operation, provenance on every proposal, no silent data egress — and stops there. Processing purpose, DPIA and risk assessment stay with the deploying organization.
  • Not a portfolio-level reallocator. It finds savings inside each project. Moving budget between projects, ranking projects against one another and portfolio governance sit above the method and are out of scope.
  • Not autonomous decision-making. The deterministic validator can only block; approving a measure is a domain expert's call (human-in-the-loop), and the framework implements nothing on the agents' say-so.
  • Not a turnkey vertical solution. The aim is a generic core with explicit extension points (data sources, cost models, personas) — not the last 10% of any one domain.
  • Not a model benchmark. The end-to-end proof runs offline against a scripted stand-in client: it shows that the loop closes, not how well a given LLM proposes or judges.

Status: the full 8-step agentic loop is wired and proven with load-bearing tests, and the end-to-end proof is an offline simulation with a scripted stand-in client — no live-model run yet. The ingest layer (real data sources) is implemented — file/CSV and SQL on both stacks with bit-identical golden extractions from the shared spec, plus HTTP as a MAF-only demonstrated extension point against a local mock — but exercised only against committed fixtures: no bundle has yet been materialized from a live source. A sibling implementation of the same method on the Claude Agents SDK is built in parallel from the same shared spec.

Disclaimer — technical framework only. Deploying organizations own their processing purposes and assessments (DPIA, risk/ROS, security review). The framework ships the technical prerequisites — local-only mode, provenance, no silent data egress — but makes no compliance guarantees.

Built on an LLM wiki: Karpathy's idea, Google's format

The knowledge architecture is the heart of the project, and it is deliberately not ours:

  • The idea is Andrej Karpathy's "LLM wiki": instead of pointing a model at documents written for people, you curate a small, versioned body of knowledge written for the model to read — concept files, explicit structure, explicit links.
  • The format is Google Cloud's Open Knowledge Format (OKF) (open spec, v0.1), which formalizes that pattern: a knowledge bundle is a directory of markdown files with YAML frontmatter (one required field, type), a reserved index.md entry point, and intra-bundle cross-links forming an emergent graph. Custom frontmatter fields are allowed and must be preserved — which is exactly where this project's own layers (expert verdicts, ingest provenance) live.

Because OKF is open and vendor-neutral, the same bundles are consumed unchanged by both reference implementations (MAF and the Claude Agents SDK sibling) — the knowledge outlives any particular agent stack.

Not RAG. Agents read a bundle by navigating it — index.md first, then its cross-links, with progressive disclosure — never by keyword retrieval or stuffing the whole bundle into a prompt. Query-time retrieval against the bundle is explicitly forbidden by the method spec: it would leak the verdict layer around the learning gate.

AI-first, humans on top

A traditional wiki is built for people — optimized for humans finding and reading information, with machine access bolted on afterwards. This project inverts that order, and is a concrete example of what that looks like:

  • The wiki (the OKF bundle) is written for the model: it is the agent's working memory and the substrate the learning loop reads from and promotes into.
  • The human affordances are layers on top: experts judge outcomes by dropping a plain JSON verdict file in an inbox folder; an explicit, fail-closed promotion gate is the only path by which an approved verdict becomes wiki knowledge; reports and reviews are rendered from the machine-readable layers.

Humans stay decisive — nothing enters the wiki without an approval — but the primary reader of every file is the model, not a person browsing.

How it works

One run, one project, eight steps — with the learning loop closing across runs:

  1. Understand — navigate the project's OKF bundle; fold the candidate's prior expert verdicts into the hypothesis prompt (ExpeL-style, retrieved structurally, never by text).
  2. Hypothesise — one typed candidate measure (strict IR, fail-fast schema).
  3. Debate — a maker-checker pair argues the reasoning (round-capped).
  4. Validate — two falsifiers on the same candidate: the deterministic validator gates the numbers (blocking, never optional) and the checker gates the reasoning. The validator is anchored to the project's declared cost baseline, so a proposal cannot invent the cost lines it claims to save against.
  5. Refine — a rejected attempt retries informed by the rejection reason, under hard attempt and token caps. Unbounded loops are forbidden everywhere.
  6. Propose or discard — a validated proposal with risk percentiles, or a typed rejection.
  7. Expert feedback — days later, an expert drops a verdict file in an inbox folder; a later run picks it up. Fully resumable; no live session assumed.
  8. Promote — an approved verdict is lifted into the wiki as a type: verdict concept file, navigable by the next run. The gate is fail-closed: raw agent output never self-promotes.

Every proposal carries provenance (citations into the bundle, model, validator decision, token usage). Every seam above is protected by a load-bearing test — a test designed to fail when the seam is detached, so the loop cannot silently degrade into theater.

How it is set up

  • One shared, framework-neutral core (shared/, a git subtree of portfolio-optimiser-commons): the business concept, the normative method spec and ingest spec, the expert-reviewer persona as an Agent Skill, and an example bundle with a golden suite as the only ground truth. Both stacks implement from the spec alone.

  • Per project: one OKF bundle — the bundled examples are hand-curated; the ingest layer that materializes a bundle from a source (file catalogues/CSV + SQL, HTTP as a MAF-only demonstrated extension point) via a deterministic, schema-validated manifest that runs before the loop is implemented and exercised against committed fixtures — no bundle has yet been materialized from a live source.

  • Run: the run.py CLI has three modes — a documented partition, since one invocation cannot exercise every flag:

    • Single-projectPROJECT_ID --docs-dir <dir>, plus optional --bundle-dir, --verdict-dir, --outbox-dir (which requires --run-id), --dimension-config, --semantic-retrieval, --decision/--rationale, --live-dry-run, and --scripted-replies <file> (the offline whole-loop door — see Walk the whole chain offline; mutually exclusive with --live-dry-run, which stops before the first model call rather than answering it).
    • Portfolio--portfolio, plus optional --goals, --ledger, --dimension-config, --semantic-retrieval; it stops early and prints a goal reached: … line when the accumulated ledger meets a goal.
    • Value report (S5.4, read-only)--report --ledger <file> rolls up the ledger's realized savings to stdout: per-project totals, the portfolio total, flagged cross-dimension overlaps (each counted once), and per-entry provenance. Add --json for deterministic JSON instead of the human table. It makes no model calls and is mode-exclusive — only --ledger/--json are permitted alongside --report; --report requires --ledger, and a stray --json without --report is refused (rc 1, never silently ignored).
    # Single-project, offline drill (builds contracts + clients, stops before the first model call):
    uv run python -m portfolio_optimiser.run FV42-GSV-E1 --docs-dir <docs> --bundle-dir <bundle> --live-dry-run
    # Portfolio run with a savings goal checked against an accumulated ledger:
    uv run python -m portfolio_optimiser.run --portfolio --goals goals.json --ledger ledger.json
    # Read-only value report over an accumulated ledger (human table; add --json for JSON):
    uv run python -m portfolio_optimiser.run --report --ledger ledger.json
    

    --semantic-retrieval (S3.1) is an opt-in ranking change, off by default. Off, prior verdicts are ranked exactly as before: a structural score over the affected cost-code set, measure type and magnitude bucket, with surface text deliberately excluded. On, that score is blended with a cosine term over the same structural triple, which lets a prior verdict on a different cost-code set outrank one that ties structurally.

    What this ships is the seam, not better retrieval. The bundled FakeEmbedder is a deterministic sha256 projection carrying no semantics, so over a structural tie the resulting order is deterministic but arbitrary. Retrieval quality depends entirely on injecting a real embedder — --embedder-config selects one from a closed registry (never an import path; a config file can never name arbitrary code to load), and docs/extending.md documents the Embedder protocol. The embedding excludes description, matching the structural score and the verdict-id hash, so a flag-on run reads no surface text either.

    The flag is accepted in both run modes, but in single-project mode it requires --bundle-dir and --verdict-dir: without them it cannot take effect, and the run is refused rather than silently ignoring the flag. Nothing about a flag-off run changes, and no savings claim depends on it.

    A bundle may hold verdicts about several candidates, while its validator-input.json describes only one. A type: verdict file therefore may declare its own retrieval key in frontmatter — affected_codes, measure_type, claimed_saving_nok — and is keyed on that; omit them and it falls back to the bundle's candidate, exactly as before. The three are all or nothing: a partial declaration is refused rather than merged with the bundle candidate, since the merge would produce a key belonging to neither. promote_verdict writes all three, so a promoted verdict about one candidate never surfaces for another.

    A bundle may also ship a cost-baseline.json — the project's actual cost lines, {code: {quantity, unit_cost}} — and when it does, the deterministic validator reconciles every affected item of a proposal against it before anything else runs. A cost code the project does not have is rejected, and so is a real code carrying a quantity or unit cost outside the configured tolerance (5% by default, relative to the baseline value). Without it, every stage of the gate reasons only about numbers the proposal supplied itself, so an internally consistent hallucination passes. The reconciliation validates; it never repairs a proposal into the baseline. A bundle that ships no baseline is simply un-anchored and runs exactly as before, while a baseline that is present but malformed is an error rather than a silent fall-back to un-anchored. On the reference-domain (non-bundle) path the project's own cost items are the baseline, so those runs are always anchored.

    The prior-verdict fold — the learning step — happens only on the --bundle-dir path; a plain --docs-dir-only run is single-shot (no fold). --decision/--rationale apply to the single-project path only and are inert in portfolio mode. --outbox-dir must differ from --verdict-dir: writing the raw outbox into a folder later read as an inbox would re-ingest raw agent output past the promotion gate (self-contamination) — documented here, deliberately not CLI-enforced. Stop criteria and budget caps are required at startup. Try the offline end-to-end proof (no model, no network): uv run python -m portfolio_optimiser.simulation.

  • A global token cap across the whole portfolio, enforced before the call. Per-run caps alone let N projects cost N times that with no ceiling over the pass. Pass a PortfolioMeter(PortfolioBudget(max_total_tokens=…, max_tokens_per_run=…)) to run_portfolio and one ledger bounds the entire pass — and, seeded from budget.read_spend, a series of passes. It bites in three places: a remainder that cannot fund one run refuses the pass at startup (BudgetRefused); a project that cannot be funded is never started, stopping the pass structurally (budget_stop, completed runs preserved); and a chat call the remainder cannot pay for is refused rather than made (the post-charge check remains, since real usage is only knowable after the response). Spend persists via budget.write_spend, which takes an explicit stamp and no wall-clock default, so the file is byte-deterministic. Python API only — not yet exposed on the CLI.

What this enables

The reference case is portfolio cost review (the example bundle is a building-energy measure), but the architecture is designed to generalize to any setting with the same shape — candidate measures inside independent projects, numbers a deterministic tool can check, and judgement only an expert has:

  • Portfolio reviews — cost savings, energy efficiency, maintenance and procurement measures, proposed per project and validated against the project's own data.
  • Compounding organizational memory — approved expert verdicts become navigable knowledge; the next run's hypotheses start from what experts actually decided, including realization gaps no solver can compute.
  • Auditable AI — an unbroken provenance chain from expert decision back through proposal, bundle file and text span, and (with ingest) to the source system, query, and timestamp.
  • Vendor-neutral knowledge — the same bundles drive two different agent stacks; switching frameworks does not orphan the organization's curated knowledge.

Docs

Stack & develop

Python ≥3.10 · MAF via the split GA packages (see pyproject.toml) · uv. Backend profiles: Azure/Foundry (full) + local (fallback).

uv sync
uv run pytest
uv run ruff check .