Generic, open framework on Microsoft Agent Framework (MAF): multi-agent cost-saving proposals gated by a mandatory deterministic validator, with HITL learning.
Find a file
Kjell Tore Guttormsen 392f8493da chore(repo): planning artifacts become local-only; fixture builders become code
Operator ruling 2026-08-05, which settles decision (g): planning documents are
generally never public, and what OUR OWN sessions generate does not go out on
the forge at all. The example itself stays public so others can run the
process.

`.claude/projects/` is the Voyage session workbench -- 25 briefs/plans/reviews
this project's own sessions produced. Untracked and gitignored, exactly as
STATE.md already is, and for the same stated reason: this repo has a public
mirror, so that class of material is local-only rather than tracked.

The line is drawn at who wrote the document, and it is drawn deliberately:
`docs/plan/`, `docs/research/` and `docs/rapport/` stay tracked. Those are
curated, dated documents written for the repo's readers, three of them linked
from the README as the decision record. Move that line if it was meant wider.

Two files were NOT process artifacts and are not deleted. Both
`build_fixture.py` scripts are cited by tracked tests
(`test_ingest_golden_sql.py`, `test_ingest_golden_http.py`) as the documented
rebuild path for byte-exact goldens -- reproduction code that had landed in the
wrong directory. Moved next to the goldens they build; both docstrings updated,
so no tracked file is left pointing into an untracked tree (verified: the only
remaining `.claude/projects` string in a tracked file is the .gitignore rule
itself). One prose reference in the dated Foundry auth recipe was dropped for
the same reason.

652 tests still pass.

Does NOT address the 27 of these already readable on open/ since the S12
release -- untracking stops future publication only. That retraction is a
separate operator decision and is deliberately not taken here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GWsexbQjPo9rsV3aUE54ZS
2026-08-05 10:08:17 +02:00
docs chore(repo): planning artifacts become local-only; fixture builders become code 2026-08-05 10:08:17 +02:00
examples chore(repo): planning artifacts become local-only; fixture builders become code 2026-08-05 10:08:17 +02:00
shared Merge commit 'e0fa223591' 2026-08-05 08:24:14 +02:00
spikes fix(fase1): spike B fan-out measures real conversation bleed, not a counter 2026-06-24 11:09:55 +02:00
src/portfolio_optimiser feat(run): a CLI door onto the offline whole-loop run (--scripted-replies) 2026-08-05 08:55:37 +02:00
tests chore(repo): planning artifacts become local-only; fixture builders become code 2026-08-05 10:08:17 +02:00
.gitignore chore(repo): planning artifacts become local-only; fixture builders become code 2026-08-05 10:08:17 +02:00
.python-version feat: initial scaffold (Python framework on Microsoft Agent Framework) 2026-06-23 22:01:22 +02:00
CHANGELOG.md docs(s31): close the review's honesty gap — narrow semantic claims to the shipped mechanism 2026-07-25 13:00:13 +02:00
CLAUDE.md docs(repo): align the public contribution text with the forge's actual doors 2026-08-05 08:29:43 +02:00
CODE_OF_CONDUCT.md Squashed 'shared/' changes from a2b57d2..ddaae5d 2026-08-05 08:24:14 +02:00
CONTRIBUTING.md docs(repo): align the public contribution text with the forge's actual doors 2026-08-05 08:29:43 +02:00
env.template docs(fase2): add env.template documenting both profiles + no-egress notes 2026-06-26 00:43:26 +02:00
LICENSE Squashed 'shared/' changes from a2b57d2..ddaae5d 2026-08-05 08:24:14 +02:00
pyproject.toml fix(ingest): narrow mcp_timeout to our OWN deadline, not the exception type (kø-z follow-up) 2026-08-03 21:22:00 +02:00
README.md docs(readme): one coherent start-to-finish walkthrough a downloader can follow 2026-08-05 08:59:37 +02:00
SECURITY.md Squashed 'shared/' changes from a2b57d2..ddaae5d 2026-08-05 08:24:14 +02:00
uv.lock fix(ingest): narrow mcp_timeout to our OWN deadline, not the exception type (kø-z follow-up) 2026-08-03 21:22:00 +02:00

portfolio-optimiser

Generic, open framework on Microsoft Agent Framework (MAF): multi-agent cost-saving proposals gated by a mandatory deterministic validator, with HITL learning.

License: MIT Python Built on Microsoft Agent Framework

A generic, open framework — built on Microsoft Agent Framework (MAF) — that finds cost savings inside each project of a portfolio of independent projects. A swarm of agents generates candidate measures; a mandatory deterministic validator (solver + Monte Carlo) decides the numbers; domain experts judge the outcomes (human-in-the-loop); and the system learns from their verdicts across runs.

Install

Python ≥3.10, with uv. The package is not published to a package index — install it from source:

git clone https://git.fromaitochitta.com/open/portfolio-optimiser.git
cd portfolio-optimiser
uv sync

Clone rather than install into an existing environment: the shared spec, the persona skill and the example bundles under shared/ are read from the working tree at run time.

Verify the install by running the whole suite from the clean clone:

uv run pytest

There is no CI runner in this organization, so nothing runs that suite automatically — the command above is the verification.

Walk the whole chain offline

Five commands, no API key, no network, no cost. They exercise the real loop — context navigation over the knowledge base, the maker/checker debate, the deterministic validator, the verdict — with scripted stand-ins for the agents' answers. Every scripted invocation prints a banner saying so, because a scripted run that reads like a model run would be worse than having no offline mode at all. What this shows is that the loop closes and the gate bites; it does not show how well a given model would propose or judge.

1 — Look at the knowledge base. It is curated markdown, not a black box:

ls shared/examples/bygg-energi-mikro/

2 — Watch the learning loop close. Two runs separated by an expert approval, with the second demonstrably informed by the first:

uv run python -m portfolio_optimiser.simulation

The trace ends with the approved verdict's marker present in Run B's prompt and absent from Run A's — knowledge crossing runs purely through the file-backed wiki (promote → re-seed → fold).

3 — Run the loop over a knowledge base, with answers you supply. Write the stand-in replies, then point the CLI at the bundle:

cat > replies.json <<'JSON'
{
  "proposer": "{\"measure\":\"LED-retrofit\",\"affected_items\":[{\"code\":\"ENERGI-TOTAL-EL\",\"quantity\":300000,\"unit_cost\":1.0}],\"claimed_saving_nok\":30000}",
  "checker": "The numbers are within a feasible range. VERDICT: APPROVE"
}
JSON

uv run python -m portfolio_optimiser.run BYGG-KONTOR-NORD \
  --docs-dir shared/examples/bygg-energi-mikro \
  --bundle-dir shared/examples/bygg-energi-mikro \
  --scripted-replies replies.json

Ends in ValidatedProposal. Swap --bundle-dir/--docs-dir for your own bundle to run it over your own data — that is the point of this door, and the reason it is not the same thing as step 2.

4 — Watch it say no. Raise claimed_saving_nok to 250000 in replies.json and run the same command again. The outcome becomes Rejection: the deterministic validator refuses a saving the project's own numbers cannot support, no matter how confidently the proposer asserted it. This is the part of the method that carries the weight — the agents propose, and something that cannot be argued with decides.

Read that summary line carefully: Rejection (verdict id=…, decision=approved) is not a contradiction. Rejection is the validator's outcome, while decision= echoes the human's recorded verdict — here the --decision default, since nobody reviewed this run. The two are deliberately separate: a machine gate that blocks, and a human judgement that approves, are different questions and are never collapsed into one field.

5 — See what it would cost with a real model, before spending anything:

uv run python -m portfolio_optimiser.costsim --projects 4 --profile local

Modelled upper bounds per role and model, with the source of each price quoted. --profile local prices the free local backend; the estimate is a ceiling, not a bill.

--live-dry-run is a different, narrower drill: it builds contracts, clients and budget against your own configuration and stops before the first model call. It verifies the setup; it does not run the loop. --scripted-replies runs the whole loop. The two are mutually exclusive and passing both is refused rather than one silently winning.

Non-goals

  • Not a compliance product. It ships the technical prerequisites — local-only operation, provenance on every proposal, no silent data egress — and stops there. Processing purpose, DPIA and risk assessment stay with the deploying organization.
  • Not a portfolio-level reallocator. It finds savings inside each project. Moving budget between projects, ranking projects against one another and portfolio governance sit above the method and are out of scope.
  • Not autonomous decision-making. The deterministic validator can only block; approving a measure is a domain expert's call (human-in-the-loop), and the framework implements nothing on the agents' say-so.
  • Not a turnkey vertical solution. The aim is a generic core with explicit extension points (data sources, cost models, personas) — not the last 10% of any one domain.
  • Not a model benchmark. The end-to-end proof runs offline against a scripted stand-in client: it shows that the loop closes, not how well a given LLM proposes or judges.

Status: the full 8-step agentic loop is wired and proven with load-bearing tests, and the end-to-end proof is an offline simulation with a scripted stand-in client — no live-model run yet. The ingest layer (real data sources) is implemented — file/CSV and SQL on both stacks with bit-identical golden extractions from the shared spec, plus HTTP as a MAF-only demonstrated extension point against a local mock — but exercised only against committed fixtures: no bundle has yet been materialized from a live source. A sibling implementation of the same method on the Claude Agents SDK is built in parallel from the same shared spec.

Disclaimer — technical framework only. Deploying organizations own their processing purposes and assessments (DPIA, risk/ROS, security review). The framework ships the technical prerequisites — local-only mode, provenance, no silent data egress — but makes no compliance guarantees.

Built on an LLM wiki: Karpathy's idea, Google's format

The knowledge architecture is the heart of the project, and it is deliberately not ours:

  • The idea is Andrej Karpathy's "LLM wiki": instead of pointing a model at documents written for people, you curate a small, versioned body of knowledge written for the model to read — concept files, explicit structure, explicit links.
  • The format is Google Cloud's Open Knowledge Format (OKF) (open spec, v0.1), which formalizes that pattern: a knowledge bundle is a directory of markdown files with YAML frontmatter (one required field, type), a reserved index.md entry point, and intra-bundle cross-links forming an emergent graph. Custom frontmatter fields are allowed and must be preserved — which is exactly where this project's own layers (expert verdicts, ingest provenance) live.

Because OKF is open and vendor-neutral, the same bundles are consumed unchanged by both reference implementations (MAF and the Claude Agents SDK sibling) — the knowledge outlives any particular agent stack.

Not RAG. Agents read a bundle by navigating it — index.md first, then its cross-links, with progressive disclosure — never by keyword retrieval or stuffing the whole bundle into a prompt. Query-time retrieval against the bundle is explicitly forbidden by the method spec: it would leak the verdict layer around the learning gate.

AI-first, humans on top

A traditional wiki is built for people — optimized for humans finding and reading information, with machine access bolted on afterwards. This project inverts that order, and is a concrete example of what that looks like:

  • The wiki (the OKF bundle) is written for the model: it is the agent's working memory and the substrate the learning loop reads from and promotes into.
  • The human affordances are layers on top: experts judge outcomes by dropping a plain JSON verdict file in an inbox folder; an explicit, fail-closed promotion gate is the only path by which an approved verdict becomes wiki knowledge; reports and reviews are rendered from the machine-readable layers.

Humans stay decisive — nothing enters the wiki without an approval — but the primary reader of every file is the model, not a person browsing.

How it works

One run, one project, eight steps — with the learning loop closing across runs:

  1. Understand — navigate the project's OKF bundle; fold the candidate's prior expert verdicts into the hypothesis prompt (ExpeL-style, retrieved structurally, never by text).
  2. Hypothesise — one typed candidate measure (strict IR, fail-fast schema).
  3. Debate — a maker-checker pair argues the reasoning (round-capped).
  4. Validate — two falsifiers on the same candidate: the deterministic validator gates the numbers (blocking, never optional) and the checker gates the reasoning. The validator is anchored to the project's declared cost baseline, so a proposal cannot invent the cost lines it claims to save against.
  5. Refine — a rejected attempt retries informed by the rejection reason, under hard attempt and token caps. Unbounded loops are forbidden everywhere.
  6. Propose or discard — a validated proposal with risk percentiles, or a typed rejection.
  7. Expert feedback — days later, an expert drops a verdict file in an inbox folder; a later run picks it up. Fully resumable; no live session assumed.
  8. Promote — an approved verdict is lifted into the wiki as a type: verdict concept file, navigable by the next run. The gate is fail-closed: raw agent output never self-promotes.

Every proposal carries provenance (citations into the bundle, model, validator decision, token usage). Every seam above is protected by a load-bearing test — a test designed to fail when the seam is detached, so the loop cannot silently degrade into theater.

How it is set up

  • One shared, framework-neutral core (shared/, a git subtree of portfolio-optimiser-commons): the business concept, the normative method spec and ingest spec, the expert-reviewer persona as an Agent Skill, and an example bundle with a golden suite as the only ground truth. Both stacks implement from the spec alone.

  • Per project: one OKF bundle — the bundled examples are hand-curated; the ingest layer that materializes a bundle from a source (file catalogues/CSV + SQL, HTTP as a MAF-only demonstrated extension point) via a deterministic, schema-validated manifest that runs before the loop is implemented and exercised against committed fixtures — no bundle has yet been materialized from a live source.

  • Run: the run.py CLI has three modes — a documented partition, since one invocation cannot exercise every flag:

    • Single-projectPROJECT_ID --docs-dir <dir>, plus optional --bundle-dir, --verdict-dir, --outbox-dir (which requires --run-id), --dimension-config, --semantic-retrieval, --decision/--rationale, --live-dry-run, and --scripted-replies <file> (the offline whole-loop door — see Walk the whole chain offline; mutually exclusive with --live-dry-run, which stops before the first model call rather than answering it).
    • Portfolio--portfolio, plus optional --goals, --ledger, --dimension-config, --semantic-retrieval; it stops early and prints a goal reached: … line when the accumulated ledger meets a goal.
    • Value report (S5.4, read-only)--report --ledger <file> rolls up the ledger's realized savings to stdout: per-project totals, the portfolio total, flagged cross-dimension overlaps (each counted once), and per-entry provenance. Add --json for deterministic JSON instead of the human table. It makes no model calls and is mode-exclusive — only --ledger/--json are permitted alongside --report; --report requires --ledger, and a stray --json without --report is refused (rc 1, never silently ignored).
    # Single-project, offline drill (builds contracts + clients, stops before the first model call):
    uv run python -m portfolio_optimiser.run FV42-GSV-E1 --docs-dir <docs> --bundle-dir <bundle> --live-dry-run
    # Portfolio run with a savings goal checked against an accumulated ledger:
    uv run python -m portfolio_optimiser.run --portfolio --goals goals.json --ledger ledger.json
    # Read-only value report over an accumulated ledger (human table; add --json for JSON):
    uv run python -m portfolio_optimiser.run --report --ledger ledger.json
    

    --semantic-retrieval (S3.1) is an opt-in ranking change, off by default. Off, prior verdicts are ranked exactly as before: a structural score over the affected cost-code set, measure type and magnitude bucket, with surface text deliberately excluded. On, that score is blended with a cosine term over the same structural triple, which lets a prior verdict on a different cost-code set outrank one that ties structurally.

    What this ships is the seam, not better retrieval. The bundled FakeEmbedder is a deterministic sha256 projection carrying no semantics, so over a structural tie the resulting order is deterministic but arbitrary. Retrieval quality depends entirely on injecting a real embedder — --embedder-config selects one from a closed registry (never an import path; a config file can never name arbitrary code to load), and docs/extending.md documents the Embedder protocol. The embedding excludes description, matching the structural score and the verdict-id hash, so a flag-on run reads no surface text either.

    The flag is accepted in both run modes, but in single-project mode it requires --bundle-dir and --verdict-dir: without them it cannot take effect, and the run is refused rather than silently ignoring the flag. Nothing about a flag-off run changes, and no savings claim depends on it.

    A bundle may hold verdicts about several candidates, while its validator-input.json describes only one. A type: verdict file therefore may declare its own retrieval key in frontmatter — affected_codes, measure_type, claimed_saving_nok — and is keyed on that; omit them and it falls back to the bundle's candidate, exactly as before. The three are all or nothing: a partial declaration is refused rather than merged with the bundle candidate, since the merge would produce a key belonging to neither. promote_verdict writes all three, so a promoted verdict about one candidate never surfaces for another.

    A bundle may also ship a cost-baseline.json — the project's actual cost lines, {code: {quantity, unit_cost}} — and when it does, the deterministic validator reconciles every affected item of a proposal against it before anything else runs. A cost code the project does not have is rejected, and so is a real code carrying a quantity or unit cost outside the configured tolerance (5% by default, relative to the baseline value). Without it, every stage of the gate reasons only about numbers the proposal supplied itself, so an internally consistent hallucination passes. The reconciliation validates; it never repairs a proposal into the baseline. A bundle that ships no baseline is simply un-anchored and runs exactly as before, while a baseline that is present but malformed is an error rather than a silent fall-back to un-anchored. On the reference-domain (non-bundle) path the project's own cost items are the baseline, so those runs are always anchored.

    The prior-verdict fold — the learning step — happens only on the --bundle-dir path; a plain --docs-dir-only run is single-shot (no fold). --decision/--rationale apply to the single-project path only and are inert in portfolio mode. --outbox-dir must differ from --verdict-dir: writing the raw outbox into a folder later read as an inbox would re-ingest raw agent output past the promotion gate (self-contamination) — documented here, deliberately not CLI-enforced. Stop criteria and budget caps are required at startup. Try the offline end-to-end proof (no model, no network): uv run python -m portfolio_optimiser.simulation.

  • A global token cap across the whole portfolio, enforced before the call. Per-run caps alone let N projects cost N times that with no ceiling over the pass. Pass a PortfolioMeter(PortfolioBudget(max_total_tokens=…, max_tokens_per_run=…)) to run_portfolio and one ledger bounds the entire pass — and, seeded from budget.read_spend, a series of passes. It bites in three places: a remainder that cannot fund one run refuses the pass at startup (BudgetRefused); a project that cannot be funded is never started, stopping the pass structurally (budget_stop, completed runs preserved); and a chat call the remainder cannot pay for is refused rather than made (the post-charge check remains, since real usage is only knowable after the response). Spend persists via budget.write_spend, which takes an explicit stamp and no wall-clock default, so the file is byte-deterministic. Python API only — not yet exposed on the CLI.

What this enables

The reference case is portfolio cost review (the example bundle is a building-energy measure), but the architecture is designed to generalize to any setting with the same shape — candidate measures inside independent projects, numbers a deterministic tool can check, and judgement only an expert has:

  • Portfolio reviews — cost savings, energy efficiency, maintenance and procurement measures, proposed per project and validated against the project's own data.
  • Compounding organizational memory — approved expert verdicts become navigable knowledge; the next run's hypotheses start from what experts actually decided, including realization gaps no solver can compute.
  • Auditable AI — an unbroken provenance chain from expert decision back through proposal, bundle file and text span, and (with ingest) to the source system, query, and timestamp.
  • Vendor-neutral knowledge — the same bundles drive two different agent stacks; switching frameworks does not orphan the organization's curated knowledge.

Docs

Stack & develop

Python ≥3.10 · MAF via the split GA packages (see pyproject.toml) · uv. Backend profiles: Azure/Foundry (full) + local (fallback).

uv sync
uv run pytest
uv run ruff check .