portfolio-optimiser/.claude/projects/2026-06-24-fase2-mvp-vertical-slice/brief.md
Kjell Tore Guttormsen 9973d9fe58 docs(fase2): /trekplan — adversarial-reviewed implementation plan [skip-docs]
14-step plan composing the four Fase 1 spikes into a src/ vertical slice
(debate -> blocking validator -> two-layer HITL + provenance -> ExpeL learning),
deterministic-core-first, on real chat clients in both profiles. Grounded in the
3 research briefs + installed-1.9.0-source introspection (7 exploration agents).

Adversarial review: plan-critic REVISE (3 blockers/7 major/4 minor) -> all
addressed; scope-guardian ALIGNED (0 creep, 9/9 criteria mapped, 6/6 Non-Goals
honored). Key revisions: in-process retriever-as-tool MVP path (mcp dep
conditional/GA-only); two-layer HITL capture + stable Verdict.id minting;
extend_instructions retired by a REAL SessionContext test; TextSpan ownership +
wave re-ordering; budget None-as-hard-fail with synthetic-usage test double;
self_repair token-bound in the generate loop (validator.py frozen); global
stop-on-failure rule. gemini-bridge Pass 2 unavailable (MCP SDK broke).
plan-validator strict 0 errors. brief research_status pending to complete.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fif1r1En5W542HbZV88yMH
2026-06-24 12:58:33 +02:00

20 KiB
Raw Blame History

type brief_version created task slug project_dir research_topics research_status auto_research interview_turns source framing phase_signals
trekbrief 2.2 2026-06-24 Fase 2 MVP vertical slice — one synthetic project end-to-end on MAF 1.9.0: debate → block-validate → HITL → learn fase2-mvp-vertical-slice .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ 3 complete false 0 interview preserve
phase effort model
research high opus
phase effort model
plan high opus
phase effort model
execute high opus
phase effort model
review high opus

Task: Fase 2 MVP vertical slice — one synthetic project end-to-end on MAF

Generated by /trekbrief on 2026-06-24. This brief is the contract between requirements and planning. /trekplan reads it to produce the implementation plan. Every decision in the plan must trace back to content in this brief.

Interview note (project driftsmodell): this brief was driven from documented operator intent — docs/plan/2026-06-23-incremental-plan.md §Fase 2, docs/research/2026-06-24-maf-capability-map.md §1+§5, Fase 1 findings, and CLAUDE.md invariants — not live Q&A, per the project's standing operating model (operator's only input is the trigger phrase; the session executes the next backbone step autonomously). interview_turns: 0 reflects that. The operator can correct any derived decision via the annotation HTML.

TL;DR

Fase 2 preserves the locked plan, composing the verified Fase 1 spikes into one MVP vertical slice: ONE synthetic project end-to-end on MAF 1.9.0 — debate → blocking hybrid validator → two-layer HITL + provenance → ExpeL learning, in both profiles. MAF GA features wired per the capability-map (adopt tokens/budget-middleware/round-caps/@tool/MCP; keep hand-rolled VerdictStore + validator; promote fresh_workflow()). Fan-out over N projects is Fase 3.

Intent

Fase 1 proved the four riskiest assumptions in isolation with throwaway spikes (maker-checker convergence, fan-out state-bleed, the hybrid validator, ExpeL retrieval) and produced a verified MAF 1.9.0 capability map that says, per need, whether to adopt a MAF feature or keep our own. We now need the first thing that is actually a system: a single project flowing through the whole method — candidate measures debated, a deterministic validator deciding the value, a domain expert judging via HITL, and the judgment fed back so the next run is smarter. This slice is the MVP backbone every later phase (generalisation in Fase 3, open-source in Fase 4, the D7 Claude-SDK sibling) builds on, so getting the seams right — token budget as middleware, fresh-workflow isolation factory, the inline validator gate, the ExpeL injection seam — matters more than breadth. The danger we are buying down is integration risk: each spike worked alone; the open question is whether they compose into one fail-fast, budget-bounded, provenance-stamped, learning loop on real (not faked) chat clients in both profiles.

Goal

A runnable vertical slice in which one synthetic "anleggskostnad" project is taken end-to-end: load its data via a local-folder MCP source with citations, run the debate (single-agent baseline, escalating to Group Chat maker-checker under a deterministic termination contract with native round caps and a shared middleware-driven token budget), pass every candidate through the obligatory blocking hybrid validator (LLM → Pydantic IR → solver → Monte-Carlo P10/P50/P90) that returns either a ValidatedProposal or a structural Rejection, emit exactly one provenance-stamped saving proposal, capture a domain expert's verdict through a two-layer HITL, persist it to the hand-rolled VerdictStore, and have the next run retrieve that verdict via the ExpeL ContextProvider seam (with the corrected two-arg extend_instructions(source_id, instructions)). Nothing runs without valid, fail-fast-validated contracts (data-source schema, model-map, termination contract, feedback schema). The slice runs on the local profile (primary) and is verified minimally on the Azure/Foundry profile. The hand-rolled Budget/ TokenMeter shrinks to a thin shared meter that a ChatMiddleware drives off real UsageDetails token counts; word-count proxies are gone.

Non-Goals

  • Fan-out over N projects (Fase 3). Fase 2 builds fresh_workflow() as the core isolation factory (the documented MS State-Isolation pattern) but runs ONE project through it. Orchestrated fan-out and adding a second project via config only are Fase 3.
  • Compliance functions (D3). No DPIA/ROS/behandlingsformål features — the deployer owns those. We ship only the technical preconditions (local-only by default, provenance, no silent egress) plus a prominent disclaimer.
  • Preview MAF features on the critical path (D2/D6). mem0/redis memory, Cosmos/Redis history, and declarative YAML agent factory are Preview — kept OFF the MVP critical path. GA core only (sessions, middleware, observability, tools, MCP, checkpointing).
  • Magentic orchestration. Experimental; Group Chat maker-checker is the debate default (CLAUDE.md invariant).
  • Replacing the hand-rolled VerdictStore or validator with MAF memory/eval. Both MAF features are the wrong shape (bag-of-words / offline quality score) — capability-map verdict is KEEP ours.
  • Chasing the last 10% (D5). Build ~90% generic core + clear extension points; competent integrators configure the last mile.
  • D7 Claude Agent SDK sibling. Starts only after this slice works end-to-end.

Constraints

  • Stack: Python ≥3.10, uv, ruff (lint+format), mypy, pytest, Pydantic.
  • MAF: agent-framework-core 1.9.0 with GA packages pinned explicitly (NOT the agent-framework[all] meta). Promote agent-framework-orchestrations 1.0.0 from dev → core (it was throwaway-dev in Fase 1).
  • API truth: verified against installed 1.9.0 source via introspection; docs via the microsoft-learn MCP (Learn pages are NOT version-pinned — where source and docs disagree, installed source wins).
  • Platform: Intel mac (no Apple Silicon). CBC solver confirmed working in Fase 1 (Spike C); PuLP via pulp[cbc] / COIN_CMD.
  • Two profiles, one core API (D2): local (fallback, primary for dev) + Azure/Foundry (full). Pluggable backend; same core API.
  • Cost discipline (D6): develop on the local profile (free); Foundry/Azure only for targeted, minimal verification; cheapest models, tiny synthetic data, hard token/round caps. No heavy test runs.
  • Determinism is blocking: the validator is obligatory and blocking — never an optional plugin; it cannot be bypassed.
  • Fail-fast at startup: no run begins without all contracts validated.
  • STATE.md is local-only (gitignored); the repo's private Forgejo remote is the only push target until Fase 4.

Preferences

  • Adopt MAF GA features instead of reinventing (capability-map §1/§5): real UsageDetails token counts (delete len(text.split())); budget cap as a shared ChatMiddleware short-circuit; native GroupChatBuilder.with_max_rounds; GA @tool/FunctionTool with explicit schema=; MCPStdioTool (local) / MCPStreamableHTTPTool (remote); observability via gen_ai.client.token.usage (derive cost = tokens × per-model pricing — there is no native cost metric).
  • Promote fresh_workflow() (Spike B) to a core fan-out factory — it IS the documented Microsoft isolation pattern; never reuse a built workflow across tasks.
  • Keep hand-rolled VerdictStore (structural Jaccard over typed cost-codes + measure-type + magnitude bucket) and validator (inline gate returning a domain object the system branches on). Promote the ExpeLContextProvider seam to core with the two-arg fix extend_instructions(source_id, instructions).
  • Debate ladder: single-agent baseline first; escalate to Group Chat maker-checker only under documented need, under a deterministic termination contract (token-disiplin).
  • Method as Agent Skill: one SKILL.md + deterministic scripts/ validator; MAF natively consumes SKILL.md via SkillsProvider.from_paths(...) — pin the experimental Skills surface and watch for breaks.
  • Roll a tiny role→deployment map (a dict/YAML → chat-client ctor); MAF declarative is Preview + not installed.
  • Solver: reuse Spike C's PuLP/CBC path in the skill's scripts/ validator (proven on Intel mac) unless research surfaces a clearly better fit.

Non-Functional Requirements

  • Hard caps enforced, fail-closed: a token budget overrun short-circuits via the ChatMiddleware; a round overrun halts via with_max_rounds. No unbounded loop can start (Fase 1 B4: max_round_count=None does NOT self-terminate).
  • No silent egress (D3 technical precondition): on the local profile no project data leaves the machine; any egress is explicit and configured.
  • Provenance on every emitted proposal (B6): source citations + model/role + validator decision + token usage traceable on the single proposal.
  • State isolation: every project run gets a fresh workflow/session — zero cross-run conversation bleed (Fase 1 B7, verified on received-message content).
  • Reproducible determinism: the validator's accept/reject and the Monte-Carlo percentiles are deterministic for a fixed seed and input.
  • Repo hygiene: uv run mypy src, uv run ruff check . clean; tests pass.

Success Criteria

  • Suite green: uv run pytest exits 0 (including new Fase 2 tests); uv run ruff check . and uv run mypy src exit 0.
  • End-to-end, one valid proposal: a single-command run of the synthetic project on the local profile produces exactly one ValidatedProposal (not Rejection) carrying a populated provenance stamp (assert provenance fields: ≥1 citation, model/role, validator decision, token usage).
  • Structural block works: an out-of-range / infeasible candidate is blocked by the validator and returns a Rejection (assert type is Rejection, with reason).
  • Verdict captured + persisted: an expert verdict entered via the two-layer HITL is written to the VerdictStore (assert a verdict record exists for the proposal).
  • Learning closes the loop: a second run on a structurally similar proposal retrieves the prior verdict via the ExpeL ContextProvider (assert the retrieved verdict id == the one persisted), exercising extend_instructions(source_id, …).
  • Fail-fast contracts: running with a malformed data-source / model-map / termination / feedback contract raises a validation error at startup BEFORE any agent call (assert it raises; assert no chat-client call was made).
  • Hard-stop respected: with a tiny budget/round cap, the run terminates via the middleware/round-cap path (assert BudgetExceeded-equivalent or round-cap halt; assert the loop did not exceed the cap).
  • Real token accounting: the core token meter reads token usage from UsageDetails (response.usage_details["total_token_count"], None-safe) — assert the meter is populated from UsageDetails on a real run (positive, load-bearing assertion), and assert no len(... .split()) word-count token proxy is introduced into src/ (the Fase 1 proxy lives in spikes/_harness.py; src/ must stay at 0 such proxies).
  • Both profiles exercised: the slice runs on the local profile (full run) and a trivial agent responds on the Azure/Foundry profile (targeted minimal check) — assert a successful response from each chat-client path.

Research Plan

Three load-bearing topics the Fase 1 capability-map did NOT resolve. Each is answerable and gates a specific set of plan steps. (research_topics: 3 in frontmatter == the three ### Topic headings below; research_status: pending is correct — these run via the manual /trekresearch calls in How to continue BEFORE /trekplan; the orchestrator must not auto-advance to planning first.)

Topic 1: Native human-in-the-loop in MAF workflows

  • Why this matters: the "expert verdict captured via HITL" success criterion needs a concrete mechanism. If MAF has a native workflow pause/resume for human input (e.g. a RequestInfoExecutor / request-response executor or a checkpoint-and-resume seam), the two-layer HITL should use it; if not, we capture the verdict out-of-band post-run and feed it forward via the VerdictStore. The choice changes the workflow graph and the HITL design substantially.
  • Research question: "Does agent-framework 1.9.0 (core + orchestrations 1.0.0) provide a native human-in-the-loop primitive to pause a workflow for external input and resume it, and how does it interact with checkpointing and session state?"
  • Suggested invocation: /trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Does Microsoft Agent Framework 1.9.0 support native human-in-the-loop pause/resume in workflows (RequestInfoExecutor / request-response), and how does it interact with checkpointing and session state?"
  • Required for plan steps: HITL design; per-project workflow graph; checkpointing integration; the verdict-capture step.
  • Confidence needed: high
  • Estimated cost: standard
  • Scope hint: both

Topic 2: Local-folder data access via MCP with citation provenance

  • Why this matters: the slice needs a local_folder data source surfaced as an MCP tool whose retrieved chunks carry enough metadata for a citation-aware AIContextProvider and the provenance stamp. We must know whether to use an existing filesystem MCP server (e.g. the official @modelcontextprotocol/ server-filesystem) via MCPStdioTool, or build a thin custom local-folder MCP server that returns citation-able chunks. This decides a build-vs-reuse step and the provenance data model.
  • Research question: "What is the best way to expose a local document folder to a MAF agent via MCPStdioTool such that retrieved content carries citation metadata (file + locator) — is there a suitable existing filesystem MCP server, or is a thin custom server required, and what citation shape does an AIContextProvider expect?"
  • Suggested invocation: /trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Best way to expose a local document folder to a Microsoft Agent Framework agent via MCPStdioTool with citation metadata — existing filesystem MCP server vs thin custom server, and the citation shape an AIContextProvider expects?"
  • Required for plan steps: data-access step; citation-aware context provider; provenance-stamping; build-vs-reuse decision for the local-folder server.
  • Confidence needed: high
  • Estimated cost: standard
  • Scope hint: both

Topic 3: Real local-profile chat client for agent-framework 1.9.0

  • Why this matters: every Fase 1 spike used FakeChatClient — no real LLM. The "both profiles" criterion needs a REAL local model on the primary (free) profile. We must know how agent-framework 1.9.0 talks to a local model (an OpenAI-compatible local endpoint via OpenAIChatClient, an Ollama integration, or another local BaseChatClient) and how UsageDetails is populated there (some providers return None — which our middleware must handle).
  • Research question: "How does agent-framework-core 1.9.0 run agents against a local model on the free/local profile (OpenAI-compatible endpoint, Ollama, or other local chat client), and does that path populate UsageDetails token counts or return None?"
  • Suggested invocation: /trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "How does Microsoft agent-framework-core 1.9.0 run agents against a local model (OpenAI-compatible endpoint, Ollama, or local chat client) on a free local profile, and does it populate UsageDetails token counts?"
  • Required for plan steps: backend-profile abstraction (local path); the budget middleware's None-handling; the end-to-end run; cost-discipline verification.
  • Confidence needed: high
  • Estimated cost: standard
  • Scope hint: both

Open Questions / Assumptions

  • [ASSUMPTION] Solver stays PuLP/CBC in the skill's scripts/ validator (Spike C proved it on Intel mac); mcp-solver is not pursued unless Topic research shows a clear win.
  • [ASSUMPTION] The role→deployment model-map is a small hand-rolled dict/YAML → chat-client constructor; the available Foundry deployment names are tenant-specific (operator-supplied), not externally researchable.
  • [ASSUMPTION] "1 turn = 1 round" matches our intent for with_max_rounds on the Group Chat (capability-map flagged this to confirm during implementation).
  • [ASSUMPTION] The synthetic reference domain from Fase 0 (small fictional "anleggskostnad" projects) is reused as the single project for this slice; no new domain data is invented.
  • [OPEN] Whether checkpointing is on the Fase 2 critical path or deferred — capability-map said "ADOPT (later)"; default is defer past MVP unless Topic 1 shows HITL needs checkpoint-resume.
  • [OPEN] Exact two-layer HITL semantics (synchronous review vs async + notification stub) — annenrangs risk in the plan; the notification stub (B11) is a stub only in Fase 2.
  • [OPEN] Which concrete local model runs on the Intel mac for the local profile. Topic 3 settles the mechanism (how 1.9.0 talks to a local model); the concrete installed model/endpoint is machine-specific and operator-supplied. The whole free-local-profile run hinges on a working local endpoint existing.
  • [OPEN] Which Foundry deployment names back the role→deployment map for the Azure/Foundry-profile check (tenant-specific, operator-supplied). The "both profiles exercised" success criterion depends on this operator input.

Prior Attempts

Fase 0 (skeleton + synthetic domain + pluggable-backend shell) and Fase 1 (four throwaway de-risk spikes AD) are complete and pushed. Fase 1 closed via /trekreview round 2 = ALLOW after a real BLOCKER was found and fixed: Spike B(b)'s fan-out "state isolation" had been confirmed with a tautological call-counter; the rebuilt experiment measured real conversation bleed on received-message content and confirmed the footgun, validated against Microsoft's own Workflows "State Isolation" doc. The MAF 1.9.0 capability map (the front-end input to this brief) was then produced from two grounded research passes (installed-source introspection + MS Learn MCP). Key Fase 1 facts carried forward: BaseChatClient drives the GA builders; max_round_count=None does NOT self-terminate (external guard required); reused workflow accumulates the conversation thread (use a fresh instance per task); ContextProvider requires source_id; extend_instructions is two-arg; real token usage lives on UsageDetails, not word-counts.

Metadata

  • Created: 2026-06-24
  • Interview turns: 0 (driven from documented operator intent per project driftsmodell)
  • Auto-research opted in: no
  • Source: trekbrief interview

How to continue

Manual (default):

# Run each research topic (order does not matter):
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Does Microsoft Agent Framework 1.9.0 support native human-in-the-loop pause/resume in workflows (RequestInfoExecutor / request-response), and how does it interact with checkpointing and session state?"
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Best way to expose a local document folder to a Microsoft Agent Framework agent via MCPStdioTool with citation metadata — existing filesystem MCP server vs thin custom server, and the citation shape an AIContextProvider expects?"
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "How does Microsoft agent-framework-core 1.9.0 run agents against a local model (OpenAI-compatible endpoint, Ollama, or local chat client) on a free local profile, and does it populate UsageDetails token counts?"

# Then plan:
/trekplan --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/

# Then execute:
/trekexecute --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/

Auto (opt-in during /trekbrief): research and planning run automatically; only execution is manual.