14-step plan composing the four Fase 1 spikes into a src/ vertical slice (debate -> blocking validator -> two-layer HITL + provenance -> ExpeL learning), deterministic-core-first, on real chat clients in both profiles. Grounded in the 3 research briefs + installed-1.9.0-source introspection (7 exploration agents). Adversarial review: plan-critic REVISE (3 blockers/7 major/4 minor) -> all addressed; scope-guardian ALIGNED (0 creep, 9/9 criteria mapped, 6/6 Non-Goals honored). Key revisions: in-process retriever-as-tool MVP path (mcp dep conditional/GA-only); two-layer HITL capture + stable Verdict.id minting; extend_instructions retired by a REAL SessionContext test; TextSpan ownership + wave re-ordering; budget None-as-hard-fail with synthetic-usage test double; self_repair token-bound in the generate loop (validator.py frozen); global stop-on-failure rule. gemini-bridge Pass 2 unavailable (MCP SDK broke). plan-validator strict 0 errors. brief research_status pending to complete. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Fif1r1En5W542HbZV88yMH
20 KiB
| type | brief_version | created | task | slug | project_dir | research_topics | research_status | auto_research | interview_turns | source | framing | phase_signals | ||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| trekbrief | 2.2 | 2026-06-24 | Fase 2 MVP vertical slice — one synthetic project end-to-end on MAF 1.9.0: debate → block-validate → HITL → learn | fase2-mvp-vertical-slice | .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ | 3 | complete | false | 0 | interview | preserve |
|
Task: Fase 2 MVP vertical slice — one synthetic project end-to-end on MAF
Generated by
/trekbriefon 2026-06-24. This brief is the contract between requirements and planning./trekplanreads it to produce the implementation plan. Every decision in the plan must trace back to content in this brief.Interview note (project driftsmodell): this brief was driven from documented operator intent —
docs/plan/2026-06-23-incremental-plan.md§Fase 2,docs/research/2026-06-24-maf-capability-map.md§1+§5, Fase 1 findings, andCLAUDE.mdinvariants — not live Q&A, per the project's standing operating model (operator's only input is the trigger phrase; the session executes the next backbone step autonomously).interview_turns: 0reflects that. The operator can correct any derived decision via the annotation HTML.
TL;DR
Fase 2 preserves the locked plan, composing the verified Fase 1 spikes into one
MVP vertical slice: ONE synthetic project end-to-end on MAF 1.9.0 — debate → blocking
hybrid validator → two-layer HITL + provenance → ExpeL learning, in both profiles. MAF
GA features wired per the capability-map (adopt tokens/budget-middleware/round-caps/@tool/MCP; keep hand-rolled VerdictStore + validator; promote fresh_workflow()). Fan-out over N projects is Fase 3.
Intent
Fase 1 proved the four riskiest assumptions in isolation with throwaway spikes (maker-checker convergence, fan-out state-bleed, the hybrid validator, ExpeL retrieval) and produced a verified MAF 1.9.0 capability map that says, per need, whether to adopt a MAF feature or keep our own. We now need the first thing that is actually a system: a single project flowing through the whole method — candidate measures debated, a deterministic validator deciding the value, a domain expert judging via HITL, and the judgment fed back so the next run is smarter. This slice is the MVP backbone every later phase (generalisation in Fase 3, open-source in Fase 4, the D7 Claude-SDK sibling) builds on, so getting the seams right — token budget as middleware, fresh-workflow isolation factory, the inline validator gate, the ExpeL injection seam — matters more than breadth. The danger we are buying down is integration risk: each spike worked alone; the open question is whether they compose into one fail-fast, budget-bounded, provenance-stamped, learning loop on real (not faked) chat clients in both profiles.
Goal
A runnable vertical slice in which one synthetic "anleggskostnad" project is taken
end-to-end: load its data via a local-folder MCP source with citations, run the
debate (single-agent baseline, escalating to Group Chat maker-checker under a
deterministic termination contract with native round caps and a shared
middleware-driven token budget), pass every candidate through the obligatory
blocking hybrid validator (LLM → Pydantic IR → solver → Monte-Carlo P10/P50/P90)
that returns either a ValidatedProposal or a structural Rejection, emit exactly
one provenance-stamped saving proposal, capture a domain expert's verdict through a
two-layer HITL, persist it to the hand-rolled VerdictStore, and have the next run
retrieve that verdict via the ExpeL ContextProvider seam (with the corrected
two-arg extend_instructions(source_id, instructions)). Nothing runs without
valid, fail-fast-validated contracts (data-source schema, model-map, termination
contract, feedback schema). The slice runs on the local profile (primary) and is
verified minimally on the Azure/Foundry profile. The hand-rolled Budget/
TokenMeter shrinks to a thin shared meter that a ChatMiddleware drives off real
UsageDetails token counts; word-count proxies are gone.
Non-Goals
- Fan-out over N projects (Fase 3). Fase 2 builds
fresh_workflow()as the core isolation factory (the documented MS State-Isolation pattern) but runs ONE project through it. Orchestrated fan-out and adding a second project via config only are Fase 3. - Compliance functions (D3). No DPIA/ROS/behandlingsformål features — the deployer owns those. We ship only the technical preconditions (local-only by default, provenance, no silent egress) plus a prominent disclaimer.
- Preview MAF features on the critical path (D2/D6). mem0/redis memory, Cosmos/Redis history, and declarative YAML agent factory are Preview — kept OFF the MVP critical path. GA core only (sessions, middleware, observability, tools, MCP, checkpointing).
- Magentic orchestration. Experimental; Group Chat maker-checker is the debate default (CLAUDE.md invariant).
- Replacing the hand-rolled VerdictStore or validator with MAF memory/eval. Both MAF features are the wrong shape (bag-of-words / offline quality score) — capability-map verdict is KEEP ours.
- Chasing the last 10% (D5). Build ~90% generic core + clear extension points; competent integrators configure the last mile.
- D7 Claude Agent SDK sibling. Starts only after this slice works end-to-end.
Constraints
- Stack: Python ≥3.10,
uv,ruff(lint+format),mypy,pytest, Pydantic. - MAF:
agent-framework-core1.9.0 with GA packages pinned explicitly (NOT theagent-framework[all]meta). Promoteagent-framework-orchestrations1.0.0 from dev → core (it was throwaway-dev in Fase 1). - API truth: verified against installed 1.9.0 source via introspection; docs
via the
microsoft-learnMCP (Learn pages are NOT version-pinned — where source and docs disagree, installed source wins). - Platform: Intel mac (no Apple Silicon). CBC solver confirmed working in Fase 1
(Spike C); PuLP via
pulp[cbc]/COIN_CMD. - Two profiles, one core API (D2): local (fallback, primary for dev) + Azure/Foundry (full). Pluggable backend; same core API.
- Cost discipline (D6): develop on the local profile (free); Foundry/Azure only for targeted, minimal verification; cheapest models, tiny synthetic data, hard token/round caps. No heavy test runs.
- Determinism is blocking: the validator is obligatory and blocking — never an optional plugin; it cannot be bypassed.
- Fail-fast at startup: no run begins without all contracts validated.
- STATE.md is local-only (gitignored); the repo's private Forgejo remote is the only push target until Fase 4.
Preferences
- Adopt MAF GA features instead of reinventing (capability-map §1/§5): real
UsageDetailstoken counts (deletelen(text.split())); budget cap as a sharedChatMiddlewareshort-circuit; nativeGroupChatBuilder.with_max_rounds; GA@tool/FunctionToolwith explicitschema=;MCPStdioTool(local) /MCPStreamableHTTPTool(remote); observability viagen_ai.client.token.usage(derive cost = tokens × per-model pricing — there is no native cost metric). - Promote
fresh_workflow()(Spike B) to a core fan-out factory — it IS the documented Microsoft isolation pattern; never reuse a built workflow across tasks. - Keep hand-rolled VerdictStore (structural Jaccard over typed cost-codes +
measure-type + magnitude bucket) and validator (inline gate returning a domain
object the system branches on). Promote the
ExpeLContextProviderseam to core with the two-arg fixextend_instructions(source_id, instructions). - Debate ladder: single-agent baseline first; escalate to Group Chat maker-checker only under documented need, under a deterministic termination contract (token-disiplin).
- Method as Agent Skill: one
SKILL.md+ deterministicscripts/validator; MAF natively consumesSKILL.mdviaSkillsProvider.from_paths(...)— pin the experimental Skills surface and watch for breaks. - Roll a tiny role→deployment map (a dict/YAML → chat-client ctor); MAF declarative is Preview + not installed.
- Solver: reuse Spike C's PuLP/CBC path in the skill's
scripts/validator (proven on Intel mac) unless research surfaces a clearly better fit.
Non-Functional Requirements
- Hard caps enforced, fail-closed: a token budget overrun short-circuits via the
ChatMiddleware; a round overrun halts viawith_max_rounds. No unbounded loop can start (Fase 1 B4:max_round_count=Nonedoes NOT self-terminate). - No silent egress (D3 technical precondition): on the local profile no project data leaves the machine; any egress is explicit and configured.
- Provenance on every emitted proposal (B6): source citations + model/role + validator decision + token usage traceable on the single proposal.
- State isolation: every project run gets a fresh workflow/session — zero cross-run conversation bleed (Fase 1 B7, verified on received-message content).
- Reproducible determinism: the validator's accept/reject and the Monte-Carlo percentiles are deterministic for a fixed seed and input.
- Repo hygiene:
uv run mypy src,uv run ruff check .clean; tests pass.
Success Criteria
- Suite green:
uv run pytestexits 0 (including new Fase 2 tests);uv run ruff check .anduv run mypy srcexit 0. - End-to-end, one valid proposal: a single-command run of the synthetic project
on the local profile produces exactly one
ValidatedProposal(notRejection) carrying a populated provenance stamp (assert provenance fields: ≥1 citation, model/role, validator decision, token usage). - Structural block works: an out-of-range / infeasible candidate is blocked by
the validator and returns a
Rejection(assert type isRejection, with reason). - Verdict captured + persisted: an expert verdict entered via the two-layer HITL is written to the VerdictStore (assert a verdict record exists for the proposal).
- Learning closes the loop: a second run on a structurally similar proposal
retrieves the prior verdict via the ExpeL
ContextProvider(assert the retrieved verdict id == the one persisted), exercisingextend_instructions(source_id, …). - Fail-fast contracts: running with a malformed data-source / model-map / termination / feedback contract raises a validation error at startup BEFORE any agent call (assert it raises; assert no chat-client call was made).
- Hard-stop respected: with a tiny budget/round cap, the run terminates via the
middleware/round-cap path (assert
BudgetExceeded-equivalent or round-cap halt; assert the loop did not exceed the cap). - Real token accounting: the core token meter reads token usage from
UsageDetails(response.usage_details["total_token_count"],None-safe) — assert the meter is populated fromUsageDetailson a real run (positive, load-bearing assertion), and assert nolen(... .split())word-count token proxy is introduced intosrc/(the Fase 1 proxy lives inspikes/_harness.py;src/must stay at 0 such proxies). - Both profiles exercised: the slice runs on the local profile (full run) and a trivial agent responds on the Azure/Foundry profile (targeted minimal check) — assert a successful response from each chat-client path.
Research Plan
Three load-bearing topics the Fase 1 capability-map did NOT resolve. Each is
answerable and gates a specific set of plan steps. (research_topics: 3 in
frontmatter == the three ### Topic headings below; research_status: pending is
correct — these run via the manual /trekresearch calls in How to continue
BEFORE /trekplan; the orchestrator must not auto-advance to planning first.)
Topic 1: Native human-in-the-loop in MAF workflows
- Why this matters: the "expert verdict captured via HITL" success criterion
needs a concrete mechanism. If MAF has a native workflow pause/resume for human
input (e.g. a
RequestInfoExecutor/ request-response executor or a checkpoint-and-resume seam), the two-layer HITL should use it; if not, we capture the verdict out-of-band post-run and feed it forward via the VerdictStore. The choice changes the workflow graph and the HITL design substantially. - Research question: "Does
agent-framework1.9.0 (core + orchestrations 1.0.0) provide a native human-in-the-loop primitive to pause a workflow for external input and resume it, and how does it interact with checkpointing and session state?" - Suggested invocation:
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Does Microsoft Agent Framework 1.9.0 support native human-in-the-loop pause/resume in workflows (RequestInfoExecutor / request-response), and how does it interact with checkpointing and session state?" - Required for plan steps: HITL design; per-project workflow graph; checkpointing integration; the verdict-capture step.
- Confidence needed: high
- Estimated cost: standard
- Scope hint: both
Topic 2: Local-folder data access via MCP with citation provenance
- Why this matters: the slice needs a
local_folderdata source surfaced as an MCP tool whose retrieved chunks carry enough metadata for a citation-awareAIContextProviderand the provenance stamp. We must know whether to use an existing filesystem MCP server (e.g. the official@modelcontextprotocol/ server-filesystem) viaMCPStdioTool, or build a thin custom local-folder MCP server that returns citation-able chunks. This decides a build-vs-reuse step and the provenance data model. - Research question: "What is the best way to expose a local document folder to
a MAF agent via
MCPStdioToolsuch that retrieved content carries citation metadata (file + locator) — is there a suitable existing filesystem MCP server, or is a thin custom server required, and what citation shape does an AIContextProvider expect?" - Suggested invocation:
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Best way to expose a local document folder to a Microsoft Agent Framework agent via MCPStdioTool with citation metadata — existing filesystem MCP server vs thin custom server, and the citation shape an AIContextProvider expects?" - Required for plan steps: data-access step; citation-aware context provider; provenance-stamping; build-vs-reuse decision for the local-folder server.
- Confidence needed: high
- Estimated cost: standard
- Scope hint: both
Topic 3: Real local-profile chat client for agent-framework 1.9.0
- Why this matters: every Fase 1 spike used
FakeChatClient— no real LLM. The "both profiles" criterion needs a REAL local model on the primary (free) profile. We must know howagent-framework1.9.0 talks to a local model (an OpenAI-compatible local endpoint viaOpenAIChatClient, an Ollama integration, or another localBaseChatClient) and howUsageDetailsis populated there (some providers returnNone— which our middleware must handle). - Research question: "How does
agent-framework-core1.9.0 run agents against a local model on the free/local profile (OpenAI-compatible endpoint, Ollama, or other local chat client), and does that path populateUsageDetailstoken counts or returnNone?" - Suggested invocation:
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "How does Microsoft agent-framework-core 1.9.0 run agents against a local model (OpenAI-compatible endpoint, Ollama, or local chat client) on a free local profile, and does it populate UsageDetails token counts?" - Required for plan steps: backend-profile abstraction (local path); the budget
middleware's
None-handling; the end-to-end run; cost-discipline verification. - Confidence needed: high
- Estimated cost: standard
- Scope hint: both
Open Questions / Assumptions
- [ASSUMPTION] Solver stays PuLP/CBC in the skill's
scripts/validator (Spike C proved it on Intel mac);mcp-solveris not pursued unless Topic research shows a clear win. - [ASSUMPTION] The role→deployment model-map is a small hand-rolled dict/YAML → chat-client constructor; the available Foundry deployment names are tenant-specific (operator-supplied), not externally researchable.
- [ASSUMPTION] "1 turn = 1 round" matches our intent for
with_max_roundson the Group Chat (capability-map flagged this to confirm during implementation). - [ASSUMPTION] The synthetic reference domain from Fase 0 (small fictional "anleggskostnad" projects) is reused as the single project for this slice; no new domain data is invented.
- [OPEN] Whether checkpointing is on the Fase 2 critical path or deferred — capability-map said "ADOPT (later)"; default is defer past MVP unless Topic 1 shows HITL needs checkpoint-resume.
- [OPEN] Exact two-layer HITL semantics (synchronous review vs async + notification stub) — annenrangs risk in the plan; the notification stub (B11) is a stub only in Fase 2.
- [OPEN] Which concrete local model runs on the Intel mac for the local profile. Topic 3 settles the mechanism (how 1.9.0 talks to a local model); the concrete installed model/endpoint is machine-specific and operator-supplied. The whole free-local-profile run hinges on a working local endpoint existing.
- [OPEN] Which Foundry deployment names back the role→deployment map for the Azure/Foundry-profile check (tenant-specific, operator-supplied). The "both profiles exercised" success criterion depends on this operator input.
Prior Attempts
Fase 0 (skeleton + synthetic domain + pluggable-backend shell) and Fase 1 (four
throwaway de-risk spikes A–D) are complete and pushed. Fase 1 closed via
/trekreview round 2 = ALLOW after a real BLOCKER was found and fixed: Spike B(b)'s
fan-out "state isolation" had been confirmed with a tautological call-counter; the
rebuilt experiment measured real conversation bleed on received-message content and
confirmed the footgun, validated against Microsoft's own Workflows "State Isolation"
doc. The MAF 1.9.0 capability map (the front-end input to this brief) was then
produced from two grounded research passes (installed-source introspection + MS
Learn MCP). Key Fase 1 facts carried forward: BaseChatClient drives the GA
builders; max_round_count=None does NOT self-terminate (external guard required);
reused workflow accumulates the conversation thread (use a fresh instance per task);
ContextProvider requires source_id; extend_instructions is two-arg; real token
usage lives on UsageDetails, not word-counts.
Metadata
- Created: 2026-06-24
- Interview turns: 0 (driven from documented operator intent per project driftsmodell)
- Auto-research opted in: no
- Source: trekbrief interview
How to continue
Manual (default):
# Run each research topic (order does not matter):
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Does Microsoft Agent Framework 1.9.0 support native human-in-the-loop pause/resume in workflows (RequestInfoExecutor / request-response), and how does it interact with checkpointing and session state?"
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "Best way to expose a local document folder to a Microsoft Agent Framework agent via MCPStdioTool with citation metadata — existing filesystem MCP server vs thin custom server, and the citation shape an AIContextProvider expects?"
/trekresearch --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/ --external "How does Microsoft agent-framework-core 1.9.0 run agents against a local model (OpenAI-compatible endpoint, Ollama, or local chat client) on a free local profile, and does it populate UsageDetails token counts?"
# Then plan:
/trekplan --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/
# Then execute:
/trekexecute --project .claude/projects/2026-06-24-fase2-mvp-vertical-slice/
Auto (opt-in during /trekbrief): research and planning run
automatically; only execution is manual.