portfolio-optimiser/.claude/projects/2026-06-23-fase1-derisk-spikes/brief.md
Kjell Tore Guttormsen 8b1ff745b9 docs(fase1): Voyage brief + plan for de-risk spikes (A-D)
Planning-only artifacts (no code yet). Plan A- after adversarial review
(critic REVISE -> revised; scope MIXED -> addressed; 19 findings, 0 overlap).
Ground truth: agent-framework-orchestrations is a separate GA 1.0.0 pkg
(-> dev dep); core is 1.9.0; MAF orchestrations are async. Next: /trekexecute.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H9FyyENxebxVThjrn9et8C
2026-06-24 01:09:35 +02:00

122 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
type: trekbrief
brief_version: "2.2"
created: 2026-06-23
task: "Fase 1 de-risk spikes (A–D): empirically validate the framework's most dangerous assumptions before the MVP vertical slice"
slug: fase1-derisk-spikes
project_dir: .claude/projects/2026-06-23-fase1-derisk-spikes/
research_topics: 0
research_status: skipped
auto_research: false
interview_turns: 3
source: interview
framing: refine
phase_signals:
- phase: research
effort: standard
- phase: plan
effort: standard
- phase: execute
effort: standard
- phase: review
effort: standard
---
# Task: Fase 1 — De-risk spikes (A–D)
> Generated by `/trekbrief` on 2026-06-23.
> This brief is the contract between requirements and planning. `/trekplan`
> reads it to produce the implementation plan. Every decision in the plan must
> trace back to content in this brief.
## TL;DR
**Refine** of the locked plan ([§Fase 1](docs/plan/2026-06-23-incremental-plan.md)) — same intent, narrowed to executable spike scope. Build four **throwaway** de-risk spikes (A–D) that convert the framework's most dangerous *documented-but-unverified* assumptions into *measured facts* before the Fase 2 vertical slice. Each spike has a falsifiable pass/fail from the plan's verification block. No new external research — reuse the research report's §15 implementation register. Runtime agent calls default to the **LOCAL profile** with hard token/round caps per D6.
## Intent
Fase 0 delivered the skeleton, locked decisions (D1–D6), and a synthetic reference domain (D4). Before committing to the full architecture in Fase 2 (vertical slice), we must empirically de-risk the four assumptions that — if wrong — would force a redesign. Specifically: (1) that a Group Chat maker-checker debate beats a single-agent baseline by enough to justify its multiplicative token cost (U3 / G7); (2) that the known MAF footguns behave as the research predicts and our guards hold — Magentic's unbounded termination when `limits=None` (G1/B4) and shared-`WorkflowBuilder` state corruption in fan-out (G2/B7); (3) that a blocking deterministic hybrid-validator (B1) can *structurally* block an out-of-range proposal from ever reaching the expert; and (4) that ExpeL retrieval (B2) actually surfaces a relevant prior verdict for a similar new proposal. These are throwaway spikes — code we expect to discard — whose only job is to turn §15 register assumptions into evidence. Getting each wrong *cheaply now* is vastly cheaper than discovering it mid-Fase-2.
## Goal
Four runnable spike modules (isolated so they are trivial to discard), each producing a **measurable pass/fail** against the plan §Fase 1 verification criteria, each running on the LOCAL profile by default with hard token and round/iteration caps, and each emitting a short findings note with the measured numbers (convergence rounds, stall frequency, token use, state-bleed observations, validator block/pass behavior, retrieval hit/miss). At the end, the framework's four most dangerous assumptions are each **confirmed or refuted with evidence**, and the findings are written down so they directly inform the Fase 2 design. The repo's quality gate (ruff + mypy(src) + pytest) stays green.
## Non-Goals
- Production code for Fase 2 (the vertical slice) — these spikes are throwaway and may be deleted after their findings are recorded.
- Compliance functions (D3 — the deployer owns DPIA/ROS/behandlingsformål; we build only technical preconditions).
- Chasing the last 10% (D5) — spikes prove the generic core, not edge cases or polish.
- Heavy Foundry/Azure runs (D6) — LOCAL profile is default; Foundry is used only for targeted, minimal verification if at all in Fase 1.
- A production sandbox for user-supplied skill scripts (B8 / G3) — out of Fase 1 scope.
- A full production VerdictStore — Spike D is a minimal 10–20 synthetic-verdict retrieval test only, not the durable store design.
- Resolving the second-rank open questions (plan §Risiko: review-latency, project topology, schema ownership, Foundry-memory Preview) — these do not block Fase 1.
## Constraints
- **D6 cost-discipline:** LOCAL profile default (OpenAI-compatible endpoint); cheapest models; tiny synthetic data; hard token + round caps; no heavy test runs. Foundry/Azure only for targeted, minimal verification.
- **Deterministic validator is obligatory and blocking** (never an optional plugin) — Spike C must demonstrate *structural* blocking, not advisory warning.
- **Stop-criteria + budget caps required at startup** — fail-fast if missing; never an unbounded loop.
- **Group Chat maker-checker is the debate default**, NOT Magentic (which is experimental, G8/A2).
- **Stack:** Python ≥3.10, MAF (`agent-framework`), `uv`, `ruff`, `mypy`, `pytest`. Type hints throughout; Pydantic for IR/validation.
- Spikes reuse the existing synthetic domain (D4: `data/reference_projects.json`, `reference_domain.py`) and the backend-profile seam (D2: `backends.py`).
## Preferences
- **90%-principle (D5):** generic core + clear extension points; do not over-fit.
- Per-agent model selection via chat-client (`FoundryChatClient` / `OpenAIChatClient`, common base `BaseChatClient`).
- Keep spikes physically isolated (e.g. a `spikes/` package or `tests/spikes/`) so discarding them leaves the core untouched.
- Findings recorded as short markdown notes inside the project dir (`.claude/projects/2026-06-23-fase1-derisk-spikes/`), not scattered.
- Prefer reusing `mcp-solver` / OR-Tools / PuLP / Z3 (R1/R2) for Spike C's solver step rather than hand-rolling.
## Non-Functional Requirements
- Each spike enforces a **hard token-budget cap** and a **round/iteration cap**; exceeding either yields a structured stop event, never a silent hang.
- Each spike **reports measured token usage** (per the plan's "every spike measures and reports token consumption").
- Runtime defaults to the LOCAL profile; **no silent egress** to Azure/Foundry without explicit profile selection.
- Spikes must not require real/sensitive data — synthetic only.
## Success Criteria
*Falsifiable, mapped 1:1 to the plan §Fase 1 verification block.*
- **Spike A (U3 / G7):** Maker-checker (proposer · critic · validator) converges in ≤ N rounds (N fixed at spike-design time) AND the hard cap is respected; a findings note documents convergence rate, stall frequency, and token use for BOTH maker-checker and single-agent baseline, with an explicit cheaper/better verdict. Verify: the spike command/test exits 0 and prints the comparison table.
- **Spike B (G1 / G2):** Unbounded Magentic (`limits=None`) does NOT self-terminate → confirms an explicit limit is required (B4); `ConcurrentBuilder` fan-out from a shared builder shows **zero state-bleed** when the fresh-instance helper (B7) is used. Verify: the spike asserts both observations and exits 0.
- **Spike C (B1):** An out-of-range proposal is **structurally blocked** (never reaches the expert/output); a valid proposal passes and yields P10/P50/P90 from the Monte Carlo step; self-repair retries are capped at N. Verify: a test asserts the blocked case raises/returns a structured rejection and the valid case returns the percentiles.
- **Spike D (B2):** Top-K retrieval fetches the relevant historical verdict for a similar new proposal from a store of 10–20 synthetic verdicts. Verify: a test asserts the expected verdict id is in the top-K for a crafted similar proposal.
- **Quality gate:** `uv run ruff check .` exits 0, `uv run ruff format --check .` clean, `uv run mypy src` exits 0, `uv run pytest` exits 0.
## Research Plan
No external research needed — the codebase and the existing research report ([docs/research/2026-06-23-prior-art-platform.md](docs/research/2026-06-23-prior-art-platform.md) §13 architecture, §15 implementation register, §15.3 footguns) plus this brief contain sufficient context for planning. The spikes themselves ARE the empirical de-risking. Any MAF API specifics (e.g. the exact current-version Python Group Chat termination API, U3) are confirmed inline via the `microsoft-learn` MCP at coding time rather than as a separate research pass.
## Open Questions / Assumptions
- **[ASSUMPTION]** The LOCAL profile resolves to an OpenAI-compatible endpoint reachable during spikes (per D6 and the `local` profile in `backends.py`). If no local endpoint is available at run time, Spike A/C/D live agent calls fall back to a minimal, capped Foundry run.
- **[ASSUMPTION — verify before Spike A]** The installed `agent-framework` version's Python Group Chat termination API matches research §15 (U3: `termination_condition` lambda). Note a version ambiguity to resolve first: project `CLAUDE.md` states `agent-framework` 1.8.0 while `STATE.md` references `agent-framework-core` 1.9.0 — confirm the actual installed version (`uv pip show agent-framework-core`) and the current termination API via the `microsoft-learn` MCP before coding Spike A.
- **[ASSUMPTION]** Spike A's convergence target "N rounds" will be fixed at spike-design time (candidate ≤ 3–5 rounds); the plan deliberately leaves N open.
- Magentic is experimental (G8); Spike B exercises it ONLY to confirm the footgun (G1), never to build core flow on it (A2).
## Prior Attempts
None for the spikes themselves — fresh. Substrate from Fase 0 is complete and committed: repo scaffold, locked decisions D1–D6, GA-slimmed dependencies (`uv.lock`), the D4 synthetic "anleggskostnad" domain (`reference_domain.py` + `data/reference_projects.json`, 3 fictional projects), and the D2 backend-profile skeleton (`backends.py`: `Profile` azure|local + `ChatBackend` protocol + stubs raising `NotImplementedError` until Fase 1). Quality gate was green at Fase 0 close (ruff + mypy(src) + 12 pytest passed).
## Metadata
- **Created:** 2026-06-23
- **Interview turns:** 3
- **Auto-research opted in:** no
- **Source:** trekbrief interview
---
## How to continue
Manual (default):
```bash
# No research topics — go straight to planning:
/trekplan --project .claude/projects/2026-06-23-fase1-derisk-spikes
# Then execute:
/trekexecute --project .claude/projects/2026-06-23-fase1-derisk-spikes
```