Ran `repo-standard` (v0.1.1, class `standalone`) and fixed everything it flagged as ERROR, plus the WARN links that were genuinely dead. README first screen: - opening line is now byte-identical to the forge description, so description == catalog == README is machine-checkable (badges moved below). - `## Install` (required for class `standalone`): clone + `uv sync`, stated as clone-only because the shared spec, persona skill and example bundles under `shared/` are read from the working tree at run time. `uv run pytest` named as the verification, with the fact that no CI runner exists said out loud rather than implied by a badge. - `## Non-goals` (required): the five limits already binding in CLAUDE.md — not a compliance product, not a portfolio-level reallocator, not autonomous decision-making, not turnkey, not a model benchmark. Dead relative links (measured, not guessed): - `docs/plan/2026-07-10-sesjonsplan-fase2-6.md` pointed at `../2026-07-14-revisjonspakke-DF-DI.md` six times; the file sits in `docs/plan/`, not `docs/`. (The sibling `../review-2026-07.md` links are correct and untouched.) - the Fase-1 spike brief linked repo-root-relative from `.claude/projects/…/`; re-anchored with `../../../`. The one remaining README ERROR was a gate false positive: `checkInternalLinks` resolves targets against `git ls-files`, which lists files only, so a link to a directory can never resolve. `[shared/](shared/)` now points at `shared/README.md` — a better target anyway, since that file carries the pull-only subtree rule. Not fixed here: the classifier lives in another repo. Remaining WARNs are all inside `shared/`, deliberately untouched: it is a pull-only commons subtree, and the nav-golden files are byte-level fixtures that gate `test_nav_golden_*` — four of them are OKF bundle-internal links, and the `/etc/passwd` ones are the negative escape fixture doing its job. Suite green: 630 passed, 4 skipped (markdown-only diff; no test touched). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ri3aVJPfynCZtHRhesCzUH
122 lines
10 KiB
Markdown
122 lines
10 KiB
Markdown
---
|
||
type: trekbrief
|
||
brief_version: "2.2"
|
||
created: 2026-06-23
|
||
task: "Fase 1 de-risk spikes (A–D): empirically validate the framework's most dangerous assumptions before the MVP vertical slice"
|
||
slug: fase1-derisk-spikes
|
||
project_dir: .claude/projects/2026-06-23-fase1-derisk-spikes/
|
||
research_topics: 0
|
||
research_status: skipped
|
||
auto_research: false
|
||
interview_turns: 3
|
||
source: interview
|
||
framing: refine
|
||
phase_signals:
|
||
- phase: research
|
||
effort: standard
|
||
- phase: plan
|
||
effort: standard
|
||
- phase: execute
|
||
effort: standard
|
||
- phase: review
|
||
effort: standard
|
||
---
|
||
|
||
# Task: Fase 1 — De-risk spikes (A–D)
|
||
|
||
> Generated by `/trekbrief` on 2026-06-23.
|
||
> This brief is the contract between requirements and planning. `/trekplan`
|
||
> reads it to produce the implementation plan. Every decision in the plan must
|
||
> trace back to content in this brief.
|
||
|
||
## TL;DR
|
||
|
||
**Refine** of the locked plan ([§Fase 1](../../../docs/plan/2026-06-23-incremental-plan.md)) — same intent, narrowed to executable spike scope. Build four **throwaway** de-risk spikes (A–D) that convert the framework's most dangerous *documented-but-unverified* assumptions into *measured facts* before the Fase 2 vertical slice. Each spike has a falsifiable pass/fail from the plan's verification block. No new external research — reuse the research report's §15 implementation register. Runtime agent calls default to the **LOCAL profile** with hard token/round caps per D6.
|
||
|
||
## Intent
|
||
|
||
Fase 0 delivered the skeleton, locked decisions (D1–D6), and a synthetic reference domain (D4). Before committing to the full architecture in Fase 2 (vertical slice), we must empirically de-risk the four assumptions that — if wrong — would force a redesign. Specifically: (1) that a Group Chat maker-checker debate beats a single-agent baseline by enough to justify its multiplicative token cost (U3 / G7); (2) that the known MAF footguns behave as the research predicts and our guards hold — Magentic's unbounded termination when `limits=None` (G1/B4) and shared-`WorkflowBuilder` state corruption in fan-out (G2/B7); (3) that a blocking deterministic hybrid-validator (B1) can *structurally* block an out-of-range proposal from ever reaching the expert; and (4) that ExpeL retrieval (B2) actually surfaces a relevant prior verdict for a similar new proposal. These are throwaway spikes — code we expect to discard — whose only job is to turn §15 register assumptions into evidence. Getting each wrong *cheaply now* is vastly cheaper than discovering it mid-Fase-2.
|
||
|
||
## Goal
|
||
|
||
Four runnable spike modules (isolated so they are trivial to discard), each producing a **measurable pass/fail** against the plan §Fase 1 verification criteria, each running on the LOCAL profile by default with hard token and round/iteration caps, and each emitting a short findings note with the measured numbers (convergence rounds, stall frequency, token use, state-bleed observations, validator block/pass behavior, retrieval hit/miss). At the end, the framework's four most dangerous assumptions are each **confirmed or refuted with evidence**, and the findings are written down so they directly inform the Fase 2 design. The repo's quality gate (ruff + mypy(src) + pytest) stays green.
|
||
|
||
## Non-Goals
|
||
|
||
- Production code for Fase 2 (the vertical slice) — these spikes are throwaway and may be deleted after their findings are recorded.
|
||
- Compliance functions (D3 — the deployer owns DPIA/ROS/behandlingsformål; we build only technical preconditions).
|
||
- Chasing the last 10% (D5) — spikes prove the generic core, not edge cases or polish.
|
||
- Heavy Foundry/Azure runs (D6) — LOCAL profile is default; Foundry is used only for targeted, minimal verification if at all in Fase 1.
|
||
- A production sandbox for user-supplied skill scripts (B8 / G3) — out of Fase 1 scope.
|
||
- A full production VerdictStore — Spike D is a minimal 10–20 synthetic-verdict retrieval test only, not the durable store design.
|
||
- Resolving the second-rank open questions (plan §Risiko: review-latency, project topology, schema ownership, Foundry-memory Preview) — these do not block Fase 1.
|
||
|
||
## Constraints
|
||
|
||
- **D6 cost-discipline:** LOCAL profile default (OpenAI-compatible endpoint); cheapest models; tiny synthetic data; hard token + round caps; no heavy test runs. Foundry/Azure only for targeted, minimal verification.
|
||
- **Deterministic validator is obligatory and blocking** (never an optional plugin) — Spike C must demonstrate *structural* blocking, not advisory warning.
|
||
- **Stop-criteria + budget caps required at startup** — fail-fast if missing; never an unbounded loop.
|
||
- **Group Chat maker-checker is the debate default**, NOT Magentic (which is experimental, G8/A2).
|
||
- **Stack:** Python ≥3.10, MAF (`agent-framework`), `uv`, `ruff`, `mypy`, `pytest`. Type hints throughout; Pydantic for IR/validation.
|
||
- Spikes reuse the existing synthetic domain (D4: `data/reference_projects.json`, `reference_domain.py`) and the backend-profile seam (D2: `backends.py`).
|
||
|
||
## Preferences
|
||
|
||
- **90%-principle (D5):** generic core + clear extension points; do not over-fit.
|
||
- Per-agent model selection via chat-client (`FoundryChatClient` / `OpenAIChatClient`, common base `BaseChatClient`).
|
||
- Keep spikes physically isolated (e.g. a `spikes/` package or `tests/spikes/`) so discarding them leaves the core untouched.
|
||
- Findings recorded as short markdown notes inside the project dir (`.claude/projects/2026-06-23-fase1-derisk-spikes/`), not scattered.
|
||
- Prefer reusing `mcp-solver` / OR-Tools / PuLP / Z3 (R1/R2) for Spike C's solver step rather than hand-rolling.
|
||
|
||
## Non-Functional Requirements
|
||
|
||
- Each spike enforces a **hard token-budget cap** and a **round/iteration cap**; exceeding either yields a structured stop event, never a silent hang.
|
||
- Each spike **reports measured token usage** (per the plan's "every spike measures and reports token consumption").
|
||
- Runtime defaults to the LOCAL profile; **no silent egress** to Azure/Foundry without explicit profile selection.
|
||
- Spikes must not require real/sensitive data — synthetic only.
|
||
|
||
## Success Criteria
|
||
|
||
*Falsifiable, mapped 1:1 to the plan §Fase 1 verification block.*
|
||
|
||
- **Spike A (U3 / G7):** Maker-checker (proposer · critic · validator) converges in ≤ N rounds (N fixed at spike-design time) AND the hard cap is respected; a findings note documents convergence rate, stall frequency, and token use for BOTH maker-checker and single-agent baseline, with an explicit cheaper/better verdict. Verify: the spike command/test exits 0 and prints the comparison table.
|
||
- **Spike B (G1 / G2):** Unbounded Magentic (`limits=None`) does NOT self-terminate → confirms an explicit limit is required (B4); `ConcurrentBuilder` fan-out from a shared builder shows **zero state-bleed** when the fresh-instance helper (B7) is used. Verify: the spike asserts both observations and exits 0.
|
||
- **Spike C (B1):** An out-of-range proposal is **structurally blocked** (never reaches the expert/output); a valid proposal passes and yields P10/P50/P90 from the Monte Carlo step; self-repair retries are capped at N. Verify: a test asserts the blocked case raises/returns a structured rejection and the valid case returns the percentiles.
|
||
- **Spike D (B2):** Top-K retrieval fetches the relevant historical verdict for a similar new proposal from a store of 10–20 synthetic verdicts. Verify: a test asserts the expected verdict id is in the top-K for a crafted similar proposal.
|
||
- **Quality gate:** `uv run ruff check .` exits 0, `uv run ruff format --check .` clean, `uv run mypy src` exits 0, `uv run pytest` exits 0.
|
||
|
||
## Research Plan
|
||
|
||
No external research needed — the codebase and the existing research report ([docs/research/2026-06-23-prior-art-platform.md](../../../docs/research/2026-06-23-prior-art-platform.md) §13 architecture, §15 implementation register, §15.3 footguns) plus this brief contain sufficient context for planning. The spikes themselves ARE the empirical de-risking. Any MAF API specifics (e.g. the exact current-version Python Group Chat termination API, U3) are confirmed inline via the `microsoft-learn` MCP at coding time rather than as a separate research pass.
|
||
|
||
## Open Questions / Assumptions
|
||
|
||
- **[ASSUMPTION]** The LOCAL profile resolves to an OpenAI-compatible endpoint reachable during spikes (per D6 and the `local` profile in `backends.py`). If no local endpoint is available at run time, Spike A/C/D live agent calls fall back to a minimal, capped Foundry run.
|
||
- **[ASSUMPTION — verify before Spike A]** The installed `agent-framework` version's Python Group Chat termination API matches research §15 (U3: `termination_condition` lambda). Note a version ambiguity to resolve first: project `CLAUDE.md` states `agent-framework` 1.8.0 while `STATE.md` references `agent-framework-core` 1.9.0 — confirm the actual installed version (`uv pip show agent-framework-core`) and the current termination API via the `microsoft-learn` MCP before coding Spike A.
|
||
- **[ASSUMPTION]** Spike A's convergence target "N rounds" will be fixed at spike-design time (candidate ≤ 3–5 rounds); the plan deliberately leaves N open.
|
||
- Magentic is experimental (G8); Spike B exercises it ONLY to confirm the footgun (G1), never to build core flow on it (A2).
|
||
|
||
## Prior Attempts
|
||
|
||
None for the spikes themselves — fresh. Substrate from Fase 0 is complete and committed: repo scaffold, locked decisions D1–D6, GA-slimmed dependencies (`uv.lock`), the D4 synthetic "anleggskostnad" domain (`reference_domain.py` + `data/reference_projects.json`, 3 fictional projects), and the D2 backend-profile skeleton (`backends.py`: `Profile` azure|local + `ChatBackend` protocol + stubs raising `NotImplementedError` until Fase 1). Quality gate was green at Fase 0 close (ruff + mypy(src) + 12 pytest passed).
|
||
|
||
## Metadata
|
||
|
||
- **Created:** 2026-06-23
|
||
- **Interview turns:** 3
|
||
- **Auto-research opted in:** no
|
||
- **Source:** trekbrief interview
|
||
|
||
---
|
||
|
||
## How to continue
|
||
|
||
Manual (default):
|
||
|
||
```bash
|
||
# No research topics — go straight to planning:
|
||
/trekplan --project .claude/projects/2026-06-23-fase1-derisk-spikes
|
||
|
||
# Then execute:
|
||
/trekexecute --project .claude/projects/2026-06-23-fase1-derisk-spikes
|
||
```
|