portfolio-optimiser/.claude/projects/2026-06-23-fase1-derisk-spikes/plan.md
Kjell Tore Guttormsen 8b1ff745b9 docs(fase1): Voyage brief + plan for de-risk spikes (A-D)
Planning-only artifacts (no code yet). Plan A- after adversarial review
(critic REVISE -> revised; scope MIXED -> addressed; 19 findings, 0 overlap).
Ground truth: agent-framework-orchestrations is a separate GA 1.0.0 pkg
(-> dev dep); core is 1.9.0; MAF orchestrations are async. Next: /trekexecute.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H9FyyENxebxVThjrn9et8C
2026-06-24 01:09:35 +02:00

473 lines
41 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Fase 1 — De-risk spikes (A–D)
> **Plan quality: A-** (post adversarial review + revision) — APPROVE_WITH_NOTES
>
> Generated by trekplan v2.0 on 2026-06-23 — `plan_version: 1.7`
## Context
Fase 0 delivered the skeleton, locked decisions (D1–D6), and a synthetic reference domain (D4). Before committing to the full architecture in Fase 2, we empirically de-risk the four assumptions that — if wrong — force a redesign (brief Intent): (A) Group Chat maker-checker beats a single-agent baseline by enough to justify its multiplicative token cost (U3 / G7); (B) the known MAF footguns behave as predicted and our guards hold — Magentic unbounded termination when `limits=None` (G1/B4) and shared-builder/fan-out state corruption (G2/B7); (C) a blocking deterministic hybrid-validator (B1) can *structurally* block an out-of-range proposal; (D) ExpeL retrieval (B2) surfaces a relevant prior verdict. These are **throwaway** spikes whose only job is to turn §15-register assumptions into measured facts, cheaply, under cost-discipline (D6: LOCAL profile default, hard token/round caps).
**Decisive planning finding (premise-verification):** the GA-slimmed Fase 0 install deliberately dropped the `agent-framework[all]` meta, so the orchestration builders (`GroupChatBuilder`, `ConcurrentBuilder`, `MagenticBuilder`) are **not importable** — they live in the separate package `agent-framework-orchestrations`, which has a clean **GA release `1.0.0`** depending on `agent-framework-core<2,>=1.9.0` (exactly our installed core). Because the spikes are **throwaway**, this package is added to the **`dev`** dependency group (not core) — exactly like the PuLP solver — deferring any core-runtime commitment to Fase 2 (revised after scope review). Also resolved: the installed `agent-framework-core` is **1.9.0** (STATE.md correct; project `CLAUDE.md` says 1.8.0 and is corrected in Step 1; the `project-core` user-memory also says 1.8.0 — that lives outside the repo and is a session-end follow-up).
**Gate-green contract (the central design choice).** Each spike is split into a **logic layer** (pure functions/classes we author — always exercised by the quality gate, no live endpoint and no full MAF workflow run) and an **integration/live layer** (drives the actual MAF builders and/or a real LOCAL LLM — runs when available, otherwise `skip`s, except Spike B where driving the builders IS the de-risk; see Step 2/Step 4). The headline empirical claims (Spike A's better/cheaper verdict; live token numbers) live in the integration/live layer and are honestly reported as endpoint-dependent. The gate stays green from the logic layer alone.
## Architecture Diagram
```mermaid
graph TD
subgraph "Reused (Fase 0, unchanged — src/ never touched)"
RD["reference_domain.py<br/>(D4 synthetic projects)"]
BC["agent_framework.BaseChatClient<br/>(subclassed by FakeChatClient)"]
end
subgraph "New — throwaway spikes/ package (dev-only)"
H["_harness.py<br/>Budget+RoundCap (B4)<br/>FakeChatClient + minimal Agent<br/>live_local_client_or_skip() -> OpenAIChatClient"]
SM["builder smoke (Step 2)<br/>proves FakeChatClient drives the GA builders"]
A["a_groupchat.py — Spike A (U3/G7)"]
B["b_footguns.py — Spike B (G1, G2/B7)"]
C["c_validator.py — Spike C (B1)"]
D["d_verdictstore.py — Spike D (B2)"]
end
subgraph "Deliverables (tracked)"
F["docs/fase1-spikes/<br/>README + findings-{a,b,c,d}.md"]
end
H --> SM --> A & B & C & D
RD --> A & C & D
BC --> H
A --> F
B --> F
C --> F
D --> F
```
## Codebase Analysis
- **Tech stack:** Python ≥3.10 (venv on 3.12), MAF GA (`agent-framework-core` 1.9.0, `-foundry` 1.8.2, `-openai` 1.8.2), `pydantic` 2.13.4 (stable), `uv`, `ruff` 0.15.18, `mypy` 2.1.0, `pytest` 9.1.1. Build: hatchling (wheel packages `src/portfolio_optimiser` only — top-level `spikes/` is never shipped).
- **Key patterns:** small typed package; `from __future__ import annotations`; frozen dataclasses for the domain; `Protocol` + `runtime_checkable` for the D2 seam; fail-fast (`ValueError`/`NotImplementedError`); tests are plain pytest functions with `-> None`, `parametrize`, `pytest.raises`, module-scope fixtures, importing from `portfolio_optimiser`.
- **Relevant files:** `src/portfolio_optimiser/backends.py` (D2 seam — `Profile`, `ChatBackend`, `get_backend`, `LocalBackend`; live-wiring of the seam is a **Fase 2** concern, NOT touched here), `src/portfolio_optimiser/reference_domain.py` (D4 — `Project`, `CostItem`, `load_reference_projects`, `.total_cost`), `src/portfolio_optimiser/data/reference_projects.json` (3 synthetic projects), `tests/test_*.py` (test style), `pyproject.toml`.
- **Reusable code:** `load_reference_projects()` → spike fixtures (A/C/D); `Project.total_cost` / `CostItem.total_cost` → savings-IR constraint source (C); `agent_framework.BaseChatClient` → subclassed by the harness `FakeChatClient`; `agent_framework.openai.OpenAIChatClient` → built **directly** by the harness for the live arm (the D2 `LocalBackend` seam is intentionally left un-wired until Fase 2, so `src/` stays untouched and the throwaway spikes add no core surface).
- **External tech (verified by introspection + PyPI, not docs):** orchestration builders in `agent_framework.orchestrations` (need `agent-framework-orchestrations`); chat clients in `agent_framework.openai` (`OpenAIChatClient`); `ContextProvider`/`MemoryContextProvider` top-level. **MAF workflows/orchestrations are async** (`await workflow.run(...)`), so the spike tests use `pytest-asyncio` (added in Step 1). Exact builder method signatures are confirmed by introspection at coding time (brief-sanctioned) and de-risked once up-front by the Step 2 builder smoke.
- **Recent git activity:** `b57aa83` Fase 0 (D4+D2+tests), `491a746` deps GA-fix + lockfile. Clean tree on `main`, synced to origin.
## Research Sources
*Ground-truth API/dependency verification — introspection + PyPI metadata are authoritative here (no web docs needed).*
| Topic | Source | Key Findings | Confidence |
|-------|--------|--------------|------------|
| Orchestration builders' real import path | `import agent_framework.orchestrations` (introspection) | `GroupChatBuilder`, `ConcurrentBuilder`, `MagenticBuilder`, `StandardMagenticManager`, `SequentialBuilder`, `TerminationCondition`, `MagenticPlanReviewRequest` live there; not top-level | high |
| Orchestrations package availability/stability | PyPI JSON | GA `1.0.0` (uploaded 2026-06-18, not yanked); `requires_dist: agent-framework-core<2,>=1.9.0`; rest b/rc | high |
| Installed core version (resolves brief ambiguity) | `uv pip show agent-framework-core` | **1.9.0** (STATE correct; CLAUDE.md 1.8.0 → corrected Step 1) | high |
| MAF orchestrations are async | introspection (`WorkflowBuilder`/orchestration `.run`) | tests need `pytest-asyncio`; integration arm is async | high |
| Solver lib for Spike C | PyPI / project knowledge | `pulp` (PuLP) MIT, pure-Python, **bundles a CBC binary in its wheel** → faithful solver-in-the-loop (R2) | high |
## Implementation Plan
Each step targets one focused change and follows the repo's TDD-ish style. Per the **gate-green contract** above, the deterministic logic layer is always exercised; integration/live layers are gated.
### Step 1: Scaffold throwaway spikes package and pin dev-only orchestration + solver + async deps
- **Files:** `pyproject.toml`, `CLAUDE.md`, `spikes/__init__.py`, `tests/spikes/__init__.py`, `tests/spikes/test_imports.py`, `docs/fase1-spikes/README.md`, `uv.lock`
- **Changes:** In `pyproject.toml` add to `[project.optional-dependencies].dev` (NOT core — spikes are throwaway): `agent-framework-orchestrations>=1.0.0` (GA; resolves with core 1.9.0), `pulp>=2.8` (CBC solver), `pytest-asyncio>=0.24`. Set `[tool.pytest.ini_options] pythonpath = ["src", "."]` (repo-root added so `import spikes` resolves; `spikes/` has no name-collision with any installed dist, so it cannot shadow them) and `asyncio_mode = "auto"`. Set `[tool.ruff] src = ["src", "tests", "spikes"]`. Create `spikes/__init__.py` and `tests/spikes/__init__.py` (docstring: "Throwaway Fase 1 de-risk spikes — dev-only, never shipped in the wheel; safe to delete after findings recorded"). Create `docs/fase1-spikes/README.md` (the four spikes, the resolved version gate, a **Disposal** subsection: `rm -rf spikes tests/spikes docs/fase1-spikes` + revert the three `dev`/`tool` pyproject edits). In `CLAUDE.md` correct the stale stack line `MAF (`agent-framework` 1.8.0)` → `MAF (`agent-framework-core` 1.9.0)` (honors the brief's Spike-A version precondition). Run `uv sync --extra dev`, then re-confirm `uv pip show agent-framework-orchestrations` resolves cleanly (premise re-verification). (new files: all except `pyproject.toml`, `CLAUDE.md`)
- **Reuses:** existing `pyproject.toml` / `CLAUDE.md` structure (Fase 0).
- **Test first:**
- File: `tests/spikes/test_imports.py` (new)
- Verifies: `import spikes`; `from agent_framework.orchestrations import GroupChatBuilder, ConcurrentBuilder, MagenticBuilder, StandardMagenticManager, TerminationCondition`; `import pulp`; `import pytest_asyncio`.
- Pattern: `tests/test_smoke.py`
- **Verify:** `uv run --extra dev pytest tests/spikes/test_imports.py -q` → expected: passed; `uv pip show agent-framework-orchestrations` → `Version: 1.0.0`; `uv run ruff check .` → exit 0
- **On failure:** retry — if the orchestrations import fails after sync, confirm the lock resolved `agent-framework-orchestrations==1.0.0`; if anything pulls a pre-release, pin tighter. Revert with `git checkout -- pyproject.toml CLAUDE.md uv.lock` if unrecoverable.
- **Checkpoint:** `git commit -m "build(fase1): add dev orchestration + solver + async deps, scaffold spikes"`
- **Manifest:**
```yaml
manifest:
expected_paths:
- pyproject.toml
- CLAUDE.md
- spikes/__init__.py
- tests/spikes/__init__.py
- tests/spikes/test_imports.py
- docs/fase1-spikes/README.md
- uv.lock
min_file_count: 7
commit_message_pattern: "^build\\(fase1\\): add dev orchestration \\+ solver \\+ async deps"
bash_syntax_check: []
forbidden_paths:
- src/portfolio_optimiser/backends.py
- src/portfolio_optimiser/reference_domain.py
must_contain:
- path: pyproject.toml
pattern: "agent-framework-orchestrations"
- path: pyproject.toml
pattern: "pytest-asyncio"
- path: CLAUDE.md
pattern: "agent-framework-core` 1.9.0"
```
### Step 2: Shared spike harness + builder smoke (budget caps, fake client, live gate, async builder de-risk)
- **Files:** `spikes/_harness.py`, `tests/spikes/test_harness.py`
- **Changes:** Implement the cost/stop invariant (B4) and the test seams shared by all spikes. (1) `@dataclass Budget(max_tokens: int, max_rounds: int)` — constructor raises `ValueError` for non-positive caps (refuse to start without caps, A4); a `TokenMeter` plus `charge(tokens)`/`tick_round()` raise `BudgetExceeded(kind, limit, observed)` the moment a cap is crossed. (Two exception types is intentional and documented in the module: `ValueError` = bad construction argument; `BudgetExceeded` = runtime breach.) (2) `FakeChatClient(BaseChatClient)` returning scripted deterministic responses and counting tokens by word-count; plus `fake_agent(client, name, instructions)` building a minimal real `agent_framework.Agent` backed by the fake client so the orchestration builders get genuine participants. (3) `live_local_client_or_skip()` → builds `agent_framework.openai.OpenAIChatClient(base_url=env PORTFOLIO_LOCAL_BASE_URL, model=env PORTFOLIO_LOCAL_MODEL)` **directly**, or `pytest.skip(...)` when env is unset (the D2 `LocalBackend` seam is deliberately not used — its live-wiring is Fase 2). (4) A **builder smoke** that front-loads the biggest risk: construct a `GroupChatBuilder` and a `ConcurrentBuilder` with two `fake_agent`s and `await` a trivial run, asserting they execute — this is the one place that proves `FakeChatClient` can drive the GA builders. If it cannot, that is itself a recorded de-risk finding (escalate), surfaced before any spike is built on the assumption. (new file)
- **Reuses:** `agent_framework.BaseChatClient`, `agent_framework.Agent`; `agent_framework.openai.OpenAIChatClient`; `agent_framework.orchestrations.{GroupChatBuilder, ConcurrentBuilder}`; the D6 cost contract.
- **Test first:**
- File: `tests/spikes/test_harness.py` (new)
- Verifies: `Budget(0, 5)` and `Budget(5, 0)` raise `ValueError`; `charge`/`tick_round` raise `BudgetExceeded` past their caps; `FakeChatClient` returns scripted replies in order and accumulates tokens; `live_local_client_or_skip()` skips when env unset; **builder smoke** (async test) constructs and runs `GroupChatBuilder`+`ConcurrentBuilder` with fake agents and asserts completion.
- Pattern: `tests/test_backends.py` (`parametrize` + `pytest.raises`)
- **Verify:** `uv run --extra dev pytest tests/spikes/test_harness.py -q` → expected: passed; `uv run ruff check spikes tests/spikes` → exit 0
- **On failure:** escalate — if the builder smoke proves `FakeChatClient` cannot drive the GA builders, STOP and record it in `docs/fase1-spikes/README.md` as a primary de-risk finding (it changes Spikes A/B); decide with the operator whether to back the builders with a tiny canned-LLM stub or restructure. Otherwise revert with `git checkout -- spikes/_harness.py tests/spikes/test_harness.py`.
- **Checkpoint:** `git commit -m "feat(fase1): shared spike harness + builder smoke (budget caps, fake client, live gate)"`
- **Manifest:**
```yaml
manifest:
expected_paths:
- spikes/_harness.py
- tests/spikes/test_harness.py
min_file_count: 2
commit_message_pattern: "^feat\\(fase1\\): shared spike harness \\+ builder smoke"
bash_syntax_check: []
forbidden_paths:
- src/portfolio_optimiser/backends.py
must_contain:
- path: spikes/_harness.py
pattern: "class BudgetExceeded"
- path: spikes/_harness.py
pattern: "def live_local_client_or_skip"
- path: tests/spikes/test_harness.py
pattern: "GroupChatBuilder"
```
### Step 3: Spike A — Group Chat maker-checker vs single-agent (U3 / G7)
- **Files:** `spikes/a_groupchat.py`, `tests/spikes/test_a_groupchat.py`, `docs/fase1-spikes/findings-a.md`
- **Changes:** **Logic layer (always green):** `Arm` result dataclass (rounds, stalls, tokens, caught_flaw); `make_termination(n_rounds=3)` (N fixed at 3 per reviewer refinement #1); `verdict(mc: Arm, single: Arm) -> dict` computing `better = mc.caught_flaw and not single.caught_flaw` and `affordable = mc.tokens <= 3 * single.tokens` and `passed = better and affordable`, plus a comparison table renderer. **Integration/live layer (gated):** `run_live(project)` builds the 3-role maker-checker (proposer · critic · validator) via `GroupChatBuilder` + `make_termination()` and a single-agent baseline, runs both against `load_reference_projects()[0]` carrying a **planted flaw** (a savings candidate violating an obvious constraint, e.g. negative residual quantity), measures real numbers, writes them + the computed verdict to `findings-a.md`. (new file)
- **Reuses:** `spikes._harness` (Budget, FakeChatClient, fake_agent, live gate); `reference_domain.load_reference_projects`; `agent_framework.orchestrations.GroupChatBuilder`/`TerminationCondition`.
- **Test first:**
- File: `tests/spikes/test_a_groupchat.py` (new)
- Verifies (logic, non-tautological — inputs are **varied**, not a single scripted constant): `verdict()` returns `passed=True` when mc caught & single missed & tokens within 3×; `passed=False` when mc tokens = 5× single (affordable=False) even if better; `passed=False` when both caught the flaw (better=False); `make_termination(3)` stops at 3 rounds. A `@pytest.mark.skipif(no endpoint)` live test runs `run_live` and asserts ≤3 rounds, cap respected, a findings table emitted.
- Pattern: `tests/test_reference_domain.py` + `tests/test_backends.py`
- **Verify:** `uv run --extra dev pytest tests/spikes/test_a_groupchat.py -q` → expected: logic tests passed (live test skipped without endpoint)
- **On failure:** retry — confirm `GroupChatBuilder` method names by introspection and adjust `run_live`; the logic layer is independent of the MAF surface. Revert with `git checkout -- spikes/a_groupchat.py tests/spikes/test_a_groupchat.py docs/fase1-spikes/findings-a.md` if unrecoverable.
- **Checkpoint:** `git commit -m "feat(fase1): spike A - group chat maker-checker vs single-agent"`
- **Manifest:**
```yaml
manifest:
expected_paths:
- spikes/a_groupchat.py
- tests/spikes/test_a_groupchat.py
- docs/fase1-spikes/findings-a.md
min_file_count: 3
commit_message_pattern: "^feat\\(fase1\\): spike A"
bash_syntax_check: []
forbidden_paths:
- src/portfolio_optimiser/reference_domain.py
must_contain:
- path: spikes/a_groupchat.py
pattern: "GroupChatBuilder"
- path: tests/spikes/test_a_groupchat.py
pattern: "skipif"
- path: docs/fase1-spikes/findings-a.md
pattern: "token"
```
### Step 4: Spike B — Magentic limits=None footgun + ConcurrentBuilder state isolation (G1 / G2)
- **Files:** `spikes/b_footguns.py`, `tests/spikes/test_b_footguns.py`, `docs/fase1-spikes/findings-b.md`
- **Changes:** Driving the builders **is** the de-risk here, so these tests run in the gate (no live LLM — `FakeChatClient`/`fake_agent`), having been de-risked by the Step 2 builder smoke. (a) **Magentic unbounded (G1):** `MagenticBuilder` + `StandardMagenticManager` with `max_round_count=None` over a fake agent that never finalizes; drive it under the harness **round/iteration guard** (the deterministic assertion — wall-clock is only a secondary safety net, not asserted) and assert it does NOT self-terminate before the guard fires → confirms an explicit limit is required (B4). With explicit limits → terminates cleanly. (b) **Fan-out state isolation (G2/B7):** `fresh_workflow()` factory + a stateful executor; assert a SHARED instance bleeds state across the 3 projects and a FRESH instance per run shows zero bleed (B7). Writes observations + a token-use line ("0 — no live LLM") to `findings-b.md`. (new file)
- **Reuses:** `spikes._harness` (round guard, FakeChatClient, fake_agent); `reference_domain.load_reference_projects`; `agent_framework.orchestrations.{MagenticBuilder, StandardMagenticManager, ConcurrentBuilder}`.
- **Test first:**
- File: `tests/spikes/test_b_footguns.py` (new)
- Verifies: unbounded Magentic hits the round guard without self-terminating (asserts guard fired, not a natural stop); bounded Magentic stops within its limit; shared-instance fan-out leaks state (assertion of bleed); fresh-instance fan-out asserts no bleed across all 3 projects.
- Pattern: `tests/test_backends.py`
- **Verify:** `uv run --extra dev pytest tests/spikes/test_b_footguns.py -q` → expected: passed
- **On failure:** escalate — if `MagenticBuilder`/`ConcurrentBuilder` cannot be driven by a fake agent despite the Step 2 smoke, record it as a genuine de-risk finding in `findings-b.md` (it answers the footgun question by other means) before reverting `git checkout -- spikes/b_footguns.py tests/spikes/test_b_footguns.py docs/fase1-spikes/findings-b.md`.
- **Checkpoint:** `git commit -m "feat(fase1): spike B - magentic unbounded + concurrent state isolation"`
- **Manifest:**
```yaml
manifest:
expected_paths:
- spikes/b_footguns.py
- tests/spikes/test_b_footguns.py
- docs/fase1-spikes/findings-b.md
min_file_count: 3
commit_message_pattern: "^feat\\(fase1\\): spike B"
bash_syntax_check: []
forbidden_paths: []
must_contain:
- path: spikes/b_footguns.py
pattern: "ConcurrentBuilder"
- path: spikes/b_footguns.py
pattern: "fresh_workflow"
- path: docs/fase1-spikes/findings-b.md
pattern: "token"
```
### Step 5: Spike C — blocking hybrid validator (IR→PuLP→Monte Carlo, self-repair) (B1)
- **Files:** `spikes/c_validator.py`, `tests/spikes/test_c_validator.py`, `docs/fase1-spikes/findings-c.md`
- **Changes:** Fully deterministic — no MAF builders. Pydantic IR `SavingsProposal` (project_id, measure, affected cost codes, claimed saving NOK, assumptions) with field validators + a cross-field `@model_validator` (claimed saving ≤ affected items' total; quantities ≥ 0). `validate_proposal(proposal) -> ValidatedProposal | Rejection`: (1) Pydantic schema validation; (2) **PuLP** feasibility/optimization over the project's cost items (uses PuLP's **bundled CBC** binary — if CBC is genuinely absent the step **escalates**; no silent LP-relaxation fallback); (3) **Monte Carlo** over uncertain unit-costs → P10/P50/P90 (stdlib `random` seeded + `statistics.quantiles`); (4) **structural block**: an out-of-range/infeasible proposal returns a `Rejection` that can never be consumed as a `ValidatedProposal`. `self_repair(generate, max_attempts=N)` capped, then hard-stops. The LLM-generation entry is a thin wrapper gated by `live_local_client_or_skip()`; all validator/solver/MC logic is tested with crafted IR directly. Writes a token-use line ("0 — validator deterministic; live generation gated") + percentiles to `findings-c.md`. (new file)
- **Reuses:** `pydantic.BaseModel`/`model_validator`; `pulp`; `reference_domain.Project`/`CostItem`/`load_reference_projects`; `spikes._harness` (Budget for the self-repair cap, live gate).
- **Test first:**
- File: `tests/spikes/test_c_validator.py` (new)
- Verifies: an out-of-range proposal is structurally blocked — `validate_proposal` returns `Rejection` and raises if treated as validated; a valid proposal returns `ValidatedProposal` with P10 ≤ P50 ≤ P90; Monte Carlo reproducible under the fixed seed; `self_repair` stops at `max_attempts`.
- Pattern: `tests/test_reference_domain.py` + `pytest.raises`
- **Verify:** `uv run --extra dev pytest tests/spikes/test_c_validator.py -q` → expected: passed
- **On failure:** escalate (if CBC absent — record platform finding) / otherwise revert `git checkout -- spikes/c_validator.py tests/spikes/test_c_validator.py docs/fase1-spikes/findings-c.md`.
- **Checkpoint:** `git commit -m "feat(fase1): spike C - blocking hybrid validator (IR/solver/monte-carlo)"`
- **Manifest:**
```yaml
manifest:
expected_paths:
- spikes/c_validator.py
- tests/spikes/test_c_validator.py
- docs/fase1-spikes/findings-c.md
min_file_count: 3
commit_message_pattern: "^feat\\(fase1\\): spike C"
bash_syntax_check: []
forbidden_paths: []
must_contain:
- path: spikes/c_validator.py
pattern: "model_validator"
- path: spikes/c_validator.py
pattern: "class Rejection"
- path: tests/spikes/test_c_validator.py
pattern: "P10"
```
### Step 6: Spike D — VerdictStore + ExpeL retrieval (B2)
- **Files:** `spikes/d_verdictstore.py`, `tests/spikes/test_d_verdictstore.py`, `docs/fase1-spikes/findings-d.md`
- **Changes:** Minimal in-memory VerdictStore: a list of `Verdict(id, proposal_features, decision, rationale)` seeded with 10–20 synthetic verdicts from the reference domain. **Similarity defined operationally (reviewer refinement #2):** a weighted score over *structured* fields — Jaccard on the affected cost-code set + match on `savings_measure_type` + a magnitude-bucket match on claimed saving — **not** raw description text. `retrieve(proposal, k) -> list[Verdict]` ranks by that score (this is the always-tested unit that satisfies SC-D). A thin custom `ContextProvider` subclass wraps `retrieve()` for ExpeL few-shot injection; its injection is asserted only against the introspected interface (if the abstract surface differs at coding time, the `retrieve()` ranking remains the deliverable and the SC-D assertion). The embedding-based similarity path is **removed** as out of scope for a throwaway spike (noted as a Fase-2 option in `findings-d.md`). Writes a token-use line ("0 — deterministic retrieval") + a top-K example to `findings-d.md`. (new file)
- **Reuses:** `agent_framework.ContextProvider` (top-level); `reference_domain.load_reference_projects`; `spikes._harness`.
- **Test first:**
- File: `tests/spikes/test_d_verdictstore.py` (new)
- Verifies (non-tautological by construction): seed a store where the true match shares the *structured* similarity fields with the query but uses **different description text**, while 2–3 **decoys** share surface description text but differ in the structured fields; assert the true match is top-1 and within top-K; assert deterministic ordering. (ContextProvider-injection assertion included only if the interface is confirmed; `retrieve()` ranking is the guaranteed SC-D check.)
- Pattern: `tests/test_reference_domain.py`
- **Verify:** `uv run --extra dev pytest tests/spikes/test_d_verdictstore.py -q` → expected: passed
- **On failure:** retry — confirm `ContextProvider` abstract methods by introspection; keep `retrieve()` ranking separable. Revert `git checkout -- spikes/d_verdictstore.py tests/spikes/test_d_verdictstore.py docs/fase1-spikes/findings-d.md` if unrecoverable.
- **Checkpoint:** `git commit -m "feat(fase1): spike D - verdictstore + expel retrieval"`
- **Manifest:**
```yaml
manifest:
expected_paths:
- spikes/d_verdictstore.py
- tests/spikes/test_d_verdictstore.py
- docs/fase1-spikes/findings-d.md
min_file_count: 3
commit_message_pattern: "^feat\\(fase1\\): spike D"
bash_syntax_check: []
forbidden_paths: []
must_contain:
- path: spikes/d_verdictstore.py
pattern: "def retrieve"
- path: tests/spikes/test_d_verdictstore.py
pattern: "decoy"
```
### Step 7: Consolidate findings, document disposal, confirm the green quality gate
- **Files:** `docs/fase1-spikes/README.md`
- **Changes:** Update `docs/fase1-spikes/README.md` to a consolidated summary table (spike → assumption → result → confirmed/refuted → implication for Fase 2), pulling each per-spike note's verdict and token-use line (each `findings-{a,b,c,d}.md` already exists from Steps 3–6). Confirm the **Disposal** subsection (from Step 1) is complete. Run the full quality gate across the whole repo. (no new findings files here — they are owned by their spike steps)
- **Reuses:** the per-spike findings written in Steps 3–6.
- **Test first:** n/a (documentation + gate; the gate IS the verification).
- **Verify:** `uv run ruff check .` → exit 0; `uv run ruff format --check .` → exit 0; `uv run mypy src` → exit 0; `uv run --extra dev pytest -q` → expected: all passed (originals + spike logic layers; live arms skipped without endpoint)
- **On failure:** escalate — a red gate means a prior step regressed; bisect by step and fix before consolidating.
- **Checkpoint:** `git commit -m "docs(fase1): consolidate spike findings + confirm green quality gate"`
- **Manifest:**
```yaml
manifest:
expected_paths:
- docs/fase1-spikes/README.md
- docs/fase1-spikes/findings-a.md
- docs/fase1-spikes/findings-b.md
- docs/fase1-spikes/findings-c.md
- docs/fase1-spikes/findings-d.md
min_file_count: 5
commit_message_pattern: "^docs\\(fase1\\): consolidate spike findings"
bash_syntax_check: []
forbidden_paths:
- src/portfolio_optimiser/backends.py
must_contain:
- path: docs/fase1-spikes/README.md
pattern: "Fase 2"
- path: docs/fase1-spikes/README.md
pattern: "Disposal"
```
### Failure recovery rules
- **revert** — undo this step's changes (`git checkout -- {files}`), do not proceed.
- **retry** — confirm the real MAF method signature by introspection, adjust, then revert if still failing.
- **escalate** — stop entirely; needs human judgment (Step 2 builder smoke fails / Step 4 builder un-fakeable / Step 5 CBC absent / Step 7 gate regression). For Steps 2 and 4, an escalate that *records the MAF limitation* is itself a valid de-risk outcome.
- **Checkpoint** — commit after each green step.
## Alternatives Considered
| Approach | Pros | Cons | Why rejected |
|----------|------|------|--------------|
| Spikes under `src/portfolio_optimiser/spikes/` | Single import root | Ships throwaway code in the wheel | Top-level `spikes/` keeps the wheel clean (hatchling packages only `src/portfolio_optimiser`) |
| `agent-framework-orchestrations` in **core** deps | One home for the locked debate default | Permanent Fase-2 commitment made under a throwaway task (scope creep, flagged in review) | Put it in `dev` like PuLP; promote to core in Fase 2 when the real flow is built |
| Wire `LocalBackend.create_chat_client()` in Step 2 for the live arm | Fills the D2 seam | Touches `src/` under a throwaway task; forces editing `test_backends.py`; non-throwaway commitment | Build `OpenAIChatClient` directly in the throwaway harness; defer seam-wiring to Fase 2; `src/` stays untouched |
| Require a live LOCAL endpoint for all spike tests | Tests the real thing | Gate red without an endpoint; violates SC-gate + D6 | Logic layer always green; integration/live arm gated |
| numpy for Monte Carlo | Familiar | New dep for trivial sampling | stdlib `random` + `statistics.quantiles` — zero new dep |
| `agent-framework[all]` to get orchestrations | One line | Re-introduces the beta/`--prerelease`/alpha-pydantic problem Fase 0 removed (`491a746`) | Pin the GA `agent-framework-orchestrations>=1.0.0` only |
| Embedding similarity (Spike D) | "Real" RAG | Needs live endpoint; non-deterministic; serves no SC; over-builds throwaway | Deterministic structured-field similarity; embeddings = Fase-2 note only |
| Hand-rolled feasibility instead of PuLP | No dep | Not a faithful solver-in-the-loop (B1/R2) | PuLP (bundled CBC) keeps the spike honest |
## Test Strategy
- **Framework:** pytest 9.1.1 + `pytest-asyncio` (`asyncio_mode = "auto"`); plain functions with `-> None`, `parametrize`, `pytest.raises`; fixtures from `load_reference_projects()`. Async tests for the MAF-builder integration.
- **Existing patterns:** `tests/test_backends.py`, `tests/test_reference_domain.py`, `tests/test_smoke.py`.
- **New tests in this plan:** 6 modules (imports, harness+builder-smoke, A, B, C, D). Determinism via `FakeChatClient` + fixed seeds; live arms isolated behind `@pytest.mark.skipif(no endpoint)`. Spike B drives the builders in-gate (no endpoint) — de-risked by the Step 2 smoke.
### Tests to write
| Type | File | Verifies | Model test |
|------|------|----------|------------|
| Import | `tests/spikes/test_imports.py` | orchestrations + pulp + pytest-asyncio + `spikes` import | `tests/test_smoke.py` |
| Unit+async | `tests/spikes/test_harness.py` | Budget caps; FakeChatClient; skip gate; builder smoke | `tests/test_backends.py` |
| Unit | `tests/spikes/test_a_groupchat.py` | varied-input verdict logic; ≤3-round termination; gated live arm | `tests/test_reference_domain.py` |
| Unit+async | `tests/spikes/test_b_footguns.py` | unbounded Magentic needs guard; zero fan-out bleed | `tests/test_backends.py` |
| Unit | `tests/spikes/test_c_validator.py` | structural block; P10≤P50≤P90; self-repair cap | `tests/test_reference_domain.py` |
| Unit | `tests/spikes/test_d_verdictstore.py` | non-tautological top-K retrieval | `tests/test_reference_domain.py` |
## Risks and Mitigations
| Priority | Risk | Location | Impact | Mitigation |
|----------|------|----------|--------|------------|
| High | `FakeChatClient` cannot drive the async GA builders | `spikes/_harness.py` (smoke), A/B | Spikes A/B integration blocked | Step 2 builder smoke front-loads this; failure is a recorded de-risk finding (escalate), not a late surprise; logic layers (A verdict, B isolation) stay green regardless |
| High | Real builder method signatures differ from planned wiring | A/B integration | Integration code won't run as written | Confirm by introspection at coding time (brief-sanctioned); per-step retry; logic layers independent of the MAF surface |
| Medium | No LOCAL endpoint during the session | A/C live arms | Live numbers not captured | Logic layers prove the machinery; live arms `skip`; findings note records "live arm not run" honestly |
| Medium | PuLP CBC binary missing on Intel mac | `spikes/c_validator.py` | Solver call fails | PuLP wheels bundle CBC; if genuinely absent → **escalate** (recorded), no silent fallback |
| Low | Async-mode misconfig breaks existing sync tests | `pyproject.toml` | Original 12 tests affected | `asyncio_mode="auto"` only auto-marks coroutine tests; sync tests unaffected — Step 7 gate catches any regression |
| Low | 2 `ExperimentalWarning`s (SKILLS/HARNESS) at import (G8) | all imports | Noise only | Harmless; do not enable `filterwarnings=error` |
| Low | `.claude/projects/**` committed to a to-be-public repo | repo | Planning scratch in history | Out of scope here; flag at Fase 4 release prep |
## Assumptions
| # | Assumption | Why unverifiable now | Impact if wrong |
|---|-----------|-----------------|-----------------|
| 1 | LOCAL profile = reachable OpenAI-compatible endpoint via `PORTFOLIO_LOCAL_*` env | Depends on operator's running Ollama/LM Studio | Live arms skip; logic layers still pass — low impact |
| 2 | `FakeChatClient` can drive the GA builders | Not yet installed/introspected to method level | Step 2 smoke decides early; failure is a documented finding, not silent breakage |
| 3 | PuLP's bundled CBC works on the Intel-mac venv | Not installed yet | Step 5 escalates with a recorded platform finding |
| 4 | N = 3 rounds is a fair maker-checker bound | Judgment from G7 token-discipline | If too tight to converge live, bump to ≤5 and re-record (documented range) |
## Verification
*Per-step manifests are checked automatically by trekexecute. These map to the brief Success Criteria.*
- [ ] Spike A: `uv run --extra dev pytest tests/spikes/test_a_groupchat.py -q` → varied-input verdict logic passes; gated live arm (when run) converges ≤3 rounds, cap respected, comparison table emitted (SC-A; empirical better/cheaper verdict is endpoint-dependent and so reported)
- [ ] Spike B: `uv run --extra dev pytest tests/spikes/test_b_footguns.py -q` → unbounded Magentic needs an external guard; fresh-instance fan-out zero state-bleed (SC-B)
- [ ] Spike C: `uv run --extra dev pytest tests/spikes/test_c_validator.py -q` → out-of-range blocked structurally; valid proposal yields P10/P50/P90; self-repair capped (SC-C)
- [ ] Spike D: `uv run --extra dev pytest tests/spikes/test_d_verdictstore.py -q` → top-K retrieval returns the structurally-similar verdict over surface-text decoys (SC-D)
- [ ] Gate: `uv run ruff check .` exit 0; `uv run ruff format --check .` clean; `uv run mypy src` exit 0; `uv run --extra dev pytest -q` all passed (SC-gate)
- [ ] Findings: `docs/fase1-spikes/` has README (summary table + Disposal) + findings-{a,b,c,d}.md, each with a confirmed/refuted verdict and a token-use line (brief Goal + NFR)
## Estimated Scope
- **Files to modify:** 2 (`pyproject.toml`, `CLAUDE.md`; `uv.lock` regenerated)
- **Files to create:** ~19 (4 spike modules + `_harness.py` + 6 test modules + 2 `__init__.py` + 5 docs notes + README owned by Step 1/7)
- **Complexity:** medium (logic is modest; the real unknown — MAF builder ergonomics with a fake client — is de-risked up-front by the Step 2 smoke)
## Execution Strategy
### Session 1: Foundation
- **Steps:** 1, 2
- **Wave:** 1
- **Depends on:** none
- **Scope fence:**
- Touch: `pyproject.toml`, `CLAUDE.md`, `spikes/__init__.py`, `spikes/_harness.py`, `tests/spikes/{__init__,test_imports,test_harness}.py`, `docs/fase1-spikes/README.md`, `uv.lock`
- Never touch: `src/**`, the four spike modules
### Session 2: Spike A
- **Steps:** 3
- **Wave:** 2
- **Depends on:** Session 1
- **Scope fence:**
- Touch: `spikes/a_groupchat.py`, `tests/spikes/test_a_groupchat.py`, `docs/fase1-spikes/findings-a.md`
- Never touch: other spike modules, `src/**`, `pyproject.toml`, `CLAUDE.md`
### Session 3: Spike B
- **Steps:** 4
- **Wave:** 2
- **Depends on:** Session 1
- **Scope fence:**
- Touch: `spikes/b_footguns.py`, `tests/spikes/test_b_footguns.py`, `docs/fase1-spikes/findings-b.md`
- Never touch: other spike modules, `src/**`, `pyproject.toml`, `CLAUDE.md`
### Session 4: Spike C
- **Steps:** 5
- **Wave:** 2
- **Depends on:** Session 1
- **Scope fence:**
- Touch: `spikes/c_validator.py`, `tests/spikes/test_c_validator.py`, `docs/fase1-spikes/findings-c.md`
- Never touch: other spike modules, `src/**`, `pyproject.toml`, `CLAUDE.md`
### Session 5: Spike D
- **Steps:** 6
- **Wave:** 2
- **Depends on:** Session 1
- **Scope fence:**
- Touch: `spikes/d_verdictstore.py`, `tests/spikes/test_d_verdictstore.py`, `docs/fase1-spikes/findings-d.md`
- Never touch: other spike modules, `src/**`, `pyproject.toml`, `CLAUDE.md`
### Session 6: Consolidation
- **Steps:** 7
- **Wave:** 3
- **Depends on:** Sessions 2–5
- **Scope fence:**
- Touch: `docs/fase1-spikes/README.md`
- Never touch: `spikes/**`, `src/**`, `pyproject.toml`, `CLAUDE.md`
### Execution Order
- **Wave 1:** Session 1 (foundation — includes the builder smoke that de-risks all spikes)
- **Wave 2:** Sessions 2, 3, 4, 5 (parallel — independent spike modules sharing only the read-only harness)
- **Wave 3:** Session 6 (after Wave 2)
### Grouping rules applied
- Steps sharing files → same session (1+2 = foundation).
- Independent spike modules → separate sessions (parallelizable in Wave 2).
- Consolidation depends on all spikes → Wave 3.
## Plan Quality Score
| Dimension | Weight | Score | Notes |
|-----------|--------|-------|-------|
| Structural integrity | 0.15 | 90 | Foundation→spikes→consolidation; src/ untouched so scope fences are accurate |
| Step quality | 0.20 | 88 | TDD-first; false LocalBackend reuse removed; findings owned per step |
| Coverage completeness | 0.20 | 90 | SC-A/B/C/D + gate + findings mapped; A's empirical verdict honestly endpoint-dependent |
| Specification quality | 0.15 | 86 | Real import paths; builder ergonomics de-risked by Step 2 smoke; CBC fallback decided |
| Risk & pre-mortem | 0.15 | 90 | Biggest risk (fake-drives-builders) front-loaded with a green-keeping logic layer |
| Headless readiness | 0.10 | 92 | On-failure + Checkpoint + Manifest per step; expected_paths ⊆ Files |
| Manifest quality | 0.05 | 90 | All manifests checkable; orphan paths fixed |
| **Weighted total** | **1.00** | **89** | **Grade: A-** |
**Adversarial review:**
- **Plan critic:** REVISE → addressed (2 blockers + 8 major + 4 minor); see Revisions.
- **Scope guardian:** MIXED → addressed (orchestrations moved to dev; embedding path dropped; token-reporting + findings-location ratified); see Revisions.
## Revisions
*Added after Phase 9 adversarial review (plan-critic + scope-guardian, deduped: 19 findings, 0 overlap).*
| # | Finding (rule_key) | Severity | Resolution |
|---|--------------------|----------|------------|
| 1 | live-arm/impossible-localbackend-wiring | blocker | Harness builds `OpenAIChatClient` directly; dropped the false `LocalBackend` reuse; D2 seam-wiring deferred to Fase 2; `src/` untouched |
| 2 | fragile-assumption/fakeclient-drives-async-builders | blocker | Added Step 2 builder smoke (front-loads the risk); added `pytest-asyncio` + `asyncio_mode=auto`; split each spike into always-green logic layer + gated integration/live layer; failure is a recorded escalate |
| 3 | structural/scope-fence-vs-live-arm | major | Resolved by #1 — `src/**` is now genuinely never touched; fences accurate |
| 4 | missing-setup/async-test-runner | major | `pytest-asyncio>=0.24` added (Step 1); async integration tests |
| 5 | verification/spike-a-verdict-tautological | major | Logic test now varies inputs to prove the verdict function; empirical U3/G7 verdict explicitly endpoint-dependent (live arm) |
| 6 | underspecified/contextprovider-inject-no-target | major | `retrieve()` ranking is the guaranteed SC-D unit; ContextProvider injection asserted only against the introspected interface |
| 7 | manifest/expected-paths-not-subset-of-files | major | All `Files:` lines now include test/findings/`uv.lock`/`CLAUDE.md` paths |
| 8 | ordering/findings-files-creation-ambiguous | major | Each spike step (3–6) owns and creates its `findings-X.md`; Step 7 only consolidates the README |
| 9 | missing-error-handling/cbc-fallback-deferred | major | Committed to PuLP's bundled CBC; absence → escalate (recorded), no silent LP-relaxation |
| 10 | coverage/spike-a-precondition-downgraded | major | Step 1 corrects `CLAUDE.md` version; installed version confirmed 1.9.0; memory fix noted as session follow-up |
| 11 | creep — orchestrations in core deps | major | Moved `agent-framework-orchestrations` to `dev` (throwaway scope), like PuLP |
| 12 | creep — Spike D embedding path | minor | Removed; deterministic similarity only; embeddings = Fase-2 note |
| 13 | gap — token reporting B/C/D | minor | Each findings note carries a token-use line (incl. "0 — no live LLM") |
| 14 | gap — findings location vs brief | minor | Kept `docs/fase1-spikes/` (tracked, survives, consistent with research/plan) — ratified as a deliberate refinement; operator can object |
| 15 | dependency — re-confirm orchestrations resolves | minor | Step 1 runs `uv pip show agent-framework-orchestrations` post-sync |
| 16 | minor/pythonpath-root-shadowing | minor | Documented: repo-root added for `import spikes`; no name collision with installed dists |
| 17 | minor/budget-two-exception-types | minor | Documented intent: `ValueError` = bad ctor arg, `BudgetExceeded` = runtime breach |
| 18 | minor/wall-clock-flaky-assertion | minor | Spike B asserts the round/iteration guard; wall-clock is a secondary safety net only |
| 19 | minor/no-teardown-for-throwaway-spikes | minor | Step 1/7 README **Disposal** subsection: `rm -rf` + revert the pyproject dev/tool edits |