portfolio-optimiser/docs/fase1-spikes/findings-d.md

2.5 KiB
Raw Blame History

Spike D findings — VerdictStore + ExpeL retrieval (B2)

Assumption: retrieval surfaces a relevant prior verdict for a similar new proposal — the substrate for the framework's learning loop.

Result — CONFIRMED (deterministic, no endpoint)

A minimal in-memory VerdictStore holds 12 synthetic verdicts seeded from the reference domain's cost codes, measure types, magnitudes, and decisions (B2's "1020").

Similarity is structural, not textual (reviewer refinement #2): a weighted score over structured fields — Jaccard on the affected cost-code set (0.60) + a measure_type match (0.25) + a magnitude-bucket match on the claimed saving (0.15). Raw description text is deliberately ignored.

retrieve(query, k) is the guaranteed SC-D unit. The test is non-tautological by construction: the true match shares the structured fields with the query but uses different wording, while two decoys share the query's surface text but differ structurally (disjoint codes, different measure type, different magnitude bucket). The structural retriever returns the true match as top-1 and ranks the surface-text decoys last — a text-matching retriever would be fooled. Ordering is deterministic (ties break by verdict id).

ExpeL injection: a thin ExpeLContextProvider subclasses the real agent_framework.ContextProvider and, in before_run, injects the retrieved verdicts as few-shot instructions via SessionContext.extend_instructions — asserted against the introspected interface. The retrieve ranking remains the deliverable regardless of the MAF session surface.

Out of scope (Fase 2 option)

The embedding-based similarity path is intentionally not built — it needs a live endpoint, is non-deterministic, and serves no SC for a throwaway spike. Structured-field similarity is sufficient to confirm B2. Embeddings (or a hybrid structured+embedding score) are a Fase 2 option for the durable VerdictStore.

Token use

0 — deterministic retrieval. No model is called; similarity is pure arithmetic over structured fields. The ExpeL provider only formats retrieved verdicts into few-shot text — the actual model call that would consume tokens is a Fase 2 concern.

Implication for Fase 2

The learning loop's retrieval is realizable with a simple, deterministic, structural similarity — good enough to surface relevant prior verdicts. Fase 2 can keep this as the baseline and add embeddings only if structured similarity proves insufficient on real data.