47 lines
2.5 KiB
Markdown
47 lines
2.5 KiB
Markdown
# Spike D findings — VerdictStore + ExpeL retrieval (B2)
|
||
|
||
**Assumption:** retrieval surfaces a relevant *prior* verdict for a similar new proposal —
|
||
the substrate for the framework's learning loop.
|
||
|
||
## Result — CONFIRMED (deterministic, no endpoint)
|
||
|
||
A minimal in-memory `VerdictStore` holds 12 synthetic verdicts seeded from the reference
|
||
domain's cost codes, measure types, magnitudes, and decisions (B2's "10–20").
|
||
|
||
**Similarity is structural, not textual (reviewer refinement #2):** a weighted score over
|
||
*structured* fields — Jaccard on the affected cost-code set (0.60) + a `measure_type`
|
||
match (0.25) + a magnitude-bucket match on the claimed saving (0.15). Raw description text
|
||
is **deliberately ignored**.
|
||
|
||
`retrieve(query, k)` is the guaranteed SC-D unit. The test is **non-tautological by
|
||
construction**: the true match shares the structured fields with the query but uses
|
||
*different wording*, while two decoys share the query's *surface text* but differ
|
||
structurally (disjoint codes, different measure type, different magnitude bucket). The
|
||
structural retriever returns the **true match as top-1** and ranks the surface-text decoys
|
||
last — a text-matching retriever would be fooled. Ordering is **deterministic** (ties
|
||
break by verdict id).
|
||
|
||
**ExpeL injection:** a thin `ExpeLContextProvider` subclasses the real
|
||
`agent_framework.ContextProvider` and, in `before_run`, injects the retrieved verdicts as
|
||
few-shot instructions via `SessionContext.extend_instructions` — asserted against the
|
||
introspected interface. The `retrieve` ranking remains the deliverable regardless of the
|
||
MAF session surface.
|
||
|
||
## Out of scope (Fase 2 option)
|
||
|
||
The **embedding-based** similarity path is intentionally not built — it needs a live
|
||
endpoint, is non-deterministic, and serves no SC for a throwaway spike. Structured-field
|
||
similarity is sufficient to confirm B2. Embeddings (or a hybrid structured+embedding score)
|
||
are a Fase 2 option for the durable VerdictStore.
|
||
|
||
## Token use
|
||
|
||
**0 — deterministic retrieval.** No model is called; similarity is pure arithmetic over
|
||
structured fields. The ExpeL provider only *formats* retrieved verdicts into few-shot text
|
||
— the actual model call that would consume tokens is a Fase 2 concern.
|
||
|
||
## Implication for Fase 2
|
||
|
||
The learning loop's retrieval is realizable with a simple, deterministic, structural
|
||
similarity — good enough to surface relevant prior verdicts. Fase 2 can keep this as the
|
||
baseline and add embeddings only if structured similarity proves insufficient on real data.
|