portfolio-optimiser/src/portfolio_optimiser/semretrieval.py
Kjell Tore Guttormsen c02c1addba fix(semretrieval): refuse a non-finite embedding instead of scoring it (kø-(l)/S3.1 MINOR)
`cosine`'s docstring claimed its guard was load-bearing because "a NaN reaching the
ranking sort key would corrupt ordering silently rather than failing loudly" — but the
guard tested `norm == 0.0` only, which a NaN or inf norm passes straight through. The
claim was prose, not behaviour.

Measured, not assumed: `cosine(unit, nan_vector)` AND `cosine(unit, inf_vector)` both
returned `nan`, and a NaN sort key made ranking INPUT-ORDER-DEPENDENT — six permutations
of the same three candidates produced four distinct orderings. That defeats the total
order `HybridRanker` documents ("`id` makes the result independent of input order").

Refuse rather than coerce, and deliberately NOT symmetric with the zero-norm branch: a
zero vector is a legitimate handled state (`FakeEmbedder` returns `np.zeros` by design),
whereas a non-finite component only ever means the INJECTED embedder is broken. Scoring
it `0.0` would launder that into "no semantic similarity" while ranking proceeded on a
forged signal — validation, never repair, mirroring `read_spend`.

Reachable via the documented `Embedder` extension point, not the shipped fake; scoped to
the norms (90% principle — a finite-normed dot-product overflow is not chased).

Also corrects `docs/extending.md`, which stated `SEMANTIC_WEIGHT_DEFAULT = 0.5` while the
code has said `0.25` since the weight was lowered.

625 -> 630 tests. Load-bearing MEASURED against the WHOLE suite, five mutations all red:
detach the guard entirely · coerce to 0.0 instead of raising · check only the first norm ·
drop "non-finite" from the message · (control) detach the zero-norm branch, which fails
ONLY the zero-norm test — the new guard does not mask the existing one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018V9vNBmxAmgJ2JMoHByiHS
2026-08-03 21:48:50 +02:00

420 lines
22 KiB
Python

"""S3.1 — semantic retrieval: brute-force cosine over embedded proposal features (D-C = numpy).
The verdict store ranks structurally by design (``verdicts.similarity``: Jaccard over affected
cost codes + measure type + magnitude bucket), and surface ``description`` text is deliberately
excluded so a true structural match beats text decoys. That ranking cannot see a semantically
related prior verdict carrying a *different* code set. This module adds the missing signal as a
**strictly opt-in** seam: an ``Embedder`` producing vectors, and a ``Retriever`` that blends
cosine with the structural score. Default retrieval is unchanged — the hybrid only ranks when a
caller explicitly asks for it (``run.py --semantic-retrieval``).
**MAF-free** (D7-portable): stdlib + numpy only. ``ProposalFeatures``/``Verdict`` are imported
ONLY under ``TYPE_CHECKING`` and every runtime access is duck-typed, and the structural score is
**injected as a callable** rather than imported — so importing (or running) this module never
pulls in ``portfolio_optimiser.verdicts`` and therefore never pulls in ``agent_framework``.
Registered in ``tests/test_okf.py``'s ``_MAF_FREE_MODULES`` (direct-import AST guard
``test_okf_is_maf_free``) and probed transitively — including *after* a ``rank()`` call, which
catches a lazy runtime import an import-time guard would miss — by
``tests/test_semretrieval_loadbearing.py``.
Determinism, stated precisely because the two halves rest on different things:
* **Ranking order** is guaranteed by the total-order key ``(-round(score, 9), id)`` — rounding
absorbs last-bit float noise and ``id`` settles exact ties. This is the guarantee that actually
holds, and it does NOT depend on the thread pins. It has to do real work: ``np.dot`` over two
separately allocated but bit-identical vectors can differ by one ulp, because the reduction path
varies with buffer alignment. Rounding is what makes that unobservable in the ranking.
* **Byte-identical vector artifacts** (the ``.npy``/``.jsonl`` store, across machines) are what the
thread pins below defend: BLAS reduction order depends on the thread count, and ``FakeEmbedder``
calls ``np.linalg.norm``, which can dispatch to BLAS. The pins are defence-in-depth for that
narrower claim — never the basis of the ranking guarantee.
Vectors are float64 and C-contiguous. This module is the only numpy importer in the package, which
is what makes an import-time pin viable at all.
"""
from __future__ import annotations
import hashlib
import json
import os
from collections.abc import Callable, Sequence
from pathlib import Path
from typing import TYPE_CHECKING, Protocol, runtime_checkable
# Pin BLAS threads BEFORE numpy is imported — every one of these is latched by the time the numpy
# import completes (measured: setting them afterwards has no effect, even before the first BLAS
# call), so placement above the import is load-bearing, not stylistic. ``setdefault`` so an
# operator can still override deliberately.
#
# One variable per backend, because which one BITES depends on what numpy links:
# VECLIB_MAXIMUM_THREADS -> Accelerate (macOS; what numpy >= 2.0 links here by default)
# OPENBLAS_NUM_THREADS -> OpenBLAS (numpy's default wheel on Linux/Windows, and numpy 1.x)
# MKL_NUM_THREADS -> Intel MKL (conda-forge / Intel distributions)
# OMP_NUM_THREADS -> the OpenMP runtime underneath several of the above
# On THIS build the effective one is VECLIB_MAXIMUM_THREADS; the other three are measured no-ops
# against Accelerate. They stay because they are the correct knobs for an OpenBLAS/MKL/OpenMP
# deployment — that is a PORTABILITY argument, not a determinism one, and the distinction matters:
# before VECLIB was added, this block pinned nothing at all on the machine it ran on.
os.environ.setdefault("VECLIB_MAXIMUM_THREADS", "1")
os.environ.setdefault("OPENBLAS_NUM_THREADS", "1")
os.environ.setdefault("OMP_NUM_THREADS", "1")
os.environ.setdefault("MKL_NUM_THREADS", "1")
import numpy as np # noqa: E402 — must follow the thread pins above; see module docstring
from pydantic import BaseModel, model_validator # noqa: E402 — same ordering constraint
if TYPE_CHECKING: # verdicts imports agent_framework — keep it out of the runtime import graph
from portfolio_optimiser.verdicts import ProposalFeatures, Verdict
# The structural score, passed in rather than imported — see module docstring.
SimilarityFn = Callable[[ProposalFeatures, ProposalFeatures], float]
# Dimension of the fake embedding. Large enough that distinct features do not collide, small
# enough that a 500+ verdict base stays cheap. Note there is no matmul on the ranking path:
# ``HybridRanker.rank`` calls ``cosine()`` once per candidate from inside the sort key — a Python
# loop over per-pair ``np.dot`` on 64-vectors — and re-invokes the embedder for every candidate on
# every call, caching nothing. Fine at this scale; the thing to change first if it ever is not.
EMBED_DIM = 64
# Blend weight for HybridRanker: score = weight * cosine + (1 - weight) * structural.
# The caller can override per construction.
#
# 0.25, and the reason is MEASURED, not aesthetic. The previous value (0.5) was justified as
# "giving the two falsifiable signals equal say" — a claim about equivalence that the measurement
# refutes. Over the 1770 pairs of a fully enumerated 60-feature family, the ``FakeEmbedder``
# cosine is a near-constant: mean 0.7500, sd 0.0380, range [0.6123, 0.8614] => spread 0.2492. A
# term with that little dynamic range is a TIE-BREAKER, not a co-equal signal, and weighting it as
# though it were an equal partner does not make it one — it just lets 0.25 of a near-constant
# outvote real structural distance.
#
# Worst-case bound: the semantic term can overturn a structural gap G only when
# w * spread > (1 - w) * G, i.e. it is safe while ``w < G / (G + spread)``. With G = one
# ``verdicts._W_MAGNITUDE`` step (0.15) and the measured spread, that ceiling is 0.3758 — so 0.5
# sat ABOVE it and the ordering was genuinely overturnable; 0.25 sits below with margin. Stated as
# a property rather than a fixture fit: w = 0.25 holds for any spread below 0.45, against 0.2492
# observed. Measured, not assumed: 9 adverse triples at w = 0.5, zero at every weight from 0.45
# down to 0.05 (see ``test_hybrid_preserves_structural_order_across_one_magnitude_step``).
#
# Lowering it costs nothing the seam is for: cosine still decides 100% of exact structural ties at
# any w > 0, which is the role SC2 pins. Only the co-equal-signal role disappears, and it was
# never defended. An injected REAL embedder has a different cosine distribution and therefore
# needs its own weight justification — this number is derived from the shipped fake one.
SEMANTIC_WEIGHT_DEFAULT = 0.25
# Mirrors ``verdicts._MAGNITUDE_BUCKETS``. DUPLICATED, not imported, on purpose: importing it
# would pull ``verdicts`` (and therefore ``agent_framework``) into this module's runtime graph
# and break the MAF-free guard. Drift is tolerable here — these edges only shape the fake
# embedder's projection, never the structural score, which stays the single source of truth.
_MAGNITUDE_BUCKETS = [(0.0, 1e5), (1e5, 5e5), (5e5, 1e6), (1e6, float("inf"))]
@runtime_checkable
class Embedder(Protocol):
"""Maps structured proposal features to a fixed-length float64 vector.
The real implementation (an embeddings client) is a CODE-LEVEL extension point and is NOT
built here — tests and the offline path use ``FakeEmbedder``. Config selects from the closed
``build_embedder`` registry; it can never name a module to import (see ``EmbedderConfig``)."""
def __call__(self, features: ProposalFeatures) -> np.ndarray: ...
def _magnitude_bucket(value: float) -> int:
for i, (low, high) in enumerate(_MAGNITUDE_BUCKETS):
if low <= value < high:
return i
return len(_MAGNITUDE_BUCKETS) - 1
def _canonical_feature_string(features: ProposalFeatures) -> str:
"""Order-independent projection of the STRUCTURAL triple: sorted ``affected_codes``, the
measure type, and the magnitude bucket. ``affected_codes`` is a set, so it is sorted; the raw
saving is bucketed so near-identical amounts land on the same point.
``description`` is deliberately EXCLUDED, matching the two invariants that already govern this
data: ``verdicts.similarity`` ("text is ignored by design") and ``verdicts._mint_id``, which
hashes the same three fields — so two proposals differing only in prose already share one id
and one structural score. Including description here made the embedding the single place where
surface text leaked into retrieval, which inverted the feature's purpose: on a mixed store the
framework's own echo of the query outranked genuine expert prose.
Honesty limit, stated at the site: over a structural tie the cosine ordering is deterministic
but semantically ARBITRARY, because the shipped ``FakeEmbedder`` is a sha256 projection with no
semantics. Retrieval *quality* arrives only with an injected embedder — the seam is the
deliverable, not better ranking.
And the limit is wider than "ties": an arbitrary term does not politely confine itself to
exact ties. It also decides NEAR-ties, and ``SEMANTIC_WEIGHT_DEFAULT`` is the only thing that
bounds how near — at weight w it can overturn any structural gap below ``w/(1-w)`` times the
cosine spread. That is why the weight is derived from a measurement rather than chosen for
symmetry, and why raising it is a correctness change, not a taste one."""
return "|".join(
(
",".join(sorted(features.affected_codes)),
features.measure_type,
str(_magnitude_bucket(features.claimed_saving_nok)),
)
)
def _expand(canonical: str, n_bytes: int) -> bytes:
"""Deterministic byte expansion via counter-prefixed sha256. Never ``hash()`` — that is
salted per process (PYTHONHASHSEED) and would make vectors differ between runs."""
out = bytearray()
counter = 0
while len(out) < n_bytes:
out += hashlib.sha256(f"{counter}:{canonical}".encode()).digest()
counter += 1
return bytes(out[:n_bytes])
class FakeEmbedder:
"""Offline stand-in embedder: a pure, non-negative, L2-normalized sha256 projection.
Non-negative components keep ``cosine`` in ``[0, 1]``, matching the structural score's range
so the two blend without a scale mismatch. It carries no real semantics — it exists so the
seam is exercisable offline at zero cost; the load-bearing proof of the seam is that removing
the cosine term flips the ranking, not that this projection is meaningful."""
def __call__(self, features: ProposalFeatures) -> np.ndarray:
raw = _expand(_canonical_feature_string(features), EMBED_DIM)
vec = np.ascontiguousarray([b / 255.0 for b in raw], dtype="<f8")
norm = float(np.linalg.norm(vec))
if norm == 0.0:
return np.zeros(EMBED_DIM, dtype="<f8")
return np.ascontiguousarray(vec / norm, dtype="<f8")
# --- Embedder config seam: a CLOSED registry, deliberately never an import path -----------------
class EmbedderConfig(BaseModel):
"""Declarative embedder choice. ``type`` membership is enforced by ``build_embedder`` (unknown
type → ``ValueError``); the validator pins the per-type required fields.
Deliberately NOT an import path. A ``"my_pkg.mod:MyEmbedder"`` string would be arbitrary code
execution at config-load time, and would hand a config file the ability to import a module that
opens a socket — precisely the capability ``NotifierConfig`` refuses to let config grant
("egress opt-in is a code-level factory kwarg, so config alone can never grant it"). It would
also route around this module's no-network guard, which is an AST/``sys.modules`` check scoped
to ``semretrieval.py`` and blind to a third module by construction.
A future networked embedder therefore enters as a NEW REGISTRY BRANCH plus a code-level
capability kwarg — never as config."""
type: str
@model_validator(mode="after")
def _required_fields_by_type(self) -> EmbedderConfig:
if not self.type:
raise ValueError("embedder config requires a non-empty type")
return self
def build_embedder(config: EmbedderConfig) -> Embedder:
"""Type-dispatch factory over a CLOSED vocabulary. Mirrors ``notify.build_notifier``.
Shipped vocabulary: ``"fake"`` → ``FakeEmbedder`` (a deterministic sha256 projection carrying
no semantics). Anything else raises ``ValueError`` naming the seam, rather than attempting to
resolve it — see ``EmbedderConfig`` for why resolution from config is refused outright."""
if config.type == "fake":
return FakeEmbedder()
raise ValueError(
f"unknown embedder type: {config.type!r} (known: 'fake'; a new embedder is added as a "
"registry branch in semretrieval.build_embedder, never as an import path in config)"
)
def load_embedder_config(path: str | Path) -> EmbedderConfig:
"""Fail-fast standalone loader for an embedder config (mirrors ``dimension.load_dimension``).
An embedder config is *authoritative startup input*, so loading is fail-fast: a missing file
raises ``FileNotFoundError`` and malformed/invalid content raises ``pydantic.ValidationError``.
Deliberate contrast to the tolerant verdict-inbox RAW layer (``load_verdicts_from_dir``), which
skips bad files rather than raising — startup configuration must never be silently degraded.
:raises FileNotFoundError: ``path`` does not point at an existing file.
:raises pydantic.ValidationError: the JSON is malformed or violates the schema.
"""
p = Path(path)
if not p.is_file():
raise FileNotFoundError(f"embedder config not found: {str(path)!r}")
return EmbedderConfig.model_validate_json(p.read_text(encoding="utf-8"))
def cosine(a: np.ndarray, b: np.ndarray) -> float:
"""Cosine similarity: ``0.0`` for a zero norm, ``ValueError`` for a non-finite one.
The guard is load-bearing: a NaN reaching the ranking sort key would corrupt ordering
silently rather than failing loudly. Measured, not assumed — NaN compares False against
everything, so ``sorted`` leaves it where the input put it, and six permutations of the same
three candidates produced four distinct orderings. That defeats the total order
``HybridRanker`` documents ("``id`` makes the result independent of input order").
The two branches are deliberately NOT symmetric. A zero vector is a legitimate, handled state
— ``FakeEmbedder`` returns ``np.zeros`` by design — so it earns a defined score. A non-finite
component only ever means the INJECTED embedder is broken (the shipped fake cannot emit one),
and coercing it to ``0.0`` would launder that into "no semantic similarity" while ranking
proceeded on a forged signal. Validation, never repair, matching ``read_spend``: reading
corrupt state as zero hands back a false answer in the caller's own units.
Scoped to the norms on purpose: a non-finite component always poisons its norm, which is the
reachable defect. A finite-normed pair whose dot product overflows is not chased here (90%
principle) — the seam this guards is a broken embedder, not float brinkmanship."""
norm_a = float(np.linalg.norm(a))
norm_b = float(np.linalg.norm(b))
if not np.isfinite(norm_a) or not np.isfinite(norm_b):
raise ValueError(
f"embedder produced a non-finite vector (norms: {norm_a}, {norm_b}); "
"a non-finite score corrupts the ranking order silently"
)
if norm_a == 0.0 or norm_b == 0.0:
return 0.0
return float(np.dot(a, b) / (norm_a * norm_b))
@runtime_checkable
class Retriever(Protocol):
"""Ranks candidate verdicts against a query and returns the top ``k``.
``VerdictStore.retrieve`` delegates to one of these. The store's default is
``StructuralRetriever``, which reproduces the pre-seam ordering exactly, so installing a
retriever is the ONLY way retrieval behaviour changes."""
def rank(
self, query: ProposalFeatures, candidates: Sequence[Verdict], k: int
) -> list[Verdict]: ...
class StructuralRetriever:
"""Today's ranking, behind the seam: weighted structural similarity, ties broken by ``id``.
``similarity`` is INJECTED rather than imported — importing it would pull ``verdicts`` (and
thus ``agent_framework``) into this module's runtime graph. This class is the store's wired
default, not a parallel implementation kept for reference."""
def __init__(self, similarity: SimilarityFn) -> None:
self._similarity = similarity
def rank(self, query: ProposalFeatures, candidates: Sequence[Verdict], k: int) -> list[Verdict]:
return sorted(
candidates,
key=lambda v: (-self._similarity(query, v.proposal_features), v.id),
)[:k]
class HybridRanker:
"""Blends semantic cosine with the structural score: ``w * cosine + (1 - w) * structural``.
Both terms live in ``[0, 1]`` (the embedder yields non-negative unit vectors), so ``weight``
means what it reads as. Ordering uses the TOTAL order ``(-round(score, 9), id)``: rounding
absorbs any last-bit float noise, and ``id`` makes the result independent of input order even
when two candidates score identically."""
def __init__(
self,
embedder: Embedder,
similarity: SimilarityFn,
weight: float = SEMANTIC_WEIGHT_DEFAULT,
) -> None:
self._embedder = embedder
self._similarity = similarity
self._weight = weight
def rank(self, query: ProposalFeatures, candidates: Sequence[Verdict], k: int) -> list[Verdict]:
query_vector = self._embedder(query)
def score(verdict: Verdict) -> float:
semantic = cosine(query_vector, self._embedder(verdict.proposal_features))
structural = self._similarity(query, verdict.proposal_features)
return self._weight * semantic + (1.0 - self._weight) * structural
return sorted(candidates, key=lambda v: (-round(score(v), 9), v.id))[:k]
# --- Vector store: an UNWIRED authoring primitive, offered to extenders -------------------------
# Exactly two files. The `.npy` holds the matrix; the `.jsonl` sidecar holds the row->verdict-id
# mapping, one object per line in the repo's deterministic on-disk idiom (`outbox._dump`). There
# is deliberately no third `meta.json`: the numpy version that fixes the `.npy` header is already
# pinned in `uv.lock`, so a provenance file would only be a second place to drift.
#
# NOTHING IN `src/` CALLS THESE, and that is the chosen disposition — the same one the repo already
# gives `write_verdict`, `promote_verdict` and `build_mcp_server`: a public primitive that is NOT
# wired into `run_project`. Wiring it would buy nothing today, because `HybridRanker.rank` re-embeds
# every candidate on every call and would bypass a persisted matrix regardless; making the cache
# load-bearing means redesigning the ranker to consult it, which is a separate piece of work.
#
# Known limitation: the empty-store branch below hardcodes `EMBED_DIM`, so a third-party `Embedder`
# of a different dimension writes a shape-inconsistent empty store. Recorded, not fixed.
_VECTORS_NPY = "vectors.npy"
_VECTORS_JSONL = "vectors.jsonl"
def save_vector_store(
directory: str | Path,
verdicts: Sequence[Verdict],
embedder: Embedder,
) -> None:
"""Embed ``verdicts`` and write the two store artifacts into ``directory``.
Byte-deterministic: verdicts are sorted by ``id`` before embedding, the matrix is C-contiguous
float64, and the sidecar uses sorted keys with LF endings. The same verdicts in any insertion
order therefore produce identical bytes. Both files are written atomically (temp + replace) so
an interrupted run leaves the previous store intact rather than a half-written one."""
path = Path(directory)
path.mkdir(parents=True, exist_ok=True)
ordered = sorted(verdicts, key=lambda v: v.id)
if ordered:
matrix = np.ascontiguousarray(
np.vstack([embedder(v.proposal_features) for v in ordered]), dtype="<f8"
)
else:
matrix = np.empty((0, EMBED_DIM), dtype="<f8")
npy_tmp = path / f"{_VECTORS_NPY}.tmp"
# Written through a file handle, not a path: np.save appends ".npy" to a path that lacks it,
# which would turn the temp file into `vectors.npy.tmp.npy` and defeat the atomic replace.
with npy_tmp.open("wb") as handle:
np.save(handle, matrix)
os.replace(npy_tmp, path / _VECTORS_NPY)
jsonl_tmp = path / f"{_VECTORS_JSONL}.tmp"
jsonl_tmp.write_text(
"".join(json.dumps({"id": v.id}, sort_keys=True) + "\n" for v in ordered),
encoding="utf-8",
newline="\n",
)
os.replace(jsonl_tmp, path / _VECTORS_JSONL)
def load_vector_store(directory: str | Path) -> tuple[list[str], np.ndarray] | None:
"""Load the store as ``(ids, matrix)``, or ``None`` when it has not been built.
Absence is tolerated because the store is a rebuildable cache. A row/line MISMATCH is not
tolerated: it would silently map every vector to the wrong verdict id and mis-rank without any
error, so it raises ``ValueError`` (the fail-fast posture of ``ledger.SavingsLedger.load``).
Unwired: nothing in ``src/`` calls this. See the section note above — the ranking path never
consults a persisted matrix, so ``None`` is not a degradation from anything."""
path = Path(directory)
npy_path, jsonl_path = path / _VECTORS_NPY, path / _VECTORS_JSONL
if not npy_path.is_file() or not jsonl_path.is_file():
return None
with npy_path.open("rb") as handle:
matrix = np.load(handle, allow_pickle=False)
ids = [
json.loads(line)["id"]
for line in jsonl_path.read_text(encoding="utf-8").splitlines()
if line.strip()
]
if matrix.shape[0] != len(ids):
raise ValueError(
f"vector store is inconsistent: {matrix.shape[0]} matrix row(s) but {len(ids)} "
f"id line(s) in {str(path)!r} — rebuild the store"
)
return ids, matrix