fix(semretrieval): refuse a non-finite embedding instead of scoring it (kø-(l)/S3.1 MINOR)
`cosine`'s docstring claimed its guard was load-bearing because "a NaN reaching the
ranking sort key would corrupt ordering silently rather than failing loudly" — but the
guard tested `norm == 0.0` only, which a NaN or inf norm passes straight through. The
claim was prose, not behaviour.
Measured, not assumed: `cosine(unit, nan_vector)` AND `cosine(unit, inf_vector)` both
returned `nan`, and a NaN sort key made ranking INPUT-ORDER-DEPENDENT — six permutations
of the same three candidates produced four distinct orderings. That defeats the total
order `HybridRanker` documents ("`id` makes the result independent of input order").
Refuse rather than coerce, and deliberately NOT symmetric with the zero-norm branch: a
zero vector is a legitimate handled state (`FakeEmbedder` returns `np.zeros` by design),
whereas a non-finite component only ever means the INJECTED embedder is broken. Scoring
it `0.0` would launder that into "no semantic similarity" while ranking proceeded on a
forged signal — validation, never repair, mirroring `read_spend`.
Reachable via the documented `Embedder` extension point, not the shipped fake; scoped to
the norms (90% principle — a finite-normed dot-product overflow is not chased).
Also corrects `docs/extending.md`, which stated `SEMANTIC_WEIGHT_DEFAULT = 0.5` while the
code has said `0.25` since the weight was lowered.
625 -> 630 tests. Load-bearing MEASURED against the WHOLE suite, five mutations all red:
detach the guard entirely · coerce to 0.0 instead of raising · check only the first norm ·
drop "non-finite" from the message · (control) detach the zero-norm branch, which fails
ONLY the zero-norm test — the new guard does not mask the existing one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018V9vNBmxAmgJ2JMoHByiHS
This commit is contained in:
parent
1b7fafe222
commit
c02c1addba
4 changed files with 132 additions and 3 deletions
|
|
@ -244,12 +244,31 @@ def load_embedder_config(path: str | Path) -> EmbedderConfig:
|
|||
|
||||
|
||||
def cosine(a: np.ndarray, b: np.ndarray) -> float:
|
||||
"""Cosine similarity, with a zero-norm guard returning ``0.0``.
|
||||
"""Cosine similarity: ``0.0`` for a zero norm, ``ValueError`` for a non-finite one.
|
||||
|
||||
The guard is load-bearing: a NaN reaching the ranking sort key would corrupt ordering
|
||||
silently rather than failing loudly."""
|
||||
silently rather than failing loudly. Measured, not assumed — NaN compares False against
|
||||
everything, so ``sorted`` leaves it where the input put it, and six permutations of the same
|
||||
three candidates produced four distinct orderings. That defeats the total order
|
||||
``HybridRanker`` documents ("``id`` makes the result independent of input order").
|
||||
|
||||
The two branches are deliberately NOT symmetric. A zero vector is a legitimate, handled state
|
||||
— ``FakeEmbedder`` returns ``np.zeros`` by design — so it earns a defined score. A non-finite
|
||||
component only ever means the INJECTED embedder is broken (the shipped fake cannot emit one),
|
||||
and coercing it to ``0.0`` would launder that into "no semantic similarity" while ranking
|
||||
proceeded on a forged signal. Validation, never repair, matching ``read_spend``: reading
|
||||
corrupt state as zero hands back a false answer in the caller's own units.
|
||||
|
||||
Scoped to the norms on purpose: a non-finite component always poisons its norm, which is the
|
||||
reachable defect. A finite-normed pair whose dot product overflows is not chased here (90%
|
||||
principle) — the seam this guards is a broken embedder, not float brinkmanship."""
|
||||
norm_a = float(np.linalg.norm(a))
|
||||
norm_b = float(np.linalg.norm(b))
|
||||
if not np.isfinite(norm_a) or not np.isfinite(norm_b):
|
||||
raise ValueError(
|
||||
f"embedder produced a non-finite vector (norms: {norm_a}, {norm_b}); "
|
||||
"a non-finite score corrupts the ranking order silently"
|
||||
)
|
||||
if norm_a == 0.0 or norm_b == 0.0:
|
||||
return 0.0
|
||||
return float(np.dot(a, b) / (norm_a * norm_b))
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue