llm-ingestion-okf/tests/test_tie_shared_rank.py
Kjell Tore Guttormsen 735468f600 feat(consume): BM25 ranking by passage and title, a large concept delivered as its passage
C1. `okf consume` and MCP's `okf_ask` now rank with BM25 (`bm25.py`) instead
of the three-signal fusion. Two signals, fused by reciprocal rank:

- passage: every body cut into 500-character windows every 250, a concept
  scored by its BEST window -- a narrow question is answered in one place;
- field: title three times, the id path and source name twice, then the body
  -- a broad question is answered by what a section is called.

The document prior and the rarity weight are gone from the default: the first
favoured big documents full of common words, the second gave its largest
weight to a word the collection does not hold. Under BM25 such a word weighs
exactly zero. A signal that scores a concept zero adds nothing to it, and ties
share a rank, so alphabetical order lifts nothing either.

Three rules carried over from the fusion, each with its own test, because the
suite showed what BM25 alone lost:
- a directory every concept shares is not read (K3-20's defect, one signal on);
- a number a section is known by (`4.2`, `10.2-2`) is kept as one token, or
  a question naming a section by its number matches nothing in it;
- a question word the collection does NOT hold is read as the collection's
  words it shares a leading word with (`consume.tokens_match`) -- Norwegian
  inflection and compounding -- at that word's idf, never at its own.

The lookup and title-covered partitions are shared with the fusion
(`_partitioned`). `ranking="fusion"` / `--ranking fusion` keeps the old order
reachable; `--cost-vocabulary` and `--rarity-weight` widen only the fusion and
are refused with the default (`ranking_flag_conflict`) rather than ignored.

C3. A concept longer than `PASSAGE_CHARS` (4 000) is delivered as the span
around its best window, snapped to whole lines, under the nearest heading
above it, with `[...]` where text was left out. `passage: {start, end, of}`
says so, `text_sha256` covers what was delivered, and `sha256` stays the
file's, so the whole can be fetched by `concept_id`. 4 000 because eight
excerpts of it stay far under a tool response's limit even with several
sub-questions merged, while a 500-character window keeps 3 500 characters of
surroundings. The budget pays for the passage, not the file.

Tests moved with the default, each stated rather than silenced:
- fusion-mechanism tests (cost vocabulary, rarity weight, reservation, shared
  rank, the reference-bundle pins) ask for `ranking="fusion"`, the order they
  were measured on; the BM25 reading of the reference bundle is a separate
  measurement, kept in local state;
- the retrieval gate still measures the shipped default. Row 1 holds. Four of
  its premises were built against the fusion (a concept forced below k that
  BM25 now delivers, a quota that no longer decides, mutants patching fusion
  code) and are `xfail(strict=True)` until the fixtures are re-measured;
- the shipped example payload is regenerated; the shipped skill is unchanged.

README's Consume section and CLAUDE.md state the new default and that the
flags described after it belong to the fusion.

The search gate's table for this commit is kept in local state: the question
sets belong to a consumer whose content does not go on a public mirror.

Suite on a clean tree after `git add`: 2390 passed, 2 skipped, 4 xfailed.
ruff, ruff format, mypy --strict clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 06:22:56 +02:00

198 lines
8.6 KiB
Python

"""The fusion's tie-break, measured rather than assumed.
`concept_scores` fuses three signals by RRF, and RRF consumes RANKS ONLY. A
rank is therefore produced for every concept in every signal -- including a
signal that does not separate them. The tie-break is declared
(`(-score, concept_id)`), so when a signal gives 269 of 270 concepts the same
score, that signal's contribution to the fusion is the concepts' own ids in
lexicographic order: a UUID, which is noise, weighted exactly as heavily as
the two signals that did the measuring.
MEASURED 2026-09-08 on the N500 bundle (270 concepts, `feae0c8`), for the
question about `vann- og frostsikring` in a subsea tunnel: the document prior
has **two** distinct values over the bundle, and 269 concepts share one of
them. The best covering concept answered **7 of 7** question tokens and led
the body signal at rank 6, and it fused to rank **14** -- outside the cut --
while concepts answering fewer tokens fused ahead of it on nothing but an
earlier id.
This file holds the two halves apart:
- The matcher is CHARACTERISED, not fixed. `normalise` already resolves the
hyphen-and-`og` coordination, so the alternative that would have widened the
tokeniser has nothing to widen. That is asserted here so the choice stays
falsifiable rather than remembered.
- The fusion gets one rule, behind one flag, off by default.
"""
from __future__ import annotations
import sys
from pathlib import Path
PROJECT_ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(PROJECT_ROOT / "tools"))
import okf_consume # noqa: E402
QUESTION = "Hvilke krav gjelder vann- og frostsikring i undersjoeisk tunnel?"
_FRONTMATTER = (
"---\ntype: reference\ntitle: {title}\nsource_file: {slug}.md\n"
"source_sha256: {digest}\ningested_at: 2026-09-01T00:00:00Z\n"
"adjudication: proposed\nbundle_id: tie-fixture\n"
"verified: [{{ by: process:okf-check, at: 2026-09-01T00:00:00Z }}]\n---\n\n"
)
def _tie_bundle(root: Path, *, fillers: int = 30) -> Path:
"""One document, so the document prior cannot separate anything.
The gold concept's id sorts LAST and the fillers' ids sort first, which is
what makes the degenerate signal's tie-break work against the concept that
answers the question. The fillers answer `krav`, `gjelder`, `vann` and
`tunnel`; only the gold answers `frostsikring` and `undersjoeisk` too.
"""
(root / "krav").mkdir(parents=True)
(root / "index.md").write_text(
"---\nokf_version: 0.2\nbundle_id: tie-fixture\n---\n\n- [krav (index)](krav/index.md)\n",
encoding="utf-8",
)
entries: list[str] = []
def add(slug: str, title: str, body: str) -> None:
entries.append(f"- [{title}]({slug}.md) — adjudication: proposed\n")
(root / "krav" / f"{slug}.md").write_text(
_FRONTMATTER.format(title=title, slug=slug, digest="1" * 64) + f"## {title}\n\n" + body,
encoding="utf-8",
)
add(
"zz-gull",
"Krav om vann- og frostsikring i undersjoeisk tunnel",
"Kravet gjelder vannsikring og frostsikring i undersjoeisk tunnel.\n" * 4,
)
for number in range(1, fillers + 1):
add(
f"aa-{number:02d}",
f"Krav om tunnel og vann {number:02d}",
"Kravet gjelder tunnel og vann i anlegget.\n" * 4,
)
(root / "krav" / "index.md").write_text("".join(entries), encoding="utf-8")
return root
def _gold_rank(root: Path, *, tie_shared_rank: bool) -> int:
concepts = [
okf_consume.read_concept(
root / f"{concept_id}.md", bundle_root=root, root_bundle_id="tie-fixture"
)
for concept_id in okf_consume.enumerate_concepts(root)
]
ranked = okf_consume.concept_scores(
concepts,
QUESTION,
okf_consume.document_scores(root, QUESTION),
tie_shared_rank=tie_shared_rank,
)
for position, (concept, _, _) in enumerate(ranked, start=1):
if concept.concept_id.endswith("zz-gull"):
return position
raise AssertionError("the gold concept is not in the ranking at all")
# --- The half that is characterised, not fixed --------------------------------
def test_the_hyphen_and_og_coordination_is_already_resolved_by_the_tokeniser() -> None:
# CHARACTERISATION, green on HEAD. This is the measurement that ruled the
# tokeniser out as the site of the fix: there is no coordination left to
# resolve, so a rule widening it could not have moved the miss.
assert okf_consume.normalise("vann- og frostsikring") == ("vann", "frostsikring")
assert okf_consume.tokens_match("frostsikring", "frostsikringen")
assert okf_consume.tokens_match("vann", "vannsikring")
def test_a_signal_that_separates_nothing_still_ranks_every_concept(tmp_path: Path) -> None:
# CHARACTERISATION of the defect's mechanism: one distinct score, and one
# rank per concept all the same. The second number is the noise.
root = _tie_bundle(tmp_path / "bundle")
prior = okf_consume.document_scores(root, QUESTION)
concepts = okf_consume.enumerate_concepts(root)
scores = {prior[concept_id.split("/", 1)[0]] for concept_id in concepts}
assert len(scores) == 1
assert len(concepts) == 31
# --- The half that gets the rule ----------------------------------------------
def test_shared_rank_lifts_the_concept_the_measuring_signals_lead(tmp_path: Path) -> None:
root = _tie_bundle(tmp_path / "bundle")
# Two DIFFERENT numbers, not one predicate two branches share: the concept
# answering every question token sits at 8 while the degenerate signal
# orders by id, and at 1 once that signal stops ordering.
assert _gold_rank(root, tie_shared_rank=False) == 8
assert _gold_rank(root, tie_shared_rank=True) == 1
def test_the_payload_is_byte_identical_with_the_flag_on(tmp_path: Path) -> None:
"""The rule became the default 2026-09-10; saying so explicitly changes nothing.
The assertion is unchanged in kind -- the implicit and the explicit value
must produce the same bytes -- only the value it names moved.
"""
root = _tie_bundle(tmp_path / "bundle")
without = okf_consume.serialise(
okf_consume.build_payload(root, question=QUESTION, ranking="fusion")
)
explicit_on = okf_consume.serialise(
okf_consume.build_payload(root, question=QUESTION, ranking="fusion", tie_shared_rank=True)
)
assert without == explicit_on
def test_the_opt_out_reproduces_the_order_the_default_used_to_give(tmp_path: Path) -> None:
"""The other half: a consumer needing the pre-2026-09-10 order can have it.
Load-bearing rather than symmetric. `--no-tie-shared-rank` is the only
thing standing between a consumer pinned to the old excerpt order and a
silent reordering, so the opt-out needs a test that goes red if it stops
being a real alternative -- which it would be if it produced the same
bytes as the default on the very fixture built to separate them.
"""
root = _tie_bundle(tmp_path / "bundle")
default = okf_consume.serialise(
okf_consume.build_payload(root, question=QUESTION, ranking="fusion", k=3)
)
opted_out = okf_consume.serialise(
okf_consume.build_payload(
root, question=QUESTION, ranking="fusion", k=3, tie_shared_rank=False
)
)
assert default != opted_out
def test_the_flag_changes_the_payload_it_is_meant_to_change(tmp_path: Path) -> None:
root = _tie_bundle(tmp_path / "bundle")
off = okf_consume.build_payload(
root, question=QUESTION, ranking="fusion", k=3, tie_shared_rank=False
)
on = okf_consume.build_payload(root, question=QUESTION, ranking="fusion", k=3)
delivered_off = [excerpt["concept_id"] for excerpt in off["excerpts"]] # type: ignore[index]
delivered_on = [excerpt["concept_id"] for excerpt in on["excerpts"]] # type: ignore[index]
assert not any(str(cid).endswith("zz-gull") for cid in delivered_off)
assert str(delivered_on[0]).endswith("zz-gull")
def test_the_cli_exposes_the_flag_and_defaults_it_on(tmp_path: Path) -> None:
root = _tie_bundle(tmp_path / "bundle")
parsed = okf_consume.parse_args([str(root), "--question", QUESTION])
assert parsed.tie_shared_rank is True
parsed_off = okf_consume.parse_args([str(root), "--question", QUESTION, "--no-tie-shared-rank"])
assert parsed_off.tie_shared_rank is False
# The withheld cap did NOT move with it, asserted here so the two are one
# measurement rather than two files' worth of trust: it is a number chosen
# for reasons of BYTES, which nothing this round touched.
assert parsed.withheld_nearest == okf_consume.WITHHELD_NEAREST_DEFAULT
assert parsed.withheld_full is False