The consumer's question about `vann- og frostsikring` in a subsea tunnel
delivered 0 of the 16 concepts covering it, best of them at fused rank 14.
Reproduced with the denominator, then decomposed per signal before anything
was built.
It is not a matcher miss. `normalise("vann- og frostsikring")` already returns
`('vann', 'frostsikring')` on HEAD, the prefix rule already bridges the
inflections, and the best covering concept already answers 7 of 7 question
tokens -- more than any delivered one. A tokeniser rule had nothing to widen.
It is the fusion, but not a weight. RRF ranks every concept in every signal,
including a signal that scored them all the same, and the declared
`(-score, concept_id)` tie-break then orders that group by id. On N500 the
document prior has TWO distinct values over 270 concepts, so the third signal
contributed alphabetical UUID order spread from 1/61 to 1/329 -- enough to put
a concept leading the body signal behind concepts sharing only `tunnel` and
`vann`.
`--tie-shared-rank` lets concepts a signal scores equally share that group's
first rank. The miss closes: best covering 14 -> 3, 2 of 16 delivered. OFF BY
DEFAULT, by the order's own rule: the three requirement lookups hold at rank 1
and the K2 digest holds, but hit@8 over the six published questions falls 5 of
6 to 4 of 6. Decomposed rather than guessed -- K2's prior is coarse (6 values
over 39 documents) rather than degenerate, and one gold sat early in its tie
group. That benefit was never a measurement, but it is a published row.
`--withheld-titles` gives each withheld entry the concept's title, so a reader
can see WHAT was withheld without reading the bundle. 11 lines of code; the
bytes are why it is off. It grows an N500 payload 37.9 % and takes the
629-concept K2 bundle's BOOKKEEPING to 122 704 B -- past the 120 000-byte limit
itself -- which would falsify the breaking point published in the tracked
`skills/okf-consume/SKILL.md` on the day it shipped.
Defaults measured, not asserted: six payload digests built from a frozen
`ff79cfa` (`git archive`, `__file__` checked) and from this tree with both
flags omitted are 6 of 6 identical, and `okf_skill.py` output is identical
apart from the paths each copy writes about itself. Contract checker exit 0 on
eight payloads, both values.
One known-positive did not reproduce and is reported rather than matched: the
order's S7 literal `2ae46f68`/169 573 B is stale by three excerpt-form commits;
HEAD measures `c759a657`/171 614 B.
Suite 1388 -> 1397. Report: docs/2026-09-08-rangeringsbom-sammensatte-ord.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
167 lines
7.2 KiB
Python
167 lines
7.2 KiB
Python
"""The fusion's tie-break, measured rather than assumed.
|
|
|
|
`concept_scores` fuses three signals by RRF, and RRF consumes RANKS ONLY. A
|
|
rank is therefore produced for every concept in every signal -- including a
|
|
signal that does not separate them. The tie-break is declared
|
|
(`(-score, concept_id)`), so when a signal gives 269 of 270 concepts the same
|
|
score, that signal's contribution to the fusion is the concepts' own ids in
|
|
lexicographic order: a UUID, which is noise, weighted exactly as heavily as
|
|
the two signals that did the measuring.
|
|
|
|
MEASURED 2026-09-08 on the N500 bundle (270 concepts, `feae0c8`), for the
|
|
question about `vann- og frostsikring` in a subsea tunnel: the document prior
|
|
has **two** distinct values over the bundle, and 269 concepts share one of
|
|
them. The best covering concept answered **7 of 7** question tokens and led
|
|
the body signal at rank 6, and it fused to rank **14** -- outside the cut --
|
|
while concepts answering fewer tokens fused ahead of it on nothing but an
|
|
earlier id.
|
|
|
|
This file holds the two halves apart:
|
|
|
|
- The matcher is CHARACTERISED, not fixed. `normalise` already resolves the
|
|
hyphen-and-`og` coordination, so the alternative that would have widened the
|
|
tokeniser has nothing to widen. That is asserted here so the choice stays
|
|
falsifiable rather than remembered.
|
|
- The fusion gets one rule, behind one flag, off by default.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
|
|
PROJECT_ROOT = Path(__file__).resolve().parents[1]
|
|
sys.path.insert(0, str(PROJECT_ROOT / "tools"))
|
|
|
|
import okf_consume # noqa: E402
|
|
|
|
QUESTION = "Hvilke krav gjelder vann- og frostsikring i undersjoeisk tunnel?"
|
|
|
|
_FRONTMATTER = (
|
|
"---\ntype: reference\ntitle: {title}\nsource_file: {slug}.md\n"
|
|
"source_sha256: {digest}\ningested_at: 2026-09-01T00:00:00Z\n"
|
|
"adjudication: proposed\nbundle_id: tie-fixture\n"
|
|
"verified: [{{ by: process:okf-check, at: 2026-09-01T00:00:00Z }}]\n---\n\n"
|
|
)
|
|
|
|
|
|
def _tie_bundle(root: Path, *, fillers: int = 30) -> Path:
|
|
"""One document, so the document prior cannot separate anything.
|
|
|
|
The gold concept's id sorts LAST and the fillers' ids sort first, which is
|
|
what makes the degenerate signal's tie-break work against the concept that
|
|
answers the question. The fillers answer `krav`, `gjelder`, `vann` and
|
|
`tunnel`; only the gold answers `frostsikring` and `undersjoeisk` too.
|
|
"""
|
|
(root / "krav").mkdir(parents=True)
|
|
(root / "index.md").write_text(
|
|
"---\nokf_version: 0.2\nbundle_id: tie-fixture\n---\n\n- [krav (index)](krav/index.md)\n",
|
|
encoding="utf-8",
|
|
)
|
|
entries: list[str] = []
|
|
|
|
def add(slug: str, title: str, body: str) -> None:
|
|
entries.append(f"- [{title}]({slug}.md) — adjudication: proposed\n")
|
|
(root / "krav" / f"{slug}.md").write_text(
|
|
_FRONTMATTER.format(title=title, slug=slug, digest="1" * 64) + f"## {title}\n\n" + body,
|
|
encoding="utf-8",
|
|
)
|
|
|
|
add(
|
|
"zz-gull",
|
|
"Krav om vann- og frostsikring i undersjoeisk tunnel",
|
|
"Kravet gjelder vannsikring og frostsikring i undersjoeisk tunnel.\n" * 4,
|
|
)
|
|
for number in range(1, fillers + 1):
|
|
add(
|
|
f"aa-{number:02d}",
|
|
f"Krav om tunnel og vann {number:02d}",
|
|
"Kravet gjelder tunnel og vann i anlegget.\n" * 4,
|
|
)
|
|
(root / "krav" / "index.md").write_text("".join(entries), encoding="utf-8")
|
|
return root
|
|
|
|
|
|
def _gold_rank(root: Path, *, tie_shared_rank: bool) -> int:
|
|
concepts = [
|
|
okf_consume.read_concept(
|
|
root / f"{concept_id}.md", bundle_root=root, root_bundle_id="tie-fixture"
|
|
)
|
|
for concept_id in okf_consume.enumerate_concepts(root)
|
|
]
|
|
ranked = okf_consume.concept_scores(
|
|
concepts,
|
|
QUESTION,
|
|
okf_consume.document_scores(root, QUESTION),
|
|
tie_shared_rank=tie_shared_rank,
|
|
)
|
|
for position, (concept, _, _) in enumerate(ranked, start=1):
|
|
if concept.concept_id.endswith("zz-gull"):
|
|
return position
|
|
raise AssertionError("the gold concept is not in the ranking at all")
|
|
|
|
|
|
# --- The half that is characterised, not fixed --------------------------------
|
|
|
|
|
|
def test_the_hyphen_and_og_coordination_is_already_resolved_by_the_tokeniser() -> None:
|
|
# CHARACTERISATION, green on HEAD. This is the measurement that ruled the
|
|
# tokeniser out as the site of the fix: there is no coordination left to
|
|
# resolve, so a rule widening it could not have moved the miss.
|
|
assert okf_consume.normalise("vann- og frostsikring") == ("vann", "frostsikring")
|
|
assert okf_consume.tokens_match("frostsikring", "frostsikringen")
|
|
assert okf_consume.tokens_match("vann", "vannsikring")
|
|
|
|
|
|
def test_a_signal_that_separates_nothing_still_ranks_every_concept(tmp_path: Path) -> None:
|
|
# CHARACTERISATION of the defect's mechanism: one distinct score, and one
|
|
# rank per concept all the same. The second number is the noise.
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
prior = okf_consume.document_scores(root, QUESTION)
|
|
concepts = okf_consume.enumerate_concepts(root)
|
|
scores = {prior[concept_id.split("/", 1)[0]] for concept_id in concepts}
|
|
assert len(scores) == 1
|
|
assert len(concepts) == 31
|
|
|
|
|
|
# --- The half that gets the rule ----------------------------------------------
|
|
|
|
|
|
def test_shared_rank_lifts_the_concept_the_measuring_signals_lead(tmp_path: Path) -> None:
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
# Two DIFFERENT numbers, not one predicate two branches share: the concept
|
|
# answering every question token sits at 8 while the degenerate signal
|
|
# orders by id, and at 1 once that signal stops ordering.
|
|
assert _gold_rank(root, tie_shared_rank=False) == 8
|
|
assert _gold_rank(root, tie_shared_rank=True) == 1
|
|
|
|
|
|
def test_the_payload_is_byte_identical_with_the_flag_off(tmp_path: Path) -> None:
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
without = okf_consume.serialise(okf_consume.build_payload(root, question=QUESTION))
|
|
explicit_off = okf_consume.serialise(
|
|
okf_consume.build_payload(root, question=QUESTION, tie_shared_rank=False)
|
|
)
|
|
assert without == explicit_off
|
|
|
|
|
|
def test_the_flag_changes_the_payload_it_is_meant_to_change(tmp_path: Path) -> None:
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
off = okf_consume.build_payload(root, question=QUESTION, k=3)
|
|
on = okf_consume.build_payload(root, question=QUESTION, k=3, tie_shared_rank=True)
|
|
delivered_off = [excerpt["concept_id"] for excerpt in off["excerpts"]] # type: ignore[index]
|
|
delivered_on = [excerpt["concept_id"] for excerpt in on["excerpts"]] # type: ignore[index]
|
|
assert not any(str(cid).endswith("zz-gull") for cid in delivered_off)
|
|
assert str(delivered_on[0]).endswith("zz-gull")
|
|
|
|
|
|
def test_the_cli_exposes_the_flag_and_defaults_it_off(tmp_path: Path) -> None:
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
parsed = okf_consume.parse_args([str(root), "--question", QUESTION])
|
|
assert parsed.tie_shared_rank is False
|
|
parsed_on = okf_consume.parse_args([str(root), "--question", QUESTION, "--tie-shared-rank"])
|
|
assert parsed_on.tie_shared_rank is True
|
|
# The other flag this session added, asserted here so "both default off"
|
|
# is one measurement rather than two files' worth of trust.
|
|
assert parsed.withheld_titles is False
|