Measured 2026-09-20 on a 2313-concept bundle of one project's own documentation: `withheld` held 2 305 entries = 186 440 B of compact JSON = **65.5 % of the 284 850-byte payload**, and not one of those bytes counted against the budget the same payload reports (`spent` was 45 192). A reader was handed 239 658 bytes the budget line did not know about, to learn 2 305 concept ids with nothing beside them -- the title being exactly what `--withheld-titles` existed to buy, and which was off because buying it for 2 305 entries cost another 37.9 %. `withheld` is now a mapping: `total` (equal to `denominators.withheld`, so SS 5.2's identity is unmoved and closes on the NUMBERS), `by_rule` (the same total decomposed over the closed rule set, so "what kind of drop" is answerable without the list), `nearest` (the best-ranked drops BY NAME, with title and source document, so a reader who sees a near miss can ask for it) and `complete`. The near misses are read off the ranking, not off `cut`'s output: `cut` sorts by id so the partition is comparable, and that order says nothing about which concept a reader might want next. Same question, same bundle, after: **52 421 bytes, 18.4 % of the old file**. The whole list stays reachable behind `--withheld-full`, and the two instruments that classify EVERY miss by its rule -- the retrieval gate and `okf_consume_measure` -- now ask for it explicitly and assert `complete` rather than assuming it. `--withheld-nearest N` sets the cap (default 20, which is `k` plus the next twelve). `--withheld-titles` is retired: a flag whose only remaining effect would be to STRIP the title from a list the caller asked for in full names no decision worth two shapes for one list. `CONTRACT_REVISION` moves to `okf-consumption/2`, because a consumer indexing the old key as a list would otherwise break silently. Three checker rules move with it, and one of them is the interesting case: `parent_unfollowable` used `excerpts` + `withheld` as the bundle's own denominator, which a truncated block is not -- so that clause now runs only where the payload SAYS it is complete, stated in SS 8.6 rather than left as a silence, with the other two clauses (shape, self-reference) running either way. `Report` carries both denominators, because a report claiming it examined 2 305 entries it never saw is the same defect one level up. The generated skill's "breaking point" section goes with it: it extrapolated a concept count from the cost of ONE withheld entry, and there is no such slope any more. It now states what this bundle's bookkeeping cost and that the block is bounded by the cap rather than by the bundle -- an extrapolation from a slope the code no longer has would be a measurement of the previous revision. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
190 lines
8.4 KiB
Python
190 lines
8.4 KiB
Python
"""The fusion's tie-break, measured rather than assumed.
|
|
|
|
`concept_scores` fuses three signals by RRF, and RRF consumes RANKS ONLY. A
|
|
rank is therefore produced for every concept in every signal -- including a
|
|
signal that does not separate them. The tie-break is declared
|
|
(`(-score, concept_id)`), so when a signal gives 269 of 270 concepts the same
|
|
score, that signal's contribution to the fusion is the concepts' own ids in
|
|
lexicographic order: a UUID, which is noise, weighted exactly as heavily as
|
|
the two signals that did the measuring.
|
|
|
|
MEASURED 2026-09-08 on the N500 bundle (270 concepts, `feae0c8`), for the
|
|
question about `vann- og frostsikring` in a subsea tunnel: the document prior
|
|
has **two** distinct values over the bundle, and 269 concepts share one of
|
|
them. The best covering concept answered **7 of 7** question tokens and led
|
|
the body signal at rank 6, and it fused to rank **14** -- outside the cut --
|
|
while concepts answering fewer tokens fused ahead of it on nothing but an
|
|
earlier id.
|
|
|
|
This file holds the two halves apart:
|
|
|
|
- The matcher is CHARACTERISED, not fixed. `normalise` already resolves the
|
|
hyphen-and-`og` coordination, so the alternative that would have widened the
|
|
tokeniser has nothing to widen. That is asserted here so the choice stays
|
|
falsifiable rather than remembered.
|
|
- The fusion gets one rule, behind one flag, off by default.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
PROJECT_ROOT = Path(__file__).resolve().parents[1]
|
|
sys.path.insert(0, str(PROJECT_ROOT / "tools"))
|
|
|
|
import okf_consume # noqa: E402
|
|
|
|
QUESTION = "Hvilke krav gjelder vann- og frostsikring i undersjoeisk tunnel?"
|
|
|
|
_FRONTMATTER = (
|
|
"---\ntype: reference\ntitle: {title}\nsource_file: {slug}.md\n"
|
|
"source_sha256: {digest}\ningested_at: 2026-09-01T00:00:00Z\n"
|
|
"adjudication: proposed\nbundle_id: tie-fixture\n"
|
|
"verified: [{{ by: process:okf-check, at: 2026-09-01T00:00:00Z }}]\n---\n\n"
|
|
)
|
|
|
|
|
|
def _tie_bundle(root: Path, *, fillers: int = 30) -> Path:
|
|
"""One document, so the document prior cannot separate anything.
|
|
|
|
The gold concept's id sorts LAST and the fillers' ids sort first, which is
|
|
what makes the degenerate signal's tie-break work against the concept that
|
|
answers the question. The fillers answer `krav`, `gjelder`, `vann` and
|
|
`tunnel`; only the gold answers `frostsikring` and `undersjoeisk` too.
|
|
"""
|
|
(root / "krav").mkdir(parents=True)
|
|
(root / "index.md").write_text(
|
|
"---\nokf_version: 0.2\nbundle_id: tie-fixture\n---\n\n- [krav (index)](krav/index.md)\n",
|
|
encoding="utf-8",
|
|
)
|
|
entries: list[str] = []
|
|
|
|
def add(slug: str, title: str, body: str) -> None:
|
|
entries.append(f"- [{title}]({slug}.md) — adjudication: proposed\n")
|
|
(root / "krav" / f"{slug}.md").write_text(
|
|
_FRONTMATTER.format(title=title, slug=slug, digest="1" * 64) + f"## {title}\n\n" + body,
|
|
encoding="utf-8",
|
|
)
|
|
|
|
add(
|
|
"zz-gull",
|
|
"Krav om vann- og frostsikring i undersjoeisk tunnel",
|
|
"Kravet gjelder vannsikring og frostsikring i undersjoeisk tunnel.\n" * 4,
|
|
)
|
|
for number in range(1, fillers + 1):
|
|
add(
|
|
f"aa-{number:02d}",
|
|
f"Krav om tunnel og vann {number:02d}",
|
|
"Kravet gjelder tunnel og vann i anlegget.\n" * 4,
|
|
)
|
|
(root / "krav" / "index.md").write_text("".join(entries), encoding="utf-8")
|
|
return root
|
|
|
|
|
|
def _gold_rank(root: Path, *, tie_shared_rank: bool) -> int:
|
|
concepts = [
|
|
okf_consume.read_concept(
|
|
root / f"{concept_id}.md", bundle_root=root, root_bundle_id="tie-fixture"
|
|
)
|
|
for concept_id in okf_consume.enumerate_concepts(root)
|
|
]
|
|
ranked = okf_consume.concept_scores(
|
|
concepts,
|
|
QUESTION,
|
|
okf_consume.document_scores(root, QUESTION),
|
|
tie_shared_rank=tie_shared_rank,
|
|
)
|
|
for position, (concept, _, _) in enumerate(ranked, start=1):
|
|
if concept.concept_id.endswith("zz-gull"):
|
|
return position
|
|
raise AssertionError("the gold concept is not in the ranking at all")
|
|
|
|
|
|
# --- The half that is characterised, not fixed --------------------------------
|
|
|
|
|
|
def test_the_hyphen_and_og_coordination_is_already_resolved_by_the_tokeniser() -> None:
|
|
# CHARACTERISATION, green on HEAD. This is the measurement that ruled the
|
|
# tokeniser out as the site of the fix: there is no coordination left to
|
|
# resolve, so a rule widening it could not have moved the miss.
|
|
assert okf_consume.normalise("vann- og frostsikring") == ("vann", "frostsikring")
|
|
assert okf_consume.tokens_match("frostsikring", "frostsikringen")
|
|
assert okf_consume.tokens_match("vann", "vannsikring")
|
|
|
|
|
|
def test_a_signal_that_separates_nothing_still_ranks_every_concept(tmp_path: Path) -> None:
|
|
# CHARACTERISATION of the defect's mechanism: one distinct score, and one
|
|
# rank per concept all the same. The second number is the noise.
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
prior = okf_consume.document_scores(root, QUESTION)
|
|
concepts = okf_consume.enumerate_concepts(root)
|
|
scores = {prior[concept_id.split("/", 1)[0]] for concept_id in concepts}
|
|
assert len(scores) == 1
|
|
assert len(concepts) == 31
|
|
|
|
|
|
# --- The half that gets the rule ----------------------------------------------
|
|
|
|
|
|
def test_shared_rank_lifts_the_concept_the_measuring_signals_lead(tmp_path: Path) -> None:
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
# Two DIFFERENT numbers, not one predicate two branches share: the concept
|
|
# answering every question token sits at 8 while the degenerate signal
|
|
# orders by id, and at 1 once that signal stops ordering.
|
|
assert _gold_rank(root, tie_shared_rank=False) == 8
|
|
assert _gold_rank(root, tie_shared_rank=True) == 1
|
|
|
|
|
|
def test_the_payload_is_byte_identical_with_the_flag_on(tmp_path: Path) -> None:
|
|
"""The rule became the default 2026-09-10; saying so explicitly changes nothing.
|
|
|
|
The assertion is unchanged in kind -- the implicit and the explicit value
|
|
must produce the same bytes -- only the value it names moved.
|
|
"""
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
without = okf_consume.serialise(okf_consume.build_payload(root, question=QUESTION))
|
|
explicit_on = okf_consume.serialise(
|
|
okf_consume.build_payload(root, question=QUESTION, tie_shared_rank=True)
|
|
)
|
|
assert without == explicit_on
|
|
|
|
|
|
def test_the_opt_out_reproduces_the_order_the_default_used_to_give(tmp_path: Path) -> None:
|
|
"""The other half: a consumer needing the pre-2026-09-10 order can have it.
|
|
|
|
Load-bearing rather than symmetric. `--no-tie-shared-rank` is the only
|
|
thing standing between a consumer pinned to the old excerpt order and a
|
|
silent reordering, so the opt-out needs a test that goes red if it stops
|
|
being a real alternative -- which it would be if it produced the same
|
|
bytes as the default on the very fixture built to separate them.
|
|
"""
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
default = okf_consume.serialise(okf_consume.build_payload(root, question=QUESTION, k=3))
|
|
opted_out = okf_consume.serialise(
|
|
okf_consume.build_payload(root, question=QUESTION, k=3, tie_shared_rank=False)
|
|
)
|
|
assert default != opted_out
|
|
|
|
|
|
def test_the_flag_changes_the_payload_it_is_meant_to_change(tmp_path: Path) -> None:
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
off = okf_consume.build_payload(root, question=QUESTION, k=3, tie_shared_rank=False)
|
|
on = okf_consume.build_payload(root, question=QUESTION, k=3)
|
|
delivered_off = [excerpt["concept_id"] for excerpt in off["excerpts"]] # type: ignore[index]
|
|
delivered_on = [excerpt["concept_id"] for excerpt in on["excerpts"]] # type: ignore[index]
|
|
assert not any(str(cid).endswith("zz-gull") for cid in delivered_off)
|
|
assert str(delivered_on[0]).endswith("zz-gull")
|
|
|
|
|
|
def test_the_cli_exposes_the_flag_and_defaults_it_on(tmp_path: Path) -> None:
|
|
root = _tie_bundle(tmp_path / "bundle")
|
|
parsed = okf_consume.parse_args([str(root), "--question", QUESTION])
|
|
assert parsed.tie_shared_rank is True
|
|
parsed_off = okf_consume.parse_args([str(root), "--question", QUESTION, "--no-tie-shared-rank"])
|
|
assert parsed_off.tie_shared_rank is False
|
|
# The withheld cap did NOT move with it, asserted here so the two are one
|
|
# measurement rather than two files' worth of trust: it is a number chosen
|
|
# for reasons of BYTES, which nothing this round touched.
|
|
assert parsed.withheld_nearest == okf_consume.WITHHELD_NEAREST_DEFAULT
|
|
assert parsed.withheld_full is False
|