`coverage` carries three lists: the terms the pre-pass read the question as, the terms no concept in the bundle answers, and the terms no delivered excerpt answers. Without it a reader holding eight excerpts cannot tell a bundle that ANSWERED its question from one that merely ranked something -- the two payloads have the same shape. FACTS, AND NO VERDICT, which is a measurement and not caution. Two readings were built and both falsified over 81 questions (16 synthetic, 65 across the three real sets, 2026-09-20): the share of question terms a delivered excerpt answers separates the synthetic controls at 0.33 against 0.50 and REVERSES on real data (covered questions down to 0.27, one genuinely uncovered question at 0.71); the share of a bundle tying the best lexical match is ~0.00 for every real question either way. Question style dominates the first, corpus size the second. The one bar this repository declares is the gate's: `UNANSWERED_BAR = 2/3` over `unanswered_in_bundle`, swept and collapsing at both ends -- at 0.50 eleven real covered questions are marked, at 0.70 the row falls to 5 of 6, at 2/3 the row is 6 of 6 and 0 of 65 real questions are marked. The margin is thin (0.6087 against 0.6667) and is published that way, together with what it does not catch: r761-sk2's own known-negative sits at 0.2857. Row 4: 3 of 6 RED -> 6 of 6 GREEN, with the 10 answered synthetic questions held unmarked as the known-negative. The contract's SS 8 gains point 7, the consumption skill is told to read the block, and the SS 7.4 known-positive moves with the document (14 721/375 -> 16 389/417). Suite 2292 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2588 lines
119 KiB
Python
2588 lines
119 KiB
Python
"""Cut an OKF bundle to one contract-conformant payload for one question.
|
|
|
|
This is the **pre-pass** `docs/consumption-contract.md` SS 1 names: the
|
|
deterministic program that reads the bundle, ranks its concepts, cuts them to a
|
|
bounded set, and emits one payload. It decides nothing about the question being
|
|
asked -- the skill does the judgement, this does the reading, the ranking and
|
|
the cut (SS 2.1).
|
|
|
|
**What it deliberately does not do, and the failure each refusal prevents:**
|
|
|
|
- **It calls no model.** A model in the run path makes the same question at the
|
|
same ref return different bytes, which is the whole property a declared cut
|
|
buys over an emergent one.
|
|
- **It opens no socket.** The repository requires an explicit per-run opt-in for
|
|
network access; this command takes no such flag, so it can never reach one.
|
|
- **It enumerates no directory.** SS 9.2 forbids that unless the named profile
|
|
says the index is derived, and measured 2026-09-07, `entries_match_directory`
|
|
is `True` for `STRICT_V1` alone -- for none of the profiles a segmented v0.2
|
|
bundle could have been built under. So the walk follows the INDEX TREE. That
|
|
costs nothing: measured on the 629-concept K2 bundle, the index walk reaches
|
|
exactly the set `rglob` finds.
|
|
- **It imports nothing outside the standard library and this repository.** The
|
|
package pins exactly one runtime dependency and a packaging test enforces it.
|
|
|
|
**It moved into the package on 2026-09-08 (O5), and the move it was written
|
|
for is the one that happened.** Until then it lived in `tools/`, on the
|
|
argument that staying outside `src/` kept it out of every wheel and left no
|
|
consumer's install surface changed by its existence. That argument held while
|
|
the only reader was this repository. It stopped holding when the generated
|
|
consumption skill became the product: a skill generated from a checkout
|
|
emitted `python3 <absolute path>/tools/okf_consume.py`, four such lines
|
|
measured in a skill generated 2026-09-08, so the skill could not be moved,
|
|
shared, or run by anyone without this clone at that exact path. A pre-pass a
|
|
consumer cannot invoke is not a reading direction they have.
|
|
|
|
The entry point is still `build_payload(...)` with the CLI a thin `main()`,
|
|
which is what made this a move rather than a rewrite. `tools/okf_consume.py`
|
|
remains as a thin wrapper, because the published reproduction blocks name it
|
|
and a measurement whose command no longer runs is a measurement nobody can
|
|
repeat.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import hashlib
|
|
import json
|
|
import math
|
|
import re
|
|
import sys
|
|
import unicodedata
|
|
from collections.abc import Mapping, Sequence
|
|
from dataclasses import dataclass, replace
|
|
from pathlib import Path
|
|
from typing import Literal
|
|
|
|
from .corpus import LOG_NAME
|
|
from .inbox import (
|
|
ADJUDICATION_ADJUDICATED,
|
|
ADJUDICATION_PROPOSED,
|
|
ADJUDICATION_STATES,
|
|
ENCLOSING_SECTION,
|
|
)
|
|
from .materialize import parse_frontmatter
|
|
from .profiles import (
|
|
RESERVED_OKF_TYPE,
|
|
SEGMENTED_OKF_V0_2,
|
|
BundleProfile,
|
|
read_block_mappings,
|
|
unquote_scalar,
|
|
)
|
|
|
|
#: The profile whose index policy reads a faceted, per-directory index -- the
|
|
#: shape both the in-repo golden bundle and the K2 corpus carry. Named as this
|
|
#: instrument's default rather than hard-coded at each call site: a caller with
|
|
#: a differently-shaped bundle passes its own.
|
|
DEFAULT_PROFILE = SEGMENTED_OKF_V0_2
|
|
|
|
#: The algorithm the `ref` names, spelled in the value itself. SS 3.3 requires
|
|
#: "a commit or equivalent content identity" and names no algorithm; a bundle
|
|
#: that is not a git checkout has no commit, so the identity is computed. Naming
|
|
#: the algorithm inline is what lets a consumer reproduce it without a document.
|
|
REF_ALGORITHM = "sha256-tree"
|
|
|
|
#: The one concept suffix this instrument reads, taken from the profile so a
|
|
#: second literal cannot drift from it.
|
|
CONCEPT_SUFFIX = DEFAULT_PROFILE.paths.concept_suffix
|
|
|
|
|
|
def _walk_index_tree(
|
|
bundle_root: Path, *, profile: BundleProfile
|
|
) -> tuple[tuple[str, ...], tuple[str, ...]]:
|
|
"""Every index file and every concept id the index tree reaches.
|
|
|
|
One walk, two results, because the alternative is two walks that can
|
|
disagree -- and a ref computed over a different set than the one that was
|
|
read is an identity for something nobody consumed.
|
|
"""
|
|
index_name = profile.index.name
|
|
suffix = profile.paths.concept_suffix
|
|
indexes: set[str] = set()
|
|
concepts: set[str] = set()
|
|
pending: list[str] = [index_name]
|
|
while pending:
|
|
relative = pending.pop()
|
|
if relative in indexes:
|
|
continue
|
|
indexes.add(relative)
|
|
index_file = bundle_root / relative
|
|
if not index_file.is_file():
|
|
continue
|
|
parent = _parent_dir(relative)
|
|
for line in index_file.read_text(encoding="utf-8").splitlines():
|
|
entry = profile.index.parse_entry(line)
|
|
# `None` is curated prose, not a defect: the index is the one file
|
|
# where this library writes beside somebody else's text.
|
|
if entry is None:
|
|
continue
|
|
target = _join(parent, entry.target)
|
|
if target is None:
|
|
continue
|
|
name = target.rsplit("/", 1)[-1]
|
|
if name == index_name:
|
|
pending.append(target)
|
|
elif name == LOG_NAME:
|
|
# A run's own log -- reachable, but metadata, never a
|
|
# concept. Counted as one it inflated the K2 walk to 630
|
|
# against a 629-concept bundle and let the log rank and cut
|
|
# like real content (measured, S7 F2).
|
|
#
|
|
# THE PRODUCER NO LONGER WRITES THE LINK (2026-09-08, after a
|
|
# consumer's navigator followed it and returned 630 where this
|
|
# counts 629). This branch is not dead: every bundle built
|
|
# between `95eb271` and that change carries it, including the
|
|
# ones consumers are reading today.
|
|
continue
|
|
elif target.endswith(suffix):
|
|
concepts.add(target[: -len(suffix)])
|
|
return _byte_sorted(indexes), _byte_sorted(concepts)
|
|
|
|
|
|
def _byte_sorted(values: set[str]) -> tuple[str, ...]:
|
|
"""Sorted by the UTF-8 encoding, never by the locale.
|
|
|
|
A locale-dependent order makes byte-identical output a claim about the
|
|
machine rather than about the bundle.
|
|
"""
|
|
return tuple(sorted(values, key=lambda value: value.encode("utf-8")))
|
|
|
|
|
|
def enumerate_concepts(
|
|
bundle_root: Path, *, profile: BundleProfile = DEFAULT_PROFILE
|
|
) -> tuple[str, ...]:
|
|
"""Every concept id in the bundle, reached through the index tree.
|
|
|
|
Ids are bundle-relative POSIX paths with the profile's concept suffix
|
|
removed -- slash-preserving, so two same-named concepts in different
|
|
documents stay two concepts.
|
|
|
|
**Not `rglob`.** SS 9.2 forbids the consumer path from enumerating a
|
|
directory unless the named profile says the index is derived. The index is
|
|
safe to rely on here because Door B recomputes it as a projection over the
|
|
whole bundle each round, which is what makes a rebuild from scratch equal an
|
|
incremental update byte for byte.
|
|
"""
|
|
_, concepts = _walk_index_tree(bundle_root, profile=profile)
|
|
return concepts
|
|
|
|
|
|
def _parent_dir(relative: str) -> str:
|
|
"""The directory part of a bundle-relative POSIX path, or `""` at the root.
|
|
|
|
Spelled here rather than as `PurePosixPath(...).parent` so the root case is
|
|
`""` and not `"."` -- joining `"."` would put a `./` segment into every id
|
|
at depth one, and two ids differing only by that segment are two ids for one
|
|
concept.
|
|
"""
|
|
head, sep, _ = relative.rpartition("/")
|
|
return head if sep else ""
|
|
|
|
|
|
def _join(parent: str, target: str) -> str | None:
|
|
"""A relative index target resolved against the index's own directory.
|
|
|
|
Returns `None` for a target that climbs above the bundle root or is
|
|
absolute. A target escaping the root is refused rather than clamped: a
|
|
clamped path names a real file the bundle never pointed at.
|
|
"""
|
|
if target.startswith("/"):
|
|
return None
|
|
parts: list[str] = parent.split("/") if parent else []
|
|
for segment in target.split("/"):
|
|
if segment in ("", "."):
|
|
continue
|
|
if segment == "..":
|
|
if not parts:
|
|
return None
|
|
parts.pop()
|
|
continue
|
|
parts.append(segment)
|
|
return "/".join(parts) if parts else None
|
|
|
|
|
|
def bundle_ref(bundle_root: Path, *, profile: BundleProfile = DEFAULT_PROFILE) -> str:
|
|
"""A content identity for the bundle: `sha256-tree:<hex>`.
|
|
|
|
The digest is taken over the LF-joined, byte-sorted lines
|
|
`<posix-relative-path>\t<sha256 of the file bytes>` across every file the
|
|
INDEX TREE reaches -- the indexes themselves and the concepts they name.
|
|
|
|
**Reachable, not `rglob`, and the cost is stated rather than hidden.** A
|
|
file in the directory that no index names is outside this identity. It is
|
|
also outside what the pre-pass may read (SS 9.2), so an identity covering it
|
|
would be an assertion about bytes this command is forbidden to look at. The
|
|
property that matters holds: every byte that can reach a payload is inside
|
|
the ref, so the ref moves whenever a delivered excerpt could.
|
|
|
|
**Path and bytes, never mtime or size.** A digest over metadata would move
|
|
on a copy and hold still on an edit that preserved length, which is the
|
|
opposite of what an identity is for.
|
|
"""
|
|
indexes, concepts = _walk_index_tree(bundle_root, profile=profile)
|
|
suffix = profile.paths.concept_suffix
|
|
lines: list[bytes] = []
|
|
for relative in (*indexes, *(f"{concept}{suffix}" for concept in concepts)):
|
|
path = bundle_root / relative
|
|
if not path.is_file():
|
|
continue
|
|
digest = hashlib.sha256(path.read_bytes()).hexdigest()
|
|
lines.append(f"{relative}\t{digest}".encode())
|
|
joined = b"\n".join(sorted(lines))
|
|
return f"{REF_ALGORITHM}:{hashlib.sha256(joined).hexdigest()}"
|
|
|
|
|
|
class ConsumeError(Exception):
|
|
"""A refusal this instrument can name.
|
|
|
|
Carries a `code` for the same reason `okf_contract_check.Finding` does: one
|
|
"invalid" verdict over a dozen defects is a diagnostic no caller can act on.
|
|
"""
|
|
|
|
def __init__(self, message: str, *, code: str) -> None:
|
|
super().__init__(message)
|
|
self.code = code
|
|
|
|
|
|
#: The three states a CONSUMER must distinguish (SS 6.1), against the two the
|
|
#: WIRE carries (`inbox.ADJUDICATION_STATES`). The difference is the whole
|
|
#: point: `unknown` is not a value a producer writes, it is what the absence of
|
|
#: the key means, and it is written explicitly here so a reader never has to
|
|
#: infer it from a missing member.
|
|
CONSUMER_ADJUDICATION_STATES = (*ADJUDICATION_STATES, "unknown")
|
|
|
|
AdjudicationState = Literal["proposed", "adjudicated", "unknown"]
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class Concept:
|
|
"""One concept read off disk, with every state written rather than implied."""
|
|
|
|
path: Path
|
|
concept_id: str
|
|
bundle_id: str
|
|
#: `True` when `bundle_id` came from the root index rather than the concept.
|
|
#: Recorded rather than silently defaulted: SS 3.1 makes identity the
|
|
#: `(bundle_id, concept_id)` tuple, so where the first half came from is
|
|
#: part of what the payload is asserting.
|
|
bundle_id_inherited: bool
|
|
sha256: str
|
|
okf_type: str
|
|
title: str
|
|
source_file: str
|
|
adjudication: AdjudicationState
|
|
#: `False` when the key was absent. `adjudication == "unknown"` already says
|
|
#: so, but a separate flag keeps the two facts from being one inference.
|
|
adjudication_present: bool
|
|
#: The identifier the producer wrote, or `""` when there is no key. The
|
|
#: number a lookup question is asked ON, and the one thing a reader needs to
|
|
#: name the concept the answer rests on.
|
|
req_number: str
|
|
#: The SS 5.1 address entries, in either YAML form.
|
|
sources: tuple[Mapping[str, str], ...]
|
|
#: `True` when a `sources` key was there, whatever this reader made of it.
|
|
#: With `sources == ()` that is the third state: present and unreadable.
|
|
sources_present: bool
|
|
#: The `LOCATOR_KEYS` this concept carries, values as written. Absent keys
|
|
#: are absent, never `""`: SS 6.4 forbids reading the absence of a
|
|
#: conditionally-written field as the negation of what it asserts.
|
|
locators: Mapping[str, str]
|
|
frontmatter: Mapping[str, str]
|
|
body: str
|
|
#: The concept this one's `parent` pointer resolves to, as its `concept_id`
|
|
#: and `title`, or `None`. Set by :func:`link_parents` over the whole set
|
|
#: and never by :func:`read_concept`: whether a pointer lands is a fact
|
|
#: about the bundle, and one file does not know its bundle.
|
|
parent: Mapping[str, str] | None = None
|
|
#: `True` when the concept carries a `parent` key that resolves to no
|
|
#: concept of its own document -- the third state, as for `sources`.
|
|
parent_unresolved: bool = False
|
|
|
|
|
|
#: The key a concept names its enclosing section under, and the key that
|
|
#: section answers to. Both are the segmented profiles' (`inbox.py` writes a
|
|
#: plan entry's `parent_id` as `parent:`, naming the other entry's `segment_id`).
|
|
PARENT_KEY = "parent"
|
|
SEGMENT_ID_KEY = "segment_id"
|
|
|
|
|
|
def link_parents(concepts: Sequence[Concept]) -> list[Concept]:
|
|
"""Every concept, with its `parent` pointer resolved or named unresolved.
|
|
|
|
A `parent` value is a `segment_id`, and a segment id is unique only inside
|
|
the plan of ONE document -- `p1` exists in every document of a
|
|
multi-document bundle -- so it is resolved among the concepts sharing this
|
|
concept's `source_file`, never across the bundle. An id two concepts of one
|
|
document claim resolves to neither, the rule `structure._lookup` applies: an
|
|
ambiguous pointer that silently picks a winner is worse than one reported
|
|
unresolved.
|
|
|
|
One pass over concepts already read; no file is opened again. A concept
|
|
carrying no `parent` comes back as it went in, which is what keeps every
|
|
payload of a bundle without the key byte-identical.
|
|
"""
|
|
claims: dict[tuple[str, str], list[Concept]] = {}
|
|
for concept in concepts:
|
|
segment_id = concept.frontmatter.get(SEGMENT_ID_KEY, "").strip()
|
|
if segment_id:
|
|
claims.setdefault((concept.source_file, segment_id), []).append(concept)
|
|
linked: list[Concept] = []
|
|
for concept in concepts:
|
|
pointer = concept.frontmatter.get(PARENT_KEY, "").strip()
|
|
if not pointer:
|
|
linked.append(concept)
|
|
continue
|
|
owners = claims.get((concept.source_file, pointer), [])
|
|
if len(owners) == 1 and owners[0] is not concept:
|
|
target = owners[0]
|
|
linked.append(
|
|
replace(concept, parent={"concept_id": target.concept_id, "title": target.title})
|
|
)
|
|
else:
|
|
linked.append(replace(concept, parent_unresolved=True))
|
|
return linked
|
|
|
|
|
|
def read_concept(path: Path, *, bundle_root: Path, root_bundle_id: str) -> Concept:
|
|
"""One concept file as a record. Reads; derives nothing about the question.
|
|
|
|
`sha256` is the digest of the CONCEPT FILE (SS 3.2), never the
|
|
`source_sha256` frontmatter key -- that one digests the source document the
|
|
concept was extracted from, and conflating them would make the payload's
|
|
content identity point at a PDF nobody in the chain reads. Both exist on
|
|
every K2 concept, which is what makes the confusion available.
|
|
"""
|
|
frontmatter = parse_frontmatter(path)
|
|
relative = path.relative_to(bundle_root).as_posix()
|
|
#: The slash-preserving id: the bundle-relative path minus the suffix. This
|
|
#: instrument's own choice, consistent with the RULE at `importer.py:244`
|
|
#: and `inbox.py:1101-1102` -- but deliberately NOT `importer.import_slug`,
|
|
#: which one line further down flattens the id to a single hyphenated
|
|
#: segment. A flattened id fails a document-prefix match in a way that looks
|
|
#: like a ranking miss rather than an id-format bug.
|
|
concept_id = relative[: -len(CONCEPT_SUFFIX)] if relative.endswith(CONCEPT_SUFFIX) else relative
|
|
raw = frontmatter.get("adjudication")
|
|
if raw is None:
|
|
adjudication: AdjudicationState = "unknown"
|
|
elif raw == ADJUDICATION_PROPOSED:
|
|
adjudication = "proposed"
|
|
elif raw == ADJUDICATION_ADJUDICATED:
|
|
adjudication = "adjudicated"
|
|
else:
|
|
raise ConsumeError(
|
|
f"{relative} carries adjudication={raw!r}, outside the wire set "
|
|
f"{ADJUDICATION_STATES}; refusing to map it to 'unknown', which "
|
|
"would report 'we cannot tell whether it was judged' where the "
|
|
"truth is that the bundle said something this consumer does not "
|
|
"understand",
|
|
code="adjudication_unknown_value",
|
|
)
|
|
declared = frontmatter.get("bundle_id")
|
|
entries, sources_present = read_sources(_frontmatter_lines(path))
|
|
return Concept(
|
|
path=path,
|
|
concept_id=concept_id,
|
|
bundle_id=declared if declared else root_bundle_id,
|
|
bundle_id_inherited=not declared,
|
|
sha256=hashlib.sha256(path.read_bytes()).hexdigest(),
|
|
okf_type=frontmatter.get("type", ""),
|
|
title=frontmatter.get("title", ""),
|
|
source_file=frontmatter.get("source_file", ""),
|
|
adjudication=adjudication,
|
|
adjudication_present=raw is not None,
|
|
req_number=frontmatter.get("req_number", ""),
|
|
sources=entries,
|
|
sources_present=sources_present,
|
|
locators={
|
|
key: value
|
|
for key, value in frontmatter.items()
|
|
if key.startswith(SOURCE_KEY_PREFIX) and value.strip()
|
|
},
|
|
frontmatter=frontmatter,
|
|
body=_body(path),
|
|
)
|
|
|
|
|
|
def _body(path: Path) -> str:
|
|
"""The text after the frontmatter block, or the whole file when there is none."""
|
|
text = path.read_text(encoding="utf-8")
|
|
lines = text.splitlines()
|
|
if not lines or lines[0].strip() != "---":
|
|
return text
|
|
for offset, line in enumerate(lines[1:], start=2):
|
|
if line.strip() == "---":
|
|
return "\n".join(lines[offset:])
|
|
return text
|
|
|
|
|
|
#: SS 6.2, from SPEC SS 5.3, lowest to highest. Imported from the checker's own
|
|
#: constant would be circular (the checker is a separate tool); spelled here and
|
|
#: held to the checker's set by the anti-drift test.
|
|
TrustTier = Literal["unverified", "machine-confirmed", "human-reviewed"]
|
|
|
|
#: The prefix that makes an actor a person. A PREFIX, never a substring: an
|
|
#: actor id `bot/human:2` contains the literal and is a machine, and promoting
|
|
#: it would be fabricated provenance produced by a matching bug.
|
|
HUMAN_ACTOR_PREFIX = "human:"
|
|
|
|
|
|
def trust_tier(verified_raw: str | None) -> TrustTier | None:
|
|
"""The tier `verified` implies, or `None` when the value cannot be read.
|
|
|
|
Three inputs, three different facts, and collapsing any two is the defect:
|
|
|
|
- `None` -- the key is ABSENT. SS 6.2: no `verified` key means
|
|
`unverified`, and SS 6.3 forbids rejecting a concept for it.
|
|
- `""` -- the key is PRESENT and this library cannot read it. Measured
|
|
2026-09-07: the line-oriented `parse_frontmatter` returns `''` for a
|
|
block-form value and the full string for a flow one, so the two are
|
|
distinguishable. Returning `None` here is not a tier; the caller withholds
|
|
the concept under a named rule. Emitting `unverified` instead would assert
|
|
a fact nobody measured, which is exactly what SS 6.4 forbids.
|
|
- a flow value -- decoded, and the tier follows the actors.
|
|
"""
|
|
if verified_raw is None:
|
|
return "unverified"
|
|
if not verified_raw.strip():
|
|
return None
|
|
entries = _parse_flow_mappings(verified_raw)
|
|
if entries is None:
|
|
return None
|
|
actors: list[str] = []
|
|
for entry in entries:
|
|
actor = entry.get("by")
|
|
if not actor:
|
|
raise ConsumeError(
|
|
f"a `verified` entry names no `by` actor ({verified_raw!r}); "
|
|
"refusing to tier it, because the tier IS a claim about who "
|
|
"checked and there is nobody to name",
|
|
code="verified_actorless",
|
|
)
|
|
actors.append(actor)
|
|
if not actors:
|
|
return None
|
|
if any(actor.startswith(HUMAN_ACTOR_PREFIX) for actor in actors):
|
|
return "human-reviewed"
|
|
return "machine-confirmed"
|
|
|
|
|
|
def _parse_flow_mappings(value: str) -> list[dict[str, str]] | None:
|
|
"""A YAML flow sequence of flow mappings, or `None` when it is not one.
|
|
|
|
Modelled on `structure._parse_flow_list` -- flow form only, because this
|
|
library's standing rule is that a value it can write is a value it can read
|
|
back. Not reused: that one splits on every comma, and `{ by: x, at: y }`
|
|
carries a comma INSIDE a mapping, so it would return four fragments where
|
|
there are two pairs.
|
|
"""
|
|
stripped = value.strip()
|
|
if not (stripped.startswith("[") and stripped.endswith("]")):
|
|
return None
|
|
body = stripped[1:-1].strip()
|
|
if not body:
|
|
return []
|
|
mappings: list[dict[str, str]] = []
|
|
for chunk in _split_top_level(body, "{", "}"):
|
|
item = chunk.strip()
|
|
if not (item.startswith("{") and item.endswith("}")):
|
|
return None
|
|
pairs: dict[str, str] = {}
|
|
# Split where a comma is not inside a `"`-quoted leaf, and unquote the
|
|
# leaf: a YAML reader reads `{ title: "a, b" }` as ONE pair (K3-22).
|
|
for field in _split_top_level(item[1:-1], "{", "}"):
|
|
key, separator, raw = field.partition(":")
|
|
if separator:
|
|
pairs[key.strip()] = unquote_scalar(raw.strip())
|
|
mappings.append(pairs)
|
|
return mappings
|
|
|
|
|
|
def _split_top_level(body: str, opener: str, closer: str) -> list[str]:
|
|
"""Split on commas that are not inside a `{...}` or a `"`-quoted scalar.
|
|
|
|
A `"` opens a quoted scalar only where a YAML value can START -- after
|
|
`:`, `,`, the opener, or at the beginning -- so a plain value carrying a
|
|
`"` in its middle is split exactly as it was before K3-22.
|
|
"""
|
|
parts: list[str] = []
|
|
depth = 0
|
|
current: list[str] = []
|
|
quoted = escaped = False
|
|
previous = ""
|
|
for character in body:
|
|
current.append(character)
|
|
if quoted:
|
|
if escaped:
|
|
escaped = False
|
|
elif character == "\\":
|
|
escaped = True
|
|
elif character == '"':
|
|
quoted = False
|
|
previous = character
|
|
continue
|
|
if character == '"' and previous in ("", ":", ",", opener):
|
|
quoted = True
|
|
continue
|
|
if character == opener:
|
|
depth += 1
|
|
elif character == closer:
|
|
depth -= 1
|
|
if character == "," and depth == 0:
|
|
current.pop()
|
|
parts.append("".join(current))
|
|
current = []
|
|
if not character.isspace():
|
|
previous = character
|
|
parts.append("".join(current))
|
|
return [part for part in parts if part.strip()]
|
|
|
|
|
|
#: The locator keys O3 (`b6a8c8b`) writes on every SEGMENTED concept, in the
|
|
#: order they are emitted so a reader comparing an excerpt against the concept
|
|
#: file reads one sequence. The ADDRESS is SPEC SS 5.1's `sources`; these are
|
|
#: this library's OWN top-level keys, because SS 5.1 has no field for a place
|
|
#: within a resource and the guard rejects every route to putting one inside a
|
|
#: `sources` entry. Their VALUES pass through as the frontmatter's own strings:
|
|
#: `source_pages: [2, 27]` reaches the payload as `"[2, 27]"`, which is what
|
|
#: makes an excerpt greppable against the file it came from.
|
|
LOCATOR_KEYS = (
|
|
"source_pages",
|
|
"source_sheet",
|
|
"source_rows",
|
|
"source_lines",
|
|
"source_offset",
|
|
)
|
|
|
|
#: What a locator key looks like to a reader that does not know the producer.
|
|
#: The pass-through rule is this PREFIX and not `LOCATOR_KEYS`, which is a list
|
|
#: of the producers someone thought of: measured 2026-09-08, 269 of 274 concepts
|
|
#: in the N500 bundle carry `source_element_id`, a locator that repository chose
|
|
#: under the same rule ("the key says what it indexes") and this library never
|
|
#: writes. An allowlist drops it, and the excerpt then names a document without
|
|
#: naming the place in it. A PREFIX, never a substring -- `resource_owner`
|
|
#: contains the literal and is not a locator, and promoting it would be
|
|
#: fabricated provenance produced by a matching bug.
|
|
SOURCE_KEY_PREFIX = "source_"
|
|
|
|
|
|
def _frontmatter_lines(path: Path) -> list[str]:
|
|
"""The raw lines between the two `---` fences, indentation intact.
|
|
|
|
`parse_frontmatter` SKIPS indented lines on purpose -- a nested `title:`
|
|
arriving later would SUBSTITUTE for the document's. That refusal is right
|
|
for a flat mapping and it is why the block form has to be read from the
|
|
raw lines instead.
|
|
"""
|
|
lines = path.read_text(encoding="utf-8").splitlines()
|
|
if not lines or lines[0].strip() != "---":
|
|
return []
|
|
block: list[str] = []
|
|
for line in lines[1:]:
|
|
if line.strip() == "---":
|
|
break
|
|
block.append(line)
|
|
return block
|
|
|
|
|
|
def read_sources(lines: Sequence[str]) -> tuple[tuple[Mapping[str, str], ...], bool]:
|
|
"""The SS 5.1 address entries, and whether the key was there at all.
|
|
|
|
Three states, kept apart because collapsing any two reports something
|
|
nobody measured: `((), False)` the concept has no `sources` key; `((), True)`
|
|
it has one this reader cannot decode; a non-empty tuple, the entries.
|
|
|
|
BOTH YAML forms are read, and that is a measurement rather than a
|
|
preference. Measured 2026-09-08: K2 writes the flow form on 629 of 629
|
|
concepts, the N500 bundle writes the block form on 270 of 270. A reader
|
|
handling one form delivers the other bundle with no address at all -- and
|
|
for N500 there is nothing else, because it carries zero locator keys.
|
|
|
|
Reading the block form is not a licence to WRITE it, and since K3-24 the
|
|
reason is a different one: the line-oriented parser DOES round-trip a
|
|
block `sources` list now (it shares this function's grammar), so what
|
|
holds the emission rule at flow is the downstream consumer that classifies
|
|
a block sequence as unreadable provenance -- measured, not remembered, in
|
|
`materialize._render_sources`.
|
|
"""
|
|
for position, line in enumerate(lines):
|
|
if line[:1] in (" ", "\t") or not line.startswith("sources:"):
|
|
continue
|
|
value = line.partition(":")[2].strip()
|
|
if value:
|
|
flow = _parse_flow_mappings(value)
|
|
if flow is None:
|
|
return (), True
|
|
return tuple(flow), True
|
|
# The same grammar the flat readers decode with (K3-24). It lives in
|
|
# `profiles` because `materialize` imports that module and this one
|
|
# imports `materialize`: one loop, or the two readers would answer the
|
|
# same question differently. An indented line before any `- ` opens no
|
|
# entry and comes back as `None`, refused rather than folded into one.
|
|
entries = read_block_mappings(lines, position)
|
|
if entries is None:
|
|
return (), True
|
|
return entries, True
|
|
return (), False
|
|
|
|
|
|
# --- The budget instrument (SS 7) --------------------------------------------
|
|
|
|
#: SS 7.5 fixes no unit deliberately -- "a token is one encoder family's unit
|
|
#: and fixing it would adopt one vendor's arithmetic as everyone's". This
|
|
#: profile chooses utf-8 bytes of the EMITTED JSON, which the repository can
|
|
#: count with no dependency at all. `tiktoken` would be runtime dependency
|
|
#: number two behind a second optional extra plus a tokenizer-version fixture
|
|
#: migration, bought to answer one comparison in its own unit.
|
|
BUDGET_UNIT = "utf-8 bytes of emitted JSON"
|
|
|
|
#: SS 7.1 requires the instrument to be NAMED, not merely used. The name is the
|
|
#: function plus the one flag that changes its answer.
|
|
BUDGET_INSTRUMENT = "okf_consume.measure (len of the ensure_ascii=False JSON encoding, utf-8)"
|
|
|
|
#: Chosen, not derived, and the reason is a measurement rather than a taste:
|
|
#: at 60 000 the largest realistic gold concept (101 313 B encoded) falls to the
|
|
#: `over_budget_alone` pre-exclusion, so a CORRECT implementation would fail its
|
|
#: own acceptance criteria. At 120 000 that concept fits with 18 424 B of
|
|
#: headroom, and 3 of the K2 corpus's 629 concepts still cannot fit alone
|
|
#: (4 at 60 000). A starting point to be moved by measurement.
|
|
DEFAULT_LIMIT = 120_000
|
|
|
|
#: The known-positive artefact (SS 7.4). A SHIPPED file rather than the bundle
|
|
#: under test, because a per-bundle known-positive can only be one of two
|
|
#: useless things: a constant that is wrong for every bundle but one, or the
|
|
#: instrument's own output, which makes `expected == measured` true by
|
|
#: construction and the rule decorative.
|
|
#:
|
|
#: The coupling is stated rather than hidden: if this document's bytes move, the
|
|
#: literal below goes stale and the pre-pass refuses until it is updated. That
|
|
#: is the intended direction -- a stale known-positive is a loud failure, and
|
|
#: the document is normative and not edited from this repository.
|
|
KNOWN_POSITIVE_CASE = "docs/consumption-contract.md, encoded as a JSON string"
|
|
|
|
#: `measure()`'s own answer for that file. Vacuous ALONE -- which is why the
|
|
#: delta below exists.
|
|
KNOWN_POSITIVE_EXPECTED = 16_389
|
|
|
|
#: The second, independent route. `wc -c` reports 15 972 raw bytes for the same
|
|
#: file; the difference is this file's JSON quoting and escaping overhead. A
|
|
#: reader can derive it without running `measure()` at all, and it moves the
|
|
#: moment `measure()` changes what it counts -- which is what stops
|
|
#: `expected == measured` from proving nothing.
|
|
KNOWN_POSITIVE_ENCODING_DELTA = 417
|
|
|
|
#: The two places that file can be, resolved in this order.
|
|
#:
|
|
#: A wheel carries `_data/consumption-contract.md`, force-included from the
|
|
#: authored document at build time; a source tree carries only the authored
|
|
#: document. The known-positive is part of the INSTRUMENT, not of the source
|
|
#: tree it was written in -- packaging the pre-pass without it would ship a
|
|
#: command that refuses every bundle on a control it cannot run. One authored
|
|
#: copy either way: a committed duplicate would let the literal above be
|
|
#: checked against a document nobody edits.
|
|
_PACKAGED_KNOWN_POSITIVE = Path(__file__).resolve().parent / "_data" / "consumption-contract.md"
|
|
_AUTHORED_KNOWN_POSITIVE = Path(__file__).resolve().parents[2] / "docs" / "consumption-contract.md"
|
|
|
|
|
|
def known_positive_path() -> Path:
|
|
"""Where the known-positive artefact is, or a coded refusal.
|
|
|
|
Refuses rather than skipping the control: SS 7.4 makes the known-positive
|
|
the thing that stops `expected == measured` from being vacuous, and a
|
|
pre-pass that quietly ran without it would emit payloads whose cost figures
|
|
rest on an instrument nobody checked.
|
|
"""
|
|
for candidate in (_PACKAGED_KNOWN_POSITIVE, _AUTHORED_KNOWN_POSITIVE):
|
|
if candidate.is_file():
|
|
return candidate
|
|
raise ConsumeError(
|
|
f"the known-positive artefact was not found at {_PACKAGED_KNOWN_POSITIVE} "
|
|
f"or {_AUTHORED_KNOWN_POSITIVE}",
|
|
code="known_positive_missing",
|
|
)
|
|
|
|
|
|
def measure(value: str) -> int:
|
|
"""The cost of `value` in the unit the gate enforces.
|
|
|
|
The ENCODED JSON form, because that is what the payload actually costs. A
|
|
knapsack weighing `stat().st_size` while the gate measures this would let a
|
|
cut computed as fitting be refused by the gate -- measured, the two differ
|
|
by 7.1 % over the K2 corpus.
|
|
"""
|
|
return len(json.dumps(value, ensure_ascii=False).encode("utf-8"))
|
|
|
|
|
|
def known_positive() -> tuple[str, int, int]:
|
|
"""The case, the figure expected of it, and the figure measured (SS 7.4).
|
|
|
|
Takes no bundle argument on purpose: see `KNOWN_POSITIVE_CASE`.
|
|
"""
|
|
measured = measure(known_positive_path().read_text(encoding="utf-8"))
|
|
return KNOWN_POSITIVE_CASE, KNOWN_POSITIVE_EXPECTED, measured
|
|
|
|
|
|
# --- Stage one: which documents are worth opening -----------------------------
|
|
|
|
#: The shortest token this instrument scores. Two characters in Norwegian are
|
|
#: almost always a function word (`og`, `er`, `en`, `av`, `de`), and a matcher
|
|
#: that scores them ranks every document equally.
|
|
MIN_TOKEN_LENGTH = 3
|
|
|
|
#: How many leading characters two tokens must share to count as a match.
|
|
#:
|
|
#: THIS INSTRUMENT'S OWN CONSTANT, and a measurement rather than a preference.
|
|
#: Token equality fails on Norwegian compounds: a question's inflected noun
|
|
#: equals none of the tokens in a concept's `title`, `source_file` or path when
|
|
#: the concept spells the same subject as a compound. Plain substring
|
|
#: containment does not save it either -- of `varene` and `varemottak`, neither
|
|
#: contains the other. A shared prefix does: `vare|ne` and `vare|mottak` share 4.
|
|
#:
|
|
#: MEASURED 2026-09-07 over a 629-concept corpus, for one question's subject
|
|
#: token: a 4-character floor matches **3** concepts -- over the concept id
|
|
#: alone AND over title + `source_file` + id together, the same 3 -- and the
|
|
#: gold concept is among them. The plan this implements recorded 6 for the same
|
|
#: measurement; 6 is not reproducible with this rule, and the number that is
|
|
#: reproducible is the one carried here. A 3-character floor over-matches
|
|
#: Norwegian function words.
|
|
MIN_SHARED_PREFIX = 4
|
|
|
|
_TOKEN_SPLIT_RE = re.compile(r"[^0-9a-zà-öø-ÿ]+")
|
|
|
|
#: Every dash a source spells an identifier's separator with, folded to the
|
|
#: ASCII hyphen. NFC folds NONE of them, so `10.2—2` from a document viewer and
|
|
#: `10.2-2` from a person typing the question are two different tokens until
|
|
#: this table runs. The set is the Unicode dash block plus the minus sign.
|
|
_DASH_TO_HYPHEN = str.maketrans(dict.fromkeys("‐‑‒–—―−", "-"))
|
|
|
|
#: An identifier: NUMERIC groups joined by `.` or `-`, with an optional letter
|
|
#: prefix that touches its digits without a separator (`R610.4`).
|
|
#:
|
|
#: THE LETTERS ARE THE POINT, and this pattern was narrowed by a measurement
|
|
#: rather than written this way. A rule that joined alphanumeric groups across
|
|
#: a separator swallowed a whole document slug -- `...bilag-3-6-premissrapport-
|
|
#: akustikk` became ONE token because `3-6` sits inside it -- and that
|
|
#: document's score for a question naming its subject fell from 0.735 to 0.0,
|
|
#: taking a hit@8 row with it. Only digits may stand on either side of a
|
|
#: separator, so a hyphenated word keeps its words.
|
|
_IDENTIFIER_RE = re.compile(r"[a-zà-öø-ÿ]*[0-9]+(?:[.-][0-9]+)+")
|
|
|
|
|
|
def normalise(text: str) -> tuple[str, ...]:
|
|
"""Text as comparable tokens: NFC, casefold, dash-fold, then split.
|
|
|
|
NFC FIRST is load-bearing and not tidiness. macOS hands filenames over
|
|
decomposed, so `å` arrives as `a` + U+030A; the combining ring is not a word
|
|
character, so an un-normalised split turns `årlig` into `a` and `rlig` and
|
|
the term is silently lost. `æ` and `ø` have no canonical decomposition, so a
|
|
test built on either passes while the bug is live -- which is why the
|
|
known-positive for this function uses `å`.
|
|
|
|
IDENTIFIERS SURVIVE THE SPLIT. Splitting on every non-alphanumeric turns a
|
|
requirement number into digit runs, and `MIN_TOKEN_LENGTH` then removes
|
|
them: `Krav 10.2—2` reached the ranker as `krav` alone, a word every concept
|
|
in a standards bundle carries, so the ranking became a corpus-wide tie and
|
|
the named requirement was withheld `below_k` (measured 2026-09-08 on three
|
|
bundles). The floor stays -- a bare `10` matches every page number in a
|
|
corpus -- and the identifier is exempted from it rather than the floor
|
|
lowered for everyone.
|
|
"""
|
|
folded = unicodedata.normalize("NFC", text).casefold().translate(_DASH_TO_HYPHEN)
|
|
tokens: list[str] = []
|
|
position = 0
|
|
for match in _IDENTIFIER_RE.finditer(folded):
|
|
tokens.extend(_split(folded[position : match.start()]))
|
|
tokens.append(match.group())
|
|
position = match.end()
|
|
tokens.extend(_split(folded[position:]))
|
|
return tuple(tokens)
|
|
|
|
|
|
def _split(folded: str) -> list[str]:
|
|
"""The generic split, on text already folded by `normalise`."""
|
|
return [token for token in _TOKEN_SPLIT_RE.split(folded) if len(token) >= MIN_TOKEN_LENGTH]
|
|
|
|
|
|
def is_identifier(token: str) -> bool:
|
|
"""Whether a token is a NUMBER a document is known by, rather than a word.
|
|
|
|
The whole token, never a part of one: `normalise` emits an identifier as
|
|
one token, so a full match is what "this token is an identifier" means. A
|
|
bare number is not one -- `2023` has no separator, and every page number in
|
|
a corpus would become an identifier if it were.
|
|
"""
|
|
return _IDENTIFIER_RE.fullmatch(token) is not None
|
|
|
|
|
|
def tokens_match(left: str, right: str, *, stems: frozenset[str] | None = None) -> bool:
|
|
"""Whether two tokens share a leading prefix of at least `MIN_SHARED_PREFIX`.
|
|
|
|
**THE SHARED PREFIX MUST BE A WORD** when `stems` is given, and that is
|
|
round 10's repair. The floor alone matched four characters that are not a
|
|
stem at all: measured on the pinned 453-concept bundle with the control run
|
|
first, `under` occurs 79 times by equality and matches 172 concepts by
|
|
prefix, while `undersjoisk` occurs 0 times and matches the same 172;
|
|
`bilateral` occurs 0 times and matches 400 of 453 through `bilag`;
|
|
`standhaftig` 0 and 219 through `standard`.
|
|
|
|
Three repairs were measured and all three failed on the same row: a longer
|
|
floor (5-8), a coverage share of the question word (0.5-0.8), and a floor
|
|
applied only to long words (>= 6, 8, 10, 12). Row 1's question token
|
|
`prisene` reaches its gold document through `pris|sammenstilling` on the
|
|
four characters `pris` -- 0.57 of the question word and 0.22 of the
|
|
document word -- so the over-match and the wanted match are the same
|
|
mechanism seen from two sides, and no threshold on length or coverage
|
|
separates them.
|
|
|
|
What separates them is that `pris` is a word and `bila` is not. With
|
|
`stems`, `bilateral` falls to 0 on both bundles and every hit@8 row keeps
|
|
its rank. `undersjoisk` still reaches 162 because it shares `under`, which
|
|
IS a word here -- a genuine Norwegian morpheme, so that residual is a
|
|
different answer rather than a ceiling.
|
|
|
|
The vocabulary is the BUNDLE's own, which makes a payload depend on the
|
|
corpus the way `rarity_weights` already does. Passing `None` reproduces the
|
|
pre-round-10 matcher exactly.
|
|
|
|
Symmetric, and it degrades to equality for short tokens: two 4-character
|
|
tokens match only if they are the same word.
|
|
|
|
**AN IDENTIFIER MATCHES BY EQUALITY ALONE**, and that is a defect fix
|
|
measured on the case it costs most rather than a preference. The prefix
|
|
rule was measured for Norwegian compounds, where `vare|ne` and
|
|
`vare|mottak` share a stem; a requirement number has no stem, and four
|
|
leading characters of `3.3.1-13` are four leading characters of every
|
|
requirement in section 3.3. MEASURED 2026-09-08 on a 446-concept bundle:
|
|
the unique identifier `3.3.1-13` reached **135** concepts under the prefix
|
|
rule and **1** under equality, which made the rarity weight rank a common
|
|
adjective as more informative than the number naming the document
|
|
(`docs/2026-09-08-sjeldenhetsvekt.md` SS 3). `54a0bc2` SS 1 named this
|
|
class -- "`df` measured over the colliding matcher measures collision
|
|
breadth, not rarity" -- and this is that sentence applied to the identifier
|
|
itself.
|
|
"""
|
|
if is_identifier(left) or is_identifier(right):
|
|
# No floor, either: `MIN_SHARED_PREFIX` made a three-character
|
|
# identifier match NOTHING, not even itself. Measured on a 629-concept
|
|
# bundle, `9.2` reached 0 concepts under the matcher while sitting
|
|
# verbatim in one title, so a document known by a short number was
|
|
# unreachable by that number.
|
|
return left == right
|
|
limit = min(len(left), len(right))
|
|
if limit < MIN_SHARED_PREFIX:
|
|
return False
|
|
shared = 0
|
|
while shared < limit and left[shared] == right[shared]:
|
|
shared += 1
|
|
if shared < MIN_SHARED_PREFIX:
|
|
return False
|
|
if stems is None:
|
|
return True
|
|
# Equality first, and only inside this branch. A token always answers to
|
|
# itself, whatever the corpus contains -- but the check must NOT move above
|
|
# the floor, where it would make `veg`/`veg` match and change the shipped
|
|
# rule for every token shorter than `MIN_SHARED_PREFIX`.
|
|
return left == right or left[:shared] in stems
|
|
|
|
|
|
#: One declared vocabulary family, spelled once: within it, any term answers to
|
|
#: any other. OFF by default and reachable only through `--cost-vocabulary`.
|
|
#:
|
|
#: WHAT IT IS FOR, and the failure it addresses. Measured 2026-09-08 over a
|
|
#: 629-concept corpus: a mandate-shaped question about cost ranked that
|
|
#: corpus's one priced table 249th of 269 lexical candidates, because the
|
|
#: question said `kostnadsbesparelser` and the document said `pris` -- two
|
|
#: words with no shared prefix. Its title score and its document score were
|
|
#: both 0. No value of `k` closes a gap in the VOCABULARY.
|
|
#:
|
|
#: WHAT IT IS NOT. Each member must be at least `MIN_SHARED_PREFIX` characters
|
|
#: or it can never match anything (`sum` is 3 and does not match `Summen`; it
|
|
#: was dropped for that reason, not by taste). The bridge is symmetric and
|
|
#: needs a family term on BOTH sides, so it can widen a cost question towards a
|
|
#: cost document and never towards an arbitrary one. Norwegian, and stated as
|
|
#: such: a corpus in another language gets nothing from it.
|
|
#:
|
|
#: HONESTY, measured rather than asserted: on that corpus the whole effect
|
|
#: rests on `kost` and `pris`. Removing either returns the priced table to rank
|
|
#: 249; removing any other member moves it not at all, and three members match
|
|
#: nothing in that corpus. They are kept because dropping a term for being
|
|
#: absent from ONE corpus fits the list to that corpus.
|
|
COST_VOCABULARY = (
|
|
"beløp",
|
|
"budsjett",
|
|
"enhet",
|
|
"honorar",
|
|
"kost",
|
|
"kroner",
|
|
"mengde",
|
|
"pris",
|
|
"utgift",
|
|
"vederlag",
|
|
)
|
|
|
|
|
|
def in_cost_vocabulary(token: str) -> bool:
|
|
"""Whether one token belongs to the declared family, by the same prefix rule."""
|
|
return any(tokens_match(token, member) for member in COST_VOCABULARY)
|
|
|
|
|
|
def question_uses_cost_vocabulary(question: str) -> bool:
|
|
"""Whether the QUESTION opens the bridge. The gate is the question, never the flag.
|
|
|
|
A question naming no term in the family gets byte-identical bytes with the
|
|
flag set, which is what keeps the flag a widening of one question class
|
|
rather than a second ranker.
|
|
"""
|
|
return any(in_cost_vocabulary(token) for token in normalise(question))
|
|
|
|
|
|
#: The DOOR's line, exactly as `inbox._link_enclosing` writes it: the two SS 6.1
|
|
#: words, a markdown link, and a bundle-absolute target. Built from the door's
|
|
#: own constant, never from prose, so a rename there breaks the recognition
|
|
#: instead of quietly loosening it.
|
|
_LINK_LINE = re.compile(rf"^{re.escape(ENCLOSING_SECTION)}: \[[^\]]*\]\(/[^)]*\)$")
|
|
|
|
|
|
def body_without_link_line(body: str) -> str:
|
|
"""`body` minus the one line the door appends to a heading-only body.
|
|
|
|
THE FILE KEEPS THE LINE AND SO DOES THE READER. Since K3-25 this is the
|
|
DEFAULT path through the ranker (`DEFAULT_LINK_IN_SIGNAL`), on the same
|
|
bytes and in the same process -- never a second build compared against a
|
|
first. `link_in_signal=True` is how a caller asks for the older reading.
|
|
|
|
Three conditions, all the door's own, and each one is a known-negative that
|
|
would otherwise be stripped: the line is LAST, a blank line stands before
|
|
it, and it matches the door's form (`_LINK_LINE`). A human sentence opening
|
|
with the same two words has no link and survives; the door's exact form with
|
|
prose under it is not last and survives. Those two are now on the default
|
|
path rather than behind a parameter, which is what makes them load-bearing
|
|
rather than illustrative. What cannot be told apart is a
|
|
human line that is byte-for-byte the door's own, last in a body -- an
|
|
ambiguity of the form itself, stated rather than hidden.
|
|
"""
|
|
lines = body.rstrip("\n").split("\n")
|
|
if len(lines) < 3 or lines[-2].strip() or not _LINK_LINE.match(lines[-1]):
|
|
return body
|
|
return "\n".join(lines[:-2]) + "\n"
|
|
|
|
|
|
#: Round 25 (K3-25). Whether the DOOR's `Enclosing section:` line is scored.
|
|
#: `False` since K3-25, on K3-23's decomposition: of the newcomers that line
|
|
#: ever added a question token to, **39 of 39** gained it from the
|
|
#: bundle-absolute PATH and **0 of 39** from the link's title, and every token
|
|
#: the path contributed is a segment of the document's own directory -- the
|
|
#: saturation `shared_id_prefix` takes OUT of the id signal, arriving back
|
|
#: through the body. Under this reading a `--shell-parent` bundle delivers
|
|
#: exactly what the unflagged build delivers: the same list, the same order and
|
|
#: the same `spent` on **16 of 16** rows, with hit@1/8/50 **6/6** and the
|
|
#: known-positive at rank 1 under every reading. It moves no byte of any bundle
|
|
#: shipped today -- **0 of 5** carry the line, so **5 of 5** payloads are
|
|
#: byte-identical -- and the day one does, it is the day the path would have
|
|
#: started costing rank instead. A constant rather than three naked literals
|
|
#: because the other eight defaults in this module are constants, and because a
|
|
#: test can then fall on the rule from the module that owns it.
|
|
#: `docs/2026-09-12-k3-runde23-stien-i-kroppssignalet.md` SS 4.
|
|
DEFAULT_LINK_IN_SIGNAL = False
|
|
|
|
|
|
def searchable_text(
|
|
concepts: Sequence["Concept"], *, link_in_signal: bool = DEFAULT_LINK_IN_SIGNAL
|
|
) -> list[str]:
|
|
"""The text a concept is scored against, one string per concept.
|
|
|
|
The same two fields the ranker's two lexical signals read -- title plus
|
|
id, and body -- joined, so a `df` counted here is a `df` over exactly what
|
|
a hit can be scored on. Counting rarity over one field and matching on
|
|
another would weight a token by how rare it is somewhere it is not read.
|
|
|
|
`link_in_signal` reads the bodies with or without the door's link line, and
|
|
the same reason binds it here: the stem vocabulary and the rarity `df` are
|
|
counted over what a hit is scored on, so a reading that took the line out of
|
|
the signal and left it in the vocabulary would measure two different texts.
|
|
Since K3-25 the default is `False` (`DEFAULT_LINK_IN_SIGNAL`).
|
|
"""
|
|
return [
|
|
f"{concept.title} {concept.concept_id.replace('/', ' ')} "
|
|
f"{concept.body if link_in_signal else body_without_link_line(concept.body)}"
|
|
for concept in concepts
|
|
]
|
|
|
|
|
|
def rarity_weights(
|
|
question_tokens: Sequence[str],
|
|
corpus: Sequence[str],
|
|
*,
|
|
stems: frozenset[str] | None = None,
|
|
) -> dict[str, float]:
|
|
"""What one hit on each question token is worth, from the bundle alone.
|
|
|
|
`log(N / df)`: `N` concepts, and `df` the number of them bearing the token
|
|
under the SAME prefix rule a hit is scored with. No constant is set by
|
|
hand and no class of token is declared anywhere -- a word every concept
|
|
carries weighs exactly `log(1) == 0` of itself, and an identifier one
|
|
concept carries takes the corpus's maximum of itself.
|
|
|
|
**What this is NOT.** `54a0bc2` swept smoothed IDF as a GATE -- a threshold
|
|
below which a concept is withheld -- and falsified it: no threshold zeroed
|
|
both known-negatives while any positive question still reached its gold
|
|
document. That result stands and is not re-litigated here. This is the
|
|
other use: an ordering inside the candidate set, with the gate untouched
|
|
and `lexical` still a count. A ranking cannot withhold anything, so the
|
|
failure mode that refuted the gate has no counterpart here.
|
|
|
|
**`df` is measured over the colliding matcher, and so measures collision
|
|
breadth as well as rarity** (`54a0bc2` § 1: every `brann*` compound shares
|
|
four leading characters). Inherited deliberately rather than fixed here:
|
|
the weight must agree with the matcher it weights, and changing the matcher
|
|
is a different change with its own measurement.
|
|
|
|
One pass over the corpus. A token borne by no concept takes the weight of
|
|
a token borne by one -- it is never consumed, because a token that matches
|
|
nothing is never a hit.
|
|
"""
|
|
total = len(corpus)
|
|
if total == 0:
|
|
return {token: 0.0 for token in question_tokens}
|
|
counts = {token: 0 for token in question_tokens}
|
|
for text in corpus:
|
|
candidate_tokens = normalise(text)
|
|
for token in counts:
|
|
if any(tokens_match(token, other, stems=stems) for other in candidate_tokens):
|
|
counts[token] += 1
|
|
return {
|
|
token: math.log(total / count) if count else math.log(total)
|
|
for token, count in counts.items()
|
|
}
|
|
|
|
|
|
def lookup_hits(concepts: Sequence["Concept"], question: str) -> tuple[str, ...]:
|
|
"""The concepts a question NAMES, rather than the ones it describes.
|
|
|
|
A question carrying an identifier that sits VERBATIM in a concept's title
|
|
or id is a lookup, not a search: the reader already knows which document
|
|
they want and is spelling its number. Returns those concepts' ids, byte
|
|
sorted so several holders of one number arrive in a declared order, and the
|
|
EMPTY tuple whenever the question carries no identifier -- which is what
|
|
makes this rule invisible to every question that is not a lookup.
|
|
|
|
**It reads the text the title-and-id signal reads, and no frontmatter key
|
|
list is declared.** Measured 2026-09-08 on three real bundles: of the 1 846
|
|
concepts carrying a `req_number`, the identifier in that key is ALSO in the
|
|
title on **1 846** of them, and on **0** does the key carry an identifier
|
|
the title lacks. A key list would therefore have bought nothing here and
|
|
would have been a constant no measurement asked for. A bundle whose
|
|
identifiers live only in frontmatter is not served by this rule, and that
|
|
is stated rather than guessed at.
|
|
|
|
**Verbatim after `normalise`, so the three spellings of one identifier are
|
|
one lookup** (`_DASH_TO_HYPHEN`) -- but `.` and `-` are NOT interchangeable,
|
|
so a question spelling `1.10` does not find a document whose id spells it
|
|
`1-10`. Measured and left open.
|
|
"""
|
|
identifiers = {token for token in normalise(question) if is_identifier(token)}
|
|
if not identifiers:
|
|
return ()
|
|
return tuple(
|
|
sorted(
|
|
concept.concept_id
|
|
for concept in concepts
|
|
if identifiers
|
|
& set(normalise(f"{concept.title} {concept.concept_id.replace('/', ' ')}"))
|
|
)
|
|
)
|
|
|
|
|
|
def _overlap(
|
|
question_tokens: Sequence[str],
|
|
candidate: str,
|
|
*,
|
|
cost_vocabulary: bool = False,
|
|
weights: Mapping[str, float] | None = None,
|
|
stems: frozenset[str] | None = None,
|
|
) -> float:
|
|
"""What the candidate text answers of the question.
|
|
|
|
A COUNT when `weights` is None -- one per question token the candidate
|
|
answers to, which is what every caller got before rarity weighting existed
|
|
and what the cut still reads. With `weights`, the sum of those tokens'
|
|
rarity weights instead.
|
|
"""
|
|
candidate_tokens = normalise(candidate)
|
|
bridged = cost_vocabulary and any(in_cost_vocabulary(token) for token in candidate_tokens)
|
|
return sum(
|
|
1 if weights is None else weights.get(token, 1.0)
|
|
for token in question_tokens
|
|
if any(tokens_match(token, other, stems=stems) for other in candidate_tokens)
|
|
or (bridged and in_cost_vocabulary(token))
|
|
)
|
|
|
|
|
|
#: How a document's prior grows with its unit count. `1.0` is a DENSITY and
|
|
#: `0.0` is a SUM; this is the classical length normalisation between them, and
|
|
#: it is here rather than inline because the value is a decision a reader should
|
|
#: find where the decision was taken.
|
|
#:
|
|
#: WHY IT MOVED (2026-09-09). A sum measures size -- that is why the density
|
|
#: replaced it -- but a density is diluted by every unit carrying none of the
|
|
#: question, so a document the segmenter split from 1 concept into 12 lost its
|
|
#: prior by a factor of 12. That put the segmentation side and the retrieval
|
|
#: side in direct competition over one number, and it is what blocked a
|
|
#: reference-improving default from shipping.
|
|
#:
|
|
#: SWEPT, not chosen: the gold document's rank under this prior over 6 questions
|
|
#: x 3 bundles = 18 rows, at 0.0, 0.25, 0.5, 0.75 and 1.0. 0.5 is at least as
|
|
#: good as the delivered 1.0 on all 18 rows and strictly better on three; 0.25
|
|
#: loses one row and 0.0 and 0.75 are measured beside it. HONESTY LIMIT: four
|
|
#: alternatives on 18 rows, one gold set, one rater -- and the rank of the
|
|
#: PRIOR is not the rank of the excerpt, because RRF fuses it with two other
|
|
#: signals. The end-to-end hit@8 measurement is the one that decided it.
|
|
DOCUMENT_PRIOR_EXPONENT = 0.5
|
|
|
|
|
|
def document_scores(
|
|
bundle_root: Path,
|
|
question: str,
|
|
*,
|
|
profile: BundleProfile = DEFAULT_PROFILE,
|
|
cost_vocabulary: bool = False,
|
|
weights: Mapping[str, float] | None = None,
|
|
stems: frozenset[str] | None = None,
|
|
) -> dict[str, float]:
|
|
"""One score per top-level document, from the indexes and the paths alone.
|
|
|
|
A "document" is a top-level entry: a directory, or -- per the corpus's own
|
|
shape -- a concept sitting at the root, which has no directory to inherit
|
|
from and is therefore its own document. Measured on K2, 11 of 629 concepts
|
|
are root-level, and they are exactly the 11 carrying neither `adjudication`
|
|
nor `bundle_id`; scoring them as members of some parent would put one bug in
|
|
three places.
|
|
|
|
**The score grows SUBLINEARLY with the unit count** -- `total /
|
|
n**DOCUMENT_PRIOR_EXPONENT`, the exponent at 0.5. Both endpoints are wrong
|
|
and each is wrong in its own direction; the constant above carries the
|
|
measurement and the sweep. The original correction, from a sum to a
|
|
density, is kept here because it is still the reason a sum is not used:
|
|
|
|
**A sum is not a score, and that is a correction rather than a
|
|
preference.** A sum over a document's units grows with the number of units,
|
|
so a large document outscores a small one on size alone. Measured on K2 for
|
|
the price question: the competition document sums to 6.0 over 79 concepts
|
|
(0.076 each) and the price document to 2.0 over 1 (2.0), so the sum ranks
|
|
the larger document three times higher while the density ranks the smaller
|
|
one twenty-six times higher. A prior that grows with size is measuring size.
|
|
|
|
**Stated because it bears on how the hit@k number should be read:** this
|
|
defect was found by running the one question whose gold was confirmed
|
|
independently, and fixing it therefore happened with that answer visible.
|
|
The fix is justified by the scoring function's own arithmetic rather than by
|
|
the answer -- but the ranker is not blind to that one case, and the
|
|
measurement document says so beside the number.
|
|
|
|
Reads the INDEX TREE only. No directory is enumerated here or anywhere else
|
|
in this command (SS 9.2).
|
|
"""
|
|
question_tokens = normalise(question)
|
|
bridge = cost_vocabulary and question_uses_cost_vocabulary(question)
|
|
indexes, concepts = _walk_index_tree(bundle_root, profile=profile)
|
|
totals: dict[str, float] = {}
|
|
units: dict[str, int] = {}
|
|
|
|
def record(document: str, overlap: float) -> None:
|
|
totals[document] = totals.get(document, 0.0) + float(overlap)
|
|
units[document] = units.get(document, 0) + 1
|
|
|
|
for concept_id in concepts:
|
|
record(
|
|
concept_id.split("/", 1)[0],
|
|
_overlap(
|
|
question_tokens,
|
|
concept_id.replace("/", " "),
|
|
cost_vocabulary=bridge,
|
|
weights=weights,
|
|
stems=stems,
|
|
),
|
|
)
|
|
for relative in indexes:
|
|
document = relative.split("/", 1)[0]
|
|
if document == profile.index.name:
|
|
continue
|
|
for line in (bundle_root / relative).read_text(encoding="utf-8").splitlines():
|
|
entry = profile.index.parse_entry(line)
|
|
if entry is None:
|
|
continue
|
|
record(
|
|
document,
|
|
_overlap(
|
|
question_tokens,
|
|
entry.label,
|
|
cost_vocabulary=bridge,
|
|
weights=weights,
|
|
stems=stems,
|
|
),
|
|
)
|
|
return {
|
|
document: totals[document] / units[document] ** DOCUMENT_PRIOR_EXPONENT
|
|
for document in totals
|
|
}
|
|
|
|
|
|
# --- Stage two: which concepts inside those documents -------------------------
|
|
|
|
#: Reciprocal Rank Fusion's smoothing constant. Lifted IN METHOD from the
|
|
#: reference implementation this repository read (`wiki-advise/rank.mjs:82`),
|
|
#: not in code and not as a quality claim: that engine's recall figures were
|
|
#: measured on an LLM candidate-generation pass this pre-pass deliberately does
|
|
#: not have, so its numbers say nothing about this one.
|
|
#:
|
|
#: RRF is chosen because it consumes RANKS ONLY. There is no score
|
|
#: normalisation to get wrong, no weight to tune against a test set, and the
|
|
#: fusion is invariant to any monotone transform of the individual signals.
|
|
RRF_K = 60
|
|
|
|
#: Whether concepts a signal scores EQUALLY share that group's first rank
|
|
#: instead of being ordered inside it by `concept_id`. ON since 2026-09-10.
|
|
#:
|
|
#: WHY IT MOVED. It shipped OFF on 2026-09-08 on a measurement -- hit@8 over
|
|
#: the six published questions fell 5 of 6 to 4 of 6 on the 629-concept bundle.
|
|
#: Round 7 re-measured that fall and it is CONDITIONAL on the document prior
|
|
#: being a sum: swept over `DOCUMENT_PRIOR_EXPONENT` x 3 bundles x 6 rows, the
|
|
#: lost row is lost at exponent 1.0 and held at 0.5. Round 6 moved that
|
|
#: exponent to 0.5 for an unrelated reason and nobody re-measured the pair, so
|
|
#: a rule was left off by a cost that had already been removed.
|
|
#:
|
|
#: WHAT IT BUYS. It is the repair for the defect that kept `--sheet-section-rows
|
|
#: --keep-table-heading` off the build default: a document split from 1 concept
|
|
#: into 12 puts its own 12 concepts in the prior signal's whole top tie group,
|
|
#: so the concept leading the body signal takes prior position 11 rather than
|
|
#: 1 and the document loses fused rank 1 to a single-concept competitor that
|
|
#: leads nothing. Fine-graining was being punished for being fine-grained.
|
|
#:
|
|
#: Measured at exponent 0.5 on three bundles, ranks per row: `[1, 1, 1, 1, 1,
|
|
#: None]` on all three with it on, against `[2, 1, 1, 1, 1, None]` on the split
|
|
#: bundle with it off. Opt-out `--no-tie-shared-rank`.
|
|
DEFAULT_TIE_SHARED_RANK = True
|
|
|
|
|
|
#: Round 10. `MIN_SHARED_PREFIX = 4` matches on four characters whether or not
|
|
#: they are a stem. Measured on the pinned 453-concept bundle, control first:
|
|
#: `bilateral` occurs 0 times by equality and matches 400 of 453 through
|
|
#: `bilag`; `standhaftig` 0 and 219 through `standard`; `undersjoisk` 0 and 172
|
|
#: through `under`. Requiring the shared prefix to occur as a token in the
|
|
#: bundle's own concepts takes the first to 0 and the second to 56 on the
|
|
#: default bundle (0 and 33 on Arm B) with every hit@8 row keeping rank 1 on
|
|
#: BOTH bundles.
|
|
#:
|
|
#: ON by measurement, not by taste, and the measurement is that the other three
|
|
#: candidates are not available: a longer floor (5-8), a coverage share
|
|
#: (0.5-0.8) and a floor for long words only (>= 8) each cost row 1 its rank on
|
|
#: the default bundle and the whole row on Arm B. Row 1 reaches its gold
|
|
#: document through `pris|sammenstilling` on the four characters `pris`, so the
|
|
#: over-match and the wanted match are one mechanism; only "is the prefix a
|
|
#: word" separates them.
|
|
#:
|
|
#: LIKE `--tie-shared-rank`, THIS MOVES A PAYLOAD WITH NO BUNDLE CHANGING. A
|
|
#: consumer pinned to the previous excerpt order needs `--no-stem-prefix`.
|
|
DEFAULT_STEM_PREFIX = True
|
|
|
|
|
|
#: Round 16. Whether a concept whose WHOLE title the question accounts for is
|
|
#: read before the concepts the fusion ranked above it. ON since 2026-09-10.
|
|
#:
|
|
#: THE DEFECT IT REPAIRS. Both lexical signals are unnormalised COVERAGE
|
|
#: COUNTS -- one per question token the candidate answers to -- so nothing in
|
|
#: the fusion measures how much of the CANDIDATE the question accounts for. A
|
|
#: section titled with the question's subject and nothing else scores exactly
|
|
#: what a narrower section titled with that subject PLUS a qualifier scores,
|
|
#: and then loses to it on the body count. Measured 2026-09-10 on a
|
|
#: 2 761-concept bundle of one standard, where the answering section carries
|
|
#: the bare term as its title on three of six scored questions: `Hovedprosesser`
|
|
#: behind `Hovedprosess 81 ...`, `Armering` behind `Armering av ...`,
|
|
#: `Inspeksjon` behind `Enkel inspeksjon`.
|
|
#:
|
|
#: WHY A PARTITION AND NOT A FOURTH SIGNAL, measured rather than argued. RRF
|
|
#: consumes RANKS ONLY, and with shared ranks a signal whose positive group is
|
|
#: SMALL separates least of all: the group takes position 1 and everyone else
|
|
#: position `len(group) + 1`, so a rule firing on 1 concept of 2 761 is worth
|
|
#: `1/61 - 1/62` to it -- an order of magnitude under the body-signal gap it
|
|
#: has to close. Measured as a signal on that bundle it moved hit@1 not at all
|
|
#: (3 of 6); as a partition it reaches 6 of 6 candidate rank 1 with the
|
|
#: known-positive still at rank 1. `lookup_hits` is the same shape and was made
|
|
#: a partition on the same arithmetic.
|
|
#:
|
|
#: LIKE `--tie-shared-rank` AND `--stem-prefix`, THIS MOVES A PAYLOAD WITH NO
|
|
#: BUNDLE CHANGING. A consumer pinned to the previous excerpt order needs
|
|
#: `--no-title-covered`.
|
|
#:
|
|
#: ROUND 17 BOUNDS THE PARTITION BY RECALL, under the same flag. The partition
|
|
#: states the covered title's PRECISION -- it says nothing the question did not
|
|
#: ask -- and round 16 let that claim override the fusion even against a title
|
|
#: answering MORE of the question. Measured 2026-09-11 on a 26-concept bundle
|
|
#: of five tender documents: a question naming a section by three of its title
|
|
#: tokens also held a neighbour's whole one-token title, and the partition
|
|
#: moved that neighbour from fusion rank 2 to rank 1 over the section the
|
|
#: question names. A covered concept now RISES only past concepts whose titles
|
|
#: answer no more question tokens, by equality, than it holds. On the
|
|
#: 2 761-concept bundle no covered concept had such a title above it, so all
|
|
#: eight payloads there are byte-identical to round 16's. Four other repairs
|
|
#: were measured against it: a minimum title length (hit@1 there back to 3 of
|
|
#: 6), a share of the question (holds only in a band set by the question's
|
|
#: word count, 1/9 < s <= 1/6), an order inside the group (the known-negative's
|
|
#: group is ONE) and a stop list (no title involved is a function word). See
|
|
#: `docs/2026-09-11-k3-runde17-dekningen-stopper-ved-en-bredere-tittel.md`.
|
|
DEFAULT_TITLE_COVERED = True
|
|
|
|
|
|
def title_covered_hits(concepts: Sequence["Concept"], question: str) -> tuple[str, ...]:
|
|
"""The concepts whose ENTIRE title the question accounts for.
|
|
|
|
Every token of the title is a token of the question, so the title says
|
|
nothing the question did not ask about. That is a statement about the
|
|
CANDIDATE -- the complement of the coverage counts, which are statements
|
|
about the question -- and it is the one thing separating a section titled
|
|
with the subject from a narrower section titled with the subject plus a
|
|
qualifier.
|
|
|
|
**BY EQUALITY, never by shared prefix, and that is measured rather than
|
|
assumed.** `tokens_match` accepts four shared leading characters, which
|
|
would admit `Anchorage` beside `Anchoring` and, on the 2 761-concept
|
|
bundle, took the group from 1 concept to 6 on one question and from 9 to 31
|
|
on another -- the answering section falling to candidate rank 6 and the
|
|
known-positive to rank 2. Under equality both hold rank 1. The precedent is
|
|
`tokens_match`'s own: an identifier matches by equality alone, for the same
|
|
reason -- a prefix rule built for compounds says nothing true about a name.
|
|
|
|
Returns the EMPTY tuple when no concept qualifies, which is what makes the
|
|
rule invisible to every question that names no section outright. Byte
|
|
sorted, so several holders of one title arrive in a declared order.
|
|
|
|
Reads `title` alone and not the concept id: an id segment is this
|
|
library's reduction of a title, so counting it would let the same words
|
|
qualify a concept twice. The directory naming a single document is not
|
|
always a uuid no question carries -- an STS document names it with its own
|
|
number (K3-19) -- which is why the fusion's first signal reads the id
|
|
through `shared_id_prefix`.
|
|
"""
|
|
question_tokens = set(normalise(question))
|
|
if not question_tokens:
|
|
return ()
|
|
return tuple(
|
|
sorted(
|
|
concept.concept_id
|
|
for concept in concepts
|
|
if (title_tokens := normalise(concept.title))
|
|
and all(token in question_tokens for token in title_tokens)
|
|
)
|
|
)
|
|
|
|
|
|
def shared_id_prefix(concept_ids: Sequence[str]) -> int:
|
|
"""How many leading DIRECTORY segments every concept id in the bundle shares.
|
|
|
|
The first fusion signal reads a concept's title together with the segments
|
|
of its id, and on a bundle holding ONE document every id starts with the
|
|
same directory. A segment every concept carries separates nothing -- the
|
|
claim `concept_scores` already makes for a whole signal that scores every
|
|
concept alike -- but here it is not even a constant: the overlap counts
|
|
each question token once, so when the question names that directory, a
|
|
concept whose TITLE already names it gains nothing from the id while every
|
|
sibling gains the token. The one concept distinguished by naming the
|
|
document loses exactly that distinction, and a concept answering nothing
|
|
but the directory stops being a guess.
|
|
|
|
MEASURED 2026-09-11 on a 2 761-concept, one-document bundle whose directory
|
|
is the document's own number (K3-19): the known-positive question names the
|
|
document and fell from rank 1 to not delivered at the default `k` (13 at
|
|
`k` = 50). With the shared prefix unread it is rank 1 at both, and the six
|
|
scored questions keep rank 1.
|
|
|
|
**Directories only, never the leaf, and only the prefix EVERY id shares.**
|
|
Where the ids share no leading directory -- the measured multi-document
|
|
bundles, whose documents sit in different top-level directories -- the
|
|
signal reads what it read before, byte for byte. That is a measurement and
|
|
not a caution: dropping each concept's own document directory instead,
|
|
shared by every concept IN that document, was measured on the pinned
|
|
43-document bundle and took a hit@8 row from rank 5 to not delivered.
|
|
Across documents that directory carries information.
|
|
|
|
`searchable_text` still reads the whole id, so the stem vocabulary does not
|
|
move, and under `--rarity-weight` a token every concept carries weighs
|
|
`log(1) == 0` there -- the same answer this rule gives.
|
|
"""
|
|
directories = [concept_id.split("/")[:-1] for concept_id in concept_ids]
|
|
shared = 0
|
|
for segments in zip(*directories):
|
|
if any(segment != segments[0] for segment in segments):
|
|
break
|
|
shared += 1
|
|
return shared
|
|
|
|
|
|
def concept_scores(
|
|
concepts: Sequence[Concept],
|
|
question: str,
|
|
document_score: Mapping[str, float],
|
|
*,
|
|
cost_vocabulary: bool = False,
|
|
weights: Mapping[str, float] | None = None,
|
|
lookup: bool = True,
|
|
tie_shared_rank: bool = DEFAULT_TIE_SHARED_RANK,
|
|
title_covered: bool = DEFAULT_TITLE_COVERED,
|
|
stems: frozenset[str] | None = None,
|
|
link_in_signal: bool = DEFAULT_LINK_IN_SIGNAL,
|
|
) -> list[tuple[Concept, float, int]]:
|
|
"""Every concept, ordered best first, fused from three signals by RRF.
|
|
|
|
The signals: (1) the question against the concept's title and the segments
|
|
of its id below the directories every id shares (`shared_id_prefix`), (2)
|
|
the question against the body, (3) the stage-one score of the document the
|
|
concept belongs to.
|
|
|
|
**Ties break lexicographically by `concept_id`, at both the per-signal sort
|
|
and the fused sort.** Declared rather than inherited from dict insertion
|
|
order, so reproducibility is a property of the input SET and not of the
|
|
order it happened to arrive in.
|
|
|
|
**No float reaches the payload.** These scores order the cut; only ranks and
|
|
whole byte counts are emitted.
|
|
|
|
`lookup=False` isolates the FUSION from the lookup partition below it, and
|
|
exists because two stages sharing one output cannot otherwise be measured
|
|
apart: the tests that state what the rarity weight does to the fusion, and
|
|
the harness that measures a lookup's effect, both need the fusion's own
|
|
order. No caller in the run path sets it and the CLI does not expose it.
|
|
|
|
The third element of each tuple is the concept's OWN lexical overlap --
|
|
signals 1 and 2 only, with the document prior excluded. The cut needs it
|
|
separately: a concept that answers nothing in the question, sitting in a
|
|
document that does, is a GUESS, and a guess is the one thing a declared cut
|
|
must not deliver.
|
|
|
|
**`tie_shared_rank` is the one rule this fusion has for a signal that does
|
|
not separate**, off by default, and it is a correction to what the declared
|
|
tie-break does rather than a weight. RRF consumes ranks, so a rank is
|
|
produced for EVERY concept in EVERY signal -- including a signal that gave
|
|
them all the same score. The tie-break then orders that group by
|
|
`concept_id`, and the fusion reads the result as if it were a measurement.
|
|
|
|
MEASURED 2026-09-08 on the N500 bundle (270 concepts): the document prior
|
|
has **two** distinct values there and 269 concepts share one, so that
|
|
signal contributed the concepts' UUIDs in alphabetical order, spread across
|
|
`1/61` to `1/329`. The best concept covering `vann- og frostsikring` in a
|
|
subsea tunnel answered **7 of 7** question tokens and led the body signal
|
|
at rank 6; it fused to rank 14, outside the cut, behind concepts sharing
|
|
only `tunnel` and `vann` whose ids sorted earlier. With shared ranks it
|
|
fuses to rank 3.
|
|
|
|
**`link_in_signal=False` is the DEFAULT reading since K3-25**
|
|
(`DEFAULT_LINK_IN_SIGNAL`), and it was K3-23's instrument before that. The
|
|
body this reads is the file's, and `inbox._link_enclosing` puts one line in
|
|
it -- `Enclosing section: [<title>](/<bundle-absolute path>)` -- whose path
|
|
repeats the document's own directory in every linked body. SS 6.1's argument
|
|
for the absolute form is about the FILE; nothing ever decided it should be
|
|
SCORED, and K3-23 measured what happens when it is: **39 of 39** newcomers
|
|
the line added a question token to gained it through the PATH and **0 of
|
|
39** through the link's title. The excerpt a reader gets is untouched either
|
|
way, so what moved here is ORDER and never bytes. `True` is still reachable
|
|
and still does the older thing; the CLI exposes neither value, because no
|
|
caller has been named that needs the older one.
|
|
|
|
The rule takes the FIRST position of a score group rather than its middle.
|
|
Both were measured on the same four cases; the middle put the same concept
|
|
at rank 5 where the first puts it at 3, and neither changed the three
|
|
known-positive lookups. First is kept because it is the reading under which
|
|
a signal that separates nothing contributes an identical constant to every
|
|
concept -- which is the whole claim -- where the middle still varies with
|
|
the size of the group a concept lands in.
|
|
"""
|
|
question_tokens = normalise(question)
|
|
bridge = cost_vocabulary and question_uses_cost_vocabulary(question)
|
|
# The id BELOW the directories every concept shares: a segment every
|
|
# concept carries separates nothing. See `shared_id_prefix`.
|
|
shared = shared_id_prefix([concept.concept_id for concept in concepts])
|
|
bodies = {
|
|
concept.concept_id: (
|
|
concept.body if link_in_signal else body_without_link_line(concept.body)
|
|
)
|
|
for concept in concepts
|
|
}
|
|
titles = {
|
|
concept.concept_id: f"{concept.title} {' '.join(concept.concept_id.split('/')[shared:])}"
|
|
for concept in concepts
|
|
}
|
|
signals: list[dict[str, float]] = [
|
|
{
|
|
concept.concept_id: float(
|
|
_overlap(
|
|
question_tokens,
|
|
titles[concept.concept_id],
|
|
cost_vocabulary=bridge,
|
|
weights=weights,
|
|
stems=stems,
|
|
)
|
|
)
|
|
for concept in concepts
|
|
},
|
|
{
|
|
concept.concept_id: float(
|
|
_overlap(
|
|
question_tokens,
|
|
bodies[concept.concept_id],
|
|
cost_vocabulary=bridge,
|
|
weights=weights,
|
|
stems=stems,
|
|
)
|
|
)
|
|
for concept in concepts
|
|
},
|
|
{
|
|
concept.concept_id: document_score.get(concept.concept_id.split("/", 1)[0], 0.0)
|
|
for concept in concepts
|
|
},
|
|
]
|
|
fused: dict[str, float] = {concept.concept_id: 0.0 for concept in concepts}
|
|
for signal in signals:
|
|
# Sort by score descending, then by id ascending -- the declared
|
|
# tie-break, applied before a rank is ever read.
|
|
order = sorted(signal, key=lambda key: (-signal[key], key))
|
|
if not tie_shared_rank:
|
|
for position, concept_id in enumerate(order, start=1):
|
|
fused[concept_id] += 1.0 / (RRF_K + position)
|
|
continue
|
|
# SHARED RANK: every concept a signal scores EQUALLY takes that score
|
|
# group's first position, so the signal contributes the same amount to
|
|
# each of them and orders none of them. See this function's docstring
|
|
# and `tests/test_tie_shared_rank.py` for what it is for.
|
|
start = 0
|
|
while start < len(order):
|
|
stop = start
|
|
while stop < len(order) and signal[order[stop]] == signal[order[start]]:
|
|
stop += 1
|
|
contribution = 1.0 / (RRF_K + start + 1)
|
|
for concept_id in order[start:stop]:
|
|
fused[concept_id] += contribution
|
|
start = stop
|
|
lexical = (
|
|
{
|
|
concept.concept_id: int(signals[0][concept.concept_id] + signals[1][concept.concept_id])
|
|
for concept in concepts
|
|
}
|
|
if weights is None
|
|
# A COUNT even when the signals are weighted. The cut reads this, and a
|
|
# word every concept carries weighs zero: were `lexical` the weighted
|
|
# sum, a concept matching only that word would fall to
|
|
# `no_lexical_match` -- turning a ranking change into the GATE
|
|
# `54a0bc2` falsified. The gate is a different axis and stays where it
|
|
# was, at the price of one more pass over the same two fields.
|
|
else {
|
|
concept.concept_id: int(
|
|
_overlap(
|
|
question_tokens, titles[concept.concept_id], cost_vocabulary=bridge, stems=stems
|
|
)
|
|
+ _overlap(
|
|
question_tokens,
|
|
bodies[concept.concept_id],
|
|
cost_vocabulary=bridge,
|
|
stems=stems,
|
|
)
|
|
)
|
|
for concept in concepts
|
|
}
|
|
)
|
|
by_id = {concept.concept_id: concept for concept in concepts}
|
|
ranked_ids = sorted(fused, key=lambda key: (-fused[key], key))
|
|
covered = set(title_covered_hits(concepts, question)) if title_covered else set()
|
|
if covered:
|
|
# A PARTITION, not a signal, and it lands BELOW the lookup partition
|
|
# so a question that NAMES a concept still reads that one first. See
|
|
# `DEFAULT_TITLE_COVERED` for the arithmetic that rules a signal out,
|
|
# and `tests/test_title_covered.py` for the mechanism on a fixture.
|
|
#
|
|
# BOUNDED BY RECALL since round 17: a covered concept RISES through the
|
|
# fusion's order and stops beneath the first concept whose title
|
|
# answers MORE question tokens, by equality, than the covered title
|
|
# holds -- or beneath a covered concept the fusion put above it. With
|
|
# nothing above it answering more, it reaches the top exactly where
|
|
# round 16's plain partition put it. `tests/test_title_covered_rise.py`
|
|
# holds the known-negative that bound exists for.
|
|
#
|
|
# STABLE: the covered concepts keep the order the fusion gave them, and
|
|
# so does everything else, so nothing here depends on dict order.
|
|
asked = frozenset(question_tokens)
|
|
answered = {
|
|
concept_id: len(set(normalise(by_id[concept_id].title)) & asked)
|
|
for concept_id in ranked_ids
|
|
}
|
|
risen: list[str] = []
|
|
for concept_id in ranked_ids:
|
|
stop = len(risen)
|
|
if concept_id in covered:
|
|
while (
|
|
stop
|
|
and risen[stop - 1] not in covered
|
|
and answered[risen[stop - 1]] <= answered[concept_id]
|
|
):
|
|
stop -= 1
|
|
risen.insert(stop, concept_id)
|
|
ranked_ids = risen
|
|
named = set(lookup_hits(concepts, question)) if lookup else set()
|
|
if named:
|
|
# THE LOOKUP LANDS BEFORE THE FUSION'S OUTPUT IS READ, and it is a
|
|
# partition rather than a fourth signal. The form was chosen by
|
|
# measurement, not by preference: a fourth RRF signal was simulated on
|
|
# the same three bundles first and put the named concept at rank
|
|
# **26 / 15 / 19** of 446 / 1 133 / 270 -- none of them delivered. RRF
|
|
# consumes RANKS ONLY, so any single signal contributes at most
|
|
# `1/(RRF_K + 1)` however certain it is, and a concept the question
|
|
# NAMES cannot outbid three signals that merely describe it
|
|
# (`docs/2026-09-08-sjeldenhetsvekt.md` SS 4 predicted exactly this).
|
|
#
|
|
# STABLE: the named concepts keep the order the fusion gave them, and
|
|
# so does everything else, so nothing here depends on dict order.
|
|
ranked_ids = [key for key in ranked_ids if key in named] + [
|
|
key for key in ranked_ids if key not in named
|
|
]
|
|
return [
|
|
(by_id[concept_id], fused[concept_id], lexical[concept_id]) for concept_id in ranked_ids
|
|
]
|
|
|
|
|
|
# --- The cut (SS 5, SS 7, SS 9.1) ---------------------------------------------
|
|
|
|
#: Every rule this instrument may drop a concept under, spelled once. CLOSED:
|
|
#: a drop with no rule is the silent cut SS 5.3 exists to forbid, and a rule
|
|
#: invented at the drop site is a vocabulary no consumer can be held to.
|
|
WITHHOLDING_RULES = (
|
|
"verdict_layer_excluded",
|
|
"verified_unreadable",
|
|
"no_lexical_match",
|
|
"over_budget_alone",
|
|
"below_k",
|
|
"over_budget_after_knapsack",
|
|
"source_quota_exceeded",
|
|
)
|
|
|
|
#: `--source-quota N` caps how many DELIVERED places one source document may
|
|
#: take, topping the shortlist back up to `k` when the bundle has no
|
|
#: alternatives to offer.
|
|
#:
|
|
#: **2 SINCE 2026-09-10, and it is the third change here that moves a payload
|
|
#: with NO bundle changing** (after `--tie-shared-rank` and `--stem-prefix`);
|
|
#: a consumer pinned to the previous excerpt order needs `--no-source-quota`.
|
|
#: The defect it repairs was measured on a 3206-concept bundle of a published
|
|
#: handbook: the code's own process overview contributes 28 of 3206 concepts
|
|
#: (0.87 %) and 8.0 % of the source characters, and took 8 of 8 delivered
|
|
#: places on one question and 7 of 8 on the known-positive -- identical at 343
|
|
#: and 1651 concepts, so the cause is the corpus's COMPOSITION (it holds its
|
|
#: own table of contents) and not its size.
|
|
#:
|
|
#: SWEPT over {2, 3, 4, off} on three bundles. At 2 and 3 hit@8 goes 5 of 6 to
|
|
#: **6 of 6 on BOTH K2 bundles** with all five standing rank-1 rows unmoved; at
|
|
#: 4 and off it is 5 of 6. On the handbook bundle hit@8 goes 2 of 6 to 4 of 6
|
|
#: and the dominant document's share of delivered places 8 of 8 to 2 of 8. 2
|
|
#: rather than 3 on rank: the recovered rows come in at 5 and 4 rather than 7
|
|
#: and 5.
|
|
DEFAULT_SOURCE_QUOTA: int | None = 2
|
|
|
|
#: The knapsack's weight granularity, in bytes. Bucketing keeps the DP table
|
|
#: small; bucketing UP the item and DOWN the capacity keeps the error one-sided,
|
|
#: so the pack may under-deliver by a bucket and can never over-spend.
|
|
WEIGHT_BUCKET = 500
|
|
|
|
|
|
def delivered_text(body: str) -> str:
|
|
"""A concept body as a payload carries it: NFC, trailing whitespace
|
|
stripped per line. One rule for an excerpt's `text` and a `parent`'s."""
|
|
return "\n".join(line.rstrip() for line in unicodedata.normalize("NFC", body).split("\n"))
|
|
|
|
|
|
def excerpt_for(concept: Concept) -> dict[str, object] | None:
|
|
"""One concept as a payload excerpt, or `None` when it cannot be tiered.
|
|
|
|
Carries the SS 8 members plus three the contract permits and this profile
|
|
needs:
|
|
|
|
- **`text`** -- the concept body. SS 8 names no content member, but SS 1
|
|
defines an excerpt as "one delivered unit of bundle CONTENT", and without
|
|
a body the budget gate would measure a two-kilobyte skeleton while the
|
|
skill went to the bundle itself, breaking SS 2.2.
|
|
- **`text_sha256`** -- the digest of the delivered bytes. `sha256` is the
|
|
digest of the WHOLE concept file (SS 3.2), so without this second digest a
|
|
consumer holds an identity that cannot verify what it was handed.
|
|
- **`bundle_id_inherited`** -- whether the first half of the SS 3.1 identity
|
|
tuple came from the concept or from the root index.
|
|
|
|
And, since C1 step 0, the keys that let a reader NAME what it is citing.
|
|
`title` is unconditional (SS 8, this revision); `req_number`, `sources` and
|
|
each locator are written only when the concept carries them, because a key
|
|
with an empty value asserts that the producer wrote one. po measured
|
|
2026-09-08 (`e7ffe9e`) that the pre-pass delivered the gold concept at rank
|
|
1 on three bundles while the model could not name it: the ranking found the
|
|
document and the delivery dropped the key.
|
|
|
|
Since K3-21, `parent` too, on the same terms: the concept that encloses
|
|
this one, as its `concept_id` and `title`, written only when the concept
|
|
names one and the pointer resolves (`link_parents`). Never the raw
|
|
`segment_id` the producer wrote -- a reader holding `p1977` can open
|
|
nothing -- and `parent_unresolved` where the pointer lands nowhere, the
|
|
`sources_unreadable` rule for the same reason.
|
|
|
|
Trailing whitespace is stripped per line: a spreadsheet render is padded to
|
|
hundreds of trailing spaces per line, and unstripped, most of a budget goes
|
|
on padding.
|
|
"""
|
|
tier = trust_tier(concept.frontmatter.get("verified"))
|
|
if tier is None:
|
|
return None
|
|
text = delivered_text(concept.body)
|
|
excerpt: dict[str, object] = {
|
|
"bundle_id": concept.bundle_id,
|
|
"concept_id": concept.concept_id,
|
|
"sha256": concept.sha256,
|
|
"adjudication": concept.adjudication,
|
|
"trust_tier": tier,
|
|
"bundle_id_inherited": concept.bundle_id_inherited,
|
|
"title": concept.title,
|
|
}
|
|
if concept.req_number:
|
|
excerpt["req_number"] = concept.req_number
|
|
if concept.sources:
|
|
excerpt["sources"] = [dict(entry) for entry in concept.sources]
|
|
elif concept.sources_present:
|
|
# The third state, written rather than silently dropped: the concept
|
|
# HAS an address and this reader could not decode it.
|
|
excerpt["sources_unreadable"] = True
|
|
excerpt.update(concept.locators)
|
|
if concept.parent is not None:
|
|
excerpt["parent"] = dict(concept.parent)
|
|
elif concept.parent_unresolved:
|
|
excerpt["parent_unresolved"] = True
|
|
excerpt["text_sha256"] = hashlib.sha256(text.encode("utf-8")).hexdigest()
|
|
excerpt["text"] = text
|
|
return excerpt
|
|
|
|
|
|
def excerpt_weight(excerpt: Mapping[str, object]) -> int:
|
|
"""What this excerpt costs by the gate's own instrument.
|
|
|
|
The ENCODED JSON object, never `stat().st_size`: measured, the two differ by
|
|
7.1 % over the K2 corpus, and a cut computed as fitting would then be
|
|
refused by a gate measuring different bytes.
|
|
"""
|
|
return len(json.dumps(excerpt, ensure_ascii=False).encode("utf-8"))
|
|
|
|
|
|
#: `--follow-parent` (K3-21 B): carry the enclosing concept's TEXT inside an
|
|
#: excerpt's `parent`, rather than the pointer alone. OFF by measurement. On
|
|
#: the one standard with shells it delivers the same set as the pointer, by
|
|
#: construction and measured (16 of 16 payloads, both k), and on six questions
|
|
#: whose answers are never heading-only sections it bought nothing a reader
|
|
#: scored -- 6 of 6 correct either way -- at +2 488 and +2 746 B on the 2 of 8
|
|
#: payloads holding a linked shell at the default k and up to +20 817 B at
|
|
#: k 50. What it does buy shows on a question AT a shell: 3 847 characters of
|
|
#: the enclosing section in the payload, where the pointer alone needs one
|
|
#: `Read`.
|
|
DEFAULT_FOLLOW_PARENT = False
|
|
|
|
|
|
def _carrying(
|
|
excerpt: Mapping[str, object], source: Concept, text: str, *, truncated: bool
|
|
) -> dict[str, object]:
|
|
parent = excerpt["parent"]
|
|
assert isinstance(parent, dict)
|
|
carried: dict[str, object] = {**parent, "sha256": source.sha256, "text": text}
|
|
if truncated:
|
|
carried["truncated"] = True
|
|
return {**excerpt, "parent": carried}
|
|
|
|
|
|
def attach_parent_text(
|
|
delivered: Sequence[dict[str, object]],
|
|
concepts: Mapping[str, Concept],
|
|
*,
|
|
limit: int,
|
|
) -> tuple[dict[str, object], ...]:
|
|
"""The delivered excerpts, each `parent` carrying its concept's text where room allows.
|
|
|
|
Runs AFTER the cut, on the room the cut left, in rank order -- so the
|
|
delivered set, its order and the withheld list are what they were without
|
|
it, and inherited text cannot displace an excerpt. That is the mechanism a
|
|
consumer measured when it copied ancestors' text into every heading-only
|
|
section of one standard: the excerpts grew until the budget held 4-8, and
|
|
the right section fell to withheld place 504 and 1 069.
|
|
|
|
A parent the payload already holds, or one a higher-ranked excerpt already
|
|
carried, is not delivered again. A text the room cannot hold whole is cut
|
|
to the longest prefix that fits and marked `truncated`; with no room for
|
|
any of it, the pointer stays a pointer. `sha256` is the enclosing concept's
|
|
own, so a claim resting on that text is cited as that concept (SS 3.2).
|
|
"""
|
|
room = limit - sum(excerpt_weight(excerpt) for excerpt in delivered)
|
|
given = {str(excerpt["concept_id"]) for excerpt in delivered}
|
|
out: list[dict[str, object]] = []
|
|
for excerpt in delivered:
|
|
parent = excerpt.get("parent")
|
|
source = concepts.get(str(parent.get("concept_id"))) if isinstance(parent, dict) else None
|
|
if source is None or source.concept_id in given:
|
|
out.append(excerpt)
|
|
continue
|
|
text = delivered_text(source.body)
|
|
base = excerpt_weight(excerpt)
|
|
chosen = _carrying(excerpt, source, text, truncated=False)
|
|
if excerpt_weight(chosen) - base > room:
|
|
# The longest prefix that fits. The encoded weight never falls as
|
|
# a prefix grows, so the search is exact.
|
|
low, high = 0, len(text)
|
|
while low < high:
|
|
middle = (low + high + 1) // 2
|
|
trial = _carrying(excerpt, source, text[:middle], truncated=True)
|
|
if excerpt_weight(trial) - base <= room:
|
|
low = middle
|
|
else:
|
|
high = middle - 1
|
|
fits = _carrying(excerpt, source, text[:low], truncated=True)
|
|
chosen = fits if low and excerpt_weight(fits) - base <= room else dict(excerpt)
|
|
if chosen is not excerpt and "text" in _mapping_of(chosen.get("parent")):
|
|
given.add(source.concept_id)
|
|
room -= excerpt_weight(chosen) - base
|
|
out.append(chosen)
|
|
return tuple(out)
|
|
|
|
|
|
def _mapping_of(value: object) -> Mapping[str, object]:
|
|
return value if isinstance(value, Mapping) else {}
|
|
|
|
|
|
def knapsack(items: Sequence[tuple[float, int]], *, capacity: int) -> tuple[int, ...]:
|
|
"""The exact 0/1 knapsack: indices of the highest-value subset that fits.
|
|
|
|
Exact rather than greedy-by-density, which has an unbounded approximation
|
|
factor -- and over at most `k` items the DP is microseconds, so the
|
|
approximation buys nothing. Deterministic: a subset replaces the incumbent
|
|
only on a STRICT improvement, so equal-value subsets resolve to the one
|
|
built from earlier items, and the caller sorts the pool by `concept_id`
|
|
first.
|
|
"""
|
|
best: list[tuple[float, tuple[int, ...]]] = [(0.0, ()) for _ in range(capacity + 1)]
|
|
for index, (value, weight) in enumerate(items):
|
|
if weight > capacity:
|
|
continue
|
|
for room in range(capacity, weight - 1, -1):
|
|
candidate_value = best[room - weight][0] + value
|
|
if candidate_value > best[room][0]:
|
|
best[room] = (candidate_value, (*best[room - weight][1], index))
|
|
return max(best, key=lambda entry: entry[0])[1]
|
|
|
|
|
|
def _pack(
|
|
shortlist: Sequence[tuple["Concept", float, dict[str, object], int]],
|
|
*,
|
|
limit: int,
|
|
reserve_top_rank: bool,
|
|
) -> tuple[list[dict[str, object]], list[str], tuple[str, int] | None]:
|
|
"""The budget step, alone: a shortlist in fused-rank order to delivered
|
|
excerpts, the ids the budget dropped, and the reservation that was made.
|
|
|
|
Lifted out of :func:`cut` so the quota-off fate of a dropped candidate is
|
|
decided by the SAME packer the run itself uses.
|
|
"""
|
|
# The DP POOL is sorted by `concept_id`, so which of two equal-value subsets
|
|
# wins is a property of the input set rather than of the order the ranker
|
|
# happened to emit. The OUTPUT is not: excerpts come back in fused-rank
|
|
# order, because the rank is what a hit@k measurement reads, and an id-sorted
|
|
# list would silently turn "position in the payload" into a different number
|
|
# from "position in the ranking".
|
|
pool = sorted(shortlist, key=lambda entry: entry[0].concept_id)
|
|
reserved: tuple[str, int] | None = None
|
|
room = limit
|
|
if reserve_top_rank and shortlist:
|
|
# `shortlist` is in fused-rank order, so its first entry IS the
|
|
# top-ranked candidate -- not the heaviest, and not the first by id.
|
|
top = shortlist[0]
|
|
reserved = (top[0].concept_id, top[3])
|
|
room = limit - top[3]
|
|
pool = [entry for entry in pool if entry[0] is not top[0]]
|
|
capacity = room // WEIGHT_BUCKET
|
|
packed = {
|
|
id(pool[index][0])
|
|
for index in knapsack(
|
|
tuple((score, -(-weight // WEIGHT_BUCKET)) for _, score, _, weight in pool),
|
|
capacity=capacity,
|
|
)
|
|
}
|
|
if reserved is not None:
|
|
packed.add(id(shortlist[0][0]))
|
|
delivered: list[dict[str, object]] = []
|
|
over_budget: list[str] = []
|
|
for concept, _, excerpt, _ in shortlist:
|
|
if id(concept) in packed:
|
|
delivered.append({**excerpt, "rank": len(delivered) + 1})
|
|
else:
|
|
over_budget.append(concept.concept_id)
|
|
return delivered, over_budget, reserved
|
|
|
|
|
|
def _fates_without_quota(
|
|
candidates: Sequence[tuple["Concept", float, dict[str, object], int]],
|
|
*,
|
|
k: int,
|
|
limit: int,
|
|
reserve_top_rank: bool,
|
|
) -> dict[str, str]:
|
|
"""What would have become of each candidate had the quota not run.
|
|
|
|
A candidate the quota-off cut would have DELIVERED is the only one the
|
|
quota can honestly claim: it is the place the quota took. Everything else
|
|
keeps the rule the quota-off cut gives it, because that is what actually
|
|
decided it.
|
|
"""
|
|
fates = {concept.concept_id: "below_k" for concept, _, _, _ in candidates[k:]}
|
|
delivered, over_budget, _ = _pack(
|
|
candidates[:k], limit=limit, reserve_top_rank=reserve_top_rank
|
|
)
|
|
fates.update({concept_id: "over_budget_after_knapsack" for concept_id in over_budget})
|
|
fates.update({str(entry["concept_id"]): "source_quota_exceeded" for entry in delivered})
|
|
return fates
|
|
|
|
|
|
def cut(
|
|
ranked: Sequence[tuple[Concept, float, int]],
|
|
*,
|
|
k: int,
|
|
limit: int,
|
|
reserve_top_rank: bool = False,
|
|
source_quota: int | None = DEFAULT_SOURCE_QUOTA,
|
|
) -> tuple[tuple[dict[str, object], ...], tuple[tuple[str, str], ...], tuple[str, int] | None]:
|
|
"""The ranked concepts split into delivered excerpts, named drops, and the
|
|
reservation that was made, if any.
|
|
|
|
**The partition is the invariant, not a consequence.** Every considered
|
|
concept lands in exactly one of the two, so `considered == withheld +
|
|
delivered` closes by construction rather than by a count computed twice.
|
|
|
|
Exclusions run before the pack, each naming its rule, because "it did not
|
|
fit" and "it could never be delivered" are different facts about the cut.
|
|
|
|
**A concept answering nothing in the question is withheld, never ranked into
|
|
the top k as filler.** Without that rule a question with no answer in the
|
|
bundle still returns eight excerpts -- a confident guess wearing a
|
|
denominator -- and hit@k over such a ranker measures the corpus's size
|
|
rather than the ranker.
|
|
|
|
**`reserve_top_rank` (default off) buys the highest-ranked candidate its
|
|
bytes before the pack runs.** The DP maximises a SUM of fused scores, so a
|
|
single candidate costing a large share of the budget loses to enough small
|
|
ones no matter how far ahead it ranks -- measured, a top-ranked excerpt
|
|
worth 56.5 % of the budget is evicted as soon as the shortlist holds enough
|
|
alternatives, which makes `k` a dial that can remove the one concept a
|
|
question was asked about. The reservation makes rank one a floor rather
|
|
than a bid, and the pack fills what is left. It runs AFTER the
|
|
`over_budget_alone` pre-exclusion, never before: a reservation for an
|
|
excerpt the budget can never hold would deliver bytes the gate refuses.
|
|
"""
|
|
withheld: list[tuple[str, str]] = []
|
|
candidates: list[tuple[Concept, float, dict[str, object], int]] = []
|
|
for concept, score, lexical in ranked:
|
|
# SS 9.1: a TYPE check at every level, never a path filter. Case-folded
|
|
# against the profile's own constant, because `type: Log` (capital L)
|
|
# really occurs in the corpus.
|
|
if concept.okf_type.casefold() == RESERVED_OKF_TYPE:
|
|
withheld.append((concept.concept_id, "verdict_layer_excluded"))
|
|
continue
|
|
if lexical == 0:
|
|
withheld.append((concept.concept_id, "no_lexical_match"))
|
|
continue
|
|
excerpt = excerpt_for(concept)
|
|
if excerpt is None:
|
|
withheld.append((concept.concept_id, "verified_unreadable"))
|
|
continue
|
|
weight = excerpt_weight(excerpt)
|
|
if weight > limit:
|
|
withheld.append((concept.concept_id, "over_budget_alone"))
|
|
continue
|
|
candidates.append((concept, score, excerpt, weight))
|
|
if source_quota is not None:
|
|
# THE QUOTA CUTS WHERE THE SHORTLIST IS CUT, never inside the pack. The
|
|
# DP maximises a sum over a set it is handed; a quota expressed there
|
|
# would be a constraint on the sum, which is a different problem and a
|
|
# slower one. Here it is a filter on the ranked candidate list, so `k`
|
|
# is still delivered in full and the freed place goes to the next
|
|
# candidate rather than being lost.
|
|
seen: dict[str, int] = {}
|
|
keep = [True] * len(candidates)
|
|
over: list[int] = []
|
|
for index, entry in enumerate(candidates):
|
|
document = entry[0].source_file
|
|
taken = seen.get(document, 0)
|
|
if taken >= source_quota:
|
|
keep[index] = False
|
|
over.append(index)
|
|
continue
|
|
seen[document] = taken + 1
|
|
# THE QUOTA NEVER SHORTENS THE PAYLOAD, and that is not a nicety. A
|
|
# bundle built from ONE document carries the same `source_file` on
|
|
# every concept, so a quota applied without this would deliver
|
|
# `source_quota` excerpts instead of `k` -- a rule against dominance
|
|
# turned into a rule against small bundles. `over` is in fused-rank
|
|
# order, so the top-up takes the BEST-ranked over-quota candidates
|
|
# back, and a bundle with no alternatives to offer is byte-identical
|
|
# to the quota being off.
|
|
for index in over[: max(k - sum(keep), 0)]:
|
|
keep[index] = True
|
|
# THE RULE A DROP CARRIES IS TRUE OF THAT DROP. The quota filters the
|
|
# WHOLE candidate list rather than the top k, so an over-quota
|
|
# candidate the RANK had already put outside k used to come back
|
|
# `source_quota_exceeded` -- the quota only reached it because it ran
|
|
# first. Measured on 25 real misses 2026-09-17: 13 were named by the
|
|
# quota and decided by the rank.
|
|
#
|
|
# The truth is the SAME cut without the quota, computed from the SAME
|
|
# candidate list by the SAME packer -- never a second implementation
|
|
# of the rule, which would be two rules with one name.
|
|
fates = _fates_without_quota(
|
|
candidates, k=k, limit=limit, reserve_top_rank=reserve_top_rank
|
|
)
|
|
withheld.extend(
|
|
(candidates[index][0].concept_id, fates[candidates[index][0].concept_id])
|
|
for index in range(len(candidates))
|
|
if not keep[index]
|
|
)
|
|
candidates = [entry for index, entry in enumerate(candidates) if keep[index]]
|
|
for concept, _, _, _ in candidates[k:]:
|
|
withheld.append((concept.concept_id, "below_k"))
|
|
shortlist = candidates[:k]
|
|
delivered, over_budget, reserved = _pack(
|
|
shortlist, limit=limit, reserve_top_rank=reserve_top_rank
|
|
)
|
|
withheld.extend((concept_id, "over_budget_after_knapsack") for concept_id in over_budget)
|
|
withheld.sort()
|
|
return tuple(delivered), tuple(withheld), reserved
|
|
|
|
|
|
# --- The payload (SS 8) -------------------------------------------------------
|
|
|
|
#: SS 8.2: present so a reader can tell which revision it is holding.
|
|
CONTRACT_REVISION = "okf-consumption/1"
|
|
|
|
#: `--k` caps the DELIVERED set. The budget is the gate; this is a second,
|
|
#: cheaper bound so a question matching half the corpus does not run a
|
|
#: 300-item DP to discover the same answer.
|
|
DEFAULT_K = 8
|
|
|
|
|
|
def root_bundle_id_of(bundle_root: Path, *, profile: BundleProfile = DEFAULT_PROFILE) -> str:
|
|
"""The `bundle_id` the root index declares, or a refusal naming which half
|
|
of SS 3.1's identity tuple is missing. Shared with the skill generator, so
|
|
the two agree on what makes a directory a readable bundle."""
|
|
root_index = bundle_root / profile.index.name
|
|
if not root_index.is_file():
|
|
raise ConsumeError(
|
|
f"{bundle_root} carries no {profile.index.name}, so there is no index "
|
|
"tree to walk and no way to read the bundle without enumerating a "
|
|
"directory, which SS 9.2 forbids",
|
|
code="bundle_unreadable",
|
|
)
|
|
declared = parse_frontmatter(root_index).get("bundle_id", "")
|
|
if not declared:
|
|
raise ConsumeError(
|
|
f"{root_index} declares no `bundle_id`; identity across bundles is "
|
|
"the (bundle_id, concept_id) tuple (SS 3.1) and half of it is missing",
|
|
code="bundle_id_missing",
|
|
)
|
|
return declared
|
|
|
|
|
|
def excerpt_text(excerpt: Mapping[str, object]) -> str:
|
|
"""The strings a delivered excerpt actually shows its reader.
|
|
|
|
Title, id and body, which is the same union the two lexical signals score
|
|
-- a coverage report counting a term the consumer cannot see would be a
|
|
report about the ranker rather than about the payload.
|
|
"""
|
|
return " ".join(str(excerpt.get(key, "")) for key in ("title", "concept_id", "text"))
|
|
|
|
|
|
def unanswered_terms(
|
|
terms: Sequence[str],
|
|
tokenised: Sequence[Sequence[str]],
|
|
*,
|
|
stems: frozenset[str] | None,
|
|
) -> list[str]:
|
|
"""The question terms none of `tokenised` answers, in the question's order.
|
|
|
|
FACTS, NEVER A VERDICT, and that is a measurement rather than caution. Two
|
|
readings were built and both were falsified against 81 questions (16
|
|
synthetic, 65 over the three real sets, 2026-09-20):
|
|
|
|
- the SHARE of question terms a delivered excerpt answers separates the
|
|
synthetic controls at 0.33 against 0.50 and then reverses on real data,
|
|
where covered questions run down to 0.27 and one real uncovered question
|
|
sits at 0.71. Question STYLE dominates it: a question carrying its own
|
|
instructions has a long tail of words no corpus answers.
|
|
- the share of the bundle that ties the best lexical match separates the
|
|
synthetic corpus and is ~0.00 for every real question, covered or not,
|
|
because a 4 369-concept bundle has no ties to speak of.
|
|
|
|
So this module states what it measured and stops there. A bar over these
|
|
lists is the READER's -- a model weighing "the bundle answers none of
|
|
`koster`, `doegn`, `gjester`" is doing something no constant here can do
|
|
across corpora, and `tools/okf_retrieval_gate.py`'s row 4 carries the one
|
|
bar this repository does declare, with its own measurement beside it.
|
|
|
|
The matcher is the ranker's own (`tokens_match`, same `stems`). The cost
|
|
vocabulary's bridge is NOT applied: it is off by default and it is a rule
|
|
about documents naming money in other words, which would make a term read
|
|
as answered by a text that never carries it.
|
|
"""
|
|
remaining = set(terms)
|
|
for tokens in tokenised:
|
|
if not remaining:
|
|
break
|
|
for term in tuple(remaining):
|
|
if any(tokens_match(term, other, stems=stems) for other in tokens):
|
|
remaining.discard(term)
|
|
return [term for term in terms if term in remaining]
|
|
|
|
|
|
def build_payload(
|
|
bundle_root: Path,
|
|
*,
|
|
question: str,
|
|
k: int = DEFAULT_K,
|
|
limit: int = DEFAULT_LIMIT,
|
|
profile: BundleProfile = DEFAULT_PROFILE,
|
|
cost_vocabulary: bool = False,
|
|
reserve_top_rank: bool = False,
|
|
rarity_weight: bool = False,
|
|
tie_shared_rank: bool = DEFAULT_TIE_SHARED_RANK,
|
|
title_covered: bool = DEFAULT_TITLE_COVERED,
|
|
withheld_titles: bool = False,
|
|
stem_prefix: bool = DEFAULT_STEM_PREFIX,
|
|
source_quota: int | None = DEFAULT_SOURCE_QUOTA,
|
|
follow_parent: bool = DEFAULT_FOLLOW_PARENT,
|
|
link_in_signal: bool = DEFAULT_LINK_IN_SIGNAL,
|
|
) -> dict[str, object]:
|
|
"""One bundle plus one question, cut to one contract-conformant payload.
|
|
|
|
Pure with respect to the clock and the network: the same
|
|
`(bundle_root, question, k, limit, cost_vocabulary, reserve_top_rank,
|
|
rarity_weight, tie_shared_rank, title_covered, withheld_titles)` at the
|
|
same bytes returns
|
|
the same object, every time.
|
|
|
|
**`withheld_titles` (default off) names what was dropped.** A `withheld`
|
|
entry carries `concept_id` and `rule` and no title, so a reader told that
|
|
262 concepts were withheld cannot tell WHAT was withheld without reading
|
|
the bundle -- which SS 2.2 forbids. The title closes that.
|
|
|
|
OFF BY DEFAULT BY MEASUREMENT, not by taste. Measured 2026-09-08: on N500
|
|
the payload grows 41 364 -> 57 023 bytes (+37.9 %), and on the 629-concept
|
|
K2 bundle the bookkeeping -- everything that is not an excerpt -- grows to
|
|
**122 704 bytes, past the 120 000-byte limit itself**. The instantiated
|
|
skill publishes that breaking point as "~75 KB at 629 concepts, reached at
|
|
roughly 8 000 concepts"; on by default would make that sentence false and
|
|
would move every consumer's bytes for a field none of them asked for.
|
|
Whether the naming is worth the bookkeeping is the caller's call, and the
|
|
flag is how it stays one.
|
|
|
|
**`link_in_signal` defaults to `False` since K3-25 and carries no CLI
|
|
flag** -- see :func:`concept_scores`. It keeps the door's `Enclosing
|
|
section:` line out of the ranking and out of the stem vocabulary; the
|
|
excerpt still carries it, so a payload built either way holds the same bytes
|
|
per excerpt and only the ORDER can differ. That is what separates a ranking
|
|
movement from a budget displacement, and it is why the default moved without
|
|
a bundle changing: **0 of 5** bundles anyone ships carries the line, so **5
|
|
of 5** payloads are byte-identical across the move.
|
|
"""
|
|
case, expected, measured = known_positive()
|
|
if expected != measured:
|
|
# SS 7.4: the instrument reports NONE of its own numbers until it has
|
|
# reproduced a known figure. Refusing here rather than emitting a
|
|
# payload with a failing known-positive is the difference between an
|
|
# instrument that has been shown to count and one that merely says so.
|
|
raise ConsumeError(
|
|
f"the budget instrument's known-positive ({case}) expected {expected} "
|
|
f"and measured {measured}; refusing to report any figure until the "
|
|
"two agree",
|
|
code="instrument_unvalidated",
|
|
)
|
|
root_bundle_id = root_bundle_id_of(bundle_root, profile=profile)
|
|
concept_ids = enumerate_concepts(bundle_root, profile=profile)
|
|
concepts = link_parents(
|
|
[
|
|
read_concept(
|
|
bundle_root / f"{concept_id}{profile.paths.concept_suffix}",
|
|
bundle_root=bundle_root,
|
|
root_bundle_id=root_bundle_id,
|
|
)
|
|
for concept_id in concept_ids
|
|
]
|
|
)
|
|
# The text the two lexical signals read, tokenised ONCE: the stem
|
|
# vocabulary, the rarity `df` and the coverage report below all count over
|
|
# the same strings, so none of them can weigh a token by how rare it is
|
|
# somewhere it is not read.
|
|
texts = searchable_text(concepts, link_in_signal=link_in_signal)
|
|
tokenised = [normalise(text) for text in texts]
|
|
# The bundle's OWN vocabulary, and the reason the rule is a set rather than
|
|
# a threshold: `pris` is a word here and `bila` is not, which is what
|
|
# separates a Norwegian compound from four coincidental characters.
|
|
stems = frozenset(token for tokens in tokenised for token in tokens) if stem_prefix else None
|
|
weights = rarity_weights(normalise(question), texts, stems=stems) if rarity_weight else None
|
|
ranked = concept_scores(
|
|
concepts,
|
|
question,
|
|
document_scores(
|
|
bundle_root,
|
|
question,
|
|
profile=profile,
|
|
cost_vocabulary=cost_vocabulary,
|
|
weights=weights,
|
|
stems=stems,
|
|
),
|
|
cost_vocabulary=cost_vocabulary,
|
|
weights=weights,
|
|
tie_shared_rank=tie_shared_rank,
|
|
title_covered=title_covered,
|
|
stems=stems,
|
|
link_in_signal=link_in_signal,
|
|
)
|
|
titles_by_id = {concept.concept_id: concept.title for concept in concepts}
|
|
matched = sum(1 for _, _, lexical in ranked if lexical > 0)
|
|
delivered, withheld, reserved = cut(
|
|
ranked,
|
|
k=k,
|
|
limit=limit,
|
|
reserve_top_rank=reserve_top_rank,
|
|
source_quota=source_quota,
|
|
)
|
|
if follow_parent:
|
|
# After the cut and never inside it: see `attach_parent_text`.
|
|
delivered = attach_parent_text(
|
|
delivered, {concept.concept_id: concept for concept in concepts}, limit=limit
|
|
)
|
|
spent = sum(excerpt_weight(excerpt) for excerpt in delivered)
|
|
if matched and not delivered:
|
|
# SS 7.3: a finding requiring a decision, never something to retry
|
|
# narrower. Distinguished from the honest empty result BY THE
|
|
# DENOMINATOR: concepts answered this question and the budget admitted
|
|
# none of them, which is a statement about the limit, not about the
|
|
# bundle.
|
|
raise ConsumeError(
|
|
f"{matched} concept(s) answered this question and the {limit}-byte "
|
|
f"budget admitted none of them; the cut strategy is wrong for this "
|
|
"bundle at this limit -- refusing rather than emitting an empty "
|
|
"payload that would read as 'nothing was found'",
|
|
code="budget_admits_nothing",
|
|
)
|
|
if spent > limit:
|
|
# Structurally unreachable while the knapsack bucket arithmetic is
|
|
# one-sided, and kept because SS 7.3 is a MUST about the emitted
|
|
# payload rather than about the algorithm that produced it.
|
|
raise ConsumeError(f"spent ({spent}) exceeds limit ({limit})", code="budget_exceeded")
|
|
question_terms = list(dict.fromkeys(normalise(question)))
|
|
return {
|
|
"contract": CONTRACT_REVISION,
|
|
"bundle": {
|
|
"bundle_id": root_bundle_id,
|
|
"ref": bundle_ref(bundle_root, profile=profile),
|
|
},
|
|
"budget": {
|
|
"unit": BUDGET_UNIT,
|
|
"instrument": BUDGET_INSTRUMENT,
|
|
"limit": limit,
|
|
"spent": spent,
|
|
"known_positive": {
|
|
"case": case,
|
|
"expected": expected,
|
|
"measured": measured,
|
|
# The second, independent route (`wc -c` on the same file), so
|
|
# `expected == measured` is not the only thing standing between
|
|
# a broken instrument and a green gate.
|
|
"raw_bytes": len(known_positive_path().read_bytes()),
|
|
"encoding_delta": KNOWN_POSITIVE_ENCODING_DELTA,
|
|
},
|
|
# Present only when a reservation was made, because a cut whose
|
|
# strategy changed without saying so is the silent cut SS 5.3
|
|
# forbids -- and absent otherwise, so the default payload keeps
|
|
# every byte it had.
|
|
**(
|
|
{"reserved": {"concept_id": reserved[0], "bytes": reserved[1]}}
|
|
if reserved is not None
|
|
else {}
|
|
),
|
|
},
|
|
"denominators": {
|
|
"considered": len(concepts),
|
|
"withheld": len(withheld),
|
|
"delivered": len(delivered),
|
|
},
|
|
"question": question,
|
|
# SS 8: what of the QUESTION this payload reaches. Facts, and no
|
|
# verdict -- see `unanswered_terms` for the two readings that were
|
|
# measured and felled.
|
|
"coverage": {
|
|
"question_terms": question_terms,
|
|
"unanswered_in_bundle": unanswered_terms(question_terms, tokenised, stems=stems),
|
|
"unanswered_in_payload": unanswered_terms(
|
|
question_terms,
|
|
[normalise(excerpt_text(excerpt)) for excerpt in delivered],
|
|
stems=stems,
|
|
),
|
|
},
|
|
"excerpts": list(delivered),
|
|
# Emitted only under the flag, and then only where the concept carries
|
|
# a title, so a bundle whose concepts have none produces the same bytes
|
|
# either way. See this function's docstring for the measurement that
|
|
# keeps the default off.
|
|
"withheld": [
|
|
{"concept_id": concept_id, "rule": rule}
|
|
| (
|
|
{"title": titles_by_id[concept_id]}
|
|
if withheld_titles and titles_by_id.get(concept_id)
|
|
else {}
|
|
)
|
|
for concept_id, rule in withheld
|
|
],
|
|
}
|
|
|
|
|
|
def serialise(payload: Mapping[str, object]) -> str:
|
|
"""The payload as the bytes that are actually emitted.
|
|
|
|
`ensure_ascii=False` is load-bearing rather than cosmetic: the default
|
|
inflates this corpus by 7.1 % (1 950 745 -> 2 089 391 B), so a gate
|
|
measuring one form and a knapsack weighing the other disagree by more than
|
|
the headroom. LF only, one trailing newline.
|
|
"""
|
|
return json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=False) + "\n"
|
|
|
|
|
|
# --- The CLI ------------------------------------------------------------------
|
|
|
|
|
|
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
|
parser = argparse.ArgumentParser(
|
|
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
|
|
)
|
|
parser.add_argument("bundle", type=Path, help="the OKF bundle directory to read")
|
|
parser.add_argument("--question", required=True, help="the question to cut the bundle for")
|
|
parser.add_argument(
|
|
"--k", type=int, default=DEFAULT_K, help=f"cap on delivered excerpts (default {DEFAULT_K})"
|
|
)
|
|
parser.add_argument(
|
|
"--limit",
|
|
type=int,
|
|
default=DEFAULT_LIMIT,
|
|
help=f"budget in {BUDGET_UNIT} (default {DEFAULT_LIMIT})",
|
|
)
|
|
parser.add_argument(
|
|
"--cost-vocabulary",
|
|
action="store_true",
|
|
help=(
|
|
"let the declared cost/price/quantity vocabulary bridge a question "
|
|
"and a document that share no word. OFF by default; a question "
|
|
"naming no term in that vocabulary is unaffected either way"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--reserve-top-rank",
|
|
action="store_true",
|
|
help=(
|
|
"give the highest-ranked candidate its bytes before the budget is "
|
|
"packed, so a large top-ranked excerpt is not out-summed by small "
|
|
"ones. OFF by default; a candidate that alone exceeds the budget is "
|
|
"still refused"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--rarity-weight",
|
|
action="store_true",
|
|
help=(
|
|
"weight each lexical hit by log(N/df) over the bundle's own "
|
|
"concepts instead of counting it as one. OFF by default, and the "
|
|
"default is a MEASUREMENT rather than a preference: measured on "
|
|
"four corpora it moved one gold rank 9->8, left one at 35 and made "
|
|
"one 96->103 worse. See docs/2026-09-08-sjeldenhetsvekt.md"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--tie-shared-rank",
|
|
action="store_true",
|
|
default=DEFAULT_TIE_SHARED_RANK,
|
|
help=(
|
|
"let concepts a signal scores EQUALLY share that score group's "
|
|
"first rank, so a signal that separates nothing contributes the "
|
|
"same constant to each of them instead of ordering them by id. ON "
|
|
"since 2026-09-10. See docs/2026-09-08-rangeringsbom-sammensatte-ord.md "
|
|
"for the rule and docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md "
|
|
"for why its published cost no longer holds"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--no-tie-shared-rank",
|
|
action="store_false",
|
|
dest="tie_shared_rank",
|
|
help="The rule's explicit opt-out, reproducing the pre-2026-09-10 order",
|
|
)
|
|
parser.add_argument(
|
|
"--stem-prefix",
|
|
action="store_true",
|
|
default=DEFAULT_STEM_PREFIX,
|
|
help=(
|
|
"require a shared PREFIX to be a word the bundle uses, so four "
|
|
"coincidental characters no longer match. Measured on the pinned "
|
|
"453-concept bundle: `bilateral` occurs 0 times and reached 400 of "
|
|
"453 through `bilag`; with this it reaches 0, and every hit@8 row "
|
|
"keeps rank 1 on both bundles. ON since 2026-09-09"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--no-stem-prefix",
|
|
action="store_false",
|
|
dest="stem_prefix",
|
|
help="The rule's explicit opt-out, reproducing the pre-round-10 matcher",
|
|
)
|
|
parser.add_argument(
|
|
"--title-covered",
|
|
action="store_true",
|
|
default=DEFAULT_TITLE_COVERED,
|
|
help=(
|
|
"read a concept whose WHOLE title the question accounts for before "
|
|
"the concepts the fusion ranked above it, stopping beneath any whose "
|
|
"title answers MORE question tokens than it holds (round 17). ON "
|
|
"since 2026-09-10. "
|
|
"Measured on a 2 761-concept bundle of one standard: hit@1 over six "
|
|
"questions 3 of 6 -> 6 of 6 with the known-positive holding rank 1, "
|
|
"where none of the six existing reading-side flags moved that "
|
|
"number at all. The title is read by EQUALITY, never by shared "
|
|
"prefix. See docs/2026-09-10-k3-runde16-hele-tittelen-tar-ruten.md"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--no-title-covered",
|
|
action="store_false",
|
|
dest="title_covered",
|
|
help="The rule's explicit opt-out, reproducing the pre-round-16 excerpt order",
|
|
)
|
|
parser.add_argument(
|
|
"--source-quota",
|
|
type=int,
|
|
default=DEFAULT_SOURCE_QUOTA,
|
|
metavar="N",
|
|
help=(
|
|
"cap how many DELIVERED places one source document may take, "
|
|
"filling the freed places from the next candidate so k is still "
|
|
"delivered in full. Default 2 since 2026-09-10. Measured on a "
|
|
"3206-concept bundle whose corpus holds its own table of contents: "
|
|
"that one document took 8 of 8 places and the answer was not "
|
|
"delivered at all; at 2 it takes 2 of 8 and the answer comes in at "
|
|
"rank 4. A bundle with no alternatives is unaffected -- the "
|
|
"shortlist is topped back up to k"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--no-source-quota",
|
|
action="store_const",
|
|
const=None,
|
|
dest="source_quota",
|
|
help="The rule's explicit opt-out, reproducing the pre-round-11 excerpt order",
|
|
)
|
|
parser.add_argument(
|
|
"--follow-parent",
|
|
action="store_true",
|
|
default=DEFAULT_FOLLOW_PARENT,
|
|
help=(
|
|
"carry the enclosing concept's text inside an excerpt's `parent`, "
|
|
"with that concept's sha256, from the room the cut LEFT and in rank "
|
|
"order -- so it never displaces an excerpt; a text that does not fit "
|
|
"is clipped and marked `truncated` (K3-21)"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--no-follow-parent",
|
|
action="store_false",
|
|
dest="follow_parent",
|
|
help="The rule's explicit opt-out: `parent` is the pointer alone",
|
|
)
|
|
parser.add_argument(
|
|
"--withheld-titles",
|
|
action="store_true",
|
|
help=(
|
|
"give each withheld entry the concept's title, so a reader can see "
|
|
"WHAT was withheld without reading the bundle. OFF by default: it "
|
|
"grew a 270-concept payload by 37.9 %% and pushed a 629-concept "
|
|
"bundle's bookkeeping past the budget limit itself"
|
|
),
|
|
)
|
|
parser.add_argument("--out", type=Path, default=None, help="write here instead of stdout")
|
|
parser.add_argument(
|
|
"--ref",
|
|
default=None,
|
|
help=(
|
|
"assert the bundle's content identity. NOT an override: the identity "
|
|
"is computed regardless and a mismatch refuses, because labelling a "
|
|
"payload with an identity its bytes do not have is what SS 3.3 exists "
|
|
"to prevent"
|
|
),
|
|
)
|
|
return parser.parse_args(argv)
|
|
|
|
|
|
def main(argv: list[str] | None = None) -> int:
|
|
"""Three exit codes, not two.
|
|
|
|
**0** a payload was written, **1** the run happened and refused, **2** the
|
|
run did not happen. Collapsing 2 into 1 would report an unread bundle as a
|
|
failed cut -- two findings with different owners under one number.
|
|
"""
|
|
args = parse_args(argv)
|
|
if not args.bundle.is_dir():
|
|
print(f"okf_consume: FAILED - {args.bundle} is not a directory", file=sys.stderr)
|
|
return 2
|
|
try:
|
|
payload = build_payload(
|
|
args.bundle,
|
|
question=args.question,
|
|
k=args.k,
|
|
limit=args.limit,
|
|
cost_vocabulary=args.cost_vocabulary,
|
|
reserve_top_rank=args.reserve_top_rank,
|
|
rarity_weight=args.rarity_weight,
|
|
tie_shared_rank=args.tie_shared_rank,
|
|
title_covered=args.title_covered,
|
|
stem_prefix=args.stem_prefix,
|
|
source_quota=args.source_quota,
|
|
withheld_titles=args.withheld_titles,
|
|
follow_parent=args.follow_parent,
|
|
)
|
|
except ConsumeError as error:
|
|
print(f"okf_consume: FAILED - {error}", file=sys.stderr)
|
|
return 1
|
|
except OSError as error:
|
|
print(f"okf_consume: FAILED - the bundle could not be read: {error}", file=sys.stderr)
|
|
return 2
|
|
except UnicodeDecodeError as error:
|
|
print(f"okf_consume: FAILED - the bundle is not utf-8: {error}", file=sys.stderr)
|
|
return 2
|
|
computed = payload["bundle"]
|
|
assert isinstance(computed, dict)
|
|
if args.ref is not None and args.ref != computed["ref"]:
|
|
# Refuses BEFORE writing: a payload on disk under an asserted ref that
|
|
# the bytes contradict is worse than no payload.
|
|
print(
|
|
f"okf_consume: FAILED - the bundle's identity is {computed['ref']}, "
|
|
f"not the asserted {args.ref}; refusing to write a payload the "
|
|
"caller would label with an identity its bytes do not have",
|
|
file=sys.stderr,
|
|
)
|
|
return 1
|
|
text = serialise(payload)
|
|
if args.out is None:
|
|
sys.stdout.write(text)
|
|
else:
|
|
args.out.write_text(text, encoding="utf-8", newline="\n")
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
raise SystemExit(main())
|