llm-ingestion-okf/tests/test_docs_promises.py
Kjell Tore Guttormsen 6ea8fcd3c7 feat(quality): okf quality, a per-file-type verdict with the denominator
G37. `okf check` is a CONTRACT check and a green one is not a quality gate:
measured 2026-09-10 by `vegnormal-okf`, three arms over one corpus all
returned 0 findings and exit 0 while their hit@k ranged from 6 of 6 to 0 of 6.
`okf quality <bundle>` asks the other question, per file type, with the
denominator on every line. A separate command rather than `okf check
--quality`, because the two answer different questions and a caller must not
be able to read one as the other. `okf check` is untouched.

Three verdicts and no fourth -- PASS, FAIL, UNMEASURED -- and a type with no
measured threshold is never PASS. Exit 0 judged and clean, 1 at least one
FAIL, 2 did not run, 3 nothing could be judged: exit 0 over a table of
unmeasured rows would be the silent pass this command exists to stop.

Two bars today, both `structure_null_share` (documents of a type yielding
exactly one concept), read off the pinned 43-document reference bundle: .pdf
8/32, .docx 2/5. Plus one definitional bar for every type, taken from the
harness's own degenerate-merge definition: 0 concepts with an empty body,
measured 0 of 8 602 concepts over four bundles. A bar needs five documents on
BOTH sides -- its own and the judged bundle's -- so .xlsx (2), .xml (1) and
every type with no corpus class in `extract._EVIDENCE` are UNMEASURED and
print their numbers without a verdict.

The floor on the judged bundle was found by RUNNING the gate, not by reading
it: one PDF cut into 2 182 concepts scored 0 of 1 against the 32-document
reference and read as PASS.

The gate walks the index tree and never a directory (SS 9.2; controlled
against the listing on four bundles, 453 / 2 761 / 3 206 / 446 either way),
and prints the bundle's own run log beside its counts -- a document rejected
at extraction leaves no row in the bundle, so the pinned corpus's 33 PDFs
show up as 32 and the two denominators must never be read as one.

Three of the order's five premises moved when re-measured, and they are in the
document rather than glossed: the four evidence corpora carry `source_file` on
0 of 446, 0 of 1 133, 0 of 270 and 0 of 2 756 concepts, so they name no file
type and cannot PASS; "41,6 %" is `vegnormal-okf`'s number and not in this
repository; and the same 828-document bundle carries two published hit@k
figures from two question sets.

Three candidate metrics measured and NOT shipped: duplicate titles within a
document (0 of 3 206 on the known-bad arm against 349 of 2 761 on the
known-good one) and short concepts (5.6 % against 14.6 %) order the two arms
the wrong way round; duplicate titles across the whole bundle order all four
correctly (37.8 / 16.3 / 12.6 / 5.7 %) and still ship without a bar, because
any bar separating them is read off the two bundles it would judge.

19 new tests, each rule exercised in both directions; the three README pins
were each driven red before being kept. Suite 1 850 passed, 1 skipped, 1 851
collected, run after `git add` -- +19 against a base of 1 832 collected,
measured on the stashed tree (STATE's 1 831 is one short of that).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 00:13:00 +02:00

255 lines
11 KiB
Python

"""The published format promise, asserted rather than trusted.
`README.md` told consumers that `docx` and `xlsx` ship no parser and always
fail fast. That was true when it was written and became false the moment the
converter seam landed -- silently, because prose has no test.
This library already learned that lesson once: a published promise without a
test goes false without anyone noticing, and a guarantee made publicly is a
test obligation. So the README's claimed format list is compared against the
registries it describes. Adding a format without touching the README, or
describing one that does not exist, fails here.
"""
from __future__ import annotations
import re
from pathlib import Path
from llm_ingestion_okf.extract import (
_CORE_EXTRACTORS,
_EVIDENCE,
_OPTIONAL_EXTRACTORS,
_PANDOC_FORMATS,
)
PROJECT_ROOT = Path(__file__).resolve().parents[1]
README = PROJECT_ROOT / "README.md"
# The line the README carries, and the one place this list is written in prose.
_FORMAT_LINE = re.compile(r"^<!-- extract-formats: (.+) -->$", re.MULTILINE)
def _declared_formats() -> set[str]:
match = _FORMAT_LINE.search(README.read_text(encoding="utf-8"))
assert match is not None, (
"README.md carries no `<!-- extract-formats: ... -->` marker; without "
"it this test cannot check the promise and the promise can drift"
)
return {token.strip() for token in match.group(1).split(",")}
def test_the_readme_names_exactly_the_formats_that_exist() -> None:
assert _declared_formats() == set(_CORE_EXTRACTORS) | set(_OPTIONAL_EXTRACTORS)
def test_the_readme_no_longer_claims_docx_and_xlsx_fail_fast() -> None:
"""The specific false sentence, pinned so it cannot come back.
Written as a search for the claim rather than for its exact wording: the
sentence could be rephrased and stay just as wrong.
"""
text = README.read_text(encoding="utf-8").lower()
for claim in (
"docx` and `xlsx` ship no parser",
"docx`/`xlsx` remain\nunimplemented",
"docx`/`xlsx` are still unimplemented",
):
assert claim.lower() not in text, f"README still claims: {claim}"
def test_the_readme_states_which_rows_are_not_measured() -> None:
"""A row that is not `measured` must not read as a supported one.
Three of the five office formats have denominator ZERO in the corpus this
work was measured on. A consumer reading the README should be able to see
that without reading the source.
Reads the CLASS from the table rather than the literal `unmeasured`: round
9 moved those three rows to `constructed`, and a test pinned to one word
would have gone green over an empty set the moment the word changed. Every
class that is not `measured` must be named in the README, whichever it is.
"""
text = README.read_text(encoding="utf-8")
weaker = {s.lstrip("."): e for s, e in _EVIDENCE.items() if e != "measured"}
assert weaker, "the evidence table lists no rows weaker than measured"
for suffix, evidence in weaker.items():
assert suffix in text, f"README does not mention the {evidence} row {suffix}"
assert evidence in text.lower(), f"README does not use the word {evidence}"
def test_the_readme_still_states_what_stays_out() -> None:
"""`.doc` (Word 97) and rastered PDFs are out, and stay named.
A format list that grows without also saying what it excludes reads as a
promise to handle anything office-shaped.
"""
text = README.read_text(encoding="utf-8")
assert ".doc`" in text or "Word 97" in text
assert ".doc" not in set(_PANDOC_FORMATS)
def test_the_readme_recursion_claim_matches_the_door() -> None:
"""The README says the drop directory is walked recursively. A sentence is
not a mechanism, so both halves are asserted here: the claim is in the
prose, and the door actually does it. Either one alone can go stale --
prose that outlived the code is the failure this whole module exists for.
"""
import tempfile
text = README.read_text(encoding="utf-8")
assert "walked **recursively**" in text
from llm_ingestion_okf.inbox import GateDecision, process_inbox
with tempfile.TemporaryDirectory() as workspace:
inbox = Path(workspace) / "inbox" / "sub"
inbox.mkdir(parents=True)
(inbox / "deep.md").write_text("Body\n", encoding="utf-8")
result = process_inbox(
Path(workspace) / "inbox",
Path(workspace) / "bundle",
"2026-09-07T08:00:00Z",
okf_type="reference",
gate=lambda body: GateDecision(sanitized_text=body, disposition="warn", reasons=()),
)
assert [item.source_file for item in result.persisted] == ["sub/deep.md"]
# --- the visible table, K3-26 ----------------------------------------------
#
# The comment marker above is machine-readable and invisible to a reader: the
# README's own prose named FIVE of the thirteen types the registry reads, and
# nothing went red, because the marker test only asks that the hidden list is
# complete. A reader does not read the marker. So the table a reader does see
# is pinned to the same registry, row for row, and to the evidence class the
# code records for each row.
_TABLE_HEADING = "## Supported file types"
# The evidence cell for a row `_EVIDENCE` does not carry. Those five are the
# stdlib rows: `_EVIDENCE` records a CORPUS class, and a row that has never
# been given one must not borrow `measured` from the row beside it.
_NO_CLASS = "stdlib, no corpus class"
def _table_rows() -> dict[str, list[str]]:
"""The table's data rows, keyed by suffix, with markup stripped per cell.
Backticks and asterisks are removed rather than matched, so the table can
be formatted freely and this test still reads what it says.
"""
text = README.read_text(encoding="utf-8")
assert _TABLE_HEADING in text, (
f"README.md carries no `{_TABLE_HEADING}` section; the format list is "
"then visible only in a hidden comment, which is what K3-26 fixed"
)
rows: dict[str, list[str]] = {}
for line in text.split(_TABLE_HEADING, 1)[1].splitlines():
stripped = line.strip()
if not stripped.startswith("|"):
if rows:
break
continue
cells = [re.sub(r"[`*]", "", cell).strip() for cell in stripped.strip("|").split("|")]
if cells and cells[0].startswith("."):
rows[cells[0]] = cells
return rows
def test_the_readme_table_names_every_type_the_registry_reads() -> None:
assert set(_table_rows()) == set(_CORE_EXTRACTORS) | set(_OPTIONAL_EXTRACTORS)
def test_the_readme_table_states_the_evidence_class_the_code_records() -> None:
rows = _table_rows()
assert rows, "no data rows found under the supported-file-types heading"
for suffix, cells in rows.items():
assert len(cells) >= 4, f"the {suffix} row has no evidence column: {cells}"
assert cells[3] == _EVIDENCE.get(suffix, _NO_CLASS), (
f"the {suffix} row says {cells[3]!r}; the code records "
f"{_EVIDENCE.get(suffix, _NO_CLASS)!r}"
)
def test_the_readme_table_separates_core_from_the_extract_extra() -> None:
"""Which rows need the optional extra is the first thing a consumer asks."""
for suffix, cells in _table_rows().items():
gated = "[extract]" in cells[2]
assert gated == (suffix in _OPTIONAL_EXTRACTORS), (
f"the {suffix} row's dependency cell reads {cells[2]!r}"
)
def test_the_readme_opening_does_not_name_five_of_thirteen() -> None:
"""The first thing a reader sees must not undersell what the code reads.
Either form passes: a pointer to the table, or the whole set spelled out.
A partial list -- the state before K3-26 -- passes neither.
"""
intro = README.read_text(encoding="utf-8").split("## Install", 1)[0]
if "#supported-file-types" in intro:
return
every = set(_CORE_EXTRACTORS) | set(_OPTIONAL_EXTRACTORS)
named = {s for s in every if re.search(rf"\b{s.lstrip('.')}\b", intro, re.IGNORECASE)}
assert named == every, f"the opening names {sorted(named)}, not all of {sorted(every)}"
def test_the_readme_carries_only_one_file_type_table() -> None:
"""A second table over the same rows is a copy nothing checks.
`### Binary extraction` carried its own six-row Format/Reader/Evidence
table until 2026-09-12. It was true when written, and it was reachable by
exactly the failure this module exists for: an evidence class copied into
prose that no test reads. Its rows now live in the pinned table alone.
"""
text = README.read_text(encoding="utf-8")
section = text.split("### Binary extraction", 1)[1].split("\n## ", 1)[0]
rows = [line for line in section.splitlines() if line.strip().startswith("|")]
assert not rows, f"a second file-type table is back under Binary extraction: {rows}"
# The `okf quality` thresholds, in the one place the README writes them. A bar
# published without a test goes false the way the format list did.
_THRESHOLD_LINE = re.compile(r"^<!-- quality-thresholds: (.+) -->$", re.MULTILINE)
THRESHOLD_DOCUMENT = PROJECT_ROOT / "docs" / "2026-09-12-g37-terskler.md"
def _declared_thresholds() -> dict[str, str]:
match = _THRESHOLD_LINE.search(README.read_text(encoding="utf-8"))
assert match is not None, (
"README.md carries no `<!-- quality-thresholds: ... -->` marker; without "
"it the published bars can drift from the ones the gate applies"
)
pairs = (token.strip().split("=") for token in match.group(1).split(","))
return {extension: share for extension, share in pairs}
def test_the_readme_names_exactly_the_thresholds_the_gate_applies() -> None:
from llm_ingestion_okf.quality import THRESHOLDS
assert _declared_thresholds() == {
extension: threshold.as_share() for extension, threshold in THRESHOLDS.items()
}
def test_the_threshold_document_carries_the_same_bars() -> None:
"""Three copies, one measurement: the code, the README and the document.
The document is where a bar's N and corpus live, so a bar that moved in the
code without moving there would publish a number nobody measured.
"""
from llm_ingestion_okf.quality import THRESHOLDS
text = THRESHOLD_DOCUMENT.read_text(encoding="utf-8")
for extension, threshold in THRESHOLDS.items():
row = f"| `{extension}` | `{threshold.metric}` | **{threshold.as_share()}** |"
assert row in text, f"{THRESHOLD_DOCUMENT.name} carries no row {row}"
def test_the_readme_quality_section_does_not_promise_a_quality_claim() -> None:
"""The one sentence that must not come back: PASS as a statement of quality."""
text = README.read_text(encoding="utf-8").lower()
assert "okf quality" in text
assert "regression bar against a pinned artifact" in text