llm-ingestion-okf/tests/fixtures
Kjell Tore Guttormsen 191de89f41 feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word
Three of round 9's four measured holes, each closed with a rule chosen on a
measurement rather than named as a limit.

`rtf` GIVES 0 SEGMENTS -> 6 of 6 AUTHORED TITLES over N = 4. The container has
no heading style, so the author's title is bold text. The grammar is markdown,
not `rtf`: the converter already writes that title as `**...**` in the same
output every office row produces, so no `rtf`-only heading form exists. Three
parameters were swept over 47 readable documents and ONE carried -- refusing a
line that ends in terminal punctuation takes false-positive lines from 9-12 to
1-2. A maximum title length (unlimited/40/60/80/120) and a
must-stand-between-blank-lines clause are both FLAT, so neither is in the rule.
The last false positive is closed by G1, the principle `_gate_outline` already
carries: recovery yields to declaration. False positives are then 0 of the 31
declaring documents by construction, and 0 of 27 on the corpus. Reach: 2 of 39
corpus documents, both `docx`, 0 of 33 `pdf` and 0 of 2 `xlsx`. Behind
`--bold-title`, default OFF pending the hit@8 measurement; the default bundle
is byte-identical without it.

BOTH ALTERNATIVES THE ORDER NAMED WERE MEASURED AND FELLED. A fourth hand-laid
fixture DECLARES heading styles in a stylesheet and the converter discards
them, emitting the same bold line -- so "read the declared headings out of the
markdown" has nothing to read. `rtf` -> `docx` -> markdown yields 0 ATX
headings on that same document, because the loss is in the `rtf` READER before
any writer sees the style. Fixtures are hand-laid in `make_k2_office.py` with
the fasit written first; they live in their own directory because Door B walks
a drop directory recursively and `k2-office/` reads its N off the listing.

THE PREFIX OVER-MATCH: THREE CANDIDATES MEASURED, ALL THREE FAILED ON ONE ROW.
Re-measured on the pinned 453-concept bundle with the control run first:
`under` occurs 79 times by equality and matches 172 by prefix, `undersjoisk` 0
and 172, `bilateral` 0 and 400 of 453, `standhaftig` 0 and 219. The two extra
known-negatives were FOUND, not chosen -- every 4-character prefix ranked by
document frequency, then a real word taken from the widest. A longer floor
(5-8), a coverage share (0.5-0.8) and a long-words-only floor (>= 8) each cost
row 1 its rank on the default bundle and the whole row on Arm B. Decomposed:
row 1's token `prisene` reaches its gold document through
`pris|sammenstilling` on four characters -- 0.57 of one word and 0.22 of the
other -- so the over-match and the wanted match are one mechanism.

THE FOURTH CANDIDATE IS THE ANSWER: the shared prefix must be a WORD the bundle
uses. `pris` is; `bila` and `stan` are not. `bilateral` 400 -> 0 and 512 -> 0,
`standhaftig` 219 -> 56 and 235 -> 33, every hit@8 row keeping rank 1 on BOTH
bundles. `undersjoisk` stops at 162 because `under` IS a word here -- a genuine
Norwegian morpheme, so that residual is a different answer, not a ceiling. ON
by default (`--no-stem-prefix`), pinned with its own known-negative on the
shipped bytes.

THE SHIM: a path importer holds the object `module_from_spec` made, and
`sys.modules[__name__] = _impl` never reaches it. Measured under both counting
methods -- 3 of 76 public names by `vars()`. One line copies the public names
into this file's globals; the dunder filter is load-bearing, because an
unfiltered copy overwrites `__name__` before the next line uses it as the alias
key. It restores attribute ACCESS and not patch-through, which is why the alias
stays. A CHANGELOG note under 0.7.0 and a shim docstring line say so, since
what the consumer asked for was the note.

Suite 1515 -> 1535.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 23:05:45 +02:00
..
consume-bundle test(consume): hit@8 over six questions against a random-ranker baseline 2026-09-07 09:37:14 +02:00
consume-provenance feat(consume): give every excerpt the name and the address an answer must cite 2026-09-08 15:01:41 +02:00
k2-office test(fixtures): a synthetic K2 denominator for pptx, odt and rtf 2026-09-07 05:17:18 +02:00
k2-rtf-variants feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word 2026-09-09 23:05:45 +02:00
font-heading-krav.pdf feat(extract,cli): typography as a PDF heading source and OCR behind an optional group, both off 2026-09-08 23:10:47 +02:00
k2-office-fasit.json feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word 2026-09-09 23:05:45 +02:00
make_fixtures.py feat(propose,cli): typography as a reserve, and the two of our own numbers it took to measure it 2026-09-09 00:25:51 +02:00
make_k2_office.py feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word 2026-09-09 23:05:45 +02:00
no-styles-krav.docx test(extract): hand-built office fixtures with frozen extracted text 2026-09-02 14:14:27 +02:00
no-text-layer.pdf feat(extract): implement pdf behind the [extract] extra with pdfplumber 2026-08-21 20:22:39 +02:00
numbered-font-krav.pdf feat(propose,cli): typography as a reserve, and the two of our own numbers it took to measure it 2026-09-09 00:25:51 +02:00
prisark.xlsx fix(extract,build): write a spreadsheet as pipe tables, stop linking the run log from the index 2026-09-08 10:06:58 +02:00
propose-golden-default.json test(propose): pin the default artifact with a committed golden 2026-09-07 01:20:31 +02:00
propose-golden-grid-default.json test(propose): pin today's grid-table split with a golden that can fire 2026-09-07 10:45:05 +02:00
README.md feat(inbox): point every concept at the document it came from, with a locator per format 2026-09-08 14:39:24 +02:00
three-page-krav.pdf feat(inbox): point every concept at the document it came from, with a locator per format 2026-09-08 14:39:24 +02:00
tomrad.xlsx feat(inbox): point every concept at the document it came from, with a locator per format 2026-09-08 14:39:24 +02:00
two-line-krav.docx test(extract): hand-built office fixtures with frozen extracted text 2026-09-02 14:14:27 +02:00
two-line-krav.pdf feat(extract): implement pdf behind the [extract] extra with pdfplumber 2026-08-21 20:22:39 +02:00
two-line-krav.xlsx test(extract): hand-built office fixtures with frozen extracted text 2026-09-02 14:14:27 +02:00

Test fixtures

The PDF fixtures

two-line-krav.pdf and no-text-layer.pdf are hand-written minimal PDFs, regenerated by make_fixtures.py in this directory:

python3 tests/fixtures/make_fixtures.py

They carry no library's output — the objects are laid out by hand and the xref offsets computed from the emitted bytes — so they are auditable byte for byte and reproducible from that one file.

Fixture What it is for
two-line-krav.pdf One heading plus one requirement row with label and value on the same line. That pairing is the property pdfplumber was chosen for.
no-text-layer.pdf A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (extractor_empty_pdf), never persist as an empty concept.
three-page-krav.pdf Three pages, one line of text each, and the middle page carries no text operators. The extractor drops empty pages, so the last page's text belongs to page 3 — which is what separates a page NUMBER from a count of the pages that produced text. Two pages could not tell those apart.

The office fixtures

two-line-krav.docx, no-styles-krav.docx and two-line-krav.xlsx are hand-laid OOXML containers, regenerated by the same make_fixtures.py. Every part is written out by hand and zipped with a fixed date_time, so they are byte-reproducible and carry no converter's output.

That last point is the whole policy, not a preference. A .docx written by the converter and then read by the converter proves only that the converter agrees with itself, and would stay green through any conversion defect that is symmetric — which is most of them.

Fixture What it is for
two-line-krav.docx A heading plus one requirement row with label and value on the same line — the docx mirror of two-line-krav.pdf.
no-styles-krav.docx The same document without word/styles.xml. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration.
two-line-krav.xlsx A sheet name that becomes a heading, plus a label/value pair on one row.
tomrad.xlsx Four rows with the third one empty. The converter renders an empty row as a pipe line of nothing but spaces, which is what a pipe table's own separator line also looks like — so a rule that reads the line rather than its position swallows the row and renumbers every row after it. Found on the K2 price sheet (8 empty rows, last row reported as 92 against a workbook that says 100); this fixture is what keeps it red.

Two things were measured while building these, and both are the same shape — structurally valid input, silently reduced output, exit code 0 and no warning:

  • Without word/styles.xml the docx extracts as flat prose with no heading. A fixture lacking that part would pin the body and pin nothing about structure, while looking exactly as convincing. Structure is the half the segment proposer reads.
  • With inline strings (t="inlineStr") rather than a shared string table, the xlsx extracts with the sheet name intact and every cell value gone. The fixture therefore uses a dimension element and a shared string table.

The K2 office fixture set (k2-office/)

krav-presentasjon.pptx, krav-tekstdokument.odt and krav-rikt-tekstformat.rtf are the synthetic denominator for the three office rows the corpus has none of. docs/2026-09-04-k2-pptx-odt-rtf.md measured that denominator at zeroK2/trinn1 holds 43 files and not one is a pptx, an odt or an rtf — so those rows were unmeasured in the sense of never having met a document at all. Regenerated by make_k2_office.py in this directory:

python3 tests/fixtures/make_k2_office.py

One document, three containers. All three carry the same authored content — a title, an intro, a 20-row label/value table, a caption and a 4x4 grid — so the only variable between the three measurements is the container and the reader that opens it. The counts are hand-counted once, in k2-office-fasit.json, and shared: 56 cells, 20 pairs, 59 distinct strings.

The generator and the fasit live one level up, and that is not tidiness. Door B walks its drop directory recursively, so anything parked inside k2-office/ would enter the run and N would stop being 3.

Same policy as the office fixtures above, for the same reason: every part is hand-laid and no converter wrote any of them. The commissioning order offered pandoc as a generator option; a file written by the converter and then read by the converter would prove only that the converter agrees with itself.

Two things were measured while building this set, both against the vendored pandoc 3.9, and both are the house shape — structurally plausible input, silently wrong output, exit code 0 and no warning:

  • RTF cell paragraphs need \pard\intbl. Without it, consecutive \trowd…\row rows are read as each row NESTED inside the previous one: five label/value rows came back as five levels of nested table, 2076 characters where 117 were expected.
  • The \uN? unicode escape loses the character after it. Measured directly: A\u248?BC reads back as AoC (ring letter present, B gone) and A\u248?xBC reads back as AoBC. The ? is taken as the control word's delimiter and \uc1 then skips a real character. The fixture writes \uN ? with an explicit space, which round-trips. This is the form Word emits, so it is a converter finding rather than a fixture quirk — recorded in docs/2026-09-07-k2-pptx-odt-rtf-fixtures.md, not worked around anywhere in src/.

Three synthetic documents in one house style are not a corpus. The rows stay unmeasured in extract._EVIDENCE and tests/test_k2_office_fixtures.py asserts that they do.

The proposer's default-profile golden

propose-golden-default.json is the artifact tools/okf_propose_segments.py produces for OUTLINE_DOCUMENT (defined in tests/test_propose_segments.py) with no flags at all, generated at commit 798f64a with --proposed-at 2026-09-03T00:00:00Z. The timestamp is an explicit argument because the artifact carries it verbatim; a wall-clock default would make the golden unreproducible by construction.

Why the fixture is OUTLINE_DOCUMENT and not DOCUMENT. The golden exists to go red if any later rule is accidentally defaulted ON. DOCUMENT was measured to contain zero bare-integer lines, so a golden over it would stay byte-identical through exactly the regression it was named to catch -- a trap written down but unable to fire. OUTLINE_DOCUMENT carries a bare-integer ascending run of three, which today's rules do not match (measured: bare 1 / 1. / 1) yield 0 candidates), so the golden pins that absence and breaks the moment it stops being true.

It transitively pins observed_extractor_version (src/llm_ingestion_okf/segmentation.py): the field is written into every artifact, so a converter or extractor bump turns this golden red. That red is legitimate -- read the diff and decide, exactly as for the frozen PDF literal below. Regenerate only after that decision, never to make a red go away.

Why the expected office text is frozen as a literal

The same reason as the PDF text below, with one addition: the literals are pinned to a named converter version. _pandoc.py refuses any binary but the vendored 3.9, and tests/test_extract.py asserts that version beside the literals. A frozen literal without a named converter pins nothing — it says "these bytes" without saying what produced them.

Why the expected PDF text is frozen as a literal

tests/test_extract.py asserts the extracted text of two-line-krav.pdf as an exact string. That is deliberate, and it is the mechanism behind a promise this library makes everywhere else:

  • Extraction is deterministic within a parser version. Measured 2026-08-21 across five configurations, two runs each, compared byte for byte (docs/2026-08-21-g2-pdf-extraction-measurement.md).
  • Extraction is not guaranteed stable across parser versions. pdfplumber pins pdfminer.six==20260107 exactly, and pdfminer.six ships date-stamped releases with no stability contract. So the real pin on extracted text is a transitive one, and it is exact.

The consequence is worth stating plainly: any golden fixture built on extracted PDF text is pinned to an exact parser version, and a parser upgrade is a fixture migration, not a routine bump. The frozen literal is what makes that upgrade break something visible instead of drifting silently. If it goes red after a dependency change, the correct response is to read the diff and decide, not to re-record the expectation.

The version range that carries this lives in pyproject.toml's [project.optional-dependencies] extract, with the same reasoning at the declaration site.

What these fixtures do not cover

Structured table recovery. Measured on real Vegnormalene, only 45 of 196 detected table objects are clean enough to hand to render_table unchanged; two independent parsers return the same wrong shape, because the breakage is in the documents' ruling geometry rather than in either library. PDFs enter this library as prose, and structured tables are out of scope until that is decided separately.

propose-golden-grid-default.json

Pins the DEFAULT proposer artifact over a document containing a pandoc grid table. Generated on unmodified code, before Arm E's rule existed, with exactly the command the test runs:

write(tmp_path, GRID_GOLDEN_DOCUMENT, "grid.md")
okf_propose_segments.main([source, "--out", out, "--proposed-at", "2026-09-03T00:00:00Z"])

It exists because propose-golden-default.json cannot pin this. That golden is taken over OUTLINE_DOCUMENT, which contains no | row and no + rule line, so no table rule -- present or future -- can move its bytes. A guard that is structurally incapable of firing is a trap written down but never armed. This one is taken over a document that has a grid table, so an accidentally default-on table rule turns it red.

Both goldens transitively pin observed_extractor_version and PROPOSER_VERSION. A red here after a dependency change is a legitimate red: read the diff and decide, do not re-record the expectation.