| .. | ||
| make_fixtures.py | ||
| no-styles-krav.docx | ||
| no-text-layer.pdf | ||
| propose-golden-default.json | ||
| README.md | ||
| two-line-krav.docx | ||
| two-line-krav.pdf | ||
| two-line-krav.xlsx | ||
Test fixtures
The PDF fixtures
two-line-krav.pdf and no-text-layer.pdf are hand-written minimal PDFs,
regenerated by make_fixtures.py in this directory:
python3 tests/fixtures/make_fixtures.py
They carry no library's output — the objects are laid out by hand and the xref offsets computed from the emitted bytes — so they are auditable byte for byte and reproducible from that one file.
| Fixture | What it is for |
|---|---|
two-line-krav.pdf |
One heading plus one requirement row with label and value on the same line. That pairing is the property pdfplumber was chosen for. |
no-text-layer.pdf |
A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (extractor_empty_pdf), never persist as an empty concept. |
The office fixtures
two-line-krav.docx, no-styles-krav.docx and two-line-krav.xlsx are
hand-laid OOXML containers, regenerated by the same make_fixtures.py. Every
part is written out by hand and zipped with a fixed date_time, so they are
byte-reproducible and carry no converter's output.
That last point is the whole policy, not a preference. A .docx written by
the converter and then read by the converter proves only that the converter
agrees with itself, and would stay green through any conversion defect that is
symmetric — which is most of them.
| Fixture | What it is for |
|---|---|
two-line-krav.docx |
A heading plus one requirement row with label and value on the same line — the docx mirror of two-line-krav.pdf. |
no-styles-krav.docx |
The same document without word/styles.xml. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
two-line-krav.xlsx |
A sheet name that becomes a heading, plus a label/value pair on one row. |
Two things were measured while building these, and both are the same shape — structurally valid input, silently reduced output, exit code 0 and no warning:
- Without
word/styles.xmlthe docx extracts as flat prose with no heading. A fixture lacking that part would pin the body and pin nothing about structure, while looking exactly as convincing. Structure is the half the segment proposer reads. - With inline strings (
t="inlineStr") rather than a shared string table, the xlsx extracts with the sheet name intact and every cell value gone. The fixture therefore uses adimensionelement and a shared string table.
The proposer's default-profile golden
propose-golden-default.json is the artifact tools/okf_propose_segments.py
produces for OUTLINE_DOCUMENT (defined in tests/test_propose_segments.py)
with no flags at all, generated at commit 798f64a with
--proposed-at 2026-09-03T00:00:00Z. The timestamp is an explicit argument
because the artifact carries it verbatim; a wall-clock default would make the
golden unreproducible by construction.
Why the fixture is OUTLINE_DOCUMENT and not DOCUMENT. The golden exists
to go red if any later rule is accidentally defaulted ON. DOCUMENT was
measured to contain zero bare-integer lines, so a golden over it would stay
byte-identical through exactly the regression it was named to catch -- a trap
written down but unable to fire. OUTLINE_DOCUMENT carries a bare-integer
ascending run of three, which today's rules do not match (measured: bare 1 /
1. / 1) yield 0 candidates), so the golden pins that absence and breaks the
moment it stops being true.
It transitively pins observed_extractor_version
(src/llm_ingestion_okf/segmentation.py): the field is written into every
artifact, so a converter or extractor bump turns this golden red. That red is
legitimate -- read the diff and decide, exactly as for the frozen PDF literal
below. Regenerate only after that decision, never to make a red go away.
Why the expected office text is frozen as a literal
The same reason as the PDF text below, with one addition: the literals are
pinned to a named converter version. _pandoc.py refuses any binary but
the vendored 3.9, and tests/test_extract.py asserts that version beside the
literals. A frozen literal without a named converter pins nothing — it says
"these bytes" without saying what produced them.
Why the expected PDF text is frozen as a literal
tests/test_extract.py asserts the extracted text of two-line-krav.pdf as an
exact string. That is deliberate, and it is the mechanism behind a promise this
library makes everywhere else:
- Extraction is deterministic within a parser version. Measured 2026-08-21
across five configurations, two runs each, compared byte for byte
(
docs/2026-08-21-g2-pdf-extraction-measurement.md). - Extraction is not guaranteed stable across parser versions.
pdfplumberpinspdfminer.six==20260107exactly, andpdfminer.sixships date-stamped releases with no stability contract. So the real pin on extracted text is a transitive one, and it is exact.
The consequence is worth stating plainly: any golden fixture built on extracted PDF text is pinned to an exact parser version, and a parser upgrade is a fixture migration, not a routine bump. The frozen literal is what makes that upgrade break something visible instead of drifting silently. If it goes red after a dependency change, the correct response is to read the diff and decide, not to re-record the expectation.
The version range that carries this lives in pyproject.toml's
[project.optional-dependencies] extract, with the same reasoning at the
declaration site.
What these fixtures do not cover
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
detected table objects are clean enough to hand to render_table unchanged;
two independent parsers return the same wrong shape, because the breakage is in
the documents' ruling geometry rather than in either library. PDFs enter this
library as prose, and structured tables are out of scope until that is
decided separately.