# Test fixtures ## The PDF fixtures `two-line-krav.pdf` and `no-text-layer.pdf` are hand-written minimal PDFs, regenerated by `make_fixtures.py` in this directory: ``` python3 tests/fixtures/make_fixtures.py ``` They carry no library's output — the objects are laid out by hand and the xref offsets computed from the emitted bytes — so they are auditable byte for byte and reproducible from that one file. | Fixture | What it is for | |---|---| | `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. | | `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. | ## The office fixtures `two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every part is written out by hand and zipped with a fixed `date_time`, so they are byte-reproducible and carry no converter's output. **That last point is the whole policy, not a preference.** A `.docx` written by the converter and then read by the converter proves only that the converter agrees with itself, and would stay green through any conversion defect that is symmetric — which is most of them. | Fixture | What it is for | |---|---| | `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. | | `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. | | `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. | Two things were measured while building these, and both are the same shape — structurally valid input, silently reduced output, exit code 0 and no warning: - **Without `word/styles.xml`** the docx extracts as flat prose with no heading. A fixture lacking that part would pin the body and pin nothing about structure, while looking exactly as convincing. Structure is the half the segment proposer reads. - **With inline strings (`t="inlineStr"`)** rather than a shared string table, the xlsx extracts with the sheet name intact and **every cell value gone**. The fixture therefore uses a `dimension` element and a shared string table. ## Why the expected office text is frozen as a literal The same reason as the PDF text below, with one addition: the literals are pinned to a **named converter version**. `_pandoc.py` refuses any binary but the vendored 3.9, and `tests/test_extract.py` asserts that version beside the literals. A frozen literal without a named converter pins nothing — it says "these bytes" without saying what produced them. ## Why the expected PDF text is frozen as a literal `tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an exact string. That is deliberate, and it is the mechanism behind a promise this library makes everywhere else: - Extraction is deterministic **within** a parser version. Measured 2026-08-21 across five configurations, two runs each, compared byte for byte (`docs/2026-08-21-g2-pdf-extraction-measurement.md`). - Extraction is **not** guaranteed stable **across** parser versions. `pdfplumber` pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped releases with no stability contract. So the real pin on extracted text is a transitive one, and it is exact. The consequence is worth stating plainly: **any golden fixture built on extracted PDF text is pinned to an exact parser version, and a parser upgrade is a fixture migration, not a routine bump.** The frozen literal is what makes that upgrade break something visible instead of drifting silently. If it goes red after a dependency change, the correct response is to read the diff and decide, not to re-record the expectation. The version range that carries this lives in `pyproject.toml`'s `[project.optional-dependencies] extract`, with the same reasoning at the declaration site. ## What these fixtures do not cover Structured table recovery. Measured on real Vegnormalene, only 45 of 196 detected table objects are clean enough to hand to `render_table` unchanged; two independent parsers return the same wrong shape, because the breakage is in the documents' ruling geometry rather than in either library. PDFs enter this library as **prose**, and structured tables are out of scope until that is decided separately.