114 lines
6 KiB
Markdown
114 lines
6 KiB
Markdown
# Test fixtures
|
|
|
|
## The PDF fixtures
|
|
|
|
`two-line-krav.pdf` and `no-text-layer.pdf` are hand-written minimal PDFs,
|
|
regenerated by `make_fixtures.py` in this directory:
|
|
|
|
```
|
|
python3 tests/fixtures/make_fixtures.py
|
|
```
|
|
|
|
They carry no library's output — the objects are laid out by hand and the xref
|
|
offsets computed from the emitted bytes — so they are auditable byte for byte
|
|
and reproducible from that one file.
|
|
|
|
| Fixture | What it is for |
|
|
|---|---|
|
|
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
|
|
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
|
|
|
|
## The office fixtures
|
|
|
|
`two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are
|
|
hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every
|
|
part is written out by hand and zipped with a fixed `date_time`, so they are
|
|
byte-reproducible and carry no converter's output.
|
|
|
|
**That last point is the whole policy, not a preference.** A `.docx` written by
|
|
the converter and then read by the converter proves only that the converter
|
|
agrees with itself, and would stay green through any conversion defect that is
|
|
symmetric — which is most of them.
|
|
|
|
| Fixture | What it is for |
|
|
|---|---|
|
|
| `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. |
|
|
| `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
|
|
| `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. |
|
|
|
|
Two things were measured while building these, and both are the same shape —
|
|
structurally valid input, silently reduced output, exit code 0 and no warning:
|
|
|
|
- **Without `word/styles.xml`** the docx extracts as flat prose with no heading.
|
|
A fixture lacking that part would pin the body and pin nothing about
|
|
structure, while looking exactly as convincing. Structure is the half the
|
|
segment proposer reads.
|
|
- **With inline strings (`t="inlineStr"`)** rather than a shared string table,
|
|
the xlsx extracts with the sheet name intact and **every cell value gone**.
|
|
The fixture therefore uses a `dimension` element and a shared string table.
|
|
|
|
## The proposer's default-profile golden
|
|
|
|
`propose-golden-default.json` is the artifact `tools/okf_propose_segments.py`
|
|
produces for `OUTLINE_DOCUMENT` (defined in `tests/test_propose_segments.py`)
|
|
with **no flags at all**, generated at commit `798f64a` with
|
|
`--proposed-at 2026-09-03T00:00:00Z`. The timestamp is an explicit argument
|
|
because the artifact carries it verbatim; a wall-clock default would make the
|
|
golden unreproducible by construction.
|
|
|
|
**Why the fixture is `OUTLINE_DOCUMENT` and not `DOCUMENT`.** The golden exists
|
|
to go red if any later rule is accidentally defaulted ON. `DOCUMENT` was
|
|
measured to contain **zero** bare-integer lines, so a golden over it would stay
|
|
byte-identical through exactly the regression it was named to catch -- a trap
|
|
written down but unable to fire. `OUTLINE_DOCUMENT` carries a bare-integer
|
|
ascending run of three, which today's rules do not match (measured: bare `1` /
|
|
`1.` / `1)` yield 0 candidates), so the golden pins that absence and breaks the
|
|
moment it stops being true.
|
|
|
|
**It transitively pins `observed_extractor_version`**
|
|
(`src/llm_ingestion_okf/segmentation.py`): the field is written into every
|
|
artifact, so a converter or extractor bump turns this golden red. That red is
|
|
legitimate -- read the diff and decide, exactly as for the frozen PDF literal
|
|
below. Regenerate only after that decision, never to make a red go away.
|
|
|
|
## Why the expected office text is frozen as a literal
|
|
|
|
The same reason as the PDF text below, with one addition: the literals are
|
|
pinned to a **named converter version**. `_pandoc.py` refuses any binary but
|
|
the vendored 3.9, and `tests/test_extract.py` asserts that version beside the
|
|
literals. A frozen literal without a named converter pins nothing — it says
|
|
"these bytes" without saying what produced them.
|
|
|
|
## Why the expected PDF text is frozen as a literal
|
|
|
|
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
|
|
exact string. That is deliberate, and it is the mechanism behind a promise this
|
|
library makes everywhere else:
|
|
|
|
- Extraction is deterministic **within** a parser version. Measured 2026-08-21
|
|
across five configurations, two runs each, compared byte for byte
|
|
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`).
|
|
- Extraction is **not** guaranteed stable **across** parser versions.
|
|
`pdfplumber` pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships
|
|
date-stamped releases with no stability contract. So the real pin on extracted
|
|
text is a transitive one, and it is exact.
|
|
|
|
The consequence is worth stating plainly: **any golden fixture built on
|
|
extracted PDF text is pinned to an exact parser version, and a parser upgrade
|
|
is a fixture migration, not a routine bump.** The frozen literal is what makes
|
|
that upgrade break something visible instead of drifting silently. If it goes
|
|
red after a dependency change, the correct response is to read the diff and
|
|
decide, not to re-record the expectation.
|
|
|
|
The version range that carries this lives in `pyproject.toml`'s
|
|
`[project.optional-dependencies] extract`, with the same reasoning at the
|
|
declaration site.
|
|
|
|
## What these fixtures do not cover
|
|
|
|
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
|
|
detected table objects are clean enough to hand to `render_table` unchanged;
|
|
two independent parsers return the same wrong shape, because the breakage is in
|
|
the documents' ruling geometry rather than in either library. PDFs enter this
|
|
library as **prose**, and structured tables are out of scope until that is
|
|
decided separately.
|