llm-ingestion-okf/tests/fixtures/README.md
Kjell Tore Guttormsen a7b050b569 test(fixtures): a synthetic K2 denominator for pptx, odt and rtf
`docs/2026-09-04-k2-pptx-odt-rtf.md` measured the corpus denominator for
these three office rows and found it ZERO: `K2/trinn1` holds 43 files and
not one of them is a `pptx`, an `odt` or an `rtf`. So `extract._EVIDENCE`
calls those rows `unmeasured` in the strongest sense available -- they
work by construction and had never met a document at all.

This is the smallest thing that changes that without inventing a corpus.
One authored document -- a title, an intro, a 20-row label/value table, a
caption and a 4x4 grid -- laid out three times in three containers, so
the container and its reader are the only variable between the three
measurements. `k2-office-fasit.json` carries the hand count taken from
the AUTHORED content rather than from any converter's output: 56 cells,
20 pairs, 59 distinct strings, shared by all three. It is committed here,
before the measurement runs, because a fasit written afterwards is a
description of a result rather than a denominator for it.

No converter wrote any of these files. `make_k2_office.py` lays every
part by hand, for the reason `make_fixtures.py` already states and this
set inherits: a file written by the converter and then read by the
converter proves only that the converter agrees with itself, and stays
green through any conversion defect that is symmetric. The commissioning
order offered pandoc as one generator option; the committed fixture
policy forbids it and the policy wins.

Two converter behaviours were measured while laying the RTF out, both of
them structurally plausible input read silently wrong, exit code 0 and no
warning. Without `\pard\intbl` on cell paragraphs, consecutive
`\trowd...\row` rows come back as each row NESTED inside the previous
one: five label/value rows read as five levels of nested table, 2076
characters where 117 were expected. And the `\uN?` unicode escape -- the
form Word emits -- loses the character after it: `A\u248?BC` reads back
as `AoC` with the `B` gone, `A\u248?xBC` reads back as `AoBC`. The
fixture writes `\uN ?` with an explicit space, which round-trips. Neither
is worked around anywhere in `src/`.

The generator and the fasit live one level above `k2-office/` and that is
not tidiness: Door B walks its drop directory recursively, so anything
parked beside the three documents would enter the run and N would stop
being 3.

Three synthetic documents in one house style are not a corpus. The rows
stay `unmeasured` and the suite asserts that they do.

Suite 1132 passed (1127 + 5), `ruff check` and `ruff format --check`
clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 05:17:18 +02:00

164 lines
8.5 KiB
Markdown

# Test fixtures
## The PDF fixtures
`two-line-krav.pdf` and `no-text-layer.pdf` are hand-written minimal PDFs,
regenerated by `make_fixtures.py` in this directory:
```
python3 tests/fixtures/make_fixtures.py
```
They carry no library's output — the objects are laid out by hand and the xref
offsets computed from the emitted bytes — so they are auditable byte for byte
and reproducible from that one file.
| Fixture | What it is for |
|---|---|
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
## The office fixtures
`two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are
hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every
part is written out by hand and zipped with a fixed `date_time`, so they are
byte-reproducible and carry no converter's output.
**That last point is the whole policy, not a preference.** A `.docx` written by
the converter and then read by the converter proves only that the converter
agrees with itself, and would stay green through any conversion defect that is
symmetric — which is most of them.
| Fixture | What it is for |
|---|---|
| `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. |
| `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
| `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. |
Two things were measured while building these, and both are the same shape —
structurally valid input, silently reduced output, exit code 0 and no warning:
- **Without `word/styles.xml`** the docx extracts as flat prose with no heading.
A fixture lacking that part would pin the body and pin nothing about
structure, while looking exactly as convincing. Structure is the half the
segment proposer reads.
- **With inline strings (`t="inlineStr"`)** rather than a shared string table,
the xlsx extracts with the sheet name intact and **every cell value gone**.
The fixture therefore uses a `dimension` element and a shared string table.
## The K2 office fixture set (`k2-office/`)
`krav-presentasjon.pptx`, `krav-tekstdokument.odt` and
`krav-rikt-tekstformat.rtf` are the synthetic denominator for the three office
rows the corpus has none of. `docs/2026-09-04-k2-pptx-odt-rtf.md` measured that
denominator at **zero**`K2/trinn1` holds 43 files and not one is a `pptx`,
an `odt` or an `rtf` — so those rows were `unmeasured` in the sense of never
having met a document at all. Regenerated by `make_k2_office.py` in this
directory:
```
python3 tests/fixtures/make_k2_office.py
```
**One document, three containers.** All three carry the same authored content —
a title, an intro, a 20-row label/value table, a caption and a 4x4 grid — so
the only variable between the three measurements is the container and the
reader that opens it. The counts are hand-counted once, in
`k2-office-fasit.json`, and shared: **56 cells, 20 pairs, 59 distinct strings.**
**The generator and the fasit live one level up, and that is not tidiness.**
Door B walks its drop directory recursively, so anything parked inside
`k2-office/` would enter the run and N would stop being 3.
Same policy as the office fixtures above, for the same reason: every part is
hand-laid and no converter wrote any of them. The commissioning order offered
pandoc as a generator option; a file written by the converter and then read by
the converter would prove only that the converter agrees with itself.
Two things were measured while building this set, both against the vendored
pandoc 3.9, and both are the house shape — structurally plausible input,
silently wrong output, exit code 0 and no warning:
- **RTF cell paragraphs need `\pard\intbl`.** Without it, consecutive
`\trowd…\row` rows are read as each row NESTED inside the previous one:
five label/value rows came back as five levels of nested table, 2076
characters where 117 were expected.
- **The `\uN?` unicode escape loses the character after it.** Measured
directly: `A\u248?BC` reads back as `AoC` (ring letter present, `B` gone) and
`A\u248?xBC` reads back as `AoBC`. The `?` is taken as the control word's
delimiter and `\uc1` then skips a real character. The fixture writes
`\uN ?` with an explicit space, which round-trips. **This is the form Word
emits**, so it is a converter finding rather than a fixture quirk — recorded
in `docs/2026-09-07-k2-pptx-odt-rtf-fixtures.md`, not worked around anywhere
in `src/`.
**Three synthetic documents in one house style are not a corpus.** The rows stay
`unmeasured` in `extract._EVIDENCE` and `tests/test_k2_office_fixtures.py`
asserts that they do.
## The proposer's default-profile golden
`propose-golden-default.json` is the artifact `tools/okf_propose_segments.py`
produces for `OUTLINE_DOCUMENT` (defined in `tests/test_propose_segments.py`)
with **no flags at all**, generated at commit `798f64a` with
`--proposed-at 2026-09-03T00:00:00Z`. The timestamp is an explicit argument
because the artifact carries it verbatim; a wall-clock default would make the
golden unreproducible by construction.
**Why the fixture is `OUTLINE_DOCUMENT` and not `DOCUMENT`.** The golden exists
to go red if any later rule is accidentally defaulted ON. `DOCUMENT` was
measured to contain **zero** bare-integer lines, so a golden over it would stay
byte-identical through exactly the regression it was named to catch -- a trap
written down but unable to fire. `OUTLINE_DOCUMENT` carries a bare-integer
ascending run of three, which today's rules do not match (measured: bare `1` /
`1.` / `1)` yield 0 candidates), so the golden pins that absence and breaks the
moment it stops being true.
**It transitively pins `observed_extractor_version`**
(`src/llm_ingestion_okf/segmentation.py`): the field is written into every
artifact, so a converter or extractor bump turns this golden red. That red is
legitimate -- read the diff and decide, exactly as for the frozen PDF literal
below. Regenerate only after that decision, never to make a red go away.
## Why the expected office text is frozen as a literal
The same reason as the PDF text below, with one addition: the literals are
pinned to a **named converter version**. `_pandoc.py` refuses any binary but
the vendored 3.9, and `tests/test_extract.py` asserts that version beside the
literals. A frozen literal without a named converter pins nothing — it says
"these bytes" without saying what produced them.
## Why the expected PDF text is frozen as a literal
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
exact string. That is deliberate, and it is the mechanism behind a promise this
library makes everywhere else:
- Extraction is deterministic **within** a parser version. Measured 2026-08-21
across five configurations, two runs each, compared byte for byte
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`).
- Extraction is **not** guaranteed stable **across** parser versions.
`pdfplumber` pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships
date-stamped releases with no stability contract. So the real pin on extracted
text is a transitive one, and it is exact.
The consequence is worth stating plainly: **any golden fixture built on
extracted PDF text is pinned to an exact parser version, and a parser upgrade
is a fixture migration, not a routine bump.** The frozen literal is what makes
that upgrade break something visible instead of drifting silently. If it goes
red after a dependency change, the correct response is to read the diff and
decide, not to re-record the expectation.
The version range that carries this lives in `pyproject.toml`'s
`[project.optional-dependencies] extract`, with the same reasoning at the
declaration site.
## What these fixtures do not cover
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
detected table objects are clean enough to hand to `render_table` unchanged;
two independent parsers return the same wrong shape, because the breakage is in
the documents' ruling geometry rather than in either library. PDFs enter this
library as **prose**, and structured tables are out of scope until that is
decided separately.