llm-ingestion-okf/tests/fixtures/README.md
Kjell Tore Guttormsen e5dc21ec2f
test(accounting): the witnesses see what the formats actually hold (M-1..M-3)
Rows 2 and 3 require the build's inventory to EQUAL the witness's, so what
the witness does not count, nothing can lose visibly. An independent review
put a header and a comment in a docx, measured 0 of either in the bundle,
and the accounting still read "2 of 2 carried".

Thirteen classes are now counted, each with a red test written first:
docx header/footer, comment, endnote and text box (a box's paragraphs are
its own, or the text is booked twice) - pptx speaker note and hidden slide
(`show="0"`, no longer counted as an ordinary slide) - xlsx formula and
hidden sheet (the state lives in `workbook.xml` and is reached through the
relationship id, so the sheet part itself says nothing about it) - odt
header/footer from `styles.xml` and annotation (counted as prose, it made
the accounting demand a reader carry a note the author wrote to themselves)
- STS `mixed-citation`, `mml:math`, `fig` and its caption, measured by the
review at 4.1 % of N200's source text and 3.9 % of N100's.

M-2: the two STS witnesses had ONE role map between them, so row 5 -- "two
witnesses agree" -- could not see a hole in it. `_sts_role_xml` and
`_sts_role_json` are written apart, each for its own delivery, and a test
holds them apart.

M-3: 20 of 63 element types had a count of ZERO in their only fixture. Seven
hand-built documents close it, every element type now occurs at least once
(a test asserts it), and ALL TWENTY documents carry a hand count read off
the fixture's own bytes (four did before). `.xlsx image` -- the operator's
own proposed exception -- could not be exercised at all until now.

Every witness also states WHAT IT STILL DOES NOT COUNT, per file type, and
the gate prints that list on every run.

THE FIXTURE ROWS ARE RED NOW, AND THAT IS THE POINT. Row 2 red on .docx,
.odt, .pptx, .xlsx and .xml; row 3 at u = 25, d = 2 over the new classes,
including a footnote and four spreadsheet cells the build genuinely drops.
`0 claimed and not found` on the same run: nothing the build DOES book as
carried failed the bundle check, so the red is the build's and not the
instrument's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 01:53:50 +02:00

268 lines
17 KiB
Markdown

# Test fixtures
## The PDF fixtures
`two-line-krav.pdf` and `no-text-layer.pdf` are hand-written minimal PDFs,
regenerated by `make_fixtures.py` in this directory:
```
python3 tests/fixtures/make_fixtures.py
```
They carry no library's output — the objects are laid out by hand and the xref
offsets computed from the emitted bytes — so they are auditable byte for byte
and reproducible from that one file.
| Fixture | What it is for |
|---|---|
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
| `outline-collision.pdf` | **Two bookmarks whose destinations resolve to the same line** — the tree's root node and a front-matter node, both on line 0, which is the shape R761 carries. The marks are collected in a dict keyed on the line index, so without this fixture the second node is dropped with nothing counting it: 2 763 nodes in, 2 762 marks out, `unresolved` at 0. A one-bookmark-per-line fixture cannot see that. |
| `three-page-krav.pdf` | Three pages, one line of text each, and **the middle page carries no text operators**. The extractor drops empty pages, so the last page's text belongs to page 3 — which is what separates a page NUMBER from a count of the pages that produced text. Two pages could not tell those apart. |
## The XML fixtures
Hand-written text, directly in this directory — **not** through
`make_fixtures.py`, which builds binary containers, and never serialised by
`ElementTree`. The reason is the same one the OOXML policy states: a library
that writes and then reads its own format proves only that it agrees with
itself, and an `ElementTree`-written fixture would stay green through any
round-trip-symmetric defect.
| Fixture | What it is for |
|---|---|
| `sts-mini.xml` | The known-positive. A `<standard>` root with `<sec>` at three nesting levels carrying `<label>`+`<title>`, two lettered points (`a)`, `b)`) with a **label and no title**, one `<table-wrap>` with a label and two rows, `<p>` bodies, a `<list>`, and **one unnumbered section** (`Forord`, `<title>` with no `<label>`) mirroring the single such section in R761. The lettered points are what the 4 954 label-only `<sec>` in that document look like: promoted to headings they would bury its own 2 761. |
| `sts-empty-label.xml` | A `<sec>` carrying a `<label>` and **nothing else**, between a lettered point that has a body and the next titled section. The label is held as a prefix for a body line that never arrives, so it was overwritten and lost: measured on R761 that is exactly one `x)`, two characters of 1 283 395, ratio 0.999998. An exact invariant does not get to be 0.999998. |
| `sts-identity.xml` | A document that **states who it is**: exactly one `<std-ident>` with a `<doc-number>` and a `<year>`, one `<title-wrap>` whose `<full>` carries a **comma** (as R761's does, which is why that title cannot be written into a `sources` flow mapping verbatim), and a `<std-ref type="dated">`. Its body carries the `sec-type="spec"` shape the `description` rule reads: a titled `<sec>` whose first spec point has one `<p>`, a second spec point that must never become the description, a titled child with no spec point of its own, and a spec point with **two** `<p>` of which only the first counts. `sts-mini.xml` is the half identity (a `<title-wrap>`, no `<doc-number>`) and `sts-empty-label.xml` the absent one. |
| `generic-feed.xml` | Known-negative: XML that is **not** STS. It must produce text and ONE plan — never zero, never a crash, and never element names promoted to headings. |
| `xml-doctype-bomb.xml` | Known-negative, security: a `<!DOCTYPE` with a small nested-entity expansion. It must be refused by `code`, and the test asserts the expansion appears in **no** output, including the error text. Small on purpose — the point is that it is never parsed, not that it detonates. |
| `xml-malformed.xml` | Known-negative: an unterminated tag must raise a typed `ExtractionError`, not leak `ParseError` and not yield zero concepts in silence. |
## The office fixtures
`two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are
hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every
part is written out by hand and zipped with a fixed `date_time`, so they are
byte-reproducible and carry no converter's output.
**That last point is the whole policy, not a preference.** A `.docx` written by
the converter and then read by the converter proves only that the converter
agrees with itself, and would stay green through any conversion defect that is
symmetric — which is most of them.
| Fixture | What it is for |
|---|---|
| `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. |
| `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
| `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. |
| `tomrad.xlsx` | Four rows with the **third one empty**. The converter renders an empty row as a pipe line of nothing but spaces, which is what a pipe table's own separator line also looks like — so a rule that reads the line rather than its position swallows the row and renumbers every row after it. Found on the K2 price sheet (8 empty rows, last row reported as 92 against a workbook that says 100); this fixture is what keeps it red. |
Two things were measured while building these, and both are the same shape —
structurally valid input, silently reduced output, exit code 0 and no warning:
- **Without `word/styles.xml`** the docx extracts as flat prose with no heading.
A fixture lacking that part would pin the body and pin nothing about
structure, while looking exactly as convincing. Structure is the half the
segment proposer reads.
- **With inline strings (`t="inlineStr"`)** rather than a shared string table,
the xlsx extracts with the sheet name intact and **every cell value gone**.
The fixture therefore uses a `dimension` element and a shared string table.
## The K2 office fixture set (`k2-office/`)
`krav-presentasjon.pptx`, `krav-tekstdokument.odt` and
`krav-rikt-tekstformat.rtf` are the synthetic denominator for the three office
rows the corpus has none of. `docs/2026-09-04-k2-pptx-odt-rtf.md` measured that
denominator at **zero**`K2/trinn1` holds 43 files and not one is a `pptx`,
an `odt` or an `rtf` — so those rows were `unmeasured` in the sense of never
having met a document at all. Regenerated by `make_k2_office.py` in this
directory:
```
python3 tests/fixtures/make_k2_office.py
```
**One document, three containers.** All three carry the same authored content —
a title, an intro, a 20-row label/value table, a caption and a 4x4 grid — so
the only variable between the three measurements is the container and the
reader that opens it. The counts are hand-counted once, in
`k2-office-fasit.json`, and shared: **56 cells, 20 pairs, 59 distinct strings.**
**The generator and the fasit live one level up, and that is not tidiness.**
Door B walks its drop directory recursively, so anything parked inside
`k2-office/` would enter the run and N would stop being 3.
Same policy as the office fixtures above, for the same reason: every part is
hand-laid and no converter wrote any of them. The commissioning order offered
pandoc as a generator option; a file written by the converter and then read by
the converter would prove only that the converter agrees with itself.
Two things were measured while building this set, both against the vendored
pandoc 3.9, and both are the house shape — structurally plausible input,
silently wrong output, exit code 0 and no warning:
- **RTF cell paragraphs need `\pard\intbl`.** Without it, consecutive
`\trowd…\row` rows are read as each row NESTED inside the previous one:
five label/value rows came back as five levels of nested table, 2076
characters where 117 were expected.
- **The `\uN?` unicode escape loses the character after it.** Measured
directly: `A\u248?BC` reads back as `AoC` (ring letter present, `B` gone) and
`A\u248?xBC` reads back as `AoBC`. The `?` is taken as the control word's
delimiter and `\uc1` then skips a real character. The fixture writes
`\uN ?` with an explicit space, which round-trips. **This is the form Word
emits**, so it is a converter finding rather than a fixture quirk — recorded
in `docs/2026-09-07-k2-pptx-odt-rtf-fixtures.md`, not worked around anywhere
in `src/`.
**Three synthetic documents in one house style are not a corpus.** The rows stay
`unmeasured` in `extract._EVIDENCE` and `tests/test_k2_office_fixtures.py`
asserts that they do.
## The proposer's default-profile golden
`propose-golden-default.json` is the artifact `tools/okf_propose_segments.py`
produces for `OUTLINE_DOCUMENT` (defined in `tests/test_propose_segments.py`)
with **no flags at all**, generated at commit `798f64a` with
`--proposed-at 2026-09-03T00:00:00Z`. The timestamp is an explicit argument
because the artifact carries it verbatim; a wall-clock default would make the
golden unreproducible by construction.
**Why the fixture is `OUTLINE_DOCUMENT` and not `DOCUMENT`.** The golden exists
to go red if any later rule is accidentally defaulted ON. `DOCUMENT` was
measured to contain **zero** bare-integer lines, so a golden over it would stay
byte-identical through exactly the regression it was named to catch -- a trap
written down but unable to fire. `OUTLINE_DOCUMENT` carries a bare-integer
ascending run of three, which today's rules do not match (measured: bare `1` /
`1.` / `1)` yield 0 candidates), so the golden pins that absence and breaks the
moment it stops being true.
**It transitively pins `observed_extractor_version`**
(`src/llm_ingestion_okf/segmentation.py`): the field is written into every
artifact, so a converter or extractor bump turns this golden red. That red is
legitimate -- read the diff and decide, exactly as for the frozen PDF literal
below. Regenerate only after that decision, never to make a red go away.
## Why the expected office text is frozen as a literal
The same reason as the PDF text below, with one addition: the literals are
pinned to a **named converter version**. `_pandoc.py` refuses any binary but
the vendored 3.9, and `tests/test_extract.py` asserts that version beside the
literals. A frozen literal without a named converter pins nothing — it says
"these bytes" without saying what produced them.
## Why the expected PDF text is frozen as a literal
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
exact string. That is deliberate, and it is the mechanism behind a promise this
library makes everywhere else:
- Extraction is deterministic **within** a parser version. Measured 2026-08-21
across five configurations, two runs each, compared byte for byte
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`).
- Extraction is **not** guaranteed stable **across** parser versions.
`pdfplumber` pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships
date-stamped releases with no stability contract. So the real pin on extracted
text is a transitive one, and it is exact.
The consequence is worth stating plainly: **any golden fixture built on
extracted PDF text is pinned to an exact parser version, and a parser upgrade
is a fixture migration, not a routine bump.** The frozen literal is what makes
that upgrade break something visible instead of drifting silently. If it goes
red after a dependency change, the correct response is to read the diff and
decide, not to re-record the expectation.
The version range that carries this lives in `pyproject.toml`'s
`[project.optional-dependencies] extract`, with the same reasoning at the
declaration site.
## The content-accounting fixtures (`accounting/`)
The fasit side of `tools/okf_accounting_gate.py`. `accounting/corpus/` holds
one document per row of README's file-type table (13 of 13) plus a `graphics/`
directory next to them that the HTML, STS and markdown documents point at --
the layout under which a picture is carried through a document AND booked as a
rejected file. `accounting/rejected/` holds one HTML document with a
zero-width space in its prose, which the guard refuses at every tier, and the
image it points at.
`inventory.json` and `rejected-inventory.json` are what `tools/okf_witness.py`
counts in those two directories, committed as data and regenerated only with
that tool:
```
python3 tools/okf_witness.py tests/fixtures/accounting/corpus > tests/fixtures/accounting/inventory.json
python3 tools/okf_witness.py tests/fixtures/accounting/rejected > tests/fixtures/accounting/rejected-inventory.json
```
Seven more documents were added 2026-09-18, one per format that had element
types it could never exercise. An independent review measured **20 of 63
element types with a count of ZERO in their only fixture**, which is why six of
seven witness mutants survived the suite: a witness cannot be caught being
wrong about something it never sees. They are written part by part by
`make_accounting_fixtures.py` in this directory, for the same reason the XML
fixtures are hand-written:
```
python3 tests/fixtures/accounting/make_accounting_fixtures.py
```
| Fixture | What it carries that nothing else did |
|---|---|
| `topptekst-og-kommentar.docx` | A header, a footer, a comment, an endnote and a **text box** -- and a footnote, a table and a heading, three types the only other docx has at 0. The header says "Utkast - gjelder ikke etter 2026-01-01" and the comment says the requirement does NOT apply in tunnels: two statements that reverse the document's meaning and that the build carries none of. |
| `notater-og-skjult.pptx` | A **speaker note** and a **hidden slide** (`show="0"`), plus a table and paragraphs. A hidden slide counted as an ordinary one is indistinguishable from one that is shown. |
| `skjult-ark-og-formel.xlsx` | A **hidden sheet**, a **formula** (`<f>B2*2</f>`) and a **picture**. The picture is what makes the operator's `.xlsx image` exception exercisable at all: the old fixture had none. |
| `liste-og-bilde.odt` | A **header and footer** (they live in `styles.xml`, so a reader of `content.xml` cannot see them), an **annotation**, a list and a picture. |
| `bilde.rtf` | A `\pict` picture: the rtf witness's image count was 0 in its only fixture. |
| `figur.html` | A picture and a table under `.html`; `side.htm` gained one too, so `.htm` and `.html` each exercise `image`. |
| `sts-rikt.xml` | A **`mixed-citation`**, an **`mml:math`**, a **`fig` with a caption**, a table with a label, cells, a list item and a footnote -- six STS roles the R761 delivery does not contain at all, which is why the gate's only real corpus could not see the hole in the role map. |
### The hand counts
Row 1's fasit is the witness's own output, so a hand count is the only number
in this loop the witness did not produce. Four of thirteen documents had one;
**all twenty have one now**, in `HAND_COUNTS` in
`tests/test_accounting_gate.py`, and `test_the_hand_counts_cover_every_document_of_the_corpus`
fails if a document is added without one. Each was counted by reading the
fixture's own bytes -- the XML parts of a zip, the control words of the rtf,
the objects of the PDF -- never by running the witness and writing down what
it said.
`witness/prosess-84-sts.twin.json` is the STS document written by hand in the
publisher's JSON node form (`standardContent`, nodes with `e`/`t`/`x`), so the
two STS witnesses can be compared on a fixture as well as on R761. Eight of
the thirteen documents are byte copies of fixtures documented above
(`image-inbox/`, `k2-office/`, `prisark.xlsx`); the other five
(`notat.md`, `logg.txt`, `mengder.csv`, `parametre.json`, `side.htm`) are
written here, and `notat.md` carries a fenced `# ...` line that is not a
heading.
## What these fixtures do not cover
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
detected table objects are clean enough to hand to `render_table` unchanged;
two independent parsers return the same wrong shape, because the breakage is in
the documents' ruling geometry rather than in either library. PDFs enter this
library as **prose**, and structured tables are out of scope until that is
decided separately.
## `propose-golden-grid-default.json`
Pins the DEFAULT proposer artifact over a document containing a pandoc **grid
table**. Generated on unmodified code, before Arm E's rule existed, with
exactly the command the test runs:
```
write(tmp_path, GRID_GOLDEN_DOCUMENT, "grid.md")
okf_propose_segments.main([source, "--out", out, "--proposed-at", "2026-09-03T00:00:00Z"])
```
It exists because `propose-golden-default.json` **cannot** pin this. That
golden is taken over `OUTLINE_DOCUMENT`, which contains no `|` row and no `+`
rule line, so no table rule -- present or future -- can move its bytes. A guard
that is structurally incapable of firing is a trap written down but never
armed. This one is taken over a document that has a grid table, so an
accidentally default-on table rule turns it red.
Both goldens transitively pin `observed_extractor_version` and `PROPOSER_VERSION`.
A red here after a dependency change is a legitimate red: read the diff and
decide, do not re-record the expectation.