test(extract): hand-built office fixtures with frozen extracted text

Three hand-laid OOXML containers, every part written out by hand and zipped
with a fixed date_time so they are byte-reproducible. No converter output
anywhere in them: a .docx written by the converter and read by the converter
proves only that the converter agrees with itself, and would stay green through
any conversion defect that is symmetric -- which is most of them.

  two-line-krav.docx    heading + label/value on one line (the docx mirror of
                        the PDF fixture)
  no-styles-krav.docx   the SAME document without word/styles.xml
  two-line-krav.xlsx    sheet name as heading + label/value on one row

THE FIXTURES FOUND A REAL DEFECT IN THE SEAM THEY WERE MEANT TO PIN. The
converter call used pypandoc's TEXT entry point, which takes an `encoding`
because it treats its source as text -- and that corrupts a zip. The xlsx
fixture failed with `Failed to unpack XLSX archive: not enough bytes` while
reading correctly from disk with the same binary. The docx of the same shape
happened to survive, which is the part worth writing down: the defect is silent
for some inputs and fatal for others, so "it worked on the file I tried" was
never evidence. Input now goes through a temporary file.

Two measurements while building, both the same shape -- structurally valid
input, silently reduced output, exit code 0, no warning:

- Without word/styles.xml the docx extracts as flat prose with no heading. A
  fixture lacking that part would pin the body and pin nothing about structure.
  Committed as a negative control that RUNS rather than a sentence in a README.
- With inline strings rather than a shared string table, the xlsx extracts with
  the sheet name intact and every cell value gone. The fixture uses a dimension
  element and a shared string table instead.

The frozen literals are pinned to a NAMED converter version, asserted beside
them: a frozen literal without one says "these bytes" without saying what
produced them.

Suite 908 -> 913. Fixtures regenerate byte-identically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-02 14:14:27 +02:00
commit 66a44f173b
7 changed files with 274 additions and 6 deletions

View file

@ -18,6 +18,43 @@ and reproducible from that one file.
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
## The office fixtures
`two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are
hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every
part is written out by hand and zipped with a fixed `date_time`, so they are
byte-reproducible and carry no converter's output.
**That last point is the whole policy, not a preference.** A `.docx` written by
the converter and then read by the converter proves only that the converter
agrees with itself, and would stay green through any conversion defect that is
symmetric — which is most of them.
| Fixture | What it is for |
|---|---|
| `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. |
| `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
| `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. |
Two things were measured while building these, and both are the same shape —
structurally valid input, silently reduced output, exit code 0 and no warning:
- **Without `word/styles.xml`** the docx extracts as flat prose with no heading.
A fixture lacking that part would pin the body and pin nothing about
structure, while looking exactly as convincing. Structure is the half the
segment proposer reads.
- **With inline strings (`t="inlineStr"`)** rather than a shared string table,
the xlsx extracts with the sheet name intact and **every cell value gone**.
The fixture therefore uses a `dimension` element and a shared string table.
## Why the expected office text is frozen as a literal
The same reason as the PDF text below, with one addition: the literals are
pinned to a **named converter version**. `_pandoc.py` refuses any binary but
the vendored 3.9, and `tests/test_extract.py` asserts that version beside the
literals. A frozen literal without a named converter pins nothing — it says
"these bytes" without saying what produced them.
## Why the expected PDF text is frozen as a literal
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an

View file

@ -1,14 +1,23 @@
"""Regenerate the committed PDF fixtures for tests/test_extract.py.
"""Regenerate the committed fixtures for tests/test_extract.py.
Hand-written minimal PDFs: objects laid out by hand, xref offsets computed
from the emitted bytes. No generator library, so the fixtures are auditable
byte for byte and reproducible from this file alone.
Hand-written minimal documents: PDF objects laid out by hand with xref offsets
computed from the emitted bytes, and OOXML containers assembled part by part.
No generator library anywhere, so every fixture is auditable byte for byte and
reproducible from this file alone.
THE POLICY IS WHAT FORBIDS THE SHORTCUT. A `.docx` written by the converter and
then read by the converter proves only that the converter agrees with itself --
it would stay green through any conversion defect that is symmetric, which is
most of them. Hand-laying the parts is what makes the fixture an independent
statement about the format rather than a recording of our own output.
Run from the repository root: python3 tests/fixtures/make_fixtures.py
"""
from __future__ import annotations
import io
import zipfile
from pathlib import Path
HERE = Path(__file__).parent
@ -55,6 +64,121 @@ def build_pdf(content: bytes) -> bytes:
return bytes(out)
# --- office containers -------------------------------------------------------
#
# A fixed timestamp on every member, because a zip records mtime and the whole
# point is a byte-reproducible file: without it the fixture would differ on
# every regeneration and `git diff --quiet` could never be the check.
_ZIP_DATE = (2020, 1, 1, 0, 0, 0)
_XML = '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
# `word/styles.xml` IS REQUIRED, not decoration. Measured during planning: the
# same document WITHOUT a styles part extracts as plain text with no heading
# marker at all, so a fixture lacking it would pin the body and silently pin
# nothing about structure -- which is the half the segment proposer reads.
_DOCX_PARTS = {
"[Content_Types].xml": _XML
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
+ '<Default Extension="xml" ContentType="application/xml"/>'
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
+ '<Override PartName="/word/styles.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.styles+xml"/>'
+ "</Types>",
"_rels/.rels": _XML
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/>'
+ "</Relationships>",
"word/_rels/document.xml.rels": _XML
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/styles" Target="styles.xml"/>'
+ "</Relationships>",
"word/styles.xml": _XML
+ '<w:styles xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">'
+ '<w:style w:type="paragraph" w:styleId="Heading1"><w:name w:val="heading 1"/></w:style>'
+ "</w:styles>",
"word/document.xml": _XML
+ '<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"><w:body>'
+ '<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr><w:r><w:t>Krav til helning</w:t></w:r></w:p>'
+ "<w:p><w:r><w:t>60 og 70 1:15</w:t></w:r></w:p>"
+ "</w:body></w:document>",
}
# The same document MINUS the styles part. A negative control, committed rather
# than described: it is what proves the styles part is load-bearing, and a
# claim of that kind that nothing runs is a claim that decays.
_DOCX_NO_STYLES_PARTS = {
"[Content_Types].xml": _XML
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
+ '<Default Extension="xml" ContentType="application/xml"/>'
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
+ "</Types>",
"_rels/.rels": _DOCX_PARTS["_rels/.rels"],
"word/document.xml": _DOCX_PARTS["word/document.xml"],
}
# A minimal SpreadsheetML workbook: one sheet, a heading row and a label/value
# row, mirroring what the PDF fixture does for its format.
#
# THE SHARED STRING TABLE IS NOT A STYLE CHOICE. The first attempt used inline
# strings (`t="inlineStr"`), which is valid SpreadsheetML and which the
# converter reads as EMPTY CELLS -- the sheet name survived and every value
# vanished, with exit code 0 and no warning. A `dimension` element and a shared
# string table are what make the values arrive. This is the same class of
# defect as the missing `styles.xml`: structurally valid input, silently
# reduced output, nothing anywhere saying so.
_XLSX_PARTS = {
"[Content_Types].xml": _XML
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
+ '<Default Extension="xml" ContentType="application/xml"/>'
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
+ '<Override PartName="/xl/workbook.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml"/>'
+ '<Override PartName="/xl/worksheets/sheet1.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml"/>'
+ '<Override PartName="/xl/sharedStrings.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sharedStrings+xml"/>'
+ "</Types>",
"_rels/.rels": _XML
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="xl/workbook.xml"/>'
+ "</Relationships>",
"xl/_rels/workbook.xml.rels": _XML
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet" Target="worksheets/sheet1.xml"/>'
+ '<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/sharedStrings" Target="sharedStrings.xml"/>'
+ "</Relationships>",
"xl/workbook.xml": _XML
+ '<workbook xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
+ ' xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">'
+ '<sheets><sheet name="Krav" sheetId="1" r:id="rId1"/></sheets></workbook>',
"xl/sharedStrings.xml": _XML
+ '<sst xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main" count="3" uniqueCount="3">'
+ "<si><t>Krav til helning</t></si><si><t>60 og 70</t></si><si><t>1:15</t></si></sst>",
"xl/worksheets/sheet1.xml": _XML
+ '<worksheet xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main">'
+ '<dimension ref="A1:B2"/><sheetData>'
+ '<row r="1"><c r="A1" t="s"><v>0</v></c></row>'
+ '<row r="2"><c r="A2" t="s"><v>1</v></c><c r="B2" t="s"><v>2</v></c></row>'
+ "</sheetData></worksheet>",
}
def build_ooxml(parts: dict[str, str]) -> bytes:
"""Zip the parts with a fixed timestamp and no compression variance.
A constant `date_time` on every member is what makes the output
reproducible: a zip records mtime, so the default `ZipFile.writestr` would
stamp the current time and the fixture would differ on every run --
which would make `git diff --quiet` useless as the regeneration check.
"""
out = io.BytesIO()
with zipfile.ZipFile(out, "w", compression=zipfile.ZIP_DEFLATED) as archive:
for name, payload in parts.items():
info = zipfile.ZipInfo(name, date_time=_ZIP_DATE)
info.compress_type = zipfile.ZIP_DEFLATED
archive.writestr(info, payload)
return out.getvalue()
if __name__ == "__main__":
for name, content in (
("two-line-krav.pdf", KRAV_CONTENT),
@ -62,3 +186,11 @@ if __name__ == "__main__":
):
(HERE / name).write_bytes(build_pdf(content))
print(f"wrote {name}")
for name, parts in (
("two-line-krav.docx", _DOCX_PARTS),
("no-styles-krav.docx", _DOCX_NO_STYLES_PARTS),
("two-line-krav.xlsx", _XLSX_PARTS),
):
(HERE / name).write_bytes(build_ooxml(parts))
print(f"wrote {name}")

BIN
tests/fixtures/no-styles-krav.docx vendored Normal file

Binary file not shown.

BIN
tests/fixtures/two-line-krav.docx vendored Normal file

Binary file not shown.

BIN
tests/fixtures/two-line-krav.xlsx vendored Normal file

Binary file not shown.