Three hand-laid OOXML containers, every part written out by hand and zipped
with a fixed date_time so they are byte-reproducible. No converter output
anywhere in them: a .docx written by the converter and read by the converter
proves only that the converter agrees with itself, and would stay green through
any conversion defect that is symmetric -- which is most of them.
two-line-krav.docx heading + label/value on one line (the docx mirror of
the PDF fixture)
no-styles-krav.docx the SAME document without word/styles.xml
two-line-krav.xlsx sheet name as heading + label/value on one row
THE FIXTURES FOUND A REAL DEFECT IN THE SEAM THEY WERE MEANT TO PIN. The
converter call used pypandoc's TEXT entry point, which takes an `encoding`
because it treats its source as text -- and that corrupts a zip. The xlsx
fixture failed with `Failed to unpack XLSX archive: not enough bytes` while
reading correctly from disk with the same binary. The docx of the same shape
happened to survive, which is the part worth writing down: the defect is silent
for some inputs and fatal for others, so "it worked on the file I tried" was
never evidence. Input now goes through a temporary file.
Two measurements while building, both the same shape -- structurally valid
input, silently reduced output, exit code 0, no warning:
- Without word/styles.xml the docx extracts as flat prose with no heading. A
fixture lacking that part would pin the body and pin nothing about structure.
Committed as a negative control that RUNS rather than a sentence in a README.
- With inline strings rather than a shared string table, the xlsx extracts with
the sheet name intact and every cell value gone. The fixture uses a dimension
element and a shared string table instead.
The frozen literals are pinned to a NAMED converter version, asserted beside
them: a frozen literal without one says "these bytes" without saying what
produced them.
Suite 908 -> 913. Fixtures regenerate byte-identically.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
196 lines
10 KiB
Python
196 lines
10 KiB
Python
"""Regenerate the committed fixtures for tests/test_extract.py.
|
|
|
|
Hand-written minimal documents: PDF objects laid out by hand with xref offsets
|
|
computed from the emitted bytes, and OOXML containers assembled part by part.
|
|
No generator library anywhere, so every fixture is auditable byte for byte and
|
|
reproducible from this file alone.
|
|
|
|
THE POLICY IS WHAT FORBIDS THE SHORTCUT. A `.docx` written by the converter and
|
|
then read by the converter proves only that the converter agrees with itself --
|
|
it would stay green through any conversion defect that is symmetric, which is
|
|
most of them. Hand-laying the parts is what makes the fixture an independent
|
|
statement about the format rather than a recording of our own output.
|
|
|
|
Run from the repository root: python3 tests/fixtures/make_fixtures.py
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import io
|
|
import zipfile
|
|
from pathlib import Path
|
|
|
|
HERE = Path(__file__).parent
|
|
|
|
# Two text lines: a heading, and one requirement row with label and value on
|
|
# the SAME line. That pairing is the property the parser choice was made on
|
|
# (see docs/2026-08-21-g2-pdf-extraction-measurement.md), so the fixture
|
|
# fails visibly if a parser upgrade ever breaks it. Byte 0xE5 is the Norwegian
|
|
# 'a-ring' in WinAnsiEncoding, which the font object below declares.
|
|
KRAV_CONTENT = (
|
|
b"BT /F1 12 Tf 20 160 Td (Krav til helning p\xe5 utkilingen) Tj ET\n"
|
|
b"BT /F1 12 Tf 20 140 Td (60 og 70 1:15) Tj ET\n"
|
|
)
|
|
|
|
# A structurally valid page carrying no text operators at all -- the shape a
|
|
# scanned or image-only PDF presents to a text extractor.
|
|
NO_TEXT_CONTENT = b"20 20 160 160 re S\n"
|
|
|
|
|
|
def build_pdf(content: bytes) -> bytes:
|
|
"""Assemble a one-page PDF around `content` as the page content stream."""
|
|
objects = [
|
|
b"<< /Type /Catalog /Pages 2 0 R >>",
|
|
b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
|
|
b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] "
|
|
b"/Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>",
|
|
b"<< /Length " + str(len(content)).encode() + b" >>\nstream\n" + content + b"endstream",
|
|
b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>",
|
|
]
|
|
|
|
out = bytearray(b"%PDF-1.4\n")
|
|
offsets = []
|
|
for number, body in enumerate(objects, start=1):
|
|
offsets.append(len(out))
|
|
out += str(number).encode() + b" 0 obj\n" + body + b"\nendobj\n"
|
|
|
|
xref_at = len(out)
|
|
size = str(len(objects) + 1).encode()
|
|
out += b"xref\n0 " + size + b"\n0000000000 65535 f \n"
|
|
for offset in offsets:
|
|
out += ("%010d 00000 n \n" % offset).encode()
|
|
out += b"trailer\n<< /Size " + size + b" /Root 1 0 R >>\n"
|
|
out += b"startxref\n" + str(xref_at).encode() + b"\n%%EOF\n"
|
|
return bytes(out)
|
|
|
|
|
|
# --- office containers -------------------------------------------------------
|
|
#
|
|
# A fixed timestamp on every member, because a zip records mtime and the whole
|
|
# point is a byte-reproducible file: without it the fixture would differ on
|
|
# every regeneration and `git diff --quiet` could never be the check.
|
|
_ZIP_DATE = (2020, 1, 1, 0, 0, 0)
|
|
|
|
_XML = '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
|
|
|
|
# `word/styles.xml` IS REQUIRED, not decoration. Measured during planning: the
|
|
# same document WITHOUT a styles part extracts as plain text with no heading
|
|
# marker at all, so a fixture lacking it would pin the body and silently pin
|
|
# nothing about structure -- which is the half the segment proposer reads.
|
|
_DOCX_PARTS = {
|
|
"[Content_Types].xml": _XML
|
|
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
|
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
|
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
|
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
|
|
+ '<Override PartName="/word/styles.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.styles+xml"/>'
|
|
+ "</Types>",
|
|
"_rels/.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/>'
|
|
+ "</Relationships>",
|
|
"word/_rels/document.xml.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/styles" Target="styles.xml"/>'
|
|
+ "</Relationships>",
|
|
"word/styles.xml": _XML
|
|
+ '<w:styles xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">'
|
|
+ '<w:style w:type="paragraph" w:styleId="Heading1"><w:name w:val="heading 1"/></w:style>'
|
|
+ "</w:styles>",
|
|
"word/document.xml": _XML
|
|
+ '<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"><w:body>'
|
|
+ '<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr><w:r><w:t>Krav til helning</w:t></w:r></w:p>'
|
|
+ "<w:p><w:r><w:t>60 og 70 1:15</w:t></w:r></w:p>"
|
|
+ "</w:body></w:document>",
|
|
}
|
|
|
|
# The same document MINUS the styles part. A negative control, committed rather
|
|
# than described: it is what proves the styles part is load-bearing, and a
|
|
# claim of that kind that nothing runs is a claim that decays.
|
|
_DOCX_NO_STYLES_PARTS = {
|
|
"[Content_Types].xml": _XML
|
|
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
|
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
|
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
|
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
|
|
+ "</Types>",
|
|
"_rels/.rels": _DOCX_PARTS["_rels/.rels"],
|
|
"word/document.xml": _DOCX_PARTS["word/document.xml"],
|
|
}
|
|
|
|
# A minimal SpreadsheetML workbook: one sheet, a heading row and a label/value
|
|
# row, mirroring what the PDF fixture does for its format.
|
|
#
|
|
# THE SHARED STRING TABLE IS NOT A STYLE CHOICE. The first attempt used inline
|
|
# strings (`t="inlineStr"`), which is valid SpreadsheetML and which the
|
|
# converter reads as EMPTY CELLS -- the sheet name survived and every value
|
|
# vanished, with exit code 0 and no warning. A `dimension` element and a shared
|
|
# string table are what make the values arrive. This is the same class of
|
|
# defect as the missing `styles.xml`: structurally valid input, silently
|
|
# reduced output, nothing anywhere saying so.
|
|
_XLSX_PARTS = {
|
|
"[Content_Types].xml": _XML
|
|
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
|
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
|
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
|
+ '<Override PartName="/xl/workbook.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml"/>'
|
|
+ '<Override PartName="/xl/worksheets/sheet1.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml"/>'
|
|
+ '<Override PartName="/xl/sharedStrings.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sharedStrings+xml"/>'
|
|
+ "</Types>",
|
|
"_rels/.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="xl/workbook.xml"/>'
|
|
+ "</Relationships>",
|
|
"xl/_rels/workbook.xml.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet" Target="worksheets/sheet1.xml"/>'
|
|
+ '<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/sharedStrings" Target="sharedStrings.xml"/>'
|
|
+ "</Relationships>",
|
|
"xl/workbook.xml": _XML
|
|
+ '<workbook xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
|
+ ' xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">'
|
|
+ '<sheets><sheet name="Krav" sheetId="1" r:id="rId1"/></sheets></workbook>',
|
|
"xl/sharedStrings.xml": _XML
|
|
+ '<sst xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main" count="3" uniqueCount="3">'
|
|
+ "<si><t>Krav til helning</t></si><si><t>60 og 70</t></si><si><t>1:15</t></si></sst>",
|
|
"xl/worksheets/sheet1.xml": _XML
|
|
+ '<worksheet xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main">'
|
|
+ '<dimension ref="A1:B2"/><sheetData>'
|
|
+ '<row r="1"><c r="A1" t="s"><v>0</v></c></row>'
|
|
+ '<row r="2"><c r="A2" t="s"><v>1</v></c><c r="B2" t="s"><v>2</v></c></row>'
|
|
+ "</sheetData></worksheet>",
|
|
}
|
|
|
|
|
|
def build_ooxml(parts: dict[str, str]) -> bytes:
|
|
"""Zip the parts with a fixed timestamp and no compression variance.
|
|
|
|
A constant `date_time` on every member is what makes the output
|
|
reproducible: a zip records mtime, so the default `ZipFile.writestr` would
|
|
stamp the current time and the fixture would differ on every run --
|
|
which would make `git diff --quiet` useless as the regeneration check.
|
|
"""
|
|
out = io.BytesIO()
|
|
with zipfile.ZipFile(out, "w", compression=zipfile.ZIP_DEFLATED) as archive:
|
|
for name, payload in parts.items():
|
|
info = zipfile.ZipInfo(name, date_time=_ZIP_DATE)
|
|
info.compress_type = zipfile.ZIP_DEFLATED
|
|
archive.writestr(info, payload)
|
|
return out.getvalue()
|
|
|
|
|
|
if __name__ == "__main__":
|
|
for name, content in (
|
|
("two-line-krav.pdf", KRAV_CONTENT),
|
|
("no-text-layer.pdf", NO_TEXT_CONTENT),
|
|
):
|
|
(HERE / name).write_bytes(build_pdf(content))
|
|
print(f"wrote {name}")
|
|
|
|
for name, parts in (
|
|
("two-line-krav.docx", _DOCX_PARTS),
|
|
("no-styles-krav.docx", _DOCX_NO_STYLES_PARTS),
|
|
("two-line-krav.xlsx", _XLSX_PARTS),
|
|
):
|
|
(HERE / name).write_bytes(build_ooxml(parts))
|
|
print(f"wrote {name}")
|