test(extract): hand-built office fixtures with frozen extracted text
Three hand-laid OOXML containers, every part written out by hand and zipped
with a fixed date_time so they are byte-reproducible. No converter output
anywhere in them: a .docx written by the converter and read by the converter
proves only that the converter agrees with itself, and would stay green through
any conversion defect that is symmetric -- which is most of them.
two-line-krav.docx heading + label/value on one line (the docx mirror of
the PDF fixture)
no-styles-krav.docx the SAME document without word/styles.xml
two-line-krav.xlsx sheet name as heading + label/value on one row
THE FIXTURES FOUND A REAL DEFECT IN THE SEAM THEY WERE MEANT TO PIN. The
converter call used pypandoc's TEXT entry point, which takes an `encoding`
because it treats its source as text -- and that corrupts a zip. The xlsx
fixture failed with `Failed to unpack XLSX archive: not enough bytes` while
reading correctly from disk with the same binary. The docx of the same shape
happened to survive, which is the part worth writing down: the defect is silent
for some inputs and fatal for others, so "it worked on the file I tried" was
never evidence. Input now goes through a temporary file.
Two measurements while building, both the same shape -- structurally valid
input, silently reduced output, exit code 0, no warning:
- Without word/styles.xml the docx extracts as flat prose with no heading. A
fixture lacking that part would pin the body and pin nothing about structure.
Committed as a negative control that RUNS rather than a sentence in a README.
- With inline strings rather than a shared string table, the xlsx extracts with
the sheet name intact and every cell value gone. The fixture uses a dimension
element and a shared string table instead.
The frozen literals are pinned to a NAMED converter version, asserted beside
them: a frozen literal without one says "these bytes" without saying what
produced them.
Suite 908 -> 913. Fixtures regenerate byte-identically.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
cd7b792aaf
commit
66a44f173b
7 changed files with 274 additions and 6 deletions
|
|
@ -26,6 +26,7 @@ from __future__ import annotations
|
|||
import csv
|
||||
import functools
|
||||
import io
|
||||
import tempfile
|
||||
import warnings
|
||||
from collections.abc import Callable, Sequence
|
||||
from html.parser import HTMLParser
|
||||
|
|
@ -250,13 +251,30 @@ def _convert_bytes(source: bytes, to: str, format: str, extra_args: Sequence[str
|
|||
would otherwise need the binary present and a real office document, which
|
||||
would make the seam's own logic untestable on a machine without the extra.
|
||||
The conversion itself is covered by the frozen-text fixtures instead.
|
||||
|
||||
THE INPUT GOES THROUGH A FILE, NOT THROUGH THE TEXT ENTRY POINT. Every
|
||||
format here is a binary container, and the converter's text entry point
|
||||
takes an `encoding` because it treats its source as text -- which corrupts
|
||||
a zip. Measured: a hand-laid `.xlsx` that pandoc reads correctly from disk
|
||||
fails through the text path with `Failed to unpack XLSX archive: not enough
|
||||
bytes`. A `.docx` of the same shape happened to survive, which is what
|
||||
makes this worth writing down: the defect is SILENT for some inputs and
|
||||
fatal for others, so "it worked on the file I tried" is not evidence here.
|
||||
|
||||
The temporary directory is removed on every path, including the failure
|
||||
one, and nothing outside it is written.
|
||||
"""
|
||||
import pypandoc
|
||||
|
||||
from ._pandoc import converter_path
|
||||
|
||||
with converter_path():
|
||||
return str(pypandoc.convert_text(source, to, format=format, extra_args=list(extra_args)))
|
||||
with tempfile.TemporaryDirectory() as staging:
|
||||
staged = Path(staging) / f"input.{format}"
|
||||
staged.write_bytes(source)
|
||||
with converter_path():
|
||||
return str(
|
||||
pypandoc.convert_file(str(staged), to, format=format, extra_args=list(extra_args))
|
||||
)
|
||||
|
||||
|
||||
def _extract_office(suffix: str, data: bytes) -> str:
|
||||
|
|
|
|||
37
tests/fixtures/README.md
vendored
37
tests/fixtures/README.md
vendored
|
|
@ -18,6 +18,43 @@ and reproducible from that one file.
|
|||
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
|
||||
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
|
||||
|
||||
## The office fixtures
|
||||
|
||||
`two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are
|
||||
hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every
|
||||
part is written out by hand and zipped with a fixed `date_time`, so they are
|
||||
byte-reproducible and carry no converter's output.
|
||||
|
||||
**That last point is the whole policy, not a preference.** A `.docx` written by
|
||||
the converter and then read by the converter proves only that the converter
|
||||
agrees with itself, and would stay green through any conversion defect that is
|
||||
symmetric — which is most of them.
|
||||
|
||||
| Fixture | What it is for |
|
||||
|---|---|
|
||||
| `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. |
|
||||
| `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
|
||||
| `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. |
|
||||
|
||||
Two things were measured while building these, and both are the same shape —
|
||||
structurally valid input, silently reduced output, exit code 0 and no warning:
|
||||
|
||||
- **Without `word/styles.xml`** the docx extracts as flat prose with no heading.
|
||||
A fixture lacking that part would pin the body and pin nothing about
|
||||
structure, while looking exactly as convincing. Structure is the half the
|
||||
segment proposer reads.
|
||||
- **With inline strings (`t="inlineStr"`)** rather than a shared string table,
|
||||
the xlsx extracts with the sheet name intact and **every cell value gone**.
|
||||
The fixture therefore uses a `dimension` element and a shared string table.
|
||||
|
||||
## Why the expected office text is frozen as a literal
|
||||
|
||||
The same reason as the PDF text below, with one addition: the literals are
|
||||
pinned to a **named converter version**. `_pandoc.py` refuses any binary but
|
||||
the vendored 3.9, and `tests/test_extract.py` asserts that version beside the
|
||||
literals. A frozen literal without a named converter pins nothing — it says
|
||||
"these bytes" without saying what produced them.
|
||||
|
||||
## Why the expected PDF text is frozen as a literal
|
||||
|
||||
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
|
||||
|
|
|
|||
140
tests/fixtures/make_fixtures.py
vendored
140
tests/fixtures/make_fixtures.py
vendored
|
|
@ -1,14 +1,23 @@
|
|||
"""Regenerate the committed PDF fixtures for tests/test_extract.py.
|
||||
"""Regenerate the committed fixtures for tests/test_extract.py.
|
||||
|
||||
Hand-written minimal PDFs: objects laid out by hand, xref offsets computed
|
||||
from the emitted bytes. No generator library, so the fixtures are auditable
|
||||
byte for byte and reproducible from this file alone.
|
||||
Hand-written minimal documents: PDF objects laid out by hand with xref offsets
|
||||
computed from the emitted bytes, and OOXML containers assembled part by part.
|
||||
No generator library anywhere, so every fixture is auditable byte for byte and
|
||||
reproducible from this file alone.
|
||||
|
||||
THE POLICY IS WHAT FORBIDS THE SHORTCUT. A `.docx` written by the converter and
|
||||
then read by the converter proves only that the converter agrees with itself --
|
||||
it would stay green through any conversion defect that is symmetric, which is
|
||||
most of them. Hand-laying the parts is what makes the fixture an independent
|
||||
statement about the format rather than a recording of our own output.
|
||||
|
||||
Run from the repository root: python3 tests/fixtures/make_fixtures.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import io
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
HERE = Path(__file__).parent
|
||||
|
|
@ -55,6 +64,121 @@ def build_pdf(content: bytes) -> bytes:
|
|||
return bytes(out)
|
||||
|
||||
|
||||
# --- office containers -------------------------------------------------------
|
||||
#
|
||||
# A fixed timestamp on every member, because a zip records mtime and the whole
|
||||
# point is a byte-reproducible file: without it the fixture would differ on
|
||||
# every regeneration and `git diff --quiet` could never be the check.
|
||||
_ZIP_DATE = (2020, 1, 1, 0, 0, 0)
|
||||
|
||||
_XML = '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
|
||||
|
||||
# `word/styles.xml` IS REQUIRED, not decoration. Measured during planning: the
|
||||
# same document WITHOUT a styles part extracts as plain text with no heading
|
||||
# marker at all, so a fixture lacking it would pin the body and silently pin
|
||||
# nothing about structure -- which is the half the segment proposer reads.
|
||||
_DOCX_PARTS = {
|
||||
"[Content_Types].xml": _XML
|
||||
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
||||
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
||||
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
||||
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
|
||||
+ '<Override PartName="/word/styles.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.styles+xml"/>'
|
||||
+ "</Types>",
|
||||
"_rels/.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/>'
|
||||
+ "</Relationships>",
|
||||
"word/_rels/document.xml.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/styles" Target="styles.xml"/>'
|
||||
+ "</Relationships>",
|
||||
"word/styles.xml": _XML
|
||||
+ '<w:styles xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">'
|
||||
+ '<w:style w:type="paragraph" w:styleId="Heading1"><w:name w:val="heading 1"/></w:style>'
|
||||
+ "</w:styles>",
|
||||
"word/document.xml": _XML
|
||||
+ '<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"><w:body>'
|
||||
+ '<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr><w:r><w:t>Krav til helning</w:t></w:r></w:p>'
|
||||
+ "<w:p><w:r><w:t>60 og 70 1:15</w:t></w:r></w:p>"
|
||||
+ "</w:body></w:document>",
|
||||
}
|
||||
|
||||
# The same document MINUS the styles part. A negative control, committed rather
|
||||
# than described: it is what proves the styles part is load-bearing, and a
|
||||
# claim of that kind that nothing runs is a claim that decays.
|
||||
_DOCX_NO_STYLES_PARTS = {
|
||||
"[Content_Types].xml": _XML
|
||||
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
||||
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
||||
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
||||
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
|
||||
+ "</Types>",
|
||||
"_rels/.rels": _DOCX_PARTS["_rels/.rels"],
|
||||
"word/document.xml": _DOCX_PARTS["word/document.xml"],
|
||||
}
|
||||
|
||||
# A minimal SpreadsheetML workbook: one sheet, a heading row and a label/value
|
||||
# row, mirroring what the PDF fixture does for its format.
|
||||
#
|
||||
# THE SHARED STRING TABLE IS NOT A STYLE CHOICE. The first attempt used inline
|
||||
# strings (`t="inlineStr"`), which is valid SpreadsheetML and which the
|
||||
# converter reads as EMPTY CELLS -- the sheet name survived and every value
|
||||
# vanished, with exit code 0 and no warning. A `dimension` element and a shared
|
||||
# string table are what make the values arrive. This is the same class of
|
||||
# defect as the missing `styles.xml`: structurally valid input, silently
|
||||
# reduced output, nothing anywhere saying so.
|
||||
_XLSX_PARTS = {
|
||||
"[Content_Types].xml": _XML
|
||||
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
||||
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
||||
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
||||
+ '<Override PartName="/xl/workbook.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml"/>'
|
||||
+ '<Override PartName="/xl/worksheets/sheet1.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml"/>'
|
||||
+ '<Override PartName="/xl/sharedStrings.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sharedStrings+xml"/>'
|
||||
+ "</Types>",
|
||||
"_rels/.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="xl/workbook.xml"/>'
|
||||
+ "</Relationships>",
|
||||
"xl/_rels/workbook.xml.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet" Target="worksheets/sheet1.xml"/>'
|
||||
+ '<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/sharedStrings" Target="sharedStrings.xml"/>'
|
||||
+ "</Relationships>",
|
||||
"xl/workbook.xml": _XML
|
||||
+ '<workbook xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
||||
+ ' xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">'
|
||||
+ '<sheets><sheet name="Krav" sheetId="1" r:id="rId1"/></sheets></workbook>',
|
||||
"xl/sharedStrings.xml": _XML
|
||||
+ '<sst xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main" count="3" uniqueCount="3">'
|
||||
+ "<si><t>Krav til helning</t></si><si><t>60 og 70</t></si><si><t>1:15</t></si></sst>",
|
||||
"xl/worksheets/sheet1.xml": _XML
|
||||
+ '<worksheet xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main">'
|
||||
+ '<dimension ref="A1:B2"/><sheetData>'
|
||||
+ '<row r="1"><c r="A1" t="s"><v>0</v></c></row>'
|
||||
+ '<row r="2"><c r="A2" t="s"><v>1</v></c><c r="B2" t="s"><v>2</v></c></row>'
|
||||
+ "</sheetData></worksheet>",
|
||||
}
|
||||
|
||||
|
||||
def build_ooxml(parts: dict[str, str]) -> bytes:
|
||||
"""Zip the parts with a fixed timestamp and no compression variance.
|
||||
|
||||
A constant `date_time` on every member is what makes the output
|
||||
reproducible: a zip records mtime, so the default `ZipFile.writestr` would
|
||||
stamp the current time and the fixture would differ on every run --
|
||||
which would make `git diff --quiet` useless as the regeneration check.
|
||||
"""
|
||||
out = io.BytesIO()
|
||||
with zipfile.ZipFile(out, "w", compression=zipfile.ZIP_DEFLATED) as archive:
|
||||
for name, payload in parts.items():
|
||||
info = zipfile.ZipInfo(name, date_time=_ZIP_DATE)
|
||||
info.compress_type = zipfile.ZIP_DEFLATED
|
||||
archive.writestr(info, payload)
|
||||
return out.getvalue()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
for name, content in (
|
||||
("two-line-krav.pdf", KRAV_CONTENT),
|
||||
|
|
@ -62,3 +186,11 @@ if __name__ == "__main__":
|
|||
):
|
||||
(HERE / name).write_bytes(build_pdf(content))
|
||||
print(f"wrote {name}")
|
||||
|
||||
for name, parts in (
|
||||
("two-line-krav.docx", _DOCX_PARTS),
|
||||
("no-styles-krav.docx", _DOCX_NO_STYLES_PARTS),
|
||||
("two-line-krav.xlsx", _XLSX_PARTS),
|
||||
):
|
||||
(HERE / name).write_bytes(build_ooxml(parts))
|
||||
print(f"wrote {name}")
|
||||
|
|
|
|||
BIN
tests/fixtures/no-styles-krav.docx
vendored
Normal file
BIN
tests/fixtures/no-styles-krav.docx
vendored
Normal file
Binary file not shown.
BIN
tests/fixtures/two-line-krav.docx
vendored
Normal file
BIN
tests/fixtures/two-line-krav.docx
vendored
Normal file
Binary file not shown.
BIN
tests/fixtures/two-line-krav.xlsx
vendored
Normal file
BIN
tests/fixtures/two-line-krav.xlsx
vendored
Normal file
Binary file not shown.
|
|
@ -310,6 +310,87 @@ def test_office_conversion_warns_that_it_is_lossy(
|
|||
# a parser upgrade break something visible instead of drifting silently.
|
||||
KRAV_TEXT = "Krav til helning på utkilingen\n60 og 70 1:15"
|
||||
|
||||
# The office fixtures are frozen the same way, and against a NAMED converter
|
||||
# version -- a frozen literal means nothing without one, because the thing it
|
||||
# pins is "this converter, on this input, produces these bytes". `_pandoc.py`
|
||||
# refuses any other version, so the two pins hold each other up.
|
||||
requires_pandoc = pytest.mark.skipif(
|
||||
importlib.util.find_spec("pypandoc") is None,
|
||||
reason="the optional [extract] extra is not installed",
|
||||
)
|
||||
|
||||
# docx: a heading and one requirement row with label and value on the SAME
|
||||
# line, mirroring the property the PDF fixture pins.
|
||||
DOCX_TEXT = "# Krav til helning\n\n60 og 70 1:15"
|
||||
|
||||
# xlsx: the sheet name becomes a heading and the rows become a table. The
|
||||
# label/value pairing survives on one row, which is the property that matters.
|
||||
XLSX_TEXT = (
|
||||
"## Krav {#sheet-1}\n\n Krav til helning \n"
|
||||
" ------------------ ------\n 60 og 70 1:15"
|
||||
)
|
||||
|
||||
# The negative control, committed rather than described: the SAME document
|
||||
# without `word/styles.xml`. The body survives and the heading marker does not.
|
||||
DOCX_NO_STYLES_TEXT = "Krav til helning\n\n60 og 70 1:15"
|
||||
|
||||
|
||||
@requires_pandoc
|
||||
def test_docx_extracts_to_its_frozen_text() -> None:
|
||||
data = (FIXTURES / "two-line-krav.docx").read_bytes()
|
||||
with pytest.warns(ExtractionWarning):
|
||||
assert extract_text("krav.docx", data) == DOCX_TEXT
|
||||
|
||||
|
||||
@requires_pandoc
|
||||
def test_xlsx_extracts_to_its_frozen_text() -> None:
|
||||
data = (FIXTURES / "two-line-krav.xlsx").read_bytes()
|
||||
with pytest.warns(ExtractionWarning):
|
||||
assert extract_text("krav.xlsx", data) == XLSX_TEXT
|
||||
|
||||
|
||||
@requires_pandoc
|
||||
def test_a_docx_without_a_styles_part_loses_its_heading() -> None:
|
||||
"""The negative control for the fixture policy, run rather than asserted.
|
||||
|
||||
`word/styles.xml` is what makes the converter see a heading. Without it the
|
||||
same document extracts as flat prose -- so a fixture built WITHOUT that
|
||||
part would pin the body and pin nothing at all about structure, while
|
||||
looking exactly as convincing.
|
||||
|
||||
Structure is the half the segment proposer reads, which is why this is a
|
||||
committed fixture and not a sentence in a README.
|
||||
"""
|
||||
data = (FIXTURES / "no-styles-krav.docx").read_bytes()
|
||||
with pytest.warns(ExtractionWarning):
|
||||
text = extract_text("krav.docx", data)
|
||||
assert text == DOCX_NO_STYLES_TEXT
|
||||
assert not text.startswith("#"), "the heading marker must be absent"
|
||||
assert DOCX_TEXT.startswith("#"), "and present in the fixture that has styles"
|
||||
|
||||
|
||||
@requires_pandoc
|
||||
def test_the_frozen_office_text_is_pinned_to_a_named_converter_version() -> None:
|
||||
"""A frozen literal without a named version pins nothing.
|
||||
|
||||
If the converter version ever moves, these literals must be re-measured
|
||||
rather than trusted -- so the version is asserted right where they live.
|
||||
"""
|
||||
from llm_ingestion_okf._pandoc import PANDOC_VERSION, resolve_pandoc
|
||||
|
||||
assert PANDOC_VERSION == "3.9"
|
||||
assert resolve_pandoc().is_file()
|
||||
|
||||
|
||||
@requires_pandoc
|
||||
def test_office_extraction_is_byte_stable_across_calls() -> None:
|
||||
data = (FIXTURES / "two-line-krav.docx").read_bytes()
|
||||
with pytest.warns(ExtractionWarning):
|
||||
first = extract_text("krav.docx", data)
|
||||
with pytest.warns(ExtractionWarning):
|
||||
second = extract_text("krav.docx", data)
|
||||
assert first == second
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_pdf_extracts_text_with_label_and_value_on_one_line() -> None:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue