test(extract): hand-built office fixtures with frozen extracted text
Three hand-laid OOXML containers, every part written out by hand and zipped
with a fixed date_time so they are byte-reproducible. No converter output
anywhere in them: a .docx written by the converter and read by the converter
proves only that the converter agrees with itself, and would stay green through
any conversion defect that is symmetric -- which is most of them.
two-line-krav.docx heading + label/value on one line (the docx mirror of
the PDF fixture)
no-styles-krav.docx the SAME document without word/styles.xml
two-line-krav.xlsx sheet name as heading + label/value on one row
THE FIXTURES FOUND A REAL DEFECT IN THE SEAM THEY WERE MEANT TO PIN. The
converter call used pypandoc's TEXT entry point, which takes an `encoding`
because it treats its source as text -- and that corrupts a zip. The xlsx
fixture failed with `Failed to unpack XLSX archive: not enough bytes` while
reading correctly from disk with the same binary. The docx of the same shape
happened to survive, which is the part worth writing down: the defect is silent
for some inputs and fatal for others, so "it worked on the file I tried" was
never evidence. Input now goes through a temporary file.
Two measurements while building, both the same shape -- structurally valid
input, silently reduced output, exit code 0, no warning:
- Without word/styles.xml the docx extracts as flat prose with no heading. A
fixture lacking that part would pin the body and pin nothing about structure.
Committed as a negative control that RUNS rather than a sentence in a README.
- With inline strings rather than a shared string table, the xlsx extracts with
the sheet name intact and every cell value gone. The fixture uses a dimension
element and a shared string table instead.
The frozen literals are pinned to a NAMED converter version, asserted beside
them: a frozen literal without one says "these bytes" without saying what
produced them.
Suite 908 -> 913. Fixtures regenerate byte-identically.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
cd7b792aaf
commit
66a44f173b
7 changed files with 274 additions and 6 deletions
37
tests/fixtures/README.md
vendored
37
tests/fixtures/README.md
vendored
|
|
@ -18,6 +18,43 @@ and reproducible from that one file.
|
|||
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
|
||||
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
|
||||
|
||||
## The office fixtures
|
||||
|
||||
`two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are
|
||||
hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every
|
||||
part is written out by hand and zipped with a fixed `date_time`, so they are
|
||||
byte-reproducible and carry no converter's output.
|
||||
|
||||
**That last point is the whole policy, not a preference.** A `.docx` written by
|
||||
the converter and then read by the converter proves only that the converter
|
||||
agrees with itself, and would stay green through any conversion defect that is
|
||||
symmetric — which is most of them.
|
||||
|
||||
| Fixture | What it is for |
|
||||
|---|---|
|
||||
| `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. |
|
||||
| `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
|
||||
| `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. |
|
||||
|
||||
Two things were measured while building these, and both are the same shape —
|
||||
structurally valid input, silently reduced output, exit code 0 and no warning:
|
||||
|
||||
- **Without `word/styles.xml`** the docx extracts as flat prose with no heading.
|
||||
A fixture lacking that part would pin the body and pin nothing about
|
||||
structure, while looking exactly as convincing. Structure is the half the
|
||||
segment proposer reads.
|
||||
- **With inline strings (`t="inlineStr"`)** rather than a shared string table,
|
||||
the xlsx extracts with the sheet name intact and **every cell value gone**.
|
||||
The fixture therefore uses a `dimension` element and a shared string table.
|
||||
|
||||
## Why the expected office text is frozen as a literal
|
||||
|
||||
The same reason as the PDF text below, with one addition: the literals are
|
||||
pinned to a **named converter version**. `_pandoc.py` refuses any binary but
|
||||
the vendored 3.9, and `tests/test_extract.py` asserts that version beside the
|
||||
literals. A frozen literal without a named converter pins nothing — it says
|
||||
"these bytes" without saying what produced them.
|
||||
|
||||
## Why the expected PDF text is frozen as a literal
|
||||
|
||||
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue