A concept named its source file by basename and, when segmented, carried a
`source_offset` into the text THIS LIBRARY extracted. Following that pointer
needed the corpus directory, the extractor and its exact transitive version --
none of which the bundle carries. Hand-walked on a real K2 concept: six steps,
four of them requiring knowledge from outside the bundle, to learn that a
requirement sits on pages 12-13 of a 20-page document.
The address is spec's: `sources: [{ resource, title }]`, where `resource` is
the dropped file's inbox-relative path (SPEC v0.2 5.1:303-306 -- "an absolute
URL, a bundle-relative path, or a path into a `references/` subdirectory").
The locator is ours, and it has to be: 5.1 has no field for a place within a
resource, and the pinned guard (1.3.0) rejects every route to putting one
inside a `sources` entry -- a non-allowlisted key by name, a nested flow list
as "scalar leaves only", and quoting as an unsupported form. So the locator is
top-level keys shaped like `source_offset`, and a path carrying a flow
terminator is refused fail-fast rather than mangled.
The unit table is built AT EXTRACTION, where the extracted text and the
original's structure are known to agree: pdf -> `source_pages` from
pdfplumber's own page numbers (a page that yielded no text does not renumber
the ones after it), xlsx -> `source_sheet` + `source_rows`, everything else ->
`source_lines`. `source_offset` stays.
Two measurements changed the design before it shipped. A `paragraphs` key for
docx would name a number the document does not have: `<w:p>` counts of
108/27/65/176/57 against converted-markdown lines of 75/33/67/144/63, not one
pair agreeing -- so the key is `source_lines` and says what it indexes. And an
empty spreadsheet row renders exactly like a table separator: the content-based
rule ate 8 empty rows on the K2 price sheet and reported its last row as 92
against a workbook that says 100. The separator is now found by position, and
`tomrad.xlsx` keeps that red.
One profile moves. `provenance` is a policy object, `None` everywhere but
`SEGMENTED_OKF_V0_2`; the other five shipped profiles are byte-identical.
K2 rebuilt from a frozen src copy: 629 concepts, 1108 files, name set identical,
0 ids moved, 479 files byte-identical, 629 changed and 0 lines removed anywhere.
629/629 now carry an address and a locator. New ref
`sha256-tree:665563a2f74423fcbcc8e4f0b0954ee73b73985ac0418de4f6987bd162a1f7c8`;
`2f82fcfe...` is stale. The pre-pass payload does not grow by one byte
(209 092 B before and after, 18 changed lines: the ref and eight per-concept
digests) -- because an excerpt carries the body, not the frontmatter, which is
also why the consumer still cannot cite "file X page 12" from a payload alone.
Report: docs/2026-09-08-proveniens-k2.md. 1339 tests, ruff and mypy clean.
Co-Authored-By: Claude <claude-opus-5>
376 lines
18 KiB
Python
376 lines
18 KiB
Python
"""Regenerate the committed fixtures for tests/test_extract.py.
|
|
|
|
Hand-written minimal documents: PDF objects laid out by hand with xref offsets
|
|
computed from the emitted bytes, and OOXML containers assembled part by part.
|
|
No generator library anywhere, so every fixture is auditable byte for byte and
|
|
reproducible from this file alone.
|
|
|
|
THE POLICY IS WHAT FORBIDS THE SHORTCUT. A `.docx` written by the converter and
|
|
then read by the converter proves only that the converter agrees with itself --
|
|
it would stay green through any conversion defect that is symmetric, which is
|
|
most of them. Hand-laying the parts is what makes the fixture an independent
|
|
statement about the format rather than a recording of our own output.
|
|
|
|
Run from the repository root: python3 tests/fixtures/make_fixtures.py
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import io
|
|
import zipfile
|
|
from pathlib import Path
|
|
|
|
HERE = Path(__file__).parent
|
|
|
|
# Two text lines: a heading, and one requirement row with label and value on
|
|
# the SAME line. That pairing is the property the parser choice was made on
|
|
# (see docs/2026-08-21-g2-pdf-extraction-measurement.md), so the fixture
|
|
# fails visibly if a parser upgrade ever breaks it. Byte 0xE5 is the Norwegian
|
|
# 'a-ring' in WinAnsiEncoding, which the font object below declares.
|
|
KRAV_CONTENT = (
|
|
b"BT /F1 12 Tf 20 160 Td (Krav til helning p\xe5 utkilingen) Tj ET\n"
|
|
b"BT /F1 12 Tf 20 140 Td (60 og 70 1:15) Tj ET\n"
|
|
)
|
|
|
|
# A structurally valid page carrying no text operators at all -- the shape a
|
|
# scanned or image-only PDF presents to a text extractor.
|
|
NO_TEXT_CONTENT = b"20 20 160 160 re S\n"
|
|
|
|
|
|
# One line of text per page, so the page a stretch of extracted text came from
|
|
# is decidable by reading the text alone. THREE pages rather than two, and the
|
|
# MIDDLE one carries no text operators: a two-page fixture cannot tell a page
|
|
# INDEX from a page NUMBER, and without a blank page in the middle it cannot
|
|
# tell either of them from a count of the pages that produced text. The
|
|
# extractor drops empty pages, so the third page's text belongs to page 3 and
|
|
# to no other number.
|
|
PAGED_CONTENTS = (
|
|
b"BT /F1 12 Tf 20 160 Td (Side en om helning) Tj ET\n",
|
|
b"20 20 160 160 re S\n",
|
|
b"BT /F1 12 Tf 20 160 Td (Side tre om utkiling) Tj ET\n",
|
|
)
|
|
|
|
|
|
def build_pdf(content: bytes) -> bytes:
|
|
"""Assemble a one-page PDF around `content` as the page content stream."""
|
|
return build_paged_pdf((content,))
|
|
|
|
|
|
def build_paged_pdf(contents: tuple[bytes, ...]) -> bytes:
|
|
"""Assemble a PDF with one page per entry of `contents`.
|
|
|
|
The object numbering is laid out first and the xref offsets computed from
|
|
the emitted bytes, exactly as the single-page form did -- the fixture stays
|
|
a hand-written statement about the format rather than a library's output.
|
|
"""
|
|
count = len(contents)
|
|
# 1 catalog, 2 pages, then one page object and one content stream per page,
|
|
# and the shared font last.
|
|
page_numbers = [3 + 2 * index for index in range(count)]
|
|
font_number = 3 + 2 * count
|
|
kids = b" ".join(str(number).encode() + b" 0 R" for number in page_numbers)
|
|
objects = [
|
|
b"<< /Type /Catalog /Pages 2 0 R >>",
|
|
b"<< /Type /Pages /Kids [" + kids + b"] /Count " + str(count).encode() + b" >>",
|
|
]
|
|
for index, content in enumerate(contents):
|
|
stream_number = page_numbers[index] + 1
|
|
objects.append(
|
|
b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Contents "
|
|
+ str(stream_number).encode()
|
|
+ b" 0 R /Resources << /Font << /F1 "
|
|
+ str(font_number).encode()
|
|
+ b" 0 R >> >> >>"
|
|
)
|
|
objects.append(
|
|
b"<< /Length " + str(len(content)).encode() + b" >>\nstream\n" + content + b"endstream"
|
|
)
|
|
objects.append(
|
|
b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>"
|
|
)
|
|
|
|
out = bytearray(b"%PDF-1.4\n")
|
|
offsets = []
|
|
for number, body in enumerate(objects, start=1):
|
|
offsets.append(len(out))
|
|
out += str(number).encode() + b" 0 obj\n" + body + b"\nendobj\n"
|
|
|
|
xref_at = len(out)
|
|
size = str(len(objects) + 1).encode()
|
|
out += b"xref\n0 " + size + b"\n0000000000 65535 f \n"
|
|
for offset in offsets:
|
|
out += ("%010d 00000 n \n" % offset).encode()
|
|
out += b"trailer\n<< /Size " + size + b" /Root 1 0 R >>\n"
|
|
out += b"startxref\n" + str(xref_at).encode() + b"\n%%EOF\n"
|
|
return bytes(out)
|
|
|
|
|
|
# --- office containers -------------------------------------------------------
|
|
#
|
|
# A fixed timestamp on every member, because a zip records mtime and the whole
|
|
# point is a byte-reproducible file: without it the fixture would differ on
|
|
# every regeneration and `git diff --quiet` could never be the check.
|
|
_ZIP_DATE = (2020, 1, 1, 0, 0, 0)
|
|
|
|
_XML = '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
|
|
|
|
# `word/styles.xml` IS REQUIRED, not decoration. Measured during planning: the
|
|
# same document WITHOUT a styles part extracts as plain text with no heading
|
|
# marker at all, so a fixture lacking it would pin the body and silently pin
|
|
# nothing about structure -- which is the half the segment proposer reads.
|
|
_DOCX_PARTS = {
|
|
"[Content_Types].xml": _XML
|
|
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
|
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
|
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
|
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
|
|
+ '<Override PartName="/word/styles.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.styles+xml"/>'
|
|
+ "</Types>",
|
|
"_rels/.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/>'
|
|
+ "</Relationships>",
|
|
"word/_rels/document.xml.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/styles" Target="styles.xml"/>'
|
|
+ "</Relationships>",
|
|
"word/styles.xml": _XML
|
|
+ '<w:styles xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">'
|
|
+ '<w:style w:type="paragraph" w:styleId="Heading1"><w:name w:val="heading 1"/></w:style>'
|
|
+ "</w:styles>",
|
|
"word/document.xml": _XML
|
|
+ '<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"><w:body>'
|
|
+ '<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr><w:r><w:t>Krav til helning</w:t></w:r></w:p>'
|
|
+ "<w:p><w:r><w:t>60 og 70 1:15</w:t></w:r></w:p>"
|
|
+ "</w:body></w:document>",
|
|
}
|
|
|
|
# The same document MINUS the styles part. A negative control, committed rather
|
|
# than described: it is what proves the styles part is load-bearing, and a
|
|
# claim of that kind that nothing runs is a claim that decays.
|
|
_DOCX_NO_STYLES_PARTS = {
|
|
"[Content_Types].xml": _XML
|
|
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
|
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
|
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
|
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
|
|
+ "</Types>",
|
|
"_rels/.rels": _DOCX_PARTS["_rels/.rels"],
|
|
"word/document.xml": _DOCX_PARTS["word/document.xml"],
|
|
}
|
|
|
|
# A minimal SpreadsheetML workbook: one sheet, a heading row and a label/value
|
|
# row, mirroring what the PDF fixture does for its format.
|
|
#
|
|
# THE SHARED STRING TABLE IS NOT A STYLE CHOICE. The first attempt used inline
|
|
# strings (`t="inlineStr"`), which is valid SpreadsheetML and which the
|
|
# converter reads as EMPTY CELLS -- the sheet name survived and every value
|
|
# vanished, with exit code 0 and no warning. A `dimension` element and a shared
|
|
# string table are what make the values arrive. This is the same class of
|
|
# defect as the missing `styles.xml`: structurally valid input, silently
|
|
# reduced output, nothing anywhere saying so.
|
|
_XLSX_PARTS = {
|
|
"[Content_Types].xml": _XML
|
|
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
|
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
|
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
|
+ '<Override PartName="/xl/workbook.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml"/>'
|
|
+ '<Override PartName="/xl/worksheets/sheet1.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml"/>'
|
|
+ '<Override PartName="/xl/sharedStrings.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sharedStrings+xml"/>'
|
|
+ "</Types>",
|
|
"_rels/.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="xl/workbook.xml"/>'
|
|
+ "</Relationships>",
|
|
"xl/_rels/workbook.xml.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet" Target="worksheets/sheet1.xml"/>'
|
|
+ '<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/sharedStrings" Target="sharedStrings.xml"/>'
|
|
+ "</Relationships>",
|
|
"xl/workbook.xml": _XML
|
|
+ '<workbook xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
|
+ ' xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">'
|
|
+ '<sheets><sheet name="Krav" sheetId="1" r:id="rId1"/></sheets></workbook>',
|
|
"xl/sharedStrings.xml": _XML
|
|
+ '<sst xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main" count="3" uniqueCount="3">'
|
|
+ "<si><t>Krav til helning</t></si><si><t>60 og 70</t></si><si><t>1:15</t></si></sst>",
|
|
"xl/worksheets/sheet1.xml": _XML
|
|
+ '<worksheet xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main">'
|
|
+ '<dimension ref="A1:B2"/><sheetData>'
|
|
+ '<row r="1"><c r="A1" t="s"><v>0</v></c></row>'
|
|
+ '<row r="2"><c r="A2" t="s"><v>1</v></c><c r="B2" t="s"><v>2</v></c></row>'
|
|
+ "</sheetData></worksheet>",
|
|
}
|
|
|
|
|
|
# A price sheet, and the negative control for it, in ONE workbook.
|
|
#
|
|
# Sheet 1 mirrors the shape measured on the K2 price sheet: row 1 carries a
|
|
# single title cell, so the table's HEADER ROW names one column while the rows
|
|
# below it carry three. That is what makes a reader see one column and a
|
|
# whitespace carpet where the source has a label and an amount. Column B also
|
|
# holds one long prose cell, which is what makes the simple-table writer pad
|
|
# every other row in that column out to its width.
|
|
#
|
|
# Sheet 2 is the negative control in the same file: one column in the SOURCE,
|
|
# so there are no columns to recover and nothing for a fix to invent.
|
|
#
|
|
# THREE NUMERIC CELLS AND ONE THAT ONLY LOOKS NUMERIC. `5647500` and `250000`
|
|
# are stored as numbers and are integral; `12.5` is stored as a number and is
|
|
# not; `92.0` is a SHARED STRING. The converter renders the first two as
|
|
# `5647500.0` and `250000.0` and the last one as `92.0` -- identical output for
|
|
# a number and for text, which is the whole reason the shared string table is
|
|
# consulted before any of them is rewritten.
|
|
_PRISARK_STRINGS = (
|
|
"Prisskjema",
|
|
"Post",
|
|
"Beskrivelse",
|
|
"Sum",
|
|
"01",
|
|
"Rigging og drift av byggeplass, medregnet alt som ikke er priset "
|
|
"spesifikt nedenfor og alt som er innkalkulert i de angitte prisene",
|
|
"02",
|
|
"Andel",
|
|
"03",
|
|
"92.0",
|
|
"Notat",
|
|
"Ingen kolonner her",
|
|
"Sum ikke oppgitt",
|
|
# A cell whose own text contains a pipe and a number. The converter escapes
|
|
# the pipe inside a pipe table, and the escape is what the rewrite's
|
|
# delimiter test has to survive: a `5.0` INSIDE a cell is not a cell.
|
|
"Kode 4 | 5.0",
|
|
"04",
|
|
)
|
|
|
|
_PRISARK_SHEET1 = (
|
|
'<row r="1"><c r="A1" t="s"><v>0</v></c></row>'
|
|
'<row r="2"><c r="A2" t="s"><v>1</v></c><c r="B2" t="s"><v>2</v></c>'
|
|
'<c r="C2" t="s"><v>3</v></c></row>'
|
|
'<row r="3"><c r="A3" t="s"><v>4</v></c><c r="B3" t="s"><v>5</v></c>'
|
|
'<c r="C3"><v>5647500</v></c></row>'
|
|
'<row r="4"><c r="A4" t="s"><v>6</v></c><c r="B4" t="s"><v>7</v></c>'
|
|
'<c r="C4"><v>12.5</v></c></row>'
|
|
'<row r="5"><c r="A5" t="s"><v>8</v></c><c r="B5" t="s"><v>9</v></c>'
|
|
'<c r="C5"><v>250000</v></c></row>'
|
|
'<row r="6"><c r="A6" t="s"><v>14</v></c><c r="B6" t="s"><v>13</v></c></row>'
|
|
)
|
|
|
|
_PRISARK_SHEET2 = (
|
|
'<row r="1"><c r="A1" t="s"><v>10</v></c></row>'
|
|
'<row r="2"><c r="A2" t="s"><v>11</v></c></row>'
|
|
'<row r="3"><c r="A3" t="s"><v>12</v></c></row>'
|
|
)
|
|
|
|
|
|
def _sheet(dimension: str, rows: str) -> str:
|
|
return (
|
|
_XML
|
|
+ '<worksheet xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main">'
|
|
+ f'<dimension ref="{dimension}"/><sheetData>'
|
|
+ rows
|
|
+ "</sheetData></worksheet>"
|
|
)
|
|
|
|
|
|
_PRISARK_PARTS = {
|
|
"[Content_Types].xml": _XML
|
|
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
|
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
|
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>'
|
|
+ '<Override PartName="/xl/workbook.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml"/>'
|
|
+ '<Override PartName="/xl/worksheets/sheet1.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml"/>'
|
|
+ '<Override PartName="/xl/worksheets/sheet2.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml"/>'
|
|
+ '<Override PartName="/xl/sharedStrings.xml" ContentType="application/vnd.openxmlformats-officedocument.spreadsheetml.sharedStrings+xml"/>'
|
|
+ "</Types>",
|
|
"_rels/.rels": _XLSX_PARTS["_rels/.rels"],
|
|
"xl/_rels/workbook.xml.rels": _XML
|
|
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
|
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet" Target="worksheets/sheet1.xml"/>'
|
|
+ '<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet" Target="worksheets/sheet2.xml"/>'
|
|
+ '<Relationship Id="rId3" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/sharedStrings" Target="sharedStrings.xml"/>'
|
|
+ "</Relationships>",
|
|
"xl/workbook.xml": _XML
|
|
+ '<workbook xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
|
+ ' xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">'
|
|
+ '<sheets><sheet name="Prisark" sheetId="1" r:id="rId1"/>'
|
|
+ '<sheet name="Enkeltkolonne" sheetId="2" r:id="rId2"/></sheets></workbook>',
|
|
"xl/sharedStrings.xml": _XML
|
|
+ '<sst xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
|
+ f' count="{len(_PRISARK_STRINGS)}" uniqueCount="{len(_PRISARK_STRINGS)}">'
|
|
+ "".join(f"<si><t>{value}</t></si>" for value in _PRISARK_STRINGS)
|
|
+ "</sst>",
|
|
"xl/worksheets/sheet1.xml": _sheet("A1:C6", _PRISARK_SHEET1),
|
|
"xl/worksheets/sheet2.xml": _sheet("A1:A3", _PRISARK_SHEET2),
|
|
}
|
|
|
|
|
|
# A sheet with an EMPTY ROW IN THE MIDDLE, which the converter renders as a
|
|
# pipe row of nothing but spaces. That row looks exactly like a table's own
|
|
# separator line to any rule that reads the line rather than its position --
|
|
# and swallowing it renumbers every row after it, silently, for the whole
|
|
# sheet. Measured on the K2 price sheet before this fixture existed: 8 empty
|
|
# rows, and the last row reported as 92 when the workbook says 100.
|
|
_TOMRAD_STRINGS = ("Rad en", "Rad to", "Rad fire")
|
|
|
|
_TOMRAD_SHEET = (
|
|
'<row r="1"><c r="A1" t="s"><v>0</v></c></row>'
|
|
'<row r="2"><c r="A2" t="s"><v>1</v></c></row>'
|
|
'<row r="3"><c r="A3"/></row>'
|
|
'<row r="4"><c r="A4" t="s"><v>2</v></c></row>'
|
|
)
|
|
|
|
_TOMRAD_PARTS = {
|
|
"[Content_Types].xml": _XLSX_PARTS["[Content_Types].xml"],
|
|
"_rels/.rels": _XLSX_PARTS["_rels/.rels"],
|
|
"xl/_rels/workbook.xml.rels": _XLSX_PARTS["xl/_rels/workbook.xml.rels"],
|
|
"xl/workbook.xml": _XML
|
|
+ '<workbook xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
|
+ ' xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">'
|
|
+ '<sheets><sheet name="Tomrad" sheetId="1" r:id="rId1"/></sheets></workbook>',
|
|
"xl/sharedStrings.xml": _XML
|
|
+ '<sst xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
|
+ f' count="{len(_TOMRAD_STRINGS)}" uniqueCount="{len(_TOMRAD_STRINGS)}">'
|
|
+ "".join(f"<si><t>{value}</t></si>" for value in _TOMRAD_STRINGS)
|
|
+ "</sst>",
|
|
"xl/worksheets/sheet1.xml": _sheet("A1:A4", _TOMRAD_SHEET),
|
|
}
|
|
|
|
|
|
def build_ooxml(parts: dict[str, str]) -> bytes:
|
|
"""Zip the parts with a fixed timestamp and no compression variance.
|
|
|
|
A constant `date_time` on every member is what makes the output
|
|
reproducible: a zip records mtime, so the default `ZipFile.writestr` would
|
|
stamp the current time and the fixture would differ on every run --
|
|
which would make `git diff --quiet` useless as the regeneration check.
|
|
"""
|
|
out = io.BytesIO()
|
|
with zipfile.ZipFile(out, "w", compression=zipfile.ZIP_DEFLATED) as archive:
|
|
for name, payload in parts.items():
|
|
info = zipfile.ZipInfo(name, date_time=_ZIP_DATE)
|
|
info.compress_type = zipfile.ZIP_DEFLATED
|
|
archive.writestr(info, payload)
|
|
return out.getvalue()
|
|
|
|
|
|
if __name__ == "__main__":
|
|
for name, content in (
|
|
("two-line-krav.pdf", KRAV_CONTENT),
|
|
("no-text-layer.pdf", NO_TEXT_CONTENT),
|
|
):
|
|
(HERE / name).write_bytes(build_pdf(content))
|
|
print(f"wrote {name}")
|
|
|
|
(HERE / "three-page-krav.pdf").write_bytes(build_paged_pdf(PAGED_CONTENTS))
|
|
print("wrote three-page-krav.pdf")
|
|
|
|
for name, parts in (
|
|
("two-line-krav.docx", _DOCX_PARTS),
|
|
("no-styles-krav.docx", _DOCX_NO_STYLES_PARTS),
|
|
("two-line-krav.xlsx", _XLSX_PARTS),
|
|
("prisark.xlsx", _PRISARK_PARTS),
|
|
("tomrad.xlsx", _TOMRAD_PARTS),
|
|
):
|
|
(HERE / name).write_bytes(build_ooxml(parts))
|
|
print(f"wrote {name}")
|