feat(inbox): point every concept at the document it came from, with a locator per format

A concept named its source file by basename and, when segmented, carried a
`source_offset` into the text THIS LIBRARY extracted. Following that pointer
needed the corpus directory, the extractor and its exact transitive version --
none of which the bundle carries. Hand-walked on a real K2 concept: six steps,
four of them requiring knowledge from outside the bundle, to learn that a
requirement sits on pages 12-13 of a 20-page document.

The address is spec's: `sources: [{ resource, title }]`, where `resource` is
the dropped file's inbox-relative path (SPEC v0.2 5.1:303-306 -- "an absolute
URL, a bundle-relative path, or a path into a `references/` subdirectory").
The locator is ours, and it has to be: 5.1 has no field for a place within a
resource, and the pinned guard (1.3.0) rejects every route to putting one
inside a `sources` entry -- a non-allowlisted key by name, a nested flow list
as "scalar leaves only", and quoting as an unsupported form. So the locator is
top-level keys shaped like `source_offset`, and a path carrying a flow
terminator is refused fail-fast rather than mangled.

The unit table is built AT EXTRACTION, where the extracted text and the
original's structure are known to agree: pdf -> `source_pages` from
pdfplumber's own page numbers (a page that yielded no text does not renumber
the ones after it), xlsx -> `source_sheet` + `source_rows`, everything else ->
`source_lines`. `source_offset` stays.

Two measurements changed the design before it shipped. A `paragraphs` key for
docx would name a number the document does not have: `<w:p>` counts of
108/27/65/176/57 against converted-markdown lines of 75/33/67/144/63, not one
pair agreeing -- so the key is `source_lines` and says what it indexes. And an
empty spreadsheet row renders exactly like a table separator: the content-based
rule ate 8 empty rows on the K2 price sheet and reported its last row as 92
against a workbook that says 100. The separator is now found by position, and
`tomrad.xlsx` keeps that red.

One profile moves. `provenance` is a policy object, `None` everywhere but
`SEGMENTED_OKF_V0_2`; the other five shipped profiles are byte-identical.

K2 rebuilt from a frozen src copy: 629 concepts, 1108 files, name set identical,
0 ids moved, 479 files byte-identical, 629 changed and 0 lines removed anywhere.
629/629 now carry an address and a locator. New ref
`sha256-tree:665563a2f74423fcbcc8e4f0b0954ee73b73985ac0418de4f6987bd162a1f7c8`;
`2f82fcfe...` is stale. The pre-pass payload does not grow by one byte
(209 092 B before and after, 18 changed lines: the ref and eight per-concept
digests) -- because an excerpt carries the body, not the frontmatter, which is
also why the consumer still cannot cite "file X page 12" from a payload alone.

Report: docs/2026-09-08-proveniens-k2.md. 1339 tests, ruff and mypy clean.

Co-Authored-By: Claude <claude-opus-5>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-08 14:39:24 +02:00
commit b6a8c8bd89
16 changed files with 1301 additions and 26 deletions

View file

@ -37,16 +37,57 @@ KRAV_CONTENT = (
NO_TEXT_CONTENT = b"20 20 160 160 re S\n"
# One line of text per page, so the page a stretch of extracted text came from
# is decidable by reading the text alone. THREE pages rather than two, and the
# MIDDLE one carries no text operators: a two-page fixture cannot tell a page
# INDEX from a page NUMBER, and without a blank page in the middle it cannot
# tell either of them from a count of the pages that produced text. The
# extractor drops empty pages, so the third page's text belongs to page 3 and
# to no other number.
PAGED_CONTENTS = (
b"BT /F1 12 Tf 20 160 Td (Side en om helning) Tj ET\n",
b"20 20 160 160 re S\n",
b"BT /F1 12 Tf 20 160 Td (Side tre om utkiling) Tj ET\n",
)
def build_pdf(content: bytes) -> bytes:
"""Assemble a one-page PDF around `content` as the page content stream."""
return build_paged_pdf((content,))
def build_paged_pdf(contents: tuple[bytes, ...]) -> bytes:
"""Assemble a PDF with one page per entry of `contents`.
The object numbering is laid out first and the xref offsets computed from
the emitted bytes, exactly as the single-page form did -- the fixture stays
a hand-written statement about the format rather than a library's output.
"""
count = len(contents)
# 1 catalog, 2 pages, then one page object and one content stream per page,
# and the shared font last.
page_numbers = [3 + 2 * index for index in range(count)]
font_number = 3 + 2 * count
kids = b" ".join(str(number).encode() + b" 0 R" for number in page_numbers)
objects = [
b"<< /Type /Catalog /Pages 2 0 R >>",
b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] "
b"/Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>",
b"<< /Length " + str(len(content)).encode() + b" >>\nstream\n" + content + b"endstream",
b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>",
b"<< /Type /Pages /Kids [" + kids + b"] /Count " + str(count).encode() + b" >>",
]
for index, content in enumerate(contents):
stream_number = page_numbers[index] + 1
objects.append(
b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Contents "
+ str(stream_number).encode()
+ b" 0 R /Resources << /Font << /F1 "
+ str(font_number).encode()
+ b" 0 R >> >> >>"
)
objects.append(
b"<< /Length " + str(len(content)).encode() + b" >>\nstream\n" + content + b"endstream"
)
objects.append(
b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>"
)
out = bytearray(b"%PDF-1.4\n")
offsets = []
@ -264,6 +305,38 @@ _PRISARK_PARTS = {
}
# A sheet with an EMPTY ROW IN THE MIDDLE, which the converter renders as a
# pipe row of nothing but spaces. That row looks exactly like a table's own
# separator line to any rule that reads the line rather than its position --
# and swallowing it renumbers every row after it, silently, for the whole
# sheet. Measured on the K2 price sheet before this fixture existed: 8 empty
# rows, and the last row reported as 92 when the workbook says 100.
_TOMRAD_STRINGS = ("Rad en", "Rad to", "Rad fire")
_TOMRAD_SHEET = (
'<row r="1"><c r="A1" t="s"><v>0</v></c></row>'
'<row r="2"><c r="A2" t="s"><v>1</v></c></row>'
'<row r="3"><c r="A3"/></row>'
'<row r="4"><c r="A4" t="s"><v>2</v></c></row>'
)
_TOMRAD_PARTS = {
"[Content_Types].xml": _XLSX_PARTS["[Content_Types].xml"],
"_rels/.rels": _XLSX_PARTS["_rels/.rels"],
"xl/_rels/workbook.xml.rels": _XLSX_PARTS["xl/_rels/workbook.xml.rels"],
"xl/workbook.xml": _XML
+ '<workbook xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
+ ' xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships">'
+ '<sheets><sheet name="Tomrad" sheetId="1" r:id="rId1"/></sheets></workbook>',
"xl/sharedStrings.xml": _XML
+ '<sst xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
+ f' count="{len(_TOMRAD_STRINGS)}" uniqueCount="{len(_TOMRAD_STRINGS)}">'
+ "".join(f"<si><t>{value}</t></si>" for value in _TOMRAD_STRINGS)
+ "</sst>",
"xl/worksheets/sheet1.xml": _sheet("A1:A4", _TOMRAD_SHEET),
}
def build_ooxml(parts: dict[str, str]) -> bytes:
"""Zip the parts with a fixed timestamp and no compression variance.
@ -289,11 +362,15 @@ if __name__ == "__main__":
(HERE / name).write_bytes(build_pdf(content))
print(f"wrote {name}")
(HERE / "three-page-krav.pdf").write_bytes(build_paged_pdf(PAGED_CONTENTS))
print("wrote three-page-krav.pdf")
for name, parts in (
("two-line-krav.docx", _DOCX_PARTS),
("no-styles-krav.docx", _DOCX_NO_STYLES_PARTS),
("two-line-krav.xlsx", _XLSX_PARTS),
("prisark.xlsx", _PRISARK_PARTS),
("tomrad.xlsx", _TOMRAD_PARTS),
):
(HERE / name).write_bytes(build_ooxml(parts))
print(f"wrote {name}")