feat(assets): a bundle carries the images its sources declare (0.10.0)

Until now no reader in this package fetched, named, described or copied a
single image. `<img>`'s attributes were never read, a NISO-STS `<graphic>`
was walked past, a PDF was opened for its text alone, the converter's
markdown writer dropped every picture, and the only writer into a bundle
took `content: str`. The two lossiness warnings said so on every run, which
made the loss honest and did not make it smaller.

Measured on R761 Prosesskoden:2025, published as a 701-page PDF and as a
NISO-STS delivery: the process text is carried in full while 12 `Tabell N-N`
and 9 `Figur N-N` captions stand over nothing, because that publisher ships
those tables as raster pictures in both. Process 84's "toleranseklasse ...
er gitt i tabell 84-2" points at empty space.

THE GATE WAS WRITTEN FIRST AND RED. `tests/test_asset_gate.py` reads its
denominator out of the source (`page.images`, `word/media/`, `ppt/media/`,
`<img`, `<graphic`), never from a constant here. Measured at 332961a, built
from `git archive` and not from the editable tree: carried 0 of 8 local
images across 5 documents (9 declared), and no `assets/` at all. After: 8 of
8, with the ninth a remote source carried as a pointer without a file.

FIVE READERS PLACE, ONE MODULE DECIDES. `assets.py` owns what an image is
(sniffed from the bytes, never from the claimed extension), what it is
called (`<sha256[:12]>-<the source's own basename>`) and how it is pointed
at (one two-line block, one regex). `.xlsx` is deliberately not a row: a
block inside its pipe tables would break the `source_rows` locator, and 0 of
4 K2 workbooks hold media.

A PDF stream that is already a file is carried VERBATIM (29 of R761's 50
objects are DCTDecode); raw samples are encoded to PNG with stdlib zlib, so
no new dependency. Rendering the page region was the alternative and was
felled on determinism: a rasterised crop's bytes, and therefore the asset's
content-addressed name and the bundle's digest, would depend on the
installed rasteriser. What the encoder cannot express exactly is refused
with a code and counted, never approximated.

NO SIZE FLOOR, and that is a measurement: over the 4 828 image objects of
the K2 corpus the size distribution is a broad spread with no gap, unlike
OCR_CID_SHARE's bimodal one, so a threshold would be a number we chose.

ON BY DEFAULT, AND THE CONTROL IS TWO WHOLE BUILDS. The 43-document
reference corpus at 332961a versus rebuilt at HEAD with `--no-assets`:
865 files on both sides, `diff -rq` reports ONE difference, the added
`Images: NOT CARRIED` line in log.md. Every concept byte-identical.
Against the default: 453 -> 454 concepts, 865 -> 867 md, 0 -> 2 964 assets
(2 964 carried of 3 145 found, 4 622 pointers), 4.7 MB -> 115 MB, 2 414 s ->
3 088 s, peak RSS 6.26 -> 8.74 GB, 422 of 865 md files differ. The one new
concept has a measured cause: the pointers are body text, so a section
holding 146 of that document's images grew from 19.0 % to 30.6 % of the
extracted text and crossed `--outline-gate`'s 0.20 share clause.

THE IMAGE BYTES ARE NOT SCREENED. The guard is text-only, the pointer block
passes the gate as body text, the picture beside it passes nothing, and
log.md says so on every run.

Also fixed, both found by measuring rather than by reading:

- a markdown image is no longer read as a cross-reference. `structure._LINK`
  never looked at the character in front of the bracket, so every pointer
  would have arrived in the index as an edge to a concept that cannot exist.
- Door C carries the assets its merged concepts point at. Before this,
  importing a bundle built with `--assets` merged 6 of 6 concepts and wrote
  no `assets/` at all, so every pointer named a missing file.

Report: docs/2026-09-17-bilder-i-bundlen-trinn1.md
Spec proposal: docs/plan/okf-assets-section-6-4.md
Suite 1 955 passed / 1 skipped (from 1 896), ruff and mypy --strict clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-17 10:01:31 +02:00
commit bc39e8091f
33 changed files with 3638 additions and 64 deletions

View file

@ -13,7 +13,9 @@ import importlib.util
import json
import sqlite3
import sys
import tempfile
import urllib.error
import warnings
from pathlib import Path
from typing import Any
@ -700,3 +702,92 @@ def test_segmentation_plan_unmatched(tmp_path: Path) -> None:
root_frontmatter_values={"bundle_id": "b-1"},
)
assert code_of(excinfo) == "segmentation_plan_unmatched"
# --- asset codes (0.10.0) --------------------------------------------------
#
# One test per code, like every code above it. These five are the only codes in
# the registry that a caller is expected to COUNT rather than to act on: an
# image a reader could not carry becomes a row in the run log and a line in the
# concept, never a failed document.
def test_asset_type_unknown() -> None:
from llm_ingestion_okf import assets
with pytest.raises(ExtractionError) as excinfo:
assets.read_image(b"%PDF-1.7\n", name="figur.png")
assert excinfo.value.code == "asset_type_unknown"
def test_asset_samples_invalid() -> None:
from llm_ingestion_okf import assets
with pytest.raises(ExtractionError) as excinfo:
assets.encode_png(8, 8, b"\x00", channels=1)
assert excinfo.value.code == "asset_samples_invalid"
def test_asset_remote() -> None:
from llm_ingestion_okf.extract import extract_document
document = extract_document(
"side.html",
b'<html><body><img src="https://example.invalid/x.png" alt="x"></body></html>',
assets=True,
)
assert [item.code for item in document.rejected] == ["asset_remote"]
def test_asset_unresolved() -> None:
from llm_ingestion_okf.extract import extract_document
document = extract_document(
"side.html",
b'<html><body><img src="mangler.png" alt="x"></body></html>',
assets=True,
)
assert [item.code for item in document.rejected] == ["asset_unresolved"]
def test_asset_pdf_unsupported() -> None:
"""A PDF image this encoder refuses rather than approximates.
A 1-bit stencil: carrying it at 8 bits would be a decision about what black
means, and a wrong one is indistinguishable from a right one in the output.
"""
pytest.importorskip("pdfplumber")
from llm_ingestion_okf.extract import extract_document
data = (
Path(__file__).parent / "fixtures" / "image-inbox" / "prosess-84-tabell.pdf"
).read_bytes()
stencil = data.replace(
b"/BitsPerComponent 8 /Filter /FlateDecode", b"/BitsPerComponent 1 /Filter /FlateDecode", 1
)
assert stencil != data
with warnings.catch_warnings():
warnings.simplefilter("ignore")
document = extract_document("krav.pdf", stencil, assets=True)
assert "asset_pdf_unsupported" in [item.code for item in document.rejected]
def test_asset_collision() -> None:
"""Two different pictures reducing to one asset name, refused in the run.
Constructed rather than found: the name carries 12 hex of the digest of its
own bytes, so reaching this by accident is a 48-bit collision. The code
exists because resolving it silently would lose one of the two pictures
while every pointer to it kept showing the other.
"""
from llm_ingestion_okf.assets import ExtractedImage, asset_name
from llm_ingestion_okf.inbox import _write_assets
image = ExtractedImage(b"AAAA", "f.png", "image/png", ".png", 1, 1)
# One name already holding DIFFERENT bytes, which is what a 48-bit digest
# collision would look like from inside the run.
seen = {asset_name(image): b"BBBB"}
with tempfile.TemporaryDirectory() as root:
with pytest.raises(MaterializationError) as excinfo:
_write_assets(Path(root) / "bundle", [image], seen)
assert excinfo.value.code == "asset_collision"