llm-ingestion-okf/tests/test_asset_viewable.py
Kjell Tore Guttormsen 9d1f4b14ed test(fixtures): replace sector-specific example material with generic, fictitious examples — green
Every fixture, test document, tool example and document now uses an invented
kitchen-and-baking handbook series, written in this repository. The package's
behaviour is unchanged; src/ changes are comments and help text only.

- Generated fixtures are regenerated from their generators. Their structural
  counts are identical before and after: elements, images, rows, cells,
  headings, bookmarks and the witness inventory's per-document totals. The
  image-inbox and accounting documents are renamed kapittel-84-*.
- tools/okf_accounting_gate.py: the two options that named one real corpus
  each are replaced by a generic, repeatable --corpus PATH with no default.
  Row 5 compares the PDF pair alone. Gate verdict unchanged: RED rows 2, 3, 6.
- tools/okf_witness.py: the STS JSON reader for one publisher's delivery is
  removed, along with its three twins and five tests. The mutation harness
  loses W09.
- docs/: 13 dated reports that documented runs on a retired reference corpus
  are removed, and 40 are neutralized. Dead links are removed, and no new
  dangling path is introduced.
- The synthetic MCP-gate corpus and the residual probe words are neutral.

Valgt: keep the `okf quality --fasit` bar value (the measured fraction, one corpus) and
rewrite only its provenance, because the verdict stays unchanged and the
number names nothing.

Term check with the local list: 0 of 411 tracked files, 0 file names, 0 of
27 binary fixtures. Suite after git add: 2457 passed, 1 skipped. The base
tree had 2460 passed and 2 skipped; five tests went with the JSON reader and
four were added by the term check. ruff, ruff format and mypy --strict src/
are clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:52:02 +02:00

1073 lines
44 KiB
Python

"""Every image a bundle carries must be one a model can actually SHOW.
0.10.0 carried what a document declared, in whatever format the publisher
happened to ship. Measured 2026-09-19 over a frozen reference delivery's own
`assets/` directory (denominator 50): 29 JPEG, 2 PNG and **19 "PC bitmap,
Windows 3.x, 8-bit, compression 1"** -- RLE8 BMP. The 19 are carried byte for
byte and are correct files; they are also files no model reads, so a figure
the bundle holds is a figure no arm can see. An absence a reader is shown is
information; a picture that is present and unreadable is worse than either,
because the count says it arrived.
THE PROPERTY, and it is a property rather than a list of formats we happen to
have met: for every asset in a built bundle, the format READ OFF ITS BYTES is
in the viewable set. A source image outside that set is converted losslessly
to PNG, or refused with a published code and a line in the concept saying what
stood there. Never carried silently in a format nothing can read.
The viewable set is stated here independently of the package, because a test
that imported the package's own constant would agree with it by construction
and measure nothing. Same for the rejection codes and the pixel grids below:
each guard is pinned against a count or a literal written out in this file, and
the lossless guard decodes through Pillow -- a decoder this package does not
use and does not ship -- so "identical pixels" is not this module agreeing with
itself.
"""
from __future__ import annotations
import hashlib
import io
import itertools
import struct
import warnings
from pathlib import Path
import pytest
from llm_ingestion_okf import cli, corpus
from llm_ingestion_okf.assets import ASSETS_DIR, read_image, sniff
from llm_ingestion_okf.errors import ExtractionError
BUNDLE_ID = "asset-viewable-fixture"
OKF_VERSION = "0.2"
#: What a model can be shown. Written out here, never imported: the four
#: formats a vision-capable model decodes. A `.bmp`, a `.tiff` and a JPEG 2000
#: codestream are all legitimate image files and none of them is on this list.
VIEWABLE = frozenset({"image/png", "image/jpeg", "image/gif", "image/webp"})
#: The codes this file requires. Stated here so a rename of the package's own
#: constant cannot make these guards pass by agreeing with it.
CODE_NOT_VIEWABLE = "asset_not_viewable"
CODE_TOO_LARGE = "asset_too_large"
# --- fixtures built byte by byte, so the expected pixels are known ----------
#: Four RGB triples. Index 0 is white, which is also what an RLE delta skips
#: over, so the expected grid below states what a skip leaves behind.
PALETTE = bytes((255, 255, 255, 0, 0, 0, 200, 16, 32, 16, 96, 200))
#: An 8x4 image as palette indices, TOP-DOWN. Written out rather than
#: generated: this is the fasit the two decodings are both measured against.
EXPECTED_INDICES = (
(1, 1, 1, 1, 1, 1, 1, 1),
(0, 1, 2, 3, 3, 2, 1, 0),
(3, 3, 3, 0, 0, 1, 1, 1),
(0, 0, 0, 0, 0, 1, 2, 3),
)
#: The same image as an RLE8 opcode stream, BOTTOM-UP as the format requires,
#: exercising all five opcodes: encoded run, absolute run (with its pad byte),
#: delta, end-of-line and end-of-bitmap. A fixture using runs alone would pass
#: with a decoder that never implements the other three.
RLE8_STREAM = bytes(
(
0x05,
0x00,
0x00,
0x03,
0x01,
0x02,
0x03,
0x00,
0x00,
0x00,
0x03,
0x03,
0x00,
0x02,
0x02,
0x00,
0x03,
0x01,
0x00,
0x00,
0x00,
0x08,
0x00,
0x01,
0x02,
0x03,
0x03,
0x02,
0x01,
0x00,
0x00,
0x00,
0x08,
0x01,
0x00,
0x01,
)
)
def expected_rgb() -> bytes:
return b"".join(PALETTE[index * 3 : index * 3 + 3] for row in EXPECTED_INDICES for index in row)
def _dib(width: int, height: int, bits: int, compression: int, body: int, entries: int) -> bytes:
return struct.pack(
"<IiiHHIIiiII", 40, width, height, 1, bits, compression, body, 3779, 3779, entries, entries
)
def _file_header(offbits: int, body: int) -> bytes:
return b"BM" + struct.pack("<IHHI", offbits + body, 0, 0, offbits)
def bmp_rle8(width: int, height: int, palette: bytes, stream: bytes) -> bytes:
"""A Windows 3.x 8-bit RLE8 BMP, the shape 19 of a real delivery's 50 assets have."""
entries = len(palette) // 3
table = b"".join(
bytes((palette[i * 3 + 2], palette[i * 3 + 1], palette[i * 3], 0)) for i in range(entries)
)
offbits = 14 + 40 + len(table)
return (
_file_header(offbits, len(stream))
+ _dib(width, height, 8, 1, len(stream), entries)
+ table
+ stream
)
def bmp_24(rows: tuple[tuple[tuple[int, int, int], ...], ...]) -> bytes:
"""An uncompressed 24-bit BMP -- the other half of the order's fixture."""
height = len(rows)
width = len(rows[0])
pad = b"\x00" * ((-width * 3) % 4)
body = b"".join(
b"".join(bytes((blue, green, red)) for (red, green, blue) in row) + pad
for row in reversed(rows)
)
offbits = 14 + 40
return _file_header(offbits, len(body)) + _dib(width, height, 24, 0, len(body), 0) + body
#: A minimal little-endian TIFF: one IFD with no usable entries. It is a real
#: TIFF by its magic and it is not a format this package can decode with the
#: stdlib, so it is the order's "cannot be converted" case.
TIFF_STUB = b"II\x2a\x00" + struct.pack("<I", 8) + struct.pack("<H", 0) + struct.pack("<I", 0)
RGB_ROWS = (
((255, 0, 0), (0, 255, 0), (0, 0, 255), (255, 255, 0), (0, 0, 0)),
((1, 2, 3), (4, 5, 6), (7, 8, 9), (10, 11, 12), (13, 14, 15)),
((255, 255, 255), (128, 128, 128), (0, 0, 0), (17, 34, 51), (68, 85, 102)),
)
def png_1x1() -> bytes:
"""A real PNG, built with zlib alone -- the already-viewable control."""
import zlib
def chunk(kind: bytes, payload: bytes) -> bytes:
return (
len(payload).to_bytes(4, "big")
+ kind
+ payload
+ zlib.crc32(kind + payload).to_bytes(4, "big")
)
ihdr = struct.pack(">IIBBBBB", 1, 1, 8, 2, 0, 0, 0)
return (
b"\x89PNG\r\n\x1a\n"
+ chunk(b"IHDR", ihdr)
+ chunk(b"IDAT", zlib.compress(b"\x00\x10\x20\x30", 9))
+ chunk(b"IEND", b"")
)
def jpeg_bytes() -> bytes:
"""A real baseline JPEG, encoded by Pillow -- the second viewable control."""
Image = pytest.importorskip("PIL.Image")
buffer = io.BytesIO()
Image.new("RGB", (4, 3), (10, 20, 30)).save(buffer, format="JPEG")
return buffer.getvalue()
def _try_decode(data: bytes) -> tuple[tuple[int, int], bytes] | None:
"""`decode_rgb`, or `None` where the independent decoder cannot read it.
Needed because this guard scans EVERY carried asset, and the point of the
round is that a bundle may hold one nothing can decode. A raised
`UnidentifiedImageError` would make the guard red for the helper's reason
instead of for the claim's.
"""
try:
return decode_rgb(data)
except Exception:
return None
def decode_rgb(data: bytes) -> tuple[tuple[int, int], bytes]:
"""An INDEPENDENT decoding: Pillow, which this package neither uses nor ships."""
Image = pytest.importorskip("PIL.Image")
with Image.open(io.BytesIO(data)) as image:
return image.size, image.convert("RGB").tobytes()
# --- the inbox and the build ------------------------------------------------
def _inbox(root: Path) -> Path:
inbox = root / "inbox"
(inbox / "graphics").mkdir(parents=True)
(inbox / "graphics" / "figur-rle8.bmp").write_bytes(bmp_rle8(8, 4, PALETTE, RLE8_STREAM))
(inbox / "graphics" / "figur-flat.bmp").write_bytes(bmp_24(RGB_ROWS))
(inbox / "graphics" / "skjema.tiff").write_bytes(TIFF_STUB)
(inbox / "kapittel.html").write_text(
"<!doctype html>\n<html><head><title>Kapittel 84</title></head>\n<body>\n"
"<h1>84 Brød og boller</h1>\n"
"<p>Hevetidsklassene staar i figuren under.</p>\n"
'<img src="graphics/figur-rle8.bmp" alt="Figur 84-1 Hevetidsklasser">\n'
"<p>Og prinsippet:</p>\n"
'<img src="graphics/figur-flat.bmp" alt="Figur 84-2 Prinsipp">\n'
"<p>Og skjemaet:</p>\n"
'<img src="graphics/skjema.tiff" alt="Skjema 84-3">\n'
"</body></html>\n",
encoding="utf-8",
)
return inbox
def _build(inbox: Path, bundle: Path, *extra: str) -> int:
with warnings.catch_warnings():
warnings.simplefilter("ignore")
return cli.main(
[
"build",
str(inbox),
"--bundle",
str(bundle),
"--bundle-id",
BUNDLE_ID,
"--okf-version",
OKF_VERSION,
*extra,
]
)
def _concepts(bundle: Path) -> list[Path]:
return [
path
for path in sorted(bundle.rglob("*.md"))
if path.name not in {"index.md", corpus.LOG_NAME}
]
def _assets(bundle: Path) -> list[Path]:
directory = bundle / ASSETS_DIR
return sorted(p for p in directory.iterdir() if p.is_file()) if directory.is_dir() else []
# --- 1. visibility, as a property of every carried asset --------------------
def test_every_carried_asset_is_a_format_a_model_can_show(tmp_path: Path) -> None:
"""Eval point 1. RED before the conversion: two BMPs are carried as BMP.
The denominator is counted in this test from the inbox itself, so a build
that silently carried nothing would fail here rather than pass over an
empty set -- the failure mode a "0 of 0 violations" assertion cannot see.
"""
pytest.importorskip("llm_ingestion_guard")
inbox = _inbox(tmp_path)
dropped = [p for p in sorted((inbox / "graphics").iterdir()) if p.is_file()]
assert len(dropped) == 3, dropped
not_viewable_in = [p for p in dropped if (sniff(p.read_bytes()) or ("", ""))[0] not in VIEWABLE]
assert len(not_viewable_in) == 3, not_viewable_in
bundle = tmp_path / "bundle"
assert _build(inbox, bundle) == 0
carried = _assets(bundle)
assert carried, "the build carried no asset at all"
offending = [
(path.name, (sniff(path.read_bytes()) or (None,))[0])
for path in carried
if (sniff(path.read_bytes()) or (None,))[0] not in VIEWABLE
]
assert offending == [], (
f"{len(offending)} of {len(carried)} carried assets are in a format no model "
f"can show: {offending}"
)
# --- 2. lossless, against a decoder this package does not use ---------------
def test_the_rle8_source_decodes_to_the_grid_this_file_states() -> None:
"""The fixture's own known-positive: without it, guard 2 could compare two
wrong decodings and agree."""
size, rgb = decode_rgb(bmp_rle8(8, 4, PALETTE, RLE8_STREAM))
assert size == (8, 4)
assert rgb == expected_rgb()
def test_the_carried_png_holds_the_source_pixels_exactly(tmp_path: Path) -> None:
"""Eval point 2. Same RGB per pixel, through the palette, 8x4 = 32 pixels."""
pytest.importorskip("llm_ingestion_guard")
bundle = tmp_path / "bundle"
assert _build(_inbox(tmp_path), bundle) == 0
sources = {
"rle8": bmp_rle8(8, 4, PALETTE, RLE8_STREAM),
"flat24": bmp_24(RGB_ROWS),
}
carried = {path.name: path.read_bytes() for path in _assets(bundle)}
assert len(carried) >= 2, carried
matched = 0
for label, source in sources.items():
source_size, source_rgb = decode_rgb(source)
# VIEWABLE is part of the claim, not decoration: a BMP carried verbatim
# trivially holds its own pixels, so a guard that did not require the
# holder to be viewable would be green before the conversion exists.
hits = [
data
for data in carried.values()
if (sniff(data) or (None,))[0] in VIEWABLE
and _try_decode(data) == (source_size, source_rgb)
]
assert len(hits) == 1, f"{label}: {len(hits)} carried assets hold these pixels"
matched += 1
assert matched == 2
assert len(expected_rgb()) == 8 * 4 * 3
# --- 3. the bundle says what happened ---------------------------------------
def test_the_bundle_records_the_conversion_per_asset(tmp_path: Path) -> None:
"""Eval point 3: original format, original checksum, new checksum."""
pytest.importorskip("llm_ingestion_guard")
bundle = tmp_path / "bundle"
assert _build(_inbox(tmp_path), bundle) == 0
text = "\n".join(path.read_text("utf-8") for path in _concepts(bundle))
sources = (bmp_rle8(8, 4, PALETTE, RLE8_STREAM), bmp_24(RGB_ROWS))
for source in sources:
before = hashlib.sha256(source).hexdigest()
assert before in text, f"the bundle never states the source checksum {before[:12]}"
assert text.count("image/bmp") >= 2, "the bundle never states the original format"
for path in _assets(bundle):
after = hashlib.sha256(path.read_bytes()).hexdigest()
assert path.name.startswith(after[:12]), path.name
converted = [p for p in _assets(bundle) if p.suffix == ".png"]
assert len(converted) == 2, converted
for path in converted:
assert hashlib.sha256(path.read_bytes()).hexdigest() in text
def test_a_converted_asset_is_one_asset_and_not_two(tmp_path: Path) -> None:
"""The accounting's rule: converting is not a second element.
Two BMPs in, two files in `assets/`, two pointers, and the concept's own
`images:` count agreeing with the pointers it holds.
"""
pytest.importorskip("llm_ingestion_guard")
bundle = tmp_path / "bundle"
assert _build(_inbox(tmp_path), bundle) == 0
assert len(_assets(bundle)) == 2, _assets(bundle)
pointers = sum(path.read_text("utf-8").count("](/assets/") for path in _concepts(bundle))
assert pointers == 2, pointers
for path in _concepts(bundle):
body = path.read_text("utf-8")
if "\nimages: " in body:
line = next(row for row in body.splitlines() if row.startswith("images: "))
assert int(line.split(":", 1)[1]) == body.count("](/assets/")
# --- 4. the ceiling is checked before the pixels exist ----------------------
def test_a_bmp_declaring_a_huge_size_is_refused_before_it_is_decoded() -> None:
"""Eval point 4. The bound is on the DECLARATION, paid before the pixels."""
import tracemalloc
# 49 MP, just over the package's 40 MP bound. Deliberately NOT 50 000 x
# 50 000: the guard has to stay able to tell "refused before the decode"
# from "refused after it", and measuring the second costs a frame of the
# declared size. At 7 000 x 7 000 that frame is 49 MB -- affordable to
# measure, and two orders of magnitude over the 4 MB this asserts.
bomb = bmp_rle8(7_000, 7_000, PALETTE, bytes([0xFF, 0x01] * 64))
assert 7_000 * 7_000 > 40_000_000
assert len(bomb) < 4096, len(bomb)
tracemalloc.start()
try:
with pytest.raises(ExtractionError) as caught:
read_image(bomb, name="bombe.bmp")
finally:
_, peak = tracemalloc.get_traced_memory()
tracemalloc.stop()
assert caught.value.code == CODE_TOO_LARGE, caught.value.code
assert peak < 4 * 1024 * 1024, peak
def test_an_rle_stream_longer_than_its_declared_image_costs_the_declaration() -> None:
"""An RLE8 run list that paints far past an 8x4 frame pays for 8x4."""
import tracemalloc
flood = bytes([0xFF, 0x02] * 400_000) + bytes((0x00, 0x01))
bomb = bmp_rle8(8, 4, PALETTE, flood)
assert len(bomb) > 800_000, len(bomb)
tracemalloc.start()
try:
try:
image = read_image(bomb, name="flom.bmp")
except ExtractionError as exc:
code = exc.code
image = None
else:
code = None
finally:
_, peak = tracemalloc.get_traced_memory()
tracemalloc.stop()
if image is not None:
assert sniff(image.data) is not None
assert (sniff(image.data) or (None,))[0] in VIEWABLE
else:
assert code is not None
# The whole file is 800 KB; decoding it must not cost a multiple of that.
assert peak < 16 * 1024 * 1024, peak
def test_a_bmp_whose_name_claims_png_is_still_converted() -> None:
"""The type is read off the BYTES. A rule that trusted the extension would
carry this file unchanged and call it viewable, and every structural check
downstream would agree with it -- the count, the suffix and the pointer.
A real delivery is why this is not hypothetical: its graphics directory
holds `.bmp`, `.jpg` and `.png` side by side and the document's own
`xlink:href` values are whatever the publisher's tool wrote.
"""
image = read_image(bmp_rle8(8, 4, PALETTE, RLE8_STREAM), name="graphics/figur.png")
assert image.media_type == "image/png"
assert image.suffix == ".png"
assert decode_rgb(image.data) == ((8, 4), expected_rgb())
# --- 5. the known-positive: what is already viewable does not move ----------
def test_a_png_is_carried_byte_for_byte() -> None:
"""Eval point 5, in-test half. Conversion must reach nothing already viewable."""
data = png_1x1()
image = read_image(data, name="figur.png")
assert image.data == data
assert image.media_type == "image/png"
assert image.suffix == ".png"
def test_a_jpeg_is_carried_byte_for_byte() -> None:
data = jpeg_bytes()
image = read_image(data, name="figur.jpg")
assert image.data == data
assert image.media_type == "image/jpeg"
assert image.suffix == ".jpg"
# --- 6. what cannot be converted is refused out loud ------------------------
def test_a_tiff_is_refused_with_a_published_code() -> None:
"""Eval point 6: never silently carried in a format nothing reads."""
with pytest.raises(ExtractionError) as caught:
read_image(TIFF_STUB, name="skjema.tiff")
assert caught.value.code == CODE_NOT_VIEWABLE, caught.value.code
def test_a_corrupt_bmp_is_refused_with_a_published_code() -> None:
"""A BMP header with a body that stops mid-stream is a refusal, not a guess."""
truncated = bmp_rle8(8, 4, PALETTE, RLE8_STREAM)[:60]
with pytest.raises(ExtractionError) as caught:
read_image(truncated, name="kuttet.bmp")
assert caught.value.code.startswith("asset_"), caught.value.code
def test_the_refused_image_leaves_a_row_the_reader_can_see(tmp_path: Path) -> None:
"""A refusal is a line in the concept and a row in the run log."""
pytest.importorskip("llm_ingestion_guard")
bundle = tmp_path / "bundle"
assert _build(_inbox(tmp_path), bundle) == 0
text = "\n".join(path.read_text("utf-8") for path in _concepts(bundle))
assert "not carried" in text
assert "skjema.tiff" in text
log = (bundle / corpus.LOG_NAME).read_text("utf-8")
assert "image" in log.lower()
# --- 7. a truncated RLE8 stream is refused, never painted -------------------
#: The code the UNCOMPRESSED path already raises for a body that stops short
#: ("refusing to pad, because a short buffer means the header was read wrong").
#: Written out here rather than imported, like the two codes above.
CODE_SAMPLES_INVALID = "asset_samples_invalid"
def rle8_indices(stream: bytes, width: int, height: int) -> bytes:
"""A PERMISSIVE RLE8 decoder, written in this file, top-down.
It paints what the stream holds and leaves the rest at index 0. That is
what the format says about a SKIPPED pixel and what every decoder does
with a stream that simply stops -- which is exactly why a truncation is
invisible without a check: no decoder disagrees, they all agree on the
wrong picture. Its job here is to MEASURE what each cut costs, so the
guard below is pinned to a pixel count this file computed rather than to
the package's opinion of itself.
"""
rows = [bytearray(width) for _ in range(height)]
position, end, x, y = 0, len(stream), 0, 0
while position + 1 < end:
count, value = stream[position], stream[position + 1]
position += 2
if count:
if 0 <= y < height and x < width:
stop = min(x + count, width)
rows[y][x:stop] = bytes((value,)) * (stop - x)
x += count
continue
if value == 0:
x, y = 0, y + 1
elif value == 1:
break
elif value == 2:
if position + 2 > end:
break
x += stream[position]
y += stream[position + 1]
position += 2
else:
run = stream[position : position + value]
position += value + (value & 1)
if 0 <= y < height and x < width:
stop = min(x + len(run), width)
rows[y][x:stop] = run[: stop - x]
x += value
rows.reverse()
return b"".join(bytes(row) for row in rows)
def test_the_cuts_below_land_on_the_opcodes_this_file_names() -> None:
"""The fixture's own known-positive: a cut that no longer truncates
anything would make every guard below pass over nothing."""
assert len(RLE8_STREAM) == 36
assert RLE8_STREAM[32:34] == bytes((0x08, 0x01)), "the last encoded run"
assert RLE8_STREAM[20:22] == bytes((0x00, 0x08)), "the absolute block"
assert RLE8_STREAM[12:16] == bytes((0x00, 0x02, 0x02, 0x00)), "the delta escape"
assert RLE8_STREAM[34:36] == bytes((0x00, 0x01)), "end-of-bitmap"
assert rle8_indices(RLE8_STREAM, 8, 4) == bytes(
index for row in EXPECTED_INDICES for index in row
)
def test_a_truncated_rle8_stream_is_refused_with_a_published_code() -> None:
"""Eval point 6 again, on the one case that was silent.
Measured by PM 2026-09-19 on a real delivered asset (352x548 = 192 896 pixels):
a stream cut to 90 % carried with 13 923 pixels wrong, to 50 % with 95 890,
to 10 % with 166 525 -- no code, no row, a partly blank PNG under a name
that says it holds the source's pixels. The uncompressed path already
refuses the same shape. NO PIXEL IS GUESSED THAT THE STREAM DID NOT SKIP
ON PURPOSE: a stream must say it is finished, and
`test_a_stream_that_stops_before_the_frame_is_refused` adds the other half
-- when it says so, the cursor must have reached the end of the frame.
Pixels a delta or an end-of-line escape passed over keep index 0, which is
the format's own value and what an independent decoder produces.
THIS GUARD IS THE TERMINATOR, and it is read off the corpus rather than
chosen: over the 25 RLE8 BMPs a frozen reference delivery ships (24 distinct;
the bundle carries 19), 25 of 25 end at an explicit end-of-bitmap escape, that escape is the stream's LAST
two bytes on 25 of 25, and `biSizeImage` equals the available bytes on
25 of 25. So a stream that runs out before its terminator is refused, and
so is one that is whole but never states it is finished -- the two are the
same bytes from a reader's side, and the measurement says no real writer
here produces the second.
"""
full = rle8_indices(RLE8_STREAM, 8, 4)
cuts = {
"cut in the middle of an encoded run": (RLE8_STREAM[:33], 8),
"cut in the middle of an absolute block": (RLE8_STREAM[:26], 11),
"cut in the middle of a delta escape": (RLE8_STREAM[:14], 17),
"cut immediately before end-of-bitmap": (RLE8_STREAM[:34], 0),
"ends with end-of-line instead of end-of-bitmap": (RLE8_STREAM[:34] + bytes((0, 0)), 0),
}
for label, (stream, lost) in cuts.items():
differ = sum(a != b for a, b in zip(full, rle8_indices(stream, 8, 4)))
assert differ == lost, f"{label}: {differ} pixels differ, this file says {lost}"
with pytest.raises(ExtractionError) as caught:
read_image(bmp_rle8(8, 4, PALETTE, stream), name="kuttet.bmp")
assert caught.value.code == CODE_SAMPLES_INVALID, f"{label}: {caught.value.code}"
# The control: the refusal is about the cut, not about the fixture.
whole = read_image(bmp_rle8(8, 4, PALETTE, RLE8_STREAM), name="hel.bmp")
assert whole.media_type == "image/png"
#: Twelve RLE8 streams for one 8x4 frame, each naming what its terminator
#: leaves behind, and whether an INDEPENDENT decoder will read the file.
#: `reaches` is the property measured below: at end-of-bitmap the cursor
#: stands at or past the end of the last row, so every row was addressed,
#: AND every end-of-line escape closed a row it had started.
#: The values are Pillow's, measured 2026-09-19 on the twelve files this
#: module builds -- not a rule read off our own code. The table is CURATED and
#: the arms are chosen so the two readers give the same verdict; the general
#: relation between them is one-way and is measured by the sweep below, not
#: here.
def _encoded(count: int, index: int) -> bytes:
return bytes((count, index))
END_OF_LINE = bytes((0, 0))
END_OF_BITMAP = bytes((0, 1))
CURSOR_CASES: tuple[tuple[str, bytes, bool], ...] = (
("end-of-bitmap before one pixel is decoded", END_OF_BITMAP, False),
("one 4-pixel run, then end-of-bitmap", _encoded(4, 1) + END_OF_BITMAP, False),
(
"the last row one pixel short, with no end-of-line",
(_encoded(8, 1) + END_OF_LINE) * 3 + _encoded(7, 1) + END_OF_BITMAP,
False,
),
(
"every row painted and closed",
(_encoded(8, 1) + END_OF_LINE) * 4 + END_OF_BITMAP,
True,
),
(
"the last row one pixel short, then end-of-line",
(_encoded(8, 1) + END_OF_LINE) * 3 + _encoded(7, 1) + END_OF_LINE + END_OF_BITMAP,
True,
),
(
"a delta skipping a whole row",
(_encoded(8, 1) + END_OF_LINE) * 2
+ bytes((0, 2, 0, 1))
+ _encoded(8, 1)
+ END_OF_LINE
+ END_OF_BITMAP,
True,
),
(
"the last row painted to its end, with no end-of-line",
(_encoded(8, 1) + END_OF_LINE) * 3 + _encoded(8, 1) + END_OF_BITMAP,
True,
),
("the shipped fixture, which uses a delta", RLE8_STREAM, True),
# One arm per CLAUSE of the cursor rule. Until 2026-09-19 the table had
# none for the ROW clause: every short stream above stops on the last row,
# so `y < height - 1` could be off by a row (PM's P8, `height - 2`) and
# survive 51 tests. This arm stops one row early with that row complete,
# which is the only shape that separates the two readings.
(
"the row before the last painted to its end, then end-of-bitmap",
(_encoded(8, 1) + END_OF_LINE) * 2 + _encoded(8, 1) + END_OF_BITMAP,
False,
),
# The cursor can REACH the end of the frame without a pixel being painted,
# and the four arms below are that class. An end-of-line escape at column 0
# does not close a row it started -- it claims a whole row nothing wrote --
# and an independent decoder does not follow it there.
("four end-of-line escapes and not one pixel painted", END_OF_LINE * 4 + END_OF_BITMAP, False),
(
"three rows painted and closed, then a bare end-of-line",
(_encoded(8, 1) + END_OF_LINE) * 3 + END_OF_LINE + END_OF_BITMAP,
False,
),
(
"three end-of-line escapes, then the last row painted",
END_OF_LINE * 3 + _encoded(8, 1) + END_OF_BITMAP,
False,
),
(
"one pixel painted, then four end-of-line escapes",
_encoded(1, 1) + END_OF_LINE * 4 + END_OF_BITMAP,
False,
),
)
def test_a_stream_that_stops_before_the_frame_is_refused() -> None:
"""A TERMINATOR IS NOT A COVERAGE PROOF, and until now this file said it was.
Measured by PM 2026-09-19: a stream whose FIRST two bytes are the
end-of-bitmap escape was carried, with 32 of 32 pixels never decoded and
all of them index 0, while an independent decoder refuses the same file
outright. The terminator rule -- added the round before to close a
truncated stream -- asks only that the stream SAY it is finished, and a
stream can say so anywhere.
THE LINE IS THE CURSOR, NOT THE PIXELS, and it is read off an independent
decoder rather than chosen. A DELTA escape states its skip as a distance,
and the pixels it passes over keep index 0 in every decoder; an
END-OF-LINE escape states only "the rest of this row", so an end-of-line
at column 0 claims a row nothing wrote, and the independent decoder does
not follow it there. That is not a refinement of the previous round's
sentence, it is a correction of it: the sentence said every decoder agreed
on an end-of-line skip, and Pillow refuses four of the twelve streams
above on exactly that construction.
Measured on the twelve streams above, Pillow reads the five whose cursor
reaches the end of the frame through painted rows and stated deltas, and
refuses the seven whose does not -- including one short by a single pixel,
one stopping a whole row early, and four that move the cursor with an
end-of-line escape that closed no row.
A PIXEL coverage count would be a different rule and a wrong one: it
refuses the delta the format defines, and 0 of the 25 RLE8 BMPs
that delivery ships would be affected either way (25 of 25 paint every
pixel, 25 of 25 reach the end of the frame, 0 of 25 use a delta), so the
corpus cannot choose between them. The independent decoder can.
"""
for label, stream, reaches in CURSOR_CASES:
data = bmp_rle8(8, 4, PALETTE, stream)
if reaches:
carried = read_image(data, name="hel.bmp")
assert carried.media_type == "image/png", label
continue
with pytest.raises(ExtractionError) as caught:
read_image(data, name="kort.bmp")
assert caught.value.code == CODE_SAMPLES_INVALID, f"{label}: {caught.value.code}"
def test_the_independent_decoder_draws_the_same_line() -> None:
"""The eight cases above, judged by a decoder this package does not own.
Without this the table is our own rule restated, and a rule that agrees
with itself measures nothing. Pillow is optional here (transitive under
`pdfplumber`), so the guard above runs everywhere and this one runs where
the extra is installed.
"""
Image = pytest.importorskip("PIL.Image")
for label, stream, reaches in CURSOR_CASES:
data = bmp_rle8(8, 4, PALETTE, stream)
try:
picture = Image.open(io.BytesIO(data))
picture.load()
read = True
except Exception:
read = False
assert read is reaches, f"{label}: the independent decoder disagrees with the table"
def test_a_delta_may_not_carry_the_cursor_out_of_its_row() -> None:
"""A delta states a distance, and a distance past the row end states nothing.
The format puts the delta's horizontal offset INSIDE the line; there is no
defined wrap. This reader's cursor keeps `x` past the row end and paints
nothing more on that row, while a flat decoder rolls the same offset into
the next row and paints there -- two different pictures from one stream.
THIS IS THE ONE DIRECTION THIS PACKAGE ALLOWS: it refuses a file the
independent decoder reads. Pillow carries this stream (its flat buffer
reaches 32 bytes); we do not, because we cannot say which of the two
pictures is the file's. The reverse direction -- carrying a file the
independent decoder refuses -- is the defect, and the sweep below measures
it at 0.
Cost measured before it was written: over every BMP reachable on this
machine (11 441 files scanned across the four raw standard deliveries and
the K2 reference corpus, 25 BMPs, all RLE8, all from one publisher), 0 of
25 use a delta escape at all, and 0 of 25 use an end-of-line at column 0.
Both clauses cost nothing measured, and the clause that WOULD have cost
something -- refusing a run that overruns its row, which closes the last
disagreement class -- is not here: 15 of those 25 files overrun a row, and
refusing them would drop 15 real figures and move a pinned bundle's bytes.
"""
stream = _encoded(4, 1) + bytes((0, 2, 8, 3)) + END_OF_BITMAP
with pytest.raises(ExtractionError) as caught:
read_image(bmp_rle8(8, 4, PALETTE, stream), name="delta.bmp")
assert caught.value.code == CODE_SAMPLES_INVALID, caught.value.code
# The known-positive on the same shape: a delta that stays inside its row
# is the shipped fixture's own construction and is still carried.
inside = _encoded(4, 1) + bytes((0, 2, 4, 3)) + END_OF_BITMAP
assert read_image(bmp_rle8(8, 4, PALETTE, inside), name="delta.bmp").media_type == "image/png"
def _overruns_a_row(stream: bytes, width: int) -> bool:
"""Does this stream paint a run that runs off the end of its row?
Written here rather than imported: it is the shape the sweep below is
allowed to disagree on, so it has to be stated independently of the reader
it judges.
"""
position, x = 0, 0
while position + 1 < len(stream):
count, value = stream[position], stream[position + 1]
position += 2
if count:
if x + count > width:
return True
x += count
elif value == 0:
x = 0
elif value == 1:
break
elif value == 2:
x += stream[position]
position += 2
else:
if x + value > width:
return True
x += value
position += value + (value & 1)
return False
def test_no_stream_this_reader_carries_is_one_the_independent_decoder_refuses() -> None:
"""The class, swept, rather than the three constructions someone thought of.
Two rounds closed this defect one construction at a time -- a truncated
stream, then an immediate end-of-bitmap -- and each time the class stayed
open one step down, because a hand-picked table can only hold the shapes
its author imagined. This enumerates every opcode sequence of length 1 to 4
over twelve opcodes on a 4x3 frame (22 620 streams: encoded runs that fit
and runs that do not, absolute blocks, end-of-line, five deltas) and asks
one question of each: if THIS reader carries it, does the independent
decoder carry it too, and draw the same picture?
Measured 2026-09-19, before the rule was written: 703 of 22 620 streams
were carried here and refused by Pillow, and 1 492 more were drawn
differently. With the end-of-line and delta clauses: 0 and 32. Every one
of the 32 is a run or absolute block that OVERRUNS its row, which this
reader clips at the row end (as the format says) and Pillow spills into the
next row; refusing those costs 15 of the 25 real files, so the residual is
stated here and not closed. The direction that matters is closed: nothing
this package carries is a file the other decoder will not read.
"""
Image = pytest.importorskip("PIL.Image")
width, height = 4, 3
ops = (
_encoded(1, 1),
_encoded(2, 2),
_encoded(4, 1),
_encoded(5, 3),
END_OF_LINE,
bytes((0, 2, 0, 1)),
bytes((0, 2, 1, 0)),
bytes((0, 2, 2, 1)),
bytes((0, 2, 0, 3)),
bytes((0, 2, 4, 0)),
bytes((0, 3, 1, 2, 3, 0)),
bytes((0, 4, 3, 2, 1, 2)),
)
carried_and_refused: list[bytes] = []
differing: list[bytes] = []
both_carry = both_refuse = 0
streams = 0
for length in (1, 2, 3, 4):
for combo in itertools.product(ops, repeat=length):
stream = b"".join(combo) + END_OF_BITMAP
data = bmp_rle8(width, height, PALETTE, stream)
streams += 1
try:
ours = read_image(data, name="sveip.bmp").data
mine: bytes | None = png_rgb(ours)[1]
except ExtractionError:
mine = None
try:
picture = Image.open(io.BytesIO(data))
picture.load()
theirs: bytes | None = picture.convert("RGB").tobytes()
except Exception:
theirs = None
if mine is not None and theirs is not None:
if mine == theirs:
both_carry += 1
else:
differing.append(stream)
elif mine is not None:
carried_and_refused.append(stream)
elif theirs is None:
both_refuse += 1
assert streams == 22620, streams
# Non-vacuity first: a sweep where nothing is carried proves nothing.
assert both_carry > 1000, both_carry
assert both_refuse > 1000, both_refuse
assert carried_and_refused == [], (
f"{len(carried_and_refused)} of {streams} streams are carried here and refused by the "
f"independent decoder, e.g. {carried_and_refused[0]!r}"
)
unexplained = [s for s in differing if not _overruns_a_row(s, width)]
assert unexplained == [], (
f"{len(unexplained)} of {streams} streams are drawn differently for a reason this file "
f"does not state, e.g. {unexplained[0]!r}"
)
def test_the_uncompressed_path_refuses_the_same_shape() -> None:
"""The twin this rule was aligned to, so the two BMP paths cannot drift."""
short = bmp_24(RGB_ROWS)[:-30]
with pytest.raises(ExtractionError) as caught:
read_image(short, name="kort.bmp")
assert caught.value.code == CODE_SAMPLES_INVALID, caught.value.code
# --- 8. the lossless guard, with no optional dependency in sight ------------
def png_rgb(data: bytes) -> tuple[tuple[int, int], bytes]:
"""An INDEPENDENT PNG decoding using the stdlib alone: `zlib` and the five
filters of PNG SS 9.2, with the palette applied.
Guard 2 above decodes through Pillow, which this package neither uses nor
ships -- the right independence, and it is an OPTIONAL dependency here
(transitive under `pdfplumber` in `[extract]`). Measured 2026-09-19 on a
core install: 4 of the 13 guards in this file, the lossless one among them,
were SKIPPED, so a plain `pip install llm-ingestion-okf` never measured the
property the round is named after. This arm always runs.
It is independent in the way that matters: the package writes a PNG by
compressing rows it filtered, and this reads one by reversing filters it
decompressed. The expected pixels come from `rle8_indices` and from
`RGB_ROWS`, both written in this file.
"""
import zlib
assert data[:8] == b"\x89PNG\r\n\x1a\n"
position, idat, palette, header = 8, b"", b"", None
while position < len(data):
length = int.from_bytes(data[position : position + 4], "big")
kind = data[position + 4 : position + 8]
payload = data[position + 8 : position + 8 + length]
position += 12 + length
if kind == b"IHDR":
header = struct.unpack(">IIBBBBB", payload)
elif kind == b"PLTE":
palette = payload
elif kind == b"IDAT":
idat += payload
elif kind == b"IEND":
break
assert header is not None
width, height, depth, colour = header[0], header[1], header[2], header[3]
assert depth == 8, depth
channels = {0: 1, 2: 3, 3: 1}[colour]
raw = zlib.decompress(idat)
stride = width * channels
out, previous, at = bytearray(), bytearray(stride), 0
for _ in range(height):
kind_byte = raw[at]
line = bytearray(raw[at + 1 : at + 1 + stride])
at += 1 + stride
for i in range(stride):
left = line[i - channels] if i >= channels else 0
up = previous[i]
upleft = previous[i - channels] if i >= channels else 0
if kind_byte == 1:
line[i] = (line[i] + left) & 0xFF
elif kind_byte == 2:
line[i] = (line[i] + up) & 0xFF
elif kind_byte == 3:
line[i] = (line[i] + (left + up) // 2) & 0xFF
elif kind_byte == 4:
estimate = left + up - upleft
distances = (abs(estimate - left), abs(estimate - up), abs(estimate - upleft))
nearest = (left, up, upleft)[distances.index(min(distances))]
line[i] = (line[i] + nearest) & 0xFF
out += line
previous = line
if colour == 3:
return (width, height), bytes(b for i in out for b in palette[i * 3 : i * 3 + 3])
return (width, height), bytes(out)
def test_the_conversion_is_lossless_measured_with_the_stdlib_alone() -> None:
"""Eval point 2 again, on a path with no optional dependency on either side.
The package encodes a PNG with `zlib`; this decodes one with `zlib` and the
filter definitions, and compares against pixels written out in THIS file.
Pillow appears nowhere, so the guard measures on a core install.
"""
indices = rle8_indices(RLE8_STREAM, 8, 4)
assert indices == bytes(index for row in EXPECTED_INDICES for index in row)
expected = b"".join(PALETTE[index * 3 : index * 3 + 3] for index in indices)
assert expected == expected_rgb()
converted = read_image(bmp_rle8(8, 4, PALETTE, RLE8_STREAM), name="figur-rle8.bmp")
assert converted.media_type == "image/png"
assert png_rgb(converted.data) == ((8, 4), expected)
flat = read_image(bmp_24(RGB_ROWS), name="figur-flat.bmp")
assert flat.media_type == "image/png"
flat_expected = b"".join(bytes(pixel) for row in RGB_ROWS for pixel in row)
assert png_rgb(flat.data) == ((5, 3), flat_expected)
# The negative control: without it, two broken decoders could agree.
assert png_rgb(converted.data)[1] != flat_expected
def test_the_bundle_holds_those_pixels_measured_with_the_stdlib_alone(tmp_path: Path) -> None:
"""The same claim about what the BUILD wrote, not about `read_image`.
Pinned exactly like guard 2 and with no importorskip for Pillow: for each
source, exactly one carried asset decodes to its pixels, and the holder
has to be in the viewable set -- a BMP carried verbatim would hold its own
pixels trivially.
"""
pytest.importorskip("llm_ingestion_guard")
bundle = tmp_path / "bundle"
assert _build(_inbox(tmp_path), bundle) == 0
indices = rle8_indices(RLE8_STREAM, 8, 4)
sources = {
"rle8": ((8, 4), b"".join(PALETTE[i * 3 : i * 3 + 3] for i in indices)),
"flat24": ((5, 3), b"".join(bytes(pixel) for row in RGB_ROWS for pixel in row)),
}
carried = [path.read_bytes() for path in _assets(bundle)]
assert len(carried) == 2, carried
for label, expected in sources.items():
hits = [
data
for data in carried
if (sniff(data) or (None,))[0] in VIEWABLE and png_rgb(data) == expected
]
assert len(hits) == 1, f"{label}: {len(hits)} carried assets hold these pixels"
def test_the_stdlib_arm_runs_with_Pillow_unimportable() -> None:
"""SHOWN, not asserted: the two guards above re-run with `PIL` blocked.
"Runs on a core install" is a claim about an environment, and this file
is usually run in one where `[extract]` -- and therefore Pillow, under
`pdfplumber` -- is installed. So the environment is made here: a finder
that raises for every `PIL` name, with a known-positive proving it fires
before the guards are re-run under it.
"""
import importlib
import sys
class NoPillow:
def find_spec(self, name: str, path: object = None, target: object = None) -> None:
if name == "PIL" or name.startswith("PIL."):
raise ImportError(f"{name} is blocked for this guard")
return None
purged = {name: module for name, module in sys.modules.items() if name.split(".")[0] == "PIL"}
for name in purged:
del sys.modules[name]
sys.meta_path.insert(0, NoPillow())
try:
# KNOWN-POSITIVE: the blocker blocks. Without this the guard below
# would pass in an environment where Pillow was simply present.
with pytest.raises(ImportError):
importlib.import_module("PIL.Image")
test_the_conversion_is_lossless_measured_with_the_stdlib_alone()
indices = rle8_indices(RLE8_STREAM, 8, 4)
expected = b"".join(PALETTE[index * 3 : index * 3 + 3] for index in indices)
assert png_rgb(read_image(bmp_rle8(8, 4, PALETTE, RLE8_STREAM), name="x.bmp").data) == (
(8, 4),
expected,
)
finally:
sys.meta_path.pop(0)
sys.modules.update(purged)