fix(assets): bound every link of the filter chain, and cover the backstop
The two findings of the 18.09 PM checkpoint of `0f308c1`. Red tests landed first in `3b587ea`; this is what turns them green. BLOCKER -- `_check_inflated` read `filters[0]`, measured that one link and returned, which is not a bound: a PDF decodes a stream through a LIST of filters. Measured in paired subprocesses from two pinned trees, idle machine: [/FlateDecode] 400 MB 408 516 B 59 232 256 -> 62 017 536 B [/FlateDecode x2] 400 MB 1 636 B 886 554 624 -> 52 367 360 B [/FlateDecode x3] 400 MB 1 070 B 889 393 152 -> 61 390 848 B [/FlateDecode x2] 1,2 GB 2 927 B 2 567 204 864 -> 60 403 712 B 542 000x the file at two links, and the picture WAS refused at the end -- by `check_payload` after `get_data()`, once the memory was spent. The single-link row is the control and does not move. It also left the 16 corpus objects behind an `[/ASCII85Decode /FlateDecode]` chain unmeasured, since `filters[0]` is not `FlateDecode` there. `_check_stream_cost` walks every link. THREE CLASSES and no fourth (`extract.bounded_pdf_filters`, pinned by a test): `FlateDecode` MEASURED, a link with another expanding link behind it inflated under the same bound and handed on; `ASCII85Decode`/`ASCIIHexDecode` bounded by their own input because they SHRINK; `DCTDecode`/`JPXDecode`/`JBIG2Decode` PASS THROUGH. Everything else -- `LZWDecode`, `RunLengthDecode`, `CCITTFaxDecode`, `/Crypt`, anything written later -- is refused UNREAD with a new code `asset_pdf_unbounded`, decided before the FIRST link is decoded so a document cannot make the run pay for the links in front of the one we cannot bound. An encrypted stream is deciphered and then measured, where `stream.decipher is not None` used to return unmeasured; 0 of 5 142 objects here are in an encrypted document, which is why nothing caught it. NOT ONE PICTURE CHANGES HANDS, AND IT IS MEASURED BY NAME. Every PDF on this machine -- 78 documents, K2 in both trinn1 and trinn2, the shipped fixtures and R761 -- run through `_pdf_images` page by page from both pinned trees: images carried 9 356 -> 9 356 documents losing one 0 of 78 documents gaining one 0 of 78 asset_pdf_unsupported 322 -> 314 asset_pdf_unbounded 0 -> 8 The 8 are the 4 `CCITTFaxDecode` stencil masks (`/ImageMask true`, `/BitsPerComponent 1`), counted twice because trinn1 and trinn2 hold the same document. They were refused before and are refused now, one step earlier and under a code that says why. MAJOR -- `check_payload(len(data))` after `get_data()` is the counted refusal four documentation surfaces point at, and deleting exactly that line passed all 2 132 tests. It is reachable through a stream pdfminer has ALREADY decoded (`decode()` sets `rawdata` to `None`), which is now the ONLY case outside the bound and has a test. Eight mutations, one line each, every one DEAD, with the unmutated tree run first as the control: first-link-only, loop dropped, inequality reversed, encrypted skipped, backstop deleted, unknown filter passed through, intermediate link not carried forward, whole check removed. `tools/okf_accounting_gate.py` gains one line, the new code in `REJECTION_CODES` -- what a rejection code requires and nothing more. Gate unchanged: exit 1, GATE RED rows 2, 3, 6. Version stays 0.10.1, untagged. Suite after `git add` against a clean tree: `uv run pytest -q` -> 2152 passed, 1 skipped (226 s). ruff, ruff format --check, mypy --strict clean. Report: docs/2026-09-18-filterkjeden-og-backstoppen.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
3b587ea567
commit
0c3c4904ee
11 changed files with 564 additions and 38 deletions
|
|
@ -750,6 +750,29 @@ def test_asset_too_large() -> None:
|
|||
assert excinfo.value.code == "asset_too_large"
|
||||
|
||||
|
||||
def test_asset_pdf_unbounded() -> None:
|
||||
"""A filter whose output cannot be measured before it is produced.
|
||||
|
||||
Its own code because it says something different from `asset_too_large`:
|
||||
that one reports a measurement that came out over the bound, this one
|
||||
reports that no measurement was possible, so the picture was refused
|
||||
unread. Reached through the real PDF path rather than a helper, because
|
||||
the decision is which FILTER the stream declares.
|
||||
"""
|
||||
from llm_ingestion_okf.extract import extract_document
|
||||
|
||||
pytest.importorskip("pdfplumber")
|
||||
sys.path.insert(0, str(Path(__file__).parent))
|
||||
from test_asset_limits import _bomb
|
||||
|
||||
extracted = extract_document(
|
||||
"lzw.pdf", _bomb(4, payload=b"\xff" * 512, filters="/LZWDecode"), assets=True
|
||||
)
|
||||
assert [item.code for item in extracted.rejected] == ["asset_pdf_unbounded"]
|
||||
registry = ExtractionError.__doc__ or ""
|
||||
assert "`asset_pdf_unbounded`" in registry
|
||||
|
||||
|
||||
def test_asset_size_invalid() -> None:
|
||||
"""A declared size that is not a size. Its own code because it says
|
||||
something different about the document than `asset_too_large` does."""
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue