llm-ingestion-okf/docs/2026-09-18-utgangsbudsjett-per-ledd.md
Kjell Tore Guttormsen 3b3b8ae0ca
fix(assets): budget every link by what its decoder COSTS (0.10.1)
Round 3 of the 0.10.1 review, and the finding is the pattern the three rounds
share: each bound an OUTPUT, and the bomb stepped one link along. The
declared size, then the first `FlateDecode`, then every `FlateDecode` -- and
then a link this package had documented as safe.

`ASCII85Decode` was classed as bounded "by its own input because it shrinks".
It quadruples: `z` is the shorthand for four zero bytes. And the output was
never the cost -- `base64.a85decode` appends one 4-byte object per group to a
list, about a hundred bytes of memory per byte of INPUT (101.4x at 1 MiB,
96.1x at 4 MiB, 94.5x at 16 MiB on CPython 3.14).

Paired subprocesses, idle machine, both sides from PINNED trees, the document
built once by a third process and read from a file because `ru_maxrss` never
falls and `b"z" * 64 MiB` alone costs 171 MB:

  [/Fl /A85]      z x 32 Mi  33 475 B   CARRIED 3 261 599 744 -> too_large 42 070 016
  [/Fl /A85]      z x 64 Mi  66 090 B   CARRIED 6 461 558 784 -> too_large 40 280 064
  [/A85]          z x  8 Mi   8.4 MB    CARRIED   933 085 184 -> too_large 62 484 480
  [/Fl /A85 /Fl]  z x 32 Mi  33 488 B  samples_invalid 3 519 180 800 -> too_large 43 438 080

The picture was CARRIED in three of the four: not a bound that fired late, no
bound at all. Doubling the `z` run trebles the old cost and leaves the new one
where it was.

WHY THIS FORM. `assets.MAX_FILTER_DECODE_BYTES` (512 MiB) is what decoding ONE
link may cost -- a separate number from `MAX_IMAGE_BYTES`, because that one
bounds the picture and this one bounds producing it. `FlateDecode` is measured
as it is paid; every other permitted filter carries a MEASURED cost ratio
(`assets.PDF_FILTER_COST_RATIO`) checked against its input BEFORE its decoder
is called, since those decoders take a whole string and return a whole string.
A filter with no ratio is refused unread. The budget TRAVELS: a deflate link
is inflated under the smaller of the picture's bound and what the next link's
decoder may be handed, or `[/Fl /A85]` pays 256 MiB for a refusal.

A chunked ASCII85 decoder written here was the alternative and was FELLED: it
would bound `_check_stream_cost` and not the run, because `stream.get_data()`
decodes the whole chain again with pdfminer's own decoder, and it would make
this package rather than pdfminer the authority on an image's bytes. The cap
is the only number that bounds that. `resource.setrlimit(RLIMIT_AS)` was
MEASURED before anything was built on it, as the order required, and is not
usable: Darwin 26.6.2 raises `ValueError: current limit exceeds maximum limit`
and does not enforce it. No child-process cap exists.

THE CAP IS READ OFF THE CORPORA, the posture `MAX_IMAGE_PIXELS` has: over the
9 668 image objects of the 77 PDFs on this machine, 16 decode through an
ASCII85 link and the largest input to one is 450 739 bytes, against a cap of
about 5.0 MB.

A PROPERTY TEST REPLACES THE LIST OF KNOWN SHAPES: every chain of length 1-3
over the ten filters pdfminer decodes, 1 110 of 1 110, both payload fills,
each delivered under the bound or refused with a published code and never paid
for on the way (`tracemalloc`, which counts allocations and is not disturbed
by load). Known-positive beside it: 258 of 258 chains over the permitted
filters still carry a small image.

MAJOR -- the backstop had no test. `check_payload` at the end of
`_check_stream_cost` could be deleted with the whole suite green, because the
second one after `get_data()` gives the same code one step later. The two
differ in whether the payment was made, so the test asserts `get_data` was
never called.

10 OF 10 MUTANTS KILLED, control green, each killer named in the report. Four
survived a first pass and two tests exist because of it.

NOT ONE PICTURE CHANGES HANDS, MEASURED BY NAME: `_pdf_images` over every PDF
on this machine from both pinned trees -- 9 306 -> 9 306 carried over 77
files, 50 -> 50 on R761, 0 of 78 files moving a count and 0 moving a code.
R761 also settles a question raised while this order was open: 50 objects, 29
[/DCTDecode], 21 [/FlateDecode], 0 ASCII85 links -- so round 2's count of 580
`[/FlateDecode /ASCII85Decode]` objects is reproducible from nothing on this
machine. It changes no decision; a bomb shape does not need a corpus.

Version stays 0.10.1, no tag. README, CHANGELOG, CLAUDE.md and errors.py
corrected TO what the code does; the round-2 report carries a correction block
rather than a rewrite. Report:
docs/2026-09-18-utgangsbudsjett-per-ledd.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 19:00:38 +02:00

14 KiB
Raw Blame History

A budget per link, and a property over every chain

A PM checkpoint of 0c3c490 — the commit that bound every link of a PDF filter chain — read the fix and found the bound still reachable. This is the third round of the same review, and the third time a bound was put on an output and the bomb stepped one link along.

The two rounds before it are docs/2026-09-17-bildestien-0-10-1.md, docs/2026-09-18-bildestien-holder-0-10-1.md and docs/2026-09-18-filterkjeden-og-backstoppen.md, which carries a correction block pointing here.

The finding is the pattern, not the filter

round what was bound where the bomb moved
1 the size the dictionary DECLARES into the stream
2 the first FlateDecode link into the second link
3 every FlateDecode link into a link documented as safe

Round 3 shipped a three-class split in which ASCII85Decode and ASCIIHexDecode were "bounded by their own input because they shrink". That sentence is false in two independent ways:

  • ASCII85 does not shrink. z is its shorthand for four zero bytes, so one input byte becomes four. The output ratio is 4, not 0.8.
  • The output is not the cost. base64.a85decode appends one 4-byte object per group to a list, so a run of z costs about a hundred bytes of memory per byte of INPUT. Measured on CPython 3.14: 101.4x at 1 MiB of input, 96.1x at 4 MiB, 94.5x at 16 MiB. Nothing in the output size says so.

The second point is the general one, and it is why this round does not add a fourth class. A bound on what a link OUTPUTS is not a bound on what producing it COSTS, and every round of this review has been an instance of that.

BLOCKER — reproduced

Measured in paired subprocesses on an idle machine, both sides from pinned trees (the before side from git archive of 0c3c490, the after side from a snapshot of the working tree), each printing the module's own __file__ as a control. The document is built ONCE by a third process and read from a file: ru_maxrss is a high-water mark that never falls, so a child that builds its own fixture reports the fixture — b"z" * (64 MiB) alone costs 171 MB, more than the bounded extraction it would be measuring. This tripped the first attempt at these figures.

chain file before after
[/FlateDecode /ASCII85Decode], z × 32 Mi 33 475 B carried, 3 261 599 744 B asset_too_large, 42 070 016 B
[/FlateDecode /ASCII85Decode], z × 64 Mi 66 090 B carried, 6 461 558 784 B asset_too_large, 40 280 064 B
[/ASCII85Decode], z × 8 Mi 8 389 449 B carried, 933 085 184 B asset_too_large, 62 484 480 B
[/Fl /A85 /Fl], z × 32 Mi 33 488 B asset_samples_invalid, 3 519 180 800 B asset_too_large, 43 438 080 B

Two things to read off it. The picture was carried in three of the four rows — this was not a bound that fired late, it was no bound at all. And the bounded cost does not follow the bomb: doubling the run of z takes the old cost from 3.26 GB to 6.46 GB and the new one from 42.1 MB to 40.3 MB.

assets.MAX_FILTER_DECODE_BYTES (512 MiB) is what decoding ONE link may cost. It is a separate number from MAX_IMAGE_BYTES (256 MiB) on purpose: that one bounds the picture this package will carry, this one bounds what producing it costs on the way.

assets.PDF_FILTER_COST_RATIO gives each permitted filter a measured worst-case peak memory per byte of input. None means the decoder is driven a chunk at a time here, so the cost is measured as it is paid — today that is FlateDecode alone.

filter cost ratio measured
FlateDecode — driven a chunk at a time (assets._inflate)
ASCII85Decode 104 101.4x / 96.1x / 94.5x at 1 / 4 / 16 MiB of z
ASCIIHexDecode 2 1.5x at 16 MiB
DCTDecode, JPXDecode, JBIG2Decode 1 pass-through in pdfminer

Everything else — LZWDecode, RunLengthDecode, CCITTFaxDecode, /Crypt, anything written later — has no measured ratio and is refused UNREAD with asset_pdf_unbounded, before any link in front of it is decoded. That is the same decision corpus.resolve_gate takes for an unknown gate name: a fallback reproduces the defect with an extra step.

The budget travels down the chain. A FlateDecode link's output is the next link's input, so it is inflated under a limit that is the smaller of MAX_IMAGE_BYTES and what the next link's decoder may be handed (assets.inflate_limit_for). Without that, [/FlateDecode /ASCII85Decode] would inflate 256 MiB of z before the link behind it was asked anything.

Why an input cap and not a bounded ASCII85 decoder

A chunked ASCII85 decoder written here would bound _check_stream_cost, and it would not bound the run: pdfminer decodes the whole chain again in stream.get_data(), with its own unbounded decoder, and that is where the memory is actually spent. The only number that bounds that is the size of the input this package allows the link to be handed. Writing our own decoder would also make this package, rather than pdfminer, the authority on what an image's bytes are.

Why not a hard backstop in a child process

The order asked for resource.setrlimit(RLIMIT_AS) to be MEASURED before anything was built on it. It was, and it is not available here: on this machine (Darwin 26.6.2, CPython 3.14) setrlimit(RLIMIT_AS, (256 MiB, hard)) raises ValueError: current limit exceeds maximum limit — a fresh CPython process has already reserved far more address space than the cap, and the hard limit reads as RLIM_INFINITY. The documented behaviour agrees: Darwin does not enforce RLIMIT_AS, RLIMIT_DATA or RLIMIT_RSS the way Linux does. So no child-process memory cap was built, and the bound is the per-link budget alone.

The cap is read off the corpora

The cap that falls out for ASCII85Decode is 512 MiB / 104 ≈ 5.0 MB of input. Measured 2026-09-18 over the 9 668 image objects of the 77 PDFs on this machine (enumerated through pdfminer's own page walk):

chain objects
[/FlateDecode] 6 235
[/DCTDecode] 2 459
[/FlateDecode /DCTDecode] 596
[/Fl] 296
unfiltered 42
[/ASCII85Decode /FlateDecode] 16
[/JPXDecode] 16
[/CCITTFaxDecode] 8

16 objects decode through an ASCII85Decode link, and the largest input any of them is handed is 450 739 bytes — more than ten times under the cap. That is the posture MAX_IMAGE_PIXELS has: a number read off the corpora and standing an order of magnitude above anything measured, so the bound costs no picture anybody has.

Two corrections to earlier published counts fall out of this table, and both are about ENUMERATION rather than about the documents. The round-2 report counted 5 142 objects over 78 PDFs and 580 behind [/FlateDecode /ASCII85Decode]. This walk finds 9 668 objects over 77 files and 0 behind that chain. The denominators differ because the two walks are different (this one recurses into LTFigure; the file R761 is not in this listing), so neither number is wrong about a document — but a chain count is only readable beside the walk that produced it, and the [/Fl /A85] group is not one this machine's corpora hold. The bomb that shape carries is real regardless: a document does not have to exist in a corpus to be handed to okf build.

The property test

tests/test_asset_limits.py::test_no_chain_of_up_to_three_filters_is_carried_over_the_bound generates every chain of length 1–3 over the ten filters pdfminer decodes — K = 1 110 — twice, once with a payload of zeros (the amplifying case at both ends: it deflates to nothing and a85encodes to a run of z) and once with a repeated non-zero byte. Each chain's stream is built by encoding the payload BACKWARDS through the chain, so every chain over the permitted filters is a valid document rather than a rejection by accident.

The requirement is one sentence: the picture is either delivered with its bytes under the bound, or refused with a code in the published vocabulary; never carried over the bound, and never paid for on the way. tracemalloc measures the paying, because it counts Python's own allocations — which is exactly where a85decode's cost lives, and unlike ru_maxrss it is not disturbed by other work on the machine.

Result: 1 110 of 1 110 chains pass, both payloads — k = K.

Beside it, test_every_bounded_chain_still_carries_a_small_image runs the 258 chains over the six permitted filters with a 64-byte image and requires none of them to be refused. A rule that refuses everything passes the property alone; it does not pass this.

And test_the_ascii85_cost_ratio_is_not_below_the_one_this_package_measured re-measures the ratio the budget rests on, in a subprocess, at two input sizes. If CPython ever changes a85decode so that it costs more, the constant is too generous and this says so before a corpus does.

MAJOR — the backstop had no test

check_payload(size, name=name) at the END of _check_stream_cost could be deleted with the whole suite still passing. It is what refuses a stream no filter in the chain expands — an unfiltered one, or one behind DCTDecode — and the SECOND check_payload, after get_data(), produces the same code and the same words one step later. A test that reads the code cannot tell the two apart.

What separates them is whether the payment was made, so the test asserts get_data was never called.

Mutants

Ten mutations, one line each, run in a scratch clone with the unmutated copy run FIRST as a control. The peak-RSS subprocess tests are deselected for these runs — they measure a high-water mark and the machine was running a corpus census — so what kills a mutant here is the property test, the code vocabulary, or an assertion about which check fired.

mutation one line killed by
backstop-deleted check_payload at the end of _check_stream_cost removed test_the_stream_bound_refuses_before_get_data_is_ever_called
cost-check-deleted the per-link check_filter_cost call removed the property test, both fills, + [/ASCII85Decode]
ascii85-budget-removed that filter's ratio set to None, so it gets no budget 9 tests, including both ratio measurements and the corpus cap
ascii85-ratio-is-one the ratio set to 1 instead of the measured 104 6 tests, including both ratio measurements
budget-does-not-travel inflate_limit_for returns MAX_IMAGE_BYTES always test_the_budget_travels_to_the_next_link
widest-output-ignored _widest_output returns its input unchanged test_a_discarded_links_size_travels_as_the_widest_it_could_become
first-flate-not-last the discard happens at the FIRST deflate link 5 tests, including both round-2 chain bombs
unknown-filter-passes a filter with no ratio is let through instead of refused test_a_filter_the_bound_cannot_measure_..., test_asset_pdf_unbounded
budget-a-hundredfold MAX_FILTER_DECODE_BYTES multiplied by 100 4 tests, including both new bombs
cost-check-off-by-a-factor the comparison allows 1 000x the limit test_an_ascii85_link_on_its_own_is_bounded

10 of 10 killed. Four of them survived a first pass and are the reason two of the tests above exist: budget-does-not-travel and widest-output-ignored had no test at all, and budget-a-hundredfold and cost-check-off-by-a-factor were only reachable through the peak-RSS subprocess tests that first pass had deselected. A mutant that survives is a test that was missing, not a mutation that was unfair.

Cost to real documents

Measured by name, not by total: _pdf_images run over every PDF on this machine from each of the two pinned trees, each printing the module file it loaded as a control, and the per-file counts compared.

corpus files carried before carried after files whose count moved
~/corpora + tests/fixtures 77 9 306 9 306 0
R761 Prosesskoden:2025 1 50 50 0
both 78 9 356 9 356 0

The rejection codes are identical too — asset_pdf_unsupported 314, asset_pdf_unbounded 8, asset_samples_invalid 40 on both sides, and 0 files where any per-file code count moved. Not one picture changes hands.

That is what the cap being read off the corpora buys: the only new refusal is check_filter_cost, and the largest ASCII85Decode input any of these documents holds is 450 739 bytes against a cap of about 5.0 MB.

The two trees are git archive of 0c3c490 and a snapshot of the working tree taken before the prose edits; diff -r between that snapshot and the committed tree touches docstrings and comments only, so what was measured is what shipped.

R761 also settles the [/FlateDecode /ASCII85Decode] question the PM raised while this order was open: 50 image objects, 29 [/DCTDecode] and 21 [/FlateDecode], and 0 ASCII85 links. So that chain is in neither the 77 corpus files nor R761, and the round-2 report's count of 580 is not reproducible from anything on this machine. It changes no decision here — the bomb that shape carries does not need a corpus to exist in — but a published count that cannot be reproduced should not stay unqualified.

What is still not bounded

Unchanged from round 2, and stated rather than implied: a stream something else has already decoded (_pdf_stream_bytes returns None), where the memory is spent before this package is asked. check_payload after get_data() COUNTS it — a counted refusal, not a bounded one.