fix(assets): budget every link by what its decoder COSTS (0.10.1)
Round 3 of the 0.10.1 review, and the finding is the pattern the three rounds share: each bound an OUTPUT, and the bomb stepped one link along. The declared size, then the first `FlateDecode`, then every `FlateDecode` -- and then a link this package had documented as safe. `ASCII85Decode` was classed as bounded "by its own input because it shrinks". It quadruples: `z` is the shorthand for four zero bytes. And the output was never the cost -- `base64.a85decode` appends one 4-byte object per group to a list, about a hundred bytes of memory per byte of INPUT (101.4x at 1 MiB, 96.1x at 4 MiB, 94.5x at 16 MiB on CPython 3.14). Paired subprocesses, idle machine, both sides from PINNED trees, the document built once by a third process and read from a file because `ru_maxrss` never falls and `b"z" * 64 MiB` alone costs 171 MB: [/Fl /A85] z x 32 Mi 33 475 B CARRIED 3 261 599 744 -> too_large 42 070 016 [/Fl /A85] z x 64 Mi 66 090 B CARRIED 6 461 558 784 -> too_large 40 280 064 [/A85] z x 8 Mi 8.4 MB CARRIED 933 085 184 -> too_large 62 484 480 [/Fl /A85 /Fl] z x 32 Mi 33 488 B samples_invalid 3 519 180 800 -> too_large 43 438 080 The picture was CARRIED in three of the four: not a bound that fired late, no bound at all. Doubling the `z` run trebles the old cost and leaves the new one where it was. WHY THIS FORM. `assets.MAX_FILTER_DECODE_BYTES` (512 MiB) is what decoding ONE link may cost -- a separate number from `MAX_IMAGE_BYTES`, because that one bounds the picture and this one bounds producing it. `FlateDecode` is measured as it is paid; every other permitted filter carries a MEASURED cost ratio (`assets.PDF_FILTER_COST_RATIO`) checked against its input BEFORE its decoder is called, since those decoders take a whole string and return a whole string. A filter with no ratio is refused unread. The budget TRAVELS: a deflate link is inflated under the smaller of the picture's bound and what the next link's decoder may be handed, or `[/Fl /A85]` pays 256 MiB for a refusal. A chunked ASCII85 decoder written here was the alternative and was FELLED: it would bound `_check_stream_cost` and not the run, because `stream.get_data()` decodes the whole chain again with pdfminer's own decoder, and it would make this package rather than pdfminer the authority on an image's bytes. The cap is the only number that bounds that. `resource.setrlimit(RLIMIT_AS)` was MEASURED before anything was built on it, as the order required, and is not usable: Darwin 26.6.2 raises `ValueError: current limit exceeds maximum limit` and does not enforce it. No child-process cap exists. THE CAP IS READ OFF THE CORPORA, the posture `MAX_IMAGE_PIXELS` has: over the 9 668 image objects of the 77 PDFs on this machine, 16 decode through an ASCII85 link and the largest input to one is 450 739 bytes, against a cap of about 5.0 MB. A PROPERTY TEST REPLACES THE LIST OF KNOWN SHAPES: every chain of length 1-3 over the ten filters pdfminer decodes, 1 110 of 1 110, both payload fills, each delivered under the bound or refused with a published code and never paid for on the way (`tracemalloc`, which counts allocations and is not disturbed by load). Known-positive beside it: 258 of 258 chains over the permitted filters still carry a small image. MAJOR -- the backstop had no test. `check_payload` at the end of `_check_stream_cost` could be deleted with the whole suite green, because the second one after `get_data()` gives the same code one step later. The two differ in whether the payment was made, so the test asserts `get_data` was never called. 10 OF 10 MUTANTS KILLED, control green, each killer named in the report. Four survived a first pass and two tests exist because of it. NOT ONE PICTURE CHANGES HANDS, MEASURED BY NAME: `_pdf_images` over every PDF on this machine from both pinned trees -- 9 306 -> 9 306 carried over 77 files, 50 -> 50 on R761, 0 of 78 files moving a count and 0 moving a code. R761 also settles a question raised while this order was open: 50 objects, 29 [/DCTDecode], 21 [/FlateDecode], 0 ASCII85 links -- so round 2's count of 580 `[/FlateDecode /ASCII85Decode]` objects is reproducible from nothing on this machine. It changes no decision; a bomb shape does not need a corpus. Version stays 0.10.1, no tag. README, CHANGELOG, CLAUDE.md and errors.py corrected TO what the code does; the round-2 report carries a correction block rather than a rewrite. Report: docs/2026-09-18-utgangsbudsjett-per-ledd.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
33d3269380
commit
3b3b8ae0ca
9 changed files with 784 additions and 126 deletions
|
|
@ -29,6 +29,7 @@ import collections
|
|||
import csv
|
||||
import functools
|
||||
import io
|
||||
import math
|
||||
import re
|
||||
import statistics
|
||||
import tempfile
|
||||
|
|
@ -43,12 +44,15 @@ from xml.etree import ElementTree
|
|||
from xml.etree.ElementTree import Element
|
||||
|
||||
from .assets import (
|
||||
PDF_FILTER_OUTPUT_RATIO,
|
||||
AssetRejection,
|
||||
ExtractedImage,
|
||||
check_filter_cost,
|
||||
check_payload,
|
||||
check_size,
|
||||
encode_png,
|
||||
inflate_bounded,
|
||||
inflate_limit_for,
|
||||
inflated_size,
|
||||
read_image,
|
||||
render_block,
|
||||
|
|
@ -1370,44 +1374,59 @@ def _pdf_alpha(attrs: dict[str, object], width: int, height: int) -> bytes | Non
|
|||
def bounded_pdf_filters() -> frozenset[str]:
|
||||
"""The PDF stream filters an image may be reached through, by NAME.
|
||||
|
||||
THREE CLASSES, and what separates them is whether a bound can be put on the
|
||||
output before the decoding is paid for.
|
||||
TWO CLASSES, and what separates them is HOW the cost of a link is bounded,
|
||||
never whether the link is safe. Every one of them is bounded.
|
||||
|
||||
* `FlateDecode` is MEASURED: inflated a chunk at a time with the output
|
||||
discarded, refused the moment the running total crosses the bound.
|
||||
* `ASCII85Decode` and `ASCIIHexDecode` SHRINK by construction -- five
|
||||
characters to four bytes, two to one -- so their output is bounded by
|
||||
their input, which is already in memory as part of the file. They are
|
||||
decoded here so that a `FlateDecode` BEHIND one can be measured.
|
||||
* `DCTDecode`, `JPXDecode` and `JBIG2Decode` are pass-through in pdfminer:
|
||||
it hands the compressed image on for the image reader to sniff, so the
|
||||
size does not change.
|
||||
* `FlateDecode` is MEASURED as it is paid: inflated a chunk at a time,
|
||||
refused the moment the running total crosses the bound, with the output
|
||||
discarded unless a link behind it has to be measured from those bytes.
|
||||
* `ASCII85Decode`, `ASCIIHexDecode`, `DCTDecode`, `JPXDecode` and
|
||||
`JBIG2Decode` are PREDICTED before they are paid: each carries a measured
|
||||
worst-case cost per byte of input (`assets.PDF_FILTER_COST_RATIO`), and a
|
||||
link whose input times that ratio is over the budget is refused before
|
||||
its decoder is called. Their decoders take a whole string and return a
|
||||
whole string, so there is no moment between the two at which a cost could
|
||||
be observed.
|
||||
|
||||
0.10.1 had a third class, and it was WRONG. `ASCII85Decode` and
|
||||
`ASCIIHexDecode` were called bounded "by their own input because they
|
||||
shrink". `z` is ASCII85's shorthand for four zero bytes, so that filter
|
||||
QUADRUPLES its input, and `base64.a85decode` appends one 4-byte object per
|
||||
group to a list, so it costs about a hundred bytes of memory per byte of
|
||||
input. Measured on the pinned tree, its own interpreter, idle machine: a
|
||||
33 475-byte PDF decoding through `[/FlateDecode /ASCII85Decode]` cost
|
||||
3 261 599 744 bytes of peak RSS and the picture was CARRIED. Under the
|
||||
ratios it is 42 070 016 bytes and `asset_too_large`, and at twice the run
|
||||
of `z` -- which trebled the old cost to 6 461 558 784 -- it is 40 280 064:
|
||||
the cost no longer follows the bomb.
|
||||
|
||||
EVERYTHING ELSE IS REFUSED with `asset_pdf_unbounded` before any of the
|
||||
stream is decoded -- `LZWDecode`, `RunLengthDecode`, `CCITTFaxDecode`,
|
||||
`/Crypt`, and any filter written after this one. They expand by an amount
|
||||
pdfminer will only reveal by producing the whole output, so there is no
|
||||
measuring them a chunk at a time, and decoding one to find out how big it
|
||||
is IS the failure this bound exists to stop. Refusing an unknown name
|
||||
rather than passing it through is the same decision `corpus.resolve_gate`
|
||||
takes for an unknown gate name: a fallback reproduces the defect with an
|
||||
extra step.
|
||||
`/Crypt`, and any filter written after this one. A filter with no measured
|
||||
ratio has no budget to be checked against, and decoding one to find out
|
||||
what it costs IS the failure this bound exists to stop. Refusing an unknown
|
||||
name rather than passing it through is the same decision
|
||||
`corpus.resolve_gate` takes for an unknown gate name: a fallback reproduces
|
||||
the defect with an extra step.
|
||||
|
||||
The cost is measured rather than assumed. Over the 5 142 image objects of
|
||||
the 78 PDFs on this machine (2026-09-18), the filter chains are 1 654
|
||||
`[/DCTDecode]`, 2 236 `[/FlateDecode]`, 596 `[/FlateDecode /DCTDecode]`,
|
||||
580 `[/FlateDecode /ASCII85Decode]`, 40 unfiltered, 16 `[/ASCII85Decode
|
||||
/FlateDecode]`, 16 `[/JPXDecode]` and 4 `[/CCITTFaxDecode]` -- so the
|
||||
refused class is those 4 objects, which are 1-bit stencil masks
|
||||
(`/ImageMask true`, `/BitsPerComponent 1`) and were already refused one
|
||||
step later by the encoder, twice over.
|
||||
The reach is measured rather than assumed. Over the 9 668 image objects of
|
||||
the 77 PDFs on this machine (2026-09-18, enumerated through pdfminer's own
|
||||
page walk), the filter chains are 6 235 `[/FlateDecode]`, 2 459
|
||||
`[/DCTDecode]`, 596 `[/FlateDecode /DCTDecode]`, 296 `[/Fl]`, 42
|
||||
unfiltered, 16 `[/ASCII85Decode /FlateDecode]`, 16 `[/JPXDecode]` and 8
|
||||
`[/CCITTFaxDecode]`. The refused class is those 8 objects, 1-bit stencil
|
||||
masks (`/ImageMask true`, `/BitsPerComponent 1`) already refused one step
|
||||
later by the encoder. The largest input any `ASCII85Decode` link is handed
|
||||
is 450 739 bytes, more than ten times under the budget's cap, which is why
|
||||
the cap costs no picture the corpora hold.
|
||||
"""
|
||||
return _BOUNDED_PDF_FILTERS
|
||||
|
||||
|
||||
#: The names in `bounded_pdf_filters`, as a constant the test suite pins. The
|
||||
#: docstring above is the published claim; this is what the code enforces, and
|
||||
#: `_pdf_filter_classes` is the same three classes as pdfminer literals.
|
||||
#: `_pdf_filter_names` maps every pdfminer spelling of these onto the canonical
|
||||
#: name that `assets.PDF_FILTER_COST_RATIO` budgets.
|
||||
_BOUNDED_PDF_FILTERS = frozenset(
|
||||
{
|
||||
"FlateDecode",
|
||||
|
|
@ -1420,12 +1439,14 @@ _BOUNDED_PDF_FILTERS = frozenset(
|
|||
)
|
||||
|
||||
|
||||
def _pdf_filter_classes() -> tuple[frozenset[object], frozenset[object], frozenset[object]]:
|
||||
"""The three classes as pdfminer literals: measured, shrinking, unchanged.
|
||||
def _pdf_filter_names() -> dict[object, str]:
|
||||
"""Every pdfminer literal this package bounds, mapped to its CANONICAL name.
|
||||
|
||||
Read from pdfminer rather than written out here, because a filter has more
|
||||
than one spelling (`/Fl` is `/FlateDecode`) and a set of names written by
|
||||
hand would refuse the abbreviation a real document uses.
|
||||
hand would refuse the abbreviation a real document uses. The canonical name
|
||||
is the key into `assets.PDF_FILTER_COST_RATIO`, so both spellings of a
|
||||
filter are budgeted by one measured number.
|
||||
"""
|
||||
from pdfminer.pdftypes import (
|
||||
LITERALS_ASCII85_DECODE,
|
||||
|
|
@ -1436,13 +1457,26 @@ def _pdf_filter_classes() -> tuple[frozenset[object], frozenset[object], frozens
|
|||
LITERALS_JPX_DECODE,
|
||||
)
|
||||
|
||||
return (
|
||||
frozenset(LITERALS_FLATE_DECODE),
|
||||
frozenset(LITERALS_ASCII85_DECODE) | frozenset(LITERALS_ASCIIHEX_DECODE),
|
||||
frozenset(LITERALS_DCT_DECODE)
|
||||
| frozenset(LITERALS_JPX_DECODE)
|
||||
| frozenset(LITERALS_JBIG2_DECODE),
|
||||
)
|
||||
families = {
|
||||
"FlateDecode": LITERALS_FLATE_DECODE,
|
||||
"ASCII85Decode": LITERALS_ASCII85_DECODE,
|
||||
"ASCIIHexDecode": LITERALS_ASCIIHEX_DECODE,
|
||||
"DCTDecode": LITERALS_DCT_DECODE,
|
||||
"JPXDecode": LITERALS_JPX_DECODE,
|
||||
"JBIG2Decode": LITERALS_JBIG2_DECODE,
|
||||
}
|
||||
return {
|
||||
literal: canonical
|
||||
for canonical, literals in families.items()
|
||||
for literal in literals
|
||||
if canonical in _BOUNDED_PDF_FILTERS
|
||||
}
|
||||
|
||||
|
||||
#: The two canonical names whose decoders produce fewer or more bytes than they
|
||||
#: were given, and which this package therefore has to run to learn the size of
|
||||
#: the link behind them. Everything else in the table is pass-through.
|
||||
_SHRINKING_FILTER_NAMES = frozenset({"ASCII85Decode", "ASCIIHexDecode"})
|
||||
|
||||
|
||||
def _pdf_stream_bytes(stream: object, name: str) -> bytes | None:
|
||||
|
|
@ -1504,13 +1538,30 @@ def _check_stream_cost(stream: object, name: str) -> None:
|
|||
corpus objects behind an `[/ASCII85Decode /FlateDecode]` chain unmeasured,
|
||||
because `filters[0]` is not `FlateDecode` there.
|
||||
|
||||
So every link is walked, in order. An unknown or unmeasurable one is
|
||||
refused BEFORE anything is decoded (`bounded_pdf_filters` says which, and
|
||||
why). A `FlateDecode` that is the last expanding link is measured and its
|
||||
output discarded, which is the common case and costs exactly what 0.10.1
|
||||
cost. A `FlateDecode` with another expanding link behind it is inflated
|
||||
under the same bound and handed on, so the link behind it can be measured
|
||||
too -- what is held is never more than the bound.
|
||||
AND THE COST OF A LINK, NOT THE SIZE OF ITS OUTPUT. Bounding every
|
||||
`FlateDecode` was still not a bound, because the bomb moved into a link
|
||||
0.10.1 had documented as safe: `ASCII85Decode`'s `z` is the shorthand for
|
||||
four zero bytes, and its decoder holds about a hundred bytes per byte of
|
||||
input. Measured on the pinned tree: 33 475 bytes of file cost
|
||||
3 261 599 744 bytes of peak RSS and the picture was CARRIED. Three rounds
|
||||
of this review each bound an OUTPUT and the bomb stepped one link along;
|
||||
what they had in common is that a decoder's working set is not its output.
|
||||
|
||||
So every link is walked, in order, and each is given a BUDGET
|
||||
(`assets.MAX_FILTER_DECODE_BYTES`) rather than a class:
|
||||
|
||||
* a filter with no measured cost ratio is refused BEFORE anything is
|
||||
decoded (`bounded_pdf_filters` says which, and why);
|
||||
* a `FlateDecode` is measured as it is paid, under a limit that is the
|
||||
smaller of the picture's own bound and what the NEXT link's decoder may
|
||||
be handed -- which is how the budget travels down the chain instead of
|
||||
being applied to each link in isolation;
|
||||
* the last `FlateDecode` in the chain has its output counted and thrown
|
||||
away, which is the common case and costs exactly what 0.10.1 cost; an
|
||||
earlier one is inflated under the same limit and handed on, so the link
|
||||
behind it can be measured from real bytes;
|
||||
* every other filter has its cost PREDICTED from its input size and its
|
||||
measured ratio, and is refused before its decoder is called.
|
||||
|
||||
WHAT THIS STILL DOES NOT BOUND, stated rather than implied: a stream
|
||||
something else has already decoded (`_pdf_stream_bytes` returns `None`),
|
||||
|
|
@ -1518,46 +1569,91 @@ def _check_stream_cost(stream: object, name: str) -> None:
|
|||
by `check_payload` AFTER `get_data()`, which makes it a counted refusal
|
||||
rather than a bounded one.
|
||||
"""
|
||||
flate, shrinking, pass_through = _pdf_filter_classes()
|
||||
names = _pdf_filter_names()
|
||||
|
||||
data = _pdf_stream_bytes(stream, name)
|
||||
if data is None:
|
||||
return
|
||||
filters = stream.get_filters() if hasattr(stream, "get_filters") else []
|
||||
raw_filters = stream.get_filters() if hasattr(stream, "get_filters") else []
|
||||
# THE WHOLE CHAIN IS READ BEFORE THE FIRST LINK IS DECODED. A filter this
|
||||
# package cannot bound must be refused without having paid for the links in
|
||||
# front of it, which is only possible if the refusal is decided up front.
|
||||
for literal, _params in filters:
|
||||
if literal not in flate and literal not in shrinking and literal not in pass_through:
|
||||
chain: list[tuple[str, object]] = []
|
||||
for literal, params in raw_filters:
|
||||
canonical = names.get(literal)
|
||||
if canonical is None:
|
||||
raise ExtractionError(
|
||||
f"the image {name!r} is decoded through {literal}, a filter whose "
|
||||
"output this package cannot measure before producing it; refused "
|
||||
"unread rather than decoded to find out what it costs",
|
||||
"cost this package has no measured ratio for; refused unread "
|
||||
"rather than decoded to find out what it costs",
|
||||
code="asset_pdf_unbounded",
|
||||
)
|
||||
for index, (literal, params) in enumerate(filters):
|
||||
if literal in flate:
|
||||
if not any(behind in flate for behind, _ in filters[index + 1 :]):
|
||||
inflated_size(data, name=name)
|
||||
return
|
||||
if _has_predictor(params):
|
||||
raise ExtractionError(
|
||||
f"the image {name!r} applies a predictor to a link that is not the "
|
||||
"last one, so the bytes this package would hand to the next filter "
|
||||
"are not the bytes pdfminer decodes; refused unread",
|
||||
code="asset_pdf_unbounded",
|
||||
)
|
||||
data = inflate_bounded(data, name=name)
|
||||
elif literal in shrinking:
|
||||
try:
|
||||
data = _shrink(literal, data)
|
||||
except Exception:
|
||||
# Not this function's problem: a stream that is not valid input
|
||||
# for its own filter is reported by the reader behind it, in
|
||||
# that reader's vocabulary. Nothing can expand from it here.
|
||||
return
|
||||
# A pass-through filter leaves the bytes exactly as they are.
|
||||
check_payload(len(data), name=name)
|
||||
chain.append((canonical, params))
|
||||
|
||||
# The LAST deflate link is the one whose bytes nothing behind has to be
|
||||
# measured from, so it is counted and thrown away; every earlier one is
|
||||
# inflated under the same bound and handed on. Deciding this by index
|
||||
# rather than by a running flag is what keeps "the bytes are gone" and "a
|
||||
# link still needs them" from ever being true at once.
|
||||
last_flate = max(
|
||||
(index for index, (canonical, _) in enumerate(chain) if canonical == "FlateDecode"),
|
||||
default=-1,
|
||||
)
|
||||
size = len(data)
|
||||
for index, (canonical, params) in enumerate(chain):
|
||||
behind = [name_behind for name_behind, _ in chain[index + 1 :]]
|
||||
if canonical == "FlateDecode":
|
||||
limit = inflate_limit_for(behind[0] if behind else None)
|
||||
if index == last_flate:
|
||||
# Nothing behind has to be measured, so the output is counted
|
||||
# and thrown away: the cheap common case, and what 0.10.1 cost.
|
||||
size = inflated_size(data, name=name, limit=limit)
|
||||
data = b""
|
||||
else:
|
||||
if _has_predictor(params):
|
||||
raise ExtractionError(
|
||||
f"the image {name!r} applies a predictor to a link that is not "
|
||||
"the last one, so the bytes this package would hand to the next "
|
||||
"filter are not the bytes pdfminer decodes; refused unread",
|
||||
code="asset_pdf_unbounded",
|
||||
)
|
||||
data = inflate_bounded(data, name=name, limit=limit)
|
||||
size = len(data)
|
||||
else:
|
||||
# PREDICTED, not measured, and predicted BEFORE the decoder is
|
||||
# called: these decoders take a whole string and return a whole
|
||||
# string, so there is no moment between the two at which the cost
|
||||
# could be observed.
|
||||
check_filter_cost(size, canonical=canonical, name=name)
|
||||
if canonical in _SHRINKING_FILTER_NAMES:
|
||||
if index > last_flate >= 0:
|
||||
# The bytes were discarded at the last deflate link, so the
|
||||
# bound travels on as the WIDEST this link could produce.
|
||||
size = _widest_output(canonical, size)
|
||||
else:
|
||||
try:
|
||||
data = _shrink(canonical, data)
|
||||
except Exception:
|
||||
# Not this function's problem: a stream that is not
|
||||
# valid input for its own filter is reported by the
|
||||
# reader behind it, in that reader's vocabulary.
|
||||
return
|
||||
size = len(data)
|
||||
# A pass-through filter leaves the bytes exactly as they are.
|
||||
check_payload(size, name=name)
|
||||
|
||||
|
||||
def _widest_output(canonical: str, size: int) -> int:
|
||||
"""The most `canonical` can produce from `size` bytes, for a link whose
|
||||
bytes were discarded and whose SIZE is all that is carried forward.
|
||||
|
||||
Rounded up rather than down, and never below one byte: a bound that is
|
||||
optimistic by a byte is not a bound.
|
||||
"""
|
||||
ratio = PDF_FILTER_OUTPUT_RATIO.get(canonical)
|
||||
if ratio is None: # pragma: no cover - only `FlateDecode`, handled above
|
||||
return size
|
||||
return max(1, math.ceil(size * ratio))
|
||||
|
||||
|
||||
def _has_predictor(params: object) -> bool:
|
||||
|
|
@ -1570,13 +1666,19 @@ def _has_predictor(params: object) -> bool:
|
|||
return isinstance(predictor, int) and predictor > 1
|
||||
|
||||
|
||||
def _shrink(literal: object, data: bytes) -> bytes:
|
||||
"""The two filters whose output is smaller than their input, decoded with
|
||||
pdfminer's own readers so both sides agree on what the bytes are."""
|
||||
from pdfminer.ascii85 import ascii85decode, asciihexdecode
|
||||
from pdfminer.pdftypes import LITERALS_ASCII85_DECODE
|
||||
def _shrink(canonical: str, data: bytes) -> bytes:
|
||||
"""The two ASCII filters, decoded with pdfminer's own readers so both sides
|
||||
agree on what the bytes are.
|
||||
|
||||
if literal in LITERALS_ASCII85_DECODE:
|
||||
Called only after `check_filter_cost` has allowed the input size, which is
|
||||
what makes handing a whole string to a decoder that returns a whole string
|
||||
a bounded thing to do. The name "shrink" is kept for the pair, but only
|
||||
`ASCIIHexDecode` actually shrinks: `ASCII85Decode` can quadruple its input,
|
||||
which is the defect this round was opened for.
|
||||
"""
|
||||
from pdfminer.ascii85 import ascii85decode, asciihexdecode
|
||||
|
||||
if canonical == "ASCII85Decode":
|
||||
return ascii85decode(data)
|
||||
return asciihexdecode(data)
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue