fix(assets): bound what the run pays, not what the document claims (0.10.1)
A second independent review read `230d1cb` -- the commit that closed the `v0.10.0` review's two MAJOR findings -- and found one of them open. The bound read `/Width` and `/Height`, which an untrusted document writes, while `get_data()` pays for the stream beside them; `/Length` is the COMPRESSED length and the two numbers are independent. Re-measured here on `ed8d9d7` before anything changed, in its own interpreter: a 408 516-byte PDF declaring 1x1 and carrying 400 MB of deflated zeros was CARRIED, no rejection, 891 904 000 B peak RSS. After: 0 carried, `asset_too_large`, 57 065 472 B. At 1,2 GB of zeros, 2 436 MB -> 64 569 344 B -- the cost no longer scales with the bomb. End to end through the CLI with the shipped defaults: 838 000 640 B and an asset written -> exit 0, 79 650 816 B, `0 carried of 1 found`, no `assets/`. Three numbers are bounded now, not one: what a container DECLARES, what a carried FILE measures (`read_image`, so a 49 MP PNG of 47 705 bytes is not passed on to a consumer), and what a PDF stream DECOMPRESSES to (`assets.inflated_size`, a chunk at a time, output discarded, before `get_data()`). The limit is stated rather than implied: the stream measurement runs where `FlateDecode` is the first filter and the document is not encrypted; every other chain is a check on the decoded length AFTER the decode, a counted refusal and not a bounded one. A non-positive declared dimension is `asset_size_invalid`, its own code, raised before the stream is read. `-1 x 40000000000` is a NEGATIVE pixel count, under which every `>` bound read as satisfied, so the check returned silently and the refusal arrived from `encode_png` as `asset_samples_invalid`. Its own code because a publisher shipping a picture bigger than this package carries and a dictionary written to be read wrong are different facts about a document. Two smaller findings in the line that says what is missing, both introduced by the first fix: the address was written twice, once bare, and a linkifying renderer autolinks a bare URL -- written once now, in one code span; and `label` became a dead parameter, so the figure's caption was dropped, a regression against 0.10.0. It is written again in the `-- <label>` form a carried pointer uses. Version bumped to 0.10.1 across all ten places. Nine were unbound and stale: four README install lines naming the previous release, two prose lines, the "current tag" entry, `uv.lock`, and a CHANGELOG whose 0.10.1 content sat under `[Unreleased]`. Two new packaging tests bind them to `__version__`, and the README's guard tag to `[tool.uv.sources]`. Every test was red first. The fate of every image is identical with and without the new bound on three K2 PDFs carrying 800 images (464/464, 311/311 with the same 12 rejections, 25/25), and the second inflate is below the noise floor there. 0 shipped artifacts move: no bundle under `examples/`, `skills/` or `tests/fixtures/` carries an image pointer at all, measured against a known-positive control. `asset_too_large` was undocumented in the error registry; both codes are there now. `tools/okf_accounting_gate.py` gains the new code in its closed list -- one string, no behaviour change, stated because that file belongs to another order. Suite 2141 passed / 1 skipped, ruff + format + mypy --strict clean. Report: docs/2026-09-18-bildestien-holder-0-10-1.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
ed8d9d709f
commit
0f308c1f56
15 changed files with 842 additions and 81 deletions
|
|
@ -76,7 +76,7 @@ from .manifest import (
|
|||
)
|
||||
from .materialize import IngestResult, materialize_bundle
|
||||
|
||||
__version__ = "0.10.0"
|
||||
__version__ = "0.10.1"
|
||||
|
||||
__all__ = [
|
||||
"BlockedFile",
|
||||
|
|
|
|||
|
|
@ -131,22 +131,50 @@ IMAGE_POINTER = re.compile(
|
|||
#: Found by an independent review of 0.10.0 before it was pushed.
|
||||
MAX_IMAGE_PIXELS = 40_000_000
|
||||
|
||||
#: The same bound in bytes, for a payload whose pixels are not yet known (an
|
||||
#: inline `data:` URI) and for raw samples with several channels.
|
||||
#: The same bound in bytes: for a payload whose pixels are not yet known (an
|
||||
#: inline `data:` URI) and for the DECOMPRESSED size of a stream, which is a
|
||||
#: different number from anything the container declares. `/Length` in a PDF
|
||||
#: image dictionary is the COMPRESSED length and `/Width` and `/Height` are a
|
||||
#: claim; nothing in the dictionary states what the decompression will cost.
|
||||
MAX_IMAGE_BYTES = 256 * 1024 * 1024
|
||||
|
||||
#: How much of a stream is inflated at a time while `inflated_size` measures
|
||||
#: it. The cap on the OUTPUT is what keeps the measurement cheaper than the
|
||||
#: bomb; the input is handed over whole because it is already in memory.
|
||||
_INFLATE_CHUNK = 1 << 20
|
||||
|
||||
|
||||
def check_size(
|
||||
width: int | None, height: int | None, *, name: str, channels: int = 1, bits: int = 8
|
||||
) -> None:
|
||||
"""Refuse a declared size beyond the bound, BEFORE anything is decoded.
|
||||
|
||||
Read off what the container DECLARES, which is the only number available
|
||||
before the cost is paid: a check after the decompression has already paid
|
||||
for the bomb it was meant to stop.
|
||||
Read off what the container DECLARES. That is the only number available
|
||||
before a cost is paid, and it is a CLAIM by an untrusted document rather
|
||||
than a measurement: this refuses a picture too large to be one we carry,
|
||||
and it says nothing about what decompressing the stream beside it costs.
|
||||
`inflated_size` is the other half, and neither substitutes for the other.
|
||||
|
||||
`None` is UNKNOWN -- a container that declares no size leaves nothing to
|
||||
bound, and inventing a number would refuse a legitimate picture. A size
|
||||
that is declared and is not positive is neither unknown nor large: it is a
|
||||
dictionary that was written wrong or written to be read wrong, and it is
|
||||
refused here rather than three steps later by the encoder. Measured on
|
||||
`230d1cb`: `/Width -1 /Height 40000000000` multiplies to a NEGATIVE pixel
|
||||
count, so every `>` below was false, the function returned silently, 400 MB
|
||||
was decompressed, and the refusal arrived from `encode_png` under
|
||||
`asset_samples_invalid` -- a code about a sample buffer for a defect in the
|
||||
declaration.
|
||||
"""
|
||||
if not width or not height:
|
||||
if width is None or height is None:
|
||||
return
|
||||
if width <= 0 or height <= 0:
|
||||
raise ExtractionError(
|
||||
f"the image {name!r} declares {width}x{height}, which is not a size; "
|
||||
"refused unread rather than multiplied out, because a non-positive "
|
||||
"dimension makes every bound below it read as satisfied",
|
||||
code="asset_size_invalid",
|
||||
)
|
||||
pixels = width * height
|
||||
expected = pixels * channels * (bits // 8 or 1)
|
||||
if pixels > MAX_IMAGE_PIXELS or expected > MAX_IMAGE_BYTES:
|
||||
|
|
@ -169,6 +197,53 @@ def check_payload(size: int, *, name: str) -> None:
|
|||
)
|
||||
|
||||
|
||||
def inflated_size(raw: bytes, *, name: str, limit: int | None = None) -> int:
|
||||
"""What a deflate stream costs to decompress, measured without paying it.
|
||||
|
||||
THE DECLARED SIZE AND THE COST ARE TWO INDEPENDENT NUMBERS, and binding
|
||||
only the first is what an independent review of 0.10.1 measured on
|
||||
`230d1cb`: a 408 516-byte PDF declaring a 1x1 picture and carrying 400 MB
|
||||
of deflated zeros was CARRIED, with no rejection, at 892 MB of peak RSS --
|
||||
about 2 100x the file size, linear, so a 10 MB document is ~21 GB and takes
|
||||
the whole batch build with it. `check_size` was reading a claim as though
|
||||
it were a cost.
|
||||
|
||||
The output is inflated a chunk at a time and DISCARDED: only the running
|
||||
total is kept, so the measurement stays bounded whatever the stream holds,
|
||||
and it stops at the first chunk that crosses `limit`. A legitimate image is
|
||||
therefore inflated twice -- once here and once by the reader that carries
|
||||
it -- which is the price of not holding an unbounded buffer to find out how
|
||||
big it is. Measured on R761 Prosesskoden:2025 (50 image objects): the
|
||||
second pass costs under a second of a 200-second extraction.
|
||||
|
||||
A stream that is not valid deflate data is not this function's problem: the
|
||||
reader behind it reports that in its own vocabulary, so a `zlib.error` ends
|
||||
the measurement at whatever was produced up to it.
|
||||
"""
|
||||
bound = MAX_IMAGE_BYTES if limit is None else limit
|
||||
decompressor = zlib.decompressobj()
|
||||
total = 0
|
||||
pending = raw
|
||||
try:
|
||||
while True:
|
||||
produced = decompressor.decompress(pending, _INFLATE_CHUNK)
|
||||
total += len(produced)
|
||||
if total > bound:
|
||||
raise ExtractionError(
|
||||
f"the stream behind {name!r} decompresses to more than {bound} "
|
||||
f"bytes from {len(raw)} bytes of input, over this package's "
|
||||
"bound; refused without being held, because the size a "
|
||||
"container declares is a claim and this is the cost",
|
||||
code="asset_too_large",
|
||||
)
|
||||
pending = decompressor.unconsumed_tail
|
||||
if decompressor.eof or not pending:
|
||||
break
|
||||
except zlib.error:
|
||||
return total
|
||||
return total
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ExtractedImage:
|
||||
"""One image a document carries, as this package will write it.
|
||||
|
|
@ -310,6 +385,16 @@ def read_image(data: bytes, *, name: str, label: str | None = None) -> Extracted
|
|||
)
|
||||
media_type, suffix = kind
|
||||
size = dimensions(data)
|
||||
# THE BOUND HOLDS FOR A FILE CARRIED VERBATIM TOO. This package does not
|
||||
# decode one, so it pays nothing for it -- but writing a 7 000 x 7 000 PNG
|
||||
# of 47 705 bytes into a bundle hands the consumer the same bomb with
|
||||
# `7000x7000 px` printed beside it, and the README's first sentence about
|
||||
# this bound says such an image is refused. Sniffed from the header rather
|
||||
# than claimed, like the type beside it. Measured: the largest of the
|
||||
# 4 828 objects in the reference corpus is 18.6 MP, so nothing anyone has
|
||||
# measured is refused here.
|
||||
if size is not None:
|
||||
check_size(size[0], size[1], name=name)
|
||||
return ExtractedImage(
|
||||
data=data,
|
||||
name=name,
|
||||
|
|
@ -527,22 +612,37 @@ def render_missing(
|
|||
absence they were never shown -- which is precisely the shape of the defect
|
||||
this whole capability exists to close.
|
||||
|
||||
The href is kept when the source had one and it survives the link grammar,
|
||||
so a remote figure says WHERE it was. Extraction never opens a socket: the
|
||||
network gate is an explicit per-run opt-in and extraction is not on that
|
||||
path, so a remote source is carried as a name and never as bytes.
|
||||
The href is kept when the source had one, so a remote figure says WHERE it
|
||||
was. Extraction never opens a socket: the network gate is an explicit
|
||||
per-run opt-in and extraction is not on that path, so a remote source is
|
||||
carried as a name and never as bytes.
|
||||
|
||||
`label` is the alt text or the figure caption, and it is written for the
|
||||
same reason the line exists at all: the reader deciding whether the missing
|
||||
picture mattered is much better served by "Figur 84-1 Tverrprofil" than by
|
||||
a file name. 0.10.1 dropped it while closing the live-link defect -- the
|
||||
parameter stayed in the signature and no branch read it -- which an
|
||||
independent review measured as a regression against 0.10.0.
|
||||
"""
|
||||
shown = _inline(name or "image")
|
||||
detail = f"Image: {shown} (not carried: {_inline(reason)})"
|
||||
# INERT, never ``. A remote reference is an address the
|
||||
# document's author chose, and 0.10.0 wrote it as a live markdown image: a
|
||||
# consumer that renders the bundle, or an agent that fetches what it
|
||||
# renders, turns "this bundle was opened" into a beacon to them -- and a
|
||||
# server-side consumer into an SSRF. This package opens no socket, which is
|
||||
# not the same as the pointer being harmless.
|
||||
#
|
||||
# The address is STATED, because a reader has to know what stood there --
|
||||
# and stated ONCE, inside a code span. 0.10.1 wrote it twice: once in the
|
||||
# span and once as the bare `name`, and a GFM/linkify renderer autolinks a
|
||||
# bare URL into `<a href="...">`. It takes a click rather than a render, so
|
||||
# it is weaker than `![..]()` -- but "inert" was half the truth, and half
|
||||
# is what this line exists not to be.
|
||||
if href:
|
||||
# INERT, never ``. A remote reference is an address the
|
||||
# document's author chose, and 0.10.0 wrote it as a live markdown
|
||||
# image: a consumer that renders the bundle, or an agent that fetches
|
||||
# what it renders, turns "this bundle was opened" into a beacon to
|
||||
# them -- and a server-side consumer into an SSRF. This package opens
|
||||
# no socket, which is not the same as the pointer being harmless.
|
||||
# The address is still STATED, in a code span, because a reader has to
|
||||
# know what stood there to judge whether the picture mattered.
|
||||
address = _inline(href).replace("`", "'")
|
||||
return f"{detail} address: `{address}`"
|
||||
shown = f"`{address}`"
|
||||
else:
|
||||
shown = _inline(name or "image")
|
||||
detail = f"Image: {shown} (not carried: {_inline(reason)})"
|
||||
if label:
|
||||
detail += f" -- {_inline(label)}"
|
||||
return detail
|
||||
|
|
|
|||
|
|
@ -133,6 +133,22 @@ class ExtractionError(IngestError):
|
|||
- `asset_samples_invalid` — the sample buffer does not fit the dimensions
|
||||
the image dictionary declares. Refused rather than padded: a short buffer
|
||||
means the dictionary was read wrong
|
||||
- `asset_too_large` — the picture is over this package's bound, either
|
||||
because it DECLARES a size beyond `MAX_IMAGE_PIXELS`, because the file
|
||||
itself is that large, or because the stream behind it DECOMPRESSES to
|
||||
more than `MAX_IMAGE_BYTES`. The three are one code because they are one
|
||||
decision — this run will not hold that picture — and because a consumer
|
||||
counting refusals wants the picture, not the mechanism. The bound is read
|
||||
off the corpora and sits an order of magnitude above anything measured
|
||||
- `asset_size_invalid` — the container declares a size that is not a size:
|
||||
a zero or negative `/Width` or `/Height`. DISTINCT from
|
||||
`asset_too_large`, because the two say different things about the
|
||||
document — one is a legitimate publisher shipping a picture bigger than
|
||||
this package carries, the other is a dictionary written wrong or written
|
||||
to be read wrong — and counting them together would make a corpus
|
||||
statistic about the first untrue. Refused before the stream is read: a
|
||||
negative dimension multiplies to a negative pixel count, under which
|
||||
every bound reads as satisfied
|
||||
"""
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -48,6 +48,7 @@ from .assets import (
|
|||
check_payload,
|
||||
check_size,
|
||||
encode_png,
|
||||
inflated_size,
|
||||
read_image,
|
||||
render_block,
|
||||
render_missing,
|
||||
|
|
@ -1365,6 +1366,40 @@ def _pdf_alpha(attrs: dict[str, object], width: int, height: int) -> bytes | Non
|
|||
return alpha if len(alpha) >= width * height else False
|
||||
|
||||
|
||||
def _check_inflated(stream: object, name: str) -> None:
|
||||
"""Refuse a PDF image stream whose DECOMPRESSED size is over the bound.
|
||||
|
||||
`check_size` reads `/Width` and `/Height`, which an untrusted document
|
||||
writes and which say nothing about what `get_data()` will return: `/Length`
|
||||
is the COMPRESSED length, and a dictionary declaring 1x1 may hang 400 MB of
|
||||
deflated zeros off it. Measured by an independent review of 0.10.1 on
|
||||
`230d1cb`: that document is 408 516 bytes, was carried without a rejection,
|
||||
and cost 892 MB of peak RSS -- the same failure mode the 0.10.0 review
|
||||
named, reached through the other number.
|
||||
|
||||
WHAT THIS BOUNDS, and what it does not. The measurement runs on the RAW
|
||||
stream, so it applies where `FlateDecode` is the first filter -- the
|
||||
overwhelming case, and the one every deflate bomb needs. It does not run
|
||||
when the stream is already decoded, when the document is encrypted (the raw
|
||||
bytes are ciphertext until pdfminer deciphers them), or when a chain puts
|
||||
another filter in front. Those are caught by `check_payload` AFTER
|
||||
`get_data()`, which makes them a counted refusal rather than a bounded one:
|
||||
the memory is spent and then the picture is dropped. Stated rather than
|
||||
implied, because the difference is exactly what the review found missing.
|
||||
"""
|
||||
from pdfminer.pdftypes import LITERALS_FLATE_DECODE
|
||||
|
||||
if getattr(stream, "decipher", None) is not None:
|
||||
return
|
||||
raw = stream.get_rawdata() if hasattr(stream, "get_rawdata") else None
|
||||
if not raw:
|
||||
return
|
||||
filters = stream.get_filters() if hasattr(stream, "get_filters") else []
|
||||
if not filters or filters[0][0] not in LITERALS_FLATE_DECODE:
|
||||
return
|
||||
inflated_size(raw, name=name)
|
||||
|
||||
|
||||
def _pdf_image(stream: object, name: str) -> ExtractedImage:
|
||||
"""One image XObject, carried verbatim where it already is a file.
|
||||
|
||||
|
|
@ -1394,16 +1429,22 @@ def _pdf_image(stream: object, name: str) -> ExtractedImage:
|
|||
|
||||
# THE DECLARED SIZE IS READ FIRST, and the stream is not touched until it
|
||||
# is within the bound: `get_data()` decompresses, so a check after it has
|
||||
# already paid for a picture of compressed zeros.
|
||||
# already paid for a picture of compressed zeros. `channels=1` because the
|
||||
# colour space is not resolved until further down and one channel is the
|
||||
# LOWEST estimate of the cost -- a bound that under-counts refuses nothing
|
||||
# it should carry, and the pixel count beside it does not depend on it.
|
||||
# What that leaves unbounded is the stream, which `_check_inflated` reads
|
||||
# for what it actually costs rather than for what it claims.
|
||||
declared = dict(getattr(stream, "attrs", {}))
|
||||
declared_width = resolve1(declared.get("Width"))
|
||||
declared_height = resolve1(declared.get("Height"))
|
||||
check_size(
|
||||
resolve1(declared.get("Width")) if isinstance(resolve1(declared.get("Width")), int) else 0,
|
||||
resolve1(declared.get("Height"))
|
||||
if isinstance(resolve1(declared.get("Height")), int)
|
||||
else 0,
|
||||
declared_width if isinstance(declared_width, int) else None,
|
||||
declared_height if isinstance(declared_height, int) else None,
|
||||
name=name,
|
||||
channels=1,
|
||||
)
|
||||
_check_inflated(stream, name)
|
||||
try:
|
||||
data = stream.get_data() # type: ignore[attr-defined]
|
||||
except Exception as exc:
|
||||
|
|
@ -1411,6 +1452,10 @@ def _pdf_image(stream: object, name: str) -> ExtractedImage:
|
|||
f"the PDF image stream behind {name!r} could not be decoded: {exc}",
|
||||
code="asset_pdf_unsupported",
|
||||
) from exc
|
||||
# THE BACKSTOP, and it is a weaker guarantee than the one above it: this
|
||||
# one COUNTS a stream that `_check_inflated` could not bound in advance
|
||||
# (see its docstring for which those are), after the memory has been spent.
|
||||
check_payload(len(data), name=name)
|
||||
if data and sniff(data) is not None:
|
||||
return read_image(data, name=name)
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue