Every fixture, test document, tool example and document now uses an invented kitchen-and-baking handbook series, written in this repository. The package's behaviour is unchanged; src/ changes are comments and help text only. - Generated fixtures are regenerated from their generators. Their structural counts are identical before and after: elements, images, rows, cells, headings, bookmarks and the witness inventory's per-document totals. The image-inbox and accounting documents are renamed kapittel-84-*. - tools/okf_accounting_gate.py: the two options that named one real corpus each are replaced by a generic, repeatable --corpus PATH with no default. Row 5 compares the PDF pair alone. Gate verdict unchanged: RED rows 2, 3, 6. - tools/okf_witness.py: the STS JSON reader for one publisher's delivery is removed, along with its three twins and five tests. The mutation harness loses W09. - docs/: 13 dated reports that documented runs on a retired reference corpus are removed, and 40 are neutralized. Dead links are removed, and no new dangling path is introduced. - The synthetic MCP-gate corpus and the residual probe words are neutral. Valgt: keep the `okf quality --fasit` bar value (the measured fraction, one corpus) and rewrite only its provenance, because the verdict stays unchanged and the number names nothing. Term check with the local list: 0 of 411 tracked files, 0 file names, 0 of 27 binary fixtures. Suite after git add: 2457 passed, 1 skipped. The base tree had 2460 passed and 2 skipped; five tests went with the JSON reader and four were added by the term check. ruff, ruff format and mypy --strict src/ are clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
254 lines
14 KiB
Markdown
254 lines
14 KiB
Markdown
# A budget per link, and a property over every chain
|
||
|
||
A PM checkpoint of `0c3c490` — the commit that bound every link of a PDF
|
||
filter chain — read the fix and found the bound still reachable. This is the
|
||
third round of the same review, and the third time a bound was put on an
|
||
**output** and the bomb stepped one link along.
|
||
|
||
The two rounds before it are
|
||
[`docs/2026-09-17-bildestien-0-10-1.md`](2026-09-17-bildestien-0-10-1.md),
|
||
[`docs/2026-09-18-bildestien-holder-0-10-1.md`](2026-09-18-bildestien-holder-0-10-1.md)
|
||
and [`docs/2026-09-18-filterkjeden-og-backstoppen.md`](2026-09-18-filterkjeden-og-backstoppen.md),
|
||
which carries a correction block pointing here.
|
||
|
||
## The finding is the pattern, not the filter
|
||
|
||
| round | what was bound | where the bomb moved |
|
||
| --- | --- | --- |
|
||
| 1 | the size the dictionary DECLARES | into the stream |
|
||
| 2 | the first `FlateDecode` link | into the second link |
|
||
| 3 | every `FlateDecode` link | into a link documented as safe |
|
||
|
||
Round 3 shipped a three-class split in which `ASCII85Decode` and
|
||
`ASCIIHexDecode` were "bounded by their own input because they shrink". That
|
||
sentence is false in two independent ways:
|
||
|
||
* **ASCII85 does not shrink.** `z` is its shorthand for four zero bytes, so
|
||
one input byte becomes four. The output ratio is 4, not 0.8.
|
||
* **The output is not the cost.** `base64.a85decode` appends one 4-byte object
|
||
per group to a list, so a run of `z` costs about a hundred bytes of memory
|
||
per byte of INPUT. Measured on CPython 3.14: **101.4x** at 1 MiB of input,
|
||
**96.1x** at 4 MiB, **94.5x** at 16 MiB. Nothing in the output size says so.
|
||
|
||
The second point is the general one, and it is why this round does not add a
|
||
fourth class. A bound on what a link OUTPUTS is not a bound on what producing
|
||
it COSTS, and every round of this review has been an instance of that.
|
||
|
||
## BLOCKER — reproduced
|
||
|
||
Measured in paired subprocesses on an idle machine, both sides from pinned
|
||
trees (the `before` side from `git archive` of `0c3c490`, the `after` side from
|
||
a snapshot of the working tree), each printing the module's own `__file__` as
|
||
a control. The document is built ONCE by a third process and read from a file:
|
||
`ru_maxrss` is a high-water mark that never falls, so a child that builds its
|
||
own fixture reports the fixture — `b"z" * (64 MiB)` alone costs 171 MB, more
|
||
than the bounded extraction it would be measuring. This tripped the first
|
||
attempt at these figures.
|
||
|
||
| chain | file | before | after |
|
||
| --- | --- | --- | --- |
|
||
| `[/FlateDecode /ASCII85Decode]`, `z` × 32 Mi | 33 475 B | **carried**, 3 261 599 744 B | `asset_too_large`, **42 070 016 B** |
|
||
| `[/FlateDecode /ASCII85Decode]`, `z` × 64 Mi | 66 090 B | **carried**, 6 461 558 784 B | `asset_too_large`, **40 280 064 B** |
|
||
| `[/ASCII85Decode]`, `z` × 8 Mi | 8 389 449 B | **carried**, 933 085 184 B | `asset_too_large`, **62 484 480 B** |
|
||
| `[/Fl /A85 /Fl]`, `z` × 32 Mi | 33 488 B | `asset_samples_invalid`, 3 519 180 800 B | `asset_too_large`, **43 438 080 B** |
|
||
|
||
Two things to read off it. The picture was **carried** in three of the four
|
||
rows — this was not a bound that fired late, it was no bound at all. And the
|
||
bounded cost does not follow the bomb: doubling the run of `z` takes the old
|
||
cost from 3.26 GB to 6.46 GB and the new one from 42.1 MB to 40.3 MB.
|
||
|
||
## The rule: a budget per link, and it travels
|
||
|
||
`assets.MAX_FILTER_DECODE_BYTES` (512 MiB) is what decoding ONE link may cost.
|
||
It is a separate number from `MAX_IMAGE_BYTES` (256 MiB) on purpose: that one
|
||
bounds the picture this package will carry, this one bounds what producing it
|
||
costs on the way.
|
||
|
||
`assets.PDF_FILTER_COST_RATIO` gives each permitted filter a **measured**
|
||
worst-case peak memory per byte of input. `None` means the decoder is driven a
|
||
chunk at a time here, so the cost is measured as it is paid — today that is
|
||
`FlateDecode` alone.
|
||
|
||
| filter | cost ratio | measured |
|
||
| --- | --- | --- |
|
||
| `FlateDecode` | — | driven a chunk at a time (`assets._inflate`) |
|
||
| `ASCII85Decode` | 104 | 101.4x / 96.1x / 94.5x at 1 / 4 / 16 MiB of `z` |
|
||
| `ASCIIHexDecode` | 2 | 1.5x at 16 MiB |
|
||
| `DCTDecode`, `JPXDecode`, `JBIG2Decode` | 1 | pass-through in pdfminer |
|
||
|
||
Everything else — `LZWDecode`, `RunLengthDecode`, `CCITTFaxDecode`, `/Crypt`,
|
||
anything written later — has no measured ratio and is refused UNREAD with
|
||
`asset_pdf_unbounded`, before any link in front of it is decoded. That is the
|
||
same decision `corpus.resolve_gate` takes for an unknown gate name: a fallback
|
||
reproduces the defect with an extra step.
|
||
|
||
The budget **travels down the chain**. A `FlateDecode` link's output is the
|
||
next link's input, so it is inflated under a limit that is the smaller of
|
||
`MAX_IMAGE_BYTES` and what the next link's decoder may be handed
|
||
(`assets.inflate_limit_for`). Without that, `[/FlateDecode /ASCII85Decode]`
|
||
would inflate 256 MiB of `z` before the link behind it was asked anything.
|
||
|
||
### Why an input cap and not a bounded ASCII85 decoder
|
||
|
||
A chunked ASCII85 decoder written here would bound `_check_stream_cost`, and
|
||
it would not bound the run: pdfminer decodes the whole chain again in
|
||
`stream.get_data()`, with its own unbounded decoder, and that is where the
|
||
memory is actually spent. The only number that bounds *that* is the size of
|
||
the input this package allows the link to be handed. Writing our own decoder
|
||
would also make this package, rather than pdfminer, the authority on what an
|
||
image's bytes are.
|
||
|
||
### Why not a hard backstop in a child process
|
||
|
||
The order asked for `resource.setrlimit(RLIMIT_AS)` to be MEASURED before
|
||
anything was built on it. It was, and it is not available here: on this
|
||
machine (Darwin 26.6.2, CPython 3.14) `setrlimit(RLIMIT_AS, (256 MiB, hard))`
|
||
raises `ValueError: current limit exceeds maximum limit` — a fresh CPython
|
||
process has already reserved far more address space than the cap, and the hard
|
||
limit reads as `RLIM_INFINITY`. The documented behaviour agrees: Darwin does
|
||
not enforce `RLIMIT_AS`, `RLIMIT_DATA` or `RLIMIT_RSS` the way Linux does. So
|
||
no child-process memory cap was built, and the bound is the per-link budget
|
||
alone.
|
||
|
||
## The cap is read off the corpora
|
||
|
||
The cap that falls out for `ASCII85Decode` is 512 MiB / 104 ≈ **5.0 MB** of
|
||
input. Measured 2026-09-18 over the **9 668 image objects of the 77 PDFs on
|
||
this machine** (enumerated through pdfminer's own page walk):
|
||
|
||
| chain | objects |
|
||
| --- | --- |
|
||
| `[/FlateDecode]` | 6 235 |
|
||
| `[/DCTDecode]` | 2 459 |
|
||
| `[/FlateDecode /DCTDecode]` | 596 |
|
||
| `[/Fl]` | 296 |
|
||
| unfiltered | 42 |
|
||
| `[/ASCII85Decode /FlateDecode]` | 16 |
|
||
| `[/JPXDecode]` | 16 |
|
||
| `[/CCITTFaxDecode]` | 8 |
|
||
|
||
**16** objects decode through an `ASCII85Decode` link, and the largest input
|
||
any of them is handed is **450 739 bytes** — more than ten times under the
|
||
cap. That is the posture `MAX_IMAGE_PIXELS` has: a number read off the corpora
|
||
and standing an order of magnitude above anything measured, so the bound costs
|
||
no picture anybody has.
|
||
|
||
Two corrections to earlier published counts fall out of this table, and both
|
||
are about ENUMERATION rather than about the documents. The round-2 report
|
||
counted 5 142 objects over 78 PDFs and **580** behind `[/FlateDecode
|
||
/ASCII85Decode]`. This walk finds 9 668 objects over 77 files and **0** behind
|
||
that chain. The denominators differ because the two walks are different (this
|
||
one recurses into `LTFigure`; the reference standard's file is not in this listing), so
|
||
neither number is wrong about a document — but a chain count is only readable
|
||
beside the walk that produced it, and the `[/Fl /A85]` group is not one this
|
||
machine's corpora hold. The bomb that shape carries is real regardless: a
|
||
document does not have to exist in a corpus to be handed to `okf build`.
|
||
|
||
## The property test
|
||
|
||
`tests/test_asset_limits.py::test_no_chain_of_up_to_three_filters_is_carried_over_the_bound`
|
||
generates **every** chain of length 1–3 over the ten filters pdfminer decodes
|
||
— `K = 1 110` — twice, once with a payload of zeros (the amplifying case at
|
||
both ends: it deflates to nothing and `a85encode`s to a run of `z`) and once
|
||
with a repeated non-zero byte. Each chain's stream is built by encoding the
|
||
payload BACKWARDS through the chain, so every chain over the permitted filters
|
||
is a valid document rather than a rejection by accident.
|
||
|
||
The requirement is one sentence: the picture is either delivered with its
|
||
bytes under the bound, or refused with a code in the published vocabulary;
|
||
never carried over the bound, and never paid for on the way. `tracemalloc`
|
||
measures the paying, because it counts Python's own allocations — which is
|
||
exactly where `a85decode`'s cost lives, and unlike `ru_maxrss` it is not
|
||
disturbed by other work on the machine.
|
||
|
||
Result: **1 110 of 1 110** chains pass, both payloads — `k = K`.
|
||
|
||
Beside it, `test_every_bounded_chain_still_carries_a_small_image` runs the
|
||
**258** chains over the six permitted filters with a 64-byte image and
|
||
requires none of them to be refused. A rule that refuses everything passes the
|
||
property alone; it does not pass this.
|
||
|
||
And `test_the_ascii85_cost_ratio_is_not_below_the_one_this_package_measured`
|
||
re-measures the ratio the budget rests on, in a subprocess, at two input
|
||
sizes. If CPython ever changes `a85decode` so that it costs more, the constant
|
||
is too generous and this says so before a corpus does.
|
||
|
||
## MAJOR — the backstop had no test
|
||
|
||
`check_payload(size, name=name)` at the END of `_check_stream_cost` could be
|
||
deleted with the whole suite still passing. It is what refuses a stream no
|
||
filter in the chain expands — an unfiltered one, or one behind `DCTDecode` —
|
||
and the SECOND `check_payload`, after `get_data()`, produces the same code and
|
||
the same words one step later. A test that reads the code cannot tell the two
|
||
apart.
|
||
|
||
What separates them is whether the payment was made, so the test asserts
|
||
`get_data` was never called.
|
||
|
||
## Mutants
|
||
|
||
Ten mutations, one line each, run in a scratch clone with the unmutated copy
|
||
run FIRST as a control. The peak-RSS subprocess tests are deselected for these
|
||
runs — they measure a high-water mark and the machine was running a corpus
|
||
census — so what kills a mutant here is the property test, the code
|
||
vocabulary, or an assertion about which check fired.
|
||
|
||
| mutation | one line | killed by |
|
||
| --- | --- | --- |
|
||
| `backstop-deleted` | `check_payload` at the end of `_check_stream_cost` removed | `test_the_stream_bound_refuses_before_get_data_is_ever_called` |
|
||
| `cost-check-deleted` | the per-link `check_filter_cost` call removed | the property test, both fills, + `[/ASCII85Decode]` |
|
||
| `ascii85-budget-removed` | that filter's ratio set to `None`, so it gets no budget | 9 tests, including both ratio measurements and the corpus cap |
|
||
| `ascii85-ratio-is-one` | the ratio set to 1 instead of the measured 104 | 6 tests, including both ratio measurements |
|
||
| `budget-does-not-travel` | `inflate_limit_for` returns `MAX_IMAGE_BYTES` always | `test_the_budget_travels_to_the_next_link` |
|
||
| `widest-output-ignored` | `_widest_output` returns its input unchanged | `test_a_discarded_links_size_travels_as_the_widest_it_could_become` |
|
||
| `first-flate-not-last` | the discard happens at the FIRST deflate link | 5 tests, including both round-2 chain bombs |
|
||
| `unknown-filter-passes` | a filter with no ratio is let through instead of refused | `test_a_filter_the_bound_cannot_measure_...`, `test_asset_pdf_unbounded` |
|
||
| `budget-a-hundredfold` | `MAX_FILTER_DECODE_BYTES` multiplied by 100 | 4 tests, including both new bombs |
|
||
| `cost-check-off-by-a-factor` | the comparison allows 1 000x the limit | `test_an_ascii85_link_on_its_own_is_bounded` |
|
||
|
||
**10 of 10 killed.** Four of them survived a first pass and are the reason two
|
||
of the tests above exist: `budget-does-not-travel` and `widest-output-ignored`
|
||
had no test at all, and `budget-a-hundredfold` and
|
||
`cost-check-off-by-a-factor` were only reachable through the peak-RSS
|
||
subprocess tests that first pass had deselected. A mutant that survives is a
|
||
test that was missing, not a mutation that was unfair.
|
||
|
||
## Cost to real documents
|
||
|
||
Measured by name, not by total: `_pdf_images` run over every PDF on this
|
||
machine from each of the two pinned trees, each printing the module file it
|
||
loaded as a control, and the per-file counts compared.
|
||
|
||
| corpus | files | carried before | carried after | files whose count moved |
|
||
| --- | --- | --- | --- | --- |
|
||
| `~/corpora` + `tests/fixtures` | 77 | 9 306 | **9 306** | **0** |
|
||
| the reference standard | 1 | 50 | **50** | **0** |
|
||
| both | 78 | 9 356 | **9 356** | **0** |
|
||
|
||
The rejection codes are identical too — `asset_pdf_unsupported` 314,
|
||
`asset_pdf_unbounded` 8, `asset_samples_invalid` 40 on both sides, and **0
|
||
files** where any per-file code count moved. Not one picture changes hands.
|
||
|
||
That is what the cap being read off the corpora buys: the only new refusal is
|
||
`check_filter_cost`, and the largest `ASCII85Decode` input any of these
|
||
documents holds is 450 739 bytes against a cap of about 5.0 MB.
|
||
|
||
The two trees are `git archive` of `0c3c490` and a snapshot of the working
|
||
tree taken before the prose edits; `diff -r` between that snapshot and the
|
||
committed tree touches docstrings and comments only, so what was measured is
|
||
what shipped.
|
||
|
||
The reference standard also settles the `[/FlateDecode /ASCII85Decode]` question the PM raised
|
||
while this order was open: **50 image objects, 29 `[/DCTDecode]` and 21
|
||
`[/FlateDecode]`, and 0 ASCII85 links.** So that chain is in neither the 77
|
||
corpus files nor the reference standard, and the round-2 report's count of 580 is not
|
||
reproducible from anything on this machine. It changes no decision here — the
|
||
bomb that shape carries does not need a corpus to exist in — but a published
|
||
count that cannot be reproduced should not stay unqualified.
|
||
|
||
## What is still not bounded
|
||
|
||
Unchanged from round 2, and stated rather than implied: a stream something
|
||
else has already decoded (`_pdf_stream_bytes` returns `None`), where the
|
||
memory is spent before this package is asked. `check_payload` after
|
||
`get_data()` COUNTS it — a counted refusal, not a bounded one.
|