feat(extract,cli): typography as a PDF heading source and OCR behind an optional group, both off

A PDF carries no notion of a heading -- a heading in a PDF is a typographic
fact -- so the text stream `pdfplumber` hands the segment proposer has already
thrown away the only evidence there was. The `docx` path never had that problem:
the converter emits ATX headings and `_ATX` cuts on them. Two readers close the
gap, and both are OFF.

`--pdf-headings font` infers a heading from the conjunction this repository
already measured (size above the document's character-weighted body median AND
a bold font name, recall 1.000 / precision 0.846) and emits it as ATX in the
SAME markdown the office path produces, so `_ATX` applies unchanged and no
PDF-only heading grammar exists.

It stays off BY MEASUREMENT, and the measurement is the point of the round:
against the operator's unit worksheet it takes `pdf` from 2 of 8 to 0 of 8,
losing two exact matches. The mechanism of the loss is stated rather than
guessed -- on those documents the outline rule already recovers the document's
own numbered chapters, so a second heading source can only add. Whole-corpus
screen: 25 of 32 `pdf` change, 0 of 5 `docx`, 0 of 2 `xlsx`. The default bundle
is byte-identical before and after this commit (`diff -r`, exit 0).

`--ocr` reads a page as an image when its own text never arrived: empty, or
`(cid:N)` placeholder codes at or above a threshold READ OFF a measured
distribution -- 834 pages over 32 files, 818 at exactly 0.0 and 16 at 0.93 or
above, nothing in between. On the one corpus document with the failure: 95.07 %
cid to 0 %, 44 to 2561 words of four or more letters, 17 to 18 pages with text.
Its engine is an optional dependency group and never a runtime dependency; a
packaging test pins both halves, and without the group every affected file is a
coded rejection (`extractor_ocr_group_missing`) rather than a crash.

Also corrects two stale published facts found while measuring: the README still
said two segmentation rules were on by default after `f6fea13` made it three,
and CLAUDE.md's K2 digest named the round-3 default. The current default is
492 concepts / 944 files, `bdefa679...`.

Report: docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md

Co-Authored-By: Claude <claude-opus-5>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-08 23:10:47 +02:00
commit 53d5c74c96
15 changed files with 1394 additions and 28 deletions

View file

@ -0,0 +1,483 @@
# K3 round 4, `pdf`: typography as a heading source, and OCR for the pages whose text never came
Fourth iteration of the per-file-type directive of 2026-09-08 13:05Z. Order
`20260908T195728Z-8160520552-from-.claude`. HEAD at the start and at the end of
the session is O5's `f6fea13` plus this session's commit; the tag on `f6fea13`
is `v0.6.0`.
Everything below is measured on the same 43-document corpus and the same
12-document sample as `docs/2026-09-08-k3-runde3-per-filtype.md`, against the
same reference: the operator's filled unit worksheet, which is local-only and
gitignored. No document title or extracted text from it is reproduced here.
## Section 0: what was measured and what was not
**Measured.** The three reproduction numbers the order names, before any edit.
Two new readers, each red-first, each with a known-negative. The per-page
`(cid:N)` distribution over every K2 PDF, with both denominators, and the
threshold read off it rather than chosen. What OCR recovers on the one document
that has the failure, through the shipped code path and the shipped instrument.
The 12-document sample re-scored with session 109's criterion under seven
configurations. The whole 43-document corpus under both renderings from one
cached extraction. Two K2 bundles differing only in the new flag.
**Not measured.** Whether a live model answers better with any of this. Whether
OCR helps a corpus that is not this one -- the `(cid:N)` denominator that
matters is **1 document of 39**. Whether the font reader helps a corpus whose
typography is unlike this one. No o200k token counts (no encoder is installed
and the single-dependency rule does not admit one).
## 1. The reproduction, before any edit
`git rev-parse HEAD` -> `f6fea13299d92b959f497652ff19a7fd79546281`,
`git status --short` empty, tag `v0.6.0` on that commit.
| # | control | expected | result |
|---|---|---|---|
| a | `pytest -q` | 1427 collected | **1427 passed** in 38 s |
| b | `okf build ~/okf-test/dokumenter`, no flags | 30 | **30** markdown files |
| c | the same folder, `--no-table-grid` | 43 | **43** markdown files |
All three reproduce. (a) differs in FORM from the PM's figure and not in
substance: on a frozen export the same suite is 1424 passed, 1 skipped and 2
export artefacts; in a working tree all 1427 run and pass.
## 2. The root cause, verified rather than inherited
The order states a diagnosis. Each half was checked against the tree before any
code was written, because a premise is an input like any other:
| claim | check | result |
|---|---|---|
| PDF reaches the proposer as a plain text stream | `_pdf_pages` called `page.extract_text()` and nothing else | confirmed |
| the `docx` path emits ATX the proposer already cuts on | `_PANDOC_WRITER = "markdown"`, never `plain` | confirmed |
| the measured best deterministic heading rule is not on the run path | `RULE_POPPLER_SIZE_AND_BOLD` occurs **twice** in `src/`, both in `propose.py` -- its definition and its entry in `RULE_NAMES` -- and **zero times** in `tools/` or `tests/` | confirmed: no candidate has ever carried it |
A PDF has no notion of a heading. A heading in a PDF is a typographic fact, and
the text stream throws that fact away before anything downstream can read it.
## 3. Deliverable A: the font reader, behind `--pdf-headings font`
### The mechanism
`pdfplumber` -- already a dependency -- reports `size` and `fontname` per
character. Per line, the reader takes the dominant size and dominant font name
over the line's non-blank characters. A line is a heading when its dominant
size is **above the document's character-weighted median size** AND its
dominant font name contains `bold`. The heading sizes of the whole document are
ranked largest-first and that rank becomes the ATX level, clamped at six.
Three choices, each with a reason rather than a taste:
- **A conjunction, not a disjunction.** `propose.py` already records the
measurement: `size AND bold` scored recall 1.000 / precision 0.846, and
adding weight as a disjunct took precision from 0.786 to 0.524.
- **The body size is the document's, not the page's.** A title page is 100 %
heading, so a per-page median compares it with itself and marks nothing.
- **The output is ATX in the same markdown the office path produces**, so
`_ATX` reads it unchanged and no PDF-only heading grammar exists. A second
grammar would be a second definition of "heading" free to drift from the one
five office formats already use. The cost is stated in Honesty limits.
### The controls
| control | result |
|---|---|
| a PDF with one font at one size, flag ON | **byte-identical** to flag OFF (test) |
| `.docx` / `.xlsx`, flag ON, whole corpus | **0 of 5** and **0 of 2** changed |
| every readable corpus document: same line count, and every differing line is the cached line with a hash run prefixed | **39 of 39** |
That last control is the load-bearing one and it states three things at once:
the round-3 cache is still valid (`extract.py` is byte-identical between
`c3b645b` and `f6fea13`), joining `extract_text_lines()` reproduces
`extract_text()` on real documents rather than on the one it was spot-checked
on, and the reader **adds markers and changes nothing else**.
**The first version of that control was wrong and said so.** It stripped
`^#{1,6} ` from the marked text and compared: 38 of 39. The one failure is a
corpus document whose own extracted text carries lines beginning with `# `, and
the stripper removed those too. The corrected control compares line by line.
### The reach
Whole corpus, 39 readable documents, one cached extraction per rendering:
| figure | value | denominator |
|---|---|---|
| documents changed | **25** | 32 `pdf` |
| documents changed | **0** | 5 `docx` |
| documents changed | **0** | 2 `xlsx` |
| ATX headings added | **1279** | — |
| `pdf` documents where the reader finds nothing | **7** | 32 |
On the order's spike document the reader reproduces the PM's numbers exactly:
**9 of 9** numbered chapters found on 11 pages in 1.1 s, with **4 extra
candidates** -- two 20 pt title lines and two 16 pt contents lines.
### Per `pdf` position, against the unit worksheet
`N` is the shipped default (`--outline-run 3 --table-grid --unit-fold`); `G` is
that default plus `--pdf-headings font`. Concepts per document, reference from
the operator's worksheet:
| pos | reference | N | **G** | verdict |
|---|---|---|---|---|
| 0 | undefined | 1 | 1 | unchanged (the reader finds nothing in `(cid:N)`) |
| 1 | 15 | 17 | **5** | worse, distance 2 -> 10 |
| 2 | 9 | **9** `treffer` | 12 | **hit lost** |
| 4 | 1 | 4 | 11 | worse, distance 3 -> 10 |
| 6 | 6 | 7 | 11 | worse, distance 1 -> 5 |
| 7 | 30 | 33 | 33 | unchanged |
| 8 | 8 | **8** `treffer` | 16 | **hit lost** |
| 9 | 4 | 10 | 12 | worse, distance 6 -> 8 |
**`pdf` goes 2 of 8 to 0 of 8.** `docx` (3 of 3) and `xlsx` (0 of 1) do not
move, which is what the whole-corpus screen predicts: the reader is PDF-only.
**The mechanism of the loss, so the row is not just a number.** At positions 2
and 8 the outline rule already recovers the document's own numbered chapters,
and the font reader ADDS the title page and the contents lines on top of them --
the four extras above. It is not that typography is wrong; it is that on these
documents the document's own numbering is already right, and a second heading
source can only add.
### The variant that asks the question the other way round
If typography competes with numbering rather than adding to it, the comparison
is the font reader with `--outline-run 0`. Measured, same cache, same criterion:
| configuration | `pdf` | `docx` | `xlsx` | whole sample |
|---|---|---|---|---|
| N, the shipped default | **2 of 8** | 3 of 3 | 0 of 1 | 5 of 12 |
| G = N + font | 0 of 8 | 3 of 3 | 0 of 1 | 3 of 12 |
| X = font instead of Arm D (`--outline-run 0`) | 1 of 8 | 3 of 3 | 0 of 1 | 4 of 12 |
| F = the three round-3 rules on top of N | **3 of 8** | 3 of 3 | 0 of 1 | **6 of 12** |
| H = F + font | 0 of 8 | 3 of 3 | 0 of 1 | 3 of 12 |
**X is the one result worth reading twice.** It is worse overall, and it is the
first configuration in four rounds to land position 7 exactly on its reference
of **30** -- a document no arm has ever reached. Position 7's worksheet answer
(c) is `I`, "no arm is closest", so the criterion's LITERAL gate declares that
any `treffer` there is an instrument failure. **Both readings are stated
because the criterion's author never faced this case:** the gate's clause at
position 7 is a statement about the three arms the operator rated, and a
configuration that did not exist then reaching the reference count is new
information rather than evidence the instrument broke. It is not enough to make
X a default -- X loses positions 2, 6 and 8 to buy it.
## 4. Deliverable B: OCR behind `--ocr`, as an optional group
### The threshold, read off a measured distribution
Every page of every K2 PDF, `(cid:N)` characters as a share of the page's
extracted characters. Denominators: **32 `pdf` files, 834 pages that produced
text** (a page that produced none is not in the table; it is caught by the
other half of the trigger).
| share | pages |
|---|---|
| exactly 0.0 | **818** |
| 0 < s < 1 % | 0 |
| 1 % <= s < 10 % | 0 |
| 10 % <= s < 50 % | 0 |
| 50 % <= s < 90 % | 0 |
| s >= 90 % | **16** |
The distribution is bimodal with **nothing between 0.0 and 0.93**. The lowest
positive share measured is 0.9300 and the highest zero is 0.0.
**`OCR_CID_SHARE = 0.10`, and what happens at plus or minus ten points is not
symmetric.** At 0.20 the rule selects the identical 16 pages -- the gap is 93
percentage points wide, so every threshold inside it is the same rule. At 0.00
the rule stops being a rule: `share >= 0.0` is true of every page, and all 834
would be read as images. The threshold is therefore bounded by an empty region
on one side and by zero on the other, and 0.10 is the middle of the useful
range rather than a calibrated value. **This number is bounded by the corpus,
not by a property of the format**, and a corpus with intermediate pages would
have to set it again.
### The trigger has two conditions because there are two failures
A page is read as an image when its extracted text is **empty** or when its
`(cid:N)` share reaches the threshold. A rule written for one of them leaves
the other exactly where it was: a page with no text layer extracts as the empty
string and has no cid codes at all.
### What it recovers
The one document in the corpus with the failure, through the shipped code path
(`extract_text(..., ocr=True)`) and measured with the shipped instrument
(`tools/okf_cid_measure.measure`, imported rather than re-implemented):
| | before | after |
|---|---|---|
| characters | 217 470 | 34 534 |
| `(cid:N)` characters | 206 758 | **0** |
| cid share | **95.07 %** | **0.00 %** |
| words of 4+ letters | **44** | **2561** |
| pages that produced text | 17 of 18 | **18 of 18** |
| wall time | — | 65.3 s, 3.6 s/page |
The before column reproduces the published known-positive of
`docs/2026-09-04-cid-andel-k2.md` exactly on both figures (95.07 %, 44 words).
**The trigger fired on 16 of the 17 pages that had text, and on the one page
that had none.** The page it did not fire on carries 16 characters and no cid
codes -- text arrived, just very little of it. A length rule would catch it and
there is no measurement for one, so it is reported rather than repaired.
### Position 0 moves from "the text never came" to segmented
| configuration | concepts |
|---|---|
| position 0, shipped default, no OCR | **1** (no plan at all; the proposer finds no boundary) |
| position 0, shipped default, `--ocr` | **4** |
| position 0, the three round-3 rules, `--ocr` | **3** |
**The cell still cannot cross, and that is the reference's doing rather than
the reader's.** The worksheet's answer for position 0 is that no unit can be
decided, so the criterion makes `treffer` unreachable there by construction and
records the dominant error instead. The label moves from `for grovt` to `for
fint`; the `pdf` row is 8 positions either way.
### Why this engine, and what it costs
`rapidocr` on `onnxruntime`, in an optional group named `ocr` and never in
`project.dependencies` -- a packaging test asserts both halves. Install
measured at 31 s into an empty venv on this machine.
Docling was measured first and is out on a platform fact rather than a
preference: it needs torch, and torch stopped publishing macOS x86_64 wheels
after 2.2.2, with no `transformers` version inside Docling's own window that
works against that one.
The models ship **inside the `rapidocr` wheel** -- the run log names files under
`site-packages/rapidocr/models/` and no download step runs -- which matters more
here than it usually would: this library's network access is an explicit
per-run opt-in, and an engine that fetched a model on first use would walk
straight through that gate.
Without the group installed, `--ocr` is a typed refusal with its own code,
`extractor_ocr_group_missing`, distinct from `extractor_extra_missing` because
the `[extract]` extra can be fully installed with only the engine absent. In a
corpus run it is a coded rejection per file, so `merged + coded rejections == N`
still holds; a test pins that.
## 5. The default question, answered on numbers
**`--pdf-headings font` does not become the default.** The order's condition is
that it be at least as good on EVERY `pdf` position and regress nothing else.
It is worse on five positions, it loses two `treffer` cells outright, and the
label row goes 2 of 8 to 0 of 8. `DEFAULT_PDF_HEADINGS = False`, with a test.
**`--ocr` can never become a default**, and that is by construction rather than
by measurement: the engine is an optional dependency, so a default that needed
it would make an ordinary install fail on the first scanned page.
## 6. What round 3 left standing
**6.1 `pdf` position 1 (17 against 15).** The font axis is the first one not
tried here, and it is worse: 17 -> **5**, distance 2 -> 10. Declined for a
third round, now on three axes.
**6.2 positions 6, 7, 9.** Position 6 unchanged (7). Position 7 unchanged at 33
under the default, and reached exactly at **30** under X -- see § 3. Position 9
unchanged at 10 under the default, 8 under X. None of them crosses under a
configuration that does not lose more elsewhere.
**6.3 position 0.** Measured in § 4: 1 -> 4 concepts with `--ocr`, and the cell
still cannot cross because the reference says no unit can be decided.
**6.4 `xlsx` position 3.** Untouched: the font reader changes 0 of 2
spreadsheets. **The reading used here is the criterion as written** -- a
`rule:sheet-section` concept is neither `Tabell linje N` nor
`rule:table-block`/`rule:table-grid`, so the eleven cost groups count as
ordinary units and the distance is **1**, `for fint`. Under the opposite
reading the same plan scores distance 11. Round 3 stated both; this round
changes neither and picks the same one, out loud.
**6.5 the `docx` cell, re-measured on the current default.** Round 3's `docx`
**0 of 3** was measured on a default that no longer exists -- `--table-grid`
entered the default in `f6fea13`. On this HEAD the shipped default scores
`docx` **3 of 3**, as expected. Reported, not reversed.
**6.6 the order's premise here is false, and it is reported rather than acted
on.** The order says K3-3 § 0 promises "a third K2 bundle on the new default"
and that the report "contains the result nowhere". It contains it:
`docs/2026-09-08-k3-runde3-per-filtype.md` has a section titled *"The K2 bundle
on the new default"* whose table carries the row **517 concepts / 969 files /
`42c8646d…`**. Nothing was struck. What IS true, and is the useful half, is
that `42c8646d…` names the round-3 default (Arm D + Arm F) and the default
moved again in `f6fea13`, so the digest is stale. Section 8 supersedes it.
## 7. Per file type, after round 4
The 12-document sample, concepts per configuration, from one cached extraction
per rendering. **The instrument was verified before it was read**: re-run on the
archive plans the operator rated, session 109's criterion returns **LITERAL
7/12, DOMINANCE 11/12** -- its published gates.
| pos | type | reference | N (default) | **G (+font)** | F (round 3) | **H (F+font)** |
|---|---|---|---|---|---|---|
| 0 | pdf | — | 1 | 1 | 1 | 1 |
| 1 | pdf | 15 | 17 | 5 | 17 | 5 |
| 2 | pdf | 9 | **9** | 12 | **9** | 11 |
| 3 | xlsx | 11 | 1 | 1 | 12 | 12 |
| 4 | pdf | 1 | 4 | 11 | **1** | 8 |
| 5 | docx | 3 | **3** | **3** | **3** | **3** |
| 6 | pdf | 6 | 7 | 11 | 7 | 11 |
| 7 | pdf | 30 | 33 | 33 | 33 | 33 |
| 8 | pdf | 8 | **8** | 16 | **8** | 16 |
| 9 | pdf | 4 | 10 | 12 | 6 | 8 |
| 10 | docx | 2 | **2** | **2** | **2** | **2** |
| 11 | docx | 1 | **1** | **1** | **1** | **1** |
**The label row:**
| type | n | Arm B | N (default) | **G (+font)** | F (round 3) | **H (F+font)** |
|---|---|---|---|---|---|---|
| `pdf` | 8 | 2 | **2** | **0** | **3** | **0** |
| `docx` | 3 | 0 | **3** | **3** | **3** | **3** |
| `xlsx` | 1 | 0 | 0 | 0 | 0 | 0 |
| whole sample | 12 | 2 | 5 | 3 | **6** | 3 |
**One divergence between the rebuilt arms and the archive, named rather than
smoothed.** Arms B, D and E were rebuilt on this tree rather than read from the
archive, and they reproduce the archive's cells on 11 of 12 positions. Position
3 differs, and the reason is a known extraction change rather than an arm: the
spreadsheet writer moved in `56ae274`, after the archive plans were written. It
does not touch the N-versus-G comparison, which reads one cache with one code.
## 8. The K2 control
Two K2 bundles from the same 43-document corpus, same `--bundle-id
k2-trinn1-20260903 --okf-version 0.2`, the default epoch stamp, differing only
in the new flag. Both exit 0 and both report the conservation identity.
| | default | `--pdf-headings font` |
|---|---|---|
| `merged + coded rejections` | `39 + 4 = 43 = N` | idem |
| rejection codes | `extractor_empty_pdf` 1, `extractor_unknown` 3 | idem |
| concepts | **492** | 359 |
| markdown files | **944** | 595 |
| content digest | **`bdefa679957edd7d…`** | `a751b48f517f3ed4…` |
| wall time | 819.9 s | 814.7 s |
| concepts reachable from the index tree | 492 of 492 | 359 of 359 |
**The digest a consumer should compare against from today is
`bdefa679957edd7d…`, at 492 concepts in 944 files.** It supersedes round 3's
`42c8646d…` (517 / 969), which named the round-3 default before `--table-grid`
joined it in `f6fea13`. The command is the one above with no flags.
And the ranking, on the S7 control question, `--k 12`:
| | default | `--pdf-headings font` |
|---|---|---|
| priced concept present | yes, title `Prissammenstilling {#sheet-1}` | yes, same title |
| candidate rank, no vocabulary | 434 of 492 | 298 of 359 |
| candidate rank, `--cost-vocabulary` | **159 of 492** | 151 of 359 |
| delivered | **no** | **no** |
| payload bytes, vocabulary on | 123 681 | 95 446 |
**Nothing gets worse and nothing gets better: the flag moves the priced concept
from rank 159 of 492 to 151 of 359 and it is delivered in neither.** The bundle
it produces is 133 concepts smaller, which is a large change to buy nothing.
**One thing in that table is not about this flag and the operator should read
it.** Round 3 published the same question against a 629-concept Arm B bundle:
candidate rank **96 of 629** with the vocabulary on, **delivered at rank 10**.
On the shipped default it is rank 159 of 492 and **not delivered at all**. That
is the default move of `47f6619`/`f6fea13` showing up on the retrieval side for
the first time, not this round's flag -- the two columns above are the
controlled comparison, and they differ only in the flag. It is one question and
one bundle, and it is written down rather than left for the next round to
rediscover.
## 9. Tests
`1427 -> 1441`. Fourteen new, and which of them were red first is stated rather
than counted -- the first run of the new file was **8 failed, 1 passed**, and
the second, after the flags reached `okf build`, **3 failed, 9 passed**:
- **Red before the implementation (11):** the ATX emission and its level
mapping; the proposer segmenting it with no new rule; the known-negative (a
PDF with no font variation is byte-identical with the flag on); the locator
indexing the marked text; the OCR refusal without the group; OCR recovering a
page with no text; OCR leaving a page that has text alone; the cid threshold
and the two-condition trigger; the flag reaching `okf build`; the build-layer
default control; the coded-rejection accounting.
- **Green by construction, and named as such (3):** the "off by default"
extraction control, the new code's entry in the error-code registry suite, and
the `ocr` group's pin. All three were written after the code they describe.
One fixture is new -- `tests/fixtures/font-heading-krav.pdf`, a hand-laid
two-font PDF regenerated by `tests/fixtures/make_fixtures.py`. It gets its own
builder rather than a parameter on the existing one: a second font object
changes the object numbering, and every other PDF fixture's value is that its
bytes have not moved. Regenerating writes all nine fixtures and `git status`
shows only the new one.
No shipped expectation moved. The README's segmentation section did move, and
that is a correction rather than a change: it still said **two** rules were on
by default after `f6fea13` made it three.
## Reproducing
```
# 1. The three reproduction numbers.
pytest -q
okf build ~/okf-test/dokumenter --bundle /tmp/b --bundle-id t --okf-version 0.2
okf build ~/okf-test/dokumenter --bundle /tmp/b2 --bundle-id t --okf-version 0.2 \
--no-table-grid
# 2. The per-page cid distribution, both denominators.
python3 /tmp/k3r4/cidpages.py # 32 pdf files, 834 pages with text
# 3. The two renderings of the whole corpus, one cached extraction each.
python3 /tmp/k3r4/font43.py # writes /tmp/k3r4/font43/
# the control is line-by-line: same line count, every differing line is the
# cached line with a hash run prefixed -> 39 of 39
# 4. The 12-sample, scored with session 109's own criterion.
python3 /tmp/k3r4/plans4.py && python3 /tmp/k3r4/score4.py
python3 /tmp/k3r4/score4b.py # the X and Y variants
# 5. What OCR recovers. Needs the optional group in a venv of its own:
# uv pip install rapidocr onnxruntime pypdfium2 pdfplumber
python3 /tmp/k3r4/ocr91.py
# 6. The K2 pair. Each build is a foreground nohup with an rc file and a
# bounded poll -- the Bash tool caps at 600 s and a build takes ~15-25 min.
/tmp/k3r4/k2.sh default
/tmp/k3r4/k2.sh font --pdf-headings font
```
## Honesty limits
- **A font-inferred heading is indistinguishable from a declared one in the
artifact.** Emitting ATX is what makes `_ATX` apply unchanged, and it is also
what makes the plan say `rule:heading` for a heading no document ever
declared. `RULE_POPPLER_SIZE_AND_BOLD` was deliberately NOT assigned to it:
that name records a poppler measurement on a path that cannot ship, and
putting it on a pdfplumber inference would make the artifact claim a
provenance it does not have. The cost is real and it is the strongest
argument against this design; it is affordable only because the flag is off.
- **n = 12, one rater, one reference.** Per file type the denominators are 8, 3
and 1. The `xlsx` row is a `0/1`, which is not a rate.
- **The OCR denominator is 1.** One document of 39 has the `(cid:N)` failure and
one has no text layer at all. Every claim about what OCR buys is a claim about
those two documents.
- **OCR output is not deterministic across versions.** It is a model's reading
of an image at a rendering resolution. It is reproducible against the model
version and `OCR_DPI` it was produced with, and no dependency pin can promise
more. The `ocr` group therefore uses ranges and says why.
- **`--ocr` and `--pdf-headings font` do not compose on an OCR'd page.** The
engine reports text, not fonts, so a recovered page carries no typography and
can contribute no heading. Both flags together are valid and each does its own
work on its own pages.
- **The cid threshold is bounded by a corpus, not by the format.** Every value
between 0.01 and 0.93 selects the same 16 pages here. A corpus with pages in
between would need the number set again, from its own distribution.
- **The 43-document corpus and the 12-document sample are the same corpus.** The
sample is drawn from it, so the whole-corpus screen is not an independent
check of the sample result.