# K3 round 4, `pdf`: typography as a heading source, and OCR for the pages whose text never came Fourth iteration of the per-file-type directive of 2026-09-08 13:05Z. Order `20260908T195728Z-8160520552-from-.claude`. HEAD at the start and at the end of the session is O5's `f6fea13` plus this session's commit; the tag on `f6fea13` is `v0.6.0`. Everything below is measured on the same 43-document corpus and the same 12-document sample as `docs/2026-09-08-k3-runde3-per-filtype.md`, against the same reference: the operator's filled unit worksheet, which is local-only and gitignored. No document title or extracted text from it is reproduced here. ## Section 0: what was measured and what was not **Measured.** The three reproduction numbers the order names, before any edit. Two new readers, each red-first, each with a known-negative. The per-page `(cid:N)` distribution over every K2 PDF, with both denominators, and the threshold read off it rather than chosen. What OCR recovers on the one document that has the failure, through the shipped code path and the shipped instrument. The 12-document sample re-scored with session 109's criterion under seven configurations. The whole 43-document corpus under both renderings from one cached extraction. Two K2 bundles differing only in the new flag. **Not measured.** Whether a live model answers better with any of this. Whether OCR helps a corpus that is not this one -- the `(cid:N)` denominator that matters is **1 document of 39**. Whether the font reader helps a corpus whose typography is unlike this one. No o200k token counts (no encoder is installed and the single-dependency rule does not admit one). ## 1. The reproduction, before any edit `git rev-parse HEAD` -> `f6fea13299d92b959f497652ff19a7fd79546281`, `git status --short` empty, tag `v0.6.0` on that commit. | # | control | expected | result | |---|---|---|---| | a | `pytest -q` | 1427 collected | **1427 passed** in 38 s | | b | `okf build ~/okf-test/dokumenter`, no flags | 30 | **30** markdown files | | c | the same folder, `--no-table-grid` | 43 | **43** markdown files | All three reproduce. (a) differs in FORM from the PM's figure and not in substance: on a frozen export the same suite is 1424 passed, 1 skipped and 2 export artefacts; in a working tree all 1427 run and pass. ## 2. The root cause, verified rather than inherited The order states a diagnosis. Each half was checked against the tree before any code was written, because a premise is an input like any other: | claim | check | result | |---|---|---| | PDF reaches the proposer as a plain text stream | `_pdf_pages` called `page.extract_text()` and nothing else | confirmed | | the `docx` path emits ATX the proposer already cuts on | `_PANDOC_WRITER = "markdown"`, never `plain` | confirmed | | the measured best deterministic heading rule is not on the run path | `RULE_POPPLER_SIZE_AND_BOLD` occurs **twice** in `src/`, both in `propose.py` -- its definition and its entry in `RULE_NAMES` -- and **zero times** in `tools/` or `tests/` | confirmed: no candidate has ever carried it | A PDF has no notion of a heading. A heading in a PDF is a typographic fact, and the text stream throws that fact away before anything downstream can read it. ## 3. Deliverable A: the font reader, behind `--pdf-headings font` ### The mechanism `pdfplumber` -- already a dependency -- reports `size` and `fontname` per character. Per line, the reader takes the dominant size and dominant font name over the line's non-blank characters. A line is a heading when its dominant size is **above the document's character-weighted median size** AND its dominant font name contains `bold`. The heading sizes of the whole document are ranked largest-first and that rank becomes the ATX level, clamped at six. Three choices, each with a reason rather than a taste: - **A conjunction, not a disjunction.** `propose.py` already records the measurement: `size AND bold` scored recall 1.000 / precision 0.846, and adding weight as a disjunct took precision from 0.786 to 0.524. - **The body size is the document's, not the page's.** A title page is 100 % heading, so a per-page median compares it with itself and marks nothing. - **The output is ATX in the same markdown the office path produces**, so `_ATX` reads it unchanged and no PDF-only heading grammar exists. A second grammar would be a second definition of "heading" free to drift from the one five office formats already use. The cost is stated in Honesty limits. ### The controls | control | result | |---|---| | a PDF with one font at one size, flag ON | **byte-identical** to flag OFF (test) | | `.docx` / `.xlsx`, flag ON, whole corpus | **0 of 5** and **0 of 2** changed | | every readable corpus document: same line count, and every differing line is the cached line with a hash run prefixed | **39 of 39** | That last control is the load-bearing one and it states three things at once: the round-3 cache is still valid (`extract.py` is byte-identical between `c3b645b` and `f6fea13`), joining `extract_text_lines()` reproduces `extract_text()` on real documents rather than on the one it was spot-checked on, and the reader **adds markers and changes nothing else**. **The first version of that control was wrong and said so.** It stripped `^#{1,6} ` from the marked text and compared: 38 of 39. The one failure is a corpus document whose own extracted text carries lines beginning with `# `, and the stripper removed those too. The corrected control compares line by line. ### The reach Whole corpus, 39 readable documents, one cached extraction per rendering: | figure | value | denominator | |---|---|---| | documents changed | **25** | 32 `pdf` | | documents changed | **0** | 5 `docx` | | documents changed | **0** | 2 `xlsx` | | ATX headings added | **1279** | — | | `pdf` documents where the reader finds nothing | **7** | 32 | On the order's spike document the reader reproduces the PM's numbers exactly: **9 of 9** numbered chapters found on 11 pages in 1.1 s, with **4 extra candidates** -- two 20 pt title lines and two 16 pt contents lines. ### Per `pdf` position, against the unit worksheet `N` is the shipped default (`--outline-run 3 --table-grid --unit-fold`); `G` is that default plus `--pdf-headings font`. Concepts per document, reference from the operator's worksheet: | pos | reference | N | **G** | verdict | |---|---|---|---|---| | 0 | undefined | 1 | 1 | unchanged (the reader finds nothing in `(cid:N)`) | | 1 | 15 | 17 | **5** | worse, distance 2 -> 10 | | 2 | 9 | **9** `treffer` | 12 | **hit lost** | | 4 | 1 | 4 | 11 | worse, distance 3 -> 10 | | 6 | 6 | 7 | 11 | worse, distance 1 -> 5 | | 7 | 30 | 33 | 33 | unchanged | | 8 | 8 | **8** `treffer` | 16 | **hit lost** | | 9 | 4 | 10 | 12 | worse, distance 6 -> 8 | **`pdf` goes 2 of 8 to 0 of 8.** `docx` (3 of 3) and `xlsx` (0 of 1) do not move, which is what the whole-corpus screen predicts: the reader is PDF-only. **The mechanism of the loss, so the row is not just a number.** At positions 2 and 8 the outline rule already recovers the document's own numbered chapters, and the font reader ADDS the title page and the contents lines on top of them -- the four extras above. It is not that typography is wrong; it is that on these documents the document's own numbering is already right, and a second heading source can only add. ### The variant that asks the question the other way round If typography competes with numbering rather than adding to it, the comparison is the font reader with `--outline-run 0`. Measured, same cache, same criterion: | configuration | `pdf` | `docx` | `xlsx` | whole sample | |---|---|---|---|---| | N, the shipped default | **2 of 8** | 3 of 3 | 0 of 1 | 5 of 12 | | G = N + font | 0 of 8 | 3 of 3 | 0 of 1 | 3 of 12 | | X = font instead of Arm D (`--outline-run 0`) | 1 of 8 | 3 of 3 | 0 of 1 | 4 of 12 | | F = the three round-3 rules on top of N | **3 of 8** | 3 of 3 | 0 of 1 | **6 of 12** | | H = F + font | 0 of 8 | 3 of 3 | 0 of 1 | 3 of 12 | **X is the one result worth reading twice.** It is worse overall, and it is the first configuration in four rounds to land position 7 exactly on its reference of **30** -- a document no arm has ever reached. Position 7's worksheet answer (c) is `I`, "no arm is closest", so the criterion's LITERAL gate declares that any `treffer` there is an instrument failure. **Both readings are stated because the criterion's author never faced this case:** the gate's clause at position 7 is a statement about the three arms the operator rated, and a configuration that did not exist then reaching the reference count is new information rather than evidence the instrument broke. It is not enough to make X a default -- X loses positions 2, 6 and 8 to buy it. ## 4. Deliverable B: OCR behind `--ocr`, as an optional group ### The threshold, read off a measured distribution Every page of every K2 PDF, `(cid:N)` characters as a share of the page's extracted characters. Denominators: **32 `pdf` files, 834 pages that produced text** (a page that produced none is not in the table; it is caught by the other half of the trigger). | share | pages | |---|---| | exactly 0.0 | **818** | | 0 < s < 1 % | 0 | | 1 % <= s < 10 % | 0 | | 10 % <= s < 50 % | 0 | | 50 % <= s < 90 % | 0 | | s >= 90 % | **16** | The distribution is bimodal with **nothing between 0.0 and 0.93**. The lowest positive share measured is 0.9300 and the highest zero is 0.0. **`OCR_CID_SHARE = 0.10`, and what happens at plus or minus ten points is not symmetric.** At 0.20 the rule selects the identical 16 pages -- the gap is 93 percentage points wide, so every threshold inside it is the same rule. At 0.00 the rule stops being a rule: `share >= 0.0` is true of every page, and all 834 would be read as images. The threshold is therefore bounded by an empty region on one side and by zero on the other, and 0.10 is the middle of the useful range rather than a calibrated value. **This number is bounded by the corpus, not by a property of the format**, and a corpus with intermediate pages would have to set it again. ### The trigger has two conditions because there are two failures A page is read as an image when its extracted text is **empty** or when its `(cid:N)` share reaches the threshold. A rule written for one of them leaves the other exactly where it was: a page with no text layer extracts as the empty string and has no cid codes at all. ### What it recovers The one document in the corpus with the failure, through the shipped code path (`extract_text(..., ocr=True)`) and measured with the shipped instrument (`tools/okf_cid_measure.measure`, imported rather than re-implemented): | | before | after | |---|---|---| | characters | 217 470 | 34 534 | | `(cid:N)` characters | 206 758 | **0** | | cid share | **95.07 %** | **0.00 %** | | words of 4+ letters | **44** | **2561** | | pages that produced text | 17 of 18 | **18 of 18** | | wall time | — | 65.3 s, 3.6 s/page | The before column reproduces the published known-positive of `docs/2026-09-04-cid-andel-k2.md` exactly on both figures (95.07 %, 44 words). **The trigger fired on 16 of the 17 pages that had text, and on the one page that had none.** The page it did not fire on carries 16 characters and no cid codes -- text arrived, just very little of it. A length rule would catch it and there is no measurement for one, so it is reported rather than repaired. ### Position 0 moves from "the text never came" to segmented | configuration | concepts | |---|---| | position 0, shipped default, no OCR | **1** (no plan at all; the proposer finds no boundary) | | position 0, shipped default, `--ocr` | **4** | | position 0, the three round-3 rules, `--ocr` | **3** | **The cell still cannot cross, and that is the reference's doing rather than the reader's.** The worksheet's answer for position 0 is that no unit can be decided, so the criterion makes `treffer` unreachable there by construction and records the dominant error instead. The label moves from `for grovt` to `for fint`; the `pdf` row is 8 positions either way. ### Why this engine, and what it costs `rapidocr` on `onnxruntime`, in an optional group named `ocr` and never in `project.dependencies` -- a packaging test asserts both halves. Install measured at 31 s into an empty venv on this machine. Docling was measured first and is out on a platform fact rather than a preference: it needs torch, and torch stopped publishing macOS x86_64 wheels after 2.2.2, with no `transformers` version inside Docling's own window that works against that one. The models ship **inside the `rapidocr` wheel** -- the run log names files under `site-packages/rapidocr/models/` and no download step runs -- which matters more here than it usually would: this library's network access is an explicit per-run opt-in, and an engine that fetched a model on first use would walk straight through that gate. Without the group installed, `--ocr` is a typed refusal with its own code, `extractor_ocr_group_missing`, distinct from `extractor_extra_missing` because the `[extract]` extra can be fully installed with only the engine absent. In a corpus run it is a coded rejection per file, so `merged + coded rejections == N` still holds; a test pins that. ## 5. The default question, answered on numbers **`--pdf-headings font` does not become the default.** The order's condition is that it be at least as good on EVERY `pdf` position and regress nothing else. It is worse on five positions, it loses two `treffer` cells outright, and the label row goes 2 of 8 to 0 of 8. `DEFAULT_PDF_HEADINGS = False`, with a test. **`--ocr` can never become a default**, and that is by construction rather than by measurement: the engine is an optional dependency, so a default that needed it would make an ordinary install fail on the first scanned page. ## 6. What round 3 left standing **6.1 `pdf` position 1 (17 against 15).** The font axis is the first one not tried here, and it is worse: 17 -> **5**, distance 2 -> 10. Declined for a third round, now on three axes. **6.2 positions 6, 7, 9.** Position 6 unchanged (7). Position 7 unchanged at 33 under the default, and reached exactly at **30** under X -- see § 3. Position 9 unchanged at 10 under the default, 8 under X. None of them crosses under a configuration that does not lose more elsewhere. **6.3 position 0.** Measured in § 4: 1 -> 4 concepts with `--ocr`, and the cell still cannot cross because the reference says no unit can be decided. **6.4 `xlsx` position 3.** Untouched: the font reader changes 0 of 2 spreadsheets. **The reading used here is the criterion as written** -- a `rule:sheet-section` concept is neither `Tabell linje N` nor `rule:table-block`/`rule:table-grid`, so the eleven cost groups count as ordinary units and the distance is **1**, `for fint`. Under the opposite reading the same plan scores distance 11. Round 3 stated both; this round changes neither and picks the same one, out loud. **6.5 the `docx` cell, re-measured on the current default.** Round 3's `docx` **0 of 3** was measured on a default that no longer exists -- `--table-grid` entered the default in `f6fea13`. On this HEAD the shipped default scores `docx` **3 of 3**, as expected. Reported, not reversed. **6.6 the order's premise here is false, and it is reported rather than acted on.** The order says K3-3 § 0 promises "a third K2 bundle on the new default" and that the report "contains the result nowhere". It contains it: `docs/2026-09-08-k3-runde3-per-filtype.md` has a section titled *"The K2 bundle on the new default"* whose table carries the row **517 concepts / 969 files / `42c8646d…`**. Nothing was struck. What IS true, and is the useful half, is that `42c8646d…` names the round-3 default (Arm D + Arm F) and the default moved again in `f6fea13`, so the digest is stale. Section 8 supersedes it. ## 7. Per file type, after round 4 The 12-document sample, concepts per configuration, from one cached extraction per rendering. **The instrument was verified before it was read**: re-run on the archive plans the operator rated, session 109's criterion returns **LITERAL 7/12, DOMINANCE 11/12** -- its published gates. | pos | type | reference | N (default) | **G (+font)** | F (round 3) | **H (F+font)** | |---|---|---|---|---|---|---| | 0 | pdf | — | 1 | 1 | 1 | 1 | | 1 | pdf | 15 | 17 | 5 | 17 | 5 | | 2 | pdf | 9 | **9** | 12 | **9** | 11 | | 3 | xlsx | 11 | 1 | 1 | 12 | 12 | | 4 | pdf | 1 | 4 | 11 | **1** | 8 | | 5 | docx | 3 | **3** | **3** | **3** | **3** | | 6 | pdf | 6 | 7 | 11 | 7 | 11 | | 7 | pdf | 30 | 33 | 33 | 33 | 33 | | 8 | pdf | 8 | **8** | 16 | **8** | 16 | | 9 | pdf | 4 | 10 | 12 | 6 | 8 | | 10 | docx | 2 | **2** | **2** | **2** | **2** | | 11 | docx | 1 | **1** | **1** | **1** | **1** | **The label row:** | type | n | Arm B | N (default) | **G (+font)** | F (round 3) | **H (F+font)** | |---|---|---|---|---|---|---| | `pdf` | 8 | 2 | **2** | **0** | **3** | **0** | | `docx` | 3 | 0 | **3** | **3** | **3** | **3** | | `xlsx` | 1 | 0 | 0 | 0 | 0 | 0 | | whole sample | 12 | 2 | 5 | 3 | **6** | 3 | **One divergence between the rebuilt arms and the archive, named rather than smoothed.** Arms B, D and E were rebuilt on this tree rather than read from the archive, and they reproduce the archive's cells on 11 of 12 positions. Position 3 differs, and the reason is a known extraction change rather than an arm: the spreadsheet writer moved in `56ae274`, after the archive plans were written. It does not touch the N-versus-G comparison, which reads one cache with one code. ## 8. The K2 control Two K2 bundles from the same 43-document corpus, same `--bundle-id k2-trinn1-20260903 --okf-version 0.2`, the default epoch stamp, differing only in the new flag. Both exit 0 and both report the conservation identity. | | default | `--pdf-headings font` | |---|---|---| | `merged + coded rejections` | `39 + 4 = 43 = N` | idem | | rejection codes | `extractor_empty_pdf` 1, `extractor_unknown` 3 | idem | | concepts | **492** | 359 | | markdown files | **944** | 595 | | content digest | **`bdefa679957edd7d…`** | `a751b48f517f3ed4…` | | wall time | 819.9 s | 814.7 s | | concepts reachable from the index tree | 492 of 492 | 359 of 359 | **The digest a consumer should compare against from today is `bdefa679957edd7d…`, at 492 concepts in 944 files.** It supersedes round 3's `42c8646d…` (517 / 969), which named the round-3 default before `--table-grid` joined it in `f6fea13`. The command is the one above with no flags. And the ranking, on the S7 control question, `--k 12`: | | default | `--pdf-headings font` | |---|---|---| | priced concept present | yes, title `Prissammenstilling {#sheet-1}` | yes, same title | | candidate rank, no vocabulary | 434 of 492 | 298 of 359 | | candidate rank, `--cost-vocabulary` | **159 of 492** | 151 of 359 | | delivered | **no** | **no** | | payload bytes, vocabulary on | 123 681 | 95 446 | **Nothing gets worse and nothing gets better: the flag moves the priced concept from rank 159 of 492 to 151 of 359 and it is delivered in neither.** The bundle it produces is 133 concepts smaller, which is a large change to buy nothing. **One thing in that table is not about this flag and the operator should read it.** Round 3 published the same question against a 629-concept Arm B bundle: candidate rank **96 of 629** with the vocabulary on, **delivered at rank 10**. On the shipped default it is rank 159 of 492 and **not delivered at all**. That is the default move of `47f6619`/`f6fea13` showing up on the retrieval side for the first time, not this round's flag -- the two columns above are the controlled comparison, and they differ only in the flag. It is one question and one bundle, and it is written down rather than left for the next round to rediscover. ## 9. Tests `1427 -> 1441`. Fourteen new, and which of them were red first is stated rather than counted -- the first run of the new file was **8 failed, 1 passed**, and the second, after the flags reached `okf build`, **3 failed, 9 passed**: - **Red before the implementation (11):** the ATX emission and its level mapping; the proposer segmenting it with no new rule; the known-negative (a PDF with no font variation is byte-identical with the flag on); the locator indexing the marked text; the OCR refusal without the group; OCR recovering a page with no text; OCR leaving a page that has text alone; the cid threshold and the two-condition trigger; the flag reaching `okf build`; the build-layer default control; the coded-rejection accounting. - **Green by construction, and named as such (3):** the "off by default" extraction control, the new code's entry in the error-code registry suite, and the `ocr` group's pin. All three were written after the code they describe. One fixture is new -- `tests/fixtures/font-heading-krav.pdf`, a hand-laid two-font PDF regenerated by `tests/fixtures/make_fixtures.py`. It gets its own builder rather than a parameter on the existing one: a second font object changes the object numbering, and every other PDF fixture's value is that its bytes have not moved. Regenerating writes all nine fixtures and `git status` shows only the new one. No shipped expectation moved. The README's segmentation section did move, and that is a correction rather than a change: it still said **two** rules were on by default after `f6fea13` made it three. ## Reproducing ``` # 1. The three reproduction numbers. pytest -q okf build ~/okf-test/dokumenter --bundle /tmp/b --bundle-id t --okf-version 0.2 okf build ~/okf-test/dokumenter --bundle /tmp/b2 --bundle-id t --okf-version 0.2 \ --no-table-grid # 2. The per-page cid distribution, both denominators. python3 /tmp/k3r4/cidpages.py # 32 pdf files, 834 pages with text # 3. The two renderings of the whole corpus, one cached extraction each. python3 /tmp/k3r4/font43.py # writes /tmp/k3r4/font43/ # the control is line-by-line: same line count, every differing line is the # cached line with a hash run prefixed -> 39 of 39 # 4. The 12-sample, scored with session 109's own criterion. python3 /tmp/k3r4/plans4.py && python3 /tmp/k3r4/score4.py python3 /tmp/k3r4/score4b.py # the X and Y variants # 5. What OCR recovers. Needs the optional group in a venv of its own: # uv pip install rapidocr onnxruntime pypdfium2 pdfplumber python3 /tmp/k3r4/ocr91.py # 6. The K2 pair. Each build is a foreground nohup with an rc file and a # bounded poll -- the Bash tool caps at 600 s and a build takes ~15-25 min. /tmp/k3r4/k2.sh default /tmp/k3r4/k2.sh font --pdf-headings font ``` ## Honesty limits - **A font-inferred heading is indistinguishable from a declared one in the artifact.** Emitting ATX is what makes `_ATX` apply unchanged, and it is also what makes the plan say `rule:heading` for a heading no document ever declared. `RULE_POPPLER_SIZE_AND_BOLD` was deliberately NOT assigned to it: that name records a poppler measurement on a path that cannot ship, and putting it on a pdfplumber inference would make the artifact claim a provenance it does not have. The cost is real and it is the strongest argument against this design; it is affordable only because the flag is off. - **n = 12, one rater, one reference.** Per file type the denominators are 8, 3 and 1. The `xlsx` row is a `0/1`, which is not a rate. - **The OCR denominator is 1.** One document of 39 has the `(cid:N)` failure and one has no text layer at all. Every claim about what OCR buys is a claim about those two documents. - **OCR output is not deterministic across versions.** It is a model's reading of an image at a rendering resolution. It is reproducible against the model version and `OCR_DPI` it was produced with, and no dependency pin can promise more. The `ocr` group therefore uses ranges and says why. - **`--ocr` and `--pdf-headings font` do not compose on an OCR'd page.** The engine reports text, not fonts, so a recovered page carries no typography and can contribute no heading. Both flags together are valid and each does its own work on its own pages. - **The cid threshold is bounded by a corpus, not by the format.** Every value between 0.01 and 0.93 selects the same 16 pages here. A corpus with pages in between would need the number set again, from its own distribution. - **The 43-document corpus and the 12-document sample are the same corpus.** The sample is drawn from it, so the whole-corpus screen is not an independent check of the sample result.