feat(extract,cli): typography as a PDF heading source and OCR behind an optional group, both off
A PDF carries no notion of a heading -- a heading in a PDF is a typographic
fact -- so the text stream `pdfplumber` hands the segment proposer has already
thrown away the only evidence there was. The `docx` path never had that problem:
the converter emits ATX headings and `_ATX` cuts on them. Two readers close the
gap, and both are OFF.
`--pdf-headings font` infers a heading from the conjunction this repository
already measured (size above the document's character-weighted body median AND
a bold font name, recall 1.000 / precision 0.846) and emits it as ATX in the
SAME markdown the office path produces, so `_ATX` applies unchanged and no
PDF-only heading grammar exists.
It stays off BY MEASUREMENT, and the measurement is the point of the round:
against the operator's unit worksheet it takes `pdf` from 2 of 8 to 0 of 8,
losing two exact matches. The mechanism of the loss is stated rather than
guessed -- on those documents the outline rule already recovers the document's
own numbered chapters, so a second heading source can only add. Whole-corpus
screen: 25 of 32 `pdf` change, 0 of 5 `docx`, 0 of 2 `xlsx`. The default bundle
is byte-identical before and after this commit (`diff -r`, exit 0).
`--ocr` reads a page as an image when its own text never arrived: empty, or
`(cid:N)` placeholder codes at or above a threshold READ OFF a measured
distribution -- 834 pages over 32 files, 818 at exactly 0.0 and 16 at 0.93 or
above, nothing in between. On the one corpus document with the failure: 95.07 %
cid to 0 %, 44 to 2561 words of four or more letters, 17 to 18 pages with text.
Its engine is an optional dependency group and never a runtime dependency; a
packaging test pins both halves, and without the group every affected file is a
coded rejection (`extractor_ocr_group_missing`) rather than a crash.
Also corrects two stale published facts found while measuring: the README still
said two segmentation rules were on by default after `f6fea13` made it three,
and CLAUDE.md's K2 digest named the round-3 default. The current default is
492 concepts / 944 files, `bdefa679...`.
Report: docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md
Co-Authored-By: Claude <claude-opus-5>
This commit is contained in:
parent
f6fea13299
commit
53d5c74c96
15 changed files with 1394 additions and 28 deletions
|
|
@ -57,6 +57,27 @@ def test_the_extract_extra_pins_exactly_what_it_ships() -> None:
|
|||
]
|
||||
|
||||
|
||||
def test_the_ocr_group_is_pinned_and_is_not_a_runtime_dependency() -> None:
|
||||
"""An inference runtime is the last thing that may arrive by accident.
|
||||
|
||||
Two claims, and the second is the one worth a test: the group's contents
|
||||
are pinned like the extra's, AND none of them appears in
|
||||
`project.dependencies`. The single-dependency test above would already
|
||||
catch that, but it reads the list and this reads the names -- so a future
|
||||
entry named differently still fails here.
|
||||
"""
|
||||
tomllib = pytest.importorskip("tomllib")
|
||||
pyproject = tomllib.loads((PROJECT_ROOT / "pyproject.toml").read_text(encoding="utf-8"))
|
||||
assert pyproject["project"]["optional-dependencies"]["ocr"] == [
|
||||
"rapidocr>=3.9,<4",
|
||||
"onnxruntime>=1.20,<2",
|
||||
"pypdfium2>=4,<6",
|
||||
]
|
||||
runtime = " ".join(pyproject["project"]["dependencies"])
|
||||
for package in ("rapidocr", "onnxruntime", "pypdfium2"):
|
||||
assert package not in runtime
|
||||
|
||||
|
||||
def test_the_declared_version_agrees_with_the_packaged_one() -> None:
|
||||
"""The two places a version is written must not drift apart.
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue