feat(extract): implement pdf behind the [extract] extra with pdfplumber
Order G2a. Populates the optional `[extract]` extra for the first time with
one parser, `pdfplumber>=0.11.10,<0.12` (MIT), and wires `pdf` through it.
The default install is untouched: exactly one runtime dependency, stdlib
otherwise, enforced by test_packaging.py.
The gate for `pdf` becomes an import probe rather than a frozenset membership
test, exactly as extract.py's docstring had promised. The rejection does not
change: without the extra, `pdf` still raises `extractor_extra_missing` with
the same message. That behaviour is asserted UNCONDITIONALLY via a sys.modules
monkeypatch, so it holds on machines where the parser is installed too — a
skip would have preserved nothing there. Verified in a clean venv without the
extra: 589 passed, 7 skipped; with it, 596 passed.
`docx`/`xlsx` are unchanged and still fail fast — the extra names exactly what
it ships.
The parser choice was forced by measurement, not preference (b73dd9d,
docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real requirement table
pdfplumber keeps 4 of 4 rows with label and value on one line, where pypdf,
pdfminer.six and pymupdf each keep 0 of 4. pymupdf is additionally out on
licence (AGPL-3.0), which an MIT package must not push onto a consumer.
Three facts from that measurement are now carried in code rather than in a
report:
- Extracted text is pinned to an exact transitive parser version
(pdfplumber pins pdfminer.six==20260107; date-stamped, no stability
contract). tests/test_extract.py freezes the expected text of a committed
hand-written fixture so a parser upgrade breaks something visible instead of
drifting silently. Reasoning at the declaration site and in
tests/fixtures/README.md.
- Determinism within a version is now held by a test, not only measured once.
- Drawn content does not survive extraction. Every pdf extraction emits the
new `ExtractionWarning`: figures have no text to recover, so a bundle built
from drawn documents is incomplete by construction. Stated categorically
rather than detected — deciding "is there a figure here" is the layout
heuristic G2b declined.
Two new error codes, both mirroring existing patterns: `extractor_empty_pdf`
(a scanned/image-only PDF, refused rather than persisted as an empty concept)
and `extractor_pdf_error` (parser failure wrapped, never leaked).
Structured table recovery (G2b) is NOT implemented and is documented as out of
scope: two independent parsers return the same wrong shape, so the breakage is
document geometry, not a library choice. PDFs enter as prose.
Also corrects an install promise this change would otherwise have published:
the README no longer presents a bare `pip install 'llm-ingestion-okf[extract]'`
as working, because the package is not on an index.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtNhsdHnMGtMi7U2mvMU8z
This commit is contained in:
parent
b73dd9d6a4
commit
658b7aafe0
14 changed files with 1046 additions and 37 deletions
|
|
@ -25,9 +25,33 @@ classifiers = [
|
|||
dependencies = ["llm-ingestion-guard>=0.3,<0.4"]
|
||||
|
||||
[project.optional-dependencies]
|
||||
# Reserved for binary file-type extraction parsers (pdf/docx/xlsx).
|
||||
# Populated when Door B's binary extraction is implemented.
|
||||
extract = []
|
||||
# Binary file-type extraction parsers. OPT-IN ONLY: this extra pulls binary
|
||||
# wheels (pillow, pypdfium2) and a transitive tree that core must never have —
|
||||
# the "exactly one runtime dependency" rule above covers the default install,
|
||||
# and this extra is outside it by construction.
|
||||
#
|
||||
# `pdf` only. `docx`/`xlsx` remain fail-fast: the extra names the parsers it
|
||||
# actually ships, so a consumer installing it gets what the error message
|
||||
# promised and nothing else.
|
||||
#
|
||||
# WHY pdfplumber, and why the floor is not free (measured 2026-08-21,
|
||||
# docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real Vegnormalene
|
||||
# requirement table pdfplumber keeps 4 of 4 rows with label and value on the
|
||||
# same line; pypdf, pdfminer.six and pymupdf each keep 0 of 4, emitting all
|
||||
# labels then all values, which a downstream reader can only re-pair by
|
||||
# guessing. In a `krav` document that is a wrong answer that looks right.
|
||||
# pymupdf is additionally out on LICENSE (AGPL-3.0 or commercial) — this
|
||||
# package is MIT and an extra must not hand a consumer copyleft they did not
|
||||
# choose.
|
||||
#
|
||||
# PARSER VERSION IS PART OF THE OUTPUT CONTRACT. pdfplumber pins
|
||||
# `pdfminer.six==20260107` exactly, and pdfminer.six ships date-stamped
|
||||
# releases with no stability contract. Extraction is deterministic WITHIN a
|
||||
# parser version (measured, 5 configurations) and NOT guaranteed across one.
|
||||
# `tests/test_extract.py` holds that promise against a committed fixture, so
|
||||
# widening this range makes a test go red instead of letting extracted text
|
||||
# drift silently. See tests/fixtures/README.md.
|
||||
extract = ["pdfplumber>=0.11.10,<0.12"]
|
||||
|
||||
[dependency-groups]
|
||||
dev = ["pytest>=8", "mypy>=1.14", "ruff>=0.9"]
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue