feat(extract): implement pdf behind the [extract] extra with pdfplumber

Order G2a. Populates the optional `[extract]` extra for the first time with
one parser, `pdfplumber>=0.11.10,<0.12` (MIT), and wires `pdf` through it.
The default install is untouched: exactly one runtime dependency, stdlib
otherwise, enforced by test_packaging.py.

The gate for `pdf` becomes an import probe rather than a frozenset membership
test, exactly as extract.py's docstring had promised. The rejection does not
change: without the extra, `pdf` still raises `extractor_extra_missing` with
the same message. That behaviour is asserted UNCONDITIONALLY via a sys.modules
monkeypatch, so it holds on machines where the parser is installed too — a
skip would have preserved nothing there. Verified in a clean venv without the
extra: 589 passed, 7 skipped; with it, 596 passed.

`docx`/`xlsx` are unchanged and still fail fast — the extra names exactly what
it ships.

The parser choice was forced by measurement, not preference (b73dd9d,
docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real requirement table
pdfplumber keeps 4 of 4 rows with label and value on one line, where pypdf,
pdfminer.six and pymupdf each keep 0 of 4. pymupdf is additionally out on
licence (AGPL-3.0), which an MIT package must not push onto a consumer.

Three facts from that measurement are now carried in code rather than in a
report:

- Extracted text is pinned to an exact transitive parser version
  (pdfplumber pins pdfminer.six==20260107; date-stamped, no stability
  contract). tests/test_extract.py freezes the expected text of a committed
  hand-written fixture so a parser upgrade breaks something visible instead of
  drifting silently. Reasoning at the declaration site and in
  tests/fixtures/README.md.
- Determinism within a version is now held by a test, not only measured once.
- Drawn content does not survive extraction. Every pdf extraction emits the
  new `ExtractionWarning`: figures have no text to recover, so a bundle built
  from drawn documents is incomplete by construction. Stated categorically
  rather than detected — deciding "is there a figure here" is the layout
  heuristic G2b declined.

Two new error codes, both mirroring existing patterns: `extractor_empty_pdf`
(a scanned/image-only PDF, refused rather than persisted as an empty concept)
and `extractor_pdf_error` (parser failure wrapped, never leaked).

Structured table recovery (G2b) is NOT implemented and is documented as out of
scope: two independent parsers return the same wrong shape, so the breakage is
document geometry, not a library choice. PDFs enter as prose.

Also corrects an install promise this change would otherwise have published:
the README no longer presents a bare `pip install 'llm-ingestion-okf[extract]'`
as working, because the package is not on an index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtNhsdHnMGtMi7U2mvMU8z
This commit is contained in:
Kjell Tore Guttormsen 2026-08-21 20:22:39 +02:00
commit 658b7aafe0
14 changed files with 1046 additions and 37 deletions

View file

@ -8,6 +8,7 @@ conformance suite.
from __future__ import annotations
import importlib.util
import json
import sqlite3
import urllib.error
@ -48,6 +49,16 @@ from llm_ingestion_okf.render import sql_value_to_text
INGESTED_AT = "2026-07-17T12:00:00Z"
FIXTURES = Path(__file__).parent / "fixtures"
# The two pdf-parser codes are only reachable with the optional extra
# installed; the rejection code that replaces them without it is asserted
# unconditionally in tests/test_extract.py.
requires_extract = pytest.mark.skipif(
importlib.util.find_spec("pdfplumber") is None,
reason="the optional [extract] extra is not installed",
)
def inbox_concept(**overrides: Any) -> str:
"""A valid inbox concept render, one field at a time made invalid."""
@ -330,11 +341,28 @@ def test_extractor_unknown() -> None:
def test_extractor_extra_missing() -> None:
# `docx` rather than `pdf`: the extra now ships a pdf parser, so the type
# that still has none is what proves this code is reachable. The pdf
# import-probe path is covered in tests/test_extract.py.
with pytest.raises(ExtractionError) as excinfo:
extract_text("doc.pdf", b"binary")
extract_text("doc.docx", b"binary")
assert code_of(excinfo) == "extractor_extra_missing"
@requires_extract
def test_extractor_empty_pdf() -> None:
with pytest.raises(ExtractionError) as excinfo:
extract_text("scan.pdf", (FIXTURES / "no-text-layer.pdf").read_bytes())
assert code_of(excinfo) == "extractor_empty_pdf"
@requires_extract
def test_extractor_pdf_error() -> None:
with pytest.raises(ExtractionError) as excinfo:
extract_text("broken.pdf", b"not a pdf at all")
assert code_of(excinfo) == "extractor_pdf_error"
def test_extractor_decode_error() -> None:
with pytest.raises(ExtractionError) as excinfo:
extract_text("note.txt", b"\xffbad")