feat(extract): implement pdf behind the [extract] extra with pdfplumber

Order G2a. Populates the optional `[extract]` extra for the first time with
one parser, `pdfplumber>=0.11.10,<0.12` (MIT), and wires `pdf` through it.
The default install is untouched: exactly one runtime dependency, stdlib
otherwise, enforced by test_packaging.py.

The gate for `pdf` becomes an import probe rather than a frozenset membership
test, exactly as extract.py's docstring had promised. The rejection does not
change: without the extra, `pdf` still raises `extractor_extra_missing` with
the same message. That behaviour is asserted UNCONDITIONALLY via a sys.modules
monkeypatch, so it holds on machines where the parser is installed too — a
skip would have preserved nothing there. Verified in a clean venv without the
extra: 589 passed, 7 skipped; with it, 596 passed.

`docx`/`xlsx` are unchanged and still fail fast — the extra names exactly what
it ships.

The parser choice was forced by measurement, not preference (b73dd9d,
docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real requirement table
pdfplumber keeps 4 of 4 rows with label and value on one line, where pypdf,
pdfminer.six and pymupdf each keep 0 of 4. pymupdf is additionally out on
licence (AGPL-3.0), which an MIT package must not push onto a consumer.

Three facts from that measurement are now carried in code rather than in a
report:

- Extracted text is pinned to an exact transitive parser version
  (pdfplumber pins pdfminer.six==20260107; date-stamped, no stability
  contract). tests/test_extract.py freezes the expected text of a committed
  hand-written fixture so a parser upgrade breaks something visible instead of
  drifting silently. Reasoning at the declaration site and in
  tests/fixtures/README.md.
- Determinism within a version is now held by a test, not only measured once.
- Drawn content does not survive extraction. Every pdf extraction emits the
  new `ExtractionWarning`: figures have no text to recover, so a bundle built
  from drawn documents is incomplete by construction. Stated categorically
  rather than detected — deciding "is there a figure here" is the layout
  heuristic G2b declined.

Two new error codes, both mirroring existing patterns: `extractor_empty_pdf`
(a scanned/image-only PDF, refused rather than persisted as an empty concept)
and `extractor_pdf_error` (parser failure wrapped, never leaked).

Structured table recovery (G2b) is NOT implemented and is documented as out of
scope: two independent parsers return the same wrong shape, so the breakage is
document geometry, not a library choice. PDFs enter as prose.

Also corrects an install promise this change would otherwise have published:
the README no longer presents a bare `pip install 'llm-ingestion-okf[extract]'`
as working, because the package is not on an index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtNhsdHnMGtMi7U2mvMU8z
This commit is contained in:
Kjell Tore Guttormsen 2026-08-21 20:22:39 +02:00
commit 658b7aafe0
14 changed files with 1046 additions and 37 deletions

View file

@ -7,8 +7,69 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- **Door B extracts `pdf` behind the optional `[extract]` extra.** The extra is
populated for the first time, with one parser: `pdfplumber>=0.11.10,<0.12`
(MIT). The default install is unchanged — still exactly one runtime
dependency, still stdlib otherwise — and a packaging test enforces that.
**The parser choice was forced by a measurement, not by preference**
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`). On a real Vegnormalene
requirement table, `pdfplumber` keeps 4 of 4 rows with label and value on the
same line; `pypdf`, `pdfminer.six` and `pymupdf` each keep 0 of 4, emitting
all labels and then all values. A downstream reader can only re-pair those by
guessing, and in a requirements document a wrong pairing looks right.
`pymupdf` was additionally excluded on licence (AGPL-3.0 or commercial):
this package is MIT and an optional extra must not hand a consumer copyleft
they did not choose.
**`docx` and `xlsx` are unchanged.** They ship no parser and still fail fast
with `extractor_extra_missing`, so the extra names exactly what it delivers.
- **`ExtractionWarning`**, exported from the package. Every `pdf` extraction
emits one. Text extraction recovers text; anything a PDF *draws* — figures,
diagrams, images — has no text to recover, so only captions survive and a
bundle built from drawn documents is **incomplete by construction**. That is
categorically true rather than document-specific, so it is stated rather than
detected: deciding "is there a figure on this page" is a layout heuristic this
library does not own. A named class so it can be filtered deliberately.
- **Two error codes**, both mirroring existing patterns rather than inventing
behaviour: `extractor_empty_pdf` (a PDF yielded no text on any page — a
scanned or image-only document; refused rather than persisted as an empty
concept, which would be the silent skip this registry exists to prevent) and
`extractor_pdf_error` (the parser failed on the bytes; the third-party
exception is wrapped, never leaked).
### Changed
- **The `[extract]` gate for `pdf` is now an import probe rather than a
membership test.** The rejection did not change: without the extra installed,
`pdf` still raises `ExtractionError` with code `extractor_extra_missing` and
the same message naming the remedy. A consumer that has not installed the
extra sees no difference at all. That behaviour is asserted unconditionally,
including on machines where the parser *is* installed, so it cannot rot into
a skipped test.
**Known limitation, stated rather than worked around:** structured table
recovery is out of scope. `pdfplumber.extract_tables()` and
`PyMuPDF.find_tables()` — two independent implementations — return the *same*
wrong shape for the measured requirement table, and across the whole handbook
only 45 of 196 detected table objects are clean enough to hand to
`render_table` unchanged. The breakage is in the documents' ruling geometry,
not in either library. PDFs therefore enter this library as **prose**, with
table lines correctly paired.
**Extracted PDF text is pinned to an exact parser version.** `pdfplumber`
pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
releases with no stability contract. Extraction is deterministic within a
parser version (measured across five configurations) and not guaranteed
across one, so any golden fixture built on extracted PDF text is a fixture
migration away from a parser upgrade. `tests/test_extract.py` freezes the
expected text of a committed fixture so that upgrade breaks something visible
instead of drifting silently; see `tests/fixtures/README.md`.
- **`DEFAULT` now stamps `generated: { by: process:okf-ingest, at: <ingested_at> }`
instead of `generated: true`.** This is a byte change in every concept file Door A
writes under `DEFAULT`, so a consumer's own golden fixtures will show one changed