feat(extract): implement pdf behind the [extract] extra with pdfplumber
Order G2a. Populates the optional `[extract]` extra for the first time with
one parser, `pdfplumber>=0.11.10,<0.12` (MIT), and wires `pdf` through it.
The default install is untouched: exactly one runtime dependency, stdlib
otherwise, enforced by test_packaging.py.
The gate for `pdf` becomes an import probe rather than a frozenset membership
test, exactly as extract.py's docstring had promised. The rejection does not
change: without the extra, `pdf` still raises `extractor_extra_missing` with
the same message. That behaviour is asserted UNCONDITIONALLY via a sys.modules
monkeypatch, so it holds on machines where the parser is installed too — a
skip would have preserved nothing there. Verified in a clean venv without the
extra: 589 passed, 7 skipped; with it, 596 passed.
`docx`/`xlsx` are unchanged and still fail fast — the extra names exactly what
it ships.
The parser choice was forced by measurement, not preference (b73dd9d,
docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real requirement table
pdfplumber keeps 4 of 4 rows with label and value on one line, where pypdf,
pdfminer.six and pymupdf each keep 0 of 4. pymupdf is additionally out on
licence (AGPL-3.0), which an MIT package must not push onto a consumer.
Three facts from that measurement are now carried in code rather than in a
report:
- Extracted text is pinned to an exact transitive parser version
(pdfplumber pins pdfminer.six==20260107; date-stamped, no stability
contract). tests/test_extract.py freezes the expected text of a committed
hand-written fixture so a parser upgrade breaks something visible instead of
drifting silently. Reasoning at the declaration site and in
tests/fixtures/README.md.
- Determinism within a version is now held by a test, not only measured once.
- Drawn content does not survive extraction. Every pdf extraction emits the
new `ExtractionWarning`: figures have no text to recover, so a bundle built
from drawn documents is incomplete by construction. Stated categorically
rather than detected — deciding "is there a figure here" is the layout
heuristic G2b declined.
Two new error codes, both mirroring existing patterns: `extractor_empty_pdf`
(a scanned/image-only PDF, refused rather than persisted as an empty concept)
and `extractor_pdf_error` (parser failure wrapped, never leaked).
Structured table recovery (G2b) is NOT implemented and is documented as out of
scope: two independent parsers return the same wrong shape, so the breakage is
document geometry, not a library choice. PDFs enter as prose.
Also corrects an install promise this change would otherwise have published:
the README no longer presents a bare `pip install 'llm-ingestion-okf[extract]'`
as working, because the package is not on an index.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtNhsdHnMGtMi7U2mvMU8z
This commit is contained in:
parent
b73dd9d6a4
commit
658b7aafe0
14 changed files with 1046 additions and 37 deletions
14
CLAUDE.md
14
CLAUDE.md
|
|
@ -17,7 +17,13 @@ one boundary rule:
|
|||
file-type→text extraction lives HERE (the guard is text-only). v1 core:
|
||||
`md`, `txt`, `csv`, `json`, `html` (stdlib). `pdf`/`docx`/`xlsx` only via
|
||||
the optional `[extract]` extra; without it those types are rejected
|
||||
fail-fast.
|
||||
fail-fast. The extra ships `pdfplumber` for `pdf` (chosen on ONE measured
|
||||
property: it keeps a requirement table's label and value on the same line
|
||||
where three alternatives do not); `docx`/`xlsx` still ship no parser.
|
||||
Structured table recovery is **out of scope** — two independent parsers
|
||||
return the same wrong shape, so the breakage is document geometry, not a
|
||||
library choice. PDFs enter as prose, and drawn content (figures) does not
|
||||
survive extraction at all, which every `pdf` extraction warns about.
|
||||
- **Door C — external bundle import:** third-party OKF bundles are assessed
|
||||
per concept via the guard's `okf.import_bundle`; only concepts clearing the
|
||||
guard's non-blocking floor are merged/indexed here. Two invariants, both
|
||||
|
|
@ -176,7 +182,11 @@ stdlib, and a packaging test enforces it. Only `guard_adapter.py` imports the
|
|||
guard; importing the package does not. Install channel until the package
|
||||
index exists (a direct reference is a channel, not the pin):
|
||||
`pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v0.3.4"`.
|
||||
Binary extraction parsers live behind the `[extract]` extra only.
|
||||
Binary extraction parsers live behind the `[extract]` extra only — today
|
||||
`pdfplumber>=0.11.10,<0.12` for `pdf`. Extracted PDF text is pinned to an
|
||||
exact transitive parser version (`pdfminer.six==20260107`), so widening that
|
||||
range is a fixture migration, guarded by a frozen literal in
|
||||
`tests/test_extract.py`; see `tests/fixtures/README.md`.
|
||||
|
||||
Phase 4 adds a `node/` half: Node/ESM with zero npm dependencies
|
||||
(`node:` builtins only), both importable and CLI-invokable, consumed by
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue