feat(extract): implement pdf behind the [extract] extra with pdfplumber

Order G2a. Populates the optional `[extract]` extra for the first time with
one parser, `pdfplumber>=0.11.10,<0.12` (MIT), and wires `pdf` through it.
The default install is untouched: exactly one runtime dependency, stdlib
otherwise, enforced by test_packaging.py.

The gate for `pdf` becomes an import probe rather than a frozenset membership
test, exactly as extract.py's docstring had promised. The rejection does not
change: without the extra, `pdf` still raises `extractor_extra_missing` with
the same message. That behaviour is asserted UNCONDITIONALLY via a sys.modules
monkeypatch, so it holds on machines where the parser is installed too — a
skip would have preserved nothing there. Verified in a clean venv without the
extra: 589 passed, 7 skipped; with it, 596 passed.

`docx`/`xlsx` are unchanged and still fail fast — the extra names exactly what
it ships.

The parser choice was forced by measurement, not preference (b73dd9d,
docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real requirement table
pdfplumber keeps 4 of 4 rows with label and value on one line, where pypdf,
pdfminer.six and pymupdf each keep 0 of 4. pymupdf is additionally out on
licence (AGPL-3.0), which an MIT package must not push onto a consumer.

Three facts from that measurement are now carried in code rather than in a
report:

- Extracted text is pinned to an exact transitive parser version
  (pdfplumber pins pdfminer.six==20260107; date-stamped, no stability
  contract). tests/test_extract.py freezes the expected text of a committed
  hand-written fixture so a parser upgrade breaks something visible instead of
  drifting silently. Reasoning at the declaration site and in
  tests/fixtures/README.md.
- Determinism within a version is now held by a test, not only measured once.
- Drawn content does not survive extraction. Every pdf extraction emits the
  new `ExtractionWarning`: figures have no text to recover, so a bundle built
  from drawn documents is incomplete by construction. Stated categorically
  rather than detected — deciding "is there a figure here" is the layout
  heuristic G2b declined.

Two new error codes, both mirroring existing patterns: `extractor_empty_pdf`
(a scanned/image-only PDF, refused rather than persisted as an empty concept)
and `extractor_pdf_error` (parser failure wrapped, never leaked).

Structured table recovery (G2b) is NOT implemented and is documented as out of
scope: two independent parsers return the same wrong shape, so the breakage is
document geometry, not a library choice. PDFs enter as prose.

Also corrects an install promise this change would otherwise have published:
the README no longer presents a bare `pip install 'llm-ingestion-okf[extract]'`
as working, because the package is not on an index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtNhsdHnMGtMi7U2mvMU8z
This commit is contained in:
Kjell Tore Guttormsen 2026-08-21 20:22:39 +02:00
commit 658b7aafe0
14 changed files with 1046 additions and 37 deletions

53
tests/fixtures/README.md vendored Normal file
View file

@ -0,0 +1,53 @@
# Test fixtures
## The PDF fixtures
`two-line-krav.pdf` and `no-text-layer.pdf` are hand-written minimal PDFs,
regenerated by `make_fixtures.py` in this directory:
```
python3 tests/fixtures/make_fixtures.py
```
They carry no library's output — the objects are laid out by hand and the xref
offsets computed from the emitted bytes — so they are auditable byte for byte
and reproducible from that one file.
| Fixture | What it is for |
|---|---|
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
## Why the expected PDF text is frozen as a literal
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
exact string. That is deliberate, and it is the mechanism behind a promise this
library makes everywhere else:
- Extraction is deterministic **within** a parser version. Measured 2026-08-21
across five configurations, two runs each, compared byte for byte
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`).
- Extraction is **not** guaranteed stable **across** parser versions.
`pdfplumber` pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships
date-stamped releases with no stability contract. So the real pin on extracted
text is a transitive one, and it is exact.
The consequence is worth stating plainly: **any golden fixture built on
extracted PDF text is pinned to an exact parser version, and a parser upgrade
is a fixture migration, not a routine bump.** The frozen literal is what makes
that upgrade break something visible instead of drifting silently. If it goes
red after a dependency change, the correct response is to read the diff and
decide, not to re-record the expectation.
The version range that carries this lives in `pyproject.toml`'s
`[project.optional-dependencies] extract`, with the same reasoning at the
declaration site.
## What these fixtures do not cover
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
detected table objects are clean enough to hand to `render_table` unchanged;
two independent parsers return the same wrong shape, because the breakage is in
the documents' ruling geometry rather than in either library. PDFs enter this
library as **prose**, and structured tables are out of scope until that is
decided separately.

64
tests/fixtures/make_fixtures.py vendored Normal file
View file

@ -0,0 +1,64 @@
"""Regenerate the committed PDF fixtures for tests/test_extract.py.
Hand-written minimal PDFs: objects laid out by hand, xref offsets computed
from the emitted bytes. No generator library, so the fixtures are auditable
byte for byte and reproducible from this file alone.
Run from the repository root: python3 tests/fixtures/make_fixtures.py
"""
from __future__ import annotations
from pathlib import Path
HERE = Path(__file__).parent
# Two text lines: a heading, and one requirement row with label and value on
# the SAME line. That pairing is the property the parser choice was made on
# (see docs/2026-08-21-g2-pdf-extraction-measurement.md), so the fixture
# fails visibly if a parser upgrade ever breaks it. Byte 0xE5 is the Norwegian
# 'a-ring' in WinAnsiEncoding, which the font object below declares.
KRAV_CONTENT = (
b"BT /F1 12 Tf 20 160 Td (Krav til helning p\xe5 utkilingen) Tj ET\n"
b"BT /F1 12 Tf 20 140 Td (60 og 70 1:15) Tj ET\n"
)
# A structurally valid page carrying no text operators at all -- the shape a
# scanned or image-only PDF presents to a text extractor.
NO_TEXT_CONTENT = b"20 20 160 160 re S\n"
def build_pdf(content: bytes) -> bytes:
"""Assemble a one-page PDF around `content` as the page content stream."""
objects = [
b"<< /Type /Catalog /Pages 2 0 R >>",
b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] "
b"/Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>",
b"<< /Length " + str(len(content)).encode() + b" >>\nstream\n" + content + b"endstream",
b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>",
]
out = bytearray(b"%PDF-1.4\n")
offsets = []
for number, body in enumerate(objects, start=1):
offsets.append(len(out))
out += str(number).encode() + b" 0 obj\n" + body + b"\nendobj\n"
xref_at = len(out)
size = str(len(objects) + 1).encode()
out += b"xref\n0 " + size + b"\n0000000000 65535 f \n"
for offset in offsets:
out += ("%010d 00000 n \n" % offset).encode()
out += b"trailer\n<< /Size " + size + b" /Root 1 0 R >>\n"
out += b"startxref\n" + str(xref_at).encode() + b"\n%%EOF\n"
return bytes(out)
if __name__ == "__main__":
for name, content in (
("two-line-krav.pdf", KRAV_CONTENT),
("no-text-layer.pdf", NO_TEXT_CONTENT),
):
(HERE / name).write_bytes(build_pdf(content))
print(f"wrote {name}")

32
tests/fixtures/no-text-layer.pdf vendored Normal file
View file

@ -0,0 +1,32 @@
%PDF-1.4
1 0 obj
<< /Type /Catalog /Pages 2 0 R >>
endobj
2 0 obj
<< /Type /Pages /Kids [3 0 R] /Count 1 >>
endobj
3 0 obj
<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>
endobj
4 0 obj
<< /Length 19 >>
stream
20 20 160 160 re S
endstream
endobj
5 0 obj
<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>
endobj
xref
0 6
0000000000 65535 f
0000000009 00000 n
0000000058 00000 n
0000000115 00000 n
0000000241 00000 n
0000000309 00000 n
trailer
<< /Size 6 /Root 1 0 R >>
startxref
406
%%EOF

33
tests/fixtures/two-line-krav.pdf vendored Normal file
View file

@ -0,0 +1,33 @@
%PDF-1.4
1 0 obj
<< /Type /Catalog /Pages 2 0 R >>
endobj
2 0 obj
<< /Type /Pages /Kids [3 0 R] /Count 1 >>
endobj
3 0 obj
<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>
endobj
4 0 obj
<< /Length 107 >>
stream
BT /F1 12 Tf 20 160 Td (Krav til helning på utkilingen) Tj ET
BT /F1 12 Tf 20 140 Td (60 og 70 1:15) Tj ET
endstream
endobj
5 0 obj
<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>
endobj
xref
0 6
0000000000 65535 f
0000000009 00000 n
0000000058 00000 n
0000000115 00000 n
0000000241 00000 n
0000000398 00000 n
trailer
<< /Size 6 /Root 1 0 R >>
startxref
495
%%EOF

View file

@ -8,6 +8,7 @@ conformance suite.
from __future__ import annotations
import importlib.util
import json
import sqlite3
import urllib.error
@ -48,6 +49,16 @@ from llm_ingestion_okf.render import sql_value_to_text
INGESTED_AT = "2026-07-17T12:00:00Z"
FIXTURES = Path(__file__).parent / "fixtures"
# The two pdf-parser codes are only reachable with the optional extra
# installed; the rejection code that replaces them without it is asserted
# unconditionally in tests/test_extract.py.
requires_extract = pytest.mark.skipif(
importlib.util.find_spec("pdfplumber") is None,
reason="the optional [extract] extra is not installed",
)
def inbox_concept(**overrides: Any) -> str:
"""A valid inbox concept render, one field at a time made invalid."""
@ -330,11 +341,28 @@ def test_extractor_unknown() -> None:
def test_extractor_extra_missing() -> None:
# `docx` rather than `pdf`: the extra now ships a pdf parser, so the type
# that still has none is what proves this code is reachable. The pdf
# import-probe path is covered in tests/test_extract.py.
with pytest.raises(ExtractionError) as excinfo:
extract_text("doc.pdf", b"binary")
extract_text("doc.docx", b"binary")
assert code_of(excinfo) == "extractor_extra_missing"
@requires_extract
def test_extractor_empty_pdf() -> None:
with pytest.raises(ExtractionError) as excinfo:
extract_text("scan.pdf", (FIXTURES / "no-text-layer.pdf").read_bytes())
assert code_of(excinfo) == "extractor_empty_pdf"
@requires_extract
def test_extractor_pdf_error() -> None:
with pytest.raises(ExtractionError) as excinfo:
extract_text("broken.pdf", b"not a pdf at all")
assert code_of(excinfo) == "extractor_pdf_error"
def test_extractor_decode_error() -> None:
with pytest.raises(ExtractionError) as excinfo:
extract_text("note.txt", b"\xffbad")

View file

@ -9,9 +9,24 @@ call — it is pure, deterministic plumbing.
from __future__ import annotations
import importlib.util
import sys
from pathlib import Path
import pytest
from llm_ingestion_okf import ExtractionError, extract_text
from llm_ingestion_okf import ExtractionError, ExtractionWarning, extract_text
FIXTURES = Path(__file__).parent / "fixtures"
# The parser-dependent tests below need the optional extra, so they are skipped
# without it — an optional extra that forced its parser on every test run would
# not be optional. The behaviour that must survive EITHER WAY is the rejection,
# and that one is asserted unconditionally above via the import probe.
requires_extract = pytest.mark.skipif(
importlib.util.find_spec("pdfplumber") is None,
reason="the optional [extract] extra is not installed",
)
# --- core extractors: happy path (a fixture per type) ---
@ -74,11 +89,11 @@ def test_missing_extension_fails_fast() -> None:
assert excinfo.value.code == "extractor_unknown"
# --- fail-fast: [extract]-gated binary types without the extra ---
# --- fail-fast: [extract]-gated binary types the extra ships no parser for ---
@pytest.mark.parametrize("name", ["doc.pdf", "doc.docx", "sheet.xlsx"])
def test_optional_type_without_extra_fails_fast(name: str) -> None:
@pytest.mark.parametrize("name", ["doc.docx", "sheet.xlsx"])
def test_optional_type_without_a_parser_fails_fast(name: str) -> None:
with pytest.raises(ExtractionError) as excinfo:
extract_text(name, b"binary")
assert excinfo.value.code == "extractor_extra_missing"
@ -87,6 +102,80 @@ def test_optional_type_without_extra_fails_fast(name: str) -> None:
assert "extract" in str(excinfo.value)
def test_pdf_without_the_extra_installed_keeps_the_same_rejection(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The gate became an import probe; the rejection did not change.
Setting the module to None in sys.modules is what CPython treats as a
failed import, so this exercises the uninstalled path REGARDLESS of
whether the extra is installed in the running environment a skip would
have preserved nothing on the machine where the parser is present.
"""
monkeypatch.setitem(sys.modules, "pdfplumber", None)
with pytest.raises(ExtractionError) as excinfo:
extract_text("doc.pdf", b"%PDF-1.4")
assert excinfo.value.code == "extractor_extra_missing"
assert "extract" in str(excinfo.value)
# --- pdf: the [extract] parser (pdfplumber) ---
# The expected text is frozen against a committed fixture ON PURPOSE. Extraction
# is deterministic within a parser version and NOT guaranteed across one
# (pdfplumber pins pdfminer.six==20260107 exactly; pdfminer.six ships
# date-stamped releases with no stability contract). This literal is what makes
# a parser upgrade break something visible instead of drifting silently.
KRAV_TEXT = "Krav til helning på utkilingen\n60 og 70 1:15"
@requires_extract
def test_pdf_extracts_text_with_label_and_value_on_one_line() -> None:
data = (FIXTURES / "two-line-krav.pdf").read_bytes()
with pytest.warns(ExtractionWarning):
assert extract_text("krav.pdf", data) == KRAV_TEXT
@requires_extract
def test_pdf_extraction_is_byte_stable_across_calls() -> None:
"""Determinism is a promise this library already makes; hold it here too."""
data = (FIXTURES / "two-line-krav.pdf").read_bytes()
with pytest.warns(ExtractionWarning):
first = extract_text("krav.pdf", data)
second = extract_text("krav.pdf", data)
assert first == second
@requires_extract
def test_pdf_extraction_warns_that_drawn_content_is_not_recovered() -> None:
"""Figures are vector drawings: only the caption survives leg 2.
Categorically true of text extraction, so it is stated as a warning on
every PDF rather than guessed at per document detecting "is there a
figure here" would be exactly the layout heuristic this order declined.
"""
data = (FIXTURES / "two-line-krav.pdf").read_bytes()
with pytest.warns(ExtractionWarning, match="figures"):
extract_text("krav.pdf", data)
@requires_extract
def test_pdf_with_no_text_layer_fails_fast() -> None:
"""A scanned PDF yields nothing; an empty concept would be a silent skip."""
data = (FIXTURES / "no-text-layer.pdf").read_bytes()
with pytest.raises(ExtractionError) as excinfo:
extract_text("scan.pdf", data)
assert excinfo.value.code == "extractor_empty_pdf"
@requires_extract
def test_corrupt_pdf_fails_fast_typed() -> None:
"""Never a leaked pdfminer exception — the "always typed" doctrine holds."""
with pytest.raises(ExtractionError) as excinfo:
extract_text("broken.pdf", b"not a pdf at all")
assert excinfo.value.code == "extractor_pdf_error"
# --- fail-fast: corrupt (non-UTF-8) bytes on a text type ---