feat(extract): implement pdf behind the [extract] extra with pdfplumber
Order G2a. Populates the optional `[extract]` extra for the first time with
one parser, `pdfplumber>=0.11.10,<0.12` (MIT), and wires `pdf` through it.
The default install is untouched: exactly one runtime dependency, stdlib
otherwise, enforced by test_packaging.py.
The gate for `pdf` becomes an import probe rather than a frozenset membership
test, exactly as extract.py's docstring had promised. The rejection does not
change: without the extra, `pdf` still raises `extractor_extra_missing` with
the same message. That behaviour is asserted UNCONDITIONALLY via a sys.modules
monkeypatch, so it holds on machines where the parser is installed too — a
skip would have preserved nothing there. Verified in a clean venv without the
extra: 589 passed, 7 skipped; with it, 596 passed.
`docx`/`xlsx` are unchanged and still fail fast — the extra names exactly what
it ships.
The parser choice was forced by measurement, not preference (b73dd9d,
docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real requirement table
pdfplumber keeps 4 of 4 rows with label and value on one line, where pypdf,
pdfminer.six and pymupdf each keep 0 of 4. pymupdf is additionally out on
licence (AGPL-3.0), which an MIT package must not push onto a consumer.
Three facts from that measurement are now carried in code rather than in a
report:
- Extracted text is pinned to an exact transitive parser version
(pdfplumber pins pdfminer.six==20260107; date-stamped, no stability
contract). tests/test_extract.py freezes the expected text of a committed
hand-written fixture so a parser upgrade breaks something visible instead of
drifting silently. Reasoning at the declaration site and in
tests/fixtures/README.md.
- Determinism within a version is now held by a test, not only measured once.
- Drawn content does not survive extraction. Every pdf extraction emits the
new `ExtractionWarning`: figures have no text to recover, so a bundle built
from drawn documents is incomplete by construction. Stated categorically
rather than detected — deciding "is there a figure here" is the layout
heuristic G2b declined.
Two new error codes, both mirroring existing patterns: `extractor_empty_pdf`
(a scanned/image-only PDF, refused rather than persisted as an empty concept)
and `extractor_pdf_error` (parser failure wrapped, never leaked).
Structured table recovery (G2b) is NOT implemented and is documented as out of
scope: two independent parsers return the same wrong shape, so the breakage is
document geometry, not a library choice. PDFs enter as prose.
Also corrects an install promise this change would otherwise have published:
the README no longer presents a bare `pip install 'llm-ingestion-okf[extract]'`
as working, because the package is not on an index.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtNhsdHnMGtMi7U2mvMU8z
This commit is contained in:
parent
b73dd9d6a4
commit
658b7aafe0
14 changed files with 1046 additions and 37 deletions
53
tests/fixtures/README.md
vendored
Normal file
53
tests/fixtures/README.md
vendored
Normal file
|
|
@ -0,0 +1,53 @@
|
|||
# Test fixtures
|
||||
|
||||
## The PDF fixtures
|
||||
|
||||
`two-line-krav.pdf` and `no-text-layer.pdf` are hand-written minimal PDFs,
|
||||
regenerated by `make_fixtures.py` in this directory:
|
||||
|
||||
```
|
||||
python3 tests/fixtures/make_fixtures.py
|
||||
```
|
||||
|
||||
They carry no library's output — the objects are laid out by hand and the xref
|
||||
offsets computed from the emitted bytes — so they are auditable byte for byte
|
||||
and reproducible from that one file.
|
||||
|
||||
| Fixture | What it is for |
|
||||
|---|---|
|
||||
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
|
||||
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
|
||||
|
||||
## Why the expected PDF text is frozen as a literal
|
||||
|
||||
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
|
||||
exact string. That is deliberate, and it is the mechanism behind a promise this
|
||||
library makes everywhere else:
|
||||
|
||||
- Extraction is deterministic **within** a parser version. Measured 2026-08-21
|
||||
across five configurations, two runs each, compared byte for byte
|
||||
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`).
|
||||
- Extraction is **not** guaranteed stable **across** parser versions.
|
||||
`pdfplumber` pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships
|
||||
date-stamped releases with no stability contract. So the real pin on extracted
|
||||
text is a transitive one, and it is exact.
|
||||
|
||||
The consequence is worth stating plainly: **any golden fixture built on
|
||||
extracted PDF text is pinned to an exact parser version, and a parser upgrade
|
||||
is a fixture migration, not a routine bump.** The frozen literal is what makes
|
||||
that upgrade break something visible instead of drifting silently. If it goes
|
||||
red after a dependency change, the correct response is to read the diff and
|
||||
decide, not to re-record the expectation.
|
||||
|
||||
The version range that carries this lives in `pyproject.toml`'s
|
||||
`[project.optional-dependencies] extract`, with the same reasoning at the
|
||||
declaration site.
|
||||
|
||||
## What these fixtures do not cover
|
||||
|
||||
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
|
||||
detected table objects are clean enough to hand to `render_table` unchanged;
|
||||
two independent parsers return the same wrong shape, because the breakage is in
|
||||
the documents' ruling geometry rather than in either library. PDFs enter this
|
||||
library as **prose**, and structured tables are out of scope until that is
|
||||
decided separately.
|
||||
64
tests/fixtures/make_fixtures.py
vendored
Normal file
64
tests/fixtures/make_fixtures.py
vendored
Normal file
|
|
@ -0,0 +1,64 @@
|
|||
"""Regenerate the committed PDF fixtures for tests/test_extract.py.
|
||||
|
||||
Hand-written minimal PDFs: objects laid out by hand, xref offsets computed
|
||||
from the emitted bytes. No generator library, so the fixtures are auditable
|
||||
byte for byte and reproducible from this file alone.
|
||||
|
||||
Run from the repository root: python3 tests/fixtures/make_fixtures.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
HERE = Path(__file__).parent
|
||||
|
||||
# Two text lines: a heading, and one requirement row with label and value on
|
||||
# the SAME line. That pairing is the property the parser choice was made on
|
||||
# (see docs/2026-08-21-g2-pdf-extraction-measurement.md), so the fixture
|
||||
# fails visibly if a parser upgrade ever breaks it. Byte 0xE5 is the Norwegian
|
||||
# 'a-ring' in WinAnsiEncoding, which the font object below declares.
|
||||
KRAV_CONTENT = (
|
||||
b"BT /F1 12 Tf 20 160 Td (Krav til helning p\xe5 utkilingen) Tj ET\n"
|
||||
b"BT /F1 12 Tf 20 140 Td (60 og 70 1:15) Tj ET\n"
|
||||
)
|
||||
|
||||
# A structurally valid page carrying no text operators at all -- the shape a
|
||||
# scanned or image-only PDF presents to a text extractor.
|
||||
NO_TEXT_CONTENT = b"20 20 160 160 re S\n"
|
||||
|
||||
|
||||
def build_pdf(content: bytes) -> bytes:
|
||||
"""Assemble a one-page PDF around `content` as the page content stream."""
|
||||
objects = [
|
||||
b"<< /Type /Catalog /Pages 2 0 R >>",
|
||||
b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
|
||||
b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] "
|
||||
b"/Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>",
|
||||
b"<< /Length " + str(len(content)).encode() + b" >>\nstream\n" + content + b"endstream",
|
||||
b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>",
|
||||
]
|
||||
|
||||
out = bytearray(b"%PDF-1.4\n")
|
||||
offsets = []
|
||||
for number, body in enumerate(objects, start=1):
|
||||
offsets.append(len(out))
|
||||
out += str(number).encode() + b" 0 obj\n" + body + b"\nendobj\n"
|
||||
|
||||
xref_at = len(out)
|
||||
size = str(len(objects) + 1).encode()
|
||||
out += b"xref\n0 " + size + b"\n0000000000 65535 f \n"
|
||||
for offset in offsets:
|
||||
out += ("%010d 00000 n \n" % offset).encode()
|
||||
out += b"trailer\n<< /Size " + size + b" /Root 1 0 R >>\n"
|
||||
out += b"startxref\n" + str(xref_at).encode() + b"\n%%EOF\n"
|
||||
return bytes(out)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
for name, content in (
|
||||
("two-line-krav.pdf", KRAV_CONTENT),
|
||||
("no-text-layer.pdf", NO_TEXT_CONTENT),
|
||||
):
|
||||
(HERE / name).write_bytes(build_pdf(content))
|
||||
print(f"wrote {name}")
|
||||
32
tests/fixtures/no-text-layer.pdf
vendored
Normal file
32
tests/fixtures/no-text-layer.pdf
vendored
Normal file
|
|
@ -0,0 +1,32 @@
|
|||
%PDF-1.4
|
||||
1 0 obj
|
||||
<< /Type /Catalog /Pages 2 0 R >>
|
||||
endobj
|
||||
2 0 obj
|
||||
<< /Type /Pages /Kids [3 0 R] /Count 1 >>
|
||||
endobj
|
||||
3 0 obj
|
||||
<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>
|
||||
endobj
|
||||
4 0 obj
|
||||
<< /Length 19 >>
|
||||
stream
|
||||
20 20 160 160 re S
|
||||
endstream
|
||||
endobj
|
||||
5 0 obj
|
||||
<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>
|
||||
endobj
|
||||
xref
|
||||
0 6
|
||||
0000000000 65535 f
|
||||
0000000009 00000 n
|
||||
0000000058 00000 n
|
||||
0000000115 00000 n
|
||||
0000000241 00000 n
|
||||
0000000309 00000 n
|
||||
trailer
|
||||
<< /Size 6 /Root 1 0 R >>
|
||||
startxref
|
||||
406
|
||||
%%EOF
|
||||
33
tests/fixtures/two-line-krav.pdf
vendored
Normal file
33
tests/fixtures/two-line-krav.pdf
vendored
Normal file
|
|
@ -0,0 +1,33 @@
|
|||
%PDF-1.4
|
||||
1 0 obj
|
||||
<< /Type /Catalog /Pages 2 0 R >>
|
||||
endobj
|
||||
2 0 obj
|
||||
<< /Type /Pages /Kids [3 0 R] /Count 1 >>
|
||||
endobj
|
||||
3 0 obj
|
||||
<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>
|
||||
endobj
|
||||
4 0 obj
|
||||
<< /Length 107 >>
|
||||
stream
|
||||
BT /F1 12 Tf 20 160 Td (Krav til helning på utkilingen) Tj ET
|
||||
BT /F1 12 Tf 20 140 Td (60 og 70 1:15) Tj ET
|
||||
endstream
|
||||
endobj
|
||||
5 0 obj
|
||||
<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>
|
||||
endobj
|
||||
xref
|
||||
0 6
|
||||
0000000000 65535 f
|
||||
0000000009 00000 n
|
||||
0000000058 00000 n
|
||||
0000000115 00000 n
|
||||
0000000241 00000 n
|
||||
0000000398 00000 n
|
||||
trailer
|
||||
<< /Size 6 /Root 1 0 R >>
|
||||
startxref
|
||||
495
|
||||
%%EOF
|
||||
|
|
@ -8,6 +8,7 @@ conformance suite.
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import json
|
||||
import sqlite3
|
||||
import urllib.error
|
||||
|
|
@ -48,6 +49,16 @@ from llm_ingestion_okf.render import sql_value_to_text
|
|||
|
||||
INGESTED_AT = "2026-07-17T12:00:00Z"
|
||||
|
||||
FIXTURES = Path(__file__).parent / "fixtures"
|
||||
|
||||
# The two pdf-parser codes are only reachable with the optional extra
|
||||
# installed; the rejection code that replaces them without it is asserted
|
||||
# unconditionally in tests/test_extract.py.
|
||||
requires_extract = pytest.mark.skipif(
|
||||
importlib.util.find_spec("pdfplumber") is None,
|
||||
reason="the optional [extract] extra is not installed",
|
||||
)
|
||||
|
||||
|
||||
def inbox_concept(**overrides: Any) -> str:
|
||||
"""A valid inbox concept render, one field at a time made invalid."""
|
||||
|
|
@ -330,11 +341,28 @@ def test_extractor_unknown() -> None:
|
|||
|
||||
|
||||
def test_extractor_extra_missing() -> None:
|
||||
# `docx` rather than `pdf`: the extra now ships a pdf parser, so the type
|
||||
# that still has none is what proves this code is reachable. The pdf
|
||||
# import-probe path is covered in tests/test_extract.py.
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text("doc.pdf", b"binary")
|
||||
extract_text("doc.docx", b"binary")
|
||||
assert code_of(excinfo) == "extractor_extra_missing"
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_extractor_empty_pdf() -> None:
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text("scan.pdf", (FIXTURES / "no-text-layer.pdf").read_bytes())
|
||||
assert code_of(excinfo) == "extractor_empty_pdf"
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_extractor_pdf_error() -> None:
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text("broken.pdf", b"not a pdf at all")
|
||||
assert code_of(excinfo) == "extractor_pdf_error"
|
||||
|
||||
|
||||
def test_extractor_decode_error() -> None:
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text("note.txt", b"\xffbad")
|
||||
|
|
|
|||
|
|
@ -9,9 +9,24 @@ call — it is pure, deterministic plumbing.
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from llm_ingestion_okf import ExtractionError, extract_text
|
||||
from llm_ingestion_okf import ExtractionError, ExtractionWarning, extract_text
|
||||
|
||||
FIXTURES = Path(__file__).parent / "fixtures"
|
||||
|
||||
# The parser-dependent tests below need the optional extra, so they are skipped
|
||||
# without it — an optional extra that forced its parser on every test run would
|
||||
# not be optional. The behaviour that must survive EITHER WAY is the rejection,
|
||||
# and that one is asserted unconditionally above via the import probe.
|
||||
requires_extract = pytest.mark.skipif(
|
||||
importlib.util.find_spec("pdfplumber") is None,
|
||||
reason="the optional [extract] extra is not installed",
|
||||
)
|
||||
|
||||
# --- core extractors: happy path (a fixture per type) ---
|
||||
|
||||
|
|
@ -74,11 +89,11 @@ def test_missing_extension_fails_fast() -> None:
|
|||
assert excinfo.value.code == "extractor_unknown"
|
||||
|
||||
|
||||
# --- fail-fast: [extract]-gated binary types without the extra ---
|
||||
# --- fail-fast: [extract]-gated binary types the extra ships no parser for ---
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", ["doc.pdf", "doc.docx", "sheet.xlsx"])
|
||||
def test_optional_type_without_extra_fails_fast(name: str) -> None:
|
||||
@pytest.mark.parametrize("name", ["doc.docx", "sheet.xlsx"])
|
||||
def test_optional_type_without_a_parser_fails_fast(name: str) -> None:
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text(name, b"binary")
|
||||
assert excinfo.value.code == "extractor_extra_missing"
|
||||
|
|
@ -87,6 +102,80 @@ def test_optional_type_without_extra_fails_fast(name: str) -> None:
|
|||
assert "extract" in str(excinfo.value)
|
||||
|
||||
|
||||
def test_pdf_without_the_extra_installed_keeps_the_same_rejection(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""The gate became an import probe; the rejection did not change.
|
||||
|
||||
Setting the module to None in sys.modules is what CPython treats as a
|
||||
failed import, so this exercises the uninstalled path REGARDLESS of
|
||||
whether the extra is installed in the running environment — a skip would
|
||||
have preserved nothing on the machine where the parser is present.
|
||||
"""
|
||||
monkeypatch.setitem(sys.modules, "pdfplumber", None)
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text("doc.pdf", b"%PDF-1.4")
|
||||
assert excinfo.value.code == "extractor_extra_missing"
|
||||
assert "extract" in str(excinfo.value)
|
||||
|
||||
|
||||
# --- pdf: the [extract] parser (pdfplumber) ---
|
||||
|
||||
# The expected text is frozen against a committed fixture ON PURPOSE. Extraction
|
||||
# is deterministic within a parser version and NOT guaranteed across one
|
||||
# (pdfplumber pins pdfminer.six==20260107 exactly; pdfminer.six ships
|
||||
# date-stamped releases with no stability contract). This literal is what makes
|
||||
# a parser upgrade break something visible instead of drifting silently.
|
||||
KRAV_TEXT = "Krav til helning på utkilingen\n60 og 70 1:15"
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_pdf_extracts_text_with_label_and_value_on_one_line() -> None:
|
||||
data = (FIXTURES / "two-line-krav.pdf").read_bytes()
|
||||
with pytest.warns(ExtractionWarning):
|
||||
assert extract_text("krav.pdf", data) == KRAV_TEXT
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_pdf_extraction_is_byte_stable_across_calls() -> None:
|
||||
"""Determinism is a promise this library already makes; hold it here too."""
|
||||
data = (FIXTURES / "two-line-krav.pdf").read_bytes()
|
||||
with pytest.warns(ExtractionWarning):
|
||||
first = extract_text("krav.pdf", data)
|
||||
second = extract_text("krav.pdf", data)
|
||||
assert first == second
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_pdf_extraction_warns_that_drawn_content_is_not_recovered() -> None:
|
||||
"""Figures are vector drawings: only the caption survives leg 2.
|
||||
|
||||
Categorically true of text extraction, so it is stated as a warning on
|
||||
every PDF rather than guessed at per document — detecting "is there a
|
||||
figure here" would be exactly the layout heuristic this order declined.
|
||||
"""
|
||||
data = (FIXTURES / "two-line-krav.pdf").read_bytes()
|
||||
with pytest.warns(ExtractionWarning, match="figures"):
|
||||
extract_text("krav.pdf", data)
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_pdf_with_no_text_layer_fails_fast() -> None:
|
||||
"""A scanned PDF yields nothing; an empty concept would be a silent skip."""
|
||||
data = (FIXTURES / "no-text-layer.pdf").read_bytes()
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text("scan.pdf", data)
|
||||
assert excinfo.value.code == "extractor_empty_pdf"
|
||||
|
||||
|
||||
@requires_extract
|
||||
def test_corrupt_pdf_fails_fast_typed() -> None:
|
||||
"""Never a leaked pdfminer exception — the "always typed" doctrine holds."""
|
||||
with pytest.raises(ExtractionError) as excinfo:
|
||||
extract_text("broken.pdf", b"not a pdf at all")
|
||||
assert excinfo.value.code == "extractor_pdf_error"
|
||||
|
||||
|
||||
# --- fail-fast: corrupt (non-UTF-8) bytes on a text type ---
|
||||
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue