llm-ingestion-okf/docs
Kjell Tore Guttormsen b73dd9d6a4 docs(extract): measure one Vegnormalene PDF page against the extraction registry
Order 20260821T170054Z-486638087-from-.claude (gap G2). Measurement only: no
parser implemented, no version bump, no pin move.

Measured on Handbok N200 Vegbygging (juli 2018), 308 pages, page index 150:
the registry rejects .pdf with extractor_extra_missing while .md/.csv controls
pass in the same call, and process_inbox reports the file as failed without
aborting the run. pdfplumber recovers Tabell 524.1 as 4/4 correctly paired text
lines where pypdf, pdfminer.six and pymupdf all score 0/4; both structural
extractors return the same wrong 2x6 grid, so table structure is the document's
geometry rather than a library defect. Whole book: 308/308 pages yield text,
45 of 196 detected tables are clean enough for render_table.

Verdict: text extraction is a small, bounded job; structured table recovery is a
separate project that nothing currently waits on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xTq1nbpz9x34udpDDExWM
2026-08-21 19:16:15 +02:00
..
plan docs(okf-v0.2): a third axis on the exposure question — where the producer lives 2026-08-09 21:58:04 +02:00
2026-08-21-g2-pdf-extraction-measurement.md docs(extract): measure one Vegnormalene PDF page against the extraction registry 2026-08-21 19:16:15 +02:00
phase-3-split-table.md docs(phase-3): keep the index-entry grammar on one line 2026-08-04 12:37:02 +02:00
upstream-okf-upgrade-runbook.md docs(okf-v0.2): V-A8 executed green, and the two premises it falsified 2026-07-31 21:12:52 +02:00