docs(extract): the extra now ships an office converter

The README told consumers that `docx` and `xlsx` ship no parser and always fail
fast. True when written; false the moment the converter seam landed -- and
false SILENTLY, because prose has no test. This repository has been bitten by
that exact shape before: a published guarantee is a test obligation.

So the correction comes with `tests/test_docs_promises.py`, which compares the
README's declared format list against the registries it describes and fails on
a format added without touching the README, on the old claim reappearing in any
wording, on an unmeasured row going unnamed, and on the exclusions being
dropped. Negative control: removing one format from the README's marker turns
it red.

The README now states which rows are measured and which are not. Three of the
five office rows have denominator ZERO in the corpus -- they work by
construction and have never met a document anyone wrote. They are not known to
be broken and not known to be right, and a reader should not have to open the
source to learn which.

The CHANGELOG's shipped entry is left as written, because a changelog records
what a release did; the correction is stated at that line instead so a reader
arriving there is not misled.

Suite 913 -> 917.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-02 14:15:55 +02:00
commit 0170c526ad
3 changed files with 159 additions and 12 deletions

View file

@ -11,10 +11,11 @@ gate, with `llm_ingestion_okf.guard_adapter` wiring that gate to the real
guard (see below). Phase 3 makes the bundle contract configurable, so types,
layers, frontmatter sets, index shape, and reserved-file policy are carried by
a profile rather than by constants (see [Upstream OKF
versions](#upstream-okf-versions)). Binary extraction is partial: `pdf` is
implemented behind the optional `[extract]` extra, while `docx`/`xlsx` remain
unimplemented and are rejected fail-fast. Phase 4 (the Node half) is planned
(see `docs/plan/`).
versions](#upstream-okf-versions)). Binary extraction runs behind the
optional `[extract]` extra: `pdf` through a PDF parser, and five office
formats through a vendored document converter. Three of those five office
rows are **unmeasured** — see [Binary extraction](#binary-extraction). Phase 4
(the Node half) is planned (see `docs/plan/`).
## Install
@ -64,9 +65,11 @@ bundle:
2. **Bundle inbox.** A drop directory where common file types are converted
to OKF concept files. All file-type→text extraction lives in this library:
`md`, `txt`, `csv`, `json`, and `html` are handled by the stdlib core;
`pdf` requires the optional `[extract]` extra and is rejected fail-fast
without it; `docx` and `xlsx` ship no parser yet and are always rejected.
Extracted text passes the security gate before anything is persisted.
`pdf` and the five office formats (`docx`, `xlsx`, `pptx`, `odt`, `rtf`)
require the optional `[extract]` extra and are rejected fail-fast without
it. Extracted text passes the security gate before anything is persisted.
<!-- extract-formats: .md, .txt, .csv, .json, .html, .htm, .pdf, .docx, .xlsx, .pptx, .odt, .rtf -->
3. **External bundle import.** Import and merge of third-party OKF bundles:
each concept is assessed via the security gate, and only concepts that
pass are merged, materialized, and linked into the index.
@ -277,9 +280,38 @@ llm-ingestion-guard<2.0,>=1.2`, measured 2026-08-23 — and resolves normally
once the package index exists. A wheel built from a *tag* carries that tag's
range instead, which is why the install commands pair tag with tag.
The optional `[extract]` extra ships one parser, `pdfplumber` (MIT), for `pdf`;
`docx`/`xlsx` are still unimplemented. It is opt-in because it pulls binary
wheels (`pillow`, `pypdfium2`), which the default install must never do.
### Binary extraction
The optional `[extract]` extra ships two things: `pdfplumber` (MIT) for `pdf`,
and `pypandoc-binary` for five office formats. It is opt-in because it pulls
binary wheels, which the default install must never do — the single runtime
dependency rule covers the default install and this extra sits outside it.
The converter **binary travels inside the wheel** and is resolved by path
rather than found on `PATH`, with its version asserted against a pin. A host
carrying a different converter is refused, not silently used: extraction is
deterministic within a converter version and not across one.
| Format | Reader | Evidence |
|---|---|---|
| `pdf` | `pdfplumber` | measured |
| `docx` | converter | measured |
| `xlsx` | converter | measured |
| `pptx` | converter | **unmeasured** |
| `odt` | converter | **unmeasured** |
| `rtf` | converter | **unmeasured** |
**`unmeasured` means what it says.** The corpus this work was measured on
contains **zero** `pptx`, `odt` and `rtf` files, so those three rows work by
construction and have never been checked against a document anyone wrote.
They are not known to be broken; they are not known to be right either, and
the distinction is the point.
**What stays out.** `.doc` (Word 97) is not supported — the converter does not
read it. Rastered or scanned PDFs are refused rather than persisted as empty
concepts, because this library does not do OCR. Drawn content — figures,
diagrams, shapes — does not survive extraction in any format here, and every
extraction says so with a warning. Structured table recovery is out of scope.
Request it by appending `[extract]` to the package name in whichever install
command from [Install](#install) you are using — this package is not on an