llm-ingestion-okf/README.md
Kjell Tore Guttormsen 171798ed32 docs(consume): the Claude Code recipe, measured end to end on two bundles
Four questions, two bundles, one run each, in a scratch project outside this
repository with a generated skill per bundle. All four passed, and zero numbers
or identifiers appeared in any answer that were not in the delivered set or in
the payload's own identities (62, 45 and 35 unique numeric tokens checked).

The skill triggered WITHOUT being named in the prompt and selected the right one
of two installed skills from the question alone, so no special invocation syntax
is needed: the generated `description`, which carries the bundle id, the concept
count and the ref, is enough to route on.

One defect the runs found, and it was in the prose rather than the payload. The
citation guidance listed the four locator keys this library writes, so on the
270-concept third-party bundle the model reported "no page locator, the address
is at document level" while the excerpt in front of it carried
`source_element_id` - that bundle's own locator, correctly delivered by the
prefix rule. The guidance now tells the reader to cite whichever `source_*` keys
are present. On the re-run the same question returned the element id. Two runs
of one question, the second measuring a changed artefact and not retrying the
first.

One finding that is not a defect in this chain: the first attempt at a
known-negative was not one. The bundle covers water and frost protection on 17
of its 270 concepts and the ranker put none of them in the cut. The consumer
behaved exactly as the contract asks - refused, named its denominator, reported
its own zero as unmeasured because `withheld` entries carry no titles, and did
not go around the cut. Recorded as a retrieval miss rather than replaced, and
it is the same shape as the open fusion finding.

A correction to this session's own measurement is in the record too: a first
sweep used `grep -rhoE "^source_[a-z_]+:"`, whose character class excludes
digits, and so missed `source_sha256` on 270 of 270 concepts. A pattern that
cannot match what it is looking for returns a zero that reads like a fact.

README gains "Consume in Claude Code": folder to answer in three commands, every
one of them run in this session. A test holds that the recipe invokes only
scripts this repository ships, at the paths it names.

Suite 1373 (1339 at the session baseline), ruff clean, mypy src clean. No
version bump, no tag, no push.

Co-Authored-By: Claude <claude-opus-5>
2026-09-08 15:32:32 +02:00

554 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# llm-ingestion-okf
Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import. Security delegated to llm-ingestion-guard.
Status: phases 13 are implemented. Phase 1 (spec-based ingestion) covers
manifest validation, the `file`/`sql`/`http` connectors, deterministic
materialization, index generation, and the golden fixture suite under
`examples/`. Phase 2 adds the bundle inbox (`process_inbox`) and
external-bundle import (`import_bundle`), both against an **injected** persist
gate, with `llm_ingestion_okf.guard_adapter` wiring that gate to the real
guard (see below). Phase 3 makes the bundle contract configurable, so types,
layers, frontmatter sets, index shape, and reserved-file policy are carried by
a profile rather than by constants (see [Upstream OKF
versions](#upstream-okf-versions)). Binary extraction runs behind the
optional `[extract]` extra: `pdf` through a PDF parser, and five office
formats through a vendored document converter. Three of those five office
rows are **unmeasured** — see [Binary extraction](#binary-extraction). Phase 4
(the Node half) is planned (see `docs/plan/`).
## Install
Python 3.10+. Neither this package nor the guard it depends on is on a package
index yet. With uv, one command is enough:
```
uv pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
```
uv resolves the guard on its own, because it reads the `[tool.uv.sources]`
entry in the `pyproject.toml` **of the tag it is installing**, and `v0.4.0`
points that entry at the guard tag below. Measured 2026-07-25 and re-measured
2026-08-20 with an empty `uv` cache; both runs installed
`llm-ingestion-guard==0.2.0` + `llm-ingestion-okf==0.4.0` and imported clean.
With plain pip, the transitive git dependency does not resolve on its own —
**install the guard first**, or installing this package fails with
`No matching distribution found for llm-ingestion-guard`:
```
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v0.2.0"
pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
```
The guard tag is paired to the okf tag, not to this branch: `v0.4.0` declares
`llm-ingestion-guard>=0.2,<0.3`, which `v0.2.0` satisfies and later guard tags
do not. `main` has since moved its own pin to `>=1.2,<2.0` (see
[Requirements](#requirements)); that pin reaches you in the next stable tag,
not in the commands above. Reading a pin off this branch and installing it
against `v0.4.0` is the one combination that fails.
`v0.4.0` is the current stable tag. `v0.5.0a2` is a pre-release for the named
OKF v0.2 pilot set only; pin it only if you are one of them (see
[Upstream OKF versions](#upstream-okf-versions)).
## Build
Installing the package installs one command. A folder of documents in, an OKF
bundle out:
```
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
```
It walks the folder recursively, proposes a segmentation for each document with
the mechanical rules, replays those proposals through the bundle inbox, writes
the bundle and its `log.md`, and prints the run's numbers. Every proposal is
marked `PROPOSED` and `adjudicated: false` — the command segments nothing a
human has approved, and says so in the artifact.
The last line that matters is the conservation identity: `merged + coded
rejections == N`, where `N` is the folder's file count read at run time. **The
run exits non-zero when it does not hold**, and names the unaccounted files, so
a pipeline cannot mistake a partial bundle for a complete one.
Flags worth knowing: `--segments off` ingests each document as one concept and
asks for no root values; `--plans-dir` keeps the proposals instead of
discarding them; `--report` writes the full report to a file as well as stdout.
`--ingested-at` and `--proposed-at` default to `1970-01-01T00:00:00Z` rather
than the clock, so two builds of the same folder are byte-identical — a
wall-clock default would break rebuild-equals-incremental for every caller who
did not pass them.
Measured 2026-09-08 on a 43-file corpus (33 `pdf`, 5 `docx`, 2 `xlsx`, and
three files no reader accepts), one
`okf build` invocation replacing the shell loop over `tools/` that produced the
same corpus's bundle on 2026-09-03:
| figure | value |
|---|---|
| `N` (folder file count, computed) | 43 |
| merged | 39/43 |
| coded rejections | 4/43 (`extractor_unknown` 3, `extractor_empty_pdf` 1) |
| K1b | `39 + 4 = 43 = N`, exit `0` |
| files written | 1108 |
| identical to the 2026-09-03 bundle | 1104/1108 |
| wall time | 842.82 s total, 19.600 s per file (re-measured 2026-09-08) |
The six files that differ are all in the corpus's two spreadsheet documents, and
they are the change reported in `docs/2026-09-08-prisform-og-loggen-k2.md`: a
spreadsheet's tables are now written as pipe tables, so each row is one line
with its cells delimited rather than padded out to the widest cell in the
column. Two concept files are renamed by it, two are removed under their old
names, and the two documents' own `index.md` follow. The root `index.md` is
identical to the stored one again, because this library no longer links the
bundle's `log.md` from it. The command's own byte-identity test compares it
against the two scripts at the current commit, where the two agree over the
whole tree.
## Consume
The other direction: a bundle plus one question in, one bounded, contract-shaped
payload out.
```
python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json
```
`tools/okf_consume.py` is the **pre-pass** `docs/consumption-contract.md` § 1
defines — the deterministic program that reads the bundle, ranks its concepts,
cuts them to a bounded set and emits one payload. It decides nothing about the
question; the skill that reads the payload does the judgement. It calls no
model, opens no socket, imports nothing outside the standard library and this
package, and takes no clock: the same bundle bytes and the same
`(question, k, limit, cost_vocabulary, reserve_top_rank, rarity_weight)`
produce byte-identical output.
`--cost-vocabulary` is off by default and widens one question class: it lets a
declared list of cost/price/quantity terms bridge a question and a document that
name money with different words. The gate is the question — one naming no such
term gets byte-identical bytes either way — and what it does and does not close
is measured in `docs/2026-09-08-blindsone-below-k-k2.md`.
`--reserve-top-rank` is off by default and answers a different objection: the
budget is packed by an exact knapsack, which maximises a SUM of scores and
therefore has no opinion about rank, so a top-ranked excerpt costing a large
share of the budget is out-summed by many small ones. Measured, that made `--k`
a dial that could EVICT the concept a question was asked about. The flag gives
rank one its bytes before the pack runs — after the `over_budget_alone`
pre-exclusion, never before — and the payload then declares
`budget.reserved`. On a 629-concept corpus it changed the delivered set in 2 of
24 measured combinations, both of them that eviction:
`docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
`--rarity-weight` is off by default and weights each lexical hit by
`log(N/df)` over the bundle's own concepts instead of counting it as one, so a
requirement number is not worth what a common verb is worth. The default being
off is a measurement rather than a preference: on four corpora it took one gold
concept from withheld to delivered and a priced sheet from candidate rank 10 to
2, left one gold rank unmoved, and cost another seven rank positions — because
the four-character prefix matcher makes a unique identifier read as
135-of-446 common on that bundle. Where it cannot help is decomposed rather
than guessed: RRF fuses RANKS, so a weight moves nothing on a signal the gold
already leads. `docs/2026-09-08-sjeldenhetsvekt.md`.
It emits the § 8 shape — `contract`, `bundle` (`bundle_id` plus a
`sha256-tree:` content identity), `budget` (unit, instrument, limit, spent and a
validated known-positive), `denominators`, `excerpts` and `withheld` — and every
withheld concept names the rule that dropped it, from a closed set of six.
Every excerpt carries the concept's `title`, and — when the producer wrote them
`req_number`, the SPEC § 5.1 address `sources`, and **every top-level
`source_*` key**, by prefix rather than by allowlist: a fixed list names the
locators its author thought of, and one real bundle locates by
`source_element_id` on 269 of its 274 concepts. A key the producer did not write
stays absent rather than arriving empty, and an address this reader cannot
decode is named (`sources_unreadable`) rather than dropped into the same
silence. The reason is a measurement: with `concept_id` and body text alone, a
delivered gold concept at rank 1 still left the answer unable to name the
document it was quoting.
`considered == withheld + delivered` closes by construction, and the payload is
refused rather than reported when it does not.
Three exit codes, not two: **0** a payload was written, **1** the run happened
and refused (the budget admitted none of the concepts that answered the
question, or an asserted `--ref` contradicted the bytes), **2** the run did not
happen. Collapsing 2 into 1 would report an unread bundle as a failed cut.
`--ref` is an **assertion**, never an override — the identity is always computed
from the bytes, because labelling a payload with an identity its bytes do not
have is the one thing § 3.3 exists to prevent.
Check any payload against the skill that will read it:
```
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json
```
`skills/okf-consume/` is the first instantiated consumption skill: a filled copy
of `skills/okf-consume-template/` naming this pre-pass, with every per-corpus
hole replaced by a measured value. Measured 2026-09-07 on a 629-concept bundle,
hit@8 was **5 of 6** questions at rank 1 against a chance baseline of **1.35 of
6** — with one control that failed, and both are in
`docs/2026-09-07-okf-konsumskill-maaling.md` with the honesty limits stated.
## Consume in Claude Code
A folder of documents to an answer a model can cite, in three commands. Every
command below was run end to end on 2026-09-08 against a nine-document folder
and a 270-concept third-party bundle; nothing here is untested.
```sh
SRC=/tmp/c1-fresh-src # the folder of documents
BUNDLE=/tmp/c1-fresh-bundle # where the OKF bundle goes
PROJECT=/tmp/c1-scratch # the project you will ask the question from
```
**1. Build the bundle.**
```sh
okf build "$SRC" --bundle "$BUNDLE" --bundle-id c1-fresh-20260908 --okf-version 0.2 --ingested-at 2026-09-08T00:00:00Z
```
**2. Generate a skill for that bundle**, straight into the project's skill
directory. The skill is instantiated for these bytes: its id, ref, concept
count, per-field denominators, whole-bundle cost and breaking point are all
measured from the bundle, and it ships a reference payload the checker accepts.
```sh
python3 tools/okf_skill.py "$BUNDLE" --out "$PROJECT/.claude/skills/c1-fresh-20260908-consume"
```
Repeat for every bundle you want reachable; each one gets its own skill named
after its `bundle_id`, which is what lets a model pick between them. A bundle
you only have read access to is fine — the generator only reads it.
**3. Ask.** From `$PROJECT`, in Claude Code:
```sh
claude -p "Hvordan skal prisene fylles ut?"
```
Measured with two bundles installed side by side: the model selected the right
skill from the question alone, ran the pre-pass and the contract check itself,
quoted the requirement verbatim, and named the document, the requirement number,
the source resource and the locator inside it. Across three questions, **0**
numbers or identifiers appeared in an answer that were not in the delivered set
or in the payload's own identities. On a question the bundle does not cover it
answered `[sourced-not-sufficient]` and reported the denominator rather than
inventing an answer.
The pre-pass and the checker are the same two commands the skill runs for you,
if you want to see the payload first:
```sh
python3 tools/okf_consume.py "$BUNDLE" --question "your question" --out /tmp/payload.json
python3 tools/okf_contract_check.py --skill "$PROJECT/.claude/skills/c1-fresh-20260908-consume/SKILL.md" --payload /tmp/payload.json
```
The honest limits: this was measured on **four questions** across two bundles,
which is a demonstration and not a hit rate. The ranking is lexical, and one of
the four found a topic the bundle **does** cover and did not rank it into the
cut — the skill then said so with its denominator instead of answering, which is
the behaviour the contract asks for, but a miss is still a miss.
`docs/2026-09-08-claude-code-skill-vilkaarlig-bundle.md` has the runs.
## Implemented scope (v1)
The library provides three entry points for getting content into an OKF
bundle:
1. **Spec-based ingestion.** An implementation of the normative ingest
specification owned by `portfolio-optimiser-commons`: manifest →
`file`/`sql`/`http` connector → deterministic materialization of
`ingest-{id}.md` concept files → index generation. Zero model calls in the
run path; output is reproducible byte-for-byte against golden fixtures.
2. **Bundle inbox.** A drop directory where common file types are converted
to OKF concept files. All file-type→text extraction lives in this library:
`md`, `txt`, `csv`, `json`, and `html` are handled by the stdlib core;
`pdf` and the five office formats (`docx`, `xlsx`, `pptx`, `odt`, `rtf`)
require the optional `[extract]` extra and are rejected fail-fast without
it. Extracted text passes the security gate before anything is persisted.
The drop directory is walked **recursively**, in sorted relative-path order:
a file at any depth is ingested and records its path relative to the inbox
root as its `source_file`, while dot-directories and a bundle directory
sitting inside the inbox are skipped with a reported code.
Under the segmented v0.2 profile a concept also points back at the document
it was extracted from, so an agent citing it can open the original at the
right place: `sources: [{ resource, title }]` in the spec's own §5.1 form,
where `resource` is the inbox-relative path, plus a locator per format —
`source_pages` for a PDF, `source_sheet` and `source_rows` for a
spreadsheet, `source_lines` otherwise. The locator keys are this library's
own, because §5.1 has no field for a place *within* a resource; the line
numbers index the extracted text and say so. Measurements:
[`docs/2026-09-08-proveniens-k2.md`](docs/2026-09-08-proveniens-k2.md).
<!-- extract-formats: .md, .txt, .csv, .json, .html, .htm, .pdf, .docx, .xlsx, .pptx, .odt, .rtf -->
3. **External bundle import.** Import and merge of third-party OKF bundles:
each concept is assessed via the security gate, and only concepts that
pass are merged, materialized, and linked into the index.
## Boundary: security is delegated
Security is owned by the sibling package
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
(pinned `>=1.2,<2.0`). The division is strict:
- **guard** answers "is this content safe to persist?" — scan, sanitize,
quarantine, fail-secure, provenance stamping.
- **this library** does the plumbing — connect a source, materialize a
deterministic OKF bundle, generate the index.
No security functionality is reimplemented here.
### What is gated today: read this before trusting a door
- **Door A (`materialize_bundle`) is ungated.** It calls nothing before
writing to disk and writes what it is given. A caller materializing
untrusted content is responsible for gating it.
- **Doors B and C (`process_inbox`, `import_bundle`) gate through an adapter
you pass in.** Each takes a `gate` argument; the flow hands it the content
and obeys the verdict, refusing to persist anything that does not clear the
guard's non-blocking floor — including a disposition it does not recognise,
and (at Door C) a concept the gate returned no verdict for. What it cannot
do is check that your adapter is a real guard: a permissive stub approves
everything, and the flow will believe it.
`llm_ingestion_okf.guard_adapter` is the adapter over the real guard, and the
only module here that imports it — importing the package itself does not:
```python
from llm_ingestion_okf import process_inbox
from llm_ingestion_okf.guard_adapter import inbox_gate
result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
okf_type="reference", gate=inbox_gate)
```
Two properties of that adapter are worth knowing before you rely on it.
It screens the **exact bytes it persists** — the guard's `prepare_input`
bookend prepares text for a model call, which this library never makes, so
only `screen_output` is used and the screened string is the written string.
And it **refuses rather than repairs**: a file carrying an invisible
zero-width or bidi character is rejected, not silently stripped and written.
Door B screens under the untrusted-upload policy, so any finding at all is
held back rather than persisted.
This section is stated plainly because earlier wording ("calls the guard at
every persist gate") described the intended end state in the present tense,
and a consumer reasonably read it as safe-by-default.
## Roadmap
The library is built in four phases so that every known OKF surface in the
ecosystem is eventually covered. Each phase has a detailed plan with
verification criteria:
1. Spec-based ingestion (Python) with byte-exact golden fixtures —
[plan](docs/plan/phase-1-door-a.md).
2. Bundle inbox and external-bundle import (Python), guard-gated —
[plan](docs/plan/phase-2-doors-b-c.md).
3. Configurable bundle contract (types, layers, frontmatter sets, index
shape, and reserved-file policy as configuration), enabling stricter
bundle profiles such as `strict-v1`
[plan](docs/plan/phase-3-configurable-contract.md).
4. A `node/` half: a zero-dependency Node/ESM package (importable and
CLI-invokable, vendored per consumer) providing bundle checking, index
generation, inbox processing, and document conversion for the OKF
second-brain plugin ecosystem. The Python and Node halves share the OKF
contract and fixture suite, not code —
[plan](docs/plan/phase-4-node-half.md).
## Upstream OKF versions
The library targets the current latest version of Google's OKF. Support is
**additive** — a new upstream version arrives as a new profile, never as a
migration of an existing one — so an *upstream* release does not change the
bytes an existing profile emits.
That guarantee is about upstream, and one profile tracks a second contract as
well. `DEFAULT` states the ingest-spec owned by `portfolio-optimiser-commons`,
so when they change that spec, `DEFAULT` follows them. It happened on
2026-08-09: `generated` moved from `true` to
`{ by: process:okf-ingest, at: <ingested_at> }`, one changed line per generated
file. Upgrading across it costs a re-run and nothing more — a profile still
recognises bundles stamped by earlier versions, so re-running writes in place
instead of refusing. `DEFAULT` remains OKF v0.1 on every axis upstream owns.
| Profile | Contract | Status |
|---|---|---|
| `DEFAULT` | commons' ingest-spec layer (OKF v0.1 semantics) | stable |
| `STRICT_V1` | a consumer's ratified v0.1 contract | stable |
| `OKF_V0_2` | OKF v0.2 | **provisional**, pre-release only |
| `STRUCTURED_V1` | `DEFAULT` plus a faceted, derived index | stable |
| `OKF_LATEST` | alias for the latest version supported as *stable* | currently `DEFAULT` |
`STRUCTURED_V1` is `DEFAULT` in every respect but the index. Under it, Door B
derives each dropped document's title, number, hierarchy and cross-references,
writes them into the concept's own frontmatter, and carries them into the index
entry — so a consumer can reason over the bundle rather than only look things
up in it. Every inferred field is named in a `derived` list, because an
unmarked heuristic is worse than no heuristic: the consumer cannot know when to
doubt it. A pointer to a document not dropped yet is rendered `N200?` rather
than omitted, since a bundle is built up over several drops and an absence that
leaves no trace is the dangerous kind. Carrying the metadata costs index
characters — roughly 3x to 6x the flat index, depending on how many facets the
profile names — and the facet key set is the dial. Design record and
measurements: [`docs/plan/structure-derivation.md`](docs/plan/structure-derivation.md).
`OKF_V0_2` ships first as a pre-release to a named pilot set and may change on
their feedback without a deprecation cycle. Pin the versioned constant rather
than `OKF_LATEST` unless you have explicitly opted into tracking; `OKF_LATEST`
moves at general availability, which is a deliberate release event rather than
a side effect of an upgrade.
Selecting a profile is keyword-only, so existing call sites are unaffected:
```python
materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)
```
A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and
puts the declaration in the bundle-root `index.md`'s frontmatter block. The
profile names the key; the **caller supplies the value**, because that value
tracks the upstream version and is not this library's to decide:
```python
materialize_bundle(
manifest, bundle_dir, ingested_at,
profile=OKF_V0_2,
root_frontmatter_values={"okf_version": "0.2"},
)
```
Omit the argument and no frontmatter block is written. Offering a key the
profile does not name is refused before anything is written to disk.
### Attested computations (v0.2 §10)
`OKF_V0_2` supports the `Attested Computation` type as a **format**: its five
contract fields — `runtime`, `parameters`, `computation`, `executor`,
`attester` — are emitted in canonical position, judged, and round-tripped.
`runtime` is required for that type and for no other, which the profile
expresses through `FrontmatterSchema.required_by_type`; a type the mapping does
not name carries no extra requirement, because §14 forbids a consumer to reject
on an unknown `type`.
Nothing here executes a computation or checks an attestation. Upstream defers
the receipt and verdict wire formats, so there is no contract to implement, and
the question an attestation answers — was this value produced the sanctioned
way — is not this library's. It re-enters scope when upstream specifies the
protocol.
On the import side, a third-party concept may name an `executor` or `attester`
resource pointing at executable code. Door C imports the **pointer** and never
the code — it writes concepts verbatim and skips every non-`.md` file — so such
a reference may not resolve, or may resolve to a file the destination tree
already holds under that path. Each one is reported in
`ImportResult.unverified_references`; the concept still merges, because §14
forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer
to surface rather than silently drop. The report names the pointer key, not the
resource it points at: recovering the resource needs the structured reader.
One limit worth knowing before you write such a concept: §10.2 presents
`executor` and `attester` as nested block mappings, and this library's
frontmatter parser is line-oriented. It reads inline **flow** mappings
(`executor: { resource: …, receipt: [ … ] }`) as opaque values that round-trip
unchanged, but it cannot read the block form — two block mappings that both
carry a `resource` collapse into one namespace and the first is lost. Write the
flow form; both are valid YAML, and a real YAML consumer recovers the same
structure from either.
## Non-goals
- Verdict/feedback machinery from the method specification (stays in the
consuming repositories).
- Embedding- or retrieval-layer functionality.
- Security functionality, in either runtime — that is always
`llm-ingestion-guard`'s domain.
## Requirements
Python 3.10+, and exactly one runtime dependency — the security boundary,
`llm-ingestion-guard>=1.2,<2.0`. Everything else is stdlib. The commands are
under [Install](#install); what follows is why they look the way they do.
From a checkout, the test suite runs with:
```
.venv/bin/python -m pytest
```
The suite is the verification surface for everything above: 596 tests, run on
2026-08-21 against this branch with the `[extract]` extra installed. Without
the extra the same suite is 589 passed and 7 skipped, measured the same day:
the seven cover the parser path, and the tests holding the fail-fast rejection
for an uninstalled extra run in both. It is not shipped in an installed
distribution — `tests/` lives at the repository root, so this command needs a
clone rather than a `pip install`.
A git URL is a PEP 508 direct reference and pins one exact tag, so it is an
install-time *channel*, not the pin: the range above stays the declared
dependency — a wheel built from this branch carries `Requires-Dist:
llm-ingestion-guard<2.0,>=1.2`, measured 2026-08-23 — and resolves normally
once the package index exists. A wheel built from a *tag* carries that tag's
range instead, which is why the install commands pair tag with tag.
### Binary extraction
The optional `[extract]` extra ships two things: `pdfplumber` (MIT) for `pdf`,
and `pypandoc-binary` for five office formats. It is opt-in because it pulls
binary wheels, which the default install must never do — the single runtime
dependency rule covers the default install and this extra sits outside it.
The converter **binary travels inside the wheel** and is resolved by path
rather than found on `PATH`, with its version asserted against a pin. A host
carrying a different converter is refused, not silently used: extraction is
deterministic within a converter version and not across one.
| Format | Reader | Evidence |
|---|---|---|
| `pdf` | `pdfplumber` | measured |
| `docx` | converter | measured |
| `xlsx` | converter | measured |
| `pptx` | converter | **unmeasured** |
| `odt` | converter | **unmeasured** |
| `rtf` | converter | **unmeasured** |
**`unmeasured` means what it says.** The corpus this work was measured on
contains **zero** `pptx`, `odt` and `rtf` files, so those three rows work by
construction and have never been checked against a document anyone wrote.
They are not known to be broken; they are not known to be right either, and
the distinction is the point.
**What stays out.** `.doc` (Word 97) is not supported — the converter does not
read it. Rastered or scanned PDFs are refused rather than persisted as empty
concepts, because this library does not do OCR. Drawn content — figures,
diagrams, shapes — does not survive extraction in any format here, and every
extraction says so with a warning. Structured table recovery is out of scope.
Request it by appending `[extract]` to the package name in whichever install
command from [Install](#install) you are using — this package is not on an
index, so a bare `pip install 'llm-ingestion-okf[extract]'` does **not** work
today, and the error message naming that command is written for the day it
does. The extra is unreleased: it reaches a consumer through a tag that
contains it, and no such tag exists yet.
Two properties of the extra are worth knowing before depending on its output:
- **Extracted text is pinned to an exact parser version.** `pdfplumber` pins
`pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
releases with no stability contract. Extraction is deterministic within a
parser version and not guaranteed across one, so a golden fixture built on
extracted PDF text is a fixture migration away from any parser upgrade.
- **Text extraction recovers text, and nothing that is drawn.** Figures,
diagrams and images have no text to recover — only their captions survive —
so a bundle built from drawn documents is incomplete by construction. The
library says so itself: every `pdf` extraction emits an `ExtractionWarning`.
Structured table recovery is separately out of scope; PDFs enter as prose.
The planned Node half targets Node/ESM with zero npm dependencies.
## License
MIT — see [LICENSE](LICENSE).