llm-ingestion-okf/README.md
Kjell Tore Guttormsen 4d1f9d3a5a chore(release): 0.8.0
Round 13 (the `.xml` core file type read as NISO-STS, and the PDF arm's
collision counter) and round 14 (a section the source DECLARES takes the
declared-structure route: `.xml` goes from 15 of 2 761 to 2 761 of 2 761
boundaries and from 23 to 2 761 concepts at the shipped defaults, hit@1/8/50
0/6 - 0/6 - 0/6 to 3/6 - 5/6 - 6/6) are both landed. This commit adds no
functionality: it sets the version, closes the CHANGELOG entry, and points
every install line at the new tag.

- `pyproject.toml` and `src/llm_ingestion_okf/__init__.py`: 0.7.0 -> 0.8.0.
  The second is the only line in `src/` this release touches. It is not a
  code change but the other half of the version, written without a `v`
  prefix, so a search for `v0.7.0` cannot find it;
  `test_packaging.py::test_the_declared_version_agrees_with_the_packaged_one`
  is what did. Left alone, the tag would report the previous release to every
  consumer that installs it.
- `CHANGELOG.md`: `[Unreleased]` becomes `[0.8.0] - 2026-09-10`, with a new
  empty `[Unreleased]` above it. The entries are round 13's and round 14's own
  words, unchanged. No compare link is added: this file has carried none since
  `[0.6.0]`, and inventing one here would be a claim about a URL nobody checked.
- The five install lines and the two prose lines naming the current tag move to
  `v0.8.0`. The tag history list gains a `v0.8.0` row as the current tag and
  KEEPS the `v0.7.0` row: that list states it is not install lines, so a
  rewrite would delete history rather than update it.
- README's test count was 1515, measured 2026-09-09; this tree measures 1575
  passed / 1 skipped with ruff 0.16.6. The surrounding sentence about the
  earlier figure is repaired too, because changing the date alone would have
  made it false.

`v0.7.0` stays on 1260fac. No lock change, no history rewrite in `docs/`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 21:09:22 +02:00

883 lines
56 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# llm-ingestion-okf
Turn a folder of documents (PDF, DOCX, XLSX, PPTX, MD) into a bundle a model can
answer from **with a source on every claim** — offline, deterministic, no model
call anywhere in the run path.
## Install
Python 3.10+ and [uv](https://docs.astral.sh/uv/). One line:
```sh
uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.8.0"
```
## Use it
```sh
okf project ~/my-documents # folder in: bundle + Claude Code skill, in this directory
claude # start Claude Code here
```
Then ask in plain language. Three shapes of request work, and the skill states
the rules for each:
- **a question** — "hva er kravene til pris?"
- **a hypothesis** — "stemmer det at leverandoeren baerer risikoen for grunnforhold?"
Answered per premise as `confirmed` / `refuted` / `undecidable-from-bundle`.
- **a task whose answer is a document** — "lag `krav-pris.md` med alle krav til
pris, ett avsnitt per krav, med dokument og kravnummer." Every claim in the
written file carries its source; a paragraph with no ground is written
and marked, never dropped.
Everything below is detail: [Consume in Claude
Code](#consume-in-claude-code) for the same thing in steps and with several
bundles at once, [Build](#build) for the flags, [Requirements](#requirements)
for the pip fallback and the guard pairing.
## What this library is
Status: phases 13 are implemented. Phase 1 (spec-based ingestion) covers
manifest validation, the `file`/`sql`/`http` connectors, deterministic
materialization, index generation, and the golden fixture suite under
`examples/`. Phase 2 adds the bundle inbox (`process_inbox`) and
external-bundle import (`import_bundle`), both against an **injected** persist
gate, with `llm_ingestion_okf.guard_adapter` wiring that gate to the real
guard (see below). Phase 3 makes the bundle contract configurable, so types,
layers, frontmatter sets, index shape, and reserved-file policy are carried by
a profile rather than by constants (see [Upstream OKF
versions](#upstream-okf-versions)). Binary extraction runs behind the
optional `[extract]` extra: `pdf` through a PDF parser, and five office
formats through a vendored document converter. Three of those five office
rows are **constructed** rather than measured — see
[Binary extraction](#binary-extraction). Phase 4
(the Node half) is planned (see `docs/plan/`).
## Install in detail
Neither this package nor the guard it depends on is on a package index yet, so
both install by direct reference. With uv, one command resolves both:
```sh
uv pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.8.0"
```
uv resolves the guard on its own, because it reads the `[tool.uv.sources]`
entry in the `pyproject.toml` **of the tag it is installing**, and `v0.8.0`
points that entry at `llm-ingestion-guard` `v1.3.0`. Use `uv tool install`
instead of `uv pip install` when you want the `okf` command on `PATH` without an
active virtualenv — that is the form the first screen shows.
With plain pip, the transitive git dependency does not resolve on its own —
**install the guard first**, or installing this package fails with
`No matching distribution found for llm-ingestion-guard`:
```sh
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.3.0"
pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.8.0"
```
The guard tag is paired to the okf tag, not to this branch. `v0.8.0` declares
`llm-ingestion-guard>=1.2,<2.0`, which `v1.3.0` satisfies; the pairing above is
read off that tag's own `[tool.uv.sources]`, not off this branch. Reading a pin
off `main` and installing it against an older okf tag is the one combination
that fails.
### Earlier tags, as history
These are not install lines. They record what each earlier tag was, so a reader
who meets one in an older document knows what they are looking at.
- `v0.8.0` — the current tag: `.xml` is a core file type, read as NISO-STS
through the stdlib parser, and a section the source DECLARES takes the
declared-structure route — one publisher's process code segments at 2 761 of
2 761 of its own declared sections at the shipped defaults. No other file
type changes one byte, measured on the bytes.
- `v0.7.0``okf project` builds the bundle `okf build` builds (they were one
flag apart before it), and the generated skill states the question,
hypothesis and document-task modes with relative paths.
- `v0.6.0` — the first tag carrying the `okf project`, `okf consume`,
`okf check` and `okf skill` subcommands.
- `v0.5.0a2` — a pre-release for the named OKF v0.2 pilot set only.
- `v0.4.0` — the last tag before the OKF v0.2 work; it declares
`llm-ingestion-guard>=0.2,<0.3`, which only guard `v0.2.0` satisfies.
**No tag yet makes OKF v0.2 generally available.** `OKF_LATEST` is unchanged and
still points at `DEFAULT`; flipping that alias is the GA event and none of the
tags above is it (see [Upstream OKF versions](#upstream-okf-versions)).
## Build
Installing the package installs one command. A folder of documents in, an OKF
bundle out:
```
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
```
It walks the folder recursively, proposes a segmentation for each document with
the mechanical rules, replays those proposals through the bundle inbox, writes
the bundle and its `log.md`, and prints the run's numbers. Every proposal is
marked `PROPOSED` and `adjudicated: false` — the command segments nothing a
human has approved, and says so in the artifact.
The last line that matters is the conservation identity: `merged + coded
rejections == N`, where `N` is the folder's file count read at run time. **The
run exits non-zero when it does not hold**, and names the unaccounted files, so
a pipeline cannot mistake a partial bundle for a complete one.
Flags worth knowing: `--segments off` ingests each document as one concept and
asks for no root values; `--plans-dir` keeps the proposals instead of
discarding them; `--report` writes the full report to a file as well as stdout.
`--ingested-at` and `--proposed-at` default to `1970-01-01T00:00:00Z` rather
than the clock, so two builds of the same folder are byte-identical — a
wall-clock default would break rebuild-equals-incremental for every caller who
did not pass them.
### The segmentation flags
Nine rules are reachable from `okf build`, and since 2026-09-11 **all nine
are ON by default** — `--outline-run 3`, `--table-grid` and `--unit-fold` since
2026-09-08, `--drop-wrapped-outline` and `--outline-gate` since 2026-09-09,
`--sheet-section-rows`, `--keep-table-heading` and `--first-span-from-zero`
since 2026-09-10, and `--close-span-gaps` since 2026-09-11 — each an operator
decision, and each with an explicit
opt-out: `--outline-run 0`, `--no-table-grid`, `--no-unit-fold`,
`--keep-wrapped-outline`, `--no-outline-gate`, `--no-sheet-section-rows`,
`--no-keep-table-heading`, `--no-first-span-from-zero`,
`--no-close-span-gaps`. Passing all nine
reproduces the pre-2026-09-08 bytes exactly, and the last three reproduce the
pre-2026-09-10 bundle byte for byte — measured with `diff -rq`, 0 differences,
not asserted. Each line below carries the number it was measured at, and
nothing beyond it.
An **eleventh** flag, `--pdf-outline`, is **off** by default, and it is the one
rule here that does not read the extracted text at all: it cuts a PDF at the
boundaries the file's own `/Outlines` bookmark tree declares. It is not
`--outline-run` under another name — that one is a text heuristic over numbered
lines in the extracted text, while this one opens a structure index the PDF
already carries and the build had never looked at.
It is a segmentation arm and not a reader option: the extracted text is byte
for byte the same either way, and a PDF that carries no bookmark tree builds
byte-identically with the flag on. Two consequences come free with it. The
concept title comes from the BOOKMARK, so it is not cut short where the page
wrapped the heading across two lines; and a page that lies before the first
bookmark destination is the contents listing rather than a second copy of the
body, so a contents entry and the section it lists stop landing as two concepts
under one id.
The measurement is one 701-page process code whose publisher also ships a
NISO-STS structure for it, so the fasit is the publisher's own. Under the
shipped default that document gives 1967 of 2761 boundaries, none of its 28
chapters, and 794 of 794 misses have their heading text present in the text the
build read. With the arm it gives 2759 of 2761 and 28 of 28. The flag stays off
because reach is the open question, not quality: **1 of the 8** reference PDFs
in this repository's own sample carries a usable tree, and a bookmark tree is
the publisher's *claim* about its own structure — a stale or wrongly pointing
one carries that error straight into the segmentation.
A **tenth** flag, `--bold-title`, is **off** by default. It is the rule for the
type whose container declares nothing: `rtf` has no heading style, so an
author's title is bold text, and the row was measured at 0 of 0 declared
headings, 0 concepts and 1368 of 1368 characters in no segment. The grammar is
markdown, not `rtf` — the converter already writes that title as `**…**` in the
same output every office row produces — and it is gated by the principle
`--outline-gate` already carries: recovery yields to declaration. Measured over
47 readable documents, three parameters were swept and one carried; false
positives are **0 of the 31 documents that declare**, and the rule reaches
**2 of 39** corpus documents, both `docx`, **0 of 33 `pdf`** and **0 of 2
`xlsx`**. On the `rtf` fixture set it recovers **6 of 6 authored titles over
N = 4** with **0** false titles and **0 of 1994** characters in no segment. On
a five-document folder it moves 26 → 27 concepts, replacing a mechanical
`tabell-linje-30` with two named concepts.
**A re-run is what this costs a consumer, and it is not a small one:** on the
43-document reference corpus the default bundle goes from **629 concepts in
1108 files** (the 2026-09-03 tree) to **492 in 944** after the 2026-09-08 move,
to **425 in 810** after the 2026-09-09 one and to **436 in 832** after the
2026-09-10 one (digest `8dff8a8e6c15d2f7…`). On a five-document folder the last
move is **15 concepts in 30 files → 26 in 52**. The proposer's own defaults
(`tools/okf_propose_segments.py`) did NOT move, so every published reproduction
block still runs as written.
**What the 2026-09-09 move had to clear, stated because it is the bar every
later move is held to:** the twelve-position reference improves (`pdf` **2 of 8
→ 7 of 8**, the sheet **5 of 12 → 10 of 12**, `docx` unchanged at 3 of 3) AND
hit@8 on a K2 bundle built with it holds **5 of 6 at ranks 1,1,1,1,1,**, with
no row losing rank 1. A configuration that improved the reference and cost a
rank was measured in the same session and did NOT ship; see
`docs/2026-09-09-k3-runde6-outline-gaten-og-prioren.md` § 9.1.
| flag | what it does | measured |
|---|---|---|
| `--outline-run N` (default **3**) | also propose a boundary where the document's own bare-integer numbering sustains an ascending run of at least `N`; `0` is this arm's opt-out | a tender PDF whose headings are bare integers: **no boundary** at `0`, **9 concepts** at `3`, against a reference of 9 |
| `--table-grid` (**on** by default; opt out with `--no-table-grid`) | a pandoc grid-table rule line no longer closes an open table block, so one grid table is one concept | a `.docx` experience list: **21 → 6** concepts |
| `--unit-fold` (**on** by default; opt out with `--no-unit-fold`) | discard a contents-list run, fold a deeper heading into its parent, fold a table into the shorter heading that introduces it. Adds no boundary, so it can only reduce a plan | on a 12-document sample scored against an operator's unit worksheet: **5 of 12** match — but that figure was measured with `--table-grid` ON, and the shipped default does not include it. Measured without it the same sample scores **2 of 12**, `docx` **0 of 3**, because the fold's table clause has no joined table to fold |
| `--keep-table-heading` (**on** by default since 2026-09-10; opt out with `--no-keep-table-heading`) | keep a heading whose body is empty only because a table opens under it, and absorb that table into its span | the two spreadsheets in that corpus, and **0 of 32 `pdf` and 0 of 5 `docx`**: the concept count does not move (1 → 1), its first byte does — the concept gains the heading line it was missing |
| `--sheet-section-rows` (**on** by default since 2026-09-10; opt out with `--no-sheet-section-rows`) | cut an open table block at the rows that label its sections — a run of at least three rows whose first cell is a bare numeric label. The opposite direction from `--table-grid`, which decides how far a block extends | a tender price sheet whose whole body is one table block: **1 → 12 concepts**, against a reference of 11 cost groups plus the sheet's preamble. Whole corpus: **1 of 39** readable documents changes, **0 of 32 `pdf`, 0 of 5 `docx`, 1 of 2 `xlsx`**. It reached 11 of 12 on the reference two rounds before it shipped, and was held back both times by a RETRIEVAL cost that turned out not to be its own: on a bundle built with it the gold document splits 1 → 12 concepts and row 1 of the hit@8 set fell rank 1 → 2. The repair is on the reading side (`--tie-shared-rank`, now the default), and with it in place the sheet reaches 11 of 12 with hit@8 holding **5 of 6 at ranks 1,1,1,1,1,** |
| `--drop-wrapped-outline` (**on** by default since 2026-09-09; opt out with `--keep-wrapped-outline`) | do not admit an `--outline-run` candidate whose line continues onto the next one. Judges recovered candidates only, never a heading the document declares | quoted regulation text, whose numbered paragraphs match the outline grammar exactly: **4 → 1 concepts**, the reference. Whole corpus: **5 of 39**, all `pdf`; on the 12-document sample **8 of 34** outline candidates wrap, and none of the 26 the operator kept. On the reference it carries `pdf` from **5 of 8 to 6 of 8** together with the gate below, and neither reaches 7 of 8 without the other |
| `--outline-gate` (**on** by default since 2026-09-09; opt out with `--no-outline-gate`) | admit `--outline-run`'s RECOVERED headings only where the document declares none of its own, plus any one recovered heading whose span covers `OUTLINE_SHARE` (0.20) of the text. Applied at admission, before spans are closed, so the text a removed mark opened is carried by the mark above it rather than lost | on the 12-document sample: `pdf` **2 of 8 → 5 of 8** alone and **7 of 8** with the rule above, `docx` unchanged at **3 of 3**. Whole corpus: it fires on **25 of 39** readable documents, changes the plan in **15 of 39**, and removes **64 of 485** proposed entries. No plan disappears (32 → 32) |
| `--first-span-from-zero` (**on** by default since 2026-09-10; opt out with `--no-first-span-from-zero`) | start the first concept at character 0, so the text above it belongs to a segment instead of to none. Adds no boundary and removes none | Measured over the 39-document corpus, the default before this rule left **207 435 characters — 11.92 %** — in no segment at all: **163 804 above the first entry** (in **32 of the 32** documents that get a plan), 26 041 *between* entries and 17 590 after the last. This rule closes the first part entirely, 79 % of the whole, leaving **43 631 characters (2.51 %) over 8 of 32 documents** with two named mechanisms of their own. It adds no boundary and the K2 concept count is identical with and without it (**425 = 425**); on the 12-position reference it changes **not one cell**, and hit@8 on a K2 bundle built with it holds **5 of 6 at ranks 1,1,1,1,1,** under both tie-breaks |
| `--close-span-gaps` (**on** by default since 2026-09-11; opt out with `--no-close-span-gaps`) | close a concept's span against the next SURVIVING concept, and the last against the end of the text. Three steps remove a candidate AFTER its neighbour's span was already closed against it — the orphan check, and `fold_units` clause 1 both between entries and on the last run — and the removed mark's text then belongs to no segment. Adds no boundary and removes none; only spans' ends move | It closes the whole remainder the rule above left: **43 631 characters, 2.51 % of the corpus over 8 of the 32 documents with a plan, to 0** — both the 26 041 between entries and the 17 590 after the last. Decomposed: orphan check **18 527** over 15 of 39 documents, clause 1 **7 514** between entries, clause 1 on the last run **all 17 590** of the tail (with `unit_fold=False` the corpus tail gap is 0). The entry count is identical (**429 = 429** on the corpus, **436 = 436** concepts on K2, **52 = 52** md on a five-document folder); on the 12-position reference it changes **not one cell** (11 of 12 under `|F|`[3]=12, 10 of 12 under `|F|`[3]=11), and hit@8 holds **5 of 6 at ranks 1,1,1,1,1,** on the new bundle, the previous default and Arm B alike |
| `--pdf-outline` (**off**; opt out is the default, opt in with the flag) | cut a PDF at the boundaries its own `/Outlines` bookmark tree declares, instead of at the ones the text rules recover. A SEGMENTATION arm, not a reader option: the extracted text is byte for byte the same either way, and a PDF that carries no tree builds byte-identically with the flag on. The title comes from the BOOKMARK, so it is not cut short at the page's line break, and a page before the first bookmark destination is the contents listing rather than a second copy of the body | one 701-page process code whose publisher also ships a NISO-STS structure for it, so the fasit is the publisher's own: boundaries **1967 of 2761 → 2759 of 2761**, chapter level **0 of 28 → 28 of 28**, concept titles identical to the source title after normalisation **2761 of 2761**, false positives **163 of 2182 → 3 of 2762**, directories carrying two concept files **132 of 2050 → 2 of 2738**, contents-copy pairs **65 → 0**. Consumption on the same eight questions: fasit present in the bundle **4 of 7 → 7 of 7**, hit@1/8/50 **1/6 · 2/6 · 4/6 → 3/6 · 5/6 · 6/6**. Cost 119.22 s → 183.31 s wall, peak RSS 3252 → 3251 MiB, no new dependency and no second parse of the pages. **Off, and the reach is why:** **1 of the 8** reference PDFs carries a usable tree at all, and a bookmark tree is the publisher's CLAIM about its own structure — a stale or wrongly pointing one carries that error straight into the segmentation |
They compose, and the order above is the order they apply in. Measured on a
five-document tender folder (2 `pdf`, 2 `docx`, 1 `xlsx`), concepts per
document:
| document | default | `--outline-run 3` | `+ --table-grid` | `+ --unit-fold` | `+ --keep-table-heading` | `+ --sheet-section-rows --drop-wrapped-outline` |
|---|---|---|---|---|---|---|
| tender PDF, technical requirements | 1 (no boundary) | 9 | 9 | 9 | 9 | 9 |
| tender PDF, technical layout | 1 (no boundary) | 1 | 1 | 1 | 1 | 1 |
| price sheet `.xlsx` | 1 | 1 | 1 | 1 | 1 | **12** |
| experience list `.docx` | 21 | 21 | 6 | 3 | 3 | 3 |
| agreement `.docx` | 2 | 2 | 1 | 1 | 1 | 1 |
| markdown files in the bundle | 31 | 49 | 33 | 30 | 30 | 52 |
Every column merged 5 of 5 with 0 rejections. The reference for the first row
is 9, so the default is a full arm behind what the proposer can do on that
document — which is a statement about the default, not a licence to change it
here.
### The two PDF reader flags
Separate from the six above, and they sit before every one of them: a
segmentation flag changes how the proposer cuts a text, these change what the
text says. **All three are off by default.**
| flag | what it does | measured |
|---|---|---|
| `--pdf-headings font` | a PDF carries no heading markup, so one is inferred from typography — a line whose dominant font size is above the document's character-weighted median AND whose dominant font name says bold — and emitted as an ATX heading in the same markdown the office path produces, so the existing heading rule reads it | on a tender PDF: **9 of 9** numbered chapters found, plus 4 extra candidates. Whole corpus: **25 of 32 `pdf`** change, **0 of 5 `docx`**, **0 of 2 `xlsx`**. **Off by measurement:** against the operator's unit worksheet it takes `pdf` from **2 of 8 to 0 of 8**, losing two exact matches, because on those documents the outline rule already found the chapters and a second heading source can only add |
| `--pdf-headings font-reserve` | the same typographic rule, applied ONLY to a document whose own numbering the outline arm finds no run of — typography as a second heading source where there is no first one, never on top of one. Three values of one option (`none`, `font`, `font-reserve`), so no caller can ask for two at once | reaches **4 of 39** readable corpus documents (10 of 32 `pdf` admit no outline run; 4 of those render differently at all). **Off by measurement, and the measurement is that it changes nothing measurable:** on the operator's twelve-position unit worksheet it alters **not one cell** — the five positions where it fires are one PDF whose glyphs carry no ToUnicode mapping and four office documents the PDF reader never touches. The position it was built for numbers its own chapters, so the reserve is silent there by construction |
| `--ocr` | read a PDF page as an image when its own text never arrived: the page extracts empty, or as `(cid:N)` placeholder codes at or above 10 % of its characters. Needs the optional `ocr` group | on the one corpus document with the failure: **95.07 % → 0 %** cid, **44 → 2561** words of four or more letters, 17 → **18** pages with text, 3.6 s/page. Whole corpus: **16 of 834** pages qualify, in **1 of 32** files |
```
pip install "llm-ingestion-okf[extract,ocr]"
```
Without that group `--ocr` is a typed refusal (`extractor_ocr_group_missing`)
per file, never a crash, and the corpus run still reports
`merged + coded rejections == N`. OCR text is a model's reading of an image: it
is reproducible against the model version and rendering resolution it was
produced with, and no dependency pin can promise more. The full measurement,
including the per-page distribution the 10 % threshold was read off, is
`docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md`.
Measured 2026-09-08 on a 43-file corpus (33 `pdf`, 5 `docx`, 2 `xlsx`, and
three files no reader accepts), one
`okf build` invocation replacing the shell loop over `tools/` that produced the
same corpus's bundle on 2026-09-03:
| figure | value |
|---|---|
| `N` (folder file count, computed) | 43 |
| merged | 39/43 |
| coded rejections | 4/43 (`extractor_unknown` 3, `extractor_empty_pdf` 1) |
| K1b | `39 + 4 = 43 = N`, exit `0` |
| files written | 1108 |
| identical to the 2026-09-03 bundle | 1104/1108 |
| wall time | 842.82 s total, 19.600 s per file (re-measured 2026-09-08) |
The six files that differ are all in the corpus's two spreadsheet documents, and
they are the change reported in `docs/2026-09-08-prisform-og-loggen-k2.md`: a
spreadsheet's tables are now written as pipe tables, so each row is one line
with its cells delimited rather than padded out to the widest cell in the
column. Two concept files are renamed by it, two are removed under their old
names, and the two documents' own `index.md` follow. The root `index.md` is
identical to the stored one again, because this library no longer links the
bundle's `log.md` from it. The command's own byte-identity test compares it
against the two scripts at the current commit, where the two agree over the
whole tree.
## Consume
The other direction: a bundle plus one question in, one bounded, contract-shaped
payload out.
```
python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json
```
`tools/okf_consume.py` is the **pre-pass** `docs/consumption-contract.md` § 1
defines — the deterministic program that reads the bundle, ranks its concepts,
cuts them to a bounded set and emits one payload. It decides nothing about the
question; the skill that reads the payload does the judgement. It calls no
model, opens no socket, imports nothing outside the standard library and this
package, and takes no clock: the same bundle bytes and the same
`(question, k, limit, cost_vocabulary, reserve_top_rank, rarity_weight)`
produce byte-identical output.
`--cost-vocabulary` is off by default and widens one question class: it lets a
declared list of cost/price/quantity terms bridge a question and a document that
name money with different words. The gate is the question — one naming no such
term gets byte-identical bytes either way — and what it does and does not close
is measured in `docs/2026-09-08-blindsone-below-k-k2.md`.
`--reserve-top-rank` is off by default and answers a different objection: the
budget is packed by an exact knapsack, which maximises a SUM of scores and
therefore has no opinion about rank, so a top-ranked excerpt costing a large
share of the budget is out-summed by many small ones. Measured, that made `--k`
a dial that could EVICT the concept a question was asked about. The flag gives
rank one its bytes before the pack runs — after the `over_budget_alone`
pre-exclusion, never before — and the payload then declares
`budget.reserved`. On a 629-concept corpus it changed the delivered set in 2 of
24 measured combinations, both of them that eviction:
`docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
`--source-quota N` is **on** by default at **2** since 2026-09-10 (opt out with
`--no-source-quota`), and it is the third widening here that alters a payload
with no bundle changing. It caps how many DELIVERED places one source document
may take, cutting where the shortlist is cut so the freed place goes to the next
candidate and `k` is still delivered in full. The defect it repairs was measured
outside this repository on a 3206-concept bundle of a published handbook: the
code's own process overview contributes **28 of 3206 concepts (0.87 %)** and
**8.0 % of the source characters**, and took **8 of 8** delivered places on one
question and **7 of 8** on the known-positive, which was not delivered at all.
Identical at 343 and 1651 concepts, so the cause is the corpus's COMPOSITION —
that it holds its own table of contents — and not its size; any corpus with a
contents list, a project overview or a summary document has the same property.
Swept over {2, 3, 4, off} on three bundles: at 2 and 3 hit@8 goes **5 of 6 to
6 of 6 on both K2 bundles** with all five standing rank-1 rows unmoved, and at 4
and off it stays 5 of 6. On the handbook bundle hit@8 goes **2 of 6 to 4 of 6**
and the dominant document's share of delivered places **8 of 8 to 2 of 8**. 2
rather than 3 on rank: the recovered rows come in at 5 and 4 rather than 7 and
5. What the gain is NOT: hit@8 asks whether the gold DOCUMENT was delivered, and
a document quota directly raises how many distinct documents a payload holds, so
that metric is not neutral with respect to this rule — the five rows that were
already rank 1 are, and they did not move. The adverse case is named rather than
found later: a bundle built from ONE document has one `source_file` on every
concept, so the quota would deliver 2 excerpts instead of `k`; the shortlist is
topped back up from the best-ranked over-quota candidates, which makes such a
bundle byte-identical to the quota being off. `--rarity-weight` was measured
against the same defect in the same session and does **not** repair it: on the
handbook bundle it leaves the dominant document at 8 of 8 places on the
question it floods and delivers neither that answer nor the known-positive.
`--rarity-weight` is off by default and weights each lexical hit by
`log(N/df)` over the bundle's own concepts instead of counting it as one, so a
requirement number is not worth what a common verb is worth. The default being
off is a measurement rather than a preference: on four corpora it took one gold
concept from withheld to delivered and a priced sheet from candidate rank 10 to
2, left one gold rank unmoved, and cost another seven rank positions — because
the four-character prefix matcher makes a unique identifier read as
135-of-446 common on that bundle. Where it cannot help is decomposed rather
than guessed: RRF fuses RANKS, so a weight moves nothing on a signal the gold
already leads. `docs/2026-09-08-sjeldenhetsvekt.md`. Its published figures were
measured under the pre-2026-09-10 tie-break and are not re-measured.
`--stem-prefix` is **on** by default since 2026-09-09 (opt out with
`--no-stem-prefix`), and like `--tie-shared-rank` below it alters a payload
with no bundle changing. `MIN_SHARED_PREFIX = 4` exists for Norwegian
compounding, and it also matches four characters that are not a stem: measured
on the pinned 453-concept bundle with the control run first, `under` occurs 79
times by equality and matches 172 concepts by prefix, while `bilateral` occurs
**0** times and matched **400 of 453** through `bilag`, and `standhaftig` 0 and
219 through `standard`. Three repairs were measured and all three failed on the
same row — a longer floor (58), a coverage share (0.50.8) and a
long-words-only floor (≥ 8) each cost row 1 its rank on the default bundle and
the whole row on Arm B, because row 1's token `prisene` reaches its gold
document through `pris|sammenstilling` on four characters. The rule that works
asks whether the shared prefix is a WORD the bundle uses: `bilateral` 400 → 0
and 512 → 0, `standhaftig` 219 → 56 and 235 → 33, **every hit@8 row keeping
rank 1 on both bundles**. `undersjøisk` stops at 162 because `under` is a word
here — a genuine Norwegian morpheme, so the residual is a different answer and
not a ceiling.
`--tie-shared-rank` is **on** by default since 2026-09-10 (opt out with
`--no-tie-shared-rank`), and it is the one change in this library that alters a
payload with no bundle changing — a consumer pinned to the previous excerpt
order needs the opt-out. RRF emits a rank for every concept in every signal,
including a signal that scored them all the same, and the declared tie-break
then orders that group by `concept_id`; the fusion reads alphabetical order as
if it were a measurement. Shared ranks make a signal that separates nothing
contribute the same constant to each concept in the group. What it buys is
general rather than cosmetic: a document the segmenter splits from 1 concept
into 12 fills that signal's whole top tie group with its own concepts, so the
one that leads the body signal takes position 11 instead of 1 and the document
loses fused rank 1 to a single-concept competitor leading nothing — **the
fusion was punishing fine-graining for being fine-grained**, which put the
segmentation side and the retrieval side in competition over one number.
It shipped OFF on 2026-09-08 because hit@8 fell 5 of 6 to 4 of 6, and that
figure is real and **conditional**: swept over 2 document-prior exponents x 3
bundles x 6 rows, the lost row is lost only at exponent 1.0. The exponent moved
to 0.5 on 2026-09-09 for an unrelated reason, correctly reported as moving no
hit@8 row, and nobody measured the pair — so a rule sat behind a published
number that had stopped being true in the same commit. A flag's "off by
measurement" is a measurement of a *configuration*, not a property of the flag.
`docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md`.
It emits the § 8 shape — `contract`, `bundle` (`bundle_id` plus a
`sha256-tree:` content identity), `budget` (unit, instrument, limit, spent and a
validated known-positive), `denominators`, `excerpts` and `withheld` — and every
withheld concept names the rule that dropped it, from a closed set of seven.
Every excerpt carries the concept's `title`, and — when the producer wrote them
`req_number`, the SPEC § 5.1 address `sources`, and **every top-level
`source_*` key**, by prefix rather than by allowlist: a fixed list names the
locators its author thought of, and one real bundle locates by
`source_element_id` on 269 of its 274 concepts. A key the producer did not write
stays absent rather than arriving empty, and an address this reader cannot
decode is named (`sources_unreadable`) rather than dropped into the same
silence. The reason is a measurement: with `concept_id` and body text alone, a
delivered gold concept at rank 1 still left the answer unable to name the
document it was quoting.
`considered == withheld + delivered` closes by construction, and the payload is
refused rather than reported when it does not.
Three exit codes, not two: **0** a payload was written, **1** the run happened
and refused (the budget admitted none of the concepts that answered the
question, or an asserted `--ref` contradicted the bytes), **2** the run did not
happen. Collapsing 2 into 1 would report an unread bundle as a failed cut.
`--ref` is an **assertion**, never an override — the identity is always computed
from the bytes, because labelling a payload with an identity its bytes do not
have is the one thing § 3.3 exists to prevent.
Check any payload against the skill that will read it:
```
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json
```
`skills/okf-consume/` is the first instantiated consumption skill: a filled copy
of `skills/okf-consume-template/` naming this pre-pass, with every per-corpus
hole replaced by a measured value. Measured 2026-09-07 on a 629-concept bundle,
hit@8 was **5 of 6** questions at rank 1 against a chance baseline of **1.35 of
6** — with one control that failed, and both are in
`docs/2026-09-07-okf-konsumskill-maaling.md` with the honesty limits stated.
## Consume in Claude Code
A folder of documents to an answer a model can cite, in **three lines**. You do
not need this repository — the first line installs the command, the second
builds the bundle and writes a skill beside it, the third asks.
```sh
uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.8.0"
okf project ~/my-documents
claude
```
`okf project` writes the bundle to `.okf/<id>/` and a skill to
`.claude/skills/<id>-consume/` in the **current directory**, then prints what it
read, what it wrote, and which documents a question cannot reach. Start `claude`
in that directory and ask in plain language; the generated skill runs the
pre-pass and the contract check itself and marks every claim with its source.
**Running non-interactively.** In print mode the skill needs its tools named
explicitly, or the model answers without ever reading the bundle and marks
every premise `undecidable-from-bundle`:
```sh
claude -p --allowedTools=Bash,Read,Grep,Glob "<the question>"
```
Add `Write,Edit` for the mode that produces a document. `--permission-mode
acceptEdits` alone does **not** do it — measured 2026-09-09 over four runs, the
`okf consume` call is refused without the explicit tool list. The interactive
`claude` above needs none of this.
`<id>` is the folder's name reduced to `[a-z0-9-]`. Run it once per folder with
`--id <name>` to have several bundles reachable at once — each skill carries its
own `bundle_id`, which is what lets a model pick between them. `--out <dir>`
puts the project somewhere other than the current directory.
Measured 2026-09-09 from a fresh `uv tool install` with this repository nowhere
on the path: 5 documents in, **26** concepts out, a skill carrying **0** paths
into any checkout, and `okf check` conformant on its own payload (15 rules, 0
findings). The 2026-09-08 run of the same measurement reported 15 concepts, and
that number was the defect rather than the result: `okf project` was calling
`build()` as a function and reading its signature's defaults, which disagreed
with argparse's on two flags. Two tests now hold the two default sets equal. Before that day the same result took a `PYTHONPATH`, a snapshot of a
clone, and a generated skill that named that clone by absolute path on four
lines — so it could not be moved, shared, or run by anyone else.
### The same thing in steps, if you want to see the payload
```sh
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
okf skill ./bundle --out ./project/.claude/skills/my-bundle-consume
okf consume ./bundle --question "your question" --out /tmp/payload.json
okf check --skill ./project/.claude/skills/my-bundle-consume/SKILL.md --payload /tmp/payload.json
```
A bundle you only have read access to is fine — the generator only reads it.
### The honest limits
Measured on **four questions** across two bundles, which is a demonstration and
not a hit rate. The ranking is lexical, and one of the four found a topic the
bundle **does** cover and did not rank it into the cut — the skill then said so
with its denominator instead of answering, which is the behaviour the contract
asks for, but a miss is still a miss.
`docs/2026-09-08-claude-code-skill-vilkaarlig-bundle.md` has the runs.
Two more things the summary tells you and this paragraph will not repeat: a
document that landed **whole** (no heading, table or numbered outline to cut it
on) comes back as one excerpt, which the budget often refuses and which often
does not carry the answer at the place you asked about; and a document that is
in the folder but **not** in the bundle cannot be quoted at all. Both cases are
answered `[sourced-not-sufficient]`, and `okf project` names the documents.
### The skill that runs this for you
`skills/okf-prosjekt/` in this repository is a Claude Code skill (Norwegian)
that wraps the command above: it takes a folder, runs `okf project`, and reads
the summary back. Install it for your user account after cloning:
```sh
mkdir -p ~/.claude/skills && cp -R skills/okf-prosjekt ~/.claude/skills/
```
## Implemented scope (v1)
The library provides three entry points for getting content into an OKF
bundle:
1. **Spec-based ingestion.** An implementation of the normative ingest
specification owned by `portfolio-optimiser-commons`: manifest →
`file`/`sql`/`http` connector → deterministic materialization of
`ingest-{id}.md` concept files → index generation. Zero model calls in the
run path; output is reproducible byte-for-byte against golden fixtures.
2. **Bundle inbox.** A drop directory where common file types are converted
to OKF concept files. All file-type→text extraction lives in this library:
`md`, `txt`, `csv`, `json`, `html` and `xml` are handled by the stdlib core;
`pdf` and the five office formats (`docx`, `xlsx`, `pptx`, `odt`, `rtf`)
require the optional `[extract]` extra and are rejected fail-fast without
it. Extracted text passes the security gate before anything is persisted.
An `xml` file that declares NISO-STS structure (`<standard>` root, or any
`<sec>`) becomes one heading per titled section, at the section's own
nesting depth, with the section's `<label>` and `<title>` on one line; a
`<sec>` carrying only a label is a body line and never a heading, and a
`<table-wrap>` becomes one markdown table. Any other XML keeps its text in
document order and gets no invented structure. XML carrying a
`<!DOCTYPE` is refused unparsed.
The drop directory is walked **recursively**, in sorted relative-path order:
a file at any depth is ingested and records its path relative to the inbox
root as its `source_file`, while dot-directories and a bundle directory
sitting inside the inbox are skipped with a reported code.
Under the segmented v0.2 profile a concept also points back at the document
it was extracted from, so an agent citing it can open the original at the
right place: `sources: [{ resource, title }]` in the spec's own §5.1 form,
where `resource` is the inbox-relative path, plus a locator per format —
`source_pages` for a PDF, `source_sheet` and `source_rows` for a
spreadsheet, `source_lines` otherwise. The locator keys are this library's
own, because §5.1 has no field for a place *within* a resource; the line
numbers index the extracted text and say so. Measurements:
[`docs/2026-09-08-proveniens-k2.md`](docs/2026-09-08-proveniens-k2.md).
<!-- extract-formats: .md, .txt, .csv, .json, .html, .htm, .xml, .pdf, .docx, .xlsx, .pptx, .odt, .rtf -->
3. **External bundle import.** Import and merge of third-party OKF bundles:
each concept is assessed via the security gate, and only concepts that
pass are merged, materialized, and linked into the index.
## Boundary: security is delegated
Security is owned by the sibling package
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
(pinned `>=1.2,<2.0`). The division is strict:
- **guard** answers "is this content safe to persist?" — scan, sanitize,
quarantine, fail-secure, provenance stamping.
- **this library** does the plumbing — connect a source, materialize a
deterministic OKF bundle, generate the index.
No security functionality is reimplemented here.
### What is gated today: read this before trusting a door
- **Door A (`materialize_bundle`) is ungated.** It calls nothing before
writing to disk and writes what it is given. A caller materializing
untrusted content is responsible for gating it.
- **Doors B and C (`process_inbox`, `import_bundle`) gate through an adapter
you pass in.** Each takes a `gate` argument; the flow hands it the content
and obeys the verdict, refusing to persist anything that does not clear the
guard's non-blocking floor — including a disposition it does not recognise,
and (at Door C) a concept the gate returned no verdict for. What it cannot
do is check that your adapter is a real guard: a permissive stub approves
everything, and the flow will believe it.
`llm_ingestion_okf.guard_adapter` is the adapter over the real guard, and the
only module here that imports it — importing the package itself does not:
```python
from llm_ingestion_okf import process_inbox
from llm_ingestion_okf.guard_adapter import inbox_gate
result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
okf_type="reference", gate=inbox_gate)
```
Two properties of that adapter are worth knowing before you rely on it.
It screens the **exact bytes it persists** — the guard's `prepare_input`
bookend prepares text for a model call, which this library never makes, so
only `screen_output` is used and the screened string is the written string.
And it **refuses rather than repairs**: a file carrying an invisible
zero-width or bidi character is rejected, not silently stripped and written.
Door B screens under the untrusted-upload policy, so any finding at all is
held back rather than persisted.
This section is stated plainly because earlier wording ("calls the guard at
every persist gate") described the intended end state in the present tense,
and a consumer reasonably read it as safe-by-default.
## Roadmap
The library is built in four phases so that every known OKF surface in the
ecosystem is eventually covered. Each phase has a detailed plan with
verification criteria:
1. Spec-based ingestion (Python) with byte-exact golden fixtures —
[plan](docs/plan/phase-1-door-a.md).
2. Bundle inbox and external-bundle import (Python), guard-gated —
[plan](docs/plan/phase-2-doors-b-c.md).
3. Configurable bundle contract (types, layers, frontmatter sets, index
shape, and reserved-file policy as configuration), enabling stricter
bundle profiles such as `strict-v1`
[plan](docs/plan/phase-3-configurable-contract.md).
4. A `node/` half: a zero-dependency Node/ESM package (importable and
CLI-invokable, vendored per consumer) providing bundle checking, index
generation, inbox processing, and document conversion for the OKF
second-brain plugin ecosystem. The Python and Node halves share the OKF
contract and fixture suite, not code —
[plan](docs/plan/phase-4-node-half.md).
## Upstream OKF versions
The library targets the current latest version of Google's OKF. Support is
**additive** — a new upstream version arrives as a new profile, never as a
migration of an existing one — so an *upstream* release does not change the
bytes an existing profile emits.
That guarantee is about upstream, and one profile tracks a second contract as
well. `DEFAULT` states the ingest-spec owned by `portfolio-optimiser-commons`,
so when they change that spec, `DEFAULT` follows them. It happened on
2026-08-09: `generated` moved from `true` to
`{ by: process:okf-ingest, at: <ingested_at> }`, one changed line per generated
file. Upgrading across it costs a re-run and nothing more — a profile still
recognises bundles stamped by earlier versions, so re-running writes in place
instead of refusing. `DEFAULT` remains OKF v0.1 on every axis upstream owns.
| Profile | Contract | Status |
|---|---|---|
| `DEFAULT` | commons' ingest-spec layer (OKF v0.1 semantics) | stable |
| `STRICT_V1` | a consumer's ratified v0.1 contract | stable |
| `OKF_V0_2` | OKF v0.2 | **provisional**, pre-release only |
| `STRUCTURED_V1` | `DEFAULT` plus a faceted, derived index | stable |
| `OKF_LATEST` | alias for the latest version supported as *stable* | currently `DEFAULT` |
`STRUCTURED_V1` is `DEFAULT` in every respect but the index. Under it, Door B
derives each dropped document's title, number, hierarchy and cross-references,
writes them into the concept's own frontmatter, and carries them into the index
entry — so a consumer can reason over the bundle rather than only look things
up in it. Every inferred field is named in a `derived` list, because an
unmarked heuristic is worse than no heuristic: the consumer cannot know when to
doubt it. A pointer to a document not dropped yet is rendered `N200?` rather
than omitted, since a bundle is built up over several drops and an absence that
leaves no trace is the dangerous kind. Carrying the metadata costs index
characters — roughly 3x to 6x the flat index, depending on how many facets the
profile names — and the facet key set is the dial. Design record and
measurements: [`docs/plan/structure-derivation.md`](docs/plan/structure-derivation.md).
`OKF_V0_2` ships first as a pre-release to a named pilot set and may change on
their feedback without a deprecation cycle. Pin the versioned constant rather
than `OKF_LATEST` unless you have explicitly opted into tracking; `OKF_LATEST`
moves at general availability, which is a deliberate release event rather than
a side effect of an upgrade.
Selecting a profile is keyword-only, so existing call sites are unaffected:
```python
materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)
```
A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and
puts the declaration in the bundle-root `index.md`'s frontmatter block. The
profile names the key; the **caller supplies the value**, because that value
tracks the upstream version and is not this library's to decide:
```python
materialize_bundle(
manifest, bundle_dir, ingested_at,
profile=OKF_V0_2,
root_frontmatter_values={"okf_version": "0.2"},
)
```
Omit the argument and no frontmatter block is written. Offering a key the
profile does not name is refused before anything is written to disk.
### Attested computations (v0.2 §10)
`OKF_V0_2` supports the `Attested Computation` type as a **format**: its five
contract fields — `runtime`, `parameters`, `computation`, `executor`,
`attester` — are emitted in canonical position, judged, and round-tripped.
`runtime` is required for that type and for no other, which the profile
expresses through `FrontmatterSchema.required_by_type`; a type the mapping does
not name carries no extra requirement, because §14 forbids a consumer to reject
on an unknown `type`.
Nothing here executes a computation or checks an attestation. Upstream defers
the receipt and verdict wire formats, so there is no contract to implement, and
the question an attestation answers — was this value produced the sanctioned
way — is not this library's. It re-enters scope when upstream specifies the
protocol.
On the import side, a third-party concept may name an `executor` or `attester`
resource pointing at executable code. Door C imports the **pointer** and never
the code — it writes concepts verbatim and skips every non-`.md` file — so such
a reference may not resolve, or may resolve to a file the destination tree
already holds under that path. Each one is reported in
`ImportResult.unverified_references`; the concept still merges, because §14
forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer
to surface rather than silently drop. The report names the pointer key, not the
resource it points at: recovering the resource needs the structured reader.
One limit worth knowing before you write such a concept: §10.2 presents
`executor` and `attester` as nested block mappings, and this library's
frontmatter parser is line-oriented. It reads inline **flow** mappings
(`executor: { resource: …, receipt: [ … ] }`) as opaque values that round-trip
unchanged, but it cannot read the block form — two block mappings that both
carry a `resource` collapse into one namespace and the first is lost. Write the
flow form; both are valid YAML, and a real YAML consumer recovers the same
structure from either.
## Non-goals
- Verdict/feedback machinery from the method specification (stays in the
consuming repositories).
- Embedding- or retrieval-layer functionality.
- Security functionality, in either runtime — that is always
`llm-ingestion-guard`'s domain.
## Requirements
Python 3.10+, and exactly one runtime dependency — the security boundary,
`llm-ingestion-guard>=1.2,<2.0`. Everything else is stdlib. The commands are
under [Install](#install); what follows is why they look the way they do.
From a checkout, the test suite runs with:
```
.venv/bin/python -m pytest
```
The suite is the verification surface for everything above: **1575 tests**,
run on 2026-09-10 against the `v0.8.0` release commit with the `[extract]`
extra installed.
(The figure stood at 596 until 2026-09-09 — measured 2026-08-21 and never
updated as the suite grew — and at 1515 until this release: a count is a
measurement with a date on it.)
Without the extra the same suite skips the tests covering the parser path;
that split was last counted on 2026-08-21 as 589 passed and 7 skipped and has
**not** been re-measured since. The tests holding the fail-fast rejection for
an uninstalled extra run in both. The suite is not shipped in an installed
distribution — `tests/` lives at the repository root, so this command needs a
clone rather than a `pip install`.
**The lint acceptance is BOTH of ruff's gates, and the rule set is declared.**
`ruff check src tests tools` *and* `ruff format --check .`, both named in a
report with the version they ran under. Neither was true before 2026-09-09:
the formatter gate was not in the acceptance and had gone red unseen, and
`[tool.ruff]` set only `line-length` and `target-version`, so the acceptance
was whatever ruff's default happened to be — which is why the tree read green
only as long as `uv.lock` froze ruff at 0.15.22. Under 0.16.6 the same
untouched code reported **148** findings, all of them new rules rather than new
defects, because 0.16 widened the default set to whole families. `select` is
now written down (`E4`, `E7`, `E9`, `F`, `I`, `RUF100`), the dev pin is
`ruff>=0.16.6,<0.17`, and the wider families are a separate decision with 148
as its starting number. `S` is measured out rather than assumed out: it reports
**2657** `S101` on a suite whose every assertion is an `assert`. Add
`--extra extract` to the sync or `mypy src` cannot find `pdfplumber`.
**Markdown is excluded from `ruff format`.** ruff 0.16 formats fenced Python
inside markdown, and the two files it would change here are records rather than
source — a README call example and a published measurement's *quotation* of
`COST_VOCABULARY` as it stood when that measurement was taken. Reformatting a
quotation makes it stop being one.
A git URL is a PEP 508 direct reference and pins one exact tag, so it is an
install-time *channel*, not the pin: the range above stays the declared
dependency — a wheel built from this branch carries `Requires-Dist:
llm-ingestion-guard<2.0,>=1.2`, measured 2026-08-23 — and resolves normally
once the package index exists. A wheel built from a *tag* carries that tag's
range instead, which is why the install commands pair tag with tag.
### Binary extraction
The optional `[extract]` extra ships two things: `pdfplumber` (MIT) for `pdf`,
and `pypandoc-binary` for five office formats. It is opt-in because it pulls
binary wheels, which the default install must never do — the single runtime
dependency rule covers the default install and this extra sits outside it.
The converter **binary travels inside the wheel** and is resolved by path
rather than found on `PATH`, with its version asserted against a pin. A host
carrying a different converter is refused, not silently used: extraction is
deterministic within a converter version and not across one.
| Format | Reader | Evidence |
|---|---|---|
| `pdf` | `pdfplumber` | measured |
| `docx` | converter | measured |
| `xlsx` | converter | measured |
| `pptx` | converter | **constructed** |
| `odt` | converter | **constructed** |
| `rtf` | converter | **constructed** |
**`constructed` means what it says, and it is a weaker word than `measured`
on purpose.** The corpus this work was measured on contains **zero** `pptx`,
`odt` and `rtf` files. Until 2026-09-09 those three rows were `unmeasured`
they worked by construction and had never been checked against a document
anyone wrote. They have now each been put through end to end on a hand-built
document with a hand-written fasit, which is more than nothing and is not a
corpus:
- `odt`**1 of 1** declared headings recovered, 1 concept, 0 characters in
no segment. N = 1 document.
- `pptx`**2 of 2** declared slide titles recovered on a deck that declares
them (a real `<p:ph type="title"/>` placeholder); **0 of 2** on a deck that
does not, where the converter writes `Slide 1` / `Slide 2` because it has no
title to use. That is the converter naming an unnamed slide, not a
segmentation failure. N = 2 decks.
- `rtf`**0** declared headings, because the container has no heading style
and the author's title is bold text. The proposer therefore proposes
nothing, and the document reaches the bundle inbox as one whole concept:
content preserved, structure zero. N = 1 document. This is the one open
finding of the three.
They are not known to be broken; a single constructed document is not a
denominator, and the distinction is the point.
**What stays out.** `.doc` (Word 97) is not supported — the converter does not
read it. Rastered or scanned PDFs are refused rather than persisted as empty
concepts, because this library does not do OCR. Drawn content — figures,
diagrams, shapes — does not survive extraction in any format here, and every
extraction says so with a warning. Structured table recovery is out of scope.
Request it by appending `[extract]` to the package name in whichever install
command from [Install](#install) you are using — this package is not on an
index, so a bare `pip install 'llm-ingestion-okf[extract]'` does **not** work
today, and the error message naming that command is written for the day it
does. The extra is unreleased: it reaches a consumer through a tag that
contains it, and no such tag exists yet.
Two properties of the extra are worth knowing before depending on its output:
- **Extracted text is pinned to an exact parser version.** `pdfplumber` pins
`pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
releases with no stability contract. Extraction is deterministic within a
parser version and not guaranteed across one, so a golden fixture built on
extracted PDF text is a fixture migration away from any parser upgrade.
- **Text extraction recovers text, and nothing that is drawn.** Figures,
diagrams and images have no text to recover — only their captions survive —
so a bundle built from drawn documents is incomplete by construction. The
library says so itself: every `pdf` extraction emits an `ExtractionWarning`.
Structured table recovery is separately out of scope; PDFs enter as prose.
The planned Node half targets Node/ESM with zero npm dependencies.
## License
MIT — see [LICENSE](LICENSE).