A PDF carries no notion of a heading -- a heading in a PDF is a typographic
fact -- so the text stream `pdfplumber` hands the segment proposer has already
thrown away the only evidence there was. The `docx` path never had that problem:
the converter emits ATX headings and `_ATX` cuts on them. Two readers close the
gap, and both are OFF.
`--pdf-headings font` infers a heading from the conjunction this repository
already measured (size above the document's character-weighted body median AND
a bold font name, recall 1.000 / precision 0.846) and emits it as ATX in the
SAME markdown the office path produces, so `_ATX` applies unchanged and no
PDF-only heading grammar exists.
It stays off BY MEASUREMENT, and the measurement is the point of the round:
against the operator's unit worksheet it takes `pdf` from 2 of 8 to 0 of 8,
losing two exact matches. The mechanism of the loss is stated rather than
guessed -- on those documents the outline rule already recovers the document's
own numbered chapters, so a second heading source can only add. Whole-corpus
screen: 25 of 32 `pdf` change, 0 of 5 `docx`, 0 of 2 `xlsx`. The default bundle
is byte-identical before and after this commit (`diff -r`, exit 0).
`--ocr` reads a page as an image when its own text never arrived: empty, or
`(cid:N)` placeholder codes at or above a threshold READ OFF a measured
distribution -- 834 pages over 32 files, 818 at exactly 0.0 and 16 at 0.93 or
above, nothing in between. On the one corpus document with the failure: 95.07 %
cid to 0 %, 44 to 2561 words of four or more letters, 17 to 18 pages with text.
Its engine is an optional dependency group and never a runtime dependency; a
packaging test pins both halves, and without the group every affected file is a
coded rejection (`extractor_ocr_group_missing`) rather than a crash.
Also corrects two stale published facts found while measuring: the README still
said two segmentation rules were on by default after `f6fea13` made it three,
and CLAUDE.md's K2 digest named the round-3 default. The current default is
492 concepts / 944 files, `bdefa679...`.
Report: docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md
Co-Authored-By: Claude <claude-opus-5>
631 lines
35 KiB
Markdown
631 lines
35 KiB
Markdown
# llm-ingestion-okf
|
||
|
||
Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import. Security delegated to llm-ingestion-guard.
|
||
|
||
Status: phases 1–3 are implemented. Phase 1 (spec-based ingestion) covers
|
||
manifest validation, the `file`/`sql`/`http` connectors, deterministic
|
||
materialization, index generation, and the golden fixture suite under
|
||
`examples/`. Phase 2 adds the bundle inbox (`process_inbox`) and
|
||
external-bundle import (`import_bundle`), both against an **injected** persist
|
||
gate, with `llm_ingestion_okf.guard_adapter` wiring that gate to the real
|
||
guard (see below). Phase 3 makes the bundle contract configurable, so types,
|
||
layers, frontmatter sets, index shape, and reserved-file policy are carried by
|
||
a profile rather than by constants (see [Upstream OKF
|
||
versions](#upstream-okf-versions)). Binary extraction runs behind the
|
||
optional `[extract]` extra: `pdf` through a PDF parser, and five office
|
||
formats through a vendored document converter. Three of those five office
|
||
rows are **unmeasured** — see [Binary extraction](#binary-extraction). Phase 4
|
||
(the Node half) is planned (see `docs/plan/`).
|
||
|
||
## Install
|
||
|
||
Python 3.10+. Neither this package nor the guard it depends on is on a package
|
||
index yet. With uv, one command is enough:
|
||
|
||
```
|
||
uv pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
|
||
```
|
||
|
||
uv resolves the guard on its own, because it reads the `[tool.uv.sources]`
|
||
entry in the `pyproject.toml` **of the tag it is installing**, and `v0.4.0`
|
||
points that entry at the guard tag below. Measured 2026-07-25 and re-measured
|
||
2026-08-20 with an empty `uv` cache; both runs installed
|
||
`llm-ingestion-guard==0.2.0` + `llm-ingestion-okf==0.4.0` and imported clean.
|
||
|
||
With plain pip, the transitive git dependency does not resolve on its own —
|
||
**install the guard first**, or installing this package fails with
|
||
`No matching distribution found for llm-ingestion-guard`:
|
||
|
||
```
|
||
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v0.2.0"
|
||
pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
|
||
```
|
||
|
||
The guard tag is paired to the okf tag, not to this branch: `v0.4.0` declares
|
||
`llm-ingestion-guard>=0.2,<0.3`, which `v0.2.0` satisfies and later guard tags
|
||
do not. `main` has since moved its own pin to `>=1.2,<2.0` (see
|
||
[Requirements](#requirements)); that pin reaches you in the next stable tag,
|
||
not in the commands above. Reading a pin off this branch and installing it
|
||
against `v0.4.0` is the one combination that fails.
|
||
|
||
`v0.6.0` is the current tag and the one the three-line form under [Consume in
|
||
Claude Code](#consume-in-claude-code) installs: it is the first tag carrying the
|
||
`okf project`, `okf consume`, `okf check` and `okf skill` subcommands, without
|
||
which that form does not exist. `v0.4.0` is the last tag before the OKF v0.2
|
||
work. `v0.5.0a2` is a pre-release for the named OKF v0.2 pilot set only.
|
||
|
||
**`v0.6.0` does not make OKF v0.2 generally available.** `OKF_LATEST` is
|
||
unchanged and still points at `DEFAULT`; flipping that alias is the GA event and
|
||
this tag is not it (see [Upstream OKF versions](#upstream-okf-versions)).
|
||
|
||
## Build
|
||
|
||
Installing the package installs one command. A folder of documents in, an OKF
|
||
bundle out:
|
||
|
||
```
|
||
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
|
||
```
|
||
|
||
It walks the folder recursively, proposes a segmentation for each document with
|
||
the mechanical rules, replays those proposals through the bundle inbox, writes
|
||
the bundle and its `log.md`, and prints the run's numbers. Every proposal is
|
||
marked `PROPOSED` and `adjudicated: false` — the command segments nothing a
|
||
human has approved, and says so in the artifact.
|
||
|
||
The last line that matters is the conservation identity: `merged + coded
|
||
rejections == N`, where `N` is the folder's file count read at run time. **The
|
||
run exits non-zero when it does not hold**, and names the unaccounted files, so
|
||
a pipeline cannot mistake a partial bundle for a complete one.
|
||
|
||
Flags worth knowing: `--segments off` ingests each document as one concept and
|
||
asks for no root values; `--plans-dir` keeps the proposals instead of
|
||
discarding them; `--report` writes the full report to a file as well as stdout.
|
||
`--ingested-at` and `--proposed-at` default to `1970-01-01T00:00:00Z` rather
|
||
than the clock, so two builds of the same folder are byte-identical — a
|
||
wall-clock default would break rebuild-equals-incremental for every caller who
|
||
did not pass them.
|
||
|
||
### The segmentation flags
|
||
|
||
Six rules are reachable from `okf build`. **Three of them are ON by default
|
||
since 2026-09-08** — `--outline-run 3`, `--table-grid` and `--unit-fold`, an
|
||
operator decision taken in two steps — and each has an explicit opt-out,
|
||
`--outline-run 0`, `--no-table-grid` and `--no-unit-fold`. Passing all three
|
||
opt-outs reproduces the pre-2026-09-08 bytes exactly. The other three are off.
|
||
Each line below carries the number it was measured at, and nothing beyond it.
|
||
|
||
**A re-run is what this costs a consumer, and it is not a small one:** on the
|
||
43-document reference corpus the default bundle goes from **629 concepts in
|
||
1108 files** to **517 in 969**. The proposer's own defaults
|
||
(`tools/okf_propose_segments.py`) did NOT move, so every published reproduction
|
||
block still runs as written.
|
||
|
||
| flag | what it does | measured |
|
||
|---|---|---|
|
||
| `--outline-run N` (default **3**) | also propose a boundary where the document's own bare-integer numbering sustains an ascending run of at least `N`; `0` is this arm's opt-out | a tender PDF whose headings are bare integers: **no boundary** at `0`, **9 concepts** at `3`, against a reference of 9 |
|
||
| `--table-grid` (**on** by default; opt out with `--no-table-grid`) | a pandoc grid-table rule line no longer closes an open table block, so one grid table is one concept | a `.docx` experience list: **21 → 6** concepts |
|
||
| `--unit-fold` (**on** by default; opt out with `--no-unit-fold`) | discard a contents-list run, fold a deeper heading into its parent, fold a table into the shorter heading that introduces it. Adds no boundary, so it can only reduce a plan | on a 12-document sample scored against an operator's unit worksheet: **5 of 12** match — but that figure was measured with `--table-grid` ON, and the shipped default does not include it. Measured without it the same sample scores **2 of 12**, `docx` **0 of 3**, because the fold's table clause has no joined table to fold |
|
||
| `--keep-table-heading` | keep a heading whose body is empty only because a table opens under it, and absorb that table into its span | the two spreadsheets in that corpus, and **0 of 32 `pdf` and 0 of 5 `docx`**: the concept count does not move (1 → 1), its first byte does — the concept gains the heading line it was missing |
|
||
| `--sheet-section-rows` | cut an open table block at the rows that label its sections — a run of at least three rows whose first cell is a bare numeric label. The opposite direction from `--table-grid`, which decides how far a block extends | a tender price sheet whose whole body is one table block: **1 → 12 concepts**, against a reference of 11 cost groups plus the sheet's preamble. Whole corpus: **1 of 39** readable documents changes, **0 of 32 `pdf`, 0 of 5 `docx`, 1 of 2 `xlsx`** |
|
||
| `--drop-wrapped-outline` | do not admit an `--outline-run` candidate whose line continues onto the next one. Judges recovered candidates only, never a heading the document declares | quoted regulation text, whose numbered paragraphs match the outline grammar exactly: **4 → 1 concepts**, the reference. Whole corpus: **5 of 39**, all `pdf`; on the 12-document sample **8 of 34** outline candidates wrap, and none of the 26 the operator kept |
|
||
|
||
They compose, and the order above is the order they apply in. Measured on a
|
||
five-document tender folder (2 `pdf`, 2 `docx`, 1 `xlsx`), concepts per
|
||
document:
|
||
|
||
| document | default | `--outline-run 3` | `+ --table-grid` | `+ --unit-fold` | `+ --keep-table-heading` | `+ --sheet-section-rows --drop-wrapped-outline` |
|
||
|---|---|---|---|---|---|---|
|
||
| tender PDF, technical requirements | 1 (no boundary) | 9 | 9 | 9 | 9 | 9 |
|
||
| tender PDF, technical layout | 1 (no boundary) | 1 | 1 | 1 | 1 | 1 |
|
||
| price sheet `.xlsx` | 1 | 1 | 1 | 1 | 1 | **12** |
|
||
| experience list `.docx` | 21 | 21 | 6 | 3 | 3 | 3 |
|
||
| agreement `.docx` | 2 | 2 | 1 | 1 | 1 | 1 |
|
||
| markdown files in the bundle | 31 | 49 | 33 | 30 | 30 | 52 |
|
||
|
||
Every column merged 5 of 5 with 0 rejections. The reference for the first row
|
||
is 9, so the default is a full arm behind what the proposer can do on that
|
||
document — which is a statement about the default, not a licence to change it
|
||
here.
|
||
|
||
### The two PDF reader flags
|
||
|
||
Separate from the six above, and they sit before every one of them: a
|
||
segmentation flag changes how the proposer cuts a text, these change what the
|
||
text says. **Both are off by default.**
|
||
|
||
| flag | what it does | measured |
|
||
|---|---|---|
|
||
| `--pdf-headings font` | a PDF carries no heading markup, so one is inferred from typography — a line whose dominant font size is above the document's character-weighted median AND whose dominant font name says bold — and emitted as an ATX heading in the same markdown the office path produces, so the existing heading rule reads it | on a tender PDF: **9 of 9** numbered chapters found, plus 4 extra candidates. Whole corpus: **25 of 32 `pdf`** change, **0 of 5 `docx`**, **0 of 2 `xlsx`**. **Off by measurement:** against the operator's unit worksheet it takes `pdf` from **2 of 8 to 0 of 8**, losing two exact matches, because on those documents the outline rule already found the chapters and a second heading source can only add |
|
||
| `--ocr` | read a PDF page as an image when its own text never arrived: the page extracts empty, or as `(cid:N)` placeholder codes at or above 10 % of its characters. Needs the optional `ocr` group | on the one corpus document with the failure: **95.07 % → 0 %** cid, **44 → 2561** words of four or more letters, 17 → **18** pages with text, 3.6 s/page. Whole corpus: **16 of 834** pages qualify, in **1 of 32** files |
|
||
|
||
```
|
||
pip install "llm-ingestion-okf[extract,ocr]"
|
||
```
|
||
|
||
Without that group `--ocr` is a typed refusal (`extractor_ocr_group_missing`)
|
||
per file, never a crash, and the corpus run still reports
|
||
`merged + coded rejections == N`. OCR text is a model's reading of an image: it
|
||
is reproducible against the model version and rendering resolution it was
|
||
produced with, and no dependency pin can promise more. The full measurement,
|
||
including the per-page distribution the 10 % threshold was read off, is
|
||
`docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md`.
|
||
|
||
Measured 2026-09-08 on a 43-file corpus (33 `pdf`, 5 `docx`, 2 `xlsx`, and
|
||
three files no reader accepts), one
|
||
`okf build` invocation replacing the shell loop over `tools/` that produced the
|
||
same corpus's bundle on 2026-09-03:
|
||
|
||
| figure | value |
|
||
|---|---|
|
||
| `N` (folder file count, computed) | 43 |
|
||
| merged | 39/43 |
|
||
| coded rejections | 4/43 (`extractor_unknown` 3, `extractor_empty_pdf` 1) |
|
||
| K1b | `39 + 4 = 43 = N`, exit `0` |
|
||
| files written | 1108 |
|
||
| identical to the 2026-09-03 bundle | 1104/1108 |
|
||
| wall time | 842.82 s total, 19.600 s per file (re-measured 2026-09-08) |
|
||
|
||
The six files that differ are all in the corpus's two spreadsheet documents, and
|
||
they are the change reported in `docs/2026-09-08-prisform-og-loggen-k2.md`: a
|
||
spreadsheet's tables are now written as pipe tables, so each row is one line
|
||
with its cells delimited rather than padded out to the widest cell in the
|
||
column. Two concept files are renamed by it, two are removed under their old
|
||
names, and the two documents' own `index.md` follow. The root `index.md` is
|
||
identical to the stored one again, because this library no longer links the
|
||
bundle's `log.md` from it. The command's own byte-identity test compares it
|
||
against the two scripts at the current commit, where the two agree over the
|
||
whole tree.
|
||
|
||
## Consume
|
||
|
||
The other direction: a bundle plus one question in, one bounded, contract-shaped
|
||
payload out.
|
||
|
||
```
|
||
python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json
|
||
```
|
||
|
||
`tools/okf_consume.py` is the **pre-pass** `docs/consumption-contract.md` § 1
|
||
defines — the deterministic program that reads the bundle, ranks its concepts,
|
||
cuts them to a bounded set and emits one payload. It decides nothing about the
|
||
question; the skill that reads the payload does the judgement. It calls no
|
||
model, opens no socket, imports nothing outside the standard library and this
|
||
package, and takes no clock: the same bundle bytes and the same
|
||
`(question, k, limit, cost_vocabulary, reserve_top_rank, rarity_weight)`
|
||
produce byte-identical output.
|
||
|
||
`--cost-vocabulary` is off by default and widens one question class: it lets a
|
||
declared list of cost/price/quantity terms bridge a question and a document that
|
||
name money with different words. The gate is the question — one naming no such
|
||
term gets byte-identical bytes either way — and what it does and does not close
|
||
is measured in `docs/2026-09-08-blindsone-below-k-k2.md`.
|
||
|
||
`--reserve-top-rank` is off by default and answers a different objection: the
|
||
budget is packed by an exact knapsack, which maximises a SUM of scores and
|
||
therefore has no opinion about rank, so a top-ranked excerpt costing a large
|
||
share of the budget is out-summed by many small ones. Measured, that made `--k`
|
||
a dial that could EVICT the concept a question was asked about. The flag gives
|
||
rank one its bytes before the pack runs — after the `over_budget_alone`
|
||
pre-exclusion, never before — and the payload then declares
|
||
`budget.reserved`. On a 629-concept corpus it changed the delivered set in 2 of
|
||
24 measured combinations, both of them that eviction:
|
||
`docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
|
||
|
||
`--rarity-weight` is off by default and weights each lexical hit by
|
||
`log(N/df)` over the bundle's own concepts instead of counting it as one, so a
|
||
requirement number is not worth what a common verb is worth. The default being
|
||
off is a measurement rather than a preference: on four corpora it took one gold
|
||
concept from withheld to delivered and a priced sheet from candidate rank 10 to
|
||
2, left one gold rank unmoved, and cost another seven rank positions — because
|
||
the four-character prefix matcher makes a unique identifier read as
|
||
135-of-446 common on that bundle. Where it cannot help is decomposed rather
|
||
than guessed: RRF fuses RANKS, so a weight moves nothing on a signal the gold
|
||
already leads. `docs/2026-09-08-sjeldenhetsvekt.md`.
|
||
|
||
It emits the § 8 shape — `contract`, `bundle` (`bundle_id` plus a
|
||
`sha256-tree:` content identity), `budget` (unit, instrument, limit, spent and a
|
||
validated known-positive), `denominators`, `excerpts` and `withheld` — and every
|
||
withheld concept names the rule that dropped it, from a closed set of six.
|
||
|
||
Every excerpt carries the concept's `title`, and — when the producer wrote them
|
||
— `req_number`, the SPEC § 5.1 address `sources`, and **every top-level
|
||
`source_*` key**, by prefix rather than by allowlist: a fixed list names the
|
||
locators its author thought of, and one real bundle locates by
|
||
`source_element_id` on 269 of its 274 concepts. A key the producer did not write
|
||
stays absent rather than arriving empty, and an address this reader cannot
|
||
decode is named (`sources_unreadable`) rather than dropped into the same
|
||
silence. The reason is a measurement: with `concept_id` and body text alone, a
|
||
delivered gold concept at rank 1 still left the answer unable to name the
|
||
document it was quoting.
|
||
`considered == withheld + delivered` closes by construction, and the payload is
|
||
refused rather than reported when it does not.
|
||
|
||
Three exit codes, not two: **0** a payload was written, **1** the run happened
|
||
and refused (the budget admitted none of the concepts that answered the
|
||
question, or an asserted `--ref` contradicted the bytes), **2** the run did not
|
||
happen. Collapsing 2 into 1 would report an unread bundle as a failed cut.
|
||
`--ref` is an **assertion**, never an override — the identity is always computed
|
||
from the bytes, because labelling a payload with an identity its bytes do not
|
||
have is the one thing § 3.3 exists to prevent.
|
||
|
||
Check any payload against the skill that will read it:
|
||
|
||
```
|
||
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json
|
||
```
|
||
|
||
`skills/okf-consume/` is the first instantiated consumption skill: a filled copy
|
||
of `skills/okf-consume-template/` naming this pre-pass, with every per-corpus
|
||
hole replaced by a measured value. Measured 2026-09-07 on a 629-concept bundle,
|
||
hit@8 was **5 of 6** questions at rank 1 against a chance baseline of **1.35 of
|
||
6** — with one control that failed, and both are in
|
||
`docs/2026-09-07-okf-konsumskill-maaling.md` with the honesty limits stated.
|
||
|
||
## Consume in Claude Code
|
||
|
||
A folder of documents to an answer a model can cite, in **three lines**. You do
|
||
not need this repository — the first line installs the command, the second
|
||
builds the bundle and writes a skill beside it, the third asks.
|
||
|
||
```sh
|
||
uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.6.0"
|
||
okf project ~/my-documents
|
||
claude
|
||
```
|
||
|
||
`okf project` writes the bundle to `.okf/<id>/` and a skill to
|
||
`.claude/skills/<id>-consume/` in the **current directory**, then prints what it
|
||
read, what it wrote, and which documents a question cannot reach. Start `claude`
|
||
in that directory and ask in plain language; the generated skill runs the
|
||
pre-pass and the contract check itself and marks every claim with its source.
|
||
|
||
`<id>` is the folder's name reduced to `[a-z0-9-]`. Run it once per folder with
|
||
`--id <name>` to have several bundles reachable at once — each skill carries its
|
||
own `bundle_id`, which is what lets a model pick between them. `--out <dir>`
|
||
puts the project somewhere other than the current directory.
|
||
|
||
Measured 2026-09-08 from a fresh `uv tool install` with this repository nowhere
|
||
on the path: 5 documents in, 15 concepts out, a skill carrying **0** paths into
|
||
any checkout, and `okf check` conformant on its own payload (15 rules, 0
|
||
findings). Before that day the same result took a `PYTHONPATH`, a snapshot of a
|
||
clone, and a generated skill that named that clone by absolute path on four
|
||
lines — so it could not be moved, shared, or run by anyone else.
|
||
|
||
### The same thing in steps, if you want to see the payload
|
||
|
||
```sh
|
||
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
|
||
okf skill ./bundle --out ./project/.claude/skills/my-bundle-consume
|
||
okf consume ./bundle --question "your question" --out /tmp/payload.json
|
||
okf check --skill ./project/.claude/skills/my-bundle-consume/SKILL.md --payload /tmp/payload.json
|
||
```
|
||
|
||
A bundle you only have read access to is fine — the generator only reads it.
|
||
|
||
### The honest limits
|
||
|
||
Measured on **four questions** across two bundles, which is a demonstration and
|
||
not a hit rate. The ranking is lexical, and one of the four found a topic the
|
||
bundle **does** cover and did not rank it into the cut — the skill then said so
|
||
with its denominator instead of answering, which is the behaviour the contract
|
||
asks for, but a miss is still a miss.
|
||
`docs/2026-09-08-claude-code-skill-vilkaarlig-bundle.md` has the runs.
|
||
|
||
Two more things the summary tells you and this paragraph will not repeat: a
|
||
document that landed **whole** (no heading, table or numbered outline to cut it
|
||
on) comes back as one excerpt, which the budget often refuses and which often
|
||
does not carry the answer at the place you asked about; and a document that is
|
||
in the folder but **not** in the bundle cannot be quoted at all. Both cases are
|
||
answered `[sourced-not-sufficient]`, and `okf project` names the documents.
|
||
|
||
### The skill that runs this for you
|
||
|
||
`skills/okf-prosjekt/` in this repository is a Claude Code skill (Norwegian)
|
||
that wraps the command above: it takes a folder, runs `okf project`, and reads
|
||
the summary back. Install it for your user account after cloning:
|
||
|
||
```sh
|
||
mkdir -p ~/.claude/skills && cp -R skills/okf-prosjekt ~/.claude/skills/
|
||
```
|
||
|
||
## Implemented scope (v1)
|
||
|
||
The library provides three entry points for getting content into an OKF
|
||
bundle:
|
||
|
||
1. **Spec-based ingestion.** An implementation of the normative ingest
|
||
specification owned by `portfolio-optimiser-commons`: manifest →
|
||
`file`/`sql`/`http` connector → deterministic materialization of
|
||
`ingest-{id}.md` concept files → index generation. Zero model calls in the
|
||
run path; output is reproducible byte-for-byte against golden fixtures.
|
||
2. **Bundle inbox.** A drop directory where common file types are converted
|
||
to OKF concept files. All file-type→text extraction lives in this library:
|
||
`md`, `txt`, `csv`, `json`, and `html` are handled by the stdlib core;
|
||
`pdf` and the five office formats (`docx`, `xlsx`, `pptx`, `odt`, `rtf`)
|
||
require the optional `[extract]` extra and are rejected fail-fast without
|
||
it. Extracted text passes the security gate before anything is persisted.
|
||
The drop directory is walked **recursively**, in sorted relative-path order:
|
||
a file at any depth is ingested and records its path relative to the inbox
|
||
root as its `source_file`, while dot-directories and a bundle directory
|
||
sitting inside the inbox are skipped with a reported code.
|
||
|
||
Under the segmented v0.2 profile a concept also points back at the document
|
||
it was extracted from, so an agent citing it can open the original at the
|
||
right place: `sources: [{ resource, title }]` in the spec's own §5.1 form,
|
||
where `resource` is the inbox-relative path, plus a locator per format —
|
||
`source_pages` for a PDF, `source_sheet` and `source_rows` for a
|
||
spreadsheet, `source_lines` otherwise. The locator keys are this library's
|
||
own, because §5.1 has no field for a place *within* a resource; the line
|
||
numbers index the extracted text and say so. Measurements:
|
||
[`docs/2026-09-08-proveniens-k2.md`](docs/2026-09-08-proveniens-k2.md).
|
||
|
||
<!-- extract-formats: .md, .txt, .csv, .json, .html, .htm, .pdf, .docx, .xlsx, .pptx, .odt, .rtf -->
|
||
3. **External bundle import.** Import and merge of third-party OKF bundles:
|
||
each concept is assessed via the security gate, and only concepts that
|
||
pass are merged, materialized, and linked into the index.
|
||
|
||
## Boundary: security is delegated
|
||
|
||
Security is owned by the sibling package
|
||
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
|
||
(pinned `>=1.2,<2.0`). The division is strict:
|
||
|
||
- **guard** answers "is this content safe to persist?" — scan, sanitize,
|
||
quarantine, fail-secure, provenance stamping.
|
||
- **this library** does the plumbing — connect a source, materialize a
|
||
deterministic OKF bundle, generate the index.
|
||
|
||
No security functionality is reimplemented here.
|
||
|
||
### What is gated today: read this before trusting a door
|
||
|
||
- **Door A (`materialize_bundle`) is ungated.** It calls nothing before
|
||
writing to disk and writes what it is given. A caller materializing
|
||
untrusted content is responsible for gating it.
|
||
- **Doors B and C (`process_inbox`, `import_bundle`) gate through an adapter
|
||
you pass in.** Each takes a `gate` argument; the flow hands it the content
|
||
and obeys the verdict, refusing to persist anything that does not clear the
|
||
guard's non-blocking floor — including a disposition it does not recognise,
|
||
and (at Door C) a concept the gate returned no verdict for. What it cannot
|
||
do is check that your adapter is a real guard: a permissive stub approves
|
||
everything, and the flow will believe it.
|
||
|
||
`llm_ingestion_okf.guard_adapter` is the adapter over the real guard, and the
|
||
only module here that imports it — importing the package itself does not:
|
||
|
||
```python
|
||
from llm_ingestion_okf import process_inbox
|
||
from llm_ingestion_okf.guard_adapter import inbox_gate
|
||
|
||
result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
|
||
okf_type="reference", gate=inbox_gate)
|
||
```
|
||
|
||
Two properties of that adapter are worth knowing before you rely on it.
|
||
It screens the **exact bytes it persists** — the guard's `prepare_input`
|
||
bookend prepares text for a model call, which this library never makes, so
|
||
only `screen_output` is used and the screened string is the written string.
|
||
And it **refuses rather than repairs**: a file carrying an invisible
|
||
zero-width or bidi character is rejected, not silently stripped and written.
|
||
Door B screens under the untrusted-upload policy, so any finding at all is
|
||
held back rather than persisted.
|
||
|
||
This section is stated plainly because earlier wording ("calls the guard at
|
||
every persist gate") described the intended end state in the present tense,
|
||
and a consumer reasonably read it as safe-by-default.
|
||
|
||
## Roadmap
|
||
|
||
The library is built in four phases so that every known OKF surface in the
|
||
ecosystem is eventually covered. Each phase has a detailed plan with
|
||
verification criteria:
|
||
|
||
1. Spec-based ingestion (Python) with byte-exact golden fixtures —
|
||
[plan](docs/plan/phase-1-door-a.md).
|
||
2. Bundle inbox and external-bundle import (Python), guard-gated —
|
||
[plan](docs/plan/phase-2-doors-b-c.md).
|
||
3. Configurable bundle contract (types, layers, frontmatter sets, index
|
||
shape, and reserved-file policy as configuration), enabling stricter
|
||
bundle profiles such as `strict-v1` —
|
||
[plan](docs/plan/phase-3-configurable-contract.md).
|
||
4. A `node/` half: a zero-dependency Node/ESM package (importable and
|
||
CLI-invokable, vendored per consumer) providing bundle checking, index
|
||
generation, inbox processing, and document conversion for the OKF
|
||
second-brain plugin ecosystem. The Python and Node halves share the OKF
|
||
contract and fixture suite, not code —
|
||
[plan](docs/plan/phase-4-node-half.md).
|
||
|
||
## Upstream OKF versions
|
||
|
||
The library targets the current latest version of Google's OKF. Support is
|
||
**additive** — a new upstream version arrives as a new profile, never as a
|
||
migration of an existing one — so an *upstream* release does not change the
|
||
bytes an existing profile emits.
|
||
|
||
That guarantee is about upstream, and one profile tracks a second contract as
|
||
well. `DEFAULT` states the ingest-spec owned by `portfolio-optimiser-commons`,
|
||
so when they change that spec, `DEFAULT` follows them. It happened on
|
||
2026-08-09: `generated` moved from `true` to
|
||
`{ by: process:okf-ingest, at: <ingested_at> }`, one changed line per generated
|
||
file. Upgrading across it costs a re-run and nothing more — a profile still
|
||
recognises bundles stamped by earlier versions, so re-running writes in place
|
||
instead of refusing. `DEFAULT` remains OKF v0.1 on every axis upstream owns.
|
||
|
||
| Profile | Contract | Status |
|
||
|---|---|---|
|
||
| `DEFAULT` | commons' ingest-spec layer (OKF v0.1 semantics) | stable |
|
||
| `STRICT_V1` | a consumer's ratified v0.1 contract | stable |
|
||
| `OKF_V0_2` | OKF v0.2 | **provisional**, pre-release only |
|
||
| `STRUCTURED_V1` | `DEFAULT` plus a faceted, derived index | stable |
|
||
| `OKF_LATEST` | alias for the latest version supported as *stable* | currently `DEFAULT` |
|
||
|
||
`STRUCTURED_V1` is `DEFAULT` in every respect but the index. Under it, Door B
|
||
derives each dropped document's title, number, hierarchy and cross-references,
|
||
writes them into the concept's own frontmatter, and carries them into the index
|
||
entry — so a consumer can reason over the bundle rather than only look things
|
||
up in it. Every inferred field is named in a `derived` list, because an
|
||
unmarked heuristic is worse than no heuristic: the consumer cannot know when to
|
||
doubt it. A pointer to a document not dropped yet is rendered `N200?` rather
|
||
than omitted, since a bundle is built up over several drops and an absence that
|
||
leaves no trace is the dangerous kind. Carrying the metadata costs index
|
||
characters — roughly 3x to 6x the flat index, depending on how many facets the
|
||
profile names — and the facet key set is the dial. Design record and
|
||
measurements: [`docs/plan/structure-derivation.md`](docs/plan/structure-derivation.md).
|
||
|
||
`OKF_V0_2` ships first as a pre-release to a named pilot set and may change on
|
||
their feedback without a deprecation cycle. Pin the versioned constant rather
|
||
than `OKF_LATEST` unless you have explicitly opted into tracking; `OKF_LATEST`
|
||
moves at general availability, which is a deliberate release event rather than
|
||
a side effect of an upgrade.
|
||
|
||
Selecting a profile is keyword-only, so existing call sites are unaffected:
|
||
|
||
```python
|
||
materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)
|
||
```
|
||
|
||
A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and
|
||
puts the declaration in the bundle-root `index.md`'s frontmatter block. The
|
||
profile names the key; the **caller supplies the value**, because that value
|
||
tracks the upstream version and is not this library's to decide:
|
||
|
||
```python
|
||
materialize_bundle(
|
||
manifest, bundle_dir, ingested_at,
|
||
profile=OKF_V0_2,
|
||
root_frontmatter_values={"okf_version": "0.2"},
|
||
)
|
||
```
|
||
|
||
Omit the argument and no frontmatter block is written. Offering a key the
|
||
profile does not name is refused before anything is written to disk.
|
||
|
||
### Attested computations (v0.2 §10)
|
||
|
||
`OKF_V0_2` supports the `Attested Computation` type as a **format**: its five
|
||
contract fields — `runtime`, `parameters`, `computation`, `executor`,
|
||
`attester` — are emitted in canonical position, judged, and round-tripped.
|
||
`runtime` is required for that type and for no other, which the profile
|
||
expresses through `FrontmatterSchema.required_by_type`; a type the mapping does
|
||
not name carries no extra requirement, because §14 forbids a consumer to reject
|
||
on an unknown `type`.
|
||
|
||
Nothing here executes a computation or checks an attestation. Upstream defers
|
||
the receipt and verdict wire formats, so there is no contract to implement, and
|
||
the question an attestation answers — was this value produced the sanctioned
|
||
way — is not this library's. It re-enters scope when upstream specifies the
|
||
protocol.
|
||
|
||
On the import side, a third-party concept may name an `executor` or `attester`
|
||
resource pointing at executable code. Door C imports the **pointer** and never
|
||
the code — it writes concepts verbatim and skips every non-`.md` file — so such
|
||
a reference may not resolve, or may resolve to a file the destination tree
|
||
already holds under that path. Each one is reported in
|
||
`ImportResult.unverified_references`; the concept still merges, because §14
|
||
forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer
|
||
to surface rather than silently drop. The report names the pointer key, not the
|
||
resource it points at: recovering the resource needs the structured reader.
|
||
|
||
One limit worth knowing before you write such a concept: §10.2 presents
|
||
`executor` and `attester` as nested block mappings, and this library's
|
||
frontmatter parser is line-oriented. It reads inline **flow** mappings
|
||
(`executor: { resource: …, receipt: [ … ] }`) as opaque values that round-trip
|
||
unchanged, but it cannot read the block form — two block mappings that both
|
||
carry a `resource` collapse into one namespace and the first is lost. Write the
|
||
flow form; both are valid YAML, and a real YAML consumer recovers the same
|
||
structure from either.
|
||
|
||
## Non-goals
|
||
|
||
- Verdict/feedback machinery from the method specification (stays in the
|
||
consuming repositories).
|
||
- Embedding- or retrieval-layer functionality.
|
||
- Security functionality, in either runtime — that is always
|
||
`llm-ingestion-guard`'s domain.
|
||
|
||
## Requirements
|
||
|
||
Python 3.10+, and exactly one runtime dependency — the security boundary,
|
||
`llm-ingestion-guard>=1.2,<2.0`. Everything else is stdlib. The commands are
|
||
under [Install](#install); what follows is why they look the way they do.
|
||
|
||
From a checkout, the test suite runs with:
|
||
|
||
```
|
||
.venv/bin/python -m pytest
|
||
```
|
||
|
||
The suite is the verification surface for everything above: 596 tests, run on
|
||
2026-08-21 against this branch with the `[extract]` extra installed. Without
|
||
the extra the same suite is 589 passed and 7 skipped, measured the same day:
|
||
the seven cover the parser path, and the tests holding the fail-fast rejection
|
||
for an uninstalled extra run in both. It is not shipped in an installed
|
||
distribution — `tests/` lives at the repository root, so this command needs a
|
||
clone rather than a `pip install`.
|
||
|
||
A git URL is a PEP 508 direct reference and pins one exact tag, so it is an
|
||
install-time *channel*, not the pin: the range above stays the declared
|
||
dependency — a wheel built from this branch carries `Requires-Dist:
|
||
llm-ingestion-guard<2.0,>=1.2`, measured 2026-08-23 — and resolves normally
|
||
once the package index exists. A wheel built from a *tag* carries that tag's
|
||
range instead, which is why the install commands pair tag with tag.
|
||
|
||
### Binary extraction
|
||
|
||
The optional `[extract]` extra ships two things: `pdfplumber` (MIT) for `pdf`,
|
||
and `pypandoc-binary` for five office formats. It is opt-in because it pulls
|
||
binary wheels, which the default install must never do — the single runtime
|
||
dependency rule covers the default install and this extra sits outside it.
|
||
|
||
The converter **binary travels inside the wheel** and is resolved by path
|
||
rather than found on `PATH`, with its version asserted against a pin. A host
|
||
carrying a different converter is refused, not silently used: extraction is
|
||
deterministic within a converter version and not across one.
|
||
|
||
| Format | Reader | Evidence |
|
||
|---|---|---|
|
||
| `pdf` | `pdfplumber` | measured |
|
||
| `docx` | converter | measured |
|
||
| `xlsx` | converter | measured |
|
||
| `pptx` | converter | **unmeasured** |
|
||
| `odt` | converter | **unmeasured** |
|
||
| `rtf` | converter | **unmeasured** |
|
||
|
||
**`unmeasured` means what it says.** The corpus this work was measured on
|
||
contains **zero** `pptx`, `odt` and `rtf` files, so those three rows work by
|
||
construction and have never been checked against a document anyone wrote.
|
||
They are not known to be broken; they are not known to be right either, and
|
||
the distinction is the point.
|
||
|
||
**What stays out.** `.doc` (Word 97) is not supported — the converter does not
|
||
read it. Rastered or scanned PDFs are refused rather than persisted as empty
|
||
concepts, because this library does not do OCR. Drawn content — figures,
|
||
diagrams, shapes — does not survive extraction in any format here, and every
|
||
extraction says so with a warning. Structured table recovery is out of scope.
|
||
|
||
Request it by appending `[extract]` to the package name in whichever install
|
||
command from [Install](#install) you are using — this package is not on an
|
||
index, so a bare `pip install 'llm-ingestion-okf[extract]'` does **not** work
|
||
today, and the error message naming that command is written for the day it
|
||
does. The extra is unreleased: it reaches a consumer through a tag that
|
||
contains it, and no such tag exists yet.
|
||
|
||
Two properties of the extra are worth knowing before depending on its output:
|
||
|
||
- **Extracted text is pinned to an exact parser version.** `pdfplumber` pins
|
||
`pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
|
||
releases with no stability contract. Extraction is deterministic within a
|
||
parser version and not guaranteed across one, so a golden fixture built on
|
||
extracted PDF text is a fixture migration away from any parser upgrade.
|
||
- **Text extraction recovers text, and nothing that is drawn.** Figures,
|
||
diagrams and images have no text to recover — only their captions survive —
|
||
so a bundle built from drawn documents is incomplete by construction. The
|
||
library says so itself: every `pdf` extraction emits an `ExtractionWarning`.
|
||
Structured table recovery is separately out of scope; PDFs enter as prose.
|
||
|
||
The planned Node half targets Node/ESM with zero npm dependencies.
|
||
|
||
## License
|
||
|
||
MIT — see [LICENSE](LICENSE).
|