The reading direction existed only for someone standing in a clone. `consume`, `contract_check` and `skill` moved from `tools/` into the package and are reachable as `okf consume`, `okf check` and `okf skill`; `okf project` is new and does the whole thing in one command. The red measurement: a consumption skill generated from a checkout carried 4 lines naming that checkout by absolute path, 2 of them the commands the skill tells a reader to run. It now names `okf consume` and `okf check`, and a test asserts this repository appears in it nowhere, with a known-positive so the zero is a measurement rather than a search that could not find. The `tools/` files stay as ALIASES, not re-exports: a re-export binds copies of the names into a second module object, so a caller patching one patches a binding the implementation never reads. Two tests that monkeypatch okf_consume went green again only under the alias. Every published reproduction block runs unchanged. The template and docs/consumption-contract.md (the section 7.4 known-positive) are force-included into the wheel from the file they are authored in, so both travel with the commands that cannot run without them and there is still one authored copy of each. Step 0, before any of it: okf build's default gained Arm E (--table-grid), with --no-table-grid as its opt-out. The default moved to D plus F earlier the same day on Arm F's published 5 of 12 -- a figure measured with Arm E ON. Without it the fold has no joined table to fold, and the shipped default scored 2 of 12 with docx 0 of 3. Measured on the operator's folder: 30 md / 15 concepts on the new default against 43 / 28 without Arm E. Install measurement from a fresh uv tool install, empty folder, this repository nowhere on PYTHONPATH: 5 documents in, 15 concepts out, 0 references to tools/ in the generated skill, okf check conformant (15 rules, 0 findings). Deviation stated rather than hidden: the order asked that tests/test_okf_consume.py be left untouched. Two assertions in it read a PATH, which is the one thing this work changes. Both were moved and the second made stronger -- it now asserts every command the README recipe names is a subcommand the CLI registers, which a file existing on disk never proved. Suite 1414 -> 1427. ruff clean, mypy --strict clean over 21 files. Record: docs/2026-09-08-o5-okf-project.md Co-Authored-By: Claude <claude-opus-5>
608 lines
33 KiB
Markdown
608 lines
33 KiB
Markdown
# llm-ingestion-okf
|
||
|
||
Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import. Security delegated to llm-ingestion-guard.
|
||
|
||
Status: phases 1–3 are implemented. Phase 1 (spec-based ingestion) covers
|
||
manifest validation, the `file`/`sql`/`http` connectors, deterministic
|
||
materialization, index generation, and the golden fixture suite under
|
||
`examples/`. Phase 2 adds the bundle inbox (`process_inbox`) and
|
||
external-bundle import (`import_bundle`), both against an **injected** persist
|
||
gate, with `llm_ingestion_okf.guard_adapter` wiring that gate to the real
|
||
guard (see below). Phase 3 makes the bundle contract configurable, so types,
|
||
layers, frontmatter sets, index shape, and reserved-file policy are carried by
|
||
a profile rather than by constants (see [Upstream OKF
|
||
versions](#upstream-okf-versions)). Binary extraction runs behind the
|
||
optional `[extract]` extra: `pdf` through a PDF parser, and five office
|
||
formats through a vendored document converter. Three of those five office
|
||
rows are **unmeasured** — see [Binary extraction](#binary-extraction). Phase 4
|
||
(the Node half) is planned (see `docs/plan/`).
|
||
|
||
## Install
|
||
|
||
Python 3.10+. Neither this package nor the guard it depends on is on a package
|
||
index yet. With uv, one command is enough:
|
||
|
||
```
|
||
uv pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
|
||
```
|
||
|
||
uv resolves the guard on its own, because it reads the `[tool.uv.sources]`
|
||
entry in the `pyproject.toml` **of the tag it is installing**, and `v0.4.0`
|
||
points that entry at the guard tag below. Measured 2026-07-25 and re-measured
|
||
2026-08-20 with an empty `uv` cache; both runs installed
|
||
`llm-ingestion-guard==0.2.0` + `llm-ingestion-okf==0.4.0` and imported clean.
|
||
|
||
With plain pip, the transitive git dependency does not resolve on its own —
|
||
**install the guard first**, or installing this package fails with
|
||
`No matching distribution found for llm-ingestion-guard`:
|
||
|
||
```
|
||
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v0.2.0"
|
||
pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
|
||
```
|
||
|
||
The guard tag is paired to the okf tag, not to this branch: `v0.4.0` declares
|
||
`llm-ingestion-guard>=0.2,<0.3`, which `v0.2.0` satisfies and later guard tags
|
||
do not. `main` has since moved its own pin to `>=1.2,<2.0` (see
|
||
[Requirements](#requirements)); that pin reaches you in the next stable tag,
|
||
not in the commands above. Reading a pin off this branch and installing it
|
||
against `v0.4.0` is the one combination that fails.
|
||
|
||
`v0.6.0` is the current tag and the one the three-line form under [Consume in
|
||
Claude Code](#consume-in-claude-code) installs: it is the first tag carrying the
|
||
`okf project`, `okf consume`, `okf check` and `okf skill` subcommands, without
|
||
which that form does not exist. `v0.4.0` is the last tag before the OKF v0.2
|
||
work. `v0.5.0a2` is a pre-release for the named OKF v0.2 pilot set only.
|
||
|
||
**`v0.6.0` does not make OKF v0.2 generally available.** `OKF_LATEST` is
|
||
unchanged and still points at `DEFAULT`; flipping that alias is the GA event and
|
||
this tag is not it (see [Upstream OKF versions](#upstream-okf-versions)).
|
||
|
||
## Build
|
||
|
||
Installing the package installs one command. A folder of documents in, an OKF
|
||
bundle out:
|
||
|
||
```
|
||
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
|
||
```
|
||
|
||
It walks the folder recursively, proposes a segmentation for each document with
|
||
the mechanical rules, replays those proposals through the bundle inbox, writes
|
||
the bundle and its `log.md`, and prints the run's numbers. Every proposal is
|
||
marked `PROPOSED` and `adjudicated: false` — the command segments nothing a
|
||
human has approved, and says so in the artifact.
|
||
|
||
The last line that matters is the conservation identity: `merged + coded
|
||
rejections == N`, where `N` is the folder's file count read at run time. **The
|
||
run exits non-zero when it does not hold**, and names the unaccounted files, so
|
||
a pipeline cannot mistake a partial bundle for a complete one.
|
||
|
||
Flags worth knowing: `--segments off` ingests each document as one concept and
|
||
asks for no root values; `--plans-dir` keeps the proposals instead of
|
||
discarding them; `--report` writes the full report to a file as well as stdout.
|
||
`--ingested-at` and `--proposed-at` default to `1970-01-01T00:00:00Z` rather
|
||
than the clock, so two builds of the same folder are byte-identical — a
|
||
wall-clock default would break rebuild-equals-incremental for every caller who
|
||
did not pass them.
|
||
|
||
### The segmentation flags
|
||
|
||
Six rules are reachable from `okf build`. **Two of them are ON by default since
|
||
2026-09-08** — `--outline-run 3` and `--unit-fold`, an operator decision — and
|
||
each has an explicit opt-out, `--outline-run 0` and `--no-unit-fold`. Passing
|
||
both opt-outs reproduces the pre-2026-09-08 bytes exactly. The other four are
|
||
off. Each line below carries the number it was measured at, and nothing beyond
|
||
it.
|
||
|
||
**A re-run is what this costs a consumer, and it is not a small one:** on the
|
||
43-document reference corpus the default bundle goes from **629 concepts in
|
||
1108 files** to **517 in 969**. The proposer's own defaults
|
||
(`tools/okf_propose_segments.py`) did NOT move, so every published reproduction
|
||
block still runs as written.
|
||
|
||
| flag | what it does | measured |
|
||
|---|---|---|
|
||
| `--outline-run N` (default **3**) | also propose a boundary where the document's own bare-integer numbering sustains an ascending run of at least `N`; `0` is this arm's opt-out | a tender PDF whose headings are bare integers: **no boundary** at `0`, **9 concepts** at `3`, against a reference of 9 |
|
||
| `--table-grid` | a pandoc grid-table rule line no longer closes an open table block, so one grid table is one concept | a `.docx` experience list: **21 → 6** concepts |
|
||
| `--unit-fold` (**on** by default; opt out with `--no-unit-fold`) | discard a contents-list run, fold a deeper heading into its parent, fold a table into the shorter heading that introduces it. Adds no boundary, so it can only reduce a plan | on a 12-document sample scored against an operator's unit worksheet: **5 of 12** match — but that figure was measured with `--table-grid` ON, and the shipped default does not include it. Measured without it the same sample scores **2 of 12**, `docx` **0 of 3**, because the fold's table clause has no joined table to fold |
|
||
| `--keep-table-heading` | keep a heading whose body is empty only because a table opens under it, and absorb that table into its span | the two spreadsheets in that corpus, and **0 of 32 `pdf` and 0 of 5 `docx`**: the concept count does not move (1 → 1), its first byte does — the concept gains the heading line it was missing |
|
||
| `--sheet-section-rows` | cut an open table block at the rows that label its sections — a run of at least three rows whose first cell is a bare numeric label. The opposite direction from `--table-grid`, which decides how far a block extends | a tender price sheet whose whole body is one table block: **1 → 12 concepts**, against a reference of 11 cost groups plus the sheet's preamble. Whole corpus: **1 of 39** readable documents changes, **0 of 32 `pdf`, 0 of 5 `docx`, 1 of 2 `xlsx`** |
|
||
| `--drop-wrapped-outline` | do not admit an `--outline-run` candidate whose line continues onto the next one. Judges recovered candidates only, never a heading the document declares | quoted regulation text, whose numbered paragraphs match the outline grammar exactly: **4 → 1 concepts**, the reference. Whole corpus: **5 of 39**, all `pdf`; on the 12-document sample **8 of 34** outline candidates wrap, and none of the 26 the operator kept |
|
||
|
||
They compose, and the order above is the order they apply in. Measured on a
|
||
five-document tender folder (2 `pdf`, 2 `docx`, 1 `xlsx`), concepts per
|
||
document:
|
||
|
||
| document | default | `--outline-run 3` | `+ --table-grid` | `+ --unit-fold` | `+ --keep-table-heading` | `+ --sheet-section-rows --drop-wrapped-outline` |
|
||
|---|---|---|---|---|---|---|
|
||
| tender PDF, technical requirements | 1 (no boundary) | 9 | 9 | 9 | 9 | 9 |
|
||
| tender PDF, technical layout | 1 (no boundary) | 1 | 1 | 1 | 1 | 1 |
|
||
| price sheet `.xlsx` | 1 | 1 | 1 | 1 | 1 | **12** |
|
||
| experience list `.docx` | 21 | 21 | 6 | 3 | 3 | 3 |
|
||
| agreement `.docx` | 2 | 2 | 1 | 1 | 1 | 1 |
|
||
| markdown files in the bundle | 31 | 49 | 33 | 30 | 30 | 52 |
|
||
|
||
Every column merged 5 of 5 with 0 rejections. The reference for the first row
|
||
is 9, so the default is a full arm behind what the proposer can do on that
|
||
document — which is a statement about the default, not a licence to change it
|
||
here.
|
||
|
||
Measured 2026-09-08 on a 43-file corpus (33 `pdf`, 5 `docx`, 2 `xlsx`, and
|
||
three files no reader accepts), one
|
||
`okf build` invocation replacing the shell loop over `tools/` that produced the
|
||
same corpus's bundle on 2026-09-03:
|
||
|
||
| figure | value |
|
||
|---|---|
|
||
| `N` (folder file count, computed) | 43 |
|
||
| merged | 39/43 |
|
||
| coded rejections | 4/43 (`extractor_unknown` 3, `extractor_empty_pdf` 1) |
|
||
| K1b | `39 + 4 = 43 = N`, exit `0` |
|
||
| files written | 1108 |
|
||
| identical to the 2026-09-03 bundle | 1104/1108 |
|
||
| wall time | 842.82 s total, 19.600 s per file (re-measured 2026-09-08) |
|
||
|
||
The six files that differ are all in the corpus's two spreadsheet documents, and
|
||
they are the change reported in `docs/2026-09-08-prisform-og-loggen-k2.md`: a
|
||
spreadsheet's tables are now written as pipe tables, so each row is one line
|
||
with its cells delimited rather than padded out to the widest cell in the
|
||
column. Two concept files are renamed by it, two are removed under their old
|
||
names, and the two documents' own `index.md` follow. The root `index.md` is
|
||
identical to the stored one again, because this library no longer links the
|
||
bundle's `log.md` from it. The command's own byte-identity test compares it
|
||
against the two scripts at the current commit, where the two agree over the
|
||
whole tree.
|
||
|
||
## Consume
|
||
|
||
The other direction: a bundle plus one question in, one bounded, contract-shaped
|
||
payload out.
|
||
|
||
```
|
||
python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json
|
||
```
|
||
|
||
`tools/okf_consume.py` is the **pre-pass** `docs/consumption-contract.md` § 1
|
||
defines — the deterministic program that reads the bundle, ranks its concepts,
|
||
cuts them to a bounded set and emits one payload. It decides nothing about the
|
||
question; the skill that reads the payload does the judgement. It calls no
|
||
model, opens no socket, imports nothing outside the standard library and this
|
||
package, and takes no clock: the same bundle bytes and the same
|
||
`(question, k, limit, cost_vocabulary, reserve_top_rank, rarity_weight)`
|
||
produce byte-identical output.
|
||
|
||
`--cost-vocabulary` is off by default and widens one question class: it lets a
|
||
declared list of cost/price/quantity terms bridge a question and a document that
|
||
name money with different words. The gate is the question — one naming no such
|
||
term gets byte-identical bytes either way — and what it does and does not close
|
||
is measured in `docs/2026-09-08-blindsone-below-k-k2.md`.
|
||
|
||
`--reserve-top-rank` is off by default and answers a different objection: the
|
||
budget is packed by an exact knapsack, which maximises a SUM of scores and
|
||
therefore has no opinion about rank, so a top-ranked excerpt costing a large
|
||
share of the budget is out-summed by many small ones. Measured, that made `--k`
|
||
a dial that could EVICT the concept a question was asked about. The flag gives
|
||
rank one its bytes before the pack runs — after the `over_budget_alone`
|
||
pre-exclusion, never before — and the payload then declares
|
||
`budget.reserved`. On a 629-concept corpus it changed the delivered set in 2 of
|
||
24 measured combinations, both of them that eviction:
|
||
`docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
|
||
|
||
`--rarity-weight` is off by default and weights each lexical hit by
|
||
`log(N/df)` over the bundle's own concepts instead of counting it as one, so a
|
||
requirement number is not worth what a common verb is worth. The default being
|
||
off is a measurement rather than a preference: on four corpora it took one gold
|
||
concept from withheld to delivered and a priced sheet from candidate rank 10 to
|
||
2, left one gold rank unmoved, and cost another seven rank positions — because
|
||
the four-character prefix matcher makes a unique identifier read as
|
||
135-of-446 common on that bundle. Where it cannot help is decomposed rather
|
||
than guessed: RRF fuses RANKS, so a weight moves nothing on a signal the gold
|
||
already leads. `docs/2026-09-08-sjeldenhetsvekt.md`.
|
||
|
||
It emits the § 8 shape — `contract`, `bundle` (`bundle_id` plus a
|
||
`sha256-tree:` content identity), `budget` (unit, instrument, limit, spent and a
|
||
validated known-positive), `denominators`, `excerpts` and `withheld` — and every
|
||
withheld concept names the rule that dropped it, from a closed set of six.
|
||
|
||
Every excerpt carries the concept's `title`, and — when the producer wrote them
|
||
— `req_number`, the SPEC § 5.1 address `sources`, and **every top-level
|
||
`source_*` key**, by prefix rather than by allowlist: a fixed list names the
|
||
locators its author thought of, and one real bundle locates by
|
||
`source_element_id` on 269 of its 274 concepts. A key the producer did not write
|
||
stays absent rather than arriving empty, and an address this reader cannot
|
||
decode is named (`sources_unreadable`) rather than dropped into the same
|
||
silence. The reason is a measurement: with `concept_id` and body text alone, a
|
||
delivered gold concept at rank 1 still left the answer unable to name the
|
||
document it was quoting.
|
||
`considered == withheld + delivered` closes by construction, and the payload is
|
||
refused rather than reported when it does not.
|
||
|
||
Three exit codes, not two: **0** a payload was written, **1** the run happened
|
||
and refused (the budget admitted none of the concepts that answered the
|
||
question, or an asserted `--ref` contradicted the bytes), **2** the run did not
|
||
happen. Collapsing 2 into 1 would report an unread bundle as a failed cut.
|
||
`--ref` is an **assertion**, never an override — the identity is always computed
|
||
from the bytes, because labelling a payload with an identity its bytes do not
|
||
have is the one thing § 3.3 exists to prevent.
|
||
|
||
Check any payload against the skill that will read it:
|
||
|
||
```
|
||
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json
|
||
```
|
||
|
||
`skills/okf-consume/` is the first instantiated consumption skill: a filled copy
|
||
of `skills/okf-consume-template/` naming this pre-pass, with every per-corpus
|
||
hole replaced by a measured value. Measured 2026-09-07 on a 629-concept bundle,
|
||
hit@8 was **5 of 6** questions at rank 1 against a chance baseline of **1.35 of
|
||
6** — with one control that failed, and both are in
|
||
`docs/2026-09-07-okf-konsumskill-maaling.md` with the honesty limits stated.
|
||
|
||
## Consume in Claude Code
|
||
|
||
A folder of documents to an answer a model can cite, in **three lines**. You do
|
||
not need this repository — the first line installs the command, the second
|
||
builds the bundle and writes a skill beside it, the third asks.
|
||
|
||
```sh
|
||
uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.6.0"
|
||
okf project ~/my-documents
|
||
claude
|
||
```
|
||
|
||
`okf project` writes the bundle to `.okf/<id>/` and a skill to
|
||
`.claude/skills/<id>-consume/` in the **current directory**, then prints what it
|
||
read, what it wrote, and which documents a question cannot reach. Start `claude`
|
||
in that directory and ask in plain language; the generated skill runs the
|
||
pre-pass and the contract check itself and marks every claim with its source.
|
||
|
||
`<id>` is the folder's name reduced to `[a-z0-9-]`. Run it once per folder with
|
||
`--id <name>` to have several bundles reachable at once — each skill carries its
|
||
own `bundle_id`, which is what lets a model pick between them. `--out <dir>`
|
||
puts the project somewhere other than the current directory.
|
||
|
||
Measured 2026-09-08 from a fresh `uv tool install` with this repository nowhere
|
||
on the path: 5 documents in, 15 concepts out, a skill carrying **0** paths into
|
||
any checkout, and `okf check` conformant on its own payload (15 rules, 0
|
||
findings). Before that day the same result took a `PYTHONPATH`, a snapshot of a
|
||
clone, and a generated skill that named that clone by absolute path on four
|
||
lines — so it could not be moved, shared, or run by anyone else.
|
||
|
||
### The same thing in steps, if you want to see the payload
|
||
|
||
```sh
|
||
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
|
||
okf skill ./bundle --out ./project/.claude/skills/my-bundle-consume
|
||
okf consume ./bundle --question "your question" --out /tmp/payload.json
|
||
okf check --skill ./project/.claude/skills/my-bundle-consume/SKILL.md --payload /tmp/payload.json
|
||
```
|
||
|
||
A bundle you only have read access to is fine — the generator only reads it.
|
||
|
||
### The honest limits
|
||
|
||
Measured on **four questions** across two bundles, which is a demonstration and
|
||
not a hit rate. The ranking is lexical, and one of the four found a topic the
|
||
bundle **does** cover and did not rank it into the cut — the skill then said so
|
||
with its denominator instead of answering, which is the behaviour the contract
|
||
asks for, but a miss is still a miss.
|
||
`docs/2026-09-08-claude-code-skill-vilkaarlig-bundle.md` has the runs.
|
||
|
||
Two more things the summary tells you and this paragraph will not repeat: a
|
||
document that landed **whole** (no heading, table or numbered outline to cut it
|
||
on) comes back as one excerpt, which the budget often refuses and which often
|
||
does not carry the answer at the place you asked about; and a document that is
|
||
in the folder but **not** in the bundle cannot be quoted at all. Both cases are
|
||
answered `[sourced-not-sufficient]`, and `okf project` names the documents.
|
||
|
||
### The skill that runs this for you
|
||
|
||
`skills/okf-prosjekt/` in this repository is a Claude Code skill (Norwegian)
|
||
that wraps the command above: it takes a folder, runs `okf project`, and reads
|
||
the summary back. Install it for your user account after cloning:
|
||
|
||
```sh
|
||
mkdir -p ~/.claude/skills && cp -R skills/okf-prosjekt ~/.claude/skills/
|
||
```
|
||
|
||
## Implemented scope (v1)
|
||
|
||
The library provides three entry points for getting content into an OKF
|
||
bundle:
|
||
|
||
1. **Spec-based ingestion.** An implementation of the normative ingest
|
||
specification owned by `portfolio-optimiser-commons`: manifest →
|
||
`file`/`sql`/`http` connector → deterministic materialization of
|
||
`ingest-{id}.md` concept files → index generation. Zero model calls in the
|
||
run path; output is reproducible byte-for-byte against golden fixtures.
|
||
2. **Bundle inbox.** A drop directory where common file types are converted
|
||
to OKF concept files. All file-type→text extraction lives in this library:
|
||
`md`, `txt`, `csv`, `json`, and `html` are handled by the stdlib core;
|
||
`pdf` and the five office formats (`docx`, `xlsx`, `pptx`, `odt`, `rtf`)
|
||
require the optional `[extract]` extra and are rejected fail-fast without
|
||
it. Extracted text passes the security gate before anything is persisted.
|
||
The drop directory is walked **recursively**, in sorted relative-path order:
|
||
a file at any depth is ingested and records its path relative to the inbox
|
||
root as its `source_file`, while dot-directories and a bundle directory
|
||
sitting inside the inbox are skipped with a reported code.
|
||
|
||
Under the segmented v0.2 profile a concept also points back at the document
|
||
it was extracted from, so an agent citing it can open the original at the
|
||
right place: `sources: [{ resource, title }]` in the spec's own §5.1 form,
|
||
where `resource` is the inbox-relative path, plus a locator per format —
|
||
`source_pages` for a PDF, `source_sheet` and `source_rows` for a
|
||
spreadsheet, `source_lines` otherwise. The locator keys are this library's
|
||
own, because §5.1 has no field for a place *within* a resource; the line
|
||
numbers index the extracted text and say so. Measurements:
|
||
[`docs/2026-09-08-proveniens-k2.md`](docs/2026-09-08-proveniens-k2.md).
|
||
|
||
<!-- extract-formats: .md, .txt, .csv, .json, .html, .htm, .pdf, .docx, .xlsx, .pptx, .odt, .rtf -->
|
||
3. **External bundle import.** Import and merge of third-party OKF bundles:
|
||
each concept is assessed via the security gate, and only concepts that
|
||
pass are merged, materialized, and linked into the index.
|
||
|
||
## Boundary: security is delegated
|
||
|
||
Security is owned by the sibling package
|
||
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
|
||
(pinned `>=1.2,<2.0`). The division is strict:
|
||
|
||
- **guard** answers "is this content safe to persist?" — scan, sanitize,
|
||
quarantine, fail-secure, provenance stamping.
|
||
- **this library** does the plumbing — connect a source, materialize a
|
||
deterministic OKF bundle, generate the index.
|
||
|
||
No security functionality is reimplemented here.
|
||
|
||
### What is gated today: read this before trusting a door
|
||
|
||
- **Door A (`materialize_bundle`) is ungated.** It calls nothing before
|
||
writing to disk and writes what it is given. A caller materializing
|
||
untrusted content is responsible for gating it.
|
||
- **Doors B and C (`process_inbox`, `import_bundle`) gate through an adapter
|
||
you pass in.** Each takes a `gate` argument; the flow hands it the content
|
||
and obeys the verdict, refusing to persist anything that does not clear the
|
||
guard's non-blocking floor — including a disposition it does not recognise,
|
||
and (at Door C) a concept the gate returned no verdict for. What it cannot
|
||
do is check that your adapter is a real guard: a permissive stub approves
|
||
everything, and the flow will believe it.
|
||
|
||
`llm_ingestion_okf.guard_adapter` is the adapter over the real guard, and the
|
||
only module here that imports it — importing the package itself does not:
|
||
|
||
```python
|
||
from llm_ingestion_okf import process_inbox
|
||
from llm_ingestion_okf.guard_adapter import inbox_gate
|
||
|
||
result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
|
||
okf_type="reference", gate=inbox_gate)
|
||
```
|
||
|
||
Two properties of that adapter are worth knowing before you rely on it.
|
||
It screens the **exact bytes it persists** — the guard's `prepare_input`
|
||
bookend prepares text for a model call, which this library never makes, so
|
||
only `screen_output` is used and the screened string is the written string.
|
||
And it **refuses rather than repairs**: a file carrying an invisible
|
||
zero-width or bidi character is rejected, not silently stripped and written.
|
||
Door B screens under the untrusted-upload policy, so any finding at all is
|
||
held back rather than persisted.
|
||
|
||
This section is stated plainly because earlier wording ("calls the guard at
|
||
every persist gate") described the intended end state in the present tense,
|
||
and a consumer reasonably read it as safe-by-default.
|
||
|
||
## Roadmap
|
||
|
||
The library is built in four phases so that every known OKF surface in the
|
||
ecosystem is eventually covered. Each phase has a detailed plan with
|
||
verification criteria:
|
||
|
||
1. Spec-based ingestion (Python) with byte-exact golden fixtures —
|
||
[plan](docs/plan/phase-1-door-a.md).
|
||
2. Bundle inbox and external-bundle import (Python), guard-gated —
|
||
[plan](docs/plan/phase-2-doors-b-c.md).
|
||
3. Configurable bundle contract (types, layers, frontmatter sets, index
|
||
shape, and reserved-file policy as configuration), enabling stricter
|
||
bundle profiles such as `strict-v1` —
|
||
[plan](docs/plan/phase-3-configurable-contract.md).
|
||
4. A `node/` half: a zero-dependency Node/ESM package (importable and
|
||
CLI-invokable, vendored per consumer) providing bundle checking, index
|
||
generation, inbox processing, and document conversion for the OKF
|
||
second-brain plugin ecosystem. The Python and Node halves share the OKF
|
||
contract and fixture suite, not code —
|
||
[plan](docs/plan/phase-4-node-half.md).
|
||
|
||
## Upstream OKF versions
|
||
|
||
The library targets the current latest version of Google's OKF. Support is
|
||
**additive** — a new upstream version arrives as a new profile, never as a
|
||
migration of an existing one — so an *upstream* release does not change the
|
||
bytes an existing profile emits.
|
||
|
||
That guarantee is about upstream, and one profile tracks a second contract as
|
||
well. `DEFAULT` states the ingest-spec owned by `portfolio-optimiser-commons`,
|
||
so when they change that spec, `DEFAULT` follows them. It happened on
|
||
2026-08-09: `generated` moved from `true` to
|
||
`{ by: process:okf-ingest, at: <ingested_at> }`, one changed line per generated
|
||
file. Upgrading across it costs a re-run and nothing more — a profile still
|
||
recognises bundles stamped by earlier versions, so re-running writes in place
|
||
instead of refusing. `DEFAULT` remains OKF v0.1 on every axis upstream owns.
|
||
|
||
| Profile | Contract | Status |
|
||
|---|---|---|
|
||
| `DEFAULT` | commons' ingest-spec layer (OKF v0.1 semantics) | stable |
|
||
| `STRICT_V1` | a consumer's ratified v0.1 contract | stable |
|
||
| `OKF_V0_2` | OKF v0.2 | **provisional**, pre-release only |
|
||
| `STRUCTURED_V1` | `DEFAULT` plus a faceted, derived index | stable |
|
||
| `OKF_LATEST` | alias for the latest version supported as *stable* | currently `DEFAULT` |
|
||
|
||
`STRUCTURED_V1` is `DEFAULT` in every respect but the index. Under it, Door B
|
||
derives each dropped document's title, number, hierarchy and cross-references,
|
||
writes them into the concept's own frontmatter, and carries them into the index
|
||
entry — so a consumer can reason over the bundle rather than only look things
|
||
up in it. Every inferred field is named in a `derived` list, because an
|
||
unmarked heuristic is worse than no heuristic: the consumer cannot know when to
|
||
doubt it. A pointer to a document not dropped yet is rendered `N200?` rather
|
||
than omitted, since a bundle is built up over several drops and an absence that
|
||
leaves no trace is the dangerous kind. Carrying the metadata costs index
|
||
characters — roughly 3x to 6x the flat index, depending on how many facets the
|
||
profile names — and the facet key set is the dial. Design record and
|
||
measurements: [`docs/plan/structure-derivation.md`](docs/plan/structure-derivation.md).
|
||
|
||
`OKF_V0_2` ships first as a pre-release to a named pilot set and may change on
|
||
their feedback without a deprecation cycle. Pin the versioned constant rather
|
||
than `OKF_LATEST` unless you have explicitly opted into tracking; `OKF_LATEST`
|
||
moves at general availability, which is a deliberate release event rather than
|
||
a side effect of an upgrade.
|
||
|
||
Selecting a profile is keyword-only, so existing call sites are unaffected:
|
||
|
||
```python
|
||
materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)
|
||
```
|
||
|
||
A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and
|
||
puts the declaration in the bundle-root `index.md`'s frontmatter block. The
|
||
profile names the key; the **caller supplies the value**, because that value
|
||
tracks the upstream version and is not this library's to decide:
|
||
|
||
```python
|
||
materialize_bundle(
|
||
manifest, bundle_dir, ingested_at,
|
||
profile=OKF_V0_2,
|
||
root_frontmatter_values={"okf_version": "0.2"},
|
||
)
|
||
```
|
||
|
||
Omit the argument and no frontmatter block is written. Offering a key the
|
||
profile does not name is refused before anything is written to disk.
|
||
|
||
### Attested computations (v0.2 §10)
|
||
|
||
`OKF_V0_2` supports the `Attested Computation` type as a **format**: its five
|
||
contract fields — `runtime`, `parameters`, `computation`, `executor`,
|
||
`attester` — are emitted in canonical position, judged, and round-tripped.
|
||
`runtime` is required for that type and for no other, which the profile
|
||
expresses through `FrontmatterSchema.required_by_type`; a type the mapping does
|
||
not name carries no extra requirement, because §14 forbids a consumer to reject
|
||
on an unknown `type`.
|
||
|
||
Nothing here executes a computation or checks an attestation. Upstream defers
|
||
the receipt and verdict wire formats, so there is no contract to implement, and
|
||
the question an attestation answers — was this value produced the sanctioned
|
||
way — is not this library's. It re-enters scope when upstream specifies the
|
||
protocol.
|
||
|
||
On the import side, a third-party concept may name an `executor` or `attester`
|
||
resource pointing at executable code. Door C imports the **pointer** and never
|
||
the code — it writes concepts verbatim and skips every non-`.md` file — so such
|
||
a reference may not resolve, or may resolve to a file the destination tree
|
||
already holds under that path. Each one is reported in
|
||
`ImportResult.unverified_references`; the concept still merges, because §14
|
||
forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer
|
||
to surface rather than silently drop. The report names the pointer key, not the
|
||
resource it points at: recovering the resource needs the structured reader.
|
||
|
||
One limit worth knowing before you write such a concept: §10.2 presents
|
||
`executor` and `attester` as nested block mappings, and this library's
|
||
frontmatter parser is line-oriented. It reads inline **flow** mappings
|
||
(`executor: { resource: …, receipt: [ … ] }`) as opaque values that round-trip
|
||
unchanged, but it cannot read the block form — two block mappings that both
|
||
carry a `resource` collapse into one namespace and the first is lost. Write the
|
||
flow form; both are valid YAML, and a real YAML consumer recovers the same
|
||
structure from either.
|
||
|
||
## Non-goals
|
||
|
||
- Verdict/feedback machinery from the method specification (stays in the
|
||
consuming repositories).
|
||
- Embedding- or retrieval-layer functionality.
|
||
- Security functionality, in either runtime — that is always
|
||
`llm-ingestion-guard`'s domain.
|
||
|
||
## Requirements
|
||
|
||
Python 3.10+, and exactly one runtime dependency — the security boundary,
|
||
`llm-ingestion-guard>=1.2,<2.0`. Everything else is stdlib. The commands are
|
||
under [Install](#install); what follows is why they look the way they do.
|
||
|
||
From a checkout, the test suite runs with:
|
||
|
||
```
|
||
.venv/bin/python -m pytest
|
||
```
|
||
|
||
The suite is the verification surface for everything above: 596 tests, run on
|
||
2026-08-21 against this branch with the `[extract]` extra installed. Without
|
||
the extra the same suite is 589 passed and 7 skipped, measured the same day:
|
||
the seven cover the parser path, and the tests holding the fail-fast rejection
|
||
for an uninstalled extra run in both. It is not shipped in an installed
|
||
distribution — `tests/` lives at the repository root, so this command needs a
|
||
clone rather than a `pip install`.
|
||
|
||
A git URL is a PEP 508 direct reference and pins one exact tag, so it is an
|
||
install-time *channel*, not the pin: the range above stays the declared
|
||
dependency — a wheel built from this branch carries `Requires-Dist:
|
||
llm-ingestion-guard<2.0,>=1.2`, measured 2026-08-23 — and resolves normally
|
||
once the package index exists. A wheel built from a *tag* carries that tag's
|
||
range instead, which is why the install commands pair tag with tag.
|
||
|
||
### Binary extraction
|
||
|
||
The optional `[extract]` extra ships two things: `pdfplumber` (MIT) for `pdf`,
|
||
and `pypandoc-binary` for five office formats. It is opt-in because it pulls
|
||
binary wheels, which the default install must never do — the single runtime
|
||
dependency rule covers the default install and this extra sits outside it.
|
||
|
||
The converter **binary travels inside the wheel** and is resolved by path
|
||
rather than found on `PATH`, with its version asserted against a pin. A host
|
||
carrying a different converter is refused, not silently used: extraction is
|
||
deterministic within a converter version and not across one.
|
||
|
||
| Format | Reader | Evidence |
|
||
|---|---|---|
|
||
| `pdf` | `pdfplumber` | measured |
|
||
| `docx` | converter | measured |
|
||
| `xlsx` | converter | measured |
|
||
| `pptx` | converter | **unmeasured** |
|
||
| `odt` | converter | **unmeasured** |
|
||
| `rtf` | converter | **unmeasured** |
|
||
|
||
**`unmeasured` means what it says.** The corpus this work was measured on
|
||
contains **zero** `pptx`, `odt` and `rtf` files, so those three rows work by
|
||
construction and have never been checked against a document anyone wrote.
|
||
They are not known to be broken; they are not known to be right either, and
|
||
the distinction is the point.
|
||
|
||
**What stays out.** `.doc` (Word 97) is not supported — the converter does not
|
||
read it. Rastered or scanned PDFs are refused rather than persisted as empty
|
||
concepts, because this library does not do OCR. Drawn content — figures,
|
||
diagrams, shapes — does not survive extraction in any format here, and every
|
||
extraction says so with a warning. Structured table recovery is out of scope.
|
||
|
||
Request it by appending `[extract]` to the package name in whichever install
|
||
command from [Install](#install) you are using — this package is not on an
|
||
index, so a bare `pip install 'llm-ingestion-okf[extract]'` does **not** work
|
||
today, and the error message naming that command is written for the day it
|
||
does. The extra is unreleased: it reaches a consumer through a tag that
|
||
contains it, and no such tag exists yet.
|
||
|
||
Two properties of the extra are worth knowing before depending on its output:
|
||
|
||
- **Extracted text is pinned to an exact parser version.** `pdfplumber` pins
|
||
`pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
|
||
releases with no stability contract. Extraction is deterministic within a
|
||
parser version and not guaranteed across one, so a golden fixture built on
|
||
extracted PDF text is a fixture migration away from any parser upgrade.
|
||
- **Text extraction recovers text, and nothing that is drawn.** Figures,
|
||
diagrams and images have no text to recover — only their captions survive —
|
||
so a bundle built from drawn documents is incomplete by construction. The
|
||
library says so itself: every `pdf` extraction emits an `ExtractionWarning`.
|
||
Structured table recovery is separately out of scope; PDFs enter as prose.
|
||
|
||
The planned Node half targets Node/ESM with zero npm dependencies.
|
||
|
||
## License
|
||
|
||
MIT — see [LICENSE](LICENSE).
|