Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import. Security delegated to llm-ingestion-guard.
Find a file
Kjell Tore Guttormsen 732f84df6e fix(extract): HTML collapsed to one line, so 828 of 828 sections had no boundary
`_HTMLTextExtractor.text()` was `" ".join("".join(parts).split())`. `str.split()`
with no argument splits on newlines too, so extraction of ANY HTML file returned
unconditionally one line. Every boundary grammar in `propose` is line-anchored
(`_ATX`, `_NUMBERED`, `_TABLE_ROW`, `_GRID_RULE`, `_OUTLINE`, each with `^`), and
on one line at most the first can match while a match at line 0 opens no interior
boundary. Measured outside this repo on a consumer's export of a published
handbook: 83 / 414 / 828 `.html` files gave 0 plans, N documents with no boundary
and exit 2 at every point, and a coarser 145-document cut gave 145 of 145. The
same sections as markdown gave 828 of 828 plans -- so the instrument was fine and
`.html` was the one core-supported type that had never met a real document.

Block tags now open lines of their own, `h1`-`h6` carry the ATX marker for their
own level (not a flat `#`, which would hand `_ATX` three top-level boundaries
where the document declares one section and two subsections), `br` breaks the
line, and every other tag stays the word boundary it already was. The output
grammar is MARKDOWN and deliberately the same markdown the office rows reach the
proposer through, so no HTML-only heading grammar exists.

NOT via the converter: `.html` stays out of `_PANDOC_FORMATS` because routing it
there would add CVE-2025-51591 (SSRF via an iframe in HTML input), unpatched in
every converter version. The test asserting that exclusion is untouched and green.

The block set is wider than the five tags the corpus exercises, on purpose:
block versus inline is a property of HTML, not of one corpus, and a `div`-
structured page carries its prose in containers this corpus never uses.

Measured after, all with denominators:
- 828 of 828 plans, exit 0, `merged + coded rejections = 828; N = 828`; 3206
  concepts / 6015 md files, which is the markdown path's count EXACTLY -- 0.0 %
  deviation against the +/-2 % bar, and the same at 50 % (1651) and 10 % (343).
  The coarser 145-document cut goes 145 of 145 with no boundary to 145 plans /
  953 concepts.
- Text preservation as an EXACT invariant, not a percentage: strip the added ATX
  markers and the non-whitespace sequence is identical to the old extractor's for
  the same bytes. 828 of 828 files exact, character ratio 1.000000 against the
  >= 99.8 % bar. 7600 markers added; 31 141 lines produced where the old
  extractor produced 828, one per file.
- `_SKIP_TAGS` unchanged at {script, style}. Dropping nav/header/footer is a
  different change with a different guarantee and is not made here.
- No other file type moved, measured rather than argued: 0 of 86 K2 corpus files
  and 0 of 5 smoke-folder files are HTML, and the smoke bundle is byte-identical
  before and after (`diff -r` empty, 52 md / 26 concepts, 0 of 5 rejected).
  `okf project` stays byte-equal to `okf build` (`diff -r` empty).

Provenance moves with it: `source_units` routed `.html` through `_line_units`
already, but the table was trivial -- every offset resolved to line 1. The
numbers now mean something, and what they mean is a line of OUR extraction (a
BLOCK), never a line of the original markup.

`_EVIDENCE` gains a `.html` row at `measured`, chosen against the class
definitions: the files are a consumer's own export of a real published handbook,
produced for their ingestion and not to exercise this row. What the class does
not claim travels with it -- one product, one format, one publisher, and a
generator's cut.

One existing test changed because the behaviour changed, and it says so:
`test_html_text_via_htmlparser` asserted the collapsed form. The other two
(`test_html_skips_script_and_style`, `test_htm_is_an_html_alias`) were re-read
and hold unchanged -- the order expected three to move; only one did.

The corpus-wide invariant runs in the suite behind `OKF_HTML_CORPUS`: a corpus
path names a consumer's export and this repository is public.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 23:57:47 +02:00
docs docs(k3): round 10, and the fasit that held a working rule off the default [skip-docs] 2026-09-09 23:27:03 +02:00
examples feat(inbox): point every concept at the document it came from, with a locator per format 2026-09-08 14:39:24 +02:00
skills feat(readme,skill,cli): the first screen an agent reads, three modes, and one flag that made two builds 2026-09-09 18:12:02 +02:00
src/llm_ingestion_okf fix(extract): HTML collapsed to one line, so 828 of 828 sections had no boundary 2026-09-09 23:57:47 +02:00
tests fix(extract): HTML collapsed to one line, so 828 of 828 sections had no boundary 2026-09-09 23:57:47 +02:00
tools chore(ruff): the acceptance was whatever the default happened to be [skip-docs] 2026-09-09 23:15:17 +02:00
.gitignore chore(gitignore): keep .claude/projects/ local-only 2026-08-31 12:46:43 +02:00
CHANGELOG.md feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word 2026-09-09 23:05:45 +02:00
CLAUDE.md feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word 2026-09-09 23:05:45 +02:00
LICENSE feat: initial commit — repo scaffold and v1 scope 2026-07-16 10:12:59 +02:00
llms.txt feat(readme,skill,cli): the first screen an agent reads, three modes, and one flag that made two builds 2026-09-09 18:12:02 +02:00
pyproject.toml chore(ruff): the acceptance was whatever the default happened to be [skip-docs] 2026-09-09 23:15:17 +02:00
README.md chore(ruff): the acceptance was whatever the default happened to be [skip-docs] 2026-09-09 23:15:17 +02:00
SECURITY.md docs(security): add SECURITY.md with reporting contact and process 2026-08-16 21:15:10 +02:00
uv.lock chore(ruff): the acceptance was whatever the default happened to be [skip-docs] 2026-09-09 23:15:17 +02:00

llm-ingestion-okf

Turn a folder of documents (PDF, DOCX, XLSX, PPTX, MD) into a bundle a model can answer from with a source on every claim — offline, deterministic, no model call anywhere in the run path.

Install

Python 3.10+ and uv. One line:

uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.7.0"

Use it

okf project ~/my-documents   # folder in: bundle + Claude Code skill, in this directory
claude                       # start Claude Code here

Then ask in plain language. Three shapes of request work, and the skill states the rules for each:

  • a question — "hva er kravene til pris?"
  • a hypothesis — "stemmer det at leverandoeren baerer risikoen for grunnforhold?" Answered per premise as confirmed / refuted / undecidable-from-bundle.
  • a task whose answer is a document — "lag krav-pris.md med alle krav til pris, ett avsnitt per krav, med dokument og kravnummer." Every claim in the written file carries its source; a paragraph with no ground is written and marked, never dropped.

Everything below is detail: Consume in Claude Code for the same thing in steps and with several bundles at once, Build for the flags, Requirements for the pip fallback and the guard pairing.

What this library is

Status: phases 13 are implemented. Phase 1 (spec-based ingestion) covers manifest validation, the file/sql/http connectors, deterministic materialization, index generation, and the golden fixture suite under examples/. Phase 2 adds the bundle inbox (process_inbox) and external-bundle import (import_bundle), both against an injected persist gate, with llm_ingestion_okf.guard_adapter wiring that gate to the real guard (see below). Phase 3 makes the bundle contract configurable, so types, layers, frontmatter sets, index shape, and reserved-file policy are carried by a profile rather than by constants (see Upstream OKF versions). Binary extraction runs behind the optional [extract] extra: pdf through a PDF parser, and five office formats through a vendored document converter. Three of those five office rows are constructed rather than measured — see Binary extraction. Phase 4 (the Node half) is planned (see docs/plan/).

Install in detail

Neither this package nor the guard it depends on is on a package index yet, so both install by direct reference. With uv, one command resolves both:

uv pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.7.0"

uv resolves the guard on its own, because it reads the [tool.uv.sources] entry in the pyproject.toml of the tag it is installing, and v0.7.0 points that entry at llm-ingestion-guard v1.3.0. Use uv tool install instead of uv pip install when you want the okf command on PATH without an active virtualenv — that is the form the first screen shows.

With plain pip, the transitive git dependency does not resolve on its own — install the guard first, or installing this package fails with No matching distribution found for llm-ingestion-guard:

pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.3.0"
pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.7.0"

The guard tag is paired to the okf tag, not to this branch. v0.7.0 declares llm-ingestion-guard>=1.2,<2.0, which v1.3.0 satisfies; the pairing above is read off that tag's own [tool.uv.sources], not off this branch. Reading a pin off main and installing it against an older okf tag is the one combination that fails.

Earlier tags, as history

These are not install lines. They record what each earlier tag was, so a reader who meets one in an older document knows what they are looking at.

  • v0.7.0 — the current tag: okf project builds the bundle okf build builds (they were one flag apart before it), and the generated skill states the question, hypothesis and document-task modes with relative paths.
  • v0.6.0 — the first tag carrying the okf project, okf consume, okf check and okf skill subcommands.
  • v0.5.0a2 — a pre-release for the named OKF v0.2 pilot set only.
  • v0.4.0 — the last tag before the OKF v0.2 work; it declares llm-ingestion-guard>=0.2,<0.3, which only guard v0.2.0 satisfies.

No tag yet makes OKF v0.2 generally available. OKF_LATEST is unchanged and still points at DEFAULT; flipping that alias is the GA event and none of the tags above is it (see Upstream OKF versions).

Build

Installing the package installs one command. A folder of documents in, an OKF bundle out:

okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2

It walks the folder recursively, proposes a segmentation for each document with the mechanical rules, replays those proposals through the bundle inbox, writes the bundle and its log.md, and prints the run's numbers. Every proposal is marked PROPOSED and adjudicated: false — the command segments nothing a human has approved, and says so in the artifact.

The last line that matters is the conservation identity: merged + coded rejections == N, where N is the folder's file count read at run time. The run exits non-zero when it does not hold, and names the unaccounted files, so a pipeline cannot mistake a partial bundle for a complete one.

Flags worth knowing: --segments off ingests each document as one concept and asks for no root values; --plans-dir keeps the proposals instead of discarding them; --report writes the full report to a file as well as stdout. --ingested-at and --proposed-at default to 1970-01-01T00:00:00Z rather than the clock, so two builds of the same folder are byte-identical — a wall-clock default would break rebuild-equals-incremental for every caller who did not pass them.

The segmentation flags

Nine rules are reachable from okf build, and since 2026-09-11 all nine are ON by default--outline-run 3, --table-grid and --unit-fold since 2026-09-08, --drop-wrapped-outline and --outline-gate since 2026-09-09, --sheet-section-rows, --keep-table-heading and --first-span-from-zero since 2026-09-10, and --close-span-gaps since 2026-09-11 — each an operator decision, and each with an explicit opt-out: --outline-run 0, --no-table-grid, --no-unit-fold, --keep-wrapped-outline, --no-outline-gate, --no-sheet-section-rows, --no-keep-table-heading, --no-first-span-from-zero, --no-close-span-gaps. Passing all nine reproduces the pre-2026-09-08 bytes exactly, and the last three reproduce the pre-2026-09-10 bundle byte for byte — measured with diff -rq, 0 differences, not asserted. Each line below carries the number it was measured at, and nothing beyond it.

A tenth flag, --bold-title, is off by default. It is the rule for the type whose container declares nothing: rtf has no heading style, so an author's title is bold text, and the row was measured at 0 of 0 declared headings, 0 concepts and 1368 of 1368 characters in no segment. The grammar is markdown, not rtf — the converter already writes that title as **…** in the same output every office row produces — and it is gated by the principle --outline-gate already carries: recovery yields to declaration. Measured over 47 readable documents, three parameters were swept and one carried; false positives are 0 of the 31 documents that declare, and the rule reaches 2 of 39 corpus documents, both docx, 0 of 33 pdf and 0 of 2 xlsx. On the rtf fixture set it recovers 6 of 6 authored titles over N = 4 with 0 false titles and 0 of 1994 characters in no segment. On a five-document folder it moves 26 → 27 concepts, replacing a mechanical tabell-linje-30 with two named concepts.

A re-run is what this costs a consumer, and it is not a small one: on the 43-document reference corpus the default bundle goes from 629 concepts in 1108 files (the 2026-09-03 tree) to 492 in 944 after the 2026-09-08 move, to 425 in 810 after the 2026-09-09 one and to 436 in 832 after the 2026-09-10 one (digest 8dff8a8e6c15d2f7…). On a five-document folder the last move is 15 concepts in 30 files → 26 in 52. The proposer's own defaults (tools/okf_propose_segments.py) did NOT move, so every published reproduction block still runs as written.

What the 2026-09-09 move had to clear, stated because it is the bar every later move is held to: the twelve-position reference improves (pdf 2 of 8 → 7 of 8, the sheet 5 of 12 → 10 of 12, docx unchanged at 3 of 3) AND hit@8 on a K2 bundle built with it holds 5 of 6 at ranks 1,1,1,1,1,, with no row losing rank 1. A configuration that improved the reference and cost a rank was measured in the same session and did NOT ship; see docs/2026-09-09-k3-runde6-outline-gaten-og-prioren.md § 9.1.

flag what it does measured
--outline-run N (default 3) also propose a boundary where the document's own bare-integer numbering sustains an ascending run of at least N; 0 is this arm's opt-out a tender PDF whose headings are bare integers: no boundary at 0, 9 concepts at 3, against a reference of 9
--table-grid (on by default; opt out with --no-table-grid) a pandoc grid-table rule line no longer closes an open table block, so one grid table is one concept a .docx experience list: 21 → 6 concepts
--unit-fold (on by default; opt out with --no-unit-fold) discard a contents-list run, fold a deeper heading into its parent, fold a table into the shorter heading that introduces it. Adds no boundary, so it can only reduce a plan on a 12-document sample scored against an operator's unit worksheet: 5 of 12 match — but that figure was measured with --table-grid ON, and the shipped default does not include it. Measured without it the same sample scores 2 of 12, docx 0 of 3, because the fold's table clause has no joined table to fold
--keep-table-heading (on by default since 2026-09-10; opt out with --no-keep-table-heading) keep a heading whose body is empty only because a table opens under it, and absorb that table into its span the two spreadsheets in that corpus, and 0 of 32 pdf and 0 of 5 docx: the concept count does not move (1 → 1), its first byte does — the concept gains the heading line it was missing
--sheet-section-rows (on by default since 2026-09-10; opt out with --no-sheet-section-rows) cut an open table block at the rows that label its sections — a run of at least three rows whose first cell is a bare numeric label. The opposite direction from --table-grid, which decides how far a block extends a tender price sheet whose whole body is one table block: 1 → 12 concepts, against a reference of 11 cost groups plus the sheet's preamble. Whole corpus: 1 of 39 readable documents changes, 0 of 32 pdf, 0 of 5 docx, 1 of 2 xlsx. It reached 11 of 12 on the reference two rounds before it shipped, and was held back both times by a RETRIEVAL cost that turned out not to be its own: on a bundle built with it the gold document splits 1 → 12 concepts and row 1 of the hit@8 set fell rank 1 → 2. The repair is on the reading side (--tie-shared-rank, now the default), and with it in place the sheet reaches 11 of 12 with hit@8 holding 5 of 6 at ranks 1,1,1,1,1,
--drop-wrapped-outline (on by default since 2026-09-09; opt out with --keep-wrapped-outline) do not admit an --outline-run candidate whose line continues onto the next one. Judges recovered candidates only, never a heading the document declares quoted regulation text, whose numbered paragraphs match the outline grammar exactly: 4 → 1 concepts, the reference. Whole corpus: 5 of 39, all pdf; on the 12-document sample 8 of 34 outline candidates wrap, and none of the 26 the operator kept. On the reference it carries pdf from 5 of 8 to 6 of 8 together with the gate below, and neither reaches 7 of 8 without the other
--outline-gate (on by default since 2026-09-09; opt out with --no-outline-gate) admit --outline-run's RECOVERED headings only where the document declares none of its own, plus any one recovered heading whose span covers OUTLINE_SHARE (0.20) of the text. Applied at admission, before spans are closed, so the text a removed mark opened is carried by the mark above it rather than lost on the 12-document sample: pdf 2 of 8 → 5 of 8 alone and 7 of 8 with the rule above, docx unchanged at 3 of 3. Whole corpus: it fires on 25 of 39 readable documents, changes the plan in 15 of 39, and removes 64 of 485 proposed entries. No plan disappears (32 → 32)
--first-span-from-zero (on by default since 2026-09-10; opt out with --no-first-span-from-zero) start the first concept at character 0, so the text above it belongs to a segment instead of to none. Adds no boundary and removes none Measured over the 39-document corpus, the default before this rule left 207 435 characters — 11.92 % — in no segment at all: 163 804 above the first entry (in 32 of the 32 documents that get a plan), 26 041 between entries and 17 590 after the last. This rule closes the first part entirely, 79 % of the whole, leaving 43 631 characters (2.51 %) over 8 of 32 documents with two named mechanisms of their own. It adds no boundary and the K2 concept count is identical with and without it (425 = 425); on the 12-position reference it changes not one cell, and hit@8 on a K2 bundle built with it holds 5 of 6 at ranks 1,1,1,1,1, under both tie-breaks
--close-span-gaps (on by default since 2026-09-11; opt out with --no-close-span-gaps) close a concept's span against the next SURVIVING concept, and the last against the end of the text. Three steps remove a candidate AFTER its neighbour's span was already closed against it — the orphan check, and fold_units clause 1 both between entries and on the last run — and the removed mark's text then belongs to no segment. Adds no boundary and removes none; only spans' ends move It closes the whole remainder the rule above left: 43 631 characters, 2.51 % of the corpus over 8 of the 32 documents with a plan, to 0 — both the 26 041 between entries and the 17 590 after the last. Decomposed: orphan check 18 527 over 15 of 39 documents, clause 1 7 514 between entries, clause 1 on the last run all 17 590 of the tail (with unit_fold=False the corpus tail gap is 0). The entry count is identical (429 = 429 on the corpus, 436 = 436 concepts on K2, 52 = 52 md on a five-document folder); on the 12-position reference it changes not one cell (11 of 12 under `

They compose, and the order above is the order they apply in. Measured on a five-document tender folder (2 pdf, 2 docx, 1 xlsx), concepts per document:

document default --outline-run 3 + --table-grid + --unit-fold + --keep-table-heading + --sheet-section-rows --drop-wrapped-outline
tender PDF, technical requirements 1 (no boundary) 9 9 9 9 9
tender PDF, technical layout 1 (no boundary) 1 1 1 1 1
price sheet .xlsx 1 1 1 1 1 12
experience list .docx 21 21 6 3 3 3
agreement .docx 2 2 1 1 1 1
markdown files in the bundle 31 49 33 30 30 52

Every column merged 5 of 5 with 0 rejections. The reference for the first row is 9, so the default is a full arm behind what the proposer can do on that document — which is a statement about the default, not a licence to change it here.

The two PDF reader flags

Separate from the six above, and they sit before every one of them: a segmentation flag changes how the proposer cuts a text, these change what the text says. All three are off by default.

flag what it does measured
--pdf-headings font a PDF carries no heading markup, so one is inferred from typography — a line whose dominant font size is above the document's character-weighted median AND whose dominant font name says bold — and emitted as an ATX heading in the same markdown the office path produces, so the existing heading rule reads it on a tender PDF: 9 of 9 numbered chapters found, plus 4 extra candidates. Whole corpus: 25 of 32 pdf change, 0 of 5 docx, 0 of 2 xlsx. Off by measurement: against the operator's unit worksheet it takes pdf from 2 of 8 to 0 of 8, losing two exact matches, because on those documents the outline rule already found the chapters and a second heading source can only add
--pdf-headings font-reserve the same typographic rule, applied ONLY to a document whose own numbering the outline arm finds no run of — typography as a second heading source where there is no first one, never on top of one. Three values of one option (none, font, font-reserve), so no caller can ask for two at once reaches 4 of 39 readable corpus documents (10 of 32 pdf admit no outline run; 4 of those render differently at all). Off by measurement, and the measurement is that it changes nothing measurable: on the operator's twelve-position unit worksheet it alters not one cell — the five positions where it fires are one PDF whose glyphs carry no ToUnicode mapping and four office documents the PDF reader never touches. The position it was built for numbers its own chapters, so the reserve is silent there by construction
--ocr read a PDF page as an image when its own text never arrived: the page extracts empty, or as (cid:N) placeholder codes at or above 10 % of its characters. Needs the optional ocr group on the one corpus document with the failure: 95.07 % → 0 % cid, 44 → 2561 words of four or more letters, 17 → 18 pages with text, 3.6 s/page. Whole corpus: 16 of 834 pages qualify, in 1 of 32 files
pip install "llm-ingestion-okf[extract,ocr]"

Without that group --ocr is a typed refusal (extractor_ocr_group_missing) per file, never a crash, and the corpus run still reports merged + coded rejections == N. OCR text is a model's reading of an image: it is reproducible against the model version and rendering resolution it was produced with, and no dependency pin can promise more. The full measurement, including the per-page distribution the 10 % threshold was read off, is docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md.

Measured 2026-09-08 on a 43-file corpus (33 pdf, 5 docx, 2 xlsx, and three files no reader accepts), one okf build invocation replacing the shell loop over tools/ that produced the same corpus's bundle on 2026-09-03:

figure value
N (folder file count, computed) 43
merged 39/43
coded rejections 4/43 (extractor_unknown 3, extractor_empty_pdf 1)
K1b 39 + 4 = 43 = N, exit 0
files written 1108
identical to the 2026-09-03 bundle 1104/1108
wall time 842.82 s total, 19.600 s per file (re-measured 2026-09-08)

The six files that differ are all in the corpus's two spreadsheet documents, and they are the change reported in docs/2026-09-08-prisform-og-loggen-k2.md: a spreadsheet's tables are now written as pipe tables, so each row is one line with its cells delimited rather than padded out to the widest cell in the column. Two concept files are renamed by it, two are removed under their old names, and the two documents' own index.md follow. The root index.md is identical to the stored one again, because this library no longer links the bundle's log.md from it. The command's own byte-identity test compares it against the two scripts at the current commit, where the two agree over the whole tree.

Consume

The other direction: a bundle plus one question in, one bounded, contract-shaped payload out.

python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json

tools/okf_consume.py is the pre-pass docs/consumption-contract.md § 1 defines — the deterministic program that reads the bundle, ranks its concepts, cuts them to a bounded set and emits one payload. It decides nothing about the question; the skill that reads the payload does the judgement. It calls no model, opens no socket, imports nothing outside the standard library and this package, and takes no clock: the same bundle bytes and the same (question, k, limit, cost_vocabulary, reserve_top_rank, rarity_weight) produce byte-identical output.

--cost-vocabulary is off by default and widens one question class: it lets a declared list of cost/price/quantity terms bridge a question and a document that name money with different words. The gate is the question — one naming no such term gets byte-identical bytes either way — and what it does and does not close is measured in docs/2026-09-08-blindsone-below-k-k2.md.

--reserve-top-rank is off by default and answers a different objection: the budget is packed by an exact knapsack, which maximises a SUM of scores and therefore has no opinion about rank, so a top-ranked excerpt costing a large share of the budget is out-summed by many small ones. Measured, that made --k a dial that could EVICT the concept a question was asked about. The flag gives rank one its bytes before the pack runs — after the over_budget_alone pre-exclusion, never before — and the payload then declares budget.reserved. On a 629-concept corpus it changed the delivered set in 2 of 24 measured combinations, both of them that eviction: docs/2026-09-08-blindsone-laas2-budsjett-k2.md.

--rarity-weight is off by default and weights each lexical hit by log(N/df) over the bundle's own concepts instead of counting it as one, so a requirement number is not worth what a common verb is worth. The default being off is a measurement rather than a preference: on four corpora it took one gold concept from withheld to delivered and a priced sheet from candidate rank 10 to 2, left one gold rank unmoved, and cost another seven rank positions — because the four-character prefix matcher makes a unique identifier read as 135-of-446 common on that bundle. Where it cannot help is decomposed rather than guessed: RRF fuses RANKS, so a weight moves nothing on a signal the gold already leads. docs/2026-09-08-sjeldenhetsvekt.md. Its published figures were measured under the pre-2026-09-10 tie-break and are not re-measured.

--stem-prefix is on by default since 2026-09-09 (opt out with --no-stem-prefix), and like --tie-shared-rank below it alters a payload with no bundle changing. MIN_SHARED_PREFIX = 4 exists for Norwegian compounding, and it also matches four characters that are not a stem: measured on the pinned 453-concept bundle with the control run first, under occurs 79 times by equality and matches 172 concepts by prefix, while bilateral occurs 0 times and matched 400 of 453 through bilag, and standhaftig 0 and 219 through standard. Three repairs were measured and all three failed on the same row — a longer floor (58), a coverage share (0.50.8) and a long-words-only floor (≥ 8) each cost row 1 its rank on the default bundle and the whole row on Arm B, because row 1's token prisene reaches its gold document through pris|sammenstilling on four characters. The rule that works asks whether the shared prefix is a WORD the bundle uses: bilateral 400 → 0 and 512 → 0, standhaftig 219 → 56 and 235 → 33, every hit@8 row keeping rank 1 on both bundles. undersjøisk stops at 162 because under is a word here — a genuine Norwegian morpheme, so the residual is a different answer and not a ceiling.

--tie-shared-rank is on by default since 2026-09-10 (opt out with --no-tie-shared-rank), and it is the one change in this library that alters a payload with no bundle changing — a consumer pinned to the previous excerpt order needs the opt-out. RRF emits a rank for every concept in every signal, including a signal that scored them all the same, and the declared tie-break then orders that group by concept_id; the fusion reads alphabetical order as if it were a measurement. Shared ranks make a signal that separates nothing contribute the same constant to each concept in the group. What it buys is general rather than cosmetic: a document the segmenter splits from 1 concept into 12 fills that signal's whole top tie group with its own concepts, so the one that leads the body signal takes position 11 instead of 1 and the document loses fused rank 1 to a single-concept competitor leading nothing — the fusion was punishing fine-graining for being fine-grained, which put the segmentation side and the retrieval side in competition over one number.

It shipped OFF on 2026-09-08 because hit@8 fell 5 of 6 to 4 of 6, and that figure is real and conditional: swept over 2 document-prior exponents x 3 bundles x 6 rows, the lost row is lost only at exponent 1.0. The exponent moved to 0.5 on 2026-09-09 for an unrelated reason, correctly reported as moving no hit@8 row, and nobody measured the pair — so a rule sat behind a published number that had stopped being true in the same commit. A flag's "off by measurement" is a measurement of a configuration, not a property of the flag. docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md.

It emits the § 8 shape — contract, bundle (bundle_id plus a sha256-tree: content identity), budget (unit, instrument, limit, spent and a validated known-positive), denominators, excerpts and withheld — and every withheld concept names the rule that dropped it, from a closed set of six.

Every excerpt carries the concept's title, and — when the producer wrote them — req_number, the SPEC § 5.1 address sources, and every top-level source_* key, by prefix rather than by allowlist: a fixed list names the locators its author thought of, and one real bundle locates by source_element_id on 269 of its 274 concepts. A key the producer did not write stays absent rather than arriving empty, and an address this reader cannot decode is named (sources_unreadable) rather than dropped into the same silence. The reason is a measurement: with concept_id and body text alone, a delivered gold concept at rank 1 still left the answer unable to name the document it was quoting. considered == withheld + delivered closes by construction, and the payload is refused rather than reported when it does not.

Three exit codes, not two: 0 a payload was written, 1 the run happened and refused (the budget admitted none of the concepts that answered the question, or an asserted --ref contradicted the bytes), 2 the run did not happen. Collapsing 2 into 1 would report an unread bundle as a failed cut. --ref is an assertion, never an override — the identity is always computed from the bytes, because labelling a payload with an identity its bytes do not have is the one thing § 3.3 exists to prevent.

Check any payload against the skill that will read it:

python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json

skills/okf-consume/ is the first instantiated consumption skill: a filled copy of skills/okf-consume-template/ naming this pre-pass, with every per-corpus hole replaced by a measured value. Measured 2026-09-07 on a 629-concept bundle, hit@8 was 5 of 6 questions at rank 1 against a chance baseline of 1.35 of 6 — with one control that failed, and both are in docs/2026-09-07-okf-konsumskill-maaling.md with the honesty limits stated.

Consume in Claude Code

A folder of documents to an answer a model can cite, in three lines. You do not need this repository — the first line installs the command, the second builds the bundle and writes a skill beside it, the third asks.

uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.7.0"
okf project ~/my-documents
claude

okf project writes the bundle to .okf/<id>/ and a skill to .claude/skills/<id>-consume/ in the current directory, then prints what it read, what it wrote, and which documents a question cannot reach. Start claude in that directory and ask in plain language; the generated skill runs the pre-pass and the contract check itself and marks every claim with its source.

Running non-interactively. In print mode the skill needs its tools named explicitly, or the model answers without ever reading the bundle and marks every premise undecidable-from-bundle:

claude -p --allowedTools=Bash,Read,Grep,Glob "<the question>"

Add Write,Edit for the mode that produces a document. --permission-mode acceptEdits alone does not do it — measured 2026-09-09 over four runs, the okf consume call is refused without the explicit tool list. The interactive claude above needs none of this.

<id> is the folder's name reduced to [a-z0-9-]. Run it once per folder with --id <name> to have several bundles reachable at once — each skill carries its own bundle_id, which is what lets a model pick between them. --out <dir> puts the project somewhere other than the current directory.

Measured 2026-09-09 from a fresh uv tool install with this repository nowhere on the path: 5 documents in, 26 concepts out, a skill carrying 0 paths into any checkout, and okf check conformant on its own payload (15 rules, 0 findings). The 2026-09-08 run of the same measurement reported 15 concepts, and that number was the defect rather than the result: okf project was calling build() as a function and reading its signature's defaults, which disagreed with argparse's on two flags. Two tests now hold the two default sets equal. Before that day the same result took a PYTHONPATH, a snapshot of a clone, and a generated skill that named that clone by absolute path on four lines — so it could not be moved, shared, or run by anyone else.

The same thing in steps, if you want to see the payload

okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
okf skill ./bundle --out ./project/.claude/skills/my-bundle-consume
okf consume ./bundle --question "your question" --out /tmp/payload.json
okf check --skill ./project/.claude/skills/my-bundle-consume/SKILL.md --payload /tmp/payload.json

A bundle you only have read access to is fine — the generator only reads it.

The honest limits

Measured on four questions across two bundles, which is a demonstration and not a hit rate. The ranking is lexical, and one of the four found a topic the bundle does cover and did not rank it into the cut — the skill then said so with its denominator instead of answering, which is the behaviour the contract asks for, but a miss is still a miss. docs/2026-09-08-claude-code-skill-vilkaarlig-bundle.md has the runs.

Two more things the summary tells you and this paragraph will not repeat: a document that landed whole (no heading, table or numbered outline to cut it on) comes back as one excerpt, which the budget often refuses and which often does not carry the answer at the place you asked about; and a document that is in the folder but not in the bundle cannot be quoted at all. Both cases are answered [sourced-not-sufficient], and okf project names the documents.

The skill that runs this for you

skills/okf-prosjekt/ in this repository is a Claude Code skill (Norwegian) that wraps the command above: it takes a folder, runs okf project, and reads the summary back. Install it for your user account after cloning:

mkdir -p ~/.claude/skills && cp -R skills/okf-prosjekt ~/.claude/skills/

Implemented scope (v1)

The library provides three entry points for getting content into an OKF bundle:

  1. Spec-based ingestion. An implementation of the normative ingest specification owned by portfolio-optimiser-commons: manifest → file/sql/http connector → deterministic materialization of ingest-{id}.md concept files → index generation. Zero model calls in the run path; output is reproducible byte-for-byte against golden fixtures.

  2. Bundle inbox. A drop directory where common file types are converted to OKF concept files. All file-type→text extraction lives in this library: md, txt, csv, json, and html are handled by the stdlib core; pdf and the five office formats (docx, xlsx, pptx, odt, rtf) require the optional [extract] extra and are rejected fail-fast without it. Extracted text passes the security gate before anything is persisted. The drop directory is walked recursively, in sorted relative-path order: a file at any depth is ingested and records its path relative to the inbox root as its source_file, while dot-directories and a bundle directory sitting inside the inbox are skipped with a reported code.

    Under the segmented v0.2 profile a concept also points back at the document it was extracted from, so an agent citing it can open the original at the right place: sources: [{ resource, title }] in the spec's own §5.1 form, where resource is the inbox-relative path, plus a locator per format — source_pages for a PDF, source_sheet and source_rows for a spreadsheet, source_lines otherwise. The locator keys are this library's own, because §5.1 has no field for a place within a resource; the line numbers index the extracted text and say so. Measurements: docs/2026-09-08-proveniens-k2.md.

  1. External bundle import. Import and merge of third-party OKF bundles: each concept is assessed via the security gate, and only concepts that pass are merged, materialized, and linked into the index.

Boundary: security is delegated

Security is owned by the sibling package llm-ingestion-guard (pinned >=1.2,<2.0). The division is strict:

  • guard answers "is this content safe to persist?" — scan, sanitize, quarantine, fail-secure, provenance stamping.
  • this library does the plumbing — connect a source, materialize a deterministic OKF bundle, generate the index.

No security functionality is reimplemented here.

What is gated today: read this before trusting a door

  • Door A (materialize_bundle) is ungated. It calls nothing before writing to disk and writes what it is given. A caller materializing untrusted content is responsible for gating it.
  • Doors B and C (process_inbox, import_bundle) gate through an adapter you pass in. Each takes a gate argument; the flow hands it the content and obeys the verdict, refusing to persist anything that does not clear the guard's non-blocking floor — including a disposition it does not recognise, and (at Door C) a concept the gate returned no verdict for. What it cannot do is check that your adapter is a real guard: a permissive stub approves everything, and the flow will believe it.

llm_ingestion_okf.guard_adapter is the adapter over the real guard, and the only module here that imports it — importing the package itself does not:

from llm_ingestion_okf import process_inbox
from llm_ingestion_okf.guard_adapter import inbox_gate

result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
                       okf_type="reference", gate=inbox_gate)

Two properties of that adapter are worth knowing before you rely on it. It screens the exact bytes it persists — the guard's prepare_input bookend prepares text for a model call, which this library never makes, so only screen_output is used and the screened string is the written string. And it refuses rather than repairs: a file carrying an invisible zero-width or bidi character is rejected, not silently stripped and written. Door B screens under the untrusted-upload policy, so any finding at all is held back rather than persisted.

This section is stated plainly because earlier wording ("calls the guard at every persist gate") described the intended end state in the present tense, and a consumer reasonably read it as safe-by-default.

Roadmap

The library is built in four phases so that every known OKF surface in the ecosystem is eventually covered. Each phase has a detailed plan with verification criteria:

  1. Spec-based ingestion (Python) with byte-exact golden fixtures — plan.
  2. Bundle inbox and external-bundle import (Python), guard-gated — plan.
  3. Configurable bundle contract (types, layers, frontmatter sets, index shape, and reserved-file policy as configuration), enabling stricter bundle profiles such as strict-v1plan.
  4. A node/ half: a zero-dependency Node/ESM package (importable and CLI-invokable, vendored per consumer) providing bundle checking, index generation, inbox processing, and document conversion for the OKF second-brain plugin ecosystem. The Python and Node halves share the OKF contract and fixture suite, not code — plan.

Upstream OKF versions

The library targets the current latest version of Google's OKF. Support is additive — a new upstream version arrives as a new profile, never as a migration of an existing one — so an upstream release does not change the bytes an existing profile emits.

That guarantee is about upstream, and one profile tracks a second contract as well. DEFAULT states the ingest-spec owned by portfolio-optimiser-commons, so when they change that spec, DEFAULT follows them. It happened on 2026-08-09: generated moved from true to { by: process:okf-ingest, at: <ingested_at> }, one changed line per generated file. Upgrading across it costs a re-run and nothing more — a profile still recognises bundles stamped by earlier versions, so re-running writes in place instead of refusing. DEFAULT remains OKF v0.1 on every axis upstream owns.

Profile Contract Status
DEFAULT commons' ingest-spec layer (OKF v0.1 semantics) stable
STRICT_V1 a consumer's ratified v0.1 contract stable
OKF_V0_2 OKF v0.2 provisional, pre-release only
STRUCTURED_V1 DEFAULT plus a faceted, derived index stable
OKF_LATEST alias for the latest version supported as stable currently DEFAULT

STRUCTURED_V1 is DEFAULT in every respect but the index. Under it, Door B derives each dropped document's title, number, hierarchy and cross-references, writes them into the concept's own frontmatter, and carries them into the index entry — so a consumer can reason over the bundle rather than only look things up in it. Every inferred field is named in a derived list, because an unmarked heuristic is worse than no heuristic: the consumer cannot know when to doubt it. A pointer to a document not dropped yet is rendered N200? rather than omitted, since a bundle is built up over several drops and an absence that leaves no trace is the dangerous kind. Carrying the metadata costs index characters — roughly 3x to 6x the flat index, depending on how many facets the profile names — and the facet key set is the dial. Design record and measurements: docs/plan/structure-derivation.md.

OKF_V0_2 ships first as a pre-release to a named pilot set and may change on their feedback without a deprecation cycle. Pin the versioned constant rather than OKF_LATEST unless you have explicitly opted into tracking; OKF_LATEST moves at general availability, which is a deliberate release event rather than a side effect of an upgrade.

Selecting a profile is keyword-only, so existing call sites are unaffected:

materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)

A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and puts the declaration in the bundle-root index.md's frontmatter block. The profile names the key; the caller supplies the value, because that value tracks the upstream version and is not this library's to decide:

materialize_bundle(
    manifest, bundle_dir, ingested_at,
    profile=OKF_V0_2,
    root_frontmatter_values={"okf_version": "0.2"},
)

Omit the argument and no frontmatter block is written. Offering a key the profile does not name is refused before anything is written to disk.

Attested computations (v0.2 §10)

OKF_V0_2 supports the Attested Computation type as a format: its five contract fields — runtime, parameters, computation, executor, attester — are emitted in canonical position, judged, and round-tripped. runtime is required for that type and for no other, which the profile expresses through FrontmatterSchema.required_by_type; a type the mapping does not name carries no extra requirement, because §14 forbids a consumer to reject on an unknown type.

Nothing here executes a computation or checks an attestation. Upstream defers the receipt and verdict wire formats, so there is no contract to implement, and the question an attestation answers — was this value produced the sanctioned way — is not this library's. It re-enters scope when upstream specifies the protocol.

On the import side, a third-party concept may name an executor or attester resource pointing at executable code. Door C imports the pointer and never the code — it writes concepts verbatim and skips every non-.md file — so such a reference may not resolve, or may resolve to a file the destination tree already holds under that path. Each one is reported in ImportResult.unverified_references; the concept still merges, because §14 forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer to surface rather than silently drop. The report names the pointer key, not the resource it points at: recovering the resource needs the structured reader.

One limit worth knowing before you write such a concept: §10.2 presents executor and attester as nested block mappings, and this library's frontmatter parser is line-oriented. It reads inline flow mappings (executor: { resource: …, receipt: [ … ] }) as opaque values that round-trip unchanged, but it cannot read the block form — two block mappings that both carry a resource collapse into one namespace and the first is lost. Write the flow form; both are valid YAML, and a real YAML consumer recovers the same structure from either.

Non-goals

  • Verdict/feedback machinery from the method specification (stays in the consuming repositories).
  • Embedding- or retrieval-layer functionality.
  • Security functionality, in either runtime — that is always llm-ingestion-guard's domain.

Requirements

Python 3.10+, and exactly one runtime dependency — the security boundary, llm-ingestion-guard>=1.2,<2.0. Everything else is stdlib. The commands are under Install; what follows is why they look the way they do.

From a checkout, the test suite runs with:

.venv/bin/python -m pytest

The suite is the verification surface for everything above: 1515 tests, run on 2026-09-09 against this branch with the [extract] extra installed. (The figure published here until that day was 596, measured 2026-08-21 and never updated as the suite grew — a count is a measurement with a date on it.) Without the extra the same suite skips the tests covering the parser path; that split was last counted on 2026-08-21 as 589 passed and 7 skipped and has not been re-measured since. The tests holding the fail-fast rejection for an uninstalled extra run in both. The suite is not shipped in an installed distribution — tests/ lives at the repository root, so this command needs a clone rather than a pip install.

The lint acceptance is BOTH of ruff's gates, and the rule set is declared. ruff check src tests tools and ruff format --check ., both named in a report with the version they ran under. Neither was true before 2026-09-09: the formatter gate was not in the acceptance and had gone red unseen, and [tool.ruff] set only line-length and target-version, so the acceptance was whatever ruff's default happened to be — which is why the tree read green only as long as uv.lock froze ruff at 0.15.22. Under 0.16.6 the same untouched code reported 148 findings, all of them new rules rather than new defects, because 0.16 widened the default set to whole families. select is now written down (E4, E7, E9, F, I, RUF100), the dev pin is ruff>=0.16.6,<0.17, and the wider families are a separate decision with 148 as its starting number. S is measured out rather than assumed out: it reports 2657 S101 on a suite whose every assertion is an assert. Add --extra extract to the sync or mypy src cannot find pdfplumber.

Markdown is excluded from ruff format. ruff 0.16 formats fenced Python inside markdown, and the two files it would change here are records rather than source — a README call example and a published measurement's quotation of COST_VOCABULARY as it stood when that measurement was taken. Reformatting a quotation makes it stop being one.

A git URL is a PEP 508 direct reference and pins one exact tag, so it is an install-time channel, not the pin: the range above stays the declared dependency — a wheel built from this branch carries Requires-Dist: llm-ingestion-guard<2.0,>=1.2, measured 2026-08-23 — and resolves normally once the package index exists. A wheel built from a tag carries that tag's range instead, which is why the install commands pair tag with tag.

Binary extraction

The optional [extract] extra ships two things: pdfplumber (MIT) for pdf, and pypandoc-binary for five office formats. It is opt-in because it pulls binary wheels, which the default install must never do — the single runtime dependency rule covers the default install and this extra sits outside it.

The converter binary travels inside the wheel and is resolved by path rather than found on PATH, with its version asserted against a pin. A host carrying a different converter is refused, not silently used: extraction is deterministic within a converter version and not across one.

Format Reader Evidence
pdf pdfplumber measured
docx converter measured
xlsx converter measured
pptx converter constructed
odt converter constructed
rtf converter constructed

constructed means what it says, and it is a weaker word than measured on purpose. The corpus this work was measured on contains zero pptx, odt and rtf files. Until 2026-09-09 those three rows were unmeasured — they worked by construction and had never been checked against a document anyone wrote. They have now each been put through end to end on a hand-built document with a hand-written fasit, which is more than nothing and is not a corpus:

  • odt1 of 1 declared headings recovered, 1 concept, 0 characters in no segment. N = 1 document.
  • pptx2 of 2 declared slide titles recovered on a deck that declares them (a real <p:ph type="title"/> placeholder); 0 of 2 on a deck that does not, where the converter writes Slide 1 / Slide 2 because it has no title to use. That is the converter naming an unnamed slide, not a segmentation failure. N = 2 decks.
  • rtf0 declared headings, because the container has no heading style and the author's title is bold text. The proposer therefore proposes nothing, and the document reaches the bundle inbox as one whole concept: content preserved, structure zero. N = 1 document. This is the one open finding of the three.

They are not known to be broken; a single constructed document is not a denominator, and the distinction is the point.

What stays out. .doc (Word 97) is not supported — the converter does not read it. Rastered or scanned PDFs are refused rather than persisted as empty concepts, because this library does not do OCR. Drawn content — figures, diagrams, shapes — does not survive extraction in any format here, and every extraction says so with a warning. Structured table recovery is out of scope.

Request it by appending [extract] to the package name in whichever install command from Install you are using — this package is not on an index, so a bare pip install 'llm-ingestion-okf[extract]' does not work today, and the error message naming that command is written for the day it does. The extra is unreleased: it reaches a consumer through a tag that contains it, and no such tag exists yet.

Two properties of the extra are worth knowing before depending on its output:

  • Extracted text is pinned to an exact parser version. pdfplumber pins pdfminer.six==20260107 exactly, and pdfminer.six ships date-stamped releases with no stability contract. Extraction is deterministic within a parser version and not guaranteed across one, so a golden fixture built on extracted PDF text is a fixture migration away from any parser upgrade.
  • Text extraction recovers text, and nothing that is drawn. Figures, diagrams and images have no text to recover — only their captions survive — so a bundle built from drawn documents is incomplete by construction. The library says so itself: every pdf extraction emits an ExtractionWarning. Structured table recovery is separately out of scope; PDFs enter as prose.

The planned Node half targets Node/ESM with zero npm dependencies.

License

MIT — see LICENSE.