C1. `okf consume` and MCP's `okf_ask` now rank with BM25 (`bm25.py`) instead
of the three-signal fusion. Two signals, fused by reciprocal rank:
- passage: every body cut into 500-character windows every 250, a concept
scored by its BEST window -- a narrow question is answered in one place;
- field: title three times, the id path and source name twice, then the body
-- a broad question is answered by what a section is called.
The document prior and the rarity weight are gone from the default: the first
favoured big documents full of common words, the second gave its largest
weight to a word the collection does not hold. Under BM25 such a word weighs
exactly zero. A signal that scores a concept zero adds nothing to it, and ties
share a rank, so alphabetical order lifts nothing either.
Three rules carried over from the fusion, each with its own test, because the
suite showed what BM25 alone lost:
- a directory every concept shares is not read (K3-20's defect, one signal on);
- a number a section is known by (`4.2`, `10.2-2`) is kept as one token, or
a question naming a section by its number matches nothing in it;
- a question word the collection does NOT hold is read as the collection's
words it shares a leading word with (`consume.tokens_match`) -- Norwegian
inflection and compounding -- at that word's idf, never at its own.
The lookup and title-covered partitions are shared with the fusion
(`_partitioned`). `ranking="fusion"` / `--ranking fusion` keeps the old order
reachable; `--cost-vocabulary` and `--rarity-weight` widen only the fusion and
are refused with the default (`ranking_flag_conflict`) rather than ignored.
C3. A concept longer than `PASSAGE_CHARS` (4 000) is delivered as the span
around its best window, snapped to whole lines, under the nearest heading
above it, with `[...]` where text was left out. `passage: {start, end, of}`
says so, `text_sha256` covers what was delivered, and `sha256` stays the
file's, so the whole can be fetched by `concept_id`. 4 000 because eight
excerpts of it stay far under a tool response's limit even with several
sub-questions merged, while a 500-character window keeps 3 500 characters of
surroundings. The budget pays for the passage, not the file.
Tests moved with the default, each stated rather than silenced:
- fusion-mechanism tests (cost vocabulary, rarity weight, reservation, shared
rank, the reference-bundle pins) ask for `ranking="fusion"`, the order they
were measured on; the BM25 reading of the reference bundle is a separate
measurement, kept in local state;
- the retrieval gate still measures the shipped default. Row 1 holds. Four of
its premises were built against the fusion (a concept forced below k that
BM25 now delivers, a quota that no longer decides, mutants patching fusion
code) and are `xfail(strict=True)` until the fixtures are re-measured;
- the shipped example payload is regenerated; the shipped skill is unchanged.
README's Consume section and CLAUDE.md state the new default and that the
flags described after it belong to the fusion.
The search gate's table for this commit is kept in local state: the question
sets belong to a consumer whose content does not go on a public mirror.
Suite on a clean tree after `git add`: 2390 passed, 2 skipped, 4 xfailed.
ruff, ruff format, mypy --strict clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1789 lines
113 KiB
Markdown
1789 lines
113 KiB
Markdown
# llm-ingestion-okf
|
||
|
||
Turn a folder of documents into a bundle a model can answer from **with a
|
||
source on every claim** — offline, deterministic, no model call anywhere in the
|
||
run path. Thirteen file types are read; [Supported file
|
||
types](#supported-file-types) lists each one with the evidence behind it.
|
||
|
||
## Install
|
||
|
||
Python 3.10+ and [uv](https://docs.astral.sh/uv/). One line:
|
||
|
||
```sh
|
||
uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v1.0.0"
|
||
```
|
||
|
||
## Use it
|
||
|
||
```sh
|
||
okf project ~/my-documents # folder in: bundle + Claude Code skill, in this directory
|
||
claude # start Claude Code here
|
||
```
|
||
|
||
**Or connect every bundle at once.** Register the server once, on user scope,
|
||
and every project you open can ask any bundle under that directory — no skill
|
||
to install per project, and nothing to regenerate when a bundle is rebuilt.
|
||
You run this line; `okf` never starts Claude Code:
|
||
|
||
```sh
|
||
claude mcp add --scope user okf -- okf mcp --root ~/okf
|
||
```
|
||
|
||
Subagents inherit MCP tools and do not inherit skills, so the server is also
|
||
the only way the same working method reaches an arm running below the main
|
||
thread.
|
||
|
||
Then ask in plain language. Three shapes of request work, and the skill states
|
||
the rules for each:
|
||
|
||
- **a question** — "hva er kravene til pris?"
|
||
- **a hypothesis** — "stemmer det at leverandoeren baerer risikoen for grunnforhold?"
|
||
Answered per premise as `confirmed` / `refuted` / `undecidable-from-bundle`.
|
||
- **a task whose answer is a document** — "lag `krav-pris.md` med alle krav til
|
||
pris, ett avsnitt per krav, med dokument og kravnummer." Every claim in the
|
||
written file carries its source; a paragraph with no ground is written
|
||
and marked, never dropped.
|
||
|
||
Everything below is detail: [Consume in Claude
|
||
Code](#consume-in-claude-code) for the same thing in steps and with several
|
||
bundles at once, [Build](#build) for the flags, [Requirements](#requirements)
|
||
for the pip fallback and the guard pairing.
|
||
|
||
## Known limitations
|
||
|
||
Read this before pointing the tool at documents you depend on. Every number
|
||
here was measured; none of it is a plan.
|
||
|
||
- **The default gate refuses whole documents, and they are documents you may
|
||
want.** Measured 2026-09-20 against a real corpus of official documentation:
|
||
`guard-trusted-source`, the shipped default, refused a minority of sources
|
||
outright, under `fail_secure` and `quarantine_review`, and most of those were
|
||
ordinary reference pages. Not one element of a refused document reaches the
|
||
bundle. Rebuilt with `--gate none`, every one of them went through
|
||
untouched, so the refusal is the gate and not the readers: a page of
|
||
official documentation naturally carries commands and instruction-shaped
|
||
text, and the guard reads that as something to hold for review. **The
|
||
corpus, its size and the per-page counts are deliberately not published
|
||
here** — it belongs to a consumer whose material this repository does not
|
||
republish — so this bullet carries no denominator. Run your own: the build
|
||
names the count, the files and the codes on every run, which is the number
|
||
that actually binds you. The build says so now — it names the count, the files, the codes and
|
||
that command — and exits 0, because the bundle is a true record of what the
|
||
gate allowed. **If you vouch for the source yourself, build with `--gate
|
||
none`;** the bundle then records that nothing was screened. The default was
|
||
chosen on one measurement over one pinned bundle's 453 concept bodies, which
|
||
is a thin denominator for a decision this consequential.
|
||
- **Nothing bounds what one run pays for images.** Each decode link is capped
|
||
(`MAX_FILTER_DECODE_BYTES`, 512 MiB) and an oversized picture is refused with
|
||
its own code, but the run as a whole has no ceiling: measured, a 70 KB PDF
|
||
carrying 16 images each under the declared limit reached **851 MB peak RSS**
|
||
and every picture was carried. A hard cap outside Python was measured and is
|
||
not available here — `resource.setrlimit(RLIMIT_AS)` raises on Darwin 26.6.2
|
||
and is not enforced — so the per-link budget is the whole bound.
|
||
`--no-assets` takes the image path out entirely.
|
||
- **Three of this repository's own gates are RED, and each red row is a stated
|
||
finding rather than a bug to be surprised by.** The retrieval gate is red on
|
||
rows 5, 7, 8 and 9, the MCP gate on row 2, and the content accounting's judge
|
||
on rows 2, 3 and 6. For a user that means: retrieval quality is measured but
|
||
not yet green on a held-out set (rows 5, 8), two mechanical mutants of the
|
||
ranking survive with 0 ranks and 0 deliveries moved (row 7), no gold set
|
||
exists for the K2 corpus (row 9), MCP anchors and concept ids are different
|
||
vocabularies so `okf_fetch` cannot be addressed with a set's anchor (row 2),
|
||
and the accounting still reports real losses on the reference corpus (rows 2,
|
||
3, 6). The rows and their numbers are under [Judge the
|
||
retrieval](#judge-the-retrieval-python3-toolsokf_retrieval_gatepy) and
|
||
[Serve a bundle over MCP](#serve-a-bundle-over-mcp-okf-mcp).
|
||
- **The content accounting counts the element classes its vocabulary names, and
|
||
no others.** `0 unaccounted` is a statement about those classes, not about the
|
||
document: a file whose suffix has no reader is accounted at file level only,
|
||
parts no vocabulary names (headers, footers, endnotes, comments, speaker
|
||
notes, cell formulas) are outside it, and an image in `xlsx`, `md`, `txt`,
|
||
`csv`, `json`, `odt` or `rtf` is unaccounted and therefore red. It is opt-in
|
||
(`--accounting PATH`) for that reason. The full list is under
|
||
[Build](#build).
|
||
- **There is no context graph and no visualisation.** Nothing in this package
|
||
draws a bundle.
|
||
|
||
## What this library is
|
||
|
||
Status: phases 1–3 are implemented. Phase 1 (spec-based ingestion) covers
|
||
manifest validation, the `file`/`sql`/`http` connectors, deterministic
|
||
materialization, index generation, and the golden fixture suite under
|
||
`examples/`. Phase 2 adds the bundle inbox (`process_inbox`) and
|
||
external-bundle import (`import_bundle`), both against an **injected** persist
|
||
gate, with `llm_ingestion_okf.guard_adapter` wiring that gate to the real
|
||
guard (see below). Phase 3 makes the bundle contract configurable, so types,
|
||
layers, frontmatter sets, index shape, and reserved-file policy are carried by
|
||
a profile rather than by constants (see [Upstream OKF
|
||
versions](#upstream-okf-versions)). Binary extraction runs behind the
|
||
optional `[extract]` extra: `pdf` through a PDF parser, and five office
|
||
formats through a vendored document converter. Three of those five office
|
||
rows are **constructed** rather than measured — see
|
||
[Binary extraction](#binary-extraction). Phase 4
|
||
(the Node half) is planned (see `docs/plan/`).
|
||
|
||
## Supported file types
|
||
|
||
Thirteen extensions are read. The table below is pinned to the extraction
|
||
registry by `tests/test_docs_promises.py` — row for row, and cell for cell on
|
||
the evidence class the code records — so a type added without a row here turns
|
||
that test red.
|
||
|
||
| File type | Read by | Dependency | Evidence | Note |
|
||
|---|---|---|---|---|
|
||
| `.md` | `_extract_passthrough` | core | stdlib, no corpus class | The decoded bytes are the concept text, so `source_lines` are the original's own lines. |
|
||
| `.txt` | `_extract_passthrough` | core | stdlib, no corpus class | As `.md`. A document with no headings yields no segments, which is a failed build rather than a flat bundle. |
|
||
| `.csv` | `_extract_csv` | core | stdlib, no corpus class | Parsed with the stdlib reader and rendered as one markdown table; a file with no header row is refused. |
|
||
| `.json` | `_extract_json` | core | stdlib, no corpus class | Fenced verbatim. No structure is derived from the keys. |
|
||
| `.html` | `_extract_html` | core | measured | Block tags open their own lines and `h1`–`h6` carry the ATX marker for their own level. It does **not** go through the converter, although the converter reads HTML: that route would add CVE-2025-51591 (SSRF via an iframe in HTML input) and buy nothing. The denominator is 828 files — one product, one format, one publisher. |
|
||
| `.htm` | `_extract_html` | core | stdlib, no corpus class | The same reader as `.html`. The 828-file class is recorded for `.html` alone, and this row does not borrow it. |
|
||
| `.xml` | `_extract_xml` | core | measured | A NISO-STS document (`<standard>` root, or any `<sec>`) becomes one heading per titled section at the section's own nesting depth; any other XML keeps its text in document order and gets no invented structure. A `<!DOCTYPE` is refused unparsed. The denominator is one file, one publisher, one schema — 2 761 titled sections. |
|
||
| `.pdf` | `_extract_pdf` | `[extract]`: pdfplumber | measured | Eight corpus documents with a hand-counted fasit, plus a 701-page process code whose publisher also ships its structure. Prose only: drawn content has no text to recover, and every extraction warns. OCR lives here as a reading mode for a PDF page whose own text never arrived (`--ocr`, `OCR_CID_SHARE`), never as an entry for image files. |
|
||
| `.docx` | `_extract_office` | `[extract]`: pypandoc-binary | measured | Five corpus documents. `source_lines` index the extracted text and not the original's paragraphs — the two counts agree on none of the five. |
|
||
| `.xlsx` | `_extract_office` | `[extract]`: pypandoc-binary | measured | Written as pipe tables, one source row per line; a sheet name becomes a heading, and a row is located by `source_sheet` and `source_rows`. |
|
||
| `.pptx` | `_extract_office` | `[extract]`: pypandoc-binary | constructed | N = 2 decks. 2 of 2 slide titles recovered on a deck that declares them, 0 of 2 on a deck that does not, where the converter writes `Slide 1` / `Slide 2` because it has no title to use. |
|
||
| `.odt` | `_extract_office` | `[extract]`: pypandoc-binary | constructed | N = 1 document. 1 of 1 declared headings recovered, 1 concept, 0 characters in no segment. |
|
||
| `.rtf` | `_extract_office` | `[extract]`: pypandoc-binary | constructed | N = 1 document, and the weakest row here. 0 declared headings: the container carries no heading style, so the author's title is bold text and the document lands as one concept — content preserved, structure zero. The `--bold-title` flag reads that bold line and is off by default. |
|
||
|
||
**Since 0.10.0 five of those rows also carry IMAGES** — `.pdf`, `.docx`,
|
||
`.pptx`, `.html`/`.htm` and `.xml`. The bytes go to `assets/` at the bundle
|
||
root under a content-addressed name, and the concept carries a pointer where
|
||
the picture stood. `.xlsx` is deliberately not among them: its converter writes
|
||
one pipe table per sheet and a two-line block inside one would break the row
|
||
locator `source_rows` is read back out of; measured 2026-09-16, 0 of 4 K2
|
||
workbooks hold any media at all, so the row is a stated limit and not a loss
|
||
taken. `.csv`, `.json`, `.md`, `.txt`, `.odt` and `.rtf` are absent because
|
||
nothing has measured an image reaching them.
|
||
|
||
The three classes are the code's own and are not interchangeable. `measured`
|
||
means real documents someone wrote for their own purposes, counted against a
|
||
fasit written before the lookup. `constructed` means the row has been through
|
||
end to end on a hand-built document with a hand-written fasit and has met no
|
||
document anyone else wrote. `stdlib, no corpus class` means the code records no
|
||
class for the row at all. The office rows are set out in full under [Binary
|
||
extraction](#binary-extraction).
|
||
|
||
### Not read today
|
||
|
||
Facts about the registry as it stands, not a queue — nothing here is planned.
|
||
|
||
- `.doc` (Word 97) — the converter does not read it.
|
||
- `.epub` — the converter does read it, and the row is deliberately absent: it
|
||
would buy nothing today, the same half of the reason that keeps HTML out of
|
||
the converter.
|
||
- `.eml` and `.msg` — no reader. An email file arrives as an unregistered
|
||
extension.
|
||
- Image files — no reader, and OCR is a PDF page mode rather than an entry for
|
||
them.
|
||
- Source files (`.py`, `.ts`, and the rest) — no reader. The text rows are
|
||
`.md`, `.txt`, `.csv` and `.json`.
|
||
- `.one` (OneNote) and `.vsd` (Visio) — no reader.
|
||
|
||
An unregistered extension is a coded rejection (`extractor_unknown`), never a
|
||
silent skip.
|
||
|
||
## Install in detail
|
||
|
||
Neither this package nor the guard it depends on is on a package index yet, so
|
||
both install by direct reference. With uv, one command resolves both:
|
||
|
||
```sh
|
||
uv pip install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v1.0.0"
|
||
```
|
||
|
||
uv resolves the guard on its own, because it reads the `[tool.uv.sources]`
|
||
entry in the `pyproject.toml` **of the tag it is installing**, and `v1.0.0`
|
||
points that entry at `llm-ingestion-guard` `v1.4.0`. Use `uv tool install`
|
||
instead of `uv pip install` when you want the `okf` command on `PATH` without an
|
||
active virtualenv — that is the form the first screen shows.
|
||
|
||
With plain pip, the transitive git dependency does not resolve on its own —
|
||
**install the guard first**, or installing this package fails with
|
||
`No matching distribution found for llm-ingestion-guard`:
|
||
|
||
```sh
|
||
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.4.0"
|
||
pip install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v1.0.0"
|
||
```
|
||
|
||
The guard tag is paired to the okf tag, not to this branch. `v1.0.0` declares
|
||
`llm-ingestion-guard>=1.2,<2.0`, which `v1.4.0` satisfies; the pairing above is
|
||
read off that tag's own `[tool.uv.sources]`, not off this branch. Reading a pin
|
||
off `main` and installing it against an older okf tag is the one combination
|
||
that fails.
|
||
|
||
### Earlier tags, as history
|
||
|
||
These are not install lines. They record what each earlier tag was, so a reader
|
||
who meets one in an older document knows what they are looking at.
|
||
|
||
- `v1.0.0` — the current tag, and the version this tool is finished at. It
|
||
adds no capability to `v0.10.1`: a document the gate refuses whole is now
|
||
named in the run's own summary with its code and with the one command that
|
||
carries it anyway, and the front page states what this tool does not do.
|
||
Read [Known limitations](#known-limitations) before you depend on it. After
|
||
this tag the library is touched for defects found in its own use, and the
|
||
next round is Google OKF v0.3.
|
||
- `v0.10.1` — the image path of `v0.10.0`, with the two
|
||
defects an independent review found in it closed. A remote `<img src>` or
|
||
`xlink:href` is inert text with the address in one code span, never a live
|
||
markdown image link, and no longer loses the figure's caption. An image is
|
||
bounded in three places rather than one: the size a container DECLARES, the
|
||
size a carried file has, and — new in this tag — what the stream behind a
|
||
PDF image actually DECOMPRESSES to, which is an independent number. A
|
||
declared size that is not positive is refused with its own code,
|
||
`asset_size_invalid`, before the stream is read.
|
||
- `v0.10.0` — a bundle carries the IMAGES its sources
|
||
declare. Five readers place them (`pdf`, the converter's office rows,
|
||
`html`, `xml`), `assets/` at the bundle root holds the bytes under a
|
||
content-addressed name, and the concept carries a two-line pointer where the
|
||
picture stood. ON by default; `--no-assets` reproduces the pre-0.10.0 bytes,
|
||
measured on the 43-document reference corpus as a one-line difference in
|
||
`log.md`. The image bytes are not screened — the gate reads text — and the
|
||
log says so. Door C carries the assets its merged concepts point at.
|
||
- `v0.9.0` — `okf quality <bundle>`, a per-file-type
|
||
verdict on a bundle with the denominator on every line. Three verdicts
|
||
(`PASS` / `FAIL` / `UNMEASURED`) and a type with no measured threshold is
|
||
never `PASS`; two bars ship, `.pdf` 8/32 and `.docx` 2/5, both regression
|
||
bars against the pinned 43-document reference rather than quality claims.
|
||
`okf check` is untouched and still has seventeen rules.
|
||
- `v0.8.5` — a block `sources:` sequence is decoded by all
|
||
three of this library's flat frontmatter readers, where they returned the
|
||
key with an empty value. `okf.parse_frontmatter` is public API, so this
|
||
changes what an outside caller reads: it returns a flow string where it
|
||
returned an empty one. That string is a READING projection — PyYAML reads
|
||
it back on 0 of the 4 605 block files measured, because the `?` opening a
|
||
query string in the source URL ends the flow scalar — and the emitter still
|
||
writes flow, so no bundle bytes move. `okf consume` also stops scoring the
|
||
door's own `Enclosing section:` link line, which is now the default reading;
|
||
the excerpt still carries the line, so what moves is order and never an
|
||
excerpt's bytes. No new functionality; `okf check` has seventeen rules.
|
||
- `v0.8.4` — frontmatter this library writes is YAML a YAML
|
||
reader reads back the same (a block scalar it would misread is written
|
||
double-quoted; a flow leaf with no form both readers accept is refused), and
|
||
`okf consume` resolves a concept's `parent:` pointer into the excerpt
|
||
(`parent: { concept_id, title }`) with one enclosing-section link in a
|
||
heading-only body, a pointer the index now resolves too. `okf consume
|
||
--follow-parent` and `okf build --shell-parent` ship off; `okf check` has
|
||
seventeen rules (`parent_unfollowable`). The guard pin moves to `v1.4.0`,
|
||
which reads a flow sequence of plain scalars (`source_offset: [1, 24]`) that
|
||
the previous pin refused.
|
||
- `v0.8.3` — a NISO-STS document's own `<doc-number>` names
|
||
its directory and titles its `sources` entry, `okf build --frontmatter
|
||
KEY=VALUE` stamps a key on every concept of a run, and an STS section's
|
||
`description` comes from its own first spec point. `okf consume` keeps a
|
||
leading directory every concept id shares out of its first fusion signal,
|
||
which is what keeps that identity from costing the answering section its
|
||
rank. `okf build --shell-parent` ships off; `okf check` still has sixteen
|
||
rules.
|
||
- `v0.8.2` — `okf check` refuses a skill and a payload that
|
||
name different bundles (`bundle_mismatch`, the checker's sixteenth rule;
|
||
`v0.8.1` has fifteen), and a covered title no longer rises above a title
|
||
that answers more of the question. No new command or flag; the shipped
|
||
`skills/okf-consume/` is regenerated so it passes that check.
|
||
- `v0.8.1` — a question that accounts for a concept's WHOLE
|
||
title reads that concept first (`--title-covered`, on by default, opt out
|
||
with `--no-title-covered`). A ranking fix, no new functionality: on one
|
||
publisher's 2 761-concept bundle the answering section was delivered at
|
||
rank 1 on 3 of 6 scored questions before it and 6 of 6 after, and no other
|
||
measured bundle's payload changed one byte.
|
||
- `v0.8.0` — `.xml` is a core file type, read as NISO-STS through the stdlib
|
||
parser, and a section the source DECLARES takes the
|
||
declared-structure route — one publisher's process code segments at 2 761 of
|
||
2 761 of its own declared sections at the shipped defaults. No other file
|
||
type changes one byte, measured on the bytes.
|
||
- `v0.7.0` — `okf project` builds the bundle `okf build` builds (they were one
|
||
flag apart before it), and the generated skill states the question,
|
||
hypothesis and document-task modes with relative paths.
|
||
- `v0.6.0` — the first tag carrying the `okf project`, `okf consume`,
|
||
`okf check` and `okf skill` subcommands.
|
||
- `v0.5.0a2` — a pre-release for the named OKF v0.2 pilot set only.
|
||
- `v0.4.0` — the last tag before the OKF v0.2 work; it declares
|
||
`llm-ingestion-guard>=0.2,<0.3`, which only guard `v0.2.0` satisfies.
|
||
|
||
**No tag yet makes OKF v0.2 generally available.** `OKF_LATEST` is unchanged and
|
||
still points at `DEFAULT`; flipping that alias is the GA event and none of the
|
||
tags above is it (see [Upstream OKF versions](#upstream-okf-versions)).
|
||
|
||
## Build
|
||
|
||
Installing the package installs one command. A folder of documents in, an OKF
|
||
bundle out:
|
||
|
||
```
|
||
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
|
||
```
|
||
|
||
It walks the folder recursively, proposes a segmentation for each document with
|
||
the mechanical rules, replays those proposals through the bundle inbox, writes
|
||
the bundle and its `log.md`, and prints the run's numbers. Every proposal is
|
||
marked `PROPOSED` and `adjudicated: false` — the command segments nothing a
|
||
human has approved, and says so in the artifact.
|
||
|
||
The last line that matters is the conservation identity: `merged + coded
|
||
rejections == N`, where `N` is the folder's file count read at run time. **The
|
||
run exits non-zero when it does not hold**, and names the unaccounted files, so
|
||
a pipeline cannot mistake a partial bundle for a complete one.
|
||
|
||
Flags worth knowing: `--segments off` ingests each document as one concept and
|
||
asks for no root values; `--plans-dir` keeps the proposals instead of
|
||
discarding them; `--report` writes the full report to a file as well as stdout.
|
||
`--ingested-at` and `--proposed-at` default to `1970-01-01T00:00:00Z` rather
|
||
than the clock, so two builds of the same folder are byte-identical — a
|
||
wall-clock default would break rebuild-equals-incremental for every caller who
|
||
did not pass them.
|
||
|
||
### Images: what the source draws, carried (0.10.0)
|
||
|
||
`okf build` carries the images its sources declare. The bytes go to `assets/`
|
||
at the bundle root, named `<sha256[:12]>-<the source's own base name>`, and the
|
||
concept carries a two-line pointer where the picture stood:
|
||
|
||
```markdown
|
||

|
||
Image: graphics/tabell-84-2.png (120x90 px) -- Tabell 84-2 Toleranseklasser
|
||
```
|
||
|
||
The first line is markdown, so a reader that renders the concept sees the
|
||
picture. The second states what the first cannot — the name the SOURCE gave the
|
||
file and the size in pixels — which are the two facts a person checking the
|
||
bundle against the original needs.
|
||
|
||
**Why it exists.** Measured on R761 Prosesskoden:2025, a process code published
|
||
as a 701-page PDF and as a NISO-STS delivery: the process text is carried in
|
||
full, and 12 `Tabell N-N` and 9 `Figur N-N` captions stand over nothing,
|
||
because the publisher ships those tables as raster images in both deliveries.
|
||
Process 84 says "toleranseklasse ... er gitt i tabell 84-2" and table 84-2 is a
|
||
JPEG. A bundle like that reads as complete and is not.
|
||
|
||
**What it costs, measured on the 43-document reference corpus** (`K2/trinn1`,
|
||
two builds of the same commit, 2026-09-17):
|
||
|
||
| | `--no-assets` | default |
|
||
|---|---|---|
|
||
| concept files | 453 | 454 |
|
||
| markdown files | 865 | 867 |
|
||
| assets | 0 | 2 964 |
|
||
| bundle size | 4.7 MB | 115 MB |
|
||
| wall time | 2 414 s | 3 088 s |
|
||
| peak RSS | 6.26 GB | 8.74 GB |
|
||
|
||
2 964 carried of 3 145 found; 4 622 pointers, so content de-duplication folds
|
||
1 658 repeats into the files they already are. 422 of the 865 markdown files
|
||
differ. **One concept appears, and the mechanism is measured rather than
|
||
guessed:** the pointers are body text, so a section holding 146 of that
|
||
document's images grew from 19.0 % to 30.6 % of the extracted text and crossed
|
||
`--outline-gate`'s 0.20 share clause.
|
||
|
||
**`--no-assets` reproduces the pre-0.10.0 bytes**, and `log.md` then says
|
||
`NOT CARRIED` rather than falling silent — a bundle nobody looked for figures
|
||
in must not read like a bundle of documents that had none.
|
||
|
||
**The image bytes are not screened.** The gate reads text; a picture is not
|
||
text. The pointer block passes the gate like any other body line, and the file
|
||
beside it passes nothing. `log.md` says so on every run that carries one.
|
||
|
||
**Every carried image is one a model can be SHOWN.** A bundle that holds a
|
||
picture nothing can read is worse than one that says the picture is missing:
|
||
the count reports that it arrived. Measured over the frozen R761 delivery's own
|
||
`assets/` (denominator 50): 29 JPEG, 2 PNG and **19 RLE8 BMP** — correct files
|
||
that no model decodes. Every asset's type is read off its bytes and tested
|
||
against the viewable set; a BMP is converted losslessly to PNG (8-bit
|
||
uncompressed, 8-bit RLE8, 24-bit uncompressed), and anything else outside the
|
||
set is refused with `asset_not_viewable` and a line in the concept saying what
|
||
stood there. A BMP variant this reader does not express — RLE4, BITFIELDS,
|
||
16- or 32-bit samples, a 12-byte BITMAPCOREHEADER — is `asset_bmp_unsupported`,
|
||
a different fact about the document and a different thing to go and fix.
|
||
|
||
The reader is stdlib and adds no dependency. Pillow, which this tree already
|
||
carries transitively under `pdfplumber`, was measured first and rejected on two
|
||
counts: images are carried on the CORE path, where `.html` and `.xml` need no
|
||
`[extract]` extra, and an asset's name is its content digest — encoding through
|
||
an installed library would make a bundle's identity move with that library's
|
||
version, which is the property page rasterisation was felled over. Pillow is
|
||
the independent decoder in the tests instead, and against it **19 of 19** of
|
||
R761's real RLE8 assets convert with identical RGB, 2 366 365 pixels compared.
|
||
|
||
A converted asset is ONE asset: one file in `assets/`, one pointer, one row in
|
||
the accounting. The pointer's second line — where the source's own file name
|
||
and the size in pixels already live — states the original media type, the
|
||
original sha256 in full and the new one, so a reader can take the original
|
||
delivery, run `shasum -a 256` and find the row.
|
||
|
||
The cost is measured per image rather than per bundle, with a committed
|
||
script (`tools/okf_asset_census.py`) run from two pinned trees over 9 714
|
||
image rows: exactly **35 rows moved**. Nineteen are the BMPs, now PNG. The
|
||
other **16 are JPEG 2000 objects** carried out of PDF streams — a format no
|
||
model decodes either, and one no stdlib route converts, so they are refused
|
||
with `asset_not_viewable` and stated in the concept instead of being carried
|
||
unreadably. **9 321 of 9 321** JPEG and PNG rows are byte-identical across the
|
||
move.
|
||
|
||
<!-- asset-viewable-media-types: image/gif,image/jpeg,image/png,image/webp -->
|
||
|
||
**A size CEILING, read off the same corpora (0.10.1).** An image over
|
||
`MAX_IMAGE_PIXELS` (40 000 000 pixels) or `MAX_IMAGE_BYTES` (256 MiB) is
|
||
refused with `asset_too_large`, counted like every other refusal. The largest
|
||
image in the 43-document reference corpus is 4 515 x 4 128 (18.6 MP) and the
|
||
largest of R761's 109 pictures is 2 072 x 656 (1.4 MP), so the bound is an
|
||
order of magnitude above anything measured.
|
||
|
||
It exists because a few kilobytes can declare an enormous picture: a 9.6 KB
|
||
PDF declaring 3 000 x 3 000 grayscale zeros took 83 MB of peak RSS, a 63 KB
|
||
one declaring 8 000 x 8 000 took 276 MB, and the cost is linear in the pixel
|
||
count, so one document could take a whole batch build with it — before any
|
||
gate, because the guard never sees image bytes.
|
||
|
||
**Three numbers are bounded, not one, because a claim is not a cost.** What a
|
||
container DECLARES (`/Width` x `/Height`, a PNG header, a `data:` payload's
|
||
encoded length) is read before anything is decoded. What a carried FILE
|
||
measures is read the same way — this package never decodes such a file, so it
|
||
pays nothing for it, but writing a 7 000 x 7 000 PNG of 47 705 bytes into a
|
||
bundle would hand the consumer the same bomb with `7000x7000 px` printed
|
||
beside it. And what a PDF image's STREAM decompresses to is measured before it
|
||
is held, a chunk at a time and discarded, because `/Length` is the COMPRESSED
|
||
length and a dictionary declaring 1x1 may hang 400 MB of deflated zeros off
|
||
it. Measured: that document is 408 516 bytes and cost 892 MB of peak RSS with
|
||
only the declared size bounded; with the stream bounded it is refused at
|
||
54 MB, and a three-times-larger bomb costs 62 MB rather than 2 436 MB.
|
||
|
||
**Every link of the filter chain is bounded, not only the first.** A PDF
|
||
decodes a stream through a *list* of filters, and `/Filter [/FlateDecode
|
||
/FlateDecode]` puts the whole expansion in the second one: measured, 1 636
|
||
bytes of file cost 886 554 624 bytes of peak RSS when only the first link was
|
||
measured (52 367 360 with every link measured), and the picture was still refused at the end — after the memory had
|
||
been spent. An encrypted stream is deciphered first and then measured like any
|
||
other.
|
||
|
||
**And what a link COSTS is bounded, not the size of its output.** Bounding
|
||
every `FlateDecode` was still not a bound, because `ASCII85Decode` had been
|
||
classed as safe "because it shrinks" and it does not: `z` is that encoding's
|
||
shorthand for four zero bytes, so the filter quadruples its input, and
|
||
`base64.a85decode` holds about a hundred bytes of memory per byte of input.
|
||
Measured in paired subprocesses on an idle machine, the document built once
|
||
and read from a file so the fixture is not what is measured: a 33 475-byte PDF
|
||
decoding through `[/FlateDecode /ASCII85Decode]` cost 3 261 599 744 bytes of
|
||
peak RSS and the picture was CARRIED; bounded it is 42 070 016 and
|
||
`asset_too_large`. Doubling the run of `z` takes the old cost to
|
||
6 461 558 784 and the bounded one to 40 280 064 — the cost no longer follows
|
||
the bomb. So `FlateDecode` is measured a chunk at a time as it is paid, and
|
||
every other permitted filter carries a MEASURED worst-case cost per byte of
|
||
input which is checked against the budget *before* its decoder is called.
|
||
Any other filter — `LZWDecode`, `RunLengthDecode`, `CCITTFaxDecode`,
|
||
`/Crypt`, anything written after this — has no measured ratio, so an image
|
||
behind one is refused UNREAD with its own code, `asset_pdf_unbounded`, rather
|
||
than decoded to find out what it costs. The cap that falls out for
|
||
`ASCII85Decode` is read off the corpora the way the pixel bound is: over the
|
||
9 668 image objects of the 77 PDFs measured, 16 decode through such a link and
|
||
the largest input to one is 450 739 bytes, more than ten times under the cap.
|
||
A property test runs **every** chain of length 1–3 over the ten filters
|
||
pdfminer decodes — 1 110 of them — and requires each to be delivered under the
|
||
bound or refused with a published code, never paid for on the way.
|
||
|
||
**What the stream bound does NOT reach**, stated because the difference
|
||
matters: a stream something else has already decoded, where the memory was
|
||
spent before this package was asked. That one is caught by a check on the
|
||
decoded length AFTER the decode, which makes it a counted refusal rather than
|
||
a bounded one — the picture is dropped by count, not by bound.
|
||
|
||
**A declared size that is not a size** — a zero or negative `/Width` or
|
||
`/Height` — is refused with its own code, `asset_size_invalid`, before the
|
||
stream is read. Distinct from `asset_too_large` on purpose: one is a publisher
|
||
shipping a picture bigger than this package carries, the other is a dictionary
|
||
written wrong or written to be read wrong.
|
||
|
||
<!-- asset-max-pixels: 40000000 -->
|
||
|
||
**A remote reference is INERT (0.10.1).** `<img src="https://...">` and an STS
|
||
`xlink:href` to an address off this machine are written as text with the
|
||
address in one code span, never as `` and never as a bare
|
||
URL a linkifying renderer would autolink. Extraction opens no socket, but a
|
||
markdown renderer or an agent that fetches what it renders does, which would
|
||
turn "this bundle was opened" into a beacon to whoever wrote the document. The
|
||
address is still stated, and so is the figure's caption, because a reader has
|
||
to know what stood there.
|
||
|
||
**`images: N` in a concept counts POINTER BLOCKS, not unique pictures.** One
|
||
image referenced twelve times in one concept is `images: 12` and one file in
|
||
`assets/`. The key is a count of the places a picture stands, and dedup is on
|
||
content.
|
||
|
||
**No size floor, and that is a measurement too.** The obvious filter is "ignore
|
||
anything under N pixels", and the distribution offers no N: over the 4 828
|
||
image objects in that corpus, 149 declare no size, 162 are under 32x32, 92
|
||
under 64x64, 406 under 128x128, 498 under 256x256, 590 under 512x512 and 2 931
|
||
are larger — a broad spread with no gap, unlike `OCR_CID_SHARE`'s, which is
|
||
bimodal with nothing between the modes. A threshold read off no gap is a number
|
||
this package chose, and it would silently drop somebody's small table.
|
||
|
||
<!-- cli-default-assets: on -->
|
||
|
||
`--accounting PATH` accounts for the CONTENT, not only the files. Before
|
||
extraction, every source document is inventoried in a per-format element
|
||
vocabulary: headings, paragraphs, tables, cells, images, pages, and so on.
|
||
After the run, each element gets exactly one fate:
|
||
|
||
- **carried:** all of its text is in the concepts written for the document,
|
||
or its image is in `assets/`;
|
||
- **pointer:** a remote image, which is never fetched, or a markdown image
|
||
reference kept verbatim;
|
||
- **a coded rejection:** the gate's or the reader's code.
|
||
|
||
The result goes to PATH as JSON and into `log.md`. The build exits 1 when an
|
||
element has no fate or has two. A document the gate refuses is logged as
|
||
`<file>: <M> elements found in the source, 0 carried: document rejected
|
||
`<code>``, and the `Images` bullet then counts what the sources declare.
|
||
|
||
"Carried" means the text is present, not that it is in the right place. A
|
||
short element such as a section label can be found elsewhere in the same
|
||
document. The judge is `tools/okf_accounting_gate.py`, which compares the
|
||
inventory against an independent witness.
|
||
|
||
**The account covers the element classes the vocabulary knows, and no others.**
|
||
`accounting._READERS` names twelve suffixes, each with its own tuple of classes
|
||
(`.md`: heading, paragraph, table, table_row, image, code_block; `.pdf`: image
|
||
and page only, which is the approved exception below). Two consequences are
|
||
stated here rather than left to be discovered, because "0 unaccounted" reads
|
||
like a statement about the document and is a statement about those classes:
|
||
|
||
- **A file whose suffix has no reader is accounted at FILE level only** —
|
||
carried, merged or rejected — never element by element.
|
||
- **Parts of a document that no vocabulary names are not counted, so content
|
||
there can go missing under exit 0 and `0 unaccounted`.** Verified against the
|
||
readers: `.docx` reads `word/document.xml` and `word/footnotes.xml`, so
|
||
headers, footers, endnotes and comments are outside; `.pptx` reads
|
||
`ppt/slides/slideN.xml`, so speaker notes, masters and layouts are outside;
|
||
`.xlsx` reads the worksheets, the shared strings and the drawings, so cell
|
||
comments are outside and a cell contributes its cached value or inline
|
||
string, never its formula; `.rtf` skips the `header`, `footer`, `info`,
|
||
`pict`, `stylesheet`, `fonttbl` and `colortbl` groups. A hidden slide or
|
||
sheet IS counted — it lives in the same part as a visible one. Nothing here
|
||
is built for now: the list is what the account does not claim.
|
||
|
||
**Two operator decisions, 2026-09-17.** The accounting stays OPT-IN until the
|
||
losses it reports on the reference corpus are fixed, because a default-on door
|
||
would fail builds that pass today. And of the three exceptions the gate
|
||
proposed, only the PDF one is approved: a PDF without a structure tree
|
||
declares no heading, paragraph or table, so no witness can count them. An
|
||
image in a workbook, or in md, txt, csv, json, odt or rtf, stays unaccounted
|
||
and therefore stays red.
|
||
|
||
A run that refused a document whole says so in both places: the accounting
|
||
carries `refused` and each document's own `status`, and `log.md` carries
|
||
`R of D document(s) refused whole`. The exit code does not move for it — it
|
||
belongs to the whole run, and a corpus holding one unreadable file among many
|
||
is ordinary — so the count is what keeps a partial refusal from being silent.
|
||
The judge treats such a document as never clean, with its elements in their own
|
||
`refused` column: every one of them is booked honestly as a coded rejection, so
|
||
u and d both stay 0 and nothing else could see the loss.
|
||
|
||
**The soft hyphen is removed before the persist gate, and counted** (operator
|
||
decision 2026-09-18). U+00AD is in `llm-ingestion-guard`'s zero-width set, and
|
||
`output:zero-width-present` is an any-tier carrier: a document carrying one is
|
||
`fail_secure` at every trust level. Measured on R761 Prosesskoden:2025 — 71
|
||
U+00AD, and 0 of U+200B, U+200C, U+200D, U+FEFF and U+2060 — those 71 are
|
||
Norwegian hyphenation points inside words (`ar[SHY]beider`, `bitu[SHY]men`),
|
||
so a 701-page process code was unreadable for the whole chain over typography.
|
||
`extract.normalise_extracted` removes that one character from every extracted
|
||
text and reports the count as `normalised_soft_hyphen`, per document and for
|
||
the run, in the accounting JSON and in a `**Normalisation**` bullet in
|
||
`log.md`. The guard is not touched and the other four characters are not
|
||
touched: they carry no typographic job in running text, so removing one would
|
||
be a decision about what the guard screens for, taken in the wrong repository.
|
||
U+00A0 NBSP is not in the guard's set and is not touched either. Reach,
|
||
measured 2026-09-19: **0 of the 78** readable documents of the reference
|
||
corpus carry any of the six characters, so no bundle measured here moves.
|
||
|
||
Two things hold with or without the flag:
|
||
|
||
- `okf build` exits 1 when it extracted at least one document and persisted
|
||
none.
|
||
- An image file beside a document is counted once. If a persisted document
|
||
carried it, it is in the conservation identity's own column
|
||
(`merged + files carried through a document + coded rejections = N`);
|
||
otherwise it is a coded rejection.
|
||
|
||
`--frontmatter KEY=VALUE` (repeatable) stamps a key on every concept of the
|
||
run, for what the operator knows and the document does not say — an edition,
|
||
a publisher's address. It splits on the first `=` and writes the value
|
||
verbatim on one line, so `'sources=[{ resource: <url>, title: <t> }]'` survives
|
||
whole. It adds any key and replaces only `sources` and `description`, the two
|
||
with a layer below them (what the document declares, else the file name);
|
||
every other key the door writes itself is refused before anything is read.
|
||
|
||
`--shell-parent` (**off**; opt out explicitly with `--no-shell-parent`) gives a
|
||
concept whose body is its heading alone a `parent:` naming the `segment_id` of
|
||
the nearest ancestor that holds text — the nearest preceding plan entry at a
|
||
smaller level, passing over an ancestor that is empty too. It copies no text
|
||
and moves no boundary. It exists for a document that states its points once and
|
||
lets every nested section inherit them: measured on one process code, **710 of
|
||
2 761** concepts are heading-only, and the plan's level and order name the same
|
||
ancestor as the document's own nesting on **710 of 710** since K3-21 (708
|
||
before: the two others sit at depth 7, and the reader clipped their level to 6
|
||
in the plan as well as in the markdown heading, so they pointed one level too
|
||
high). It
|
||
was off because `okf consume` did not read `parent`. Since K3-21 it does: an
|
||
excerpt carries `parent` as the enclosing concept's `concept_id` and `title`
|
||
(resolved inside the concept's own document, never the raw segment id), and a
|
||
heading-only body gains one line, `Enclosing section: [title](/path)`, in the
|
||
bundle-relative form SPEC § 6.1 recommends. `okf consume --follow-parent`
|
||
carries the enclosing section's text inside `parent` as well, with that
|
||
concept's own `sha256`, and only from the room the cut left — so it never
|
||
displaces an excerpt. Whether either default moves is measured separately.
|
||
**The ranking does not score that line (K3-25).** It reaches the excerpt and
|
||
the file exactly as before; `consume.DEFAULT_LINK_IN_SIGNAL` is `False`, so the
|
||
body signal and the stem vocabulary read a heading-only body without it. The
|
||
measurement: of the newcomers the line ever added a question token to, **39 of
|
||
39** gained it from the bundle-absolute PATH and **0 of 39** from the link's
|
||
title, every such token being a segment of the document's own directory — so a
|
||
bundle built with `--shell-parent` now delivers what the unflagged build
|
||
delivers, on **16 of 16** rows. It cost no shipped bundle a byte: **0 of 5**
|
||
carry the line, so **5 of 5** payloads are byte-identical across the move. The index resolves the pointer too: a segment id is looked up among the
|
||
concepts of its own document, so it renders `parent: p1` when the concept it
|
||
names is in the bundle and keeps the `?` only when nothing answers to it.
|
||
|
||
### The segmentation flags
|
||
|
||
Nine rules are reachable from `okf build`, and since 2026-09-11 **all nine
|
||
are ON by default** — `--outline-run 3`, `--table-grid` and `--unit-fold` since
|
||
2026-09-08, `--drop-wrapped-outline` and `--outline-gate` since 2026-09-09,
|
||
`--sheet-section-rows`, `--keep-table-heading` and `--first-span-from-zero`
|
||
since 2026-09-10, and `--close-span-gaps` since 2026-09-11 — each an operator
|
||
decision, and each with an explicit
|
||
opt-out: `--outline-run 0`, `--no-table-grid`, `--no-unit-fold`,
|
||
`--keep-wrapped-outline`, `--no-outline-gate`, `--no-sheet-section-rows`,
|
||
`--no-keep-table-heading`, `--no-first-span-from-zero`,
|
||
`--no-close-span-gaps`. Passing all nine
|
||
reproduces the pre-2026-09-08 bytes exactly, and the last three reproduce the
|
||
pre-2026-09-10 bundle byte for byte — measured with `diff -rq`, 0 differences,
|
||
not asserted. Each line below carries the number it was measured at, and
|
||
nothing beyond it.
|
||
|
||
An **eleventh** flag, `--pdf-outline`, is **off** by default, and it is the one
|
||
rule here that does not read the extracted text at all: it cuts a PDF at the
|
||
boundaries the file's own `/Outlines` bookmark tree declares. It is not
|
||
`--outline-run` under another name — that one is a text heuristic over numbered
|
||
lines in the extracted text, while this one opens a structure index the PDF
|
||
already carries and the build had never looked at.
|
||
|
||
It is a segmentation arm and not a reader option: the extracted text is byte
|
||
for byte the same either way, and a PDF that carries no bookmark tree builds
|
||
byte-identically with the flag on. Two consequences come free with it. The
|
||
concept title comes from the BOOKMARK, so it is not cut short where the page
|
||
wrapped the heading across two lines; and a page that lies before the first
|
||
bookmark destination is the contents listing rather than a second copy of the
|
||
body, so a contents entry and the section it lists stop landing as two concepts
|
||
under one id.
|
||
|
||
The measurement is one 701-page process code whose publisher also ships a
|
||
NISO-STS structure for it, so the fasit is the publisher's own. Under the
|
||
shipped default that document gives 1967 of 2761 boundaries, none of its 28
|
||
chapters, and 794 of 794 misses have their heading text present in the text the
|
||
build read. With the arm it gives 2759 of 2761 and 28 of 28. The flag stays off
|
||
because reach is the open question, not quality: **1 of the 8** reference PDFs
|
||
in this repository's own sample carries a usable tree, and a bookmark tree is
|
||
the publisher's *claim* about its own structure — a stale or wrongly pointing
|
||
one carries that error straight into the segmentation.
|
||
|
||
A **tenth** flag, `--bold-title`, is **off** by default. It is the rule for the
|
||
type whose container declares nothing: `rtf` has no heading style, so an
|
||
author's title is bold text, and the row was measured at 0 of 0 declared
|
||
headings, 0 concepts and 1368 of 1368 characters in no segment. The grammar is
|
||
markdown, not `rtf` — the converter already writes that title as `**…**` in the
|
||
same output every office row produces — and it is gated by the principle
|
||
`--outline-gate` already carries: recovery yields to declaration. Measured over
|
||
47 readable documents, three parameters were swept and one carried; false
|
||
positives are **0 of the 31 documents that declare**, and the rule reaches
|
||
**2 of 39** corpus documents, both `docx`, **0 of 33 `pdf`** and **0 of 2
|
||
`xlsx`**. On the `rtf` fixture set it recovers **6 of 6 authored titles over
|
||
N = 4** with **0** false titles and **0 of 1994** characters in no segment. On
|
||
a five-document folder it moves 26 → 27 concepts, replacing a mechanical
|
||
`tabell-linje-30` with two named concepts.
|
||
|
||
**A re-run is what this costs a consumer, and it is not a small one:** on the
|
||
43-document reference corpus the default bundle goes from **629 concepts in
|
||
1108 files** (the 2026-09-03 tree) to **492 in 944** after the 2026-09-08 move,
|
||
to **425 in 810** after the 2026-09-09 one and to **436 in 832** after the
|
||
2026-09-10 one (digest `8dff8a8e6c15d2f7…`). On a five-document folder the last
|
||
move is **15 concepts in 30 files → 26 in 52**. The proposer's own defaults
|
||
(`tools/okf_propose_segments.py`) did NOT move, so every published reproduction
|
||
block still runs as written.
|
||
|
||
**What the 2026-09-09 move had to clear, stated because it is the bar every
|
||
later move is held to:** the twelve-position reference improves (`pdf` **2 of 8
|
||
→ 7 of 8**, the sheet **5 of 12 → 10 of 12**, `docx` unchanged at 3 of 3) AND
|
||
hit@8 on a K2 bundle built with it holds **5 of 6 at ranks 1,1,1,1,1,–**, with
|
||
no row losing rank 1. A configuration that improved the reference and cost a
|
||
rank was measured in the same session and did NOT ship; see
|
||
`docs/2026-09-09-k3-runde6-outline-gaten-og-prioren.md` § 9.1.
|
||
|
||
| flag | what it does | measured |
|
||
|---|---|---|
|
||
| `--outline-run N` (default **3**) | also propose a boundary where the document's own bare-integer numbering sustains an ascending run of at least `N`; `0` is this arm's opt-out | a tender PDF whose headings are bare integers: **no boundary** at `0`, **9 concepts** at `3`, against a reference of 9 |
|
||
| `--table-grid` (**on** by default; opt out with `--no-table-grid`) | a pandoc grid-table rule line no longer closes an open table block, so one grid table is one concept | a `.docx` experience list: **21 → 6** concepts |
|
||
| `--unit-fold` (**on** by default; opt out with `--no-unit-fold`) | discard a contents-list run, fold a deeper heading into its parent, fold a table into the shorter heading that introduces it. Adds no boundary, so it can only reduce a plan | on a 12-document sample scored against an operator's unit worksheet: **5 of 12** match — but that figure was measured with `--table-grid` ON, and the shipped default does not include it. Measured without it the same sample scores **2 of 12**, `docx` **0 of 3**, because the fold's table clause has no joined table to fold |
|
||
| `--keep-table-heading` (**on** by default since 2026-09-10; opt out with `--no-keep-table-heading`) | keep a heading whose body is empty only because a table opens under it, and absorb that table into its span | the two spreadsheets in that corpus, and **0 of 32 `pdf` and 0 of 5 `docx`**: the concept count does not move (1 → 1), its first byte does — the concept gains the heading line it was missing |
|
||
| `--sheet-section-rows` (**on** by default since 2026-09-10; opt out with `--no-sheet-section-rows`) | cut an open table block at the rows that label its sections — a run of at least three rows whose first cell is a bare numeric label. The opposite direction from `--table-grid`, which decides how far a block extends | a tender price sheet whose whole body is one table block: **1 → 12 concepts**, against a reference of 11 cost groups plus the sheet's preamble. Whole corpus: **1 of 39** readable documents changes, **0 of 32 `pdf`, 0 of 5 `docx`, 1 of 2 `xlsx`**. It reached 11 of 12 on the reference two rounds before it shipped, and was held back both times by a RETRIEVAL cost that turned out not to be its own: on a bundle built with it the gold document splits 1 → 12 concepts and row 1 of the hit@8 set fell rank 1 → 2. The repair is on the reading side (`--tie-shared-rank`, now the default), and with it in place the sheet reaches 11 of 12 with hit@8 holding **5 of 6 at ranks 1,1,1,1,1,–** |
|
||
| `--drop-wrapped-outline` (**on** by default since 2026-09-09; opt out with `--keep-wrapped-outline`) | do not admit an `--outline-run` candidate whose line continues onto the next one. Judges recovered candidates only, never a heading the document declares | quoted regulation text, whose numbered paragraphs match the outline grammar exactly: **4 → 1 concepts**, the reference. Whole corpus: **5 of 39**, all `pdf`; on the 12-document sample **8 of 34** outline candidates wrap, and none of the 26 the operator kept. On the reference it carries `pdf` from **5 of 8 to 6 of 8** together with the gate below, and neither reaches 7 of 8 without the other |
|
||
| `--outline-gate` (**on** by default since 2026-09-09; opt out with `--no-outline-gate`) | admit `--outline-run`'s RECOVERED headings only where the document declares none of its own, plus any one recovered heading whose span covers `OUTLINE_SHARE` (0.20) of the text. Applied at admission, before spans are closed, so the text a removed mark opened is carried by the mark above it rather than lost | on the 12-document sample: `pdf` **2 of 8 → 5 of 8** alone and **7 of 8** with the rule above, `docx` unchanged at **3 of 3**. Whole corpus: it fires on **25 of 39** readable documents, changes the plan in **15 of 39**, and removes **64 of 485** proposed entries. No plan disappears (32 → 32) |
|
||
| `--first-span-from-zero` (**on** by default since 2026-09-10; opt out with `--no-first-span-from-zero`) | start the first concept at character 0, so the text above it belongs to a segment instead of to none. Adds no boundary and removes none | Measured over the 39-document corpus, the default before this rule left **207 435 characters — 11.92 %** — in no segment at all: **163 804 above the first entry** (in **32 of the 32** documents that get a plan), 26 041 *between* entries and 17 590 after the last. This rule closes the first part entirely, 79 % of the whole, leaving **43 631 characters (2.51 %) over 8 of 32 documents** with two named mechanisms of their own. It adds no boundary and the K2 concept count is identical with and without it (**425 = 425**); on the 12-position reference it changes **not one cell**, and hit@8 on a K2 bundle built with it holds **5 of 6 at ranks 1,1,1,1,1,–** under both tie-breaks |
|
||
| `--close-span-gaps` (**on** by default since 2026-09-11; opt out with `--no-close-span-gaps`) | close a concept's span against the next SURVIVING concept, and the last against the end of the text. Three steps remove a candidate AFTER its neighbour's span was already closed against it — the orphan check, and `fold_units` clause 1 both between entries and on the last run — and the removed mark's text then belongs to no segment. Adds no boundary and removes none; only spans' ends move | It closes the whole remainder the rule above left: **43 631 characters, 2.51 % of the corpus over 8 of the 32 documents with a plan, to 0** — both the 26 041 between entries and the 17 590 after the last. Decomposed: orphan check **18 527** over 15 of 39 documents, clause 1 **7 514** between entries, clause 1 on the last run **all 17 590** of the tail (with `unit_fold=False` the corpus tail gap is 0). The entry count is identical (**429 = 429** on the corpus, **436 = 436** concepts on K2, **52 = 52** md on a five-document folder); on the 12-position reference it changes **not one cell** (11 of 12 under `|F|`[3]=12, 10 of 12 under `|F|`[3]=11), and hit@8 holds **5 of 6 at ranks 1,1,1,1,1,–** on the new bundle, the previous default and Arm B alike |
|
||
| `--pdf-outline` (**off**; opt out is the default, opt in with the flag) | cut a PDF at the boundaries its own `/Outlines` bookmark tree declares, instead of at the ones the text rules recover. A SEGMENTATION arm, not a reader option: the extracted text is byte for byte the same either way, and a PDF that carries no tree builds byte-identically with the flag on. The title comes from the BOOKMARK, so it is not cut short at the page's line break, and a page before the first bookmark destination is the contents listing rather than a second copy of the body | one 701-page process code whose publisher also ships a NISO-STS structure for it, so the fasit is the publisher's own: boundaries **1967 of 2761 → 2759 of 2761**, chapter level **0 of 28 → 28 of 28**, concept titles identical to the source title after normalisation **2761 of 2761**, false positives **163 of 2182 → 3 of 2762**, directories carrying two concept files **132 of 2050 → 2 of 2738**, contents-copy pairs **65 → 0**. Consumption on the same eight questions: fasit present in the bundle **4 of 7 → 7 of 7**, hit@1/8/50 **1/6 · 2/6 · 4/6 → 3/6 · 5/6 · 6/6**. Cost 119.22 s → 183.31 s wall, peak RSS 3252 → 3251 MiB, no new dependency and no second parse of the pages. **Off, and the reach is why:** **1 of the 8** reference PDFs carries a usable tree at all, and a bookmark tree is the publisher's CLAIM about its own structure — a stale or wrongly pointing one carries that error straight into the segmentation |
|
||
|
||
They compose, and the order above is the order they apply in. Measured on a
|
||
five-document tender folder (2 `pdf`, 2 `docx`, 1 `xlsx`), concepts per
|
||
document:
|
||
|
||
| document | default | `--outline-run 3` | `+ --table-grid` | `+ --unit-fold` | `+ --keep-table-heading` | `+ --sheet-section-rows --drop-wrapped-outline` |
|
||
|---|---|---|---|---|---|---|
|
||
| tender PDF, technical requirements | 1 (no boundary) | 9 | 9 | 9 | 9 | 9 |
|
||
| tender PDF, technical layout | 1 (no boundary) | 1 | 1 | 1 | 1 | 1 |
|
||
| price sheet `.xlsx` | 1 | 1 | 1 | 1 | 1 | **12** |
|
||
| experience list `.docx` | 21 | 21 | 6 | 3 | 3 | 3 |
|
||
| agreement `.docx` | 2 | 2 | 1 | 1 | 1 | 1 |
|
||
| markdown files in the bundle | 31 | 49 | 33 | 30 | 30 | 52 |
|
||
|
||
Every column merged 5 of 5 with 0 rejections. The reference for the first row
|
||
is 9, so the default is a full arm behind what the proposer can do on that
|
||
document — which is a statement about the default, not a licence to change it
|
||
here.
|
||
|
||
### The two PDF reader flags
|
||
|
||
Separate from the six above, and they sit before every one of them: a
|
||
segmentation flag changes how the proposer cuts a text, these change what the
|
||
text says. **All three are off by default.**
|
||
|
||
| flag | what it does | measured |
|
||
|---|---|---|
|
||
| `--pdf-headings font` | a PDF carries no heading markup, so one is inferred from typography — a line whose dominant font size is above the document's character-weighted median AND whose dominant font name says bold — and emitted as an ATX heading in the same markdown the office path produces, so the existing heading rule reads it | on a tender PDF: **9 of 9** numbered chapters found, plus 4 extra candidates. Whole corpus: **25 of 32 `pdf`** change, **0 of 5 `docx`**, **0 of 2 `xlsx`**. **Off by measurement:** against the operator's unit worksheet it takes `pdf` from **2 of 8 to 0 of 8**, losing two exact matches, because on those documents the outline rule already found the chapters and a second heading source can only add |
|
||
| `--pdf-headings font-reserve` | the same typographic rule, applied ONLY to a document whose own numbering the outline arm finds no run of — typography as a second heading source where there is no first one, never on top of one. Three values of one option (`none`, `font`, `font-reserve`), so no caller can ask for two at once | reaches **4 of 39** readable corpus documents (10 of 32 `pdf` admit no outline run; 4 of those render differently at all). **Off by measurement, and the measurement is that it changes nothing measurable:** on the operator's twelve-position unit worksheet it alters **not one cell** — the five positions where it fires are one PDF whose glyphs carry no ToUnicode mapping and four office documents the PDF reader never touches. The position it was built for numbers its own chapters, so the reserve is silent there by construction |
|
||
| `--ocr` | read a PDF page as an image when its own text never arrived: the page extracts empty, or as `(cid:N)` placeholder codes at or above 10 % of its characters. Needs the optional `ocr` group | on the one corpus document with the failure: **95.07 % → 0 %** cid, **44 → 2561** words of four or more letters, 17 → **18** pages with text, 3.6 s/page. Whole corpus: **16 of 834** pages qualify, in **1 of 32** files |
|
||
|
||
```
|
||
pip install "llm-ingestion-okf[extract,ocr]"
|
||
```
|
||
|
||
Without that group `--ocr` is a typed refusal (`extractor_ocr_group_missing`)
|
||
per file, never a crash, and the corpus run still reports
|
||
`merged + coded rejections == N`. OCR text is a model's reading of an image: it
|
||
is reproducible against the model version and rendering resolution it was
|
||
produced with, and no dependency pin can promise more. The full measurement,
|
||
including the per-page distribution the 10 % threshold was read off, is
|
||
`docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md`.
|
||
|
||
Measured 2026-09-08 on a 43-file corpus (33 `pdf`, 5 `docx`, 2 `xlsx`, and
|
||
three files no reader accepts), one
|
||
`okf build` invocation replacing the shell loop over `tools/` that produced the
|
||
same corpus's bundle on 2026-09-03:
|
||
|
||
| figure | value |
|
||
|---|---|
|
||
| `N` (folder file count, computed) | 43 |
|
||
| merged | 39/43 |
|
||
| coded rejections | 4/43 (`extractor_unknown` 3, `extractor_empty_pdf` 1) |
|
||
| K1b | `39 + 4 = 43 = N`, exit `0` |
|
||
| files written | 1108 |
|
||
| identical to the 2026-09-03 bundle | 1104/1108 |
|
||
| wall time | 842.82 s total, 19.600 s per file (re-measured 2026-09-08) |
|
||
|
||
The six files that differ are all in the corpus's two spreadsheet documents, and
|
||
they are the change reported in `docs/2026-09-08-prisform-og-loggen-k2.md`: a
|
||
spreadsheet's tables are now written as pipe tables, so each row is one line
|
||
with its cells delimited rather than padded out to the widest cell in the
|
||
column. Two concept files are renamed by it, two are removed under their old
|
||
names, and the two documents' own `index.md` follow. The root `index.md` is
|
||
identical to the stored one again, because this library no longer links the
|
||
bundle's `log.md` from it. The command's own byte-identity test compares it
|
||
against the two scripts at the current commit, where the two agree over the
|
||
whole tree.
|
||
|
||
## Consume
|
||
|
||
The other direction: a bundle plus one question in, one bounded, contract-shaped
|
||
payload out.
|
||
|
||
```
|
||
python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json
|
||
```
|
||
|
||
`tools/okf_consume.py` is the **pre-pass** `docs/consumption-contract.md` § 1
|
||
defines — the deterministic program that reads the bundle, ranks its concepts,
|
||
cuts them to a bounded set and emits one payload. It decides nothing about the
|
||
question; the skill that reads the payload does the judgement. It calls no
|
||
model, opens no socket, imports nothing outside the standard library and this
|
||
package, and takes no clock: the same bundle bytes and the same
|
||
`(question, k, limit, cost_vocabulary, reserve_top_rank, rarity_weight)`
|
||
produce byte-identical output.
|
||
|
||
**The ranking is BM25 since v1.1** (`--ranking bm25`, the default). Two signals
|
||
are fused by rank: each concept's best 500-character passage, and its title and
|
||
id path weighted above its body. A word the bundle does not hold weighs nothing
|
||
by itself; one it holds in another form — a Norwegian inflection or compound —
|
||
is read as that form. A concept longer than 4 000 characters is delivered as
|
||
the passage that answers, under the nearest heading above it, marked with
|
||
`passage: {start, end, of}`, so the whole can be fetched by its `concept_id`.
|
||
`--ranking fusion` is the older three-signal ranking; the flags below that say
|
||
they widen a signal (`--cost-vocabulary`, `--rarity-weight`) belong to it and
|
||
are refused without it. The rest of this section describes the fusion.
|
||
|
||
`--cost-vocabulary` is off by default and widens one question class: it lets a
|
||
declared list of cost/price/quantity terms bridge a question and a document that
|
||
name money with different words. The gate is the question — one naming no such
|
||
term gets byte-identical bytes either way — and what it does and does not close
|
||
is measured in `docs/2026-09-08-blindsone-below-k-k2.md`.
|
||
|
||
`--reserve-top-rank` is off by default and answers a different objection: the
|
||
budget is packed by an exact knapsack, which maximises a SUM of scores and
|
||
therefore has no opinion about rank, so a top-ranked excerpt costing a large
|
||
share of the budget is out-summed by many small ones. Measured, that made `--k`
|
||
a dial that could EVICT the concept a question was asked about. The flag gives
|
||
rank one its bytes before the pack runs — after the `over_budget_alone`
|
||
pre-exclusion, never before — and the payload then declares
|
||
`budget.reserved`. On a 629-concept corpus it changed the delivered set in 2 of
|
||
24 measured combinations, both of them that eviction:
|
||
`docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
|
||
|
||
`--source-quota N` is **on** by default at **2** since 2026-09-10 (opt out with
|
||
`--no-source-quota`), and it is the third widening here that alters a payload
|
||
with no bundle changing. It caps how many DELIVERED places one source document
|
||
may take, cutting where the shortlist is cut so the freed place goes to the next
|
||
candidate and `k` is still delivered in full. The defect it repairs was measured
|
||
outside this repository on a 3206-concept bundle of a published handbook: the
|
||
code's own process overview contributes **28 of 3206 concepts (0.87 %)** and
|
||
**8.0 % of the source characters**, and took **8 of 8** delivered places on one
|
||
question and **7 of 8** on the known-positive, which was not delivered at all.
|
||
Identical at 343 and 1651 concepts, so the cause is the corpus's COMPOSITION —
|
||
that it holds its own table of contents — and not its size; any corpus with a
|
||
contents list, a project overview or a summary document has the same property.
|
||
Swept over {2, 3, 4, off} on three bundles: at 2 and 3 hit@8 goes **5 of 6 to
|
||
6 of 6 on both K2 bundles** with all five standing rank-1 rows unmoved, and at 4
|
||
and off it stays 5 of 6. On the handbook bundle hit@8 goes **2 of 6 to 4 of 6**
|
||
and the dominant document's share of delivered places **8 of 8 to 2 of 8**. 2
|
||
rather than 3 on rank: the recovered rows come in at 5 and 4 rather than 7 and
|
||
5. What the gain is NOT: hit@8 asks whether the gold DOCUMENT was delivered, and
|
||
a document quota directly raises how many distinct documents a payload holds, so
|
||
that metric is not neutral with respect to this rule — the five rows that were
|
||
already rank 1 are, and they did not move. The adverse case is named rather than
|
||
found later: a bundle built from ONE document has one `source_file` on every
|
||
concept, so the quota would deliver 2 excerpts instead of `k`; the shortlist is
|
||
topped back up from the best-ranked over-quota candidates, which makes such a
|
||
bundle byte-identical to the quota being off. `--rarity-weight` was measured
|
||
against the same defect in the same session and does **not** repair it: on the
|
||
handbook bundle it leaves the dominant document at 8 of 8 places on the
|
||
question it floods and delivers neither that answer nor the known-positive.
|
||
|
||
`--rarity-weight` is off by default and weights each lexical hit by
|
||
`log(N/df)` over the bundle's own concepts instead of counting it as one, so a
|
||
requirement number is not worth what a common verb is worth. The default being
|
||
off is a measurement rather than a preference: on four corpora it took one gold
|
||
concept from withheld to delivered and a priced sheet from candidate rank 10 to
|
||
2, left one gold rank unmoved, and cost another seven rank positions — because
|
||
the four-character prefix matcher makes a unique identifier read as
|
||
135-of-446 common on that bundle. Where it cannot help is decomposed rather
|
||
than guessed: RRF fuses RANKS, so a weight moves nothing on a signal the gold
|
||
already leads. `docs/2026-09-08-sjeldenhetsvekt.md`. Its published figures were
|
||
measured under the pre-2026-09-10 tie-break and are not re-measured.
|
||
|
||
`--stem-prefix` is **on** by default since 2026-09-09 (opt out with
|
||
`--no-stem-prefix`), and like `--tie-shared-rank` below it alters a payload
|
||
with no bundle changing. `MIN_SHARED_PREFIX = 4` exists for Norwegian
|
||
compounding, and it also matches four characters that are not a stem: measured
|
||
on the pinned 453-concept bundle with the control run first, `under` occurs 79
|
||
times by equality and matches 172 concepts by prefix, while `bilateral` occurs
|
||
**0** times and matched **400 of 453** through `bilag`, and `standhaftig` 0 and
|
||
219 through `standard`. Three repairs were measured and all three failed on the
|
||
same row — a longer floor (5–8), a coverage share (0.5–0.8) and a
|
||
long-words-only floor (≥ 8) each cost row 1 its rank on the default bundle and
|
||
the whole row on Arm B, because row 1's token `prisene` reaches its gold
|
||
document through `pris|sammenstilling` on four characters. The rule that works
|
||
asks whether the shared prefix is a WORD the bundle uses: `bilateral` 400 → 0
|
||
and 512 → 0, `standhaftig` 219 → 56 and 235 → 33, **every hit@8 row keeping
|
||
rank 1 on both bundles**. `undersjøisk` stops at 162 because `under` is a word
|
||
here — a genuine Norwegian morpheme, so the residual is a different answer and
|
||
not a ceiling.
|
||
|
||
`--tie-shared-rank` is **on** by default since 2026-09-10 (opt out with
|
||
`--no-tie-shared-rank`), and it is the one change in this library that alters a
|
||
payload with no bundle changing — a consumer pinned to the previous excerpt
|
||
order needs the opt-out. RRF emits a rank for every concept in every signal,
|
||
including a signal that scored them all the same, and the declared tie-break
|
||
then orders that group by `concept_id`; the fusion reads alphabetical order as
|
||
if it were a measurement. Shared ranks make a signal that separates nothing
|
||
contribute the same constant to each concept in the group. What it buys is
|
||
general rather than cosmetic: a document the segmenter splits from 1 concept
|
||
into 12 fills that signal's whole top tie group with its own concepts, so the
|
||
one that leads the body signal takes position 11 instead of 1 and the document
|
||
loses fused rank 1 to a single-concept competitor leading nothing — **the
|
||
fusion was punishing fine-graining for being fine-grained**, which put the
|
||
segmentation side and the retrieval side in competition over one number.
|
||
|
||
It shipped OFF on 2026-09-08 because hit@8 fell 5 of 6 to 4 of 6, and that
|
||
figure is real and **conditional**: swept over 2 document-prior exponents x 3
|
||
bundles x 6 rows, the lost row is lost only at exponent 1.0. The exponent moved
|
||
to 0.5 on 2026-09-09 for an unrelated reason, correctly reported as moving no
|
||
hit@8 row, and nobody measured the pair — so a rule sat behind a published
|
||
number that had stopped being true in the same commit. A flag's "off by
|
||
measurement" is a measurement of a *configuration*, not a property of the flag.
|
||
`docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md`.
|
||
|
||
`--title-covered` is **on** by default since 2026-09-10 (opt out with
|
||
`--no-title-covered`), and it is the fourth widening here that alters a payload
|
||
with no bundle changing. A concept whose EVERY title token is a token of the
|
||
question is read before the concepts the fusion ranked above it. The defect it
|
||
repairs is that both lexical signals are unnormalised COVERAGE COUNTS: they
|
||
measure how much of the question a candidate answers, and nothing measures how
|
||
much of the CANDIDATE the question accounts for, so a section titled with the
|
||
question's subject alone scores exactly what a narrower section titled with that
|
||
subject plus a qualifier scores — and then loses on the body count, because a
|
||
longer title and a longer body can only reach more of the question. Measured on
|
||
a 2 761-concept bundle of one standard, where **none of the six flags above
|
||
moved the number at all**: hit@1/8/50 **3 of 6 · 5 of 6 · 5 of 6 → 6 of 6 · 6 of
|
||
6 · 6 of 6** at default `k`, the same 6 of 6 at `--k 50`, the known-positive
|
||
holding rank 1 at both and the known-negative still not a hit. The three
|
||
recovered rows go 4 → 1, not-delivered → 1 and 3 → 1.
|
||
|
||
It is a PARTITION and not a fourth RRF signal, and the arithmetic is the
|
||
reason: RRF consumes ranks alone, so with shared ranks a rule whose positive
|
||
group has `m` members is worth `1/61 − 1/(61 + m)`, and a rule firing on ONE
|
||
concept of 2 761 is worth 0.00026 against a body-signal gap of 0.0029 — **a
|
||
precise rule is worth least under this fusion, backwards from what precision is
|
||
for**. Measured as a signal the same predicate moves hit@1 not at all; as a
|
||
partition it reaches 6 of 6. `lookup_hits` is the same shape for the same
|
||
measured reason, and it still wins: this partition lands below it. The title is
|
||
read by EQUALITY, never by shared prefix — four shared leading characters take
|
||
the group from 1 to 6 on one question and 9 to 31 on another, with the answering
|
||
section falling to candidate rank 6 and the known-positive to 2.
|
||
|
||
Two candidate repairs were measured and felled first. Pivoted length
|
||
normalisation of the body signal collapses at every value swept (b = 0.25, 0.5,
|
||
0.75, 1.0 give hit@8 3, 1, 1 and 0 of 6, and at 1.0 the known-positive falls to
|
||
rank 49), because the median concept holds 22 tokens against a mean of 60 — so
|
||
length normalisation promotes thousands of tiny concepts over the section that
|
||
treats the subject. Title *precision* as a signal reaches candidate hit@1 5 of 6
|
||
and takes the known-positive from rank 1 to 4 every time it does. The reach of
|
||
what shipped is narrow and stated as such: it fires on 6 of 8 questions on that
|
||
bundle and on **0 of 21** measured cells across the pinned K2 bundle, Arm B and
|
||
the three N bundles, whose payloads are byte-identical either way.
|
||
`docs/2026-09-10-k3-runde16-hele-tittelen-tar-ruten.md`.
|
||
|
||
**Since round 17 the partition is bounded by RECALL**, under the same flag. A
|
||
covered title states its PRECISION — it says nothing the question did not ask —
|
||
and round 16 let that claim override the fusion even against a title answering
|
||
MORE of the question: on a 26-concept bundle of five tender documents, a
|
||
question naming a section by three of its title tokens also held a neighbour's
|
||
whole one-token title, and the neighbour took rank 1. A covered concept now
|
||
rises through the fusion's order and stops beneath the first concept whose
|
||
title answers more question tokens, by equality, than it holds. On the
|
||
standard's bundle nothing stood above a covered concept answering more, so its
|
||
8 payloads are byte-identical at both `k`; the 27 payloads where the rule never
|
||
fires are byte-identical too. A minimum title length, a share of the question,
|
||
an order inside the group and a stop list were measured beside it and not
|
||
taken. 0.8.1's unbounded order is reproducible by no flag;
|
||
`--no-title-covered` still gives the pre-0.8.1 order.
|
||
`docs/2026-09-11-k3-runde17-dekningen-stopper-ved-en-bredere-tittel.md`.
|
||
|
||
It emits the § 8 shape — `contract`, `bundle` (`bundle_id` plus a
|
||
`sha256-tree:` content identity), `budget` (unit, instrument, limit, spent and a
|
||
validated known-positive), `denominators`, `excerpts` and `withheld` — and every
|
||
withheld concept is accounted for by the rule that dropped it, from a
|
||
closed set of seven.
|
||
|
||
**`withheld` is counts plus names, not one entry per concept** (revision
|
||
`okf-consumption/2`). It carries the `total`, the same total decomposed
|
||
`by_rule`, the best-ranked drops by name — with title and source document, so a
|
||
reader who sees a near miss can ask for it — and `complete`, which says whether
|
||
those names ARE the whole set. `--withheld-nearest N` sets how many are named
|
||
(default 20) and `--withheld-full` names every one, which is what an instrument
|
||
classifying every miss should ask for. The default moved on a measurement: on a
|
||
large real bundle the flat list came to **65.5 % of the written payload**, none
|
||
of it counted against the budget the same payload reported, and none of it
|
||
anything a reader could act on. The same question after the change costs
|
||
**18.4 %** of what it did before. `--withheld-titles` is retired by
|
||
that change — it existed to buy the one field the near misses now carry.
|
||
|
||
Every excerpt carries the concept's `title`, and — when the producer wrote them
|
||
— `req_number`, the SPEC § 5.1 address `sources`, and **every top-level
|
||
`source_*` key**, by prefix rather than by allowlist: a fixed list names the
|
||
locators its author thought of, and one real bundle locates by
|
||
`source_element_id` on 269 of its 274 concepts. A key the producer did not write
|
||
stays absent rather than arriving empty, and an address this reader cannot
|
||
decode is named (`sources_unreadable`) rather than dropped into the same
|
||
silence. A concept naming the section that encloses it delivers `parent` — that
|
||
concept's `concept_id` and `title` — or `parent_unresolved` when the pointer
|
||
lands nowhere; `okf check` (seventeen rules) refuses a `parent` a reader could
|
||
not follow. The reason is a measurement: with `concept_id` and body text alone, a
|
||
delivered gold concept at rank 1 still left the answer unable to name the
|
||
document it was quoting.
|
||
`considered == withheld + delivered` closes by construction, and the payload is
|
||
refused rather than reported when it does not.
|
||
|
||
Three exit codes, not two: **0** a payload was written, **1** the run happened
|
||
and refused (the budget admitted none of the concepts that answered the
|
||
question, or an asserted `--ref` contradicted the bytes), **2** the run did not
|
||
happen. Collapsing 2 into 1 would report an unread bundle as a failed cut.
|
||
`--ref` is an **assertion**, never an override — the identity is always computed
|
||
from the bytes, because labelling a payload with an identity its bytes do not
|
||
have is the one thing § 3.3 exists to prevent.
|
||
|
||
Check any payload against the skill that will read it:
|
||
|
||
```
|
||
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json
|
||
```
|
||
|
||
`skills/okf-consume/` is an instantiated consumption skill generated by
|
||
`okf skill` from the golden bundle this repository ships,
|
||
`examples/ingest-golden-segmented-okf-v0-2/expected-bundle`, together with the
|
||
payload that proves it: `okf check` on the pair is conformant with 0 findings.
|
||
It is regenerated with the command in `skills/okf-consume/references/README.md`
|
||
rather than edited, and a test holds the shipped bytes to the generator's. Its
|
||
numbers describe that three-concept bundle and nothing larger. The measurement
|
||
behind the consumption skill as a form — hit@8 **5 of 6** questions at rank 1
|
||
on a 629-concept bundle against a chance baseline of **1.35 of 6**, with one
|
||
control that failed — is `docs/2026-09-07-okf-konsumskill-maaling.md`; the copy
|
||
filled by hand for that corpus cannot ship and is no longer the one here.
|
||
|
||
## Judge a bundle: `okf quality`
|
||
|
||
`okf check` is a **contract** check — it asks whether a payload carries what a
|
||
claim must rest on. It is not a quality gate, and that is measured rather than
|
||
conceded: on 2026-09-10 three arms over one corpus all returned 0 findings and
|
||
exit 0 while their hit@k ranged from 6 of 6 to 0 of 6.
|
||
|
||
`okf quality <bundle>` asks the other question, **per file type and with the
|
||
denominator on every line**:
|
||
|
||
```sh
|
||
okf quality .okf/my-bundle
|
||
```
|
||
|
||
Three verdicts and no fourth — `PASS`, `FAIL`, `UNMEASURED` — and a type with
|
||
no measured threshold is **never** `PASS`. Exit codes: **0** judged and clean,
|
||
**1** at least one `FAIL`, **2** the run did not happen, **3** nothing could be
|
||
judged (every row `UNMEASURED`), because exit 0 over a table of unmeasured rows
|
||
would be the silent pass this command exists to stop.
|
||
|
||
Two thresholds exist today, both `structure_null_share` — the share of a type's
|
||
documents that yielded exactly one concept — read off the pinned 43-document
|
||
reference bundle:
|
||
|
||
<!-- quality-thresholds: .pdf=8/32, .docx=2/5 -->
|
||
|
||
| file type | threshold | N |
|
||
|---|---|---|
|
||
| `.pdf` | 8/32 | 32 documents |
|
||
| `.docx` | 2/5 | 5 documents |
|
||
| every type | 0 concepts with an empty body | definitional |
|
||
|
||
Every other type is `UNMEASURED`, including `.xlsx` (2 documents), `.xml`
|
||
(1 document) and `.html` (no bundle measured here). A threshold needs at least
|
||
five documents on both sides — the bundle's and its own — because a `1/1` is
|
||
not a rate.
|
||
|
||
**What a `PASS` is not.** It is a regression bar against a pinned artifact, not
|
||
a claim that the cut found the document's own structure. hit@k needs a question
|
||
set and is not asked here; the measurement that says so, the corpora behind
|
||
each number, and three candidate metrics that were measured and not shipped are
|
||
in [`docs/2026-09-12-g37-terskler.md`](docs/2026-09-12-g37-terskler.md).
|
||
|
||
### Boundary recall: `--fasit`
|
||
|
||
Give the command the boundaries the source itself declares and it adds one
|
||
whole-bundle row, `boundary_share` — the share of them that became a concept:
|
||
|
||
```sh
|
||
okf quality .okf/my-bundle --fasit declared-sections.json
|
||
```
|
||
|
||
The fasit is a JSON list whose rows carry `title` and the key they are matched
|
||
on, `norm` (all whitespace stripped, lowercased). An unreadable one exits **2**,
|
||
never `UNMEASURED`. A boundary counts as recovered in either of two forms — a
|
||
concept whose normalised title equals `norm`, or the pair of the concept's own
|
||
directory and its residual title — because the numbering token a publisher glues
|
||
into a heading lands in the concept *id* on one route and in the *title* on
|
||
another: measured on one 2 761-section standard, the first form alone reaches
|
||
**22 of 2 761** where the two together reach **2 759**.
|
||
|
||
<!-- quality-boundary-threshold: 2759/2761 -->
|
||
|
||
| metric | threshold | N |
|
||
|---|---|---|
|
||
| `boundary_share` | 2 759/2 761 | 2 761 declared boundaries, **1 corpus** |
|
||
|
||
**`--fasit` is an assertion**, the way `okf consume --ref` is: it says this
|
||
bundle is a build of the document the fasit describes. A bundle of another
|
||
product scores near zero and reads `FAIL` — that is the assertion being wrong,
|
||
not the bundle. The bar itself rests on **one product**, which the output says
|
||
on every run. Both facts, the arm it separates (1 148 of 2 761 against 2 759 of
|
||
2 761) and the interval any bar could sit in are in
|
||
[`docs/2026-09-12-g37-terskler.md`](docs/2026-09-12-g37-terskler.md) § 7.
|
||
|
||
## Judge the retrieval: `python3 tools/okf_retrieval_gate.py`
|
||
|
||
A separate question from `okf quality`, and a separate command: quality asks
|
||
what a bundle looks like, this asks whether the payload for a question carries
|
||
the fasit — and whether the payload says so when it does not know.
|
||
|
||
```bash
|
||
python3 tools/okf_retrieval_gate.py # nine rows, one exit code
|
||
python3 tools/okf_retrieval_gate.py --json # the same rows as JSON
|
||
```
|
||
|
||
Nine rows, exit 0 only when every one is green, 1 otherwise, 2 for wrong
|
||
input. Rows 1–4, 6 and 7 run against a synthetic corpus this repository
|
||
generates and six question sets it ships, pinned by sha256: no network, no
|
||
private corpus, no clock. A question set is always an input — `sha256` is
|
||
checked before a byte is measured and a mismatch is exit 2 — because a gold
|
||
set names a consumer's documents and this repository is public. **The corpus
|
||
is pinned the same way** (`SPECS_SHA256`): every row counts against those
|
||
documents, so moving them without moving the pin is exit 2.
|
||
|
||
**It is RED today, on rows 5, 7, 8 and 9**, and each of those is a
|
||
finding rather than a defect in the gate:
|
||
|
||
| row | what it asks | today |
|
||
|---|---|---|
|
||
| 1 | hit@payload, one fasit entry = one unit | 10 of 10 |
|
||
| 2 | every miss carries exactly one class, each forced by its own fixture | 7 of 7 |
|
||
| 3 | the `rule` the payload prints for a withheld fasit is the true one | 5 of 5 |
|
||
| 4 | an uncovered question comes back marked, a covered one does not | 6 of 6 |
|
||
| 5 | a hold-out set, frozen and with its threshold written first | 0 of 1 |
|
||
| 6 | every delivery confirmed against the bundle's own bytes | 10 of 10 |
|
||
| 7 | mechanical mutants of the ranking and the cut, felled | 12 of 14 |
|
||
| 8 | the three real sets, from path + sha256 | 0 of 3 sets, NOT RUN without `--real` |
|
||
| 9 | K2 | 0 of 6, no gold set exists |
|
||
|
||
Rows 3 and 4 were this gate's two findings and both are closed, which is what
|
||
a gate written before the capability is for. Row 3: in a bundle built from ONE
|
||
source document, every concept past the first two carries that document's
|
||
`source_file`, so a concept the RANK had already lost came back withheld as
|
||
`source_quota_exceeded`. A drop now keeps the rule the same cut without the
|
||
quota would have given it, and only a candidate that cut would have delivered
|
||
is named as the quota's — 2 of 5 to 5 of 5. Row 4: the payload had no key a
|
||
consumer could read as "this bundle does not answer that", so an uncovered
|
||
question came back with excerpts and no statement. `coverage` states the terms
|
||
the pre-pass read, the terms no concept in the bundle answers and the terms no
|
||
delivered excerpt answers — facts, no verdict, and the one bar is the gate's
|
||
own `UNANSWERED_BAR = 2/3` — 3 of 6 to 6 of 6.
|
||
|
||
Row 7 reports two survivors with what they moved rather than with a shrug:
|
||
killing the document prior and flattening the fusion (`RRF_K`) each moved
|
||
**0 ranks and 0 deliveries** on these fixtures. Both have a mechanism —
|
||
a question that names its document reaches it through the title-and-id signal
|
||
as well, and `1/(K+r)` is strictly decreasing in `r` for every `K`.
|
||
|
||
Rows 8 and 9 are never green by leaving something out, and since 2026-09-19
|
||
that is enforced rather than stated: row 8 requires **all three** named sets
|
||
(`wiki-20`, `r761-sk2`, `vegnormal-32`) and is NOT RUN until it has them,
|
||
whatever the ones that ran scored — one set of three used to read `6 of 6
|
||
GREEN`. The sets live in other repositories and are read, never written:
|
||
`--real wiki <set.json> <sha256> <bundle>` runs one, and
|
||
`--real vegnormal <set.json> <sha256> "N100=<bundle>,N200=<bundle>"` runs one
|
||
that spans bundles. Row 9 takes `--k2 <set.json> <sha256> <bundle>` in this
|
||
gate's own set shape; without one it stays RED against its recorded
|
||
denominator of six.
|
||
|
||
Granularity is stated on every line and the two forms are never summed: a set
|
||
naming a citation is measured at citation granularity, a set naming only a
|
||
section is measured at concept granularity. **Row 8's own headline is
|
||
therefore at QUESTION granularity** — the one unit all three sets share —
|
||
with the two unit totals printed below it, each with its own denominator. The
|
||
table above reports the gate's DEFAULT run, where row 8 is `0 of 3` and NOT
|
||
RUN because the sets are not here; the last run that was given all three, on
|
||
one machine 2026-09-19, scored **44 of 64 questions**, and below it *7 of 29
|
||
at citation granularity, 38 of 50 at concept granularity*. That figure is not
|
||
reproducible from this repository alone, which is why it is labelled with the
|
||
day and the machine rather than printed as a row.
|
||
|
||
## Consume in Claude Code
|
||
|
||
A folder of documents to an answer a model can cite, in **three lines**. You do
|
||
not need this repository — the first line installs the command, the second
|
||
builds the bundle and writes a skill beside it, the third asks.
|
||
|
||
```sh
|
||
uv tool install "llm-ingestion-okf[extract] @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v1.0.0"
|
||
okf project ~/my-documents
|
||
claude
|
||
```
|
||
|
||
`okf project` writes the bundle to `.okf/<id>/` and a skill to
|
||
`.claude/skills/<id>-consume/` in the **current directory**, then prints what it
|
||
read, what it wrote, and which documents a question cannot reach. Start `claude`
|
||
in that directory and ask in plain language; the generated skill runs the
|
||
pre-pass and the contract check itself and marks every claim with its source.
|
||
|
||
**Running non-interactively.** In print mode the skill needs its tools named
|
||
explicitly, or the model answers without ever reading the bundle and marks
|
||
every premise `undecidable-from-bundle`:
|
||
|
||
```sh
|
||
claude -p --allowedTools=Bash,Read,Grep,Glob "<the question>"
|
||
```
|
||
|
||
Add `Write,Edit` for the mode that produces a document. `--permission-mode
|
||
acceptEdits` alone does **not** do it — measured 2026-09-09 over four runs, the
|
||
`okf consume` call is refused without the explicit tool list. The interactive
|
||
`claude` above needs none of this.
|
||
|
||
`<id>` is the folder's name reduced to `[a-z0-9-]`. Run it once per folder with
|
||
`--id <name>` to have several bundles reachable at once — each skill carries its
|
||
own `bundle_id`, which is what lets a model pick between them. `--out <dir>`
|
||
puts the project somewhere other than the current directory.
|
||
|
||
Measured 2026-09-09 from a fresh `uv tool install` with this repository nowhere
|
||
on the path: 5 documents in, **26** concepts out, a skill carrying **0** paths
|
||
into any checkout, and `okf check` conformant on its own payload (17 rules, 0
|
||
findings; the count was 15 until `bundle_mismatch` landed 2026-09-10 and 16
|
||
until `parent_unfollowable` landed 2026-09-11, and the pair was re-measured each
|
||
time rather than carried over -- the last time from a frozen export of this
|
||
repository, with 26 concepts and 0 checkout paths again). The 2026-09-08 run of the same measurement reported 15 concepts, and
|
||
that number was the defect rather than the result: `okf project` was calling
|
||
`build()` as a function and reading its signature's defaults, which disagreed
|
||
with argparse's on two flags. Two tests now hold the two default sets equal. Before that day the same result took a `PYTHONPATH`, a snapshot of a
|
||
clone, and a generated skill that named that clone by absolute path on four
|
||
lines — so it could not be moved, shared, or run by anyone else.
|
||
|
||
### The same thing in steps, if you want to see the payload
|
||
|
||
```sh
|
||
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
|
||
okf skill --out ./project/.claude/skills/okf-consume-any
|
||
okf consume ./bundle --question "your question" --out /tmp/payload.json
|
||
okf check --skill ./project/.claude/skills/okf-consume-any/SKILL.md --payload /tmp/payload.json
|
||
```
|
||
|
||
A bundle you only have read access to is fine — the generator only reads it.
|
||
|
||
### The honest limits
|
||
|
||
Measured on **four questions** across two bundles, which is a demonstration and
|
||
not a hit rate. The ranking is lexical, and one of the four found a topic the
|
||
bundle **does** cover and did not rank it into the cut — the skill then said so
|
||
with its denominator instead of answering, which is the behaviour the contract
|
||
asks for, but a miss is still a miss.
|
||
`docs/2026-09-08-claude-code-skill-vilkaarlig-bundle.md` has the runs.
|
||
|
||
Two more things the summary tells you and this paragraph will not repeat: a
|
||
document that landed **whole** (no heading, table or numbered outline to cut it
|
||
on) comes back as one excerpt, which the budget often refuses and which often
|
||
does not carry the answer at the place you asked about; and a document that is
|
||
in the folder but **not** in the bundle cannot be quoted at all. Both cases are
|
||
answered `[sourced-not-sufficient]`, and `okf project` names the documents.
|
||
|
||
### The skill that runs this for you
|
||
|
||
`skills/okf-prosjekt/` in this repository is a Claude Code skill (Norwegian)
|
||
that wraps the command above: it takes a folder, runs `okf project`, and reads
|
||
the summary back. Install it for your user account after cloning:
|
||
|
||
```sh
|
||
mkdir -p ~/.claude/skills && cp -R skills/okf-prosjekt ~/.claude/skills/
|
||
```
|
||
|
||
## Serve a bundle over MCP: `okf mcp`
|
||
|
||
Two shapes, one implementation, and the difference is what an agent has to be
|
||
told in advance.
|
||
|
||
```sh
|
||
okf mcp --bundle .okf/my-bundle # one server, one bundle
|
||
okf mcp --root ~/bundles --root ./.okf # one server, every bundle under the roots
|
||
```
|
||
|
||
`--bundle` serves exactly one bundle, fixed at startup; its tools take no
|
||
bundle argument, because there is nothing to choose. `--root` (repeatable)
|
||
serves every bundle found under the given directories and **knows none of them
|
||
by name**: it discovers them per call, so a bundle you add, remove or rebuild
|
||
while the server is running is picked up by the next call. No restart, no
|
||
configuration edit, no code change.
|
||
|
||
Four tools on a multi-bundle server and **three** on a single-bundle one —
|
||
`okf_list` is absent where there is nothing to list — and each one's description
|
||
says why it exists:
|
||
|
||
| tool | what it answers |
|
||
|---|---|
|
||
| `okf_list` | which bundles are reachable right now, with each one's content identity and concept count (multi-bundle servers only) |
|
||
| `okf_describe` | what one bundle is: id, ref, concept count, source documents, and how many concepts carry each conditionally-written field. Omitting `bundle_id` on a multi-bundle server describes them all, as `okf_ask` does |
|
||
| `okf_ask` | one question, one bounded payload of excerpts, each with its bundle id, concept id, title and provenance locators. Omitting `bundle_id` on a multi-bundle server asks them all and splits the budget |
|
||
| `okf_fetch` | one named concept, verbatim, with its frontmatter and locators |
|
||
|
||
**The server carries the working method, because a subagent inherits MCP tools
|
||
and not skills.** Its `instructions` and the `okf_ask` description state the
|
||
short form — read the map, put the question into the bundle's own words, split
|
||
it into sub-questions, read what lay just outside the cut and ask again with
|
||
its words, then write one answer in the questioner's language. Claude Code
|
||
truncates both at 2 KB, so the long form stays in the skill, which has no such
|
||
cap; a test holds the short one under the limit with a control, because a
|
||
truncated method is worse than a missing one.
|
||
|
||
**Nothing is cached between calls, and that is the design.** Every call
|
||
re-reads the directories and recomputes the bundle's content identity, so the
|
||
identity in an answer is a fact about the bytes at the moment of the call
|
||
rather than at startup — a server that answered from yesterday's bundle is the
|
||
one failure you cannot see from the outside. The cost is real and is paid per
|
||
call: on a 2 756-concept bundle the identity is a 0.75 s hash of the whole
|
||
concept tree, and one `okf_ask` is 5.6 s.
|
||
|
||
**Refusals are loud.** A path climbing out of the bundle, a symlink leaving the
|
||
served root, a bundle id nobody answers to, a directory whose manifest cannot be
|
||
read, and a concept above the server's size ceiling each come back as an error
|
||
with a code — never as a plausible-looking empty answer. A concept over the
|
||
ceiling is refused whole rather than truncated: a truncated concept read as
|
||
whole is a wrong answer that looks right. A directory that cannot be read as a
|
||
bundle is **reported** in `okf_list`'s `unreadable`, not skipped.
|
||
|
||
The protocol is written with the standard library only. An MCP SDK would be
|
||
this package's second runtime dependency on the default install path, for four
|
||
JSON-RPC methods and a newline framing — see
|
||
[Requirements](#requirements).
|
||
|
||
`tools/okf_mcp_gate.py` is the eval: it starts the server as a subprocess,
|
||
speaks real stdio to it, and measures six rows. It was written red before the
|
||
server existed, and it is red today on row 2. The measurements, the update
|
||
drill and the limits are in
|
||
[`docs/2026-09-20-mcp-to-varianter.md`](docs/2026-09-20-mcp-to-varianter.md).
|
||
|
||
### One skill for every bundle: `okf skill` and `okf card`
|
||
|
||
**`okf skill --out <dir>` writes one installable skill for ANY bundle. That is
|
||
the default since 2026-09-20**, and `okf project` installs the same one:
|
||
|
||
```sh
|
||
okf skill --out ~/.claude/skills/okf-consume-any
|
||
okf card .okf/my-bundle # the per-bundle numbers, as JSON, on demand
|
||
```
|
||
|
||
The generic skill carries no bundle's id, no ref and no count; it tells its
|
||
reader to run `okf card <bundle>` first. The card is **derived on every run and
|
||
never written into the bundle**, so there is no second artefact that can
|
||
disagree with the bytes beside it. It is therefore never stale, and one skill
|
||
serves every bundle a project holds.
|
||
|
||
`okf skill <bundle> --for-bundle` still writes the per-bundle form, with the
|
||
identity and the numbers measured into the text — which is exactly what makes
|
||
that file stale the moment the bundle is rebuilt. It refuses out loud when it
|
||
was not regenerated (`bundle_mismatch`), so its cost is a stopped session
|
||
rather than a wrong answer; that is why it is no longer the default.
|
||
|
||
Measured on two unrelated bundles: two per-bundle skills are identical on 281
|
||
of 313 and 311 lines. The 62 lines that differ are exactly identity, concept
|
||
count, the conditional-field table, the whole-bundle cost and the payload-cost
|
||
section — the five things a rebuild invalidates.
|
||
|
||
## Implemented scope (v1)
|
||
|
||
The library provides three entry points for getting content into an OKF
|
||
bundle:
|
||
|
||
1. **Spec-based ingestion.** An implementation of the normative ingest
|
||
specification owned by `portfolio-optimiser-commons`: manifest →
|
||
`file`/`sql`/`http` connector → deterministic materialization of
|
||
`ingest-{id}.md` concept files → index generation. Zero model calls in the
|
||
run path; output is reproducible byte-for-byte against golden fixtures.
|
||
2. **Bundle inbox.** A drop directory where common file types are converted
|
||
to OKF concept files. All file-type→text extraction lives in this library:
|
||
`md`, `txt`, `csv`, `json`, `html` and `xml` are handled by the stdlib core;
|
||
`pdf` and the five office formats (`docx`, `xlsx`, `pptx`, `odt`, `rtf`)
|
||
require the optional `[extract]` extra and are rejected fail-fast without
|
||
it. Extracted text passes the security gate before anything is persisted.
|
||
An `xml` file that declares NISO-STS structure (`<standard>` root, or any
|
||
`<sec>`) becomes one heading per titled section, at the section's own
|
||
nesting depth, with the section's `<label>` and `<title>` on one line; a
|
||
`<sec>` carrying only a label is a body line and never a heading, and a
|
||
`<table-wrap>` becomes one markdown table. Any other XML keeps its text in
|
||
document order and gets no invented structure. XML carrying a
|
||
`<!DOCTYPE` is refused unparsed.
|
||
A NISO-STS document that **states who it is** names its own directory:
|
||
`okf build` takes the directory from the document's one `<std-ident>`
|
||
`<doc-number>` (reduced to the id grammar) instead of the file's stem, and
|
||
the `sources` title from `<doc-number>` + `<year>`, then `<title-wrap>`,
|
||
then the file name, whichever is the first that can be written verbatim.
|
||
Stated more than once, or claimed by a second document in the same run, a
|
||
declared name is not used and the file name stays. Each titled section's
|
||
`description` is its own first spec point — the first `<p>` of its first
|
||
direct-child `<sec sec-type="spec">`, whole — and a section with none gets
|
||
no `description` at all; nothing is derived from the title. A point a YAML
|
||
reader could not read verbatim (a `: ` inside it) is left out rather than
|
||
quoted or cleaned up. Measurements:
|
||
[`docs/2026-09-11-k3-runde19-dokumentidentitet-og-frontmatter.md`](docs/2026-09-11-k3-runde19-dokumentidentitet-og-frontmatter.md).
|
||
The drop directory is walked **recursively**, in sorted relative-path order:
|
||
a file at any depth is ingested and records its path relative to the inbox
|
||
root as its `source_file`, while dot-directories and a bundle directory
|
||
sitting inside the inbox are skipped with a reported code.
|
||
|
||
Under the segmented v0.2 profile a concept also points back at the document
|
||
it was extracted from, so an agent citing it can open the original at the
|
||
right place: `sources: [{ resource, title }]` in the spec's own §5.1 form,
|
||
where `resource` is the inbox-relative path, plus a locator per format —
|
||
`source_pages` for a PDF, `source_sheet` and `source_rows` for a
|
||
spreadsheet, `source_lines` otherwise. The locator keys are this library's
|
||
own, because §5.1 has no field for a place *within* a resource; the line
|
||
numbers index the extracted text and say so. Measurements:
|
||
[`docs/2026-09-08-proveniens-k2.md`](docs/2026-09-08-proveniens-k2.md).
|
||
|
||
<!-- extract-formats: .md, .txt, .csv, .json, .html, .htm, .xml, .pdf, .docx, .xlsx, .pptx, .odt, .rtf -->
|
||
3. **External bundle import.** Import and merge of third-party OKF bundles:
|
||
each concept is assessed via the security gate, and only concepts that
|
||
pass are merged, materialized, and linked into the index.
|
||
|
||
## Boundary: security is delegated
|
||
|
||
Security is owned by the sibling package
|
||
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
|
||
(pinned `>=1.2,<2.0`). The division is strict:
|
||
|
||
- **guard** answers "is this content safe to persist?" — scan, sanitize,
|
||
quarantine, fail-secure, provenance stamping.
|
||
- **this library** does the plumbing — connect a source, materialize a
|
||
deterministic OKF bundle, generate the index.
|
||
|
||
No security functionality is reimplemented here.
|
||
|
||
### What is gated today: read this before trusting a door
|
||
|
||
- **Door A (`materialize_bundle`) is ungated.** It calls nothing before
|
||
writing to disk and writes what it is given. A caller materializing
|
||
untrusted content is responsible for gating it.
|
||
- **Doors B and C (`process_inbox`, `import_bundle`) gate through an adapter
|
||
you pass in.** Each takes a `gate` argument; the flow hands it the content
|
||
and obeys the verdict, refusing to persist anything that does not clear the
|
||
guard's non-blocking floor — including a disposition it does not recognise,
|
||
and (at Door C) a concept the gate returned no verdict for. What it cannot
|
||
do is check that your adapter is a real guard: a permissive stub approves
|
||
everything, and the flow will believe it.
|
||
- **`okf build` runs a real guard by default, and names it in the bundle.**
|
||
`--gate` takes `guard-trusted-source` (the default), `guard-user-upload` or
|
||
`none`, and the name is written into the bundle's `log.md` either way, so a
|
||
consumer holding a bundle can tell a screened one from an unscreened one
|
||
without asking. `okf project` takes the same `--gate` with the same default: it is the one flag there that may move a bundle's bytes, and it is there because a command that cannot reach the gate screens by a default nothing said was a choice.
|
||
|
||
That paragraph is new, and the sentence above it was true of our own command
|
||
until 2026-09-15: `okf build` injected a permissive stub and no argument
|
||
anywhere in the package named a gate, so the only path most people use
|
||
screened nothing while the guard sat in `pyproject.toml` as a mandatory
|
||
runtime dependency. It was reported from outside, reproduced here, and the
|
||
cost of each tier was measured before the default was chosen — over the 453
|
||
concept bodies of the pinned reference bundle, `guard-trusted-source` returns
|
||
the persist disposition on 453 of 453 and `guard-user-upload` holds 1 of
|
||
them. Neither waves anything through: an invisible carrier and a CRITICAL
|
||
finding fail secure at both.
|
||
|
||
<!-- cli-default-gate: guard-trusted-source -->
|
||
|
||
- **`--gate none` is still reachable, by name.** The corpus harness reproduces
|
||
published numbers with it, and a caller measuring segmentation alone has a
|
||
legitimate reason to take the gate out of the picture. What changed is that
|
||
asking for it is an act, and the log says `NOTHING WAS SCREENED`.
|
||
|
||
`llm_ingestion_okf.guard_adapter` is the adapter over the real guard, and the
|
||
only module here that imports it — importing the package itself does not:
|
||
|
||
```python
|
||
from llm_ingestion_okf import process_inbox
|
||
from llm_ingestion_okf.guard_adapter import inbox_gate
|
||
|
||
result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
|
||
okf_type="reference", gate=inbox_gate)
|
||
```
|
||
|
||
Two properties of that adapter are worth knowing before you rely on it.
|
||
It screens the **exact bytes it persists** — the guard's `prepare_input`
|
||
bookend prepares text for a model call, which this library never makes, so
|
||
only `screen_output` is used and the screened string is the written string.
|
||
And it **refuses rather than repairs**: a file carrying an invisible
|
||
zero-width or bidi character is rejected, not silently stripped and written.
|
||
Door B screens under the untrusted-upload policy, so any finding at all is
|
||
held back rather than persisted.
|
||
|
||
This section is stated plainly because earlier wording ("calls the guard at
|
||
every persist gate") described the intended end state in the present tense,
|
||
and a consumer reasonably read it as safe-by-default.
|
||
|
||
## Roadmap
|
||
|
||
The library is built in four phases so that every known OKF surface in the
|
||
ecosystem is eventually covered. Each phase has a detailed plan with
|
||
verification criteria:
|
||
|
||
1. Spec-based ingestion (Python) with byte-exact golden fixtures —
|
||
[plan](docs/plan/phase-1-door-a.md).
|
||
2. Bundle inbox and external-bundle import (Python), guard-gated —
|
||
[plan](docs/plan/phase-2-doors-b-c.md).
|
||
3. Configurable bundle contract (types, layers, frontmatter sets, index
|
||
shape, and reserved-file policy as configuration), enabling stricter
|
||
bundle profiles such as `strict-v1` —
|
||
[plan](docs/plan/phase-3-configurable-contract.md).
|
||
4. A `node/` half: a zero-dependency Node/ESM package (importable and
|
||
CLI-invokable, vendored per consumer) providing bundle checking, index
|
||
generation, inbox processing, and document conversion for the OKF
|
||
second-brain plugin ecosystem. The Python and Node halves share the OKF
|
||
contract and fixture suite, not code —
|
||
[plan](docs/plan/phase-4-node-half.md).
|
||
|
||
## Upstream OKF versions
|
||
|
||
The library targets the current latest version of Google's OKF. Support is
|
||
**additive** — a new upstream version arrives as a new profile, never as a
|
||
migration of an existing one — so an *upstream* release does not change the
|
||
bytes an existing profile emits.
|
||
|
||
That guarantee is about upstream, and one profile tracks a second contract as
|
||
well. `DEFAULT` states the ingest-spec owned by `portfolio-optimiser-commons`,
|
||
so when they change that spec, `DEFAULT` follows them. It happened on
|
||
2026-08-09: `generated` moved from `true` to
|
||
`{ by: process:okf-ingest, at: <ingested_at> }`, one changed line per generated
|
||
file. Upgrading across it costs a re-run and nothing more — a profile still
|
||
recognises bundles stamped by earlier versions, so re-running writes in place
|
||
instead of refusing. `DEFAULT` remains OKF v0.1 on every axis upstream owns.
|
||
|
||
| Profile | Contract | Status |
|
||
|---|---|---|
|
||
| `DEFAULT` | commons' ingest-spec layer (OKF v0.1 semantics) | stable |
|
||
| `STRICT_V1` | a consumer's ratified v0.1 contract | stable |
|
||
| `OKF_V0_2` | OKF v0.2 | **provisional**, pre-release only |
|
||
| `STRUCTURED_V1` | `DEFAULT` plus a faceted, derived index | stable |
|
||
| `OKF_LATEST` | alias for the latest version supported as *stable* | currently `DEFAULT` |
|
||
|
||
`STRUCTURED_V1` is `DEFAULT` in every respect but the index. Under it, Door B
|
||
derives each dropped document's title, number, hierarchy and cross-references,
|
||
writes them into the concept's own frontmatter, and carries them into the index
|
||
entry — so a consumer can reason over the bundle rather than only look things
|
||
up in it. Every inferred field is named in a `derived` list, because an
|
||
unmarked heuristic is worse than no heuristic: the consumer cannot know when to
|
||
doubt it. A pointer to a document not dropped yet is rendered `N200?` rather
|
||
than omitted, since a bundle is built up over several drops and an absence that
|
||
leaves no trace is the dangerous kind. Carrying the metadata costs index
|
||
characters — roughly 3x to 6x the flat index, depending on how many facets the
|
||
profile names — and the facet key set is the dial. Design record and
|
||
measurements: [`docs/plan/structure-derivation.md`](docs/plan/structure-derivation.md).
|
||
|
||
`OKF_V0_2` ships first as a pre-release to a named pilot set and may change on
|
||
their feedback without a deprecation cycle. Pin the versioned constant rather
|
||
than `OKF_LATEST` unless you have explicitly opted into tracking; `OKF_LATEST`
|
||
moves at general availability, which is a deliberate release event rather than
|
||
a side effect of an upgrade.
|
||
|
||
Selecting a profile is keyword-only, so existing call sites are unaffected:
|
||
|
||
```python
|
||
materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)
|
||
```
|
||
|
||
A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and
|
||
puts the declaration in the bundle-root `index.md`'s frontmatter block. The
|
||
profile names the key; the **caller supplies the value**, because that value
|
||
tracks the upstream version and is not this library's to decide:
|
||
|
||
```python
|
||
materialize_bundle(
|
||
manifest, bundle_dir, ingested_at,
|
||
profile=OKF_V0_2,
|
||
root_frontmatter_values={"okf_version": "0.2"},
|
||
)
|
||
```
|
||
|
||
Omit the argument and no frontmatter block is written. Offering a key the
|
||
profile does not name is refused before anything is written to disk.
|
||
|
||
### Attested computations (v0.2 §10)
|
||
|
||
`OKF_V0_2` supports the `Attested Computation` type as a **format**: its five
|
||
contract fields — `runtime`, `parameters`, `computation`, `executor`,
|
||
`attester` — are emitted in canonical position, judged, and round-tripped.
|
||
`runtime` is required for that type and for no other, which the profile
|
||
expresses through `FrontmatterSchema.required_by_type`; a type the mapping does
|
||
not name carries no extra requirement, because §14 forbids a consumer to reject
|
||
on an unknown `type`.
|
||
|
||
Nothing here executes a computation or checks an attestation. Upstream defers
|
||
the receipt and verdict wire formats, so there is no contract to implement, and
|
||
the question an attestation answers — was this value produced the sanctioned
|
||
way — is not this library's. It re-enters scope when upstream specifies the
|
||
protocol.
|
||
|
||
On the import side, a third-party concept may name an `executor` or `attester`
|
||
resource pointing at executable code. Door C imports the **pointer** and never
|
||
the code — it writes concepts verbatim and skips every non-`.md` file — so such
|
||
a reference may not resolve, or may resolve to a file the destination tree
|
||
already holds under that path. Each one is reported in
|
||
`ImportResult.unverified_references`; the concept still merges, because §14
|
||
forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer
|
||
to surface rather than silently drop. The report names the pointer key, not the
|
||
resource it points at: recovering the resource needs the structured reader.
|
||
|
||
One limit worth knowing before you write such a concept: §10.2 presents
|
||
`executor` and `attester` as nested block mappings, and this library's
|
||
frontmatter parser is line-oriented. Which form it reads depends on the KEY,
|
||
and the distinction is worth stating precisely, because this paragraph used to
|
||
blur it:
|
||
|
||
- **`sources` in either form.** `consume.read_sources` has always read both,
|
||
and since 2026-09-12 (K3-24) so do all three copies of the flat grammar
|
||
(`materialize.parse_frontmatter`, which is public API, and the two internal
|
||
ones). A block sequence of mappings is decoded into the flow rendering those
|
||
readers round-trip; the entries never enter the document's key namespace.
|
||
Measured against PyYAML 6.0.3 and the pinned guard 1.4.0 on 4 609 of 4 609
|
||
concept files carrying a block `sources`, all three readers return the same
|
||
entries both references do.
|
||
- **Every other key, flow only.** `executor`, `attester` and any other block
|
||
mapping are still skipped rather than parsed: two block mappings that both
|
||
carry a `resource` would collapse into one namespace and the first would be
|
||
lost, so they are refused instead. Reading them needs the structured reader
|
||
(D1b).
|
||
|
||
Write the flow form, and know its limit: it is valid YAML only while every
|
||
plain value inside it avoids what ends a flow scalar — `,`, `[`, `]`, `{`, `}`,
|
||
and for PyYAML also `?` — as well as `": "`, `" #"`, a trailing `:` and a
|
||
leading YAML indicator. Within that limit a YAML consumer recovers the same
|
||
structure from either form. Beyond it no flow form works: quoting satisfies
|
||
PyYAML, but the guard refuses a quoted value inside a flow mapping, so this
|
||
library refuses such a value rather than write frontmatter a reader cannot
|
||
parse. An earlier version of this paragraph said "both are valid YAML" without
|
||
that limit; it was measured false for an unquoted URL with a query string. That
|
||
refusal is also why K3-24 is a reading change only: this library still emits
|
||
flow, because a named downstream consumer classifies a block sequence as
|
||
unreadable provenance even though the guard reads it.
|
||
|
||
## Non-goals
|
||
|
||
- Verdict/feedback machinery from the method specification (stays in the
|
||
consuming repositories).
|
||
- Embedding- or retrieval-layer functionality.
|
||
- Security functionality, in either runtime — that is always
|
||
`llm-ingestion-guard`'s domain.
|
||
|
||
## Requirements
|
||
|
||
Python 3.10+, and exactly one runtime dependency — the security boundary,
|
||
`llm-ingestion-guard>=1.2,<2.0`. Everything else is stdlib. The commands are
|
||
under [Install](#install); what follows is why they look the way they do.
|
||
|
||
From a checkout, the test suite runs with:
|
||
|
||
```
|
||
.venv/bin/python -m pytest
|
||
```
|
||
|
||
The suite is the verification surface for everything above: **1827 tests
|
||
collected, 1826 passed and 1 skipped**, run on 2026-09-12 with the `[extract]`
|
||
extra installed. Both numbers are given because they are two measurements: the
|
||
figure published before the `v0.8.2` release was the PASSED count, and `pytest
|
||
--collect-only -q` reported one more.
|
||
(The figure stood at 596 until 2026-09-09 — measured 2026-08-21 and never
|
||
updated as the suite grew — at 1515 until the `v0.8.0` release, at 1575
|
||
through it, at 1582 through the `v0.8.1` release, at 1602 through the
|
||
`v0.8.2` release, at 1658 after K3-19, at 1667 after K3-20 and through
|
||
the `v0.8.3` release, and at 1783 through the `v0.8.4` release, after K3-22
|
||
and K3-21; the figure above is the `v0.8.5` release's, after K3-23, K3-24 and
|
||
K3-25: a count is a measurement with a date on it.)
|
||
Without the extra the same suite skips the tests covering the parser path;
|
||
that split was last counted on 2026-08-21 as 589 passed and 7 skipped and has
|
||
**not** been re-measured since. The tests holding the fail-fast rejection for
|
||
an uninstalled extra run in both. The suite is not shipped in an installed
|
||
distribution — `tests/` lives at the repository root, so this command needs a
|
||
clone rather than a `pip install`.
|
||
|
||
**The lint acceptance is BOTH of ruff's gates, and the rule set is declared.**
|
||
`ruff check src tests tools` *and* `ruff format --check .`, both named in a
|
||
report with the version they ran under. Neither was true before 2026-09-09:
|
||
the formatter gate was not in the acceptance and had gone red unseen, and
|
||
`[tool.ruff]` set only `line-length` and `target-version`, so the acceptance
|
||
was whatever ruff's default happened to be — which is why the tree read green
|
||
only as long as `uv.lock` froze ruff at 0.15.22. Under 0.16.6 the same
|
||
untouched code reported **148** findings, all of them new rules rather than new
|
||
defects, because 0.16 widened the default set to whole families. `select` is
|
||
now written down (`E4`, `E7`, `E9`, `F`, `I`, `RUF100`), the dev pin is
|
||
`ruff>=0.16.6,<0.17`, and the wider families are a separate decision with 148
|
||
as its starting number. `S` is measured out rather than assumed out: it reports
|
||
**2657** `S101` on a suite whose every assertion is an `assert`. Add
|
||
`--extra extract` to the sync or `mypy src` cannot find `pdfplumber`.
|
||
|
||
**Markdown is excluded from `ruff format`.** ruff 0.16 formats fenced Python
|
||
inside markdown, and the two files it would change here are records rather than
|
||
source — a README call example and a published measurement's *quotation* of
|
||
`COST_VOCABULARY` as it stood when that measurement was taken. Reformatting a
|
||
quotation makes it stop being one.
|
||
|
||
A git URL is a PEP 508 direct reference and pins one exact tag, so it is an
|
||
install-time *channel*, not the pin: the range above stays the declared
|
||
dependency — a wheel built from this branch carries `Requires-Dist:
|
||
llm-ingestion-guard<2.0,>=1.2`, measured 2026-08-23 — and resolves normally
|
||
once the package index exists. A wheel built from a *tag* carries that tag's
|
||
range instead, which is why the install commands pair tag with tag.
|
||
|
||
### Binary extraction
|
||
|
||
The optional `[extract]` extra ships two things: `pdfplumber` (MIT) for `pdf`,
|
||
and `pypandoc-binary` for five office formats. It is opt-in because it pulls
|
||
binary wheels, which the default install must never do — the single runtime
|
||
dependency rule covers the default install and this extra sits outside it.
|
||
|
||
The converter **binary travels inside the wheel** and is resolved by path
|
||
rather than found on `PATH`, with its version asserted against a pin. A host
|
||
carrying a different converter is refused, not silently used: extraction is
|
||
deterministic within a converter version and not across one.
|
||
|
||
These six rows have their reader and their evidence class in the one table
|
||
this README carries: [Supported file types](#supported-file-types), which is
|
||
pinned to the extraction registry cell by cell. A second table here would be a
|
||
copy nothing checks, and a copy of an evidence class is exactly the thing that
|
||
goes false quietly.
|
||
|
||
**`constructed` means what it says, and it is a weaker word than `measured`
|
||
on purpose.** The corpus this work was measured on contains **zero** `pptx`,
|
||
`odt` and `rtf` files. Until 2026-09-09 those three rows were `unmeasured` —
|
||
they worked by construction and had never been checked against a document
|
||
anyone wrote. They have now each been put through end to end on a hand-built
|
||
document with a hand-written fasit, which is more than nothing and is not a
|
||
corpus:
|
||
|
||
- `odt` — **1 of 1** declared headings recovered, 1 concept, 0 characters in
|
||
no segment. N = 1 document.
|
||
- `pptx` — **2 of 2** declared slide titles recovered on a deck that declares
|
||
them (a real `<p:ph type="title"/>` placeholder); **0 of 2** on a deck that
|
||
does not, where the converter writes `Slide 1` / `Slide 2` because it has no
|
||
title to use. That is the converter naming an unnamed slide, not a
|
||
segmentation failure. N = 2 decks.
|
||
- `rtf` — **0** declared headings, because the container has no heading style
|
||
and the author's title is bold text. The proposer therefore proposes
|
||
nothing, and the document reaches the bundle inbox as one whole concept:
|
||
content preserved, structure zero. N = 1 document. This is the one open
|
||
finding of the three.
|
||
|
||
They are not known to be broken; a single constructed document is not a
|
||
denominator, and the distinction is the point.
|
||
|
||
**What stays out.** `.doc` (Word 97) is not supported — the converter does not
|
||
read it. Rastered or scanned PDFs are refused rather than persisted as empty
|
||
concepts, because this library does not do OCR. Drawn content — figures,
|
||
diagrams, shapes — does not survive extraction in any format here, and every
|
||
extraction says so with a warning. Structured table recovery is out of scope.
|
||
|
||
Request it by appending `[extract]` to the package name in whichever install
|
||
command from [Install](#install) you are using — this package is not on an
|
||
index, so a bare `pip install 'llm-ingestion-okf[extract]'` does **not** work
|
||
today, and the error message naming that command is written for the day it
|
||
does. The extra is unreleased: it reaches a consumer through a tag that
|
||
contains it, and no such tag exists yet.
|
||
|
||
Two properties of the extra are worth knowing before depending on its output:
|
||
|
||
- **Extracted text is pinned to an exact parser version.** `pdfplumber` pins
|
||
`pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
|
||
releases with no stability contract. Extraction is deterministic within a
|
||
parser version and not guaranteed across one, so a golden fixture built on
|
||
extracted PDF text is a fixture migration away from any parser upgrade.
|
||
- **Text extraction recovers text, and nothing that is drawn.** Figures,
|
||
diagrams and images have no text to recover — only their captions survive —
|
||
so a bundle built from drawn documents is incomplete by construction. The
|
||
library says so itself: every `pdf` extraction emits an `ExtractionWarning`.
|
||
Structured table recovery is separately out of scope; PDFs enter as prose.
|
||
|
||
The planned Node half targets Node/ESM with zero npm dependencies.
|
||
|
||
## License
|
||
|
||
MIT — see [LICENSE](LICENSE).
|