Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import. Security delegated to llm-ingestion-guard.
Find a file
Kjell Tore Guttormsen dfaf3cc134 feat(propose): Arm F, one unit fold behind a flag, measured against the operator's worksheet [skip-docs]
Order 20260908T133512Z-139864689-from-.claude. First iteration of the
per-file-type directive (operator 2026-09-08 13:05Z), not the last. No
threshold is set: ratifying a bar is the operator's, and setting one inside the
work that produces the measurement would be fitting the bar to the number.

[skip-docs] covers README.md only, and it follows a precedent re-measured this
round rather than quoted: `grep -c` for outline-run, table-grid, Arm C, Arm D
and Arm E returns 0 in README.md and CHANGELOG.md, while --path-prefix, a real
interface change, has a CHANGELOG entry. The rule is "interface and behaviour
changes yes, arm flags no", and --unit-fold is an arm flag that defaults off.
CLAUDE.md IS updated, because its `okf build` bullet enumerates which arms are
off there and would otherwise become false.

FUNN 1, and step 1 asked for it: the reproduction broke. Arm E on HEAD is
byte-identical to the archive on 31 of 33 plans; the two that differ are 2 of 2
spreadsheets in the corpus. The cause is EXTRACTION, not segmentation --
56ae274 writes a workbook as pipe tables, and the sample's price sheet extracts
to 11 048 characters where the worksheet records 100 694, which is the figure
that commit's own message predicts. The consequence is a segmentation
regression against the reference: K3 position 3 went 3 concepts -> 1 under both
Arm D and Arm E, where the operator wants eleven. The mechanism is the orphan
check dropping the sheet heading once a table opens below it (propose.py:461),
already reported there as a ranking regression. Doors unchanged: 43 .err, 4
FAILED, extractable 39/43.

THE MATCH CRITERION WAS WRITTEN DOWN BEFORE ANY CELL WAS SCORED, and it stalls
at 7/12 on the literal calibration gate after three rounds, each revision
recorded. The five failures are not the criterion's: at every one it agrees
with the operator's own (a), (b) or free text and disagrees only with (c).
Column (c) is a RELATIVE judgement ("closest today"); the four K3 categories
are absolute. The only way to reach 12/12 is to define "correct" as "the
closest arm", which reads (c) back out of itself. The dominance gate, declared
in advance as the second reading, holds at 11/12.

ARM F is one rule with three clauses derived from the operator's three, not
twelve special cases, and it only MERGES or DISCARDS: a run of at least
CONTENTS_RUN same-level page-numbered headings is a contents list and goes; a
heading deeper than the unit level folds into its parent, extending the
parent's span; a table folds back into the shorter heading that introduces it,
keeping the HEADING's name. K3 first rater, n=12: 2 coarse / 5 fine / 0
duplicate / 5 correct -- best of four arms, ceiling was 4, two moved, nothing
regressed anywhere.

THE PAPER MEASUREMENT CAME FIRST AND FALSIFIED THE FIRST VERSION. Clause 2 was
letting rule:outline -- Arm D's RECOVERY of an integer numbering run -- vote on
the unit level, which took K3 positions 1, 7 and 9 to 3, 4 and 7 concepts
instead of 17, 34 and 11. A recovered numbering is a heuristic, not a level a
document declares, and the unit worksheet showed the operator ATX and dotted
headings only. Fixed with its own red test; 11 of 12 predictions correct after.

PER FILE TYPE, which is the directive: docx 3 of 3 (solved on this sample), pdf
2 of 8 (lags, unchanged by Arm F, and the remainder is decomposed per position
rather than left as one number), xlsx 0 of 1 (regressed, see FUNN 1). Outside
the corpus, n=1 each: pptx and odt byte-identical, rtf proposes nothing either
way, txt differs and exposes clause 2's fallback.

CONTENTS_RUN swept 1..5 and off. Distance prefers 1; three ships anyway,
because at 1 the body chapter "... i henhold til TEK 17" is deleted for ending
in a number, and no K3 cell differs between 1 and 4 -- the metric prefers a
value that provably deletes a chapter and cannot see the cost.

Whole corpus, all 43 through arm_run in ascending foreground chunks: 32 plans,
491 entries against Arm E's 679, 14 documents changed, 1 plan disappeared
entirely (three drawing-schedule numbers that clause 1 correctly reads as a
contents run) and that is reported rather than special-cased.

THE okf build MECHANISM IS REPRODUCED AND IT IS NOT DOOR B: cli.py calls the
proposer with no arm flag at all, so the shipped build path is Arm B. On a
tender PDF that means no boundary where Arm D finds nine and the reference says
nine. Largest per-file-type gap this round found; it is a default change and
therefore the operator's.

5 tests red first, 1373 -> 1379. ruff clean, mypy --strict clean on 17 files.
K2 consumer bundle unchanged: 1108 files, digest 9cd74519... with the flag off.
No bundle built, no version bump, no tag, no push.

Report: docs/2026-09-08-k3-arm-f-mot-enhetsarket.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 16:28:41 +02:00
docs feat(propose): Arm F, one unit fold behind a flag, measured against the operator's worksheet [skip-docs] 2026-09-08 16:28:41 +02:00
examples feat(inbox): point every concept at the document it came from, with a locator per format 2026-09-08 14:39:24 +02:00
skills docs(consume): the Claude Code recipe, measured end to end on two bundles 2026-09-08 15:32:32 +02:00
src/llm_ingestion_okf feat(propose): Arm F, one unit fold behind a flag, measured against the operator's worksheet [skip-docs] 2026-09-08 16:28:41 +02:00
tests feat(propose): Arm F, one unit fold behind a flag, measured against the operator's worksheet [skip-docs] 2026-09-08 16:28:41 +02:00
tools feat(consume): carry every source_* key by prefix, and generate a skill per bundle 2026-09-08 15:13:22 +02:00
.gitignore chore(gitignore): keep .claude/projects/ local-only 2026-08-31 12:46:43 +02:00
CHANGELOG.md fix(extract,build): write a spreadsheet as pipe tables, stop linking the run log from the index 2026-09-08 10:06:58 +02:00
CLAUDE.md feat(propose): Arm F, one unit fold behind a flag, measured against the operator's worksheet [skip-docs] 2026-09-08 16:28:41 +02:00
LICENSE feat: initial commit — repo scaffold and v1 scope 2026-07-16 10:12:59 +02:00
llms.txt docs(llms.txt): name the uv prerequisite and the pip fallback 2026-08-21 11:32:59 +02:00
pyproject.toml feat(cli): okf build, one installed command for folder in, bundle out 2026-09-07 05:06:33 +02:00
README.md docs(consume): the Claude Code recipe, measured end to end on two bundles 2026-09-08 15:32:32 +02:00
SECURITY.md docs(security): add SECURITY.md with reporting contact and process 2026-08-16 21:15:10 +02:00
uv.lock build(deps): pin llm-ingestion-guard v1.3.0 so the gate reads our own goldens 2026-09-03 20:41:03 +02:00

llm-ingestion-okf

Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import. Security delegated to llm-ingestion-guard.

Status: phases 13 are implemented. Phase 1 (spec-based ingestion) covers manifest validation, the file/sql/http connectors, deterministic materialization, index generation, and the golden fixture suite under examples/. Phase 2 adds the bundle inbox (process_inbox) and external-bundle import (import_bundle), both against an injected persist gate, with llm_ingestion_okf.guard_adapter wiring that gate to the real guard (see below). Phase 3 makes the bundle contract configurable, so types, layers, frontmatter sets, index shape, and reserved-file policy are carried by a profile rather than by constants (see Upstream OKF versions). Binary extraction runs behind the optional [extract] extra: pdf through a PDF parser, and five office formats through a vendored document converter. Three of those five office rows are unmeasured — see Binary extraction. Phase 4 (the Node half) is planned (see docs/plan/).

Install

Python 3.10+. Neither this package nor the guard it depends on is on a package index yet. With uv, one command is enough:

uv pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"

uv resolves the guard on its own, because it reads the [tool.uv.sources] entry in the pyproject.toml of the tag it is installing, and v0.4.0 points that entry at the guard tag below. Measured 2026-07-25 and re-measured 2026-08-20 with an empty uv cache; both runs installed llm-ingestion-guard==0.2.0 + llm-ingestion-okf==0.4.0 and imported clean.

With plain pip, the transitive git dependency does not resolve on its own — install the guard first, or installing this package fails with No matching distribution found for llm-ingestion-guard:

pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v0.2.0"
pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"

The guard tag is paired to the okf tag, not to this branch: v0.4.0 declares llm-ingestion-guard>=0.2,<0.3, which v0.2.0 satisfies and later guard tags do not. main has since moved its own pin to >=1.2,<2.0 (see Requirements); that pin reaches you in the next stable tag, not in the commands above. Reading a pin off this branch and installing it against v0.4.0 is the one combination that fails.

v0.4.0 is the current stable tag. v0.5.0a2 is a pre-release for the named OKF v0.2 pilot set only; pin it only if you are one of them (see Upstream OKF versions).

Build

Installing the package installs one command. A folder of documents in, an OKF bundle out:

okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2

It walks the folder recursively, proposes a segmentation for each document with the mechanical rules, replays those proposals through the bundle inbox, writes the bundle and its log.md, and prints the run's numbers. Every proposal is marked PROPOSED and adjudicated: false — the command segments nothing a human has approved, and says so in the artifact.

The last line that matters is the conservation identity: merged + coded rejections == N, where N is the folder's file count read at run time. The run exits non-zero when it does not hold, and names the unaccounted files, so a pipeline cannot mistake a partial bundle for a complete one.

Flags worth knowing: --segments off ingests each document as one concept and asks for no root values; --plans-dir keeps the proposals instead of discarding them; --report writes the full report to a file as well as stdout. --ingested-at and --proposed-at default to 1970-01-01T00:00:00Z rather than the clock, so two builds of the same folder are byte-identical — a wall-clock default would break rebuild-equals-incremental for every caller who did not pass them.

Measured 2026-09-08 on a 43-file corpus (33 pdf, 5 docx, 2 xlsx, and three files no reader accepts), one okf build invocation replacing the shell loop over tools/ that produced the same corpus's bundle on 2026-09-03:

figure value
N (folder file count, computed) 43
merged 39/43
coded rejections 4/43 (extractor_unknown 3, extractor_empty_pdf 1)
K1b 39 + 4 = 43 = N, exit 0
files written 1108
identical to the 2026-09-03 bundle 1104/1108
wall time 842.82 s total, 19.600 s per file (re-measured 2026-09-08)

The six files that differ are all in the corpus's two spreadsheet documents, and they are the change reported in docs/2026-09-08-prisform-og-loggen-k2.md: a spreadsheet's tables are now written as pipe tables, so each row is one line with its cells delimited rather than padded out to the widest cell in the column. Two concept files are renamed by it, two are removed under their old names, and the two documents' own index.md follow. The root index.md is identical to the stored one again, because this library no longer links the bundle's log.md from it. The command's own byte-identity test compares it against the two scripts at the current commit, where the two agree over the whole tree.

Consume

The other direction: a bundle plus one question in, one bounded, contract-shaped payload out.

python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json

tools/okf_consume.py is the pre-pass docs/consumption-contract.md § 1 defines — the deterministic program that reads the bundle, ranks its concepts, cuts them to a bounded set and emits one payload. It decides nothing about the question; the skill that reads the payload does the judgement. It calls no model, opens no socket, imports nothing outside the standard library and this package, and takes no clock: the same bundle bytes and the same (question, k, limit, cost_vocabulary, reserve_top_rank, rarity_weight) produce byte-identical output.

--cost-vocabulary is off by default and widens one question class: it lets a declared list of cost/price/quantity terms bridge a question and a document that name money with different words. The gate is the question — one naming no such term gets byte-identical bytes either way — and what it does and does not close is measured in docs/2026-09-08-blindsone-below-k-k2.md.

--reserve-top-rank is off by default and answers a different objection: the budget is packed by an exact knapsack, which maximises a SUM of scores and therefore has no opinion about rank, so a top-ranked excerpt costing a large share of the budget is out-summed by many small ones. Measured, that made --k a dial that could EVICT the concept a question was asked about. The flag gives rank one its bytes before the pack runs — after the over_budget_alone pre-exclusion, never before — and the payload then declares budget.reserved. On a 629-concept corpus it changed the delivered set in 2 of 24 measured combinations, both of them that eviction: docs/2026-09-08-blindsone-laas2-budsjett-k2.md.

--rarity-weight is off by default and weights each lexical hit by log(N/df) over the bundle's own concepts instead of counting it as one, so a requirement number is not worth what a common verb is worth. The default being off is a measurement rather than a preference: on four corpora it took one gold concept from withheld to delivered and a priced sheet from candidate rank 10 to 2, left one gold rank unmoved, and cost another seven rank positions — because the four-character prefix matcher makes a unique identifier read as 135-of-446 common on that bundle. Where it cannot help is decomposed rather than guessed: RRF fuses RANKS, so a weight moves nothing on a signal the gold already leads. docs/2026-09-08-sjeldenhetsvekt.md.

It emits the § 8 shape — contract, bundle (bundle_id plus a sha256-tree: content identity), budget (unit, instrument, limit, spent and a validated known-positive), denominators, excerpts and withheld — and every withheld concept names the rule that dropped it, from a closed set of six.

Every excerpt carries the concept's title, and — when the producer wrote them — req_number, the SPEC § 5.1 address sources, and every top-level source_* key, by prefix rather than by allowlist: a fixed list names the locators its author thought of, and one real bundle locates by source_element_id on 269 of its 274 concepts. A key the producer did not write stays absent rather than arriving empty, and an address this reader cannot decode is named (sources_unreadable) rather than dropped into the same silence. The reason is a measurement: with concept_id and body text alone, a delivered gold concept at rank 1 still left the answer unable to name the document it was quoting. considered == withheld + delivered closes by construction, and the payload is refused rather than reported when it does not.

Three exit codes, not two: 0 a payload was written, 1 the run happened and refused (the budget admitted none of the concepts that answered the question, or an asserted --ref contradicted the bytes), 2 the run did not happen. Collapsing 2 into 1 would report an unread bundle as a failed cut. --ref is an assertion, never an override — the identity is always computed from the bytes, because labelling a payload with an identity its bytes do not have is the one thing § 3.3 exists to prevent.

Check any payload against the skill that will read it:

python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json

skills/okf-consume/ is the first instantiated consumption skill: a filled copy of skills/okf-consume-template/ naming this pre-pass, with every per-corpus hole replaced by a measured value. Measured 2026-09-07 on a 629-concept bundle, hit@8 was 5 of 6 questions at rank 1 against a chance baseline of 1.35 of 6 — with one control that failed, and both are in docs/2026-09-07-okf-konsumskill-maaling.md with the honesty limits stated.

Consume in Claude Code

A folder of documents to an answer a model can cite, in three commands. Every command below was run end to end on 2026-09-08 against a nine-document folder and a 270-concept third-party bundle; nothing here is untested.

SRC=/tmp/c1-fresh-src            # the folder of documents
BUNDLE=/tmp/c1-fresh-bundle      # where the OKF bundle goes
PROJECT=/tmp/c1-scratch          # the project you will ask the question from

1. Build the bundle.

okf build "$SRC" --bundle "$BUNDLE" --bundle-id c1-fresh-20260908   --okf-version 0.2 --ingested-at 2026-09-08T00:00:00Z

2. Generate a skill for that bundle, straight into the project's skill directory. The skill is instantiated for these bytes: its id, ref, concept count, per-field denominators, whole-bundle cost and breaking point are all measured from the bundle, and it ships a reference payload the checker accepts.

python3 tools/okf_skill.py "$BUNDLE"   --out "$PROJECT/.claude/skills/c1-fresh-20260908-consume"

Repeat for every bundle you want reachable; each one gets its own skill named after its bundle_id, which is what lets a model pick between them. A bundle you only have read access to is fine — the generator only reads it.

3. Ask. From $PROJECT, in Claude Code:

claude -p "Hvordan skal prisene fylles ut?"

Measured with two bundles installed side by side: the model selected the right skill from the question alone, ran the pre-pass and the contract check itself, quoted the requirement verbatim, and named the document, the requirement number, the source resource and the locator inside it. Across three questions, 0 numbers or identifiers appeared in an answer that were not in the delivered set or in the payload's own identities. On a question the bundle does not cover it answered [sourced-not-sufficient] and reported the denominator rather than inventing an answer.

The pre-pass and the checker are the same two commands the skill runs for you, if you want to see the payload first:

python3 tools/okf_consume.py "$BUNDLE" --question "your question" --out /tmp/payload.json
python3 tools/okf_contract_check.py   --skill "$PROJECT/.claude/skills/c1-fresh-20260908-consume/SKILL.md"   --payload /tmp/payload.json

The honest limits: this was measured on four questions across two bundles, which is a demonstration and not a hit rate. The ranking is lexical, and one of the four found a topic the bundle does cover and did not rank it into the cut — the skill then said so with its denominator instead of answering, which is the behaviour the contract asks for, but a miss is still a miss. docs/2026-09-08-claude-code-skill-vilkaarlig-bundle.md has the runs.

Implemented scope (v1)

The library provides three entry points for getting content into an OKF bundle:

  1. Spec-based ingestion. An implementation of the normative ingest specification owned by portfolio-optimiser-commons: manifest → file/sql/http connector → deterministic materialization of ingest-{id}.md concept files → index generation. Zero model calls in the run path; output is reproducible byte-for-byte against golden fixtures.

  2. Bundle inbox. A drop directory where common file types are converted to OKF concept files. All file-type→text extraction lives in this library: md, txt, csv, json, and html are handled by the stdlib core; pdf and the five office formats (docx, xlsx, pptx, odt, rtf) require the optional [extract] extra and are rejected fail-fast without it. Extracted text passes the security gate before anything is persisted. The drop directory is walked recursively, in sorted relative-path order: a file at any depth is ingested and records its path relative to the inbox root as its source_file, while dot-directories and a bundle directory sitting inside the inbox are skipped with a reported code.

    Under the segmented v0.2 profile a concept also points back at the document it was extracted from, so an agent citing it can open the original at the right place: sources: [{ resource, title }] in the spec's own §5.1 form, where resource is the inbox-relative path, plus a locator per format — source_pages for a PDF, source_sheet and source_rows for a spreadsheet, source_lines otherwise. The locator keys are this library's own, because §5.1 has no field for a place within a resource; the line numbers index the extracted text and say so. Measurements: docs/2026-09-08-proveniens-k2.md.

  1. External bundle import. Import and merge of third-party OKF bundles: each concept is assessed via the security gate, and only concepts that pass are merged, materialized, and linked into the index.

Boundary: security is delegated

Security is owned by the sibling package llm-ingestion-guard (pinned >=1.2,<2.0). The division is strict:

  • guard answers "is this content safe to persist?" — scan, sanitize, quarantine, fail-secure, provenance stamping.
  • this library does the plumbing — connect a source, materialize a deterministic OKF bundle, generate the index.

No security functionality is reimplemented here.

What is gated today: read this before trusting a door

  • Door A (materialize_bundle) is ungated. It calls nothing before writing to disk and writes what it is given. A caller materializing untrusted content is responsible for gating it.
  • Doors B and C (process_inbox, import_bundle) gate through an adapter you pass in. Each takes a gate argument; the flow hands it the content and obeys the verdict, refusing to persist anything that does not clear the guard's non-blocking floor — including a disposition it does not recognise, and (at Door C) a concept the gate returned no verdict for. What it cannot do is check that your adapter is a real guard: a permissive stub approves everything, and the flow will believe it.

llm_ingestion_okf.guard_adapter is the adapter over the real guard, and the only module here that imports it — importing the package itself does not:

from llm_ingestion_okf import process_inbox
from llm_ingestion_okf.guard_adapter import inbox_gate

result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
                       okf_type="reference", gate=inbox_gate)

Two properties of that adapter are worth knowing before you rely on it. It screens the exact bytes it persists — the guard's prepare_input bookend prepares text for a model call, which this library never makes, so only screen_output is used and the screened string is the written string. And it refuses rather than repairs: a file carrying an invisible zero-width or bidi character is rejected, not silently stripped and written. Door B screens under the untrusted-upload policy, so any finding at all is held back rather than persisted.

This section is stated plainly because earlier wording ("calls the guard at every persist gate") described the intended end state in the present tense, and a consumer reasonably read it as safe-by-default.

Roadmap

The library is built in four phases so that every known OKF surface in the ecosystem is eventually covered. Each phase has a detailed plan with verification criteria:

  1. Spec-based ingestion (Python) with byte-exact golden fixtures — plan.
  2. Bundle inbox and external-bundle import (Python), guard-gated — plan.
  3. Configurable bundle contract (types, layers, frontmatter sets, index shape, and reserved-file policy as configuration), enabling stricter bundle profiles such as strict-v1plan.
  4. A node/ half: a zero-dependency Node/ESM package (importable and CLI-invokable, vendored per consumer) providing bundle checking, index generation, inbox processing, and document conversion for the OKF second-brain plugin ecosystem. The Python and Node halves share the OKF contract and fixture suite, not code — plan.

Upstream OKF versions

The library targets the current latest version of Google's OKF. Support is additive — a new upstream version arrives as a new profile, never as a migration of an existing one — so an upstream release does not change the bytes an existing profile emits.

That guarantee is about upstream, and one profile tracks a second contract as well. DEFAULT states the ingest-spec owned by portfolio-optimiser-commons, so when they change that spec, DEFAULT follows them. It happened on 2026-08-09: generated moved from true to { by: process:okf-ingest, at: <ingested_at> }, one changed line per generated file. Upgrading across it costs a re-run and nothing more — a profile still recognises bundles stamped by earlier versions, so re-running writes in place instead of refusing. DEFAULT remains OKF v0.1 on every axis upstream owns.

Profile Contract Status
DEFAULT commons' ingest-spec layer (OKF v0.1 semantics) stable
STRICT_V1 a consumer's ratified v0.1 contract stable
OKF_V0_2 OKF v0.2 provisional, pre-release only
STRUCTURED_V1 DEFAULT plus a faceted, derived index stable
OKF_LATEST alias for the latest version supported as stable currently DEFAULT

STRUCTURED_V1 is DEFAULT in every respect but the index. Under it, Door B derives each dropped document's title, number, hierarchy and cross-references, writes them into the concept's own frontmatter, and carries them into the index entry — so a consumer can reason over the bundle rather than only look things up in it. Every inferred field is named in a derived list, because an unmarked heuristic is worse than no heuristic: the consumer cannot know when to doubt it. A pointer to a document not dropped yet is rendered N200? rather than omitted, since a bundle is built up over several drops and an absence that leaves no trace is the dangerous kind. Carrying the metadata costs index characters — roughly 3x to 6x the flat index, depending on how many facets the profile names — and the facet key set is the dial. Design record and measurements: docs/plan/structure-derivation.md.

OKF_V0_2 ships first as a pre-release to a named pilot set and may change on their feedback without a deprecation cycle. Pin the versioned constant rather than OKF_LATEST unless you have explicitly opted into tracking; OKF_LATEST moves at general availability, which is a deliberate release event rather than a side effect of an upgrade.

Selecting a profile is keyword-only, so existing call sites are unaffected:

materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)

A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and puts the declaration in the bundle-root index.md's frontmatter block. The profile names the key; the caller supplies the value, because that value tracks the upstream version and is not this library's to decide:

materialize_bundle(
    manifest, bundle_dir, ingested_at,
    profile=OKF_V0_2,
    root_frontmatter_values={"okf_version": "0.2"},
)

Omit the argument and no frontmatter block is written. Offering a key the profile does not name is refused before anything is written to disk.

Attested computations (v0.2 §10)

OKF_V0_2 supports the Attested Computation type as a format: its five contract fields — runtime, parameters, computation, executor, attester — are emitted in canonical position, judged, and round-tripped. runtime is required for that type and for no other, which the profile expresses through FrontmatterSchema.required_by_type; a type the mapping does not name carries no extra requirement, because §14 forbids a consumer to reject on an unknown type.

Nothing here executes a computation or checks an attestation. Upstream defers the receipt and verdict wire formats, so there is no contract to implement, and the question an attestation answers — was this value produced the sanctioned way — is not this library's. It re-enters scope when upstream specifies the protocol.

On the import side, a third-party concept may name an executor or attester resource pointing at executable code. Door C imports the pointer and never the code — it writes concepts verbatim and skips every non-.md file — so such a reference may not resolve, or may resolve to a file the destination tree already holds under that path. Each one is reported in ImportResult.unverified_references; the concept still merges, because §14 forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer to surface rather than silently drop. The report names the pointer key, not the resource it points at: recovering the resource needs the structured reader.

One limit worth knowing before you write such a concept: §10.2 presents executor and attester as nested block mappings, and this library's frontmatter parser is line-oriented. It reads inline flow mappings (executor: { resource: …, receipt: [ … ] }) as opaque values that round-trip unchanged, but it cannot read the block form — two block mappings that both carry a resource collapse into one namespace and the first is lost. Write the flow form; both are valid YAML, and a real YAML consumer recovers the same structure from either.

Non-goals

  • Verdict/feedback machinery from the method specification (stays in the consuming repositories).
  • Embedding- or retrieval-layer functionality.
  • Security functionality, in either runtime — that is always llm-ingestion-guard's domain.

Requirements

Python 3.10+, and exactly one runtime dependency — the security boundary, llm-ingestion-guard>=1.2,<2.0. Everything else is stdlib. The commands are under Install; what follows is why they look the way they do.

From a checkout, the test suite runs with:

.venv/bin/python -m pytest

The suite is the verification surface for everything above: 596 tests, run on 2026-08-21 against this branch with the [extract] extra installed. Without the extra the same suite is 589 passed and 7 skipped, measured the same day: the seven cover the parser path, and the tests holding the fail-fast rejection for an uninstalled extra run in both. It is not shipped in an installed distribution — tests/ lives at the repository root, so this command needs a clone rather than a pip install.

A git URL is a PEP 508 direct reference and pins one exact tag, so it is an install-time channel, not the pin: the range above stays the declared dependency — a wheel built from this branch carries Requires-Dist: llm-ingestion-guard<2.0,>=1.2, measured 2026-08-23 — and resolves normally once the package index exists. A wheel built from a tag carries that tag's range instead, which is why the install commands pair tag with tag.

Binary extraction

The optional [extract] extra ships two things: pdfplumber (MIT) for pdf, and pypandoc-binary for five office formats. It is opt-in because it pulls binary wheels, which the default install must never do — the single runtime dependency rule covers the default install and this extra sits outside it.

The converter binary travels inside the wheel and is resolved by path rather than found on PATH, with its version asserted against a pin. A host carrying a different converter is refused, not silently used: extraction is deterministic within a converter version and not across one.

Format Reader Evidence
pdf pdfplumber measured
docx converter measured
xlsx converter measured
pptx converter unmeasured
odt converter unmeasured
rtf converter unmeasured

unmeasured means what it says. The corpus this work was measured on contains zero pptx, odt and rtf files, so those three rows work by construction and have never been checked against a document anyone wrote. They are not known to be broken; they are not known to be right either, and the distinction is the point.

What stays out. .doc (Word 97) is not supported — the converter does not read it. Rastered or scanned PDFs are refused rather than persisted as empty concepts, because this library does not do OCR. Drawn content — figures, diagrams, shapes — does not survive extraction in any format here, and every extraction says so with a warning. Structured table recovery is out of scope.

Request it by appending [extract] to the package name in whichever install command from Install you are using — this package is not on an index, so a bare pip install 'llm-ingestion-okf[extract]' does not work today, and the error message naming that command is written for the day it does. The extra is unreleased: it reaches a consumer through a tag that contains it, and no such tag exists yet.

Two properties of the extra are worth knowing before depending on its output:

  • Extracted text is pinned to an exact parser version. pdfplumber pins pdfminer.six==20260107 exactly, and pdfminer.six ships date-stamped releases with no stability contract. Extraction is deterministic within a parser version and not guaranteed across one, so a golden fixture built on extracted PDF text is a fixture migration away from any parser upgrade.
  • Text extraction recovers text, and nothing that is drawn. Figures, diagrams and images have no text to recover — only their captions survive — so a bundle built from drawn documents is incomplete by construction. The library says so itself: every pdf extraction emits an ExtractionWarning. Structured table recovery is separately out of scope; PDFs enter as prose.

The planned Node half targets Node/ESM with zero npm dependencies.

License

MIT — see LICENSE.