Order 20260908T133512Z-139864689-from-.claude. First iteration of the
per-file-type directive (operator 2026-09-08 13:05Z), not the last. No
threshold is set: ratifying a bar is the operator's, and setting one inside the
work that produces the measurement would be fitting the bar to the number.
[skip-docs] covers README.md only, and it follows a precedent re-measured this
round rather than quoted: `grep -c` for outline-run, table-grid, Arm C, Arm D
and Arm E returns 0 in README.md and CHANGELOG.md, while --path-prefix, a real
interface change, has a CHANGELOG entry. The rule is "interface and behaviour
changes yes, arm flags no", and --unit-fold is an arm flag that defaults off.
CLAUDE.md IS updated, because its `okf build` bullet enumerates which arms are
off there and would otherwise become false.
FUNN 1, and step 1 asked for it: the reproduction broke. Arm E on HEAD is
byte-identical to the archive on 31 of 33 plans; the two that differ are 2 of 2
spreadsheets in the corpus. The cause is EXTRACTION, not segmentation --
56ae274 writes a workbook as pipe tables, and the sample's price sheet extracts
to 11 048 characters where the worksheet records 100 694, which is the figure
that commit's own message predicts. The consequence is a segmentation
regression against the reference: K3 position 3 went 3 concepts -> 1 under both
Arm D and Arm E, where the operator wants eleven. The mechanism is the orphan
check dropping the sheet heading once a table opens below it (propose.py:461),
already reported there as a ranking regression. Doors unchanged: 43 .err, 4
FAILED, extractable 39/43.
THE MATCH CRITERION WAS WRITTEN DOWN BEFORE ANY CELL WAS SCORED, and it stalls
at 7/12 on the literal calibration gate after three rounds, each revision
recorded. The five failures are not the criterion's: at every one it agrees
with the operator's own (a), (b) or free text and disagrees only with (c).
Column (c) is a RELATIVE judgement ("closest today"); the four K3 categories
are absolute. The only way to reach 12/12 is to define "correct" as "the
closest arm", which reads (c) back out of itself. The dominance gate, declared
in advance as the second reading, holds at 11/12.
ARM F is one rule with three clauses derived from the operator's three, not
twelve special cases, and it only MERGES or DISCARDS: a run of at least
CONTENTS_RUN same-level page-numbered headings is a contents list and goes; a
heading deeper than the unit level folds into its parent, extending the
parent's span; a table folds back into the shorter heading that introduces it,
keeping the HEADING's name. K3 first rater, n=12: 2 coarse / 5 fine / 0
duplicate / 5 correct -- best of four arms, ceiling was 4, two moved, nothing
regressed anywhere.
THE PAPER MEASUREMENT CAME FIRST AND FALSIFIED THE FIRST VERSION. Clause 2 was
letting rule:outline -- Arm D's RECOVERY of an integer numbering run -- vote on
the unit level, which took K3 positions 1, 7 and 9 to 3, 4 and 7 concepts
instead of 17, 34 and 11. A recovered numbering is a heuristic, not a level a
document declares, and the unit worksheet showed the operator ATX and dotted
headings only. Fixed with its own red test; 11 of 12 predictions correct after.
PER FILE TYPE, which is the directive: docx 3 of 3 (solved on this sample), pdf
2 of 8 (lags, unchanged by Arm F, and the remainder is decomposed per position
rather than left as one number), xlsx 0 of 1 (regressed, see FUNN 1). Outside
the corpus, n=1 each: pptx and odt byte-identical, rtf proposes nothing either
way, txt differs and exposes clause 2's fallback.
CONTENTS_RUN swept 1..5 and off. Distance prefers 1; three ships anyway,
because at 1 the body chapter "... i henhold til TEK 17" is deleted for ending
in a number, and no K3 cell differs between 1 and 4 -- the metric prefers a
value that provably deletes a chapter and cannot see the cost.
Whole corpus, all 43 through arm_run in ascending foreground chunks: 32 plans,
491 entries against Arm E's 679, 14 documents changed, 1 plan disappeared
entirely (three drawing-schedule numbers that clause 1 correctly reads as a
contents run) and that is reported rather than special-cased.
THE okf build MECHANISM IS REPRODUCED AND IT IS NOT DOOR B: cli.py calls the
proposer with no arm flag at all, so the shipped build path is Arm B. On a
tender PDF that means no boundary where Arm D finds nine and the reference says
nine. Largest per-file-type gap this round found; it is a default change and
therefore the operator's.
5 tests red first, 1373 -> 1379. ruff clean, mypy --strict clean on 17 files.
K2 consumer bundle unchanged: 1108 files, digest 9cd74519... with the flag off.
No bundle built, no version bump, no tag, no push.
Report: docs/2026-09-08-k3-arm-f-mot-enhetsarket.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
23 KiB
llm-ingestion-okf
Context
Shared OKF (Open Knowledge Format) ingestion library. Three entry doors, one boundary rule:
- Door A — spec-based ingestion: implements the normative
ingest-spec.mdowned byportfolio-optimiser-commons(manifest →file/sql/httpconnector → deterministic materialization ofingest-{id}.md→ index generation; zero model calls). This repo IMPLEMENTS the spec; commons keeps authorship. Spec changes the library needs go via commons, never edited locally. The library ships the §11 golden fixtures (byte-exact) for the three door-A source types (ingest-golden-{file,sql,http}/, shipped in9dd86b1). - Door B — bundle inbox: converts dropped files to OKF concepts. The drop
directory is walked RECURSIVELY, sorted by relative path, and a concept's
source_fileis that relative path (/-separated) while its NAME still comes from the basename — so a nested duplicate hits the §3 collision refusal rather than vanishing. Dot-directories and a bundle nested inside the inbox are skipped with a code, never silently, because recursion makes the door's own output reachable as its own input (operator 2026-09-06; the flat listing was not a boundary, it was an absence with no denominator). All file-type→text extraction lives HERE (the guard is text-only). v1 core:md,txt,csv,json,html(stdlib).pdf/docx/xlsxonly via the optional[extract]extra; without it those types are rejected fail-fast. The extra shipspdfplumberforpdf(chosen on ONE measured property: it keeps a requirement table's label and value on the same line where three alternatives do not);docx/xlsxstill ship no parser. Structured table recovery is out of scope — two independent parsers return the same wrong shape, so the breakage is document geometry, not a library choice. PDFs enter as prose, and drawn content (figures) does not survive extraction at all, which everypdfextraction warns about. Under theSTRUCTURED_V1profile Door B additionally DERIVES structure — title (leading heading →titlekey →path.stem), document number, hierarchy, and cross-references — writes it into the concept frontmatter, and projects it into a faceted index entry. Every inferred field is named in aderivedlist; an unmarked heuristic is worse than none. Under the SEGMENTED v0.2 profile a concept additionally POINTS BACK at the original:sources: [{ resource, title }]in SPEC §5.1's form (resourceis the inbox-relative path), plus a locator per format —source_pages,source_sheet+source_rows, elsesource_lines. The locator keys are OURS and must stay top-level: §5.1 has no field for a place within a resource, and the pinned guard rejects every route to putting one inside asourcesentry (non-allowlisted key, nested flow list, quoted scalar), so a locator in the entry would emit bundles Door C could never read back. The unit table is built AT EXTRACTION — a page number cannot be recovered from joined text — andsource_offsetstays.source_linesindexes the EXTRACTED text, never the original's paragraphs: measured, docx<w:p>counts and converted-line counts do not agree on a single one of five documents. Record:docs/2026-09-08-proveniens-k2.md. The index is a PROJECTION recomputed from the whole bundle each round, which is what makes rebuild-from-scratch equal an incremental update byte for byte.DEFAULTis untouched and byte-identical. Record:docs/plan/structure-derivation.md. - Door C — external bundle import: third-party OKF bundles are assessed
per concept via the guard's
okf.import_bundle; only concepts clearing the guard's non-blocking floor are merged/indexed here. Two invariants, both load-bearing: a merged concept is written verbatim (this library's line-oriented frontmatter parser cannot round-trip the block lists the guard's parser accepts, so stamping an external concept would destroy sender data and persist bytes the guard never screened), and ownership is therefore proven by content identity — an occupied target name is re-used only when the bytes there are already identical, never overwritten otherwise.
Boundary rule (non-negotiable, zero overlap): llm-ingestion-guard
(pinned >=1.2,<2.0) answers "is this content safe to persist?" —
scan/sanitize/quarantine/fail-secure/provenance-stamp. This library is
plumbing: connect source → materialize deterministic OKF bundle → generate
index. Never reimplement security; call the guard at persist gates
(prepare_input/screen_output, okf.import_bundle). When in doubt which
side of the boundary something belongs on: ask the operator.
Implementation baseline: the stricter behaviors from
portfolio-optimiser (streaming row caps, utf-8-sig, in-memory staging with
pre-mutation collision gate, validated ingested_at, typed IngestError)
are the library baseline. First consumer: portfolio-optimiser-claude.
Roadmap (phases 1–3 shipped; what follows is demand-driven)
-
Phase 1 — Door A (Python). ingest-spec implementation + the §11 golden fixtures. Consumers:
portfolio-optimiser-claudefirst, thenportfolio-optimiser. -
Phase 2 — Doors B/C (Python). Bundle inbox and external-bundle import, guard-gated.
-
Phase 3 — Configurable bundle contract. Types, layers, frontmatter sets, index shape, and reserved-file policy become config instead of constants; proving consumer is
claude-code-llm-wiki(strict-v1profile). Two consumers hold opposite postures on whether an index is authored or directory-derived, so neither is a library invariant and nothing here enumerates a directory unless the profile says derived. -
Phase 4 — Node half (
node/). Zero-dependency Node/ESM package (importable and CLI-invokable, vendorable per plugin — matching the marketplace precedent) for the second-brain world: bundle check, index generation, inbox split/frontmatter/write, and doc conversion (docx/pdf/eml/html → md). Covers okr, linkedin-studio, ms-ai-architect, and the marketplace catalog. -
Phase 5 — MCP as a way to populate a bundle. NOT COMMITTED; needs-based (operator 2026-08-02, superseding the 2026-07-27 commitment.) No MCP work, and no data-lake or database source types, are undertaken without a stated need.
docs/plan/mcp-bundle-population.mdstays as a design record, not a queue. Its open fork — whether we are the MCP server (an agent calls our doors as tools) or an MCP client (a manifest source type pulling from someone else's server) — no longer blocks anything, because nothing waits behind it. It is a question to answer if a need arrives, not before. This is also whysqlstaying sqlite-only is not a gap: a Postgres driver would be runtime dependency number two, bought for no asked-for use.
The two halves share the OKF contract and fixture suite, not code.
Standing posture (operator 2026-08-02). Phases 1–3 shipped; the library now runs on what it has. Work is defect fixes, improvements, and features that a consumer has actually asked for or that measured feedback shows are needed — not roadmap completion for its own sake. The upstream version policy below is the one exception, and it is not a counterexample: "always latest" is a promise already made to consumers, so an upstream release is the stated need. Phase 4 keeps four named consumers with working implementations to lift, so its need is real but untriggered — it starts when one of them asks, not on a date.
Upstream version policy (standing, non-negotiable)
The library always supports the current latest version of Google OKF. Set by
the operator 2026-07-26. Phases 1–3 were built against v0.1; v0.2 shipped
2026-07-25, so v0.2 support is committed work — not contingent on a consumer
asking for it. Plan: docs/plan/okf-v0.2-alignment.md.
Support is additive, expressed as a new profile, never a migration of existing ones. This is what makes the policy sustainable instead of a recurring crisis, and it is bounded by three facts that do not yield to it:
DEFAULTstates commons' ingest-spec §5 layer — itsgeneratedshape is commons' call, raised there, never patched locally. This fired 2026-08-09: commons ratified and executed the O2 form, soDEFAULTnow stampsgenerated: { by: process:okf-ingest, at: <ingested_at> }and four goldens moved with it. It is not a counterexample to "additive, never a migration" — that rule governs upstream versions, and commons' spec is a separate axisDEFAULTtracks by definition.DEFAULTstays v0.1 on everything upstream owns. Ownership recognition is one-way, so the cost to a consumer stays a re-run: a profile carrying an actor still owns the older literal stamp.STRICT_V1mirrors the proving consumer's ratified contract — changing another repo's contract from here violates O2.okf_version's value belongs to catalog (decision E1).
Rollout is pilot-first. A new upstream version reaches a small pilot set on a
pre-release tag and is revised on their feedback before general availability —
consumers testing real data find what fixtures cannot. OKF_LATEST means the
latest version supported as stable, so flipping that alias is the GA event, not
a merge side effect.
Two invariants fall out: no profile hard-codes an upstream version, and no bundle
declares a version its shape has not earned. The first has a mechanism, not just
an intention: a profile names a key, a caller owns its value. okf_version
is declared through materialize_bundle(..., root_frontmatter_values=...)
because its value tracks the upstream Google version and belongs to catalog
(decision E1) — a constant here would claim a decision we do not own, and would
be the one thing to chase on every upstream release. Where upstream itself defers a
contract — v0.2's attestation receipt and verdict wire formats — the format is
supported and the unspecified runtime is not; it re-enters scope when upstream
specifies it. Because "always latest" decays silently, the release checklist
carries an upstream-version re-check.
Structured frontmatter values are emitted in YAML flow form, never block.
Both are valid YAML and an upstream reader recovers the same structure from
either, but this library's parser is line-oriented: it round-trips a flow
mapping as an opaque value and cannot read the block form at all — two block
mappings sharing an inner key (§10.2's executor and attester, both carrying
resource) collapse into one namespace and the first is lost silently.
Emitting block would produce bundles we cannot read back. Reading it needs the
structured reader (D1b); until then the constraint binds what we write.
Every upstream release runs docs/upstream-okf-upgrade-runbook.md. Pin the
commit, enumerate the whole okf/ tree, read the shipped example bundles and not
only SPEC.md, classify the diff, measure our exposure and each consumer's, plan
additively, pilot before GA, then inform every OKF-consuming repo. The runbook is
not optional and not a summary of good intentions: each step names the concrete
failure it prevents, and all of them are failures that happened during v0.1 → v0.2.
This repo is a black box for its consumers. The target cost of an upstream
release to a consuming repo is a re-run, nothing more: support is additive
(a new profile, never a migration), existing profiles stay byte-stable, new public
parameters are keyword-only with defaults so positional call sites stay
source-compatible, and consumer golden fixtures must not churn. The boundary is
stated every time rather than glossed — the library absorbs shape changes, not
upstream changes to content a consumer authored (v0.2's timestamp and
# Citations supersessions). For that class the deliverable is a measured exposure
report per consumer, sent before they ask.
Phase 4 preconditions (coordination, not unilateral moves):
- Lifts okr's reference implementations (
okf-check.mjs,okf-index.mjs, innboks libs) in agreement with okr and the marketplace catalog; the catalog remains the convention owner and re-pins its shared gate here. - linkedin-studio's
ingest/published/provenance-record grammar stays plugin-local by design (different lifecycle) — do not normalize it. - The Node-side persist gate remains security territory: guard-as-contract (per okr's adoption doc) until a Node guard exists in the security repo. No security reimplementation here, in either runtime.
Non-goals (all phases)
- Verdict/feedback machinery (method-spec) — stays in consumer repos.
- Embedding/RAG/retrieval layers.
- Security functionality — always the guard's domain.
Stack
Python 3.10+. Package llm_ingestion_okf (src layout, hatchling).
Exactly one runtime dependency, ever: llm-ingestion-guard>=1.2,<2.0
(itself zero-dep), landed with the Door B/C persist gates. Everything else is
stdlib, and a packaging test enforces it. Only guard_adapter.py imports the
guard; importing the package does not. Install channel until the package
index exists (a direct reference is a channel, not the pin):
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.2.0".
Binary extraction parsers live behind the [extract] extra only — today
pdfplumber>=0.11.10,<0.12 for pdf. Extracted PDF text is pinned to an
exact transitive parser version (pdfminer.six==20260107), so widening that
range is a fixture migration, guarded by a frozen literal in
tests/test_extract.py; see tests/fixtures/README.md.
Phase 4 adds a node/ half: Node/ESM with zero npm dependencies
(node: builtins only), both importable and CLI-invokable, consumed by
vendoring per plugin rather than npm publishing. The halves share contract
and fixtures, never code.
Conventions
- Type hints everywhere;
mypy --stricttarget. - Determinism is bit-exact:
ingested_atis an explicit required argument (no wall-clock defaults); LF-only output; golden fixtures compared byte-for-byte. - Filenames and titles are normalized to Unicode NFC before use
(
materialize.reduce_to_id_grammar,inbox.process_inbox): macOS/APFS hands filenames over in decomposed form, so anéarrives ase+ combining acute. Without normalizing first, the same visual name (e.g. a Norwegian slugger title like "linkedin-studio") reduces differently depending on which form it arrived in, splitting one title into two generated filenames. - No model calls anywhere in the run path.
- Credentials only as env-var references resolved at runtime; never in manifests, logs, or frontmatter.
- Network access requires an explicit per-run opt-in flag; refuse fail-fast otherwise.
- Conventional Commits:
type(scope): description. - English for all code, docs, and commit messages (public repo).
Commands
- Test:
pytest - Lint:
ruff check .+ruff format --check . - Type check:
mypy --strict src/ - Build a bundle:
okf build <folder> --bundle <dir> --bundle-id <id> --okf-version <v>— the installed console script ([project.scripts]), the packaged form of what used to be a shell loop over twotools/scripts. It is orchestration only: the proposer and the corpus harness live inllm_ingestion_okf.proposeandllm_ingestion_okf.corpus, and thetools/scripts are thin entry points to the same functions so the published reproduction blocks still run. Path scope for a document's proposals is its RELATIVE path minus the extension (the door walks recursively, and two same-named documents in different folders must not collide);--ingested-atand--proposed-atdefault to one shared epoch constant rather than the clock, because a wall-clock default takes rebuild-equals-incremental away from anyone who omits them. Arm C, Arm D, Arm E and Arm F are off and not exposed here -- which is measured, not incidental: the build path calls the proposer with no arm flag at all, so a tender PDF that Arm D splits into nine concepts lands as one (docs/2026-09-08-k3-arm-f-mot-enhetsarket.md, theokf buildsection). - Consume a bundle:
python3 tools/okf_consume.py <bundle> --question "<q>" [--k N] [--limit N] [--out PATH] [--ref IDENTITY]— the pre-passdocs/consumption-contract.md§ 1 defines, and the only reading direction this library has. It lives intools/for the reasonokf_contract_check.pystates for itself: outsidesrc/, so no consumer's install surface changes because it exists. Its entry point isbuild_payload(...)with the CLI a thinmain(), so lifting it intosrc/the day a consumer asks for a wheel-installed command is a move, not a rewrite. Deterministic and offline by construction: no model call, no socket, no clock, stdlib plus this package only. It walks the index tree, never a directory — § 9.2 forbids enumerating one unless the named profile says the index is derived, and measured,entries_match_directoryisTrueforSTRICT_V1alone; the walk loses nothing (629 = 629 on the K2 bundle, controlled in a test against the very method § 9.2 forbids).--refis an assertion, never an override: the emitted identity is always the computed one, because § 3.3 exists to stop a payload being labelled with an identity its bytes do not have. Three exit codes: 0 written, 1 refused, 2 did not run. Every excerpt carries the concept'stitle, plusreq_number, the § 5.1 addresssources, and every top-levelsource_*key by PREFIX — never an allowlist, because a list names the producers its author thought of and one bundle locates bysource_element_idon 269 of 274 concepts. A prefix, never a substring (resource_owneris not a locator). An absent key stays absent and an undecodable address is named (sources_unreadable).sourcesis READ in both YAML forms because the two real bundles disagree (flow 629/629 on one, block 270/270 on the other) — reading block is not a licence to write it, the emission rule is unchanged. Contract § 8 makestitlea MUST (checker codeexcerpt_unnamed) and the rest SHOULD, because they are conditional on the producer. The measurement behind it: rank 1 of 8 on 3 of 3 bundles, correct answer on 1 of 3. - Connect a bundle to Claude Code: `python3 tools/okf_skill.py --out
` instantiates `skills/okf-consume-template/` for THAT bundle — its id,
ref, concept count, conditional-field denominators, whole-bundle cost and
breaking point, all measured, plus a reference payload the checker accepts.
In `tools/` for the reason the other two are, and because a wheel-installed
`okf skill` would emit a command pointing at a file the wheel does not carry.
The form was chosen on a measurement: the contract checker passes the
UNFILLED template and passes a skill built for another bundle, so it cannot
tell the two apart — the choice rests on § 5/§ 6.4/§ 7.6 being per-bundle
numbers a generic skill can only leave as holes or state falsely.
The first instantiated consumption skill is `skills/okf-consume/`; the
measurement behind it, including the control that FAILED, is
`docs/2026-09-07-okf-konsumskill-maaling.md`. **The ranking is this
repository's own choice** — the contract binds a payload, not a retrieval
algorithm (§ 10) — and it has THREE optional widenings, all **off by default**
and all keeping the default payload byte-identical. `--cost-vocabulary`: a
declared cost/price/quantity vocabulary family that bridges a question and a
document naming money with different words, gated on the QUESTION carrying
such a term, so a question without one is byte-identical either way. It moves
a measured case from candidate rank 249 to 10 and does **not** deliver it —
the budget is a second, independent lock. Measured, with two rules falsified
before building and the `k`-sweep that showed a higher `k` can EVICT a gold
concept, in `docs/2026-09-08-blindsone-below-k-k2.md`.
`--reserve-top-rank` is that second lock: the pack is an exact knapsack over a
SUM, so it has no opinion about rank and out-sums a top-ranked excerpt costing
a large share of the budget. The flag gives rank one its bytes first, AFTER
the `over_budget_alone` pre-exclusion and never before, and declares
`budget.reserved` in the payload. It fixes the eviction and does **not** close
the mandate-shaped blind spot (that concept ranks 10, not 1); the budget stays
the caller's decision, because deriving a limit from the corpus was measured
and falsified — two defensible derivations, 49x apart, one of them breaking
the known-positive. `docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
`--rarity-weight` is the third: each lexical hit weighs `log(N/df)` over the
bundle's own concepts instead of 1, so an identifier is not worth what a
common verb is worth. It enters the RANKING and never the GATE — `lexical`
stays a count, because a word every concept carries weighs exactly 0 and a
weighted gate is what `
54a0bc2` falsified. Off by default BY MEASUREMENT: it delivers one of three requirement lookups and takes a priced sheet from candidate rank 10 to 2, leaves one gold unmoved and costs another seven rank positions. Two limits are decomposed rather than guessed, and both are someone else's mechanism: `MIN_SHARED_PREFIX = 4` makes a unique identifier read as 135-of-446 common, and RRF consumes RANKS, so no weighting inside a signal can move a gold that already leads it. `docs/2026-09-08-sjeldenhetsvekt.md`.
Workflow
- TDD: no production code without a failing test first.
- This repo is published PUBLICLY (
open/namespace on Forgejo).STATE.mdanddocs/oppstartsprompt.mdare LOCAL-ONLY (gitignored) — never commit session state or internal briefs. No secrets, sober English prose, no marketing language. - Consumer content stays at form level in public files. Some consumers we
read are private (
claude-code-llm-wikiis, pending an Anthropic ToS assessment). Key names, counts, gate names, profile fields and contract shapes are publishable; page bodies, full title or path lists from a consumer's bundle, and Anthropic-derived prose are not. Findings about a private consumer's data go back to them through coord, never as a file here. This costs nothing — every question this library asks of a consumer is about shapes and key sets — and it is not reversible once pushed. - After
git commit: push to Forgejo (git push origin) immediately. Never GitHub.
Communication patterns
Linking to local files
When pointing to local files in responses, always use markdown link syntax with a descriptive name:
- Use
[Human-friendly name](file:///absolute/path)— never barefile:///...URLs or autolinks<file://...>. - Always use absolute paths. Never
~/or relative paths. - For multiple files, render as a bullet list of named markdown links.
Why: bare file:// URLs only render the first as clickable across multiple
lines. Named markdown links make each entry independently clickable and look
cleaner.