llm-ingestion-okf/CLAUDE.md
Kjell Tore Guttormsen 191de89f41 feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word
Three of round 9's four measured holes, each closed with a rule chosen on a
measurement rather than named as a limit.

`rtf` GIVES 0 SEGMENTS -> 6 of 6 AUTHORED TITLES over N = 4. The container has
no heading style, so the author's title is bold text. The grammar is markdown,
not `rtf`: the converter already writes that title as `**...**` in the same
output every office row produces, so no `rtf`-only heading form exists. Three
parameters were swept over 47 readable documents and ONE carried -- refusing a
line that ends in terminal punctuation takes false-positive lines from 9-12 to
1-2. A maximum title length (unlimited/40/60/80/120) and a
must-stand-between-blank-lines clause are both FLAT, so neither is in the rule.
The last false positive is closed by G1, the principle `_gate_outline` already
carries: recovery yields to declaration. False positives are then 0 of the 31
declaring documents by construction, and 0 of 27 on the corpus. Reach: 2 of 39
corpus documents, both `docx`, 0 of 33 `pdf` and 0 of 2 `xlsx`. Behind
`--bold-title`, default OFF pending the hit@8 measurement; the default bundle
is byte-identical without it.

BOTH ALTERNATIVES THE ORDER NAMED WERE MEASURED AND FELLED. A fourth hand-laid
fixture DECLARES heading styles in a stylesheet and the converter discards
them, emitting the same bold line -- so "read the declared headings out of the
markdown" has nothing to read. `rtf` -> `docx` -> markdown yields 0 ATX
headings on that same document, because the loss is in the `rtf` READER before
any writer sees the style. Fixtures are hand-laid in `make_k2_office.py` with
the fasit written first; they live in their own directory because Door B walks
a drop directory recursively and `k2-office/` reads its N off the listing.

THE PREFIX OVER-MATCH: THREE CANDIDATES MEASURED, ALL THREE FAILED ON ONE ROW.
Re-measured on the pinned 453-concept bundle with the control run first:
`under` occurs 79 times by equality and matches 172 by prefix, `undersjoisk` 0
and 172, `bilateral` 0 and 400 of 453, `standhaftig` 0 and 219. The two extra
known-negatives were FOUND, not chosen -- every 4-character prefix ranked by
document frequency, then a real word taken from the widest. A longer floor
(5-8), a coverage share (0.5-0.8) and a long-words-only floor (>= 8) each cost
row 1 its rank on the default bundle and the whole row on Arm B. Decomposed:
row 1's token `prisene` reaches its gold document through
`pris|sammenstilling` on four characters -- 0.57 of one word and 0.22 of the
other -- so the over-match and the wanted match are one mechanism.

THE FOURTH CANDIDATE IS THE ANSWER: the shared prefix must be a WORD the bundle
uses. `pris` is; `bila` and `stan` are not. `bilateral` 400 -> 0 and 512 -> 0,
`standhaftig` 219 -> 56 and 235 -> 33, every hit@8 row keeping rank 1 on BOTH
bundles. `undersjoisk` stops at 162 because `under` IS a word here -- a genuine
Norwegian morpheme, so that residual is a different answer, not a ceiling. ON
by default (`--no-stem-prefix`), pinned with its own known-negative on the
shipped bytes.

THE SHIM: a path importer holds the object `module_from_spec` made, and
`sys.modules[__name__] = _impl` never reaches it. Measured under both counting
methods -- 3 of 76 public names by `vars()`. One line copies the public names
into this file's globals; the dunder filter is load-bearing, because an
unfiltered copy overwrites `__name__` before the next line uses it as the alias
key. It restores attribute ACCESS and not patch-through, which is why the alias
stays. A CHANGELOG note under 0.7.0 and a shim docstring line say so, since
what the consumer asked for was the note.

Suite 1515 -> 1535.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 23:05:45 +02:00

686 lines
46 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# llm-ingestion-okf
## Context
Shared OKF (Open Knowledge Format) ingestion library. Three entry doors,
one boundary rule:
- **Door A — spec-based ingestion:** implements the normative
`ingest-spec.md` owned by `portfolio-optimiser-commons` (manifest →
`file`/`sql`/`http` connector → deterministic materialization of
`ingest-{id}.md` → index generation; zero model calls). This repo
IMPLEMENTS the spec; commons keeps authorship. Spec changes the library
needs go via commons, never edited locally. The library ships the §11
golden fixtures (byte-exact) for the three door-A source types
(`ingest-golden-{file,sql,http}/`, shipped in `9dd86b1`).
- **Door B — bundle inbox:** converts dropped files to OKF concepts. The drop
directory is walked RECURSIVELY, sorted by relative path, and a concept's
`source_file` is that relative path (`/`-separated) while its NAME still
comes from the basename — so a nested duplicate hits the §3 collision
refusal rather than vanishing. Dot-directories and a bundle nested inside
the inbox are skipped with a code, never silently, because recursion makes
the door's own output reachable as its own input (operator 2026-09-06; the
flat listing was not a boundary, it was an absence with no denominator). All
file-type→text extraction lives HERE (the guard is text-only). v1 core:
`md`, `txt`, `csv`, `json`, `html` (stdlib). `pdf`/`docx`/`xlsx` only via
the optional `[extract]` extra; without it those types are rejected
fail-fast. The extra ships `pdfplumber` for `pdf` (chosen on ONE measured
property: it keeps a requirement table's label and value on the same line
where three alternatives do not); `docx`/`xlsx` still ship no parser.
`pptx` and `md` were measured end to end for the first time 2026-09-10
(`docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md` § 4) on two
hand-built documents, which is more than zero and is not a fasit: `md`
recovers 3 of 4 declared headings, and `pptx` segments per slide only where
the deck's slides carry title placeholders the converter recognises — a deck
whose slides do not lands as ONE concept. A converter attribute also leaks
into concept titles (`{#slide-N}`, `{#sheet-1}`), reaching 2 of 810 files on
the K2 default bundle and 1 of 30 on the operator's test folder; because a
filename is reduced from its title, fixing it RENAMES concept ids a consumer
has already cited, so it is an operator question and not a patch.
Structured table recovery is **out of scope** — two independent parsers
return the same wrong shape, so the breakage is document geometry, not a
library choice. PDFs enter as prose, and drawn content (figures) does not
survive extraction at all, which every `pdf` extraction warns about.
Under the `STRUCTURED_V1` profile Door B additionally DERIVES structure —
title (leading heading → `title` key → `path.stem`), document number,
hierarchy, and cross-references — writes it into the concept frontmatter, and
projects it into a faceted index entry. Every inferred field is named in a
`derived` list; an unmarked heuristic is worse than none.
Under the SEGMENTED v0.2 profile a concept additionally POINTS BACK at the
original: `sources: [{ resource, title }]` in SPEC §5.1's form (`resource` is
the inbox-relative path), plus a locator per format — `source_pages`,
`source_sheet`+`source_rows`, else `source_lines`. The locator keys are OURS
and must stay top-level: §5.1 has no field for a place within a resource, and
the pinned guard rejects every route to putting one inside a `sources` entry
(non-allowlisted key, nested flow list, quoted scalar), so a locator in the
entry would emit bundles Door C could never read back. The unit table is
built AT EXTRACTION — a page number cannot be recovered from joined text —
and `source_offset` stays. `source_lines` indexes the EXTRACTED text, never
the original's paragraphs: measured, docx `<w:p>` counts and converted-line
counts do not agree on a single one of five documents. Record:
`docs/2026-09-08-proveniens-k2.md`. The index is a
PROJECTION recomputed from the whole bundle each round, which is what makes
rebuild-from-scratch equal an incremental update byte for byte. `DEFAULT` is
untouched and byte-identical. Record: `docs/plan/structure-derivation.md`.
- **Door C — external bundle import:** third-party OKF bundles are assessed
per concept via the guard's `okf.import_bundle`; only concepts clearing the
guard's non-blocking floor are merged/indexed here. Two invariants, both
load-bearing: a merged concept is written **verbatim** (this library's
line-oriented frontmatter parser cannot round-trip the block lists the
guard's parser accepts, so stamping an external concept would destroy sender
data and persist bytes the guard never screened), and ownership is therefore
proven by **content identity** — an occupied target name is re-used only
when the bytes there are already identical, never overwritten otherwise.
**Boundary rule (non-negotiable, zero overlap):** `llm-ingestion-guard`
(pinned `>=1.2,<2.0`) answers "is this content safe to persist?" —
scan/sanitize/quarantine/fail-secure/provenance-stamp. This library is
plumbing: connect source → materialize deterministic OKF bundle → generate
index. Never reimplement security; call the guard at persist gates
(`prepare_input`/`screen_output`, `okf.import_bundle`). When in doubt which
side of the boundary something belongs on: ask the operator.
**Implementation baseline:** the stricter behaviors from
`portfolio-optimiser` (streaming row caps, utf-8-sig, in-memory staging with
pre-mutation collision gate, validated `ingested_at`, typed `IngestError`)
are the library baseline. First consumer: `portfolio-optimiser-claude`.
### Roadmap (phases 13 shipped; what follows is demand-driven)
1. **Phase 1 — Door A (Python).** ingest-spec implementation + the §11
golden fixtures. Consumers: `portfolio-optimiser-claude` first, then
`portfolio-optimiser`.
2. **Phase 2 — Doors B/C (Python).** Bundle inbox and external-bundle
import, guard-gated.
3. **Phase 3 — Configurable bundle contract.** Types, layers, frontmatter
sets, index shape, and reserved-file policy become config instead of
constants; proving consumer is `claude-code-llm-wiki` (`strict-v1`
profile). Two consumers hold opposite postures on whether an index is
authored or directory-derived, so neither is a library invariant and
nothing here enumerates a directory unless the profile says derived.
4. **Phase 4 — Node half (`node/`).** Zero-dependency Node/ESM package
(importable *and* CLI-invokable, vendorable per plugin — matching the
marketplace precedent) for the second-brain world: bundle check, index
generation, inbox split/frontmatter/write, and doc conversion
(docx/pdf/eml/html → md). Covers okr, linkedin-studio, ms-ai-architect,
and the marketplace catalog.
5. **Phase 5 — MCP as a way to populate a bundle. NOT COMMITTED; needs-based
(operator 2026-08-02, superseding the 2026-07-27 commitment.)** No MCP work,
and no data-lake or database source types, are undertaken without a stated
need. `docs/plan/mcp-bundle-population.md` stays as a design record, not a
queue. Its open fork — whether we are the MCP **server** (an agent calls our
doors as tools) or an MCP **client** (a manifest source type pulling from
someone else's server) — no longer blocks anything, because nothing waits
behind it. It is a question to answer *if* a need arrives, not before. This
is also why `sql` staying sqlite-only is not a gap: a Postgres driver would
be runtime dependency number two, bought for no asked-for use.
The two halves share the OKF contract and fixture suite, **not code**.
**Standing posture (operator 2026-08-02).** Phases 13 shipped; the library now
runs on what it has. Work is defect fixes, improvements, and features that a
consumer has actually asked for or that measured feedback shows are needed —
not roadmap completion for its own sake. The upstream version policy below is
the one exception, and it is not a counterexample: "always latest" is a promise
already made to consumers, so an upstream release *is* the stated need.
Phase 4 keeps four named consumers with working implementations to lift, so its
need is real but untriggered — it starts when one of them asks, not on a date.
### Upstream version policy (standing, non-negotiable)
**The library always supports the current latest version of Google OKF.** Set by
the operator 2026-07-26. Phases 13 were built against v0.1; v0.2 shipped
2026-07-25, so v0.2 support is committed work — not contingent on a consumer
asking for it. Plan: `docs/plan/okf-v0.2-alignment.md`.
Support is **additive, expressed as a new profile**, never a migration of
existing ones. This is what makes the policy sustainable instead of a recurring
crisis, and it is bounded by three facts that do not yield to it:
- `DEFAULT` states commons' ingest-spec §5 layer — its `generated` shape is
commons' call, raised there, never patched locally. **This fired 2026-08-09:**
commons ratified and executed the O2 form, so `DEFAULT` now stamps
`generated: { by: process:okf-ingest, at: <ingested_at> }` and four goldens
moved with it. It is not a counterexample to "additive, never a migration" —
that rule governs *upstream* versions, and commons' spec is a separate axis
`DEFAULT` tracks by definition. `DEFAULT` stays v0.1 on everything upstream
owns. Ownership recognition is one-way, so the cost to a consumer stays a
re-run: a profile carrying an actor still owns the older literal stamp.
- `STRICT_V1` mirrors the proving consumer's ratified contract — changing another
repo's contract from here violates O2.
- `okf_version`'s *value* belongs to catalog (decision E1).
**Rollout is pilot-first.** A new upstream version reaches a small pilot set on a
pre-release tag and is revised on their feedback before general availability —
consumers testing real data find what fixtures cannot. `OKF_LATEST` means the
latest version supported as *stable*, so flipping that alias is the GA event, not
a merge side effect.
Two invariants fall out: no profile hard-codes an upstream version, and no bundle
declares a version its shape has not earned. The first has a mechanism, not just
an intention: **a profile names a key, a caller owns its value.** `okf_version`
is declared through `materialize_bundle(..., root_frontmatter_values=...)`
because its value tracks the upstream Google version and belongs to catalog
(decision E1) — a constant here would claim a decision we do not own, and would
be the one thing to chase on every upstream release. Where upstream itself defers a
contract — v0.2's attestation receipt and verdict wire formats — the format is
supported and the unspecified runtime is not; it re-enters scope when upstream
specifies it. Because "always latest" decays silently, the release checklist
carries an upstream-version re-check.
**Structured frontmatter values are emitted in YAML *flow* form, never block.**
Both are valid YAML and an upstream reader recovers the same structure from
either, but this library's parser is line-oriented: it round-trips a flow
mapping as an opaque value and cannot read the block form at all — two block
mappings sharing an inner key (§10.2's `executor` and `attester`, both carrying
`resource`) collapse into one namespace and the first is lost silently.
Emitting block would produce bundles we cannot read back. Reading it needs the
structured reader (D1b); until then the constraint binds what we write.
**Every upstream release runs `docs/upstream-okf-upgrade-runbook.md`.** Pin the
commit, enumerate the whole `okf/` tree, **read the shipped example bundles and not
only `SPEC.md`**, classify the diff, measure our exposure and each consumer's, plan
additively, pilot before GA, then inform every OKF-consuming repo. The runbook is
not optional and not a summary of good intentions: each step names the concrete
failure it prevents, and all of them are failures that happened during v0.1 → v0.2.
**This repo is a black box for its consumers.** The target cost of an upstream
release to a consuming repo is **a re-run, nothing more**: support is additive
(a new profile, never a migration), existing profiles stay byte-stable, new public
parameters are keyword-only with defaults so positional call sites stay
source-compatible, and consumer golden fixtures must not churn. The boundary is
stated every time rather than glossed — the library absorbs *shape* changes, not
upstream changes to content a consumer authored (v0.2's `timestamp` and
`# Citations` supersessions). For that class the deliverable is a measured exposure
report per consumer, sent before they ask.
Phase 4 preconditions (coordination, not unilateral moves):
- Lifts okr's reference implementations (`okf-check.mjs`, `okf-index.mjs`,
innboks libs) in agreement with okr and the marketplace catalog; the
catalog remains the convention owner and re-pins its shared gate here.
- linkedin-studio's `ingest/published/` provenance-record grammar stays
plugin-local by design (different lifecycle) — do not normalize it.
- The Node-side persist gate remains security territory: guard-as-contract
(per okr's adoption doc) until a Node guard exists in the security repo.
No security reimplementation here, in either runtime.
### Non-goals (all phases)
- Verdict/feedback machinery (method-spec) — stays in consumer repos.
- Embedding/RAG/retrieval layers.
- Security functionality — always the guard's domain.
## Stack
Python 3.10+. Package `llm_ingestion_okf` (src layout, hatchling).
**Exactly one runtime dependency, ever:** `llm-ingestion-guard>=1.2,<2.0`
(itself zero-dep), landed with the Door B/C persist gates. Everything else is
stdlib, and a packaging test enforces it. Only `guard_adapter.py` imports the
guard; importing the package does not. Install channel until the package
index exists (a direct reference is a channel, not the pin):
`pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.2.0"`.
Binary extraction parsers live behind the `[extract]` extra only — today
`pdfplumber>=0.11.10,<0.12` for `pdf`. Extracted PDF text is pinned to an
exact transitive parser version (`pdfminer.six==20260107`), so widening that
range is a fixture migration, guarded by a frozen literal in
`tests/test_extract.py`; see `tests/fixtures/README.md`.
Phase 4 adds a `node/` half: Node/ESM with zero npm dependencies
(`node:` builtins only), both importable and CLI-invokable, consumed by
vendoring per plugin rather than npm publishing. The halves share contract
and fixtures, never code.
## Conventions
- Type hints everywhere; `mypy --strict` target.
- Determinism is bit-exact: `ingested_at` is an explicit required argument
(no wall-clock defaults); LF-only output; golden fixtures compared
byte-for-byte.
- Filenames and titles are normalized to Unicode NFC before use
(`materialize.reduce_to_id_grammar`, `inbox.process_inbox`): macOS/APFS
hands filenames over in decomposed form, so an `é` arrives as `e` +
combining acute. Without normalizing first, the same visual name (e.g. a
Norwegian slugger title like "linkedin-studio") reduces differently
depending on which form it arrived in, splitting one title into two
generated filenames.
- No model calls anywhere in the run path.
- Credentials only as env-var *references* resolved at runtime; never in
manifests, logs, or frontmatter.
- Network access requires an explicit per-run opt-in flag; refuse fail-fast
otherwise.
- Conventional Commits: `type(scope): description`.
- English for all code, docs, and commit messages (public repo).
## Commands
- Test: `pytest`
- Lint: `ruff check .` + `ruff format --check .`
- Type check: `mypy --strict src/`
- Folder to questionable bundle in ONE command: `okf project <folder>`
`okf build` with the package default into `<out>/.okf/<id>/` plus `okf skill`
into `<out>/.claude/skills/<id>-consume/`, `<out>` defaulting to cwd and
`<id>` to the folder name reduced to `[a-z0-9-]`. It owns NO flag that moves
a bundle's bytes and a test holds it byte-equal to `okf build`; two build
paths would leave every measurement report pinned to a bundle nobody
produces. **That invariant was FALSE from the day those two
flags became defaults until O6 measured it, and the test could not see it:** `cli.build`'s Python SIGNATURE defaulted
`keep_table_heading` and `sheet_section_rows` to `False` while argparse
defaulted both to `True`, and `project.create` calls `build()` as a function,
so it read the signature. Measured on a five-document folder, `okf project`
wrote **15 concepts / 30 files** against `okf build`'s **26 / 52**, the whole
difference in the priced spreadsheet -- the document a question about price
has to reach. The byte-equality test compared `project.create` against the
same `build()`, so both sides carried the same wrong value, and its two
fixture documents had neither a table nor a sheet: **a test and the code
agreeing over a set where the difference cannot appear.** Two tests now hold
it -- one comparing the signature's defaults against argparse's for every
same-typed parameter, one building a document whose concept count actually
moves with the two flags. `skills/okf-prosjekt/` is the Claude Code skill
over it.
- **The generated consumption skill states THREE modes and RELATIVE paths**
(O6, 2026-09-09). Question (the default), hypothesis (decomposed into
premises and answered PER PREMISE as `confirmed` / `refuted` /
`undecidable-from-bundle` -- three literals, no fourth; a weak source is
`[sourced-not-sufficient]` on that PREMISE, because four premises and one
weak source is three answers and one gap), and a task producing a document
(source per claim IN the artefact, an ungrounded paragraph written and marked
rather than dropped, the cut declared inside the document because the
document travels without the chat). The five markings are untouched -- the
modes add no sixth. In the layout `okf project` writes, the commands are
`okf consume .okf/<id>` and `okf check --skill
.claude/skills/<id>-consume/SKILL.md`, runnable from the project root, which
is where `okf project`'s own closing line tells the reader to start `claude`;
a path OUTSIDE that root stays absolute on purpose, since `../../..` is not
more portable, only harder to read. Two absolute paths to zero -- and
**O5's published "4 absolute paths -> 0" was measured with `grep -c "^/"`
against paths indented by two spaces**, a query that could not have found one
either way, so the zero was never a measurement. Every path assertion here
runs its pattern against a known-positive first.
- Build a bundle: `okf build <folder> --bundle <dir> --bundle-id <id>
--okf-version <v>` — the installed console script (`[project.scripts]`),
the packaged form of what used to be a shell loop over two `tools/`
scripts. It is orchestration only: the proposer and the corpus harness
live in `llm_ingestion_okf.propose` and `llm_ingestion_okf.corpus`, and
the `tools/` scripts are thin entry points to the same functions so the
published reproduction blocks still run. Path scope for a document's
proposals is its RELATIVE path minus the extension (the door walks
recursively, and two same-named documents in different folders must not
collide); `--ingested-at` and `--proposed-at` default to one shared epoch
constant rather than the clock, because a wall-clock default takes
rebuild-equals-incremental away from anyone who omits them.
**TEN segmentation rules are REACHABLE here, and since 2026-09-09 ALL
TEN are ON by default** -- the tenth is `--contents-name`, round 9's repair
of clause 1, which admits a title into a contents run only when a NAME
survives stripping its page number. Measured: clause 1 discarded 68
candidates over 11 of 39 readable documents, 19 of them over 5 documents
rows of a drawing's dimension chain, a schematic's labels, a door schedule,
a coordinate column and a soil-layer table. The threshold is SWEPT
(`propose.CONTENTS_NAME_RUN = 2`) and collapses at both ends: at 1 it
rescues 13 of 19, at 3 the two-letter section name `VA` stops being a name
and takes four REAL contents entries with it. At 2 it rescues 16 of 19 and 0
of 49. Corpus 429 -> 447 candidates, K2 436 -> 453 concepts / 832 -> 865 md
(`21af4a1aa98315cf...`), the 12-position reference label-identical in BOTH
readings and hit@8 `[1,1,1,1,1,None]` on the new bundle AND Arm B. Opt-out
`--no-contents-name`.
**An ELEVENTH flag, `--bold-title`, is OFF** (round 10, 2026-09-09): the rule
for the type whose container declares nothing. `rtf` measured 0 of 0 declared
headings, 0 concepts, 1368 of 1368 characters in no segment. The grammar is
MARKDOWN, not `rtf` -- the converter already writes the author's bold title
as `**...**` in the same output every office row produces, so no `rtf`-only
heading grammar exists, the same shape of decision as the PDF font reader's
ATX form. Three parameters swept over 47 readable documents and ONE carried
(refusing a line that ends in terminal punctuation: false-positive lines
9-12 -> 1-2); a maximum title length and a stand-between-blank-lines clause
are both FLAT and neither is in the rule. The last false positive is closed
by G1, `_gate_outline`'s own principle, so false positives are **0 of the 31
declaring documents**. Reach **2 of 39** corpus documents, both `docx`, **0
of 33 `pdf`**. **BOTH alternatives the order named were measured and
FELLED**: a hand-laid fixture that DECLARES a heading style has it discarded
by the converter, and `rtf` -> `docx` -> markdown yields 0 ATX headings on
that same document, because the loss is in the `rtf` READER before any
writer.
The nine below are unchanged -- `--outline-run 3`, `--table-grid` and
`--unit-fold` since 2026-09-08, `--drop-wrapped-outline` and
`--outline-gate` since 2026-09-09, `--sheet-section-rows`,
`--keep-table-heading` and `--first-span-from-zero` since 2026-09-10, and
`--close-span-gaps` since 2026-09-11, each
with an explicit opt-out (`--outline-run 0`, `--no-table-grid`,
`--no-unit-fold`, `--keep-wrapped-outline`, `--no-outline-gate`,
`--no-sheet-section-rows`, `--no-keep-table-heading`,
`--no-first-span-from-zero`, `--no-close-span-gaps`) that together reproduce
the pre-move bytes -- measured, `diff -rq` 0 differences, not asserted. **The 2026-09-09 pair is one
decision and cannot be split**: the gate takes `pdf` from 2 of 8 to 5 of 8
and the pair takes it to 7 of 8 (the sheet 5 of 12 -> 10 of 12, `docx`
unchanged at 3 of 3). **The gate is G1+G2:** Arm D's RECOVERED headings are
admitted only where the document DECLARES none of its own -- which is
`fold_units` clause 2's principle moved from voting to admission -- plus any
one recovered heading covering `propose.OUTLINE_SHARE` (0.20, swept flat
from 0.10 to 0.30 and collapsing at both ends) of the text. It filters at
ADMISSION, before spans close, so the text a removed mark opened is carried
by the mark above; the post-filter form scores identically and loses that
text, which is why only one of them shipped. **The bar it had to clear is
now the bar**: reference cells up AND hit@8 holding rank 1 on every row on
every bundle. `--sheet-section-rows --keep-table-heading` reaches 11 of 12
and SHIPPED 2026-09-10, after two rounds off. It was held back because on a
K2 bundle built with it row 1 fell rank 1 -> 2 (the gold document goes 1
concept -> 12), under both prior exponents. **That was never these rules'
defect and it is not a segmentation question**: RRF emits a distinct rank
for every concept in a signal that scored them all EQUALLY, so the gold
document's own twelve concepts fill the document-prior tie group and the one
leading the body signal takes position 11 instead of 1. The repair is the
reading side's `consume.DEFAULT_TIE_SHARED_RANK`, and with it every
acceptance condition holds at once. **The fusion was punishing fine-graining
for being fine-grained**, which put the segmentation side and the retrieval
side in competition over one number for two rounds. Arm E joined a session after the other two, on a number measured
AFTER the first move: without it Arm F's table clause has no joined table to
fold, and the shipped D+F default scored 2 of 12 with `docx` 0 of 3 against
the 5 of 12 the fold was published with. **The proposer's own defaults did NOT move** (`propose.py`'s rules stay
off): the goldens and every published reproduction block are pinned to them,
so the two layers disagree on purpose and `cli.DEFAULT_OUTLINE_RUN` /
`cli.DEFAULT_UNIT_FOLD` say where. The cost to a consumer is a re-run and it
is not small: the 43-document reference corpus goes 629 concepts / 1108 files
(the delivered 2026-09-03 tree) to 492 / 944 after the 2026-09-08 move and to
425 / 810 after the 2026-09-09 one (`bdf4977ca5a443c4...`) and to
**436 / 832** after the 2026-09-10 one (`8dff8a8e6c15d2f7...`, default flags,
default epoch stamp, measured on `38104b7` + this round). On the operator's
own five-document folder the last move is 15 concepts / 30 files -> 26 / 52. Digests published
before 2026-09-09 were computed with a path-DEPENDENT command and are not
comparable to this one; the reproducible form is `find . -type f | sort |
xargs shasum -a 256 | shasum -a 256` from inside the bundle, under which the
previous default is `862116da16e422f6...`. The pinned artifact lives at
`~/corpora/okf-telling-20260829/K2-bundle-default-20260910` and
`tests/test_default_bundle_pin.py` holds its concept count AND its per-row
hit@8 ranks -- the count alone survived a configuration that lost a rank,
which is how a previous round's regression hid. Since 2026-09-10 it also
holds the KNOWN-NEGATIVE on the same bytes: read with
`--no-tie-shared-rank`, the shipped default bundle reproduces the very fall
the rules were held back for, so the pin names its own cause instead of
being green for an unstated reason. **And the number the
decision cites belongs to another configuration:** Arm F's 5 of 12 was
measured with `--table-grid` ON; without it the same sample scores 2 of 12
and `docx` 0 of 3, because the fold's table clause has no joined table to
fold. The ten: `--contents-name` (round 9), `--outline-run N` (Arm D),
`--table-grid` (Arm E),
`--unit-fold` (Arm F), `--keep-table-heading` (D1), `--sheet-section-rows`
and `--drop-wrapped-outline` (both D3), `--outline-gate` (G1+G2),
`--first-span-from-zero` and `--close-span-gaps`, each passed to the
proposer unchanged. **The last two are the same defect at two ends and
NEITHER is a segmentation rule**: a mark removed after its neighbour's span
was closed takes that text out of the plan. Round 8 measured the whole
remainder -- 43 631 characters, 2.51 %, over 8 of 32 documents with a plan --
down to **0**, with the concept count identical at 436 and every hit@8 row
holding rank 1 on both K2 bundles and the reference sheet label-identical at
11 of 12. Three steps leak: the orphan check (18 527 characters over 15 of
39 documents), `fold_units` clause 1 between entries (7 514) and the same
clause on the last run (all 17 590 tail characters; with `unit_fold=False`
the corpus tail gap is 0). **Round 7's own § 5 does not reproduce**: it
reports `md` at 3 of 4 declared headings and a `rule:table-block` displacing
`## 3 Prising`, but D1 -- the repair for exactly that -- became the default
in the same commit, so the number describes the configuration that existed
before the move. On `a364ef4` the default recovers **4 of 4**, and
`tests/test_md_declared_headings.py` now holds the cell with its cause as a
known-negative. Report:
`docs/2026-09-11-k3-runde8-tabellblokk-og-siste-spenn.md`. That last
one is ON since 2026-09-10 and is not a segmentation rule at all -- it adds
no boundary, and the K2 concept count is identical with and without it
(425 = 425 on the 2026-09-09 default). It repairs a measured loss: **32 of
the 32** documents that get a plan left the text above their first concept
in NO segment. **The hole is bigger than that rule, and this is the number
to carry:** measured 2026-09-10, the pre-move default left **207 435
characters, 11.92 %** of the corpus in no segment -- 163 804 above the first
entry, 26 041 BETWEEN entries, 17 590 after the last. The rule closes the
first part entirely and 79 % of the whole; **43 631 characters, 2.51 %, over
8 of 32 documents remain**, and the between-part has a named mechanism (a
`rule:table-block` candidate displacing a DECLARED heading and opening below
it). Neither remainder is a ceiling; both are in STATE with their numbers. Until that day the build path called the proposer with no
arm flag at all, so a tender PDF that Arm D splits into nine concepts landed
as one -- a build path a full arm behind the proposer. Exposing them was not
the same decision as moving one, and the two were taken a session apart:
**which arm ships as the default is the operator's**, answered 2026-09-08 as
above. Adding a flag still leaves the default byte-identical (measured by
digest before and after, and by Arm E over all 43 corpus documents); MOVING
the default is the one thing that does not, which is why it took an operator
decision and carries an opt-out. Arm C
(`--max-segment-chars`) stays unexposed: no reference has ever been measured
for its cap. The two D3 rules read grammars nothing else here reads: a table
row's FIRST CELL (a run of bare numeric labels cuts the block that holds
them, which is the only way to reach a sheet whose units are rows and the
opposite direction from Arm E), and whether a RECOVERED heading's line is a
wrapped sentence (a heading is a complete line; quoted regulation and a
recovered table row are not). Both remove or add nothing anywhere else: over
the 43-document corpus they change 1 and 5 of 39 readable documents, and
**0 of 5 `docx` either way**. Reports:
`docs/2026-09-08-k3-arm-f-mot-enhetsarket.md`,
`docs/2026-09-08-k3-runde2-per-filtype.md` and
`docs/2026-09-08-k3-runde3-per-filtype.md`.
- **Three PDF READER flags, all off, and they sit BEFORE every segmentation
flag** -- an arm changes how the proposer cuts a text, these change what the
text says. `--pdf-headings font` infers a heading from typography (dominant
font size above the document's character-weighted body median AND a bold font
name -- the CONJUNCTION measured at recall 1.000 / precision 0.846, where
adding weight as a disjunct took precision 0.786 -> 0.524) and emits it as an
ATX heading in the SAME markdown the office path produces, so `_ATX` applies
unchanged and **no PDF-only heading grammar exists**. It is off **by
measurement, not by caution**: against the operator's unit worksheet it takes
`pdf` from **2 of 8 to 0 of 8**, losing two exact matches, because on those
documents the outline rule already recovers the document's own numbered
chapters and a second heading source can only add. The cost of the ATX form is
named rather than hidden: a font-inferred heading carries `rule:heading` and is
indistinguishable in the artifact from one the document declared, which is why
`RULE_POPPLER_SIZE_AND_BOLD` was deliberately NOT assigned to it -- that name
records a poppler measurement on a path that cannot ship. `--ocr` reads a page
as an IMAGE when its own text never arrived (empty, or `(cid:N)` codes at or
above `OCR_CID_SHARE = 0.10`, a threshold READ OFF the measured per-page
distribution: 834 pages over 32 files, 818 at exactly 0.0 and 16 at 0.93 or
above, nothing between). Its engine is the optional `ocr` group
(`rapidocr`/`onnxruntime`/`pypdfium2`) and **never** a runtime dependency; a
packaging test pins both halves, and without it every affected file is a coded
rejection (`extractor_ocr_group_missing`), never a crash. `--ocr` can never
become a default -- an optional dependency in the default path would make an
ordinary install fail on the first scanned page. Report:
`docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md`.
**`--pdf-headings font-reserve`** is the third value on that same option
(`none`, `font`, `font-reserve` — three answers to one question, so no caller
can ask for two at once): the same typographic rule applied ONLY where Arm D's
outline gate admits no run at all, typography as a second heading source where
there is no first one. The condition lives in ONE function
(`propose.heading_reserve_applies`) that the proposer and the door both
consult, the door receiving it as a callable the way it already receives
`gate` — a plan indexes the exact string it was proposed against, so a reserve
firing on one side only would turn every document it touches into a coded
rejection. It reads the gate AS CONFIGURED, so at `--outline-run 0` it is
unconditional and equals round 4's "font instead of Arm D". **Off, and the
measurement is that it changes nothing measurable:** on the twelve-position
reference it alters **not one cell** — the five positions where it fires are
one PDF whose glyphs carry no ToUnicode mapping and four office documents the
PDF reader never touches — and the position it was built for has **three**
outline runs, so the reserve is silent there by construction. Its reach is
real but unrated: **4 of 39** readable corpus documents, none in the sample.
Report: `docs/2026-09-08-k3-runde5-hitat8-og-skriftakse.md`, which also
corrects two of this repository's own published figures — the S7 candidate
ranks (96 of 629 / 159 of 492 were measured with the cost vocabulary reaching
only half the ranker; consistently scored they are **10 of 629** and **19 of
492**) and round 4's attribution of that concept's non-delivery to the default
move (it is not delivered on the Arm B bundle either, by a different
mechanism). hit@8 over the six published questions holds at **5 of 6 on both
K2 bundles**, so the default move cost the retrieval side nothing.
- Consume a bundle: `okf consume <bundle> --question "<q>"
[--k N] [--limit N] [--out PATH] [--ref IDENTITY]` — the **pre-pass**
`docs/consumption-contract.md` § 1 defines, and the only reading direction
this library has. **It moved into the package 2026-09-08 (O5)** and the move
it was written for is the one that happened: `build_payload(...)` was always
the entry point with the CLI a thin `main()`, so it was a move and not a
rewrite. What forced it was the generated skill — from `tools/` it emitted
`python3 <absolute path>/tools/okf_consume.py`, so the skill could not be
moved, shared or run by anyone without that clone. `tools/okf_consume.py`
remains as an ALIAS (`sys.modules[__name__] = _impl`, never a re-export: a
re-export binds copies, and a caller patching one patches a binding the
implementation never reads). Deterministic and offline by construction: no model call, no socket,
no clock, stdlib plus this package only. It **walks the index tree, never a
directory** — § 9.2 forbids enumerating one unless the named profile says the
index is derived, and measured, `entries_match_directory` is `True` for
`STRICT_V1` alone; the walk loses nothing (629 = 629 on the K2 bundle,
controlled in a test against the very method § 9.2 forbids). `--ref` is an
**assertion**, never an override: the emitted identity is always the computed
one, because § 3.3 exists to stop a payload being labelled with an identity
its bytes do not have. Three exit codes: 0 written, 1 refused, 2 did not run.
**Every excerpt carries the concept's `title`**, plus `req_number`, the § 5.1
address `sources`, and **every top-level `source_*` key by PREFIX** — never an
allowlist, because a list names the producers its author thought of and one
bundle locates by `source_element_id` on 269 of 274 concepts. A prefix, never
a substring (`resource_owner` is not a locator). An absent key stays absent
and an undecodable address is named (`sources_unreadable`). `sources` is READ
in both YAML forms because the two real bundles disagree (flow 629/629 on one,
block 270/270 on the other) — reading block is not a licence to write it, the
emission rule is unchanged. Contract § 8 makes `title` a MUST (checker code
`excerpt_unnamed`) and the rest SHOULD, because they are conditional on the
producer. The measurement behind it: rank 1 of 8 on 3 of 3 bundles, correct
answer on 1 of 3.
- Connect a bundle to Claude Code: `okf skill <bundle> --out
<dir>` instantiates `skills/okf-consume-template/` for THAT bundle — its id,
ref, concept count, conditional-field denominators, whole-bundle cost and
breaking point, all measured, plus a reference payload the checker accepts.
It was kept in `tools/` until 2026-09-08 because a wheel-installed
`okf skill` would emit a command pointing at a file the wheel does not carry.
That objection was about what the GENERATED skill NAMES, and O5 answered it
by changing that: the emitted commands are `okf consume` and `okf check`,
names on PATH. The template and `docs/consumption-contract.md` (the § 7.4
known-positive) are force-included into the wheel from the file they are
authored in — one authored copy, no committed duplicate.
The form was chosen on a measurement: the contract checker passes the
UNFILLED template and passes a skill built for another bundle, so it cannot
tell the two apart — the choice rests on § 5/§ 6.4/§ 7.6 being per-bundle
numbers a generic skill can only leave as holes or state falsely.
The first instantiated consumption skill is `skills/okf-consume/`; the
measurement behind it, including the control that FAILED, is
`docs/2026-09-07-okf-konsumskill-maaling.md`. **The ranking is this
repository's own choice** — the contract binds a payload, not a retrieval
algorithm (§ 10) — and it has FOUR optional widenings. Three are **off by
default** and keep the default payload byte-identical; the fourth
(`--tie-shared-rank`) became the default 2026-09-10 and is the one change in
this repository that alters a payload with NO bundle changing, so a consumer
pinned to the old excerpt order needs `--no-tie-shared-rank`. `--cost-vocabulary`: a
declared cost/price/quantity vocabulary family that bridges a question and a
document naming money with different words, gated on the QUESTION carrying
such a term, so a question without one is byte-identical either way. It moves
a measured case from candidate rank 249 to 10 and does **not** deliver it —
the budget is a second, independent lock. Measured, with two rules falsified
before building and the `k`-sweep that showed a higher `k` can EVICT a gold
concept, in `docs/2026-09-08-blindsone-below-k-k2.md`.
`--reserve-top-rank` is that second lock: the pack is an exact knapsack over a
SUM, so it has no opinion about rank and out-sums a top-ranked excerpt costing
a large share of the budget. The flag gives rank one its bytes first, AFTER
the `over_budget_alone` pre-exclusion and never before, and declares
`budget.reserved` in the payload. It fixes the eviction and does **not** close
the mandate-shaped blind spot (that concept ranks 10, not 1); the budget stays
the caller's decision, because deriving a limit from the corpus was measured
and falsified — two defensible derivations, 49x apart, one of them breaking
the known-positive. `docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
**A SIXTH flag, `--stem-prefix`, is ON since 2026-09-09** and is the second
change here that alters a payload with NO bundle changing (opt-out
`--no-stem-prefix`). `MIN_SHARED_PREFIX = 4` exists for Norwegian
compounding and also matches four characters that are not a stem: on the
pinned 453-concept bundle, control first, `under` occurs 79 times by equality
and matches 172 by prefix, `bilateral` occurs **0** and matched **400 of
453** through `bilag`, `standhaftig` 0 and 219 through `standard`. The two
extra known-negatives were FOUND, not chosen -- every 4-character prefix
ranked by document frequency, then a real word taken from the widest. **Three
candidates were measured and all three failed on the SAME row**: a longer
floor (5-8), a coverage share (0.5-0.8) and a long-words-only floor (>= 8).
Decomposed, row 1's token `prisene` reaches its gold document through
`pris|sammenstilling` on four characters -- 0.57 of one word and 0.22 of the
other -- so **the over-match and the wanted match are one mechanism** and no
threshold on length or coverage separates them. The fourth candidate does:
the shared prefix must be a WORD the bundle uses. `bilateral` 400 -> 0 and
512 -> 0, `standhaftig` 219 -> 56 and 235 -> 33, every hit@8 row keeping rank
1 on BOTH bundles. `undersjoisk` stops at **162** because `under` IS a word
here -- a genuine Norwegian morpheme, so that residual is a different answer,
never a ceiling. The vocabulary is the BUNDLE's own, so the rule makes a
payload corpus-dependent the way `rarity_weights` already is.
`--rarity-weight` is the third: each lexical hit weighs `log(N/df)` over the
bundle's own concepts instead of 1, so an identifier is not worth what a
common verb is worth. It enters the RANKING and never the GATE — `lexical`
stays a count, because a word every concept carries weighs exactly 0 and a
weighted gate is what `54a0bc2` falsified. Off by default BY MEASUREMENT: it
delivers one of three requirement lookups and takes a priced sheet from
candidate rank 10 to 2, leaves one gold unmoved and costs another seven rank
positions. Two limits are decomposed rather than guessed, and both are
someone else's mechanism: `MIN_SHARED_PREFIX = 4` makes a unique identifier
read as 135-of-446 common, and RRF consumes RANKS, so no weighting inside a
signal can move a gold that already leads it.
`docs/2026-09-08-sjeldenhetsvekt.md`.
`--tie-shared-rank` is the fourth and **the only one that is now ON**
(2026-09-10, opt-out `--no-tie-shared-rank`). It is a correction to the
TIE-BREAK
rather than a weight: RRF ranks every concept in every signal, including a
signal that scored them all the same, and the declared `(-score, concept_id)`
tie-break then orders that group by id. Measured on N500, whose document
prior has **two** distinct values over 270 concepts, that signal contributed
alphabetical UUID order and put a concept answering 7 of 7 question tokens at
fused rank 14 — outside the cut — behind concepts sharing only `tunnel` and
`vann`. Under shared ranks it is rank 3 and 2 of the 16 covering concepts are
delivered. **It shipped OFF on a measurement that was CONDITIONAL and stopped
being true in a commit reported as changing nothing.** The published cost —
hit@8 falling 5 of 6 to 4 of 6 — is real only at `DOCUMENT_PRIOR_EXPONENT`
1.0. Round 6 moved that exponent to 0.5 for an unrelated reason and correctly
reported it moved no hit@8 row; nobody measured the PAIR. Swept 2026-09-10
over 2 exponents x 3 bundles x 6 rows: at 0.5 the rule holds
`[1,1,1,1,1,]` on all three bundles and FIXES the split bundle's row 1
(2 -> 1), which is what let `--sheet-section-rows --keep-table-heading`
become a build default. **A flag's "off by measurement" is a measurement of a
CONFIGURATION, not a property of the flag** — when a constant it interacts
with moves, its default is unmeasured again, and nothing in the tree says so
because the two decisions live in different files. The adverse case is
recorded rather than hidden: on a synthetic 30-concept fixture where one
signal separates and two do not, shared ranks move a gold from rank 18 to 30
(`tests/test_okf_consume.py`). Note also that
`docs/2026-09-08-sjeldenhetsvekt.md`'s figures were measured under the older
tie-break and are NOT re-measured — on one fixture the change takes the
weight's gold from fused rank 18 to 1.
`docs/2026-09-08-rangeringsbom-sammensatte-ord.md` and
`docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md`.
The other three stay off. A FIFTH flag is not a ranking widening and is
listed apart: `--withheld-titles`
gives each `withheld` entry the concept's `title`, so a reader can see WHAT
was withheld without reading the bundle (§ 2.2 forbids going to look). The
code is 11 lines; the bytes are the reason it is off. Measured, it grows an
N500 payload 37.9 % and takes the 629-concept K2 bundle's BOOKKEEPING to
122 704 B — past the 120 000-byte limit itself — which would make the
breaking point published in the tracked `skills/okf-consume/SKILL.md`
("~75 KB at 629 concepts … at roughly 8 000 concepts") false on the day it
shipped.
## Workflow
- TDD: no production code without a failing test first.
- This repo is published PUBLICLY (`open/` namespace on Forgejo). `STATE.md`
and `docs/oppstartsprompt.md` are LOCAL-ONLY (gitignored) — never commit
session state or internal briefs. No secrets, sober English prose, no
marketing language.
- **Consumer content stays at form level in public files.** Some consumers we
read are private (`claude-code-llm-wiki` is, pending an Anthropic ToS
assessment). Key names, counts, gate names, profile fields and contract shapes
are publishable; page bodies, full title or path lists from a consumer's
bundle, and Anthropic-derived prose are not. Findings about a private
consumer's data go back to them through coord, never as a file here. This
costs nothing — every question this library asks of a consumer is about shapes
and key sets — and it is not reversible once pushed.
- After `git commit`: push to Forgejo (`git push origin`) immediately.
Never GitHub.
## Communication patterns
### Linking to local files
When pointing to local files in responses, always use markdown link syntax
with a descriptive name:
- Use `[Human-friendly name](file:///absolute/path)` — never bare
`file:///...` URLs or autolinks `<file://...>`.
- Always use absolute paths. Never `~/` or relative paths.
- For multiple files, render as a bullet list of named markdown links.
Why: bare `file://` URLs only render the first as clickable across multiple
lines. Named markdown links make each entry independently clickable and look
cleaner.