llm-ingestion-okf/CLAUDE.md
Kjell Tore Guttormsen bc39e8091f feat(assets): a bundle carries the images its sources declare (0.10.0)
Until now no reader in this package fetched, named, described or copied a
single image. `<img>`'s attributes were never read, a NISO-STS `<graphic>`
was walked past, a PDF was opened for its text alone, the converter's
markdown writer dropped every picture, and the only writer into a bundle
took `content: str`. The two lossiness warnings said so on every run, which
made the loss honest and did not make it smaller.

Measured on R761 Prosesskoden:2025, published as a 701-page PDF and as a
NISO-STS delivery: the process text is carried in full while 12 `Tabell N-N`
and 9 `Figur N-N` captions stand over nothing, because that publisher ships
those tables as raster pictures in both. Process 84's "toleranseklasse ...
er gitt i tabell 84-2" points at empty space.

THE GATE WAS WRITTEN FIRST AND RED. `tests/test_asset_gate.py` reads its
denominator out of the source (`page.images`, `word/media/`, `ppt/media/`,
`<img`, `<graphic`), never from a constant here. Measured at 332961a, built
from `git archive` and not from the editable tree: carried 0 of 8 local
images across 5 documents (9 declared), and no `assets/` at all. After: 8 of
8, with the ninth a remote source carried as a pointer without a file.

FIVE READERS PLACE, ONE MODULE DECIDES. `assets.py` owns what an image is
(sniffed from the bytes, never from the claimed extension), what it is
called (`<sha256[:12]>-<the source's own basename>`) and how it is pointed
at (one two-line block, one regex). `.xlsx` is deliberately not a row: a
block inside its pipe tables would break the `source_rows` locator, and 0 of
4 K2 workbooks hold media.

A PDF stream that is already a file is carried VERBATIM (29 of R761's 50
objects are DCTDecode); raw samples are encoded to PNG with stdlib zlib, so
no new dependency. Rendering the page region was the alternative and was
felled on determinism: a rasterised crop's bytes, and therefore the asset's
content-addressed name and the bundle's digest, would depend on the
installed rasteriser. What the encoder cannot express exactly is refused
with a code and counted, never approximated.

NO SIZE FLOOR, and that is a measurement: over the 4 828 image objects of
the K2 corpus the size distribution is a broad spread with no gap, unlike
OCR_CID_SHARE's bimodal one, so a threshold would be a number we chose.

ON BY DEFAULT, AND THE CONTROL IS TWO WHOLE BUILDS. The 43-document
reference corpus at 332961a versus rebuilt at HEAD with `--no-assets`:
865 files on both sides, `diff -rq` reports ONE difference, the added
`Images: NOT CARRIED` line in log.md. Every concept byte-identical.
Against the default: 453 -> 454 concepts, 865 -> 867 md, 0 -> 2 964 assets
(2 964 carried of 3 145 found, 4 622 pointers), 4.7 MB -> 115 MB, 2 414 s ->
3 088 s, peak RSS 6.26 -> 8.74 GB, 422 of 865 md files differ. The one new
concept has a measured cause: the pointers are body text, so a section
holding 146 of that document's images grew from 19.0 % to 30.6 % of the
extracted text and crossed `--outline-gate`'s 0.20 share clause.

THE IMAGE BYTES ARE NOT SCREENED. The guard is text-only, the pointer block
passes the gate as body text, the picture beside it passes nothing, and
log.md says so on every run.

Also fixed, both found by measuring rather than by reading:

- a markdown image is no longer read as a cross-reference. `structure._LINK`
  never looked at the character in front of the bracket, so every pointer
  would have arrived in the index as an edge to a concept that cannot exist.
- Door C carries the assets its merged concepts point at. Before this,
  importing a bundle built with `--assets` merged 6 of 6 concepts and wrote
  no `assets/` at all, so every pointer named a missing file.

Report: docs/2026-09-17-bilder-i-bundlen-trinn1.md
Spec proposal: docs/plan/okf-assets-section-6-4.md
Suite 1 955 passed / 1 skipped (from 1 896), ruff and mypy --strict clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 10:01:31 +02:00

85 KiB
Raw Permalink Blame History

llm-ingestion-okf

Context

Shared OKF (Open Knowledge Format) ingestion library. Three entry doors, one boundary rule:

  • Door A — spec-based ingestion: implements the normative ingest-spec.md owned by portfolio-optimiser-commons (manifest → file/sql/http connector → deterministic materialization of ingest-{id}.md → index generation; zero model calls). This repo IMPLEMENTS the spec; commons keeps authorship. Spec changes the library needs go via commons, never edited locally. The library ships the §11 golden fixtures (byte-exact) for the three door-A source types (ingest-golden-{file,sql,http}/, shipped in 9dd86b1).
  • Door B — bundle inbox: converts dropped files to OKF concepts. The drop directory is walked RECURSIVELY, sorted by relative path, and a concept's source_file is that relative path (/-separated) while its NAME still comes from the basename — so a nested duplicate hits the §3 collision refusal rather than vanishing. Dot-directories and a bundle nested inside the inbox are skipped with a code, never silently, because recursion makes the door's own output reachable as its own input (operator 2026-09-06; the flat listing was not a boundary, it was an absence with no denominator). All file-type→text extraction lives HERE (the guard is text-only). v1 core: md, txt, csv, json, html (stdlib). html got a measured segmentation row 2026-09-10. Until then _HTMLTextExtractor.text() was " ".join("".join(parts).split()), and str.split() with no argument splits on newlines too, so extraction of ANY HTML file returned unconditionally ONE line while every boundary grammar in propose is line-anchored -- measured outside this repo, 828 of 828 real sections gave 0 plans and exit 2 at every sample point, and a coarser 145-document cut gave 145 of 145. Block tags now open their own lines and h1-h6 carry the ATX marker for their OWN level (a flat # would hand _ATX three top-level boundaries where the document declares one section and two subsections). The output grammar is MARKDOWN, the same the office rows reach the proposer through, so no HTML-only heading grammar exists; the fix is in the extractor and never the converter, because .html stays out of _PANDOC_FORMATS on CVE-2025-51591. After: 828 of 828 plans, exit 0, 3206 concepts / 6015 md -- the markdown path's count EXACTLY, and the same at 414 (1651) and 83 (343). Text preservation is an EXACT invariant and not a percentage: strip the added ATX markers and the non-whitespace sequence is identical to the old extractor's, 828 of 828 files, character ratio 1.000000. _SKIP_TAGS stays {script, style}. Exposure elsewhere measured rather than argued: 0 of 86 K2 corpus files and 0 of 5 smoke-folder files are HTML, and the smoke bundle is byte-identical before and after. _EVIDENCE gains a .html row at measured -- and a .pdf row at measured since 2026-09-10, the row with the most measurement behind it and no entry in the table at all -- with the limit that travels with it -- one product, one format, one publisher, and a generator's cut, not 828 documents anyone wrote. .xml became a CORE type 2026-09-11 and it is the first row whose ceiling is structural rather than recovered. A NISO-STS zip from a publisher's own viewer was 110 of 110 unreadable, 0 plans, exit 2 -- .xml was in neither registry -- and the one xml file in it IS the whole product: R761 Prosesskoden:2025, the document round 12 met as a 701-page PDF, carrying 7 715 <sec>, 2 761 with a <title>, 4 954 with a <label> and no title, 10 <table-wrap>, root <standard>, 0 <!DOCTYPE. Its <sec>-nesting depths over the titled sections are 28/118/500/1141/868/97/9, row for row the fasit's own. The reader is stdlib (xml.etree.ElementTree) and adds NO dependency -- defusedxml and lxml are 0 occurrences in uv.lock -- so it is core beside .html rather than behind [extract], which would make a pure-stdlib type binary-dependent. The output grammar is MARKDOWN, the same the office rows reach the proposer through: <sec> with a <title> becomes one ATX line carrying <label> + space + <title> at its own nesting depth, <sec> with only a <label> becomes a body line with the label in front (never a heading -- 4 954 of 7 715 are lettered points and one heading each buries the document's own 2 761), and <table-wrap> becomes its label plus one markdown table through render_table. <label> carries the number and <title> carries the text -- 2 of 2 761 titles begin with a digit -- so emitting <title> alone scores 0 of 2 761 with nothing in the code looking wrong. Inline is an ALLOWLIST and block is the default, the inverse of the HTML reader, because block-versus-inline is a property of HTML and XML has no such universal; the allowlist is load-bearing at 1 701 <italic> and 1 396 <bold> inside that document's prose. The ATX ceiling is 6 and STS nesting reaches 7, so the depth is CLIPPED and not dropped: 9 of 2 761 sit at depth 7 and ####### matches nothing. Since K3-21 the clip is the HEADING's alone: the OutlineMark beside it carries the declared depth, so the plan reads those 9 at 7. A <!DOCTYPE is REFUSED unparsed with its own code, a guarantee about the code rather than about the machine -- measured on 3.14.0 with pyexpat 2.7.3, an external SYSTEM entity is refused by the stdlib but the billion-laughs limit comes from libexpat

    = 2.4.0 and not from Python, while pyproject.toml requires only >=3.10. XML that is not STS keeps its text in document order and gets NO invented structure, and .xml never routes through the converter -- a second parser that would never see that refusal. XML that is not STS also gets 0 plans and a FAILED build (exit 2), and that is NOT an .xml defect: a folder holding one .txt of prose with no headings gives the same three lines and the same exit, so it is general okf build behaviour for any structureless document. The gate stays -- a run replaying zero plans would emit a flat bundle and call it success -- because separating "0 plans, 0 unreadable" from "0 plans because nothing could be read" changes the outcome on 0 of the 4 reference corpora. THE READER REACHED ITS CEILING IN ROUND 13 AND THE BUILD DID NOT, AND ROUND 14 CLOSED IT AT THE SHIPPED DEFAULTS. The reader emitted 2 761 of 2 761 heading lines while the build delivered 23 concepts and 15 of 2 761 boundaries -- two steps after the reader, each measured: the orphan check took 710 of 2 761 (710 of 710 removed headings are followed immediately by another heading, 0 of 2 051 delivered ones are -- they are container sections) and Arm F took 2 066 more, 2 089 -> 23. find_candidates already skipped both for outline_marks, which is why the PDF bookmark arm reaches 2 762; an STS <sec><title> is the same class of declaration and only arrived as rule:heading. The fix is ONE new rule constant reached from ONE row: extract.xml_outline reports the marks the reader WROTE ITSELF -- no bridge, no tolerance constant, no unresolved bucket, the difference from pdf_outline whose naive nearest-line rule was wrong on 1 840 of 2 762 -- propose.RULE_XML_SECTION (rule:xml-section) is its own name in RULE_NAMES and _ORPHAN_EXEMPT, and build_plan chooses the route by the ROW (DECLARED_STRUCTURE_IDS), never by the text: the same markdown from a .md file is still a guess and still carries rule:heading. At shipped defaults, no flag: 2 761 concepts, 2 761 of 2 761 declared sections became a concept with the source's own directory AND title, 0 concepts matching no declaration, a)-points 0 of 4 954, table blocks 10 of 10, hit@1/8/50 3/6 / 5/6 / 6/6 from 0/6 / 0/6 / 0/6 with the known-positive at rank 1, and 2 761 shared concept ids with the PDF arm (100 % of this bundle, 2 761 of 2 762 of that one) against round 13's 2 022. NO other file type changes one byte and it is MEASURED on the bytes: the whole 43-document reference corpus rebuilt is diff -r-identical to the pinned bundle (865 md), the five-document folder is diff -r-identical, okf project stays byte-equal to okf build, and the PDF arm still proposes 2 762. Two directories of 2 738 still hold two concepts (11, 12) -- the publisher reuses a section number, the same 2 the PDF arm has, and 0 is not reachable without inventing an id; round 13's 14 such directories were false positives of the TEXT route reading the document's own contents listing and are gone. Report: docs/2026-09-10-k3-runde14-deklarert-struktur-tar-ruten.md. Since K3-19 an STS document's own identity names its directory (extract.declared_identity, read by cli._document_prefixes and the door): the directory was the delivery file's stem, a UUID occurring 0 times in the document, while its one <std-ident> carried <doc-number>. Only the stem is replaced, and a declared name two documents in one run claim is used by NEITHER -- the slug_owners gate would refuse both with "rename one", which a name read from inside a document cannot obey. The sources title is <doc-number> + <year>, then <title-wrap>, then the file name: R761's <full> carries a COMMA, a flow terminator, so it is never written and never cleaned up. A titled section's description is its own FIRST spec point (first <p> of the first DIRECT-child sec-type="spec", whole), carried by the plan entry, screened by the gate, and written only where a YAML reader reads it verbatim (inbox._yaml_plain): 2 026 of 2 761 titled sections on R761 carry a point, 1 807 are written (2 have no <p>, 217 carry : and PyYAML refused exactly those frontmatters), none invented. SS 4.1 sets no length, so the one-paragraph limit is ours. The directory name reached the RANKING, and K3-20 closed it in consume: signal 1 read every id segment, so on a one-document bundle every concept carried the document's own name, and a question naming the document matched all of them -- measured, the known-positive went rank 1 -> not delivered at the default k (13 at k=50) with S1-S6 unmoved. consume.shared_id_prefix now keeps the leading directories EVERY id shares out of that signal: KP rank 1 at both k, S1-S6 6/6, and K2 (12 payloads), N100/N200/N500 (15) and the five-document folder (5) byte-identical, because ids that share no prefix read exactly as before. Dropping each concept's OWN document directory instead was measured and felled -- a K2 hit@8 row went rank 5 -> not delivered. Reports: docs/2026-09-11-k3-runde19-dokumentidentitet-og-frontmatter.md and docs/2026-09-11-k3-runde20-delt-katalog-og-arvet-kontekst.md. The registries are COUPLED: a row in _CORE_EXTRACTORS and not in segmentation._STDLIB_EXTRACTOR_IDS refuses every proposal for the type, two layers away from the extractor. pdf/docx/xlsx only via the optional [extract] extra; without it those types are rejected fail-fast. The extra ships pdfplumber for pdf (chosen on ONE measured property: it keeps a requirement table's label and value on the same line where three alternatives do not); docx/xlsx still ship no parser. pptx and md were measured end to end for the first time 2026-09-10 (docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md § 4) on two hand-built documents, which is more than zero and is not a fasit: md recovers 3 of 4 declared headings, and pptx segments per slide only where the deck's slides carry title placeholders the converter recognises — a deck whose slides do not lands as ONE concept. A converter attribute also leaks into concept titles ({#slide-N}, {#sheet-1}), reaching 2 of 810 files on the K2 default bundle and 1 of 30 on the operator's test folder; because a filename is reduced from its title, fixing it RENAMES concept ids a consumer has already cited, so it is an operator question and not a patch. Structured table recovery is out of scope — two independent parsers return the same wrong shape, so the breakage is document geometry, not a library choice. PDFs enter as prose, and drawn content (figures) does not survive extraction at all, which every pdf extraction warns about. Under the STRUCTURED_V1 profile Door B additionally DERIVES structure — title (leading heading → title key → path.stem), document number, hierarchy, and cross-references — writes it into the concept frontmatter, and projects it into a faceted index entry. Every inferred field is named in a derived list; an unmarked heuristic is worse than none. Under the SEGMENTED v0.2 profile a concept additionally POINTS BACK at the original: sources: [{ resource, title }] in SPEC §5.1's form (resource is the inbox-relative path), plus a locator per format — source_pages, source_sheet+source_rows, else source_lines. The locator keys are OURS and must stay top-level: §5.1 has no field for a place within a resource, and the pinned guard rejects every route to putting one inside a sources entry (non-allowlisted key, nested flow list, quoted scalar), so a locator in the entry would emit bundles Door C could never read back. The unit table is built AT EXTRACTION — a page number cannot be recovered from joined text — and source_offset stays. source_lines indexes the EXTRACTED text, never the original's paragraphs: measured, docx <w:p> counts and converted-line counts do not agree on a single one of five documents. Record: docs/2026-09-08-proveniens-k2.md. The index is a PROJECTION recomputed from the whole bundle each round, which is what makes rebuild-from-scratch equal an incremental update byte for byte. DEFAULT is untouched and byte-identical. Record: docs/plan/structure-derivation.md. SINCE 0.10.0 DOOR B CARRIES THE IMAGES ITS SOURCES DECLARE. Until then no reader here fetched, named or copied one -- <img>'s attrs were never read, an STS <graphic> was walked past, a PDF was opened for text alone, the converter's markdown writer dropped every picture, and the only writer into a bundle was materialize.write_bytes(..., content: str). Measured on R761 Prosesskoden:2025: the process text is carried in full while 12 Tabell N-N and 9 Figur N-N captions stand over nothing, because that publisher ships those tables as raster pictures in BOTH the PDF and the NISO-STS delivery -- process 84's "toleranseklasse ... er gitt i tabell 84-2" points at empty space. Five readers PLACE and one module DECIDES: assets.py owns what an image is (sniffed from the bytes, never from the claimed extension), what it is called (<sha256[:12]>-<the source's own BASENAME>, so one image reached by two paths is one file) and how it is pointed at (one two-line block, one regex, IMAGE_POINTER, which is what okf describe will find its work with). The pointer is a markdown image at /assets/<name> -- bundle-absolute, because a segmented bundle puts concepts at different depths -- followed by one line carrying the source's own file name and the size in px. .xlsx is deliberately NOT a row: its converter writes one pipe table per sheet and a two-line block inside one would break the source_rows locator read back out of it; 0 of 4 K2 workbooks hold any media, so it is a stated limit and not a loss taken. A PDF stream that is already a file is carried VERBATIM (DCTDecode, JPXDecode -- 29 of R761's 50 objects), and raw samples are encoded to PNG with stdlib zlib. Rendering the page region was the alternative and was FELLED on determinism: a rasterised crop's bytes, and therefore the asset's content-addressed name and the bundle's digest, would depend on the installed rasteriser -- the one property OCR_DPI's docstring already admits OCR text cannot have. What the encoder cannot express EXACTLY (stencil mask, Decode array, CMYK, anything but 8-bit samples) is refused with a code and counted, never approximated. NO SIZE FLOOR, and that is a measurement: over the 4 828 image objects of the K2 corpus the distribution is 149 / 162 / 92 / 406 / 498 / 590 / 2 931 across the size buckets -- a broad spread with no gap, unlike OCR_CID_SHARE's bimodal one, so a threshold would be a number this package chose. ON by default; --no-assets reproduces the pre-0.10.0 bytes. Cost measured on the 43-document reference corpus, two builds of one commit: 453 -> 454 concepts, 865 -> 867 md, 0 -> 2 964 assets (2 964 carried of 3 145 found, 4 622 pointers), 4.7 MB -> 115 MB, 2 414 s -> 3 088 s, peak RSS 6.26 -> 8.74 GB, 422 of 865 md files differ. The ONE new concept has a measured cause and not a guessed one: the pointers are body text, so a section holding 146 of that document's images grew from 19.0 % to 30.6 % of the extracted text and crossed --outline-gate's 0.20 share clause. THE PLAN AND THE RUN MUST AGREE: a plan records text_sha256 of the exact string it was proposed against, so propose and the door take the same assets value and each computes the SAME resolver root independently -- the document's own directory, containment by connectors.safe_resolve. A reference above it is refused (asset_unresolved), a remote one is never fetched (asset_remote, extraction opens no socket) and both leave a line in the concept saying what was there. THE IMAGE BYTES ARE NOT SCREENED -- the guard is text-only, the pointer block passes the gate as body text, the picture beside it passes nothing -- and log.md says so on every run.

  • Door C — external bundle import: third-party OKF bundles are assessed per concept via the guard's okf.import_bundle; only concepts clearing the guard's non-blocking floor are merged/indexed here. Two invariants, both load-bearing: a merged concept is written verbatim (this library's line-oriented frontmatter parser cannot round-trip the block lists the guard's parser accepts, so stamping an external concept would destroy sender data and persist bytes the guard never screened), and ownership is therefore proven by content identity — an occupied target name is re-used only when the bytes there are already identical, never overwritten otherwise. Since 0.10.0 it also carries the ASSETS its merged concepts point at, by that same content-identity rule. Measured before the repair: a bundle built with --assets imported as 6 of 6 concepts and no assets/ at all, so every pointer in the imported bundle named a missing file — the same "complete and not" defect one door over. POINTED AT, never every file in the sender's assets/: an asset belonging to a concept the gate refused must not ride in on the back of one it cleared, and an asset nothing names is a file no retirement pass reaches. A pointer whose asset the sender did not ship is left alone, because SPEC §6.1 requires a consumer to tolerate a broken link and a pointer recording a figure nobody holds is information.

okf build RUNS a real guard and NAMES it in the bundle (F1, 2026-09-15). From the day the command was packaged until then, corpus.measure wired an unconditional approve-everything stub into process_inbox and 0 of 90 add_argument calls in the package named a gate — so the one path most people use screened nothing, while pyproject.toml made the guard a MANDATORY runtime dependency and the README recommended a composition the command line could not reach. Reported from outside by claude-code-llm-wiki, reproduced here first. --gate takes guard-trusted-source (default), guard-user-upload or none, corpus.resolve_gate is the ONE name→callable map (guard imported lazily, so importing the package still does not pull the dependency in), and an unknown name RAISES (gate_invalid) rather than falling back — a fallback reproduces the defect with an extra step. The default was chosen on a measurement: over the 453 concept bodies of the pinned reference bundle, PRESET_TRUSTED_SOURCE persists 453 of 453 and PRESET_USER_UPLOAD holds 1, costing that concept's whole source document (1 of 39) — and neither tier waves anything through, an invisible carrier and a CRITICAL finding are fail_secure at BOTH. Door B's library default is UNCHANGED at PRESET_USER_UPLOAD: an inbox drop is an untrusted upload, an operator pointing this command at their own folder is not. The second tier is guard_adapter.inbox_gate_trusted_source, the three-line adapter that module's docstring already described — never a preset parameter. The gate's NAME is written into the §9 log.md, because a stub is only dangerous when nothing downstream can see it; --gate none renders NOTHING WAS SCREENED. The corpus harness carries the same flag and the SAME default (a test holds the two paths byte-equal); okf project takes none, owning no flag that moves bytes. The composition process_inbox(segmentations=..., gate=inbox_gate) now has a test — before this, grep -rl inbox_gate tests/ gave 1 file with 0 occurrences of segment, which is how the defect survived.

A FENCED CODE BLOCK DECLARES NO STRUCTURE (F2, 2026-09-15). The proposer read every line with the same grammars, so # Use the opus[1m] alias inside a ```bash fence became a level-1 ATX heading. Two effects and the SMALLER one was visible: the document was REFUSED entirely when the line carried [ or ] (5 of 191 pages of the reporter's corpus, inbox_title_invalid), and the concept TITLE was silently taken from somebody's shell session on 62 of 191 (32.5 %). The fix is in the PROPOSER and never in Door B's title rule — that rule is right, and a heading that was never a heading is what has to stop being proposed. propose.fenced_lines is computed once per text and NO rule reads a fenced line: not _ATX, not the numbered grammar, not a table row, not --bold-title, and not Arm D's outline RUN, which selects from the whole line list (filtering only at admission would let a fenced install listing decide which run wins). Four CommonMark § 4.5 details are load-bearing, each a way to remove REAL boundaries: three leading spaces still open a fence; a backtick fence's info string may not contain a backtick (or a line holding only `okf build` silences the document); a closing fence must be at least as long as its opener; an unclosed fence runs to the end. It lands unconditionally, not as an eleventh flag, and the exposure is measured on the bytes: 0 of 865 concept files in the pinned default bundle and 0 of the shipped fixtures and goldens reaching the proposer carry a fence of either kind, so a rule that can only fire INSIDE one cannot have moved anything measured here. It is a defect, not a default move. BOTH CHANGES TOGETHER MOVE ONE LINE, AND IT IS MEASURED ON THE BYTES: the 43-document reference corpus built at b6da09c (from git archive, never the editable tree) and rebuilt at the shipped defaults differ in log.md alone, by the added **Gate**: bullet -- 865 concept files on both sides, every concept byte-identical. The same run found something this work did NOT cause: the pinned artifact K2-bundle-default-20260912 was written 2026-09-09 21:38, two days before ed0418f (K3-22) changed title: quoting, so it differs from what HEAD produces on 42 concept files -- and tests/test_default_bundle_pin.py stays green because it pins the count and the hit@8 ranks, not the bytes. Re-pinning it is the OPERATOR's. Report: docs/2026-09-15-f1-f2-gaten-og-kodefencen.md.

Boundary rule (non-negotiable, zero overlap): llm-ingestion-guard (pinned >=1.2,<2.0) answers "is this content safe to persist?" — scan/sanitize/quarantine/fail-secure/provenance-stamp. This library is plumbing: connect source → materialize deterministic OKF bundle → generate index. Never reimplement security; call the guard at persist gates (prepare_input/screen_output, okf.import_bundle). When in doubt which side of the boundary something belongs on: ask the operator.

Implementation baseline: the stricter behaviors from portfolio-optimiser (streaming row caps, utf-8-sig, in-memory staging with pre-mutation collision gate, validated ingested_at, typed IngestError) are the library baseline. First consumer: portfolio-optimiser-claude.

Roadmap (phases 13 shipped; what follows is demand-driven)

  1. Phase 1 — Door A (Python). ingest-spec implementation + the §11 golden fixtures. Consumers: portfolio-optimiser-claude first, then portfolio-optimiser.

  2. Phase 2 — Doors B/C (Python). Bundle inbox and external-bundle import, guard-gated.

  3. Phase 3 — Configurable bundle contract. Types, layers, frontmatter sets, index shape, and reserved-file policy become config instead of constants; proving consumer is claude-code-llm-wiki (strict-v1 profile). Two consumers hold opposite postures on whether an index is authored or directory-derived, so neither is a library invariant and nothing here enumerates a directory unless the profile says derived.

  4. Phase 4 — Node half (node/). Zero-dependency Node/ESM package (importable and CLI-invokable, vendorable per plugin — matching the marketplace precedent) for the second-brain world: bundle check, index generation, inbox split/frontmatter/write, and doc conversion (docx/pdf/eml/html → md). Covers okr, linkedin-studio, ms-ai-architect, and the marketplace catalog.

  5. Phase 5 — MCP as a way to populate a bundle. NOT COMMITTED; needs-based (operator 2026-08-02, superseding the 2026-07-27 commitment.) No MCP work, and no data-lake or database source types, are undertaken without a stated need. docs/plan/mcp-bundle-population.md stays as a design record, not a queue. Its open fork — whether we are the MCP server (an agent calls our doors as tools) or an MCP client (a manifest source type pulling from someone else's server) — no longer blocks anything, because nothing waits behind it. It is a question to answer if a need arrives, not before. This is also why sql staying sqlite-only is not a gap: a Postgres driver would be runtime dependency number two, bought for no asked-for use.

The two halves share the OKF contract and fixture suite, not code.

Standing posture (operator 2026-08-02). Phases 13 shipped; the library now runs on what it has. Work is defect fixes, improvements, and features that a consumer has actually asked for or that measured feedback shows are needed — not roadmap completion for its own sake. The upstream version policy below is the one exception, and it is not a counterexample: "always latest" is a promise already made to consumers, so an upstream release is the stated need. Phase 4 keeps four named consumers with working implementations to lift, so its need is real but untriggered — it starts when one of them asks, not on a date.

Upstream version policy (standing, non-negotiable)

The library always supports the current latest version of Google OKF. Set by the operator 2026-07-26. Phases 13 were built against v0.1; v0.2 shipped 2026-07-25, so v0.2 support is committed work — not contingent on a consumer asking for it. Plan: docs/plan/okf-v0.2-alignment.md.

Support is additive, expressed as a new profile, never a migration of existing ones. This is what makes the policy sustainable instead of a recurring crisis, and it is bounded by three facts that do not yield to it:

  • DEFAULT states commons' ingest-spec §5 layer — its generated shape is commons' call, raised there, never patched locally. This fired 2026-08-09: commons ratified and executed the O2 form, so DEFAULT now stamps generated: { by: process:okf-ingest, at: <ingested_at> } and four goldens moved with it. It is not a counterexample to "additive, never a migration" — that rule governs upstream versions, and commons' spec is a separate axis DEFAULT tracks by definition. DEFAULT stays v0.1 on everything upstream owns. Ownership recognition is one-way, so the cost to a consumer stays a re-run: a profile carrying an actor still owns the older literal stamp.
  • STRICT_V1 mirrors the proving consumer's ratified contract — changing another repo's contract from here violates O2.
  • okf_version's value belongs to catalog (decision E1).

Rollout is pilot-first. A new upstream version reaches a small pilot set on a pre-release tag and is revised on their feedback before general availability — consumers testing real data find what fixtures cannot. OKF_LATEST means the latest version supported as stable, so flipping that alias is the GA event, not a merge side effect.

Two invariants fall out: no profile hard-codes an upstream version, and no bundle declares a version its shape has not earned. The first has a mechanism, not just an intention: a profile names a key, a caller owns its value. okf_version is declared through materialize_bundle(..., root_frontmatter_values=...) because its value tracks the upstream Google version and belongs to catalog (decision E1) — a constant here would claim a decision we do not own, and would be the one thing to chase on every upstream release. Where upstream itself defers a contract — v0.2's attestation receipt and verdict wire formats — the format is supported and the unspecified runtime is not; it re-enters scope when upstream specifies it. Because "always latest" decays silently, the release checklist carries an upstream-version re-check.

Structured frontmatter values are emitted in YAML flow form, never block. Both can be valid YAML -- within the flow-scalar limit below -- and an upstream reader then recovers the same structure from either, but this library's parser is line-oriented: it round-trips a flow mapping as an opaque value and cannot read the block form at all — two block mappings sharing an inner key (§10.2's executor and attester, both carrying resource) collapse into one namespace and the first is lost silently. Emitting block would produce bundles we cannot read back. Reading it needs the structured reader (D1b); until then the constraint binds what we write.

Every value is written so a YAML reader reads it back the same (K3-22). SPEC § 11 point 1 requires "a parseable YAML frontmatter block", and before K3-22 the pinned K2 default bundle failed PyYAML on 41 of 455 blocks and each R761 build on 1 -- block scalars written verbatim. A block scalar that is not plain-safe (profiles.yaml_block_plain, K3-19's rule) is now written double-quoted with \ and " escaped; every other value keeps its bytes. A FLOW leaf has no quoted form -- the pinned guard refuses any quote inside a flow mapping -- so a leaf PyYAML would refuse or misread (?, ": ", " #", a quote, a leading indicator) is refused by profiles.yaml_flow_plain with the door's existing code, never written. Readers unquote a "-wrapped value only: 0 such values existed in any measured bundle, while 11 193 '-wrapped ones do and stay untouched. PyYAML is a dev dependency that validates the rules in tests/test_yaml_frontmatter.py; src/ imports no yaml. Report: docs/2026-09-11-k3-runde22-yaml-lesbar-frontmatter.md.

Every upstream release runs docs/upstream-okf-upgrade-runbook.md. Pin the commit, enumerate the whole okf/ tree, read the shipped example bundles and not only SPEC.md, classify the diff, measure our exposure and each consumer's, plan additively, pilot before GA, then inform every OKF-consuming repo. The runbook is not optional and not a summary of good intentions: each step names the concrete failure it prevents, and all of them are failures that happened during v0.1 → v0.2.

This repo is a black box for its consumers. The target cost of an upstream release to a consuming repo is a re-run, nothing more: support is additive (a new profile, never a migration), existing profiles stay byte-stable, new public parameters are keyword-only with defaults so positional call sites stay source-compatible, and consumer golden fixtures must not churn. The boundary is stated every time rather than glossed — the library absorbs shape changes, not upstream changes to content a consumer authored (v0.2's timestamp and # Citations supersessions). For that class the deliverable is a measured exposure report per consumer, sent before they ask.

Phase 4 preconditions (coordination, not unilateral moves):

  • Lifts okr's reference implementations (okf-check.mjs, okf-index.mjs, innboks libs) in agreement with okr and the marketplace catalog; the catalog remains the convention owner and re-pins its shared gate here.
  • linkedin-studio's ingest/published/ provenance-record grammar stays plugin-local by design (different lifecycle) — do not normalize it.
  • The Node-side persist gate remains security territory: guard-as-contract (per okr's adoption doc) until a Node guard exists in the security repo. No security reimplementation here, in either runtime.

Non-goals (all phases)

  • Verdict/feedback machinery (method-spec) — stays in consumer repos.
  • Embedding/RAG/retrieval layers.
  • Security functionality — always the guard's domain.

Stack

Python 3.10+. Package llm_ingestion_okf (src layout, hatchling). Exactly one runtime dependency, ever: llm-ingestion-guard>=1.2,<2.0 (itself zero-dep), landed with the Door B/C persist gates. Everything else is stdlib, and a packaging test enforces it. Only guard_adapter.py imports the guard; importing the package does not. Install channel until the package index exists (a direct reference is a channel, not the pin): pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.2.0". Binary extraction parsers live behind the [extract] extra only — today pdfplumber>=0.11.10,<0.12 for pdf. Extracted PDF text is pinned to an exact transitive parser version (pdfminer.six==20260107), so widening that range is a fixture migration, guarded by a frozen literal in tests/test_extract.py; see tests/fixtures/README.md.

Phase 4 adds a node/ half: Node/ESM with zero npm dependencies (node: builtins only), both importable and CLI-invokable, consumed by vendoring per plugin rather than npm publishing. The halves share contract and fixtures, never code.

Conventions

  • Type hints everywhere; mypy --strict target.
  • Determinism is bit-exact: ingested_at is an explicit required argument (no wall-clock defaults); LF-only output; golden fixtures compared byte-for-byte.
  • Filenames and titles are normalized to Unicode NFC before use (materialize.reduce_to_id_grammar, inbox.process_inbox): macOS/APFS hands filenames over in decomposed form, so an é arrives as e + combining acute. Without normalizing first, the same visual name (e.g. a Norwegian slugger title like "linkedin-studio") reduces differently depending on which form it arrived in, splitting one title into two generated filenames.
  • No model calls anywhere in the run path.
  • Credentials only as env-var references resolved at runtime; never in manifests, logs, or frontmatter.
  • Network access requires an explicit per-run opt-in flag; refuse fail-fast otherwise.
  • Conventional Commits: type(scope): description.
  • English for all code, docs, and commit messages (public repo).

Commands

  • Test: pytest
  • Lint: ruff check . + ruff format --check .
  • Type check: mypy --strict src/
  • Folder to questionable bundle in ONE command: okf project <folder>okf build with the package default into <out>/.okf/<id>/ plus okf skill into <out>/.claude/skills/<id>-consume/, <out> defaulting to cwd and <id> to the folder name reduced to [a-z0-9-]. It owns NO flag that moves a bundle's bytes and a test holds it byte-equal to okf build; two build paths would leave every measurement report pinned to a bundle nobody produces. That invariant was FALSE from the day those two flags became defaults until O6 measured it, and the test could not see it: cli.build's Python SIGNATURE defaulted keep_table_heading and sheet_section_rows to False while argparse defaulted both to True, and project.create calls build() as a function, so it read the signature. Measured on a five-document folder, okf project wrote 15 concepts / 30 files against okf build's 26 / 52, the whole difference in the priced spreadsheet -- the document a question about price has to reach. The byte-equality test compared project.create against the same build(), so both sides carried the same wrong value, and its two fixture documents had neither a table nor a sheet: a test and the code agreeing over a set where the difference cannot appear. Two tests now hold it -- one comparing the signature's defaults against argparse's for every same-typed parameter, one building a document whose concept count actually moves with the two flags. skills/okf-prosjekt/ is the Claude Code skill over it.
  • The generated consumption skill states THREE modes and RELATIVE paths (O6, 2026-09-09). Question (the default), hypothesis (decomposed into premises and answered PER PREMISE as confirmed / refuted / undecidable-from-bundle -- three literals, no fourth; a weak source is [sourced-not-sufficient] on that PREMISE, because four premises and one weak source is three answers and one gap), and a task producing a document (source per claim IN the artefact, an ungrounded paragraph written and marked rather than dropped, the cut declared inside the document because the document travels without the chat). The five markings are untouched -- the modes add no sixth. In the layout okf project writes, the commands are okf consume .okf/<id> and okf check --skill .claude/skills/<id>-consume/SKILL.md, runnable from the project root, which is where okf project's own closing line tells the reader to start claude; a path OUTSIDE that root stays absolute on purpose, since ../../.. is not more portable, only harder to read. Two absolute paths to zero -- and O5's published "4 absolute paths -> 0" was measured with grep -c "^/" against paths indented by two spaces, a query that could not have found one either way, so the zero was never a measurement. Every path assertion here runs its pattern against a known-positive first.
  • Build a bundle: okf build <folder> --bundle <dir> --bundle-id <id> --okf-version <v> — the installed console script ([project.scripts]), the packaged form of what used to be a shell loop over two tools/ scripts. It is orchestration only: the proposer and the corpus harness live in llm_ingestion_okf.propose and llm_ingestion_okf.corpus, and the tools/ scripts are thin entry points to the same functions so the published reproduction blocks still run. Path scope for a document's proposals is its RELATIVE path minus the extension (the door walks recursively, and two same-named documents in different folders must not collide); --ingested-at and --proposed-at default to one shared epoch constant rather than the clock, because a wall-clock default takes rebuild-equals-incremental away from anyone who omits them. TEN segmentation rules are REACHABLE here, and since 2026-09-09 ALL TEN are ON by default -- the tenth is --contents-name, round 9's repair of clause 1, which admits a title into a contents run only when a NAME survives stripping its page number. Measured: clause 1 discarded 68 candidates over 11 of 39 readable documents, 19 of them over 5 documents rows of a drawing's dimension chain, a schematic's labels, a door schedule, a coordinate column and a soil-layer table. The threshold is SWEPT (propose.CONTENTS_NAME_RUN = 2) and collapses at both ends: at 1 it rescues 13 of 19, at 3 the two-letter section name VA stops being a name and takes four REAL contents entries with it. At 2 it rescues 16 of 19 and 0 of 49. Corpus 429 -> 447 candidates, K2 436 -> 453 concepts / 832 -> 865 md (21af4a1aa98315cf...), the 12-position reference label-identical in BOTH readings and hit@8 [1,1,1,1,1,None] on the new bundle AND Arm B. Opt-out --no-contents-name. An ELEVENTH flag, --bold-title, is OFF (round 10, 2026-09-09): the rule for the type whose container declares nothing. rtf measured 0 of 0 declared headings, 0 concepts, 1368 of 1368 characters in no segment. The grammar is MARKDOWN, not rtf -- the converter already writes the author's bold title as **...** in the same output every office row produces, so no rtf-only heading grammar exists, the same shape of decision as the PDF font reader's ATX form. Three parameters swept over 47 readable documents and ONE carried (refusing a line that ends in terminal punctuation: false-positive lines 9-12 -> 1-2); a maximum title length and a stand-between-blank-lines clause are both FLAT and neither is in the rule. The last false positive is closed by G1, _gate_outline's own principle, so false positives are 0 of the 31 declaring documents. Reach 2 of 39 corpus documents, both docx, 0 of 33 pdf. BOTH alternatives the order named were measured and FELLED: a hand-laid fixture that DECLARES a heading style has it discarded by the converter, and rtf -> docx -> markdown yields 0 ATX headings on that same document, because the loss is in the rtf READER before any writer. The nine below are unchanged -- --outline-run 3, --table-grid and --unit-fold since 2026-09-08, --drop-wrapped-outline and --outline-gate since 2026-09-09, --sheet-section-rows, --keep-table-heading and --first-span-from-zero since 2026-09-10, and --close-span-gaps since 2026-09-11, each with an explicit opt-out (--outline-run 0, --no-table-grid, --no-unit-fold, --keep-wrapped-outline, --no-outline-gate, --no-sheet-section-rows, --no-keep-table-heading, --no-first-span-from-zero, --no-close-span-gaps) that together reproduce the pre-move bytes -- measured, diff -rq 0 differences, not asserted. The 2026-09-09 pair is one decision and cannot be split: the gate takes pdf from 2 of 8 to 5 of 8 and the pair takes it to 7 of 8 (the sheet 5 of 12 -> 10 of 12, docx unchanged at 3 of 3). The gate is G1+G2: Arm D's RECOVERED headings are admitted only where the document DECLARES none of its own -- which is fold_units clause 2's principle moved from voting to admission -- plus any one recovered heading covering propose.OUTLINE_SHARE (0.20, swept flat from 0.10 to 0.30 and collapsing at both ends) of the text. It filters at ADMISSION, before spans close, so the text a removed mark opened is carried by the mark above; the post-filter form scores identically and loses that text, which is why only one of them shipped. The bar it had to clear is now the bar: reference cells up AND hit@8 holding rank 1 on every row on every bundle. --sheet-section-rows --keep-table-heading reaches 11 of 12 and SHIPPED 2026-09-10, after two rounds off. It was held back because on a K2 bundle built with it row 1 fell rank 1 -> 2 (the gold document goes 1 concept -> 12), under both prior exponents. That was never these rules' defect and it is not a segmentation question: RRF emits a distinct rank for every concept in a signal that scored them all EQUALLY, so the gold document's own twelve concepts fill the document-prior tie group and the one leading the body signal takes position 11 instead of 1. The repair is the reading side's consume.DEFAULT_TIE_SHARED_RANK, and with it every acceptance condition holds at once. The fusion was punishing fine-graining for being fine-grained, which put the segmentation side and the retrieval side in competition over one number for two rounds. Arm E joined a session after the other two, on a number measured AFTER the first move: without it Arm F's table clause has no joined table to fold, and the shipped D+F default scored 2 of 12 with docx 0 of 3 against the 5 of 12 the fold was published with. The proposer's own defaults did NOT move (propose.py's rules stay off): the goldens and every published reproduction block are pinned to them, so the two layers disagree on purpose and cli.DEFAULT_OUTLINE_RUN / cli.DEFAULT_UNIT_FOLD say where. The cost to a consumer is a re-run and it is not small: the 43-document reference corpus goes 629 concepts / 1108 files (the delivered 2026-09-03 tree) to 492 / 944 after the 2026-09-08 move and to 425 / 810 after the 2026-09-09 one (bdf4977ca5a443c4...) and to 436 / 832 after the 2026-09-10 one (8dff8a8e6c15d2f7..., default flags, default epoch stamp, measured on 38104b7 + this round). On the operator's own five-document folder the last move is 15 concepts / 30 files -> 26 / 52. Digests published before 2026-09-09 were computed with a path-DEPENDENT command and are not comparable to this one; the reproducible form is find . -type f | sort | xargs shasum -a 256 | shasum -a 256 from inside the bundle, under which the previous default is 862116da16e422f6.... The pinned artifact lives at ~/corpora/okf-telling-20260829/K2-bundle-default-20260910 and tests/test_default_bundle_pin.py holds its concept count AND its per-row hit@8 ranks -- the count alone survived a configuration that lost a rank, which is how a previous round's regression hid. Since 2026-09-10 it also holds the KNOWN-NEGATIVE on the same bytes: read with --no-tie-shared-rank, the shipped default bundle reproduces the very fall the rules were held back for, so the pin names its own cause instead of being green for an unstated reason. And the number the decision cites belongs to another configuration: Arm F's 5 of 12 was measured with --table-grid ON; without it the same sample scores 2 of 12 and docx 0 of 3, because the fold's table clause has no joined table to fold. The ten: --contents-name (round 9), --outline-run N (Arm D), --table-grid (Arm E), --unit-fold (Arm F), --keep-table-heading (D1), --sheet-section-rows and --drop-wrapped-outline (both D3), --outline-gate (G1+G2), --first-span-from-zero and --close-span-gaps, each passed to the proposer unchanged. The last two are the same defect at two ends and NEITHER is a segmentation rule: a mark removed after its neighbour's span was closed takes that text out of the plan. Round 8 measured the whole remainder -- 43 631 characters, 2.51 %, over 8 of 32 documents with a plan -- down to 0, with the concept count identical at 436 and every hit@8 row holding rank 1 on both K2 bundles and the reference sheet label-identical at 11 of 12. Three steps leak: the orphan check (18 527 characters over 15 of 39 documents), fold_units clause 1 between entries (7 514) and the same clause on the last run (all 17 590 tail characters; with unit_fold=False the corpus tail gap is 0). Round 7's own § 5 does not reproduce: it reports md at 3 of 4 declared headings and a rule:table-block displacing ## 3 Prising, but D1 -- the repair for exactly that -- became the default in the same commit, so the number describes the configuration that existed before the move. On a364ef4 the default recovers 4 of 4, and tests/test_md_declared_headings.py now holds the cell with its cause as a known-negative. Report: docs/2026-09-11-k3-runde8-tabellblokk-og-siste-spenn.md. That last one is ON since 2026-09-10 and is not a segmentation rule at all -- it adds no boundary, and the K2 concept count is identical with and without it (425 = 425 on the 2026-09-09 default). It repairs a measured loss: 32 of the 32 documents that get a plan left the text above their first concept in NO segment. The hole is bigger than that rule, and this is the number to carry: measured 2026-09-10, the pre-move default left 207 435 characters, 11.92 % of the corpus in no segment -- 163 804 above the first entry, 26 041 BETWEEN entries, 17 590 after the last. The rule closes the first part entirely and 79 % of the whole; 43 631 characters, 2.51 %, over 8 of 32 documents remain, and the between-part has a named mechanism (a rule:table-block candidate displacing a DECLARED heading and opening below it). Neither remainder is a ceiling; both are in STATE with their numbers. Until that day the build path called the proposer with no arm flag at all, so a tender PDF that Arm D splits into nine concepts landed as one -- a build path a full arm behind the proposer. Exposing them was not the same decision as moving one, and the two were taken a session apart: which arm ships as the default is the operator's, answered 2026-09-08 as above. Adding a flag still leaves the default byte-identical (measured by digest before and after, and by Arm E over all 43 corpus documents); MOVING the default is the one thing that does not, which is why it took an operator decision and carries an opt-out. Arm C (--max-segment-chars) stays unexposed: no reference has ever been measured for its cap. The two D3 rules read grammars nothing else here reads: a table row's FIRST CELL (a run of bare numeric labels cuts the block that holds them, which is the only way to reach a sheet whose units are rows and the opposite direction from Arm E), and whether a RECOVERED heading's line is a wrapped sentence (a heading is a complete line; quoted regulation and a recovered table row are not). Both remove or add nothing anywhere else: over the 43-document corpus they change 1 and 5 of 39 readable documents, and 0 of 5 docx either way. Reports: docs/2026-09-08-k3-arm-f-mot-enhetsarket.md, docs/2026-09-08-k3-runde2-per-filtype.md and docs/2026-09-08-k3-runde3-per-filtype.md.
  • --assets / --no-assets (0.10.0) is not a segmentation flag either, and it is the first flag here that writes a NON-MARKDOWN file. ON by default. It adds no boundary rule; it changes what the extracted text SAYS, so it sits with the three PDF reader flags rather than with the twelve arms — and like them it must be given the same value on both sides of a plan. The full measurement, the layout and the refusal codes are in the Door B paragraph above; the spec proposal for the layout is docs/plan/okf-assets-section-6-4.md. okf project does not take it: it owns no flag that moves a bundle's bytes, so it gets the default. The corpus harness takes it with the SAME default, for the reason --gate does — a test holds the two paths byte-equal, and two defaults would make that equality depend on which command you ran.
  • --frontmatter KEY=VALUE (K3-19, repeatable) is not a segmentation flag and moves no byte unless given: it stamps a key on every concept of the run, split on the FIRST = and written on ONE line -- a block-form sources is invisible to parse_frontmatter, so the flow form is the only one that survives our own readers. Since K3-22 a scalar goes out double-quoted where a YAML reader would not read it plain, and a flow value goes out as given but is REFUSED (exit 2) when a leaf has no flow form both PyYAML and the guard read -- so a sources URL with a query string, the form K3-19's own flagged build wrote on 2 761 of 2 761 concepts, fails the build. It adds any key and REPLACES only sources and description, the two with a derived layer below them: precedence flag > what the document declares > file name. Every other key the door writes (inbox._door_keys, including Door A's ingest_manifest, which would make that door claim a Door B file) is refused before anything is read. okf project does not take it -- it owns no flag that moves a bundle's bytes.
  • --shell-parent (K3-20) is OFF and is not a segmentation flag either: a plan entry whose span holds its heading alone gets parent_id naming the nearest PRECEDING entry at a smaller level whose own span holds text (propose._link_shells), passing over an empty ancestor; the door writes the existing parent: key. Nothing is copied -- a consumer's own build of the same standard copied the inherited text in and took hit@1 6/6 -> 2/6. The route reads the PLAN, never the row: on R761 it names the ancestor the <sec> nesting names on 710 of 710 shells since K3-21 D (708 before: the 2 misses sit at depth 7, and the reader clipped the outline mark to 6 along with the heading), where reading section numbers gets 686 (12 begins with 1). 35 of 710 have no ancestor holding text and get none. Since K3-21 okf consume reads the key (consume.link_parents, resolved among the concepts of the concept's OWN source_file, because p1 exists in every document): an excerpt carries parent: { concept_id, title }, conditional like req_number, or parent_unresolved: true where the pointer lands nowhere, and a heading-only body gains ONE line Enclosing section: [title](/path) (SPEC SS 5.1 lineage through links, SS 6.1 the recommended absolute form), appended AFTER structure derivation -- read as body text it became a second, unresolved references edge -- and screened on its own. okf check's seventeenth rule, parent_unfollowable, holds the form. Since K3-25 the RANKING does not score that line (consume.DEFAULT_LINK_IN_SIGNAL = False on searchable_text, concept_scores and build_payload; no CLI flag, True still reachable). The line reaches the file and the excerpt exactly as before, so what moved is ORDER and never an excerpt's bytes. Measured: of the newcomers it ever added a question token to, 39 of 39 gained it from the bundle-absolute PATH and 0 of 39 from the link's TITLE, every such token being a segment of the document's own directory -- the saturation shared_id_prefix took OUT of the id signal, back in through the body. Under this reading a --shell-parent bundle delivers what the unflagged build delivers on 16 of 16 rows (list, order and spent), hit@1/8/50 6/6 at both k with the known-positive at rank 1. Exposure today is zero: 0 of 5 shipped bundles carry the line, so 5 of 5 payloads are byte-identical across the move, and the day a bundle carries it is the day the path would have started costing rank instead. okf consume --follow-parent (K3-21 B) is the second form: parent also carries the enclosing concept's text with that concept's own sha256, placed AFTER the cut from the room it left, in rank order, so the delivered set is the same with it as without it; a text that does not fit is cut to the longest prefix that does and marked truncated. Whether either default moves is K3-21 B's measurement, not this line's. Since K3-21 C the index RESOLVES such a parent: it rendered every one unresolved (parent: p1?, 675 of 675) because structure read parent as a document NUMBER and a segment id answers to none. structure._segment_lookup keys (source_file, segment_id) off the concept's own frontmatter and is asked first, inside the pointing concept's document; a value no segment answers to is a number, looked up as before, and a pointer naming nothing keeps ?. The one key keeps its two meanings (inbox.py). Report: docs/2026-09-11-k3-runde20-delt-katalog-og-arvet-kontekst.md.
  • A TWELFTH flag, --pdf-outline, is OFF (round 12, 2026-09-10) and it is the only one here that does not read the extracted text at all: it cuts a PDF at the boundaries its own /Outlines bookmark tree declares. It is NOT Arm D -- --outline-run/--outline-gate are a TEXT heuristic over numbered lines in the extracted text, and this opens a structure index the file already carries. Measured on ONE 701-page process code whose publisher also ships a NISO-STS structure for it, so the fasit is the publisher's own 2 761 titled sections: the shipped default finds 1 967 of 2 761, 0 of its 28 chapters, and 794 of 794 misses have their heading text PRESENT in the extracted text -- the line is read, the boundary is never opened. With the arm: 2 759 of 2 761 (99.9 %), chapter level 28 of 28, concept titles identical to the source after normalisation 2 761 of 2 761 (the bookmark title is complete because it does not come from the page), false positives 163 of 2 182 -> 3 of 2 762, directories carrying two concept files 132 -> 2 with the 65 contents-copy pairs at 0, front-matter concepts 72 -> 2. Consumption: fasit present in the bundle 4 of 7 -> 7 of 7, hit@1/8/50 1/6 - 2/6 - 4/6 -> 3/6 - 5/6 - 6/6; the known-positive is a real concept now and ranks 13, so it is delivered at k=50 and not at k=8 -- the segmentation half of that row is closed and the ranking half is not. The bridge from (page, /XYZ top) to a line index is the whole risk and BOTH routes are measured: extract_text_lines splits lines identically to extract_text on 701 of 701 pages and that check SHIPS per page, the y route and the title route disagree on 0 of 2 762, flat from 0 to 8 pt and collapsing at 12, so the rule carries no tolerance constant; the naive nearest-line rule was wrong on 1 840 of 2 762, one line early every time. The orphan check is NOT applied to a bookmark mark -- it asks whether anything stands under a candidate's first line, the right question for a guess and the wrong one for a publisher's declaration; 683 of 2 762 marks are container sections and applying it scores 2 079. An unresolvable /Dest is dropped and COUNTED (R761 has 0 of 2 763; one of the eight reference PDFs has 2 of 2). NO new dependency: pdfminer.six already ships under pdfplumber in [extract], so uv.lock is untouched and pypdf stays out. Cost 119.22 s -> 183.31 s wall, peak RSS 3 252 -> 3 251 MiB, pages parsed 1 -> 1. The default did not move, and the reach is why: 1 of 8 reference PDFs carries a usable tree, and a bookmark tree is the publisher's CLAIM about its own structure. On the folder where no PDF has one, diff -r is empty against both the arm off and the pre-change tree. Report: docs/2026-09-10-k3-runde12-pdf-outlines.md.
  • Three PDF READER flags, all off, and they sit BEFORE every segmentation flag -- an arm changes how the proposer cuts a text, these change what the text says. --pdf-headings font infers a heading from typography (dominant font size above the document's character-weighted body median AND a bold font name -- the CONJUNCTION measured at recall 1.000 / precision 0.846, where adding weight as a disjunct took precision 0.786 -> 0.524) and emits it as an ATX heading in the SAME markdown the office path produces, so _ATX applies unchanged and no PDF-only heading grammar exists. It is off by measurement, not by caution: against the operator's unit worksheet it takes pdf from 2 of 8 to 0 of 8, losing two exact matches, because on those documents the outline rule already recovers the document's own numbered chapters and a second heading source can only add. The cost of the ATX form is named rather than hidden: a font-inferred heading carries rule:heading and is indistinguishable in the artifact from one the document declared, which is why RULE_POPPLER_SIZE_AND_BOLD was deliberately NOT assigned to it -- that name records a poppler measurement on a path that cannot ship. --ocr reads a page as an IMAGE when its own text never arrived (empty, or (cid:N) codes at or above OCR_CID_SHARE = 0.10, a threshold READ OFF the measured per-page distribution: 834 pages over 32 files, 818 at exactly 0.0 and 16 at 0.93 or above, nothing between). Its engine is the optional ocr group (rapidocr/onnxruntime/pypdfium2) and never a runtime dependency; a packaging test pins both halves, and without it every affected file is a coded rejection (extractor_ocr_group_missing), never a crash. --ocr can never become a default -- an optional dependency in the default path would make an ordinary install fail on the first scanned page. Report: docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md. --pdf-headings font-reserve is the third value on that same option (none, font, font-reserve — three answers to one question, so no caller can ask for two at once): the same typographic rule applied ONLY where Arm D's outline gate admits no run at all, typography as a second heading source where there is no first one. The condition lives in ONE function (propose.heading_reserve_applies) that the proposer and the door both consult, the door receiving it as a callable the way it already receives gate — a plan indexes the exact string it was proposed against, so a reserve firing on one side only would turn every document it touches into a coded rejection. It reads the gate AS CONFIGURED, so at --outline-run 0 it is unconditional and equals round 4's "font instead of Arm D". Off, and the measurement is that it changes nothing measurable: on the twelve-position reference it alters not one cell — the five positions where it fires are one PDF whose glyphs carry no ToUnicode mapping and four office documents the PDF reader never touches — and the position it was built for has three outline runs, so the reserve is silent there by construction. Its reach is real but unrated: 4 of 39 readable corpus documents, none in the sample. Report: docs/2026-09-08-k3-runde5-hitat8-og-skriftakse.md, which also corrects two of this repository's own published figures — the S7 candidate ranks (96 of 629 / 159 of 492 were measured with the cost vocabulary reaching only half the ranker; consistently scored they are 10 of 629 and 19 of 492) and round 4's attribution of that concept's non-delivery to the default move (it is not delivered on the Arm B bundle either, by a different mechanism). hit@8 over the six published questions holds at 5 of 6 on both K2 bundles, so the default move cost the retrieval side nothing.
  • Judge a bundle: okf quality <bundle> (G37, 2026-09-12). A per-file-type verdict, with the denominator on every line, and it is a SEPARATE command from okf check on purpose: check is the contract check, and a green one is not a quality gate -- measured 10.09 by vegnormal-okf, three arms over one corpus all gave 0 findings and exit 0 while hit@k ranged 6 of 6 to 0 of 6. Three verdicts and no fourth (PASS / FAIL / UNMEASURED), a type with no measured threshold is never PASS, and exit 3 exists for "nothing could be judged" so exit 0 over a table of unmeasured rows cannot be a silent pass (0 clean, 1 a FAIL, 2 did not run). Two bars today, both structure_null_share off the pinned 43-document reference -- .pdf 8/32, .docx 2/5 -- plus one definitional bar for every type (0 empty bodies, measured 0 of 8 602 concepts over four bundles). A bar needs five documents on BOTH sides, its own and the judged bundle's, which was found by RUNNING the gate: a one-PDF bundle scored 0 of 1 against the 32-document reference and read as PASS. The bars are regression bars against a pinned artifact, never a quality claim, and the defect that started G37 -- the HTML arm's 1 148 of 2 761 boundaries -- is UNMEASURED here, because no bundle-only metric reaches it: three candidates were measured over the same four bundles and two order the known-bad and known-good arms the wrong way round (duplicate titles within a document 0/3 206 against 349/2 761; short concepts 5.6 % against 14.6 %), while the third (duplicate titles across the whole bundle, 37.8 / 16.3 / 12.6 / 5.7 %) orders them correctly and ships anyway WITHOUT a bar, since any bar separating them sits between the two bundles that define it. The gate's denominator is the bundle's, never the corpus's: a rejected document leaves no row at all (the pinned corpus holds 33 PDFs, the bundle shows 32), so the run log is printed beside the counts and a bundle without one says so. Three of the four evidence corpora (n100/n200/n500) carry source_file on 0 of 446 / 0 of 1 133 / 0 of 270 concepts, so they name no file type and every row is UNMEASURED -- the order expected them to PASS. Thresholds, the nine bundles and the premises re-measured: docs/2026-09-12-g37-terskler.md. README publishes the bars behind <!-- quality-thresholds: ... -->, pinned to the code AND the document by tests/test_docs_promises.py. --fasit <json> REACHES that defect (G37b, 2026-09-13) and is the only input this gate takes: one whole-bundle row, boundary_share = declared boundaries that became a concept over declared boundaries. It is whole-bundle and never per file type, because a fasit names ONE document's sections and a bundle can spread them over 828 source files -- which the known-bad arm does. The normalisation was derived before the metric was built, not guessed: strip all whitespace, lowercase, reproduces the fasit's own norm from its own title on 2 761 of 2 761 rows (alphanumerics-only scores 58 -- it eats the . in 2.1Hovedprosesser). A boundary is recovered in EITHER of two forms and neither is a fallback: the concept's normalised title equal to norm, or the pair (concept's own directory, residual title) -- the literal form wants the declared title WITH its numbering token, the pair form WITHOUT, and no bundle can offer both, because okf's default route moves that token into the concept id. Measured on the known-good arm: literal 22 of 2 761, paired 2 737, either 2 759 (99.9 %); on r761-2025-d1 the split is exactly inverted (2 727 literal, 0 paired), so a gate scoring one form alone reports a 99.9 % arm as 0.8 % and calls it a segmentation defect. The two forms are vegnormal-okf's M8 correction, which they took verbatim from THIS repository's round-14 report -- the instrument reproduces both so the two repos cannot silently measure different things. One bar, at the pinned artifact's own value: 2 759/2 761 with corpora = 1, and P2 is in the OUTPUT and not only in the document (N = 1 corpus on every boundary row). The known-bad arm is 1 148 of 2 761 (41.6 %), now FAIL + exit 1 where the bundle-only gate gave exit 3. --fasit is an ASSERTION (the posture okf consume --ref has) that this bundle is a build of the document the fasit describes: the K2 reference and n100-2023 both score 0 of 2 761 and read FAIL, which is the assertion being wrong and not the bundle -- a gate telling those apart would need a bar read off the bundles it judges. The bar is TIGHT and the cost is published: 2 of 4 R761 builds fall under it (2 752 and 2 727 of 2 761), while any bar between 41.6 % and 98.8 % separates the known-bad arm from every R761 build measured -- the shipped one is the only point in that interval read off a pinned artifact. An unreadable fasit exits 2 with its reason, never a quiet UNMEASURED, and a fasit under five rows is UNMEASURED (MIN_DECLARED_FOR_A_THRESHOLD, the document floor in the fasit's unit). Without --fasit the command is byte-for-byte what it was, held by a test. README publishes this bar behind <!-- quality-boundary-threshold: ... -->; SS 7 of the threshold document carries the seven bundles and the honesty limits.
  • Consume a bundle: okf consume <bundle> --question "<q>" [--k N] [--limit N] [--out PATH] [--ref IDENTITY] — the pre-pass docs/consumption-contract.md § 1 defines, and the only reading direction this library has. It moved into the package 2026-09-08 (O5) and the move it was written for is the one that happened: build_payload(...) was always the entry point with the CLI a thin main(), so it was a move and not a rewrite. What forced it was the generated skill — from tools/ it emitted python3 <absolute path>/tools/okf_consume.py, so the skill could not be moved, shared or run by anyone without that clone. tools/okf_consume.py remains as an ALIAS (sys.modules[__name__] = _impl, never a re-export: a re-export binds copies, and a caller patching one patches a binding the implementation never reads). Deterministic and offline by construction: no model call, no socket, no clock, stdlib plus this package only. It walks the index tree, never a directory — § 9.2 forbids enumerating one unless the named profile says the index is derived, and measured, entries_match_directory is True for STRICT_V1 alone; the walk loses nothing (629 = 629 on the K2 bundle, controlled in a test against the very method § 9.2 forbids). --ref is an assertion, never an override: the emitted identity is always the computed one, because § 3.3 exists to stop a payload being labelled with an identity its bytes do not have. Three exit codes: 0 written, 1 refused, 2 did not run. Every excerpt carries the concept's title, plus req_number, the § 5.1 address sources, and every top-level source_* key by PREFIX — never an allowlist, because a list names the producers its author thought of and one bundle locates by source_element_id on 269 of 274 concepts. A prefix, never a substring (resource_owner is not a locator). An absent key stays absent and an undecodable address is named (sources_unreadable). Since K3-21 an excerpt also carries parent, the enclosing concept's concept_id and title -- never the raw segment_id -- or parent_unresolved, and the contract's SS 8 point 6 says the checker reads it. sources is READ in both YAML forms because the two real bundles disagree (flow 629/629 on one, block 270/270 on the other) — reading block is not a licence to write it, the emission rule is unchanged. Contract § 8 makes title a MUST (checker code excerpt_unnamed) and the rest SHOULD, because they are conditional on the producer. The measurement behind it: rank 1 of 8 on 3 of 3 bundles, correct answer on 1 of 3.
  • Connect a bundle to Claude Code: `okf skill --out ` instantiates `skills/okf-consume-template/` for THAT bundle — its id, ref, concept count, conditional-field denominators, whole-bundle cost and breaking point, all measured, plus a reference payload the checker accepts. It was kept in `tools/` until 2026-09-08 because a wheel-installed `okf skill` would emit a command pointing at a file the wheel does not carry. That objection was about what the GENERATED skill NAMES, and O5 answered it by changing that: the emitted commands are `okf consume` and `okf check`, names on PATH. The template and `docs/consumption-contract.md` (the § 7.4 known-positive) are force-included into the wheel from the file they are authored in — one authored copy, no committed duplicate. The form was chosen on a measurement: the contract checker passed the UNFILLED template and passed a skill built for another bundle, so it could not tell the two apart — the choice rests on § 5/§ 6.4/§ 7.6 being per-bundle numbers a generic skill can only leave as holes or state falsely, and that argument NEVER rested on conformance, so closing the measurement leaves it standing. **That half is CLOSED 2026-09-10 by `bundle_mismatch`, the checker's SIXTEENTH rule** (seventeen since K3-21 added `parent_unfollowable`; `RULES` is a tuple and `Report.rules_evaluated` is `len(RULES)`, so every published «15 rules» line is now «16 rules» — a contract change for anyone quoting it). It compares the identity a skill DECLARES against the identity its payload declares, **both `bundle_id` AND `ref`**: three distinct builds on this machine carry the one `bundle_id` `k2-trinn1-20260903` at three refs, so an id-only rule would pass a stale skill, which is the case the generated skill warns about in its own words. SS 3.3 is the ground — «a version is the producer's assertion; a ref is a fact about bytes» — and SS 3.1 for the same question one level down, an excerpt naming a bundle its payload does not. Reproduced on `113b3f8` before any code moved, three pairs at `conformant: 15 rules over 8 excerpts and 438 withheld entries, 0 findings` and rc 0; after, all three at rc 1 with **one** finding over 16 rules — a foreign corpus's payload, the unfilled template, and (the arm that separates a whole rule from half of one) a payload sharing the id at a foreign ref. **An identity the rule cannot READ is a finding, never a silent pass**, or the template passes again. The right pair is untouched at rc 0 / 0 findings, and `{}` is unchanged at **9** findings because a payload declaring no identity stays `ref_missing`'s defect — no rule restates another. NO new field was needed: the identity already lived in the generated skill's prose, now one authored copy in `skill.identity_line` read back by `contract_check.skill_identity`, generated skill bytes byte-identical before and after on both tracked bundles. **The rule compares DECLARED against DECLARED and never opens the bundle**, so a payload lying about its own `ref` still passes — that is `okf consume --ref`'s job and needs a bundle path this command deliberately does not take; «closed» means the three measured forms now fell, not that no fourth exists. Report: `docs/2026-09-10-k3-runde15-bundle-mismatch.md`. The first instantiated consumption skill is `skills/okf-consume/`; the measurement behind the form, including the control that FAILED, is `docs/2026-09-07-okf-konsumskill-maaling.md`. **That skill was the new rule's first REAL find**: hand-filled for K2 before `okf skill` existed, it declared no bundle identity a reader can act on, so `okf check` refused it against its own shipped example payload (rc 1, 1 finding) — **1 of 1** shipped hand-made instantiated skill. **Since 2026-09-11 (K3-18) it is GENERATED** by `okf skill` from `examples/ingest-golden-segmented-okf-v0-2/expected-bundle`, the bundle its payload always came from, with the command in its `references/README.md`: `--force` and `--example-question "Hva sier veiledningen om krav?"` are both required (the payload test asserts bytes for that question), and the checkout prefix is then stripped, because `okf skill` writes the bundle root and the skill path ABSOLUTE when `--out` is not under `.claude/skills/`. The pair is rc 0, 17 rules (16 before K3-21), 0 findings, and a test holds the shipped bytes to the generator's. Its frontmatter `name` is now `b-golden-segmented-okf-v0-2-consume`: Claude Code takes a project skill's COMMAND from the directory name and uses `name` only as a display label, and nothing here named the skill `okf-consume`. Report: `docs/2026-09-11-k3-runde18-konsumskillen-regenerert.md`. **The ranking is this repository's own choice** — the contract binds a payload, not a retrieval algorithm (§ 10) — and it has FOUR optional widenings. Three are **off by default** and keep the default payload byte-identical; the fourth (`--tie-shared-rank`) became the default 2026-09-10 and is the one change in this repository that alters a payload with NO bundle changing, so a consumer pinned to the old excerpt order needs `--no-tie-shared-rank`. `--cost-vocabulary`: a declared cost/price/quantity vocabulary family that bridges a question and a document naming money with different words, gated on the QUESTION carrying such a term, so a question without one is byte-identical either way. It moves a measured case from candidate rank 249 to 10 and does **not** deliver it — the budget is a second, independent lock. Measured, with two rules falsified before building and the `k`-sweep that showed a higher `k` can EVICT a gold concept, in `docs/2026-09-08-blindsone-below-k-k2.md`. `--reserve-top-rank` is that second lock: the pack is an exact knapsack over a SUM, so it has no opinion about rank and out-sums a top-ranked excerpt costing a large share of the budget. The flag gives rank one its bytes first, AFTER the `over_budget_alone` pre-exclusion and never before, and declares `budget.reserved` in the payload. It fixes the eviction and does **not** close the mandate-shaped blind spot (that concept ranks 10, not 1); the budget stays the caller's decision, because deriving a limit from the corpus was measured and falsified — two defensible derivations, 49x apart, one of them breaking the known-positive. `docs/2026-09-08-blindsone-laas2-budsjett-k2.md`. **A SIXTH flag, `--stem-prefix`, is ON since 2026-09-09** and is the second change here that alters a payload with NO bundle changing (opt-out `--no-stem-prefix`). `MIN_SHARED_PREFIX = 4` exists for Norwegian compounding and also matches four characters that are not a stem: on the pinned 453-concept bundle, control first, `under` occurs 79 times by equality and matches 172 by prefix, `bilateral` occurs **0** and matched **400 of 453** through `bilag`, `standhaftig` 0 and 219 through `standard`. The two extra known-negatives were FOUND, not chosen -- every 4-character prefix ranked by document frequency, then a real word taken from the widest. **Three candidates were measured and all three failed on the SAME row**: a longer floor (5-8), a coverage share (0.5-0.8) and a long-words-only floor (>= 8). Decomposed, row 1's token `prisene` reaches its gold document through `pris|sammenstilling` on four characters -- 0.57 of one word and 0.22 of the other -- so **the over-match and the wanted match are one mechanism** and no threshold on length or coverage separates them. The fourth candidate does: the shared prefix must be a WORD the bundle uses. `bilateral` 400 -> 0 and 512 -> 0, `standhaftig` 219 -> 56 and 235 -> 33, every hit@8 row keeping rank 1 on BOTH bundles. `undersjoisk` stops at **162** because `under` IS a word here -- a genuine Norwegian morpheme, so that residual is a different answer, never a ceiling. The vocabulary is the BUNDLE's own, so the rule makes a payload corpus-dependent the way `rarity_weights` already is. **A SEVENTH flag, `--source-quota N`, is ON at 2 since 2026-09-10** (opt-out `--no-source-quota`) and is the third change here that alters a payload with NO bundle changing. It caps how many DELIVERED places one `source_file` may take, cutting where `shortlist = candidates[:k]` cuts, so the freed place goes to the next candidate and `k` is still delivered in full. The defect was measured OUTSIDE this repo on a 3206-concept bundle of a published handbook: the code's own process overview is **28 of 3206 concepts (0.87 %)** and **8.0 % of the source characters** yet took **8 of 8** delivered places on one question and **7 of 8** on the known-positive, which was not delivered at all -- identical at 343 and 1651 concepts, so it is the corpus's COMPOSITION (it holds its own table of contents) and not its size, and a split would move it rather than remove it. Swept over {2, 3, 4, off} on three bundles with the fasit prefixes validated against the bundle FIRST (that control caught a defect in the measuring query itself): at 2 and 3 hit@8 goes **5 of 6 to 6 of 6 on BOTH K2 bundles** with all five standing rank-1 rows unmoved -- the recovered row had missed on every bundle and every configuration measured until now -- and at 4 and off it stays 5 of 6. On the handbook bundle hit@8 goes **2 of 6 to 4 of 6** and the dominant document's share **8 of 8 to 2 of 8**. 2 rather than 3 on rank. **What the gain is NOT:** hit@8 asks whether the gold DOCUMENT was delivered and a document quota raises how many distinct documents a payload holds, so that metric is not neutral with respect to this rule; the five rows already at rank 1 are, and did not move. **The adverse case is named:** a one-document bundle has one `source_file` on every concept, so the quota would deliver 2 where `k` were asked -- the shortlist is topped back up from the best-ranked over-quota candidates, making such a bundle byte-identical to the quota being off. The `WITHHOLDING_RULES` vocabulary goes six to seven (`source_quota_exceeded`, a DIVERSITY drop and not a relevance one) and is published in the contract SS 5.3 and in the generated SKILL.md, verified by reading the generated file. Editing the contract moved the SS 7.4 known-positive (12 563 -> 13 238 encoded, delta 336 -> 345), which is that coupling working. `--rarity-weight` was measured against the same defect and does NOT repair it -- it leaves the dominant document at 8 of 8 places on the question it floods -- and stays off. `--rarity-weight` is the third: each lexical hit weighs `log(N/df)` over the bundle's own concepts instead of 1, so an identifier is not worth what a common verb is worth. It enters the RANKING and never the GATE — `lexical` stays a count, because a word every concept carries weighs exactly 0 and a weighted gate is what `54a0bc2` falsified. Off by default BY MEASUREMENT: it delivers one of three requirement lookups and takes a priced sheet from candidate rank 10 to 2, leaves one gold unmoved and costs another seven rank positions. Two limits are decomposed rather than guessed, and both are someone else's mechanism: `MIN_SHARED_PREFIX = 4` makes a unique identifier read as 135-of-446 common, and RRF consumes RANKS, so no weighting inside a signal can move a gold that already leads it. `docs/2026-09-08-sjeldenhetsvekt.md`. `--tie-shared-rank` is the fourth and **the only one that is now ON** (2026-09-10, opt-out `--no-tie-shared-rank`). It is a correction to the TIE-BREAK rather than a weight: RRF ranks every concept in every signal, including a signal that scored them all the same, and the declared `(-score, concept_id)` tie-break then orders that group by id. Measured on N500, whose document prior has **two** distinct values over 270 concepts, that signal contributed alphabetical UUID order and put a concept answering 7 of 7 question tokens at fused rank 14 — outside the cut — behind concepts sharing only `tunnel` and `vann`. Under shared ranks it is rank 3 and 2 of the 16 covering concepts are delivered. **It shipped OFF on a measurement that was CONDITIONAL and stopped being true in a commit reported as changing nothing.** The published cost — hit@8 falling 5 of 6 to 4 of 6 — is real only at `DOCUMENT_PRIOR_EXPONENT` 1.0. Round 6 moved that exponent to 0.5 for an unrelated reason and correctly reported it moved no hit@8 row; nobody measured the PAIR. Swept 2026-09-10 over 2 exponents x 3 bundles x 6 rows: at 0.5 the rule holds `[1,1,1,1,1,]` on all three bundles and FIXES the split bundle's row 1 (2 -> 1), which is what let `--sheet-section-rows --keep-table-heading` become a build default. **A flag's "off by measurement" is a measurement of a CONFIGURATION, not a property of the flag** — when a constant it interacts with moves, its default is unmeasured again, and nothing in the tree says so because the two decisions live in different files. The adverse case is recorded rather than hidden: on a synthetic 30-concept fixture where one signal separates and two do not, shared ranks move a gold from rank 18 to 30 (`tests/test_okf_consume.py`). Note also that `docs/2026-09-08-sjeldenhetsvekt.md`'s figures were measured under the older tie-break and are NOT re-measured — on one fixture the change takes the weight's gold from fused rank 18 to 1. `docs/2026-09-08-rangeringsbom-sammensatte-ord.md` and `docs/2026-09-10-k3-runde7-forste-spenn-og-rangeringen.md`. The other three stay off. A FIFTH flag is not a ranking widening and is listed apart: `--withheld-titles` gives each `withheld` entry the concept's `title`, so a reader can see WHAT was withheld without reading the bundle (§ 2.2 forbids going to look). The code is 11 lines; the bytes are the reason it is off. Measured, it grows an N500 payload 37.9 % and takes the 629-concept K2 bundle's BOOKKEEPING to 122 704 B — past the 120 000-byte limit itself — which would have made the breaking point then published in the hand-filled K2 copy of `skills/okf-consume/SKILL.md` ("~75 KB at 629 concepts … at roughly 8 000 concepts") false on the day it shipped. That copy was replaced by a generated one 2026-09-11; the measurement of the flag stands.

Workflow

  • TDD: no production code without a failing test first.
  • This repo is published PUBLICLY (open/ namespace on Forgejo). STATE.md and docs/oppstartsprompt.md are LOCAL-ONLY (gitignored) — never commit session state or internal briefs. No secrets, sober English prose, no marketing language.
  • Consumer content stays at form level in public files. Some consumers we read are private (claude-code-llm-wiki is, pending an Anthropic ToS assessment). Key names, counts, gate names, profile fields and contract shapes are publishable; page bodies, full title or path lists from a consumer's bundle, and Anthropic-derived prose are not. Findings about a private consumer's data go back to them through coord, never as a file here. This costs nothing — every question this library asks of a consumer is about shapes and key sets — and it is not reversible once pushed.
  • After git commit: push to Forgejo (git push origin) immediately. Never GitHub.

Communication patterns

Linking to local files

When pointing to local files in responses, always use markdown link syntax with a descriptive name:

  • Use [Human-friendly name](file:///absolute/path) — never bare file:///... URLs or autolinks <file://...>.
  • Always use absolute paths. Never ~/ or relative paths.
  • For multiple files, render as a bullet list of named markdown links.

Why: bare file:// URLs only render the first as clickable across multiple lines. Named markdown links make each entry independently clickable and look cleaner.