# Connecting an arbitrary OKF bundle to Claude Code **Date:** 2026-09-08. **Order:** `20260908T124346Z-4668330384` (C1). **Commits:** `17c49fc` (step 0), `c95d189` (the prefix rule and the generator), and the commit this document lands in. **Baseline:** `b6a8c8b`, suite 1339. The question this answers is the operator's, in their own words: *"That Claude Code should be able to work with an arbitrary OKF bundle that we give Claude information about is extremely important."* Three things had to hold — build a bundle from a folder, connect it to Claude Code, and ask it a question whose answer is a document or part of one. The first existed. The second existed only as one skill hand-instantiated for one corpus plus a template full of holes, installed nowhere, never run inside Claude Code. The third was therefore not testable at all. --- ## 0. What was measured and what was not | | measured | not measured | |---|---|---| | the excerpt fields | yes, on three bundles and two fixtures | — | | the `source_*` pass-through rule | yes, on the bundle that broke the allowlist | whether a fourth producer uses a key not starting with `source_` | | the form choice (generator vs generic skill) | yes, including a control that came out against the obvious gate | operator preference; only the mechanics are here | | the generator's refusals | yes, four of them, each with its code | — | | end to end in Claude Code | **four questions, two bundles, one run each** | a hit rate. Four questions is a demonstration | | the ranking | unchanged, byte-for-byte, on the K2 control | — | | cost per answer in tokens | — | no tokenizer on this machine; costs are in `BUDGET_UNIT` bytes | --- ## 1. Reproduction, first The known-positive, before anything was changed: ``` python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md \ --payload # conformant: 15 rules over 8 excerpts and 621 withheld entries, 0 findings # exit 0 ``` Then a fresh bundle, built the way an operator would build one: nine documents (two spreadsheets, three word-processor files, four PDFs) copied out of the corpus tree into a scratch folder. ``` okf build /tmp/c1-fresh-src --bundle /tmp/c1-fresh-bundle \ --bundle-id c1-fresh-20260908 --okf-version 0.2 --ingested-at 2026-09-08T00:00:00Z # merged + coded rejections = 9; N = 9 (substantive 9/9, rejected 0/9) # 5.15 s, 0.572 s per file ``` 54 concepts, ref `sha256-tree:7245851d…9685f71`. The index walk reaches 54; a directory walk finds 55, and the one extra is `log.md`, which the walk excludes by name. 54 = 54, controlled. One question with a known gold concept: ``` python3 tools/okf_consume.py /tmp/c1-fresh-bundle \ --question "Hvordan skal prisene fylles ut?" --out /tmp/c1-fresh-payload.json ``` `considered 54 = withheld 46 + delivered 8`, `spent` 82 030 of 120 000, and the gold concept — the priced spreadsheet — at **rank 1**. That is the position an operator starts from. --- ## 2. Step 0: the excerpt carries the key the question is asked ON Two independent measurements pointed at the same missing field. `portfolio-optimiser` ran three paid arms: the gold concept was delivered at **rank 1 of 8 on 3 of 3 bundles**, and the model answered correctly on **1 of 3**, because a delivered excerpt carried `concept_id`, body text and nothing the document is known by. The previous session measured it from the other side: the provenance it had just written into every concept did not reach the payload at all — nine members, not one of them provenance. `excerpt_for` now carries `title` unconditionally, and `req_number`, the SPEC § 5.1 address `sources`, and the locator keys when the producer wrote them. Four design decisions here were made by measurement rather than by taste. **(a) `sources` is read in both YAML forms.** The two real bundles disagree completely: | | flow `sources: [{...}]` | block `sources:` + indented | a locator key | |---|---|---|---| | the 629-concept procurement bundle | **629 of 629** | 0 | 618 `source_offset`, 586 `source_pages` | | the 270-concept normative bundle | 0 | **270 of 270** | 0 of the five this library writes | `parse_frontmatter` skips indented lines on purpose — a nested `title:` arriving later would substitute for the document's — so the block form arrives as `''`, present and unreadable. A flow-only reader delivers the second bundle with **no address at all**. Reading the block form is not a licence to write it; the emission rule is untouched. **(b) An undecodable address is named, not dropped.** `sources_unreadable: true` is a third state beside "an address" and "no address". Collapsing it into either reports something nobody measured. **(c) An absent key stays absent.** Writing `req_number: ""` would assert that the producer wrote an empty identifier. Contract § 6.4 forbids exactly that inference, from the other direction. **(d) The pass-through is a PREFIX, and this one came from a peer's measurement.** The first implementation used the five locator keys this library writes. The PM measured it against the normative bundle and found the excerpt still lost that bundle's locator: 269 of its 274 files carry `source_element_id`, a key that repository chose under this chain's own rule ("the key says what it indexes") and that this library never writes. An allowlist is a list of the producers its author thought of. The rule is now every top-level key beginning with `source_`, and a prefix rather than a substring — `resource_owner` contains the literal and is not a locator. *A correction to my own measurement, stated because it is the reason the peer had to send it twice:* my first sweep used `grep -rhoE "^source_[a-z_]+:"`, whose character class excludes digits, so it reported `source_element_id` and silently missed `source_sha256` on 270 of 270. The generator's own counting — over parsed frontmatter, not a regex — found both. A pattern that cannot match what it is looking for returns a zero that reads like a fact. ### The K2 control: the field moved, the ranking did not Same bundle, same question, same `k`. The "before" run is a copy of the tool taken from `git archive b6a8c8b`, run with its own `docs/` beside it, so the editable install cannot leak the new code into the baseline. | | before (`b6a8c8b`) | after | |---|---|---| | delivered concept ids, in order | 8 | **identical** | | `text_sha256` per excerpt | — | **identical** | | `withheld` list | 621 entries | **identical** | | denominators | 629 = 621 + 8 | **identical** | | payload bytes | 108 877 | 113 143 (+3.92 %) | | `spent` (`BUDGET_UNIT`) | 18 606 | 22 210 (+450.5 B per excerpt) | | excerpt members | 9 | **17** | | changed lines in the whole payload | — | 99 | | payload sha256 | `afb7e811…` | `c94919c4…` | **It is the field, not the ranking.** Nothing about which concepts were chosen, or in which order, moved by one byte. The contract document's own bytes moved twice while § 8 was being written, so the budget instrument's known-positive moved with them — by design, since a stale known-positive is meant to be a loud failure: 10 349 → 12 049 → **12 563** measured, 10 060 → 11 719 → **12 227** raw, delta 289 → 330 → **336**. Contract § 8 gains `title` as a MUST with its own checker code (`excerpt_unnamed`, so the checker now runs **15** rules, was 14), and `req_number`, `sources` and the locators as SHOULD — they are conditional on the producer, and a bundle whose concepts carry no identifier cannot deliver one. --- ## 3. The form: a generator per bundle, chosen against a control that failed Two candidates. (a) A generator that instantiates the template for one bundle. (b) One generic skill taking the bundle root as a runtime argument. **The obvious gate does not work.** The contract checker was asked to tell them apart, and cannot: ``` python3 tools/okf_contract_check.py --skill skills/okf-consume-template/SKILL.md \ --payload /tmp/c1-fresh-payload.json # conformant: 15 rules over 8 excerpts and 46 withheld entries, 0 findings — exit 0 python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md \ --payload /tmp/c1-fresh-payload.json # conformant: 15 rules over 8 excerpts and 46 withheld entries, 0 findings — exit 0 ``` The **unfilled template** passes against a real payload, and a skill built for a different corpus passes against this one's. So conformance could not decide the form, and any claim that "(b) fails the checker" would have been false. The checker reads section headings, the marking vocabulary and the payload; it does not read whether a number in the prose is true of this bundle. **What decides it is what the skill has to state.** § 5's denominators, § 6.4's conditional-field list and § 7.6's breaking point are per-bundle numbers. A generic skill can only leave them as holes — the template's own definition of *unfinished* — or carry another corpus's numbers, which is worse than a gap, because a stated cost that is false for this bundle is a measurement failure. Instantiating is what makes them true. Steps from "here is a bundle" to the first correct, marked answer: | | (a) generator | (b) generic skill | |---|---|---| | one-time setup | 1 command per bundle | 1 copy, once | | typed at question time | the question | the question **and the bundle path, every time** | | with two bundles installed | the model selects on the skill name — **measured, it did** | nothing selects; the caller must say which | | per-bundle denominators | measured into the file | absent or false | (a), and the deciding column is the third: the operator's requirement is "one *or more* bundles", and (b) has no mechanism for the "more". `tools/okf_skill.py` therefore instantiates. It measures, per bundle: id, ref, concept count, the conditional-field table with a denominator for every field (the `source_*` rows **discovered** from the bundle rather than listed), the whole-bundle cost by the gate's own instrument, the share one measured answer spent, and the concept count at which the `withheld` bookkeeping alone reaches the limit. It also runs the index-walk-against-directory control **once, at generation time, never on the question path** — § 9.2 binds a consumer reaching for context at query time, and a build-time control is what turns "the walk loses nothing" into a number the skill can quote. It lives in `tools/` for the reason `okf_consume.py` states for itself — outside `src/`, so no consumer's install surface changes — and for one more: a wheel-installed `okf skill` would emit a command pointing at `tools/okf_consume.py`, which the wheel does not contain. **Red tests first, and every gate the checker lacks is one of them**: no placeholder survives (with the pattern shown able to find, against the template); the skill names its own bundle's id and ref and **not** the other bundle's; its commands are absolute and point at files that exist; it reports a denominator per conditional field; it is byte-deterministic for a given bundle and destination. Four refusals, each with its code: a directory with no index (`bundle_unreadable`), an index with no `bundle_id` (`bundle_id_missing`), a bundle with no concepts (`bundle_empty`), and an occupied target without `--force` (`target_occupied`). No model call, no new dependency. --- ## 4. End to end in Claude Code Two skills installed side by side in a scratch project outside this repository, one per bundle. Each row is **one run**; none was repeated to get a better one. ``` python3 tools/okf_skill.py --out /tmp/c1-scratch/.claude/skills/-consume cd /tmp/c1-scratch && claude -p "" ``` | # | bundle | question | skill named in the prompt? | result | |---|---|---|---|---| | i | fresh, 54 concepts | "Hvordan skal prisene fylles ut?" | yes | **pass** | | ii | fresh, 54 concepts | a topic measured absent from the bundle | yes | **pass** — refused | | iii | third-party, 270 concepts | "Hva krever Krav 10.2—2 i N500? Gjengi det sentrale vilkåret." | **no** | **pass** | | iv | fresh, 54 concepts | "Hvordan skal prisene fylles ut?" | **no** | **pass** — and selected the right one of two | **(i) The known-positive.** Rank 1 of 8 delivered. The answer opened by naming the document and the sheet, and cited the locator the previous session put into the concept: *"kilde ``, ark `Prissammenstilling`, rader 1 til 100"*. It marked with all five required literals, reported `considered 54 / withheld 46 / delivered 8`, named `below_k` as the rule that dropped a concept it flagged as possibly relevant, reproduced the budget known-positive (12 563 = 12 563), and ran the contract check itself: 15 rules, 0 findings. **(ii) The known-negative.** The absence was measured **before** the question was asked, with the query shown able to find first (`pris` matches 15 files; `jernbane|signalanlegg` matches 0 of 54). The answer: *"dokumentene i bundelen stiller ingen krav til jernbanesignalanlegg"*, marked `[sourced-not-sufficient]`, with its own known-positive control for the regex it used to check the delivered texts, and with the 46 withheld concepts named as `[unread]` rather than as absent. It named external standards only under `[unverifiable-from-bundle]`. **(iii) The third-party bundle.** Measured on the tree on disk at the time: `vegnormal-n500-2024`, 270 concepts, ref `sha256-tree:673a0c2c…d6b5b3ad` — **not** the V1 tree the order named, because that repository landed a rebuild during the day. The gold concept came back at rank 1, and the answer quoted the requirement verbatim against the concept body, named the title and the `req_number`, the `sources` resource, the concept id and the excerpt digest, and stated `adjudication: unknown` and `trust_tier: unverified` as measured absences at the producer rather than as judgements about the source. **One defect this run found, and it was in the prose.** The first (iii) run said *"the producer wrote no page locator (no `source_pages`), so the address is at document level"* — while the excerpt in front of it carried `source_element_id`, that bundle's own locator. The excerpt was right and the **skill's citation guidance was wrong**: it listed the four keys this library writes instead of telling the reader to cite whichever `source_*` keys are present. Fixed in the template and in both instantiated copies; on the re-run the same question returned *"element `id-4ca67e2f-…`"* in the source pointer. That is two runs of one question, and the second is a measurement of a **changed artefact**, not a retry of the first. **(iv) The trigger control.** No skill named in the prompt, two skills installed. It selected `c1-fresh-20260908-consume` on the question alone, ran the pre-pass and the checker, and cited `source_sheet` and `source_rows` by name. So no special invocation syntax is needed; the generated `description` — which names the bundle id, the concept count and the ref — is enough to route on. ### Invented numbers: zero Every numeric or identifier-shaped token in each answer was extracted and checked against a denominator stated per run. | run | unique numeric/id tokens | not found in the denominator | |---|---|---| | i | 62 | **0** | | ii | 45 | **0** | | iii | 35 | **0** | For (iii) the denominator is the delivered payload itself, reproduced exactly — the model's reported `spent` of 11 941 equals the payload rebuilt here. For (i) and (ii) the model composed its own question string, so its payload is not byte-reproducible; the denominator there is the whole bundle text **plus** the payload's own identities (the bundle ref, every concept-file digest, every delivered-text digest) **plus** the figures the skill itself supplies. A first pass that omitted those identities reported eight and nine "missing" tokens, every one of them a fragment of a correctly quoted digest — the wrong denominator, corrected here rather than published. --- ## 5. A finding about the ranking, not about the skill The first attempt at (ii) asked the normative bundle about water and frost protection in a subsea tunnel. The model answered `[sourced-not-sufficient]`, said the relevant requirements "if they exist, are among these 262 withheld", and refused to go around the cut. Checked afterwards, **the bundle does cover it**: 17 files match `frostsikring`, 10 match "vann- og frostsikring", 13 match `undersjøisk`, and none of them was ranked into the cut. So it was never a known-negative — it was a **retrieval miss**, and it is recorded as one rather than quietly replaced. Two things follow, and they point in opposite directions. The consumer behaved exactly as the contract asks: it did not fabricate, it named its denominator, and it reported its own zero as unmeasured because the `withheld` entries carry only ids and rules, no titles. The ranker did not: a compound Norwegian question ("vann- og frostsikring", "undersjøisk") was out-ranked by concepts sharing only "tunnel" and "vann". That belongs with the fusion work already open, and it is the same shape as the `MIN_SHARED_PREFIX` finding. A second, smaller observation: `withheld` entries carry `concept_id` and `rule` and no title, so a reader who wants to know *what* was dropped cannot tell without reading the bundle — which § 2.2 forbids. The consumer in run (ii) identified this itself and reported it as unmeasured. Not fixed here; noted. --- ## 6. Suite | | before (`b6a8c8b`) | after | |---|---|---| | `pytest -q` | 1339 | **1373** | | `ruff check src tests tools` | clean | clean | | `ruff format --check .` | clean | clean | | `mypy src` | 17 files, clean | 17 files, clean | No version bump, no tag, no push. --- ## 7. Honesty limits - **Four questions across two bundles is a demonstration, not a hit rate.** One run each; nothing was repeated to get a better answer. The one repeat is named, and it measures a changed artefact. - The end-to-end runs allowed the model a tool set (`Bash`, `Read`, `Skill`, `Glob`, `Grep`). Without permission to run the pre-pass the skill cannot do its first step, so this is not an optional flag; it is part of what was measured. - **No token figure anywhere.** There is no tokenizer on this machine, so every cost is in `BUDGET_UNIT` bytes. Quoting an o200k number would be repeating a published figure, not measuring one. - The generated skill's prose is thinner than the hand-instantiated copy for the corpus that was measured over days. It states every number it can measure and claims nothing it cannot; a corpus that needs more — an extension marking, a measured cost distribution over many questions — needs a hand-edited copy, and the generator refuses to overwrite one without `--force` for that reason. - A scratch project inherits the machine's global session-start hooks: run (i) ended by reporting an unrelated 17-message coordination inbox. Harmless here, worth knowing before anyone reads a generated answer as self-contained.