Four questions, two bundles, one run each, in a scratch project outside this repository with a generated skill per bundle. All four passed, and zero numbers or identifiers appeared in any answer that were not in the delivered set or in the payload's own identities (62, 45 and 35 unique numeric tokens checked). The skill triggered WITHOUT being named in the prompt and selected the right one of two installed skills from the question alone, so no special invocation syntax is needed: the generated `description`, which carries the bundle id, the concept count and the ref, is enough to route on. One defect the runs found, and it was in the prose rather than the payload. The citation guidance listed the four locator keys this library writes, so on the 270-concept third-party bundle the model reported "no page locator, the address is at document level" while the excerpt in front of it carried `source_element_id` - that bundle's own locator, correctly delivered by the prefix rule. The guidance now tells the reader to cite whichever `source_*` keys are present. On the re-run the same question returned the element id. Two runs of one question, the second measuring a changed artefact and not retrying the first. One finding that is not a defect in this chain: the first attempt at a known-negative was not one. The bundle covers water and frost protection on 17 of its 270 concepts and the ranker put none of them in the cut. The consumer behaved exactly as the contract asks - refused, named its denominator, reported its own zero as unmeasured because `withheld` entries carry no titles, and did not go around the cut. Recorded as a retrieval miss rather than replaced, and it is the same shape as the open fusion finding. A correction to this session's own measurement is in the record too: a first sweep used `grep -rhoE "^source_[a-z_]+:"`, whose character class excludes digits, and so missed `source_sha256` on 270 of 270 concepts. A pattern that cannot match what it is looking for returns a zero that reads like a fact. README gains "Consume in Claude Code": folder to answer in three commands, every one of them run in this session. A test holds that the recipe invokes only scripts this repository ships, at the paths it names. Suite 1373 (1339 at the session baseline), ruff clean, mypy src clean. No version bump, no tag, no push. Co-Authored-By: Claude <claude-opus-5>
19 KiB
Connecting an arbitrary OKF bundle to Claude Code
Date: 2026-09-08. Order: 20260908T124346Z-4668330384 (C1).
Commits: 17c49fc (step 0), c95d189 (the prefix rule and the generator),
and the commit this document lands in.
Baseline: b6a8c8b, suite 1339.
The question this answers is the operator's, in their own words: "That Claude Code should be able to work with an arbitrary OKF bundle that we give Claude information about is extremely important." Three things had to hold — build a bundle from a folder, connect it to Claude Code, and ask it a question whose answer is a document or part of one. The first existed. The second existed only as one skill hand-instantiated for one corpus plus a template full of holes, installed nowhere, never run inside Claude Code. The third was therefore not testable at all.
0. What was measured and what was not
| measured | not measured | |
|---|---|---|
| the excerpt fields | yes, on three bundles and two fixtures | — |
the source_* pass-through rule |
yes, on the bundle that broke the allowlist | whether a fourth producer uses a key not starting with source_ |
| the form choice (generator vs generic skill) | yes, including a control that came out against the obvious gate | operator preference; only the mechanics are here |
| the generator's refusals | yes, four of them, each with its code | — |
| end to end in Claude Code | four questions, two bundles, one run each | a hit rate. Four questions is a demonstration |
| the ranking | unchanged, byte-for-byte, on the K2 control | — |
| cost per answer in tokens | — | no tokenizer on this machine; costs are in BUDGET_UNIT bytes |
1. Reproduction, first
The known-positive, before anything was changed:
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md \
--payload <a K2 payload>
# conformant: 15 rules over 8 excerpts and 621 withheld entries, 0 findings
# exit 0
Then a fresh bundle, built the way an operator would build one: nine documents (two spreadsheets, three word-processor files, four PDFs) copied out of the corpus tree into a scratch folder.
okf build /tmp/c1-fresh-src --bundle /tmp/c1-fresh-bundle \
--bundle-id c1-fresh-20260908 --okf-version 0.2 --ingested-at 2026-09-08T00:00:00Z
# merged + coded rejections = 9; N = 9 (substantive 9/9, rejected 0/9)
# 5.15 s, 0.572 s per file
54 concepts, ref sha256-tree:7245851d…9685f71. The index walk reaches 54; a
directory walk finds 55, and the one extra is log.md, which the walk excludes
by name. 54 = 54, controlled.
One question with a known gold concept:
python3 tools/okf_consume.py /tmp/c1-fresh-bundle \
--question "Hvordan skal prisene fylles ut?" --out /tmp/c1-fresh-payload.json
considered 54 = withheld 46 + delivered 8, spent 82 030 of 120 000, and the
gold concept — the priced spreadsheet — at rank 1. That is the position an
operator starts from.
2. Step 0: the excerpt carries the key the question is asked ON
Two independent measurements pointed at the same missing field.
portfolio-optimiser ran three paid arms: the gold concept was delivered at
rank 1 of 8 on 3 of 3 bundles, and the model answered correctly on 1 of
3, because a delivered excerpt carried concept_id, body text and nothing the
document is known by. The previous session measured it from the other side: the
provenance it had just written into every concept did not reach the payload at
all — nine members, not one of them provenance.
excerpt_for now carries title unconditionally, and req_number, the SPEC
§ 5.1 address sources, and the locator keys when the producer wrote them. Four
design decisions here were made by measurement rather than by taste.
(a) sources is read in both YAML forms. The two real bundles disagree
completely:
flow sources: [{...}] |
block sources: + indented |
a locator key | |
|---|---|---|---|
| the 629-concept procurement bundle | 629 of 629 | 0 | 618 source_offset, 586 source_pages |
| the 270-concept normative bundle | 0 | 270 of 270 | 0 of the five this library writes |
parse_frontmatter skips indented lines on purpose — a nested title: arriving
later would substitute for the document's — so the block form arrives as '',
present and unreadable. A flow-only reader delivers the second bundle with no
address at all. Reading the block form is not a licence to write it; the
emission rule is untouched.
(b) An undecodable address is named, not dropped. sources_unreadable: true
is a third state beside "an address" and "no address". Collapsing it into either
reports something nobody measured.
(c) An absent key stays absent. Writing req_number: "" would assert that
the producer wrote an empty identifier. Contract § 6.4 forbids exactly that
inference, from the other direction.
(d) The pass-through is a PREFIX, and this one came from a peer's
measurement. The first implementation used the five locator keys this library
writes. The PM measured it against the normative bundle and found the excerpt
still lost that bundle's locator: 269 of its 274 files carry
source_element_id, a key that repository chose under this chain's own rule
("the key says what it indexes") and that this library never writes. An
allowlist is a list of the producers its author thought of. The rule is now
every top-level key beginning with source_, and a prefix rather than a
substring — resource_owner contains the literal and is not a locator.
A correction to my own measurement, stated because it is the reason the peer
had to send it twice: my first sweep used grep -rhoE "^source_[a-z_]+:",
whose character class excludes digits, so it reported source_element_id and
silently missed source_sha256 on 270 of 270. The generator's own counting —
over parsed frontmatter, not a regex — found both. A pattern that cannot match
what it is looking for returns a zero that reads like a fact.
The K2 control: the field moved, the ranking did not
Same bundle, same question, same k. The "before" run is a copy of the tool
taken from git archive b6a8c8b, run with its own docs/ beside it, so the
editable install cannot leak the new code into the baseline.
before (b6a8c8b) |
after | |
|---|---|---|
| delivered concept ids, in order | 8 | identical |
text_sha256 per excerpt |
— | identical |
withheld list |
621 entries | identical |
| denominators | 629 = 621 + 8 | identical |
| payload bytes | 108 877 | 113 143 (+3.92 %) |
spent (BUDGET_UNIT) |
18 606 | 22 210 (+450.5 B per excerpt) |
| excerpt members | 9 | 17 |
| changed lines in the whole payload | — | 99 |
| payload sha256 | afb7e811… |
c94919c4… |
It is the field, not the ranking. Nothing about which concepts were chosen, or in which order, moved by one byte.
The contract document's own bytes moved twice while § 8 was being written, so the budget instrument's known-positive moved with them — by design, since a stale known-positive is meant to be a loud failure: 10 349 → 12 049 → 12 563 measured, 10 060 → 11 719 → 12 227 raw, delta 289 → 330 → 336.
Contract § 8 gains title as a MUST with its own checker code
(excerpt_unnamed, so the checker now runs 15 rules, was 14), and
req_number, sources and the locators as SHOULD — they are conditional on the
producer, and a bundle whose concepts carry no identifier cannot deliver one.
3. The form: a generator per bundle, chosen against a control that failed
Two candidates. (a) A generator that instantiates the template for one bundle. (b) One generic skill taking the bundle root as a runtime argument.
The obvious gate does not work. The contract checker was asked to tell them apart, and cannot:
python3 tools/okf_contract_check.py --skill skills/okf-consume-template/SKILL.md \
--payload /tmp/c1-fresh-payload.json
# conformant: 15 rules over 8 excerpts and 46 withheld entries, 0 findings — exit 0
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md \
--payload /tmp/c1-fresh-payload.json
# conformant: 15 rules over 8 excerpts and 46 withheld entries, 0 findings — exit 0
The unfilled template passes against a real payload, and a skill built for a different corpus passes against this one's. So conformance could not decide the form, and any claim that "(b) fails the checker" would have been false. The checker reads section headings, the marking vocabulary and the payload; it does not read whether a number in the prose is true of this bundle.
What decides it is what the skill has to state. § 5's denominators, § 6.4's conditional-field list and § 7.6's breaking point are per-bundle numbers. A generic skill can only leave them as holes — the template's own definition of unfinished — or carry another corpus's numbers, which is worse than a gap, because a stated cost that is false for this bundle is a measurement failure. Instantiating is what makes them true.
Steps from "here is a bundle" to the first correct, marked answer:
| (a) generator | (b) generic skill | |
|---|---|---|
| one-time setup | 1 command per bundle | 1 copy, once |
| typed at question time | the question | the question and the bundle path, every time |
| with two bundles installed | the model selects on the skill name — measured, it did | nothing selects; the caller must say which |
| per-bundle denominators | measured into the file | absent or false |
(a), and the deciding column is the third: the operator's requirement is "one or more bundles", and (b) has no mechanism for the "more".
tools/okf_skill.py therefore instantiates. It measures, per bundle: id, ref,
concept count, the conditional-field table with a denominator for every field
(the source_* rows discovered from the bundle rather than listed), the
whole-bundle cost by the gate's own instrument, the share one measured answer
spent, and the concept count at which the withheld bookkeeping alone reaches
the limit. It also runs the index-walk-against-directory control once, at
generation time, never on the question path — § 9.2 binds a consumer reaching
for context at query time, and a build-time control is what turns "the walk
loses nothing" into a number the skill can quote.
It lives in tools/ for the reason okf_consume.py states for itself — outside
src/, so no consumer's install surface changes — and for one more: a
wheel-installed okf skill would emit a command pointing at
tools/okf_consume.py, which the wheel does not contain.
Red tests first, and every gate the checker lacks is one of them: no
placeholder survives (with the pattern shown able to find, against the template);
the skill names its own bundle's id and ref and not the other bundle's; its
commands are absolute and point at files that exist; it reports a denominator per
conditional field; it is byte-deterministic for a given bundle and destination.
Four refusals, each with its code: a directory with no index
(bundle_unreadable), an index with no bundle_id (bundle_id_missing), a
bundle with no concepts (bundle_empty), and an occupied target without
--force (target_occupied). No model call, no new dependency.
4. End to end in Claude Code
Two skills installed side by side in a scratch project outside this repository, one per bundle. Each row is one run; none was repeated to get a better one.
python3 tools/okf_skill.py <bundle> --out /tmp/c1-scratch/.claude/skills/<id>-consume
cd /tmp/c1-scratch && claude -p "<the question>"
| # | bundle | question | skill named in the prompt? | result |
|---|---|---|---|---|
| i | fresh, 54 concepts | "Hvordan skal prisene fylles ut?" | yes | pass |
| ii | fresh, 54 concepts | a topic measured absent from the bundle | yes | pass — refused |
| iii | third-party, 270 concepts | "Hva krever Krav 10.2—2 i N500? Gjengi det sentrale vilkåret." | no | pass |
| iv | fresh, 54 concepts | "Hvordan skal prisene fylles ut?" | no | pass — and selected the right one of two |
(i) The known-positive. Rank 1 of 8 delivered. The answer opened by naming
the document and the sheet, and cited the locator the previous session put into
the concept: "kilde <the spreadsheet>, ark Prissammenstilling, rader 1 til
100". It marked with all five required literals, reported
considered 54 / withheld 46 / delivered 8, named below_k as the rule that
dropped a concept it flagged as possibly relevant, reproduced the budget
known-positive (12 563 = 12 563), and ran the contract check itself: 15 rules,
0 findings.
(ii) The known-negative. The absence was measured before the question was
asked, with the query shown able to find first (pris matches 15 files;
jernbane|signalanlegg matches 0 of 54). The answer: "dokumentene i bundelen
stiller ingen krav til jernbanesignalanlegg", marked
[sourced-not-sufficient], with its own known-positive control for the regex it
used to check the delivered texts, and with the 46 withheld concepts named as
[unread] rather than as absent. It named external standards only under
[unverifiable-from-bundle].
(iii) The third-party bundle. Measured on the tree on disk at the time:
vegnormal-n500-2024, 270 concepts, ref
sha256-tree:673a0c2c…d6b5b3ad — not the V1 tree the order named, because
that repository landed a rebuild during the day. The gold concept came back at
rank 1, and the answer quoted the requirement verbatim against the concept body,
named the title and the req_number, the sources resource, the concept id and
the excerpt digest, and stated adjudication: unknown and
trust_tier: unverified as measured absences at the producer rather than as
judgements about the source.
One defect this run found, and it was in the prose. The first (iii) run said
"the producer wrote no page locator (no source_pages), so the address is at
document level" — while the excerpt in front of it carried source_element_id,
that bundle's own locator. The excerpt was right and the skill's citation
guidance was wrong: it listed the four keys this library writes instead of
telling the reader to cite whichever source_* keys are present. Fixed in the
template and in both instantiated copies; on the re-run the same question
returned "element id-4ca67e2f-…" in the source pointer. That is two runs of
one question, and the second is a measurement of a changed artefact, not a
retry of the first.
(iv) The trigger control. No skill named in the prompt, two skills installed.
It selected c1-fresh-20260908-consume on the question alone, ran the pre-pass
and the checker, and cited source_sheet and source_rows by name. So no
special invocation syntax is needed; the generated description — which names
the bundle id, the concept count and the ref — is enough to route on.
Invented numbers: zero
Every numeric or identifier-shaped token in each answer was extracted and checked against a denominator stated per run.
| run | unique numeric/id tokens | not found in the denominator |
|---|---|---|
| i | 62 | 0 |
| ii | 45 | 0 |
| iii | 35 | 0 |
For (iii) the denominator is the delivered payload itself, reproduced exactly —
the model's reported spent of 11 941 equals the payload rebuilt here. For (i)
and (ii) the model composed its own question string, so its payload is not
byte-reproducible; the denominator there is the whole bundle text plus the
payload's own identities (the bundle ref, every concept-file digest, every
delivered-text digest) plus the figures the skill itself supplies. A first
pass that omitted those identities reported eight and nine "missing" tokens,
every one of them a fragment of a correctly quoted digest — the wrong
denominator, corrected here rather than published.
5. A finding about the ranking, not about the skill
The first attempt at (ii) asked the normative bundle about water and frost
protection in a subsea tunnel. The model answered
[sourced-not-sufficient], said the relevant requirements "if they exist, are
among these 262 withheld", and refused to go around the cut. Checked
afterwards, the bundle does cover it: 17 files match frostsikring, 10 match
"vann- og frostsikring", 13 match undersjøisk, and none of them was ranked
into the cut.
So it was never a known-negative — it was a retrieval miss, and it is
recorded as one rather than quietly replaced. Two things follow, and they point
in opposite directions. The consumer behaved exactly as the contract asks: it
did not fabricate, it named its denominator, and it reported its own zero as
unmeasured because the withheld entries carry only ids and rules, no titles.
The ranker did not: a compound Norwegian question ("vann- og frostsikring",
"undersjøisk") was out-ranked by concepts sharing only "tunnel" and "vann". That
belongs with the fusion work already open, and it is the same shape as the
MIN_SHARED_PREFIX finding.
A second, smaller observation: withheld entries carry concept_id and rule
and no title, so a reader who wants to know what was dropped cannot tell
without reading the bundle — which § 2.2 forbids. The consumer in run (ii)
identified this itself and reported it as unmeasured. Not fixed here; noted.
6. Suite
before (b6a8c8b) |
after | |
|---|---|---|
pytest -q |
1339 | 1373 |
ruff check src tests tools |
clean | clean |
ruff format --check . |
clean | clean |
mypy src |
17 files, clean | 17 files, clean |
No version bump, no tag, no push.
7. Honesty limits
- Four questions across two bundles is a demonstration, not a hit rate. One run each; nothing was repeated to get a better answer. The one repeat is named, and it measures a changed artefact.
- The end-to-end runs allowed the model a tool set (
Bash,Read,Skill,Glob,Grep). Without permission to run the pre-pass the skill cannot do its first step, so this is not an optional flag; it is part of what was measured. - No token figure anywhere. There is no tokenizer on this machine, so every
cost is in
BUDGET_UNITbytes. Quoting an o200k number would be repeating a published figure, not measuring one. - The generated skill's prose is thinner than the hand-instantiated copy for the
corpus that was measured over days. It states every number it can measure and
claims nothing it cannot; a corpus that needs more — an extension marking, a
measured cost distribution over many questions — needs a hand-edited copy, and
the generator refuses to overwrite one without
--forcefor that reason. - A scratch project inherits the machine's global session-start hooks: run (i) ended by reporting an unrelated 17-message coordination inbox. Harmless here, worth knowing before anyone reads a generated answer as self-contained.