Four claims on the front page were false on this commit, and one of them was a number no division ever produced. **The retrieval gate.** README reported it RED on rows 3, 4, 5, 7, 8 and 9, with row 3 at 2 of 5 and row 4 at 3 of 6. Run on this commit it is RED on rows 5, 7, 8 and 9, with row 3 at 5 of 5 and row 4 at 6 of 6: `f81683e` made a withheld concept carry the rule that actually decided it, and `05cb190` gave the payload a `coverage` block, and neither updated the table. Row 8 is `0 of 3 | NOT RUN` on the default run and was published as `44 of 64 questions`, which is what it scores the day all three private sets are handed to it -- now labelled with the day and the machine rather than printed as a row. The same four figures were stale in `CLAUDE.md`. **The breaking point in a generated skill.** `int(LIMIT / per_withheld) if per_withheld else 0` printed `At roughly 0 concepts the bookkeeping alone reaches the 120000-byte limit` whenever the generation run withheld nothing -- the absence of a measurement, rendered as one, and read as a bundle that breaks before it holds anything. A run with no withheld entry has no slope to extrapolate from, so the sentence is withheld with its reason. The shipped `skills/okf-consume/SKILL.md` is generated with the question its `references/README.md` names, withholds nothing, and carried exactly that `0`; it is regenerated. Two arms in the test, because one would pass on an empty set: the bundles that withhold something must still state a positive figure. The sentence for that arm also stopped saying `**4 bytes** for 3 concepts` where the 4 bytes were the cost of 0 withheld entries. It is now `for N of M concepts`, which moves two generated skills' line counts and therefore the published comparison: 280 of 312 and 310 -> 281 of 313 and 311, re-measured, with the 62 differing lines unchanged. **Four tools.** A single-bundle server exposes three: `okf_list` is absent where there is nothing to list. README's table already said so in a cell; the heading and the CHANGELOG did not. **What `--accounting` accounts for.** The account is over the element classes each format's vocabulary names, verified against `accounting._READERS` rather than against the report: a file whose suffix has no reader is accounted at file level only, `.docx` reads `document.xml` and `footnotes.xml` (so headers, footers, endnotes and comments are outside), `.pptx` reads the slides (so speaker notes are outside), `.xlsx` reads the worksheets (so cell comments are outside and a cell contributes its cached value, never its formula), and `.rtf` skips its header and footer groups. A hidden slide or sheet IS counted -- it lives in the same part as a visible one. Nothing is built for this; the list is what `0 unaccounted` does not claim. Gates re-run on the commit: retrieval `GATE RED: rows 5, 7, 8, 9` (exit 1), MCP `GATE RED: rows 2` (exit 1), both matching what is now written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
9.7 KiB
An MCP surface over OKF bundles, in two shapes
2026-09-20. Capability loop:
the eval was written RED at 5f1772e, before any server existed; the capability
follows in its own commit.
The operator's question was not "does MCP work". It was: one server per bundle or one server for many, and must these artefacts be made again every time a bundle is rebuilt or a new one appears? This round builds the three artefacts that question compares, and measures the answer.
What was measured, and against what
tools/okf_mcp_gate.py, six rows, one exit code. The server is started as a
subprocess and spoken to over newline-delimited JSON-RPC beginning at
initialize -- never imported. A client built from the server's own framing
helpers would agree with the server by construction, so the client is written
separately in the gate.
Denominators are pinned in the gate and recounted a second time in the tests: 7 required tools across the two shapes, 4 artefact classes, 3 bundles times 3 discovery checks, 3 cross-bundle checks, 6 hostile cases. A row that counted what the server happened to offer would go green by offering less.
| row | what it asks | today |
|---|---|---|
| 1 | every required tool answers over real stdio, carrying bundle id and concept id | 7 of 7 |
| 2 | every anchor the frozen graded set points at, fetched verbatim | 83 of 181 |
| 3 | one concept changes: does the stale artefact refuse, or answer quietly | 4 of 4 |
| 4 | three unknown bundles appear while the server runs | 9 of 9 |
| 5 | one documented sequence, two bundles, both sources | 3 of 3 |
| 6 | traversal, symlink, broken manifest, 10 MB concept, unknown id | 6 of 6 |
GATE RED: rows 2, exit 1.
Reproduce:
uv run python tools/okf_mcp_gate.py \
--sett <the frozen set>/sporsmal.json \
--frys <the frozen set>/frys.json \
--bundle-root <a directory holding its bundles>
Without the last three flags row 2 is 0 of 0 with the reason stated: the set
names a consumer's documents, this repository is public, and a gold set is an
input here and never a constant.
Row 3 is the operator's question, and the answer has four rows
The drill: copy a bundle, start the artefact, change one concept, ask again.
| artefact | stale answer | artefacts to remake | manual steps |
|---|---|---|---|
| one server in front of one bundle | refuses / cannot go stale | 0 | 0 |
| one server in front of many | refuses / cannot go stale | 0 | 0 |
| today's generated skill (per bundle) | refuses out loud (bundle_mismatch) |
1 | 1, per consuming project |
| the generic skill (one for all) | cannot go stale | 0 | 0 |
Neither MCP shape needs an update when a bundle is rebuilt, and neither needs one when a bundle is added. That is not luck: nothing is cached across calls. Every call re-walks the roots and recomputes the bundle's content identity, so the identity in an answer is a fact about the bytes at the moment of the call. The cost is real and is paid per call -- see the limits below.
Row 3 was 1 of 4 before any capability existed, which the order did not
predict and is worth stating: today's per-bundle skill already refuses out loud
when its bundle moves, because okf check's bundle_mismatch rule compares the
declared ref against the payload's. The skill's cost is not silence. It is that
one artefact has to be regenerated and reinstalled wherever it was installed,
and that number is not measurable from inside this machine.
The generic skill, measured rather than assumed
The order cited 227 of 285 lines identical between two generated skills,
measured 2026-09-18. Measured again here, on two different bundles
(examples/ingest-golden-segmented-okf-v0-2 and tests/fixtures/consume-bundle):
281 of 313 and 311 lines identical, 62 lines differing
(re-measured 2026-09-20 after the breaking-point sentence was repaired; it was
280 of 312 and 310, with the same 62). Neither number
contradicts the other -- they are different pairs of bundles -- and the shape of
the finding is the same: what differs is identity, concept count, the
conditional-field table, the whole-bundle cost and the breaking point.
skill.render_generic() carries none of them. The property that makes that
claim checkable rather than asserted is that the function takes no argument:
there is no bundle it could have read, and two calls return the same bytes. A
test controls it against a per-bundle skill, which must carry exactly what the
generic one does not -- without that control, an assertion about an absence
passes on an empty string.
The per-bundle half is okf card <bundle>, derived on every run and never
written into the bundle. The order proposed storing it there. Writing a card
file into every bundle would move the bytes of all six examples/*/expected-bundle
trees (23 files compared byte-for-byte) and of the pinned reference bundle, to
store something recomputable in under a second -- and a stored card is one more
artefact that can disagree with the bytes beside it, which is the defect the
generic skill exists to remove. Chosen as derived because it answers the
maintenance question more completely, not less.
Row 2 decomposed: the bundle, the ranker, and the vocabulary
83 of 181 (bundle, anchor) pairs, M = 181 counted from the set at run time.
The order's own figure of 197 is the set's atom count under a different
definition; 181 is what the pair rule below yields on the file as frozen at
version 4.
Three numbers, and the middle one is the finding:
- 99 of 181 pairs are present in the bundles at all. 82 are not: the text
the set quotes is not in the bundle, which is red for the BUNDLE and not for
the server.
r761-2025is the sharpest case at 17 of 33 present. - 83 of the 99 present were reached, so the surface reaches 83.8 % of what
is there.
r761-2025is again the outlier: 2 reached of 17 present. - 0 of 83 were met by
okf_fetchon the anchor as a concept id. The set's anchors (Krav 2.3.1—3) and this library's concept ids are different vocabularies, so the cheap route -- a true ceiling -- never fires, and every pair met was met throughokf_ask, which runs the ranker. That makes 83 a FLOOR on the ceiling, never the ceiling. A surface offering a lookup by the publisher's own anchor would separate the two, and does not exist today.
Quote comparison folds exactly two things and nothing else: U+00AD, because
okf build strips soft hyphens from extracted text while the publisher's JSON
keeps them, and whitespace runs, because a quote cut out of a paragraph carries
the line breaks of wherever it was cut. Case is not folded.
Hostile input, and why a code set rather than "was refused"
Row 6 declares, per case, the refusal CODES that count as the right refusal.
The first run of this gate had the 10 MB concept refused as concept_unknown --
the fixture had written the file without naming it in the index, so the size
ceiling never ran and the row was green for a reason unrelated to the attack.
Two checks giving the same verdict are not the same guarantee.
Containment is two independent checks: the bundle's own index must name the
concept, AND the resolved path must be inside the bundle. A mutant removing the
first one survives, and the mechanism is printed: the traversal is then
refused by the second, as path_escape instead of concept_unknown. A mutant
removing both is killed. That survival is the redundancy working and is reported
as such rather than as a kill.
A note added 2026-09-20, after this round: that sentence was true of
okf_fetch and of no other tool. okf_ask and okf_describe made only the
first of the two checks -- the index rule, which is a string rule and cannot
see a symlink -- and read whatever the joined path pointed at. The second check
now lives in consume.resolve_in_bundle and every reader here goes through it;
the tests are tests/test_read_path_containment.py, red on 8 of 11 rows before
the repair with okf_fetch's two rows green as the control.
Mutants
13 mutants, applied in a scratch copy of the tree and never in the working tree, with an unmutated control first: 12 killed, 1 survived with a mechanism, 0 errors. The control's gate rows and pytest targets are green before the first mutation, so a kill cannot be the call having failed.
Killed: a cached bundle identity (row 3), two bundles known by name in the many-shape (row 4), a fetched concept without its concept id (row 1), both containment checks removed (row 6), discovery run once at startup (row 4), row 2's denominator taken from the run (test), a symlink descended (test), the size ceiling removed (row 6), the generic skill naming a bundle (test), a broken manifest skipped silently (row 6), a listing tool on the one-shape (test), and an unknown bundle answered instead of refused (row 6).
Limits, stated rather than implied
- Nothing is cached, and it costs. On the 2 756-concept bundle the content
identity is a 0.75 s hash of the whole concept tree and one
okf_askis 5.6 s. Row 2's full run over four bundles and 181 pairs took 4 min 13 s. A cache would have to be keyed on something cheaper than the hash and still correct; no such key is shipped, and the cost is the price of the row-3 result above. - The gate measures a ceiling and a maintenance cost. Whether an arm answers
WELL is a different question, asked by
tools/okf_retrieval_gate.py. No arm was run here and no model was called. - The architecture choice is the operator's. These rows are its input.
- Row 3 counts artefacts and steps inside this machine. A project that has installed a generated skill pays one more step per project, and that number is not measurable from here.
- No MCP server was registered in any
settings.jsonor.mcp.json.