docs(front-page): the gate numbers the gate actually prints, and a breaking point that was measured
Four claims on the front page were false on this commit, and one of them was a number no division ever produced. **The retrieval gate.** README reported it RED on rows 3, 4, 5, 7, 8 and 9, with row 3 at 2 of 5 and row 4 at 3 of 6. Run on this commit it is RED on rows 5, 7, 8 and 9, with row 3 at 5 of 5 and row 4 at 6 of 6: `f81683e` made a withheld concept carry the rule that actually decided it, and `05cb190` gave the payload a `coverage` block, and neither updated the table. Row 8 is `0 of 3 | NOT RUN` on the default run and was published as `44 of 64 questions`, which is what it scores the day all three private sets are handed to it -- now labelled with the day and the machine rather than printed as a row. The same four figures were stale in `CLAUDE.md`. **The breaking point in a generated skill.** `int(LIMIT / per_withheld) if per_withheld else 0` printed `At roughly 0 concepts the bookkeeping alone reaches the 120000-byte limit` whenever the generation run withheld nothing -- the absence of a measurement, rendered as one, and read as a bundle that breaks before it holds anything. A run with no withheld entry has no slope to extrapolate from, so the sentence is withheld with its reason. The shipped `skills/okf-consume/SKILL.md` is generated with the question its `references/README.md` names, withholds nothing, and carried exactly that `0`; it is regenerated. Two arms in the test, because one would pass on an empty set: the bundles that withhold something must still state a positive figure. The sentence for that arm also stopped saying `**4 bytes** for 3 concepts` where the 4 bytes were the cost of 0 withheld entries. It is now `for N of M concepts`, which moves two generated skills' line counts and therefore the published comparison: 280 of 312 and 310 -> 281 of 313 and 311, re-measured, with the 62 differing lines unchanged. **Four tools.** A single-bundle server exposes three: `okf_list` is absent where there is nothing to list. README's table already said so in a cell; the heading and the CHANGELOG did not. **What `--accounting` accounts for.** The account is over the element classes each format's vocabulary names, verified against `accounting._READERS` rather than against the report: a file whose suffix has no reader is accounted at file level only, `.docx` reads `document.xml` and `footnotes.xml` (so headers, footers, endnotes and comments are outside), `.pptx` reads the slides (so speaker notes are outside), `.xlsx` reads the worksheets (so cell comments are outside and a cell contributes its cached value, never its formula), and `.rtf` skips its header and footer groups. A hidden slide or sheet IS counted -- it lives in the same part as a visible one. Nothing is built for this; the list is what `0 unaccounted` does not claim. Gates re-run on the commit: retrieval `GATE RED: rows 5, 7, 8, 9` (exit 1), MCP `GATE RED: rows 2` (exit 1), both matching what is now written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
bf697bfcad
commit
d300338e4d
9 changed files with 203 additions and 754 deletions
|
|
@ -1,6 +1,6 @@
|
|||
# An MCP surface over OKF bundles, in two shapes
|
||||
|
||||
2026-09-20. Order `20260918T163400Z-6303812376-from-.claude`. Capability loop:
|
||||
2026-09-20. Capability loop:
|
||||
the eval was written RED at `5f1772e`, before any server existed; the capability
|
||||
follows in its own commit.
|
||||
|
||||
|
|
@ -75,7 +75,9 @@ and that number is not measurable from inside this machine.
|
|||
The order cited 227 of 285 lines identical between two generated skills,
|
||||
measured 2026-09-18. Measured again here, on two different bundles
|
||||
(`examples/ingest-golden-segmented-okf-v0-2` and `tests/fixtures/consume-bundle`):
|
||||
**280 of 312 and 310 lines identical, 62 lines differing**. Neither number
|
||||
**281 of 313 and 311 lines identical, 62 lines differing**
|
||||
(re-measured 2026-09-20 after the breaking-point sentence was repaired; it was
|
||||
280 of 312 and 310, with the same 62). Neither number
|
||||
contradicts the other -- they are different pairs of bundles -- and the shape of
|
||||
the finding is the same: what differs is identity, concept count, the
|
||||
conditional-field table, the whole-bundle cost and the breaking point.
|
||||
|
|
@ -137,6 +139,14 @@ refused by the second, as `path_escape` instead of `concept_unknown`. A mutant
|
|||
removing both is killed. That survival is the redundancy working and is reported
|
||||
as such rather than as a kill.
|
||||
|
||||
**A note added 2026-09-20, after this round:** that sentence was true of
|
||||
`okf_fetch` and of no other tool. `okf_ask` and `okf_describe` made only the
|
||||
first of the two checks -- the index rule, which is a string rule and cannot
|
||||
see a symlink -- and read whatever the joined path pointed at. The second check
|
||||
now lives in `consume.resolve_in_bundle` and every reader here goes through it;
|
||||
the tests are `tests/test_read_path_containment.py`, red on 8 of 11 rows before
|
||||
the repair with `okf_fetch`'s two rows green as the control.
|
||||
|
||||
## Mutants
|
||||
|
||||
13 mutants, applied in a scratch copy of the tree and never in the working tree,
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue