docs(front-page): the gate numbers the gate actually prints, and a breaking point that was measured
Four claims on the front page were false on this commit, and one of them was a number no division ever produced. **The retrieval gate.** README reported it RED on rows 3, 4, 5, 7, 8 and 9, with row 3 at 2 of 5 and row 4 at 3 of 6. Run on this commit it is RED on rows 5, 7, 8 and 9, with row 3 at 5 of 5 and row 4 at 6 of 6: `f81683e` made a withheld concept carry the rule that actually decided it, and `05cb190` gave the payload a `coverage` block, and neither updated the table. Row 8 is `0 of 3 | NOT RUN` on the default run and was published as `44 of 64 questions`, which is what it scores the day all three private sets are handed to it -- now labelled with the day and the machine rather than printed as a row. The same four figures were stale in `CLAUDE.md`. **The breaking point in a generated skill.** `int(LIMIT / per_withheld) if per_withheld else 0` printed `At roughly 0 concepts the bookkeeping alone reaches the 120000-byte limit` whenever the generation run withheld nothing -- the absence of a measurement, rendered as one, and read as a bundle that breaks before it holds anything. A run with no withheld entry has no slope to extrapolate from, so the sentence is withheld with its reason. The shipped `skills/okf-consume/SKILL.md` is generated with the question its `references/README.md` names, withholds nothing, and carried exactly that `0`; it is regenerated. Two arms in the test, because one would pass on an empty set: the bundles that withhold something must still state a positive figure. The sentence for that arm also stopped saying `**4 bytes** for 3 concepts` where the 4 bytes were the cost of 0 withheld entries. It is now `for N of M concepts`, which moves two generated skills' line counts and therefore the published comparison: 280 of 312 and 310 -> 281 of 313 and 311, re-measured, with the 62 differing lines unchanged. **Four tools.** A single-bundle server exposes three: `okf_list` is absent where there is nothing to list. README's table already said so in a cell; the heading and the CHANGELOG did not. **What `--accounting` accounts for.** The account is over the element classes each format's vocabulary names, verified against `accounting._READERS` rather than against the report: a file whose suffix has no reader is accounted at file level only, `.docx` reads `document.xml` and `footnotes.xml` (so headers, footers, endnotes and comments are outside), `.pptx` reads the slides (so speaker notes are outside), `.xlsx` reads the worksheets (so cell comments are outside and a cell contributes its cached value, never its formula), and `.rtf` skips its header and footer groups. A hidden slide or sheet IS counted -- it lives in the same part as a visible one. Nothing is built for this; the list is what `0 unaccounted` does not claim. Gates re-run on the commit: retrieval `GATE RED: rows 5, 7, 8, 9` (exit 1), MCP `GATE RED: rows 2` (exit 1), both matching what is now written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
bf697bfcad
commit
d300338e4d
9 changed files with 203 additions and 754 deletions
|
|
@ -227,14 +227,15 @@ one answer delivered all 3 concepts — so that answer was about
|
|||
**101.5 %** of the corpus. One question is one measurement: a
|
||||
different question moves `spent` and this figure with it.
|
||||
|
||||
**The breaking point, stated so it can be observed to have been passed.**
|
||||
The `withheld` list carries one entry per considered concept and grows
|
||||
linearly: here it is **4 bytes** for 3 concepts. At roughly
|
||||
**0 concepts** the bookkeeping alone reaches the 120000-byte
|
||||
limit, and although it is not counted against `spent`, a payload whose
|
||||
bookkeeping dwarfs its content has stopped being a cut. The pre-pass also
|
||||
reads every concept body on every run, so the same growth is a wall-clock
|
||||
cost with no precomputed index behind it.
|
||||
**The breaking point could not be measured on this bundle.** The
|
||||
`withheld` list carries one entry per considered concept, and on this
|
||||
bundle at generation time nothing was withheld: all 3 concepts
|
||||
were delivered. There is therefore no per-entry cost to extrapolate
|
||||
from, and no concept count is stated here — a bundle large enough to
|
||||
withhold something states one. What does hold either way: the
|
||||
bookkeeping is not counted against `spent`, and the pre-pass reads
|
||||
every concept body on every run, so growth is a wall-clock cost with
|
||||
no precomputed index behind it.
|
||||
|
||||
## Denominators
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue