docs: remove what this repository published about a consumer's corpus

Operator decision 2026-09-21: nothing from that consumer's collection goes out
on the public remote. The NAME stays where it is already published -- it is a
consumer of this library, named as such, and removing it would mean rewriting
published history, which this repository does not do. What goes is everything
that describes their CONTENT.

Removed across README, CLAUDE.md, CHANGELOG, four dated reports, the
consumption contract, three source modules and three test modules: their
corpus's document and page counts, the concept count of a bundle built from
it, the byte figures of a payload built from it, the question and fasit counts
and recorded score of their evaluation set, a bundle id with two content refs,
an order id naming them, and a path into their repository.

Kept, because the argument survives without the corpus: RATIOS and
percentages. A ratio is the finding -- a withheld list that is 65.5 % of a
payload is a defect at any corpus size -- and it discloses nothing about how
large anyone's collection is. Where a claim lost its denominator it now SAYS
so rather than quietly reading as unmeasured: the gate-refusal limitation in
the README states that the corpus and its counts are deliberately withheld and
points the reader at their own build, which is the number that binds them
anyway.

One integrity pin is kept and named here rather than left to be found: the
retrieval gate still pins that set by sha256, because the pin is what refuses
a self-written file in the right shape, and a checksum discloses nothing about
what it checksums. Its recorded SCORE is gone -- that was their figure about
their own corpus, and the row now says so instead of restating it.

The known-positive constants move with the contract document, as they must.
Suite green, 2372 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-21 04:28:15 +02:00
commit cf21449ddb
15 changed files with 108 additions and 97 deletions

View file

@ -717,9 +717,9 @@ KNOWN_POSITIVE_CASE = "docs/consumption-contract.md, encoded as a JSON string"
#: `measure()`'s own answer for that file. Vacuous ALONE -- which is why the
#: delta below exists.
KNOWN_POSITIVE_EXPECTED = 19_839
KNOWN_POSITIVE_EXPECTED = 19_837
#: The second, independent route. `wc -c` reports 19 360 raw bytes for the same
#: The second, independent route. `wc -c` reports 19 358 raw bytes for the same
#: file; the difference is this file's JSON quoting and escaping overhead. A
#: reader can derive it without running `measure()` at all, and it moves the
#: moment `measure()` changes what it counts -- which is what stops
@ -2133,14 +2133,15 @@ CONTRACT_REVISION = "okf-consumption/2"
#: denominator is for; the near misses are what a reader can act on.
#:
#: **20, and the number is read off a measurement rather than chosen.**
#: Measured 2026-09-20 on a 2313-concept bundle of one project's own
#: documentation: the flat list held 2 305 entries = 186 440 B of compact JSON
#: = **65.5 % of the 284 850-byte file**, and not one of those bytes counted
#: against the budget the same payload reports (`spent` was 45 192). So a
#: reader was handed 239 658 bytes the budget line did not know about, in
#: order to learn 2 305 concept ids with nothing beside them -- the field
#: `--withheld-titles` existed to buy, and which was off because buying it for
#: 2 305 entries cost another 37.9 %. At twenty entries the title and the
#: Measured 2026-09-20 on a large real bundle: the flat list carried one entry
#: per withheld concept and came to **65.5 % of the written payload**, and not
#: one of those bytes counted against the budget the same payload reports. So
#: a reader was handed most of a file the budget line did not know about, in
#: order to learn one concept id per withheld concept with nothing beside it
#: -- the field `--withheld-titles` existed to buy, and which was off because
#: buying it for a list that long cost another 37.9 %. The corpus belongs to a
#: consumer and its counts are not restated here; the ratio is the argument
#: and discloses nothing about its size. At twenty entries the title and the
#: document are free, and the list becomes the one thing it never was: a set
#: of names a reader can ask for. `k` is 8, so twenty is the cut plus the next
#: twelve; the whole list stays reachable behind one switch.
@ -2324,7 +2325,7 @@ def build_payload(
`--withheld-titles` is retired by this change rather than kept beside it.
It existed to buy the one field the near misses now carry by default, and
it was off because buying that field for 2 305 entries cost another 37.9 %.
it was off because buying that field for one entry per withheld concept cost another 37.9 %.
A flag whose only remaining effect would be to STRIP the title from a list
the caller explicitly asked for in full names no decision worth two shapes
for one list.
@ -2663,8 +2664,8 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
action="store_true",
help=(
"name EVERY withheld concept instead of the nearest N. Measured "
"2026-09-20 on a 2313-concept bundle: the whole list is 186 440 B "
"= 65.5 %% of the written file, and none of it counts against the "
"2026-09-20 on a large real bundle: the whole list is 65.5 %% of "
"the written payload, and none of it counts against the "
"budget the payload reports"
),
)

View file

@ -110,7 +110,7 @@ REQUIRED_SECTIONS = (
# BOOKKEEPING, and a skill could carry all seven while saying nothing
# about how to read a question, whether to search twice, or what the
# answer should look like -- which is the document the operator measured
# as unusable on a 2313-concept bundle. The rule follows the template, not
# as unusable on a large real bundle. The rule follows the template, not
# the other way round: these two are required because the template now
# carries them, and a skill without them is thin in the way that mattered.
"Working method",
@ -141,8 +141,8 @@ class Report:
#: The withheld entries this report READ, which since `okf-consumption/2`
#: is the sample the payload names and not the whole withheld set. The
#: total is in the payload; this is the denominator of what was checked,
#: and conflating the two would let a report claim it examined 2 305
#: entries it never saw.
#: and conflating the two would let a report claim it examined entries it
#: never saw.
withheld_examined: int
#: What the payload says its withheld set holds. `None` when it states no
#: total -- unmeasured, never zero.

View file

@ -624,9 +624,8 @@ def _breaking_point(
Until `okf-consumption/2` this section extrapolated a concept count at
which the bookkeeping alone would fill the budget, because `withheld`
carried one entry per considered concept and grew linearly. Measured
2026-09-20 on a 2313-concept bundle, that growth had arrived: the list was
186 440 B = 65.5 % of the written file, none of it counted against
`spent`.
2026-09-20 on a large real bundle, that growth had arrived: the list came
to 65.5 % of the written file, none of it counted against `spent`.
It does not grow that way any more, so this section no longer states a
concept count -- a number extrapolated from a slope the code no longer has