feat(consume): one source document took 8 of 8 delivered places, so cap it
Measured outside this repository on a 3206-concept bundle of a published
handbook: the code's own process overview contributes 28 of 3206 concepts
(0.87 %) and 117 488 of 1 469 225 source characters (8.0 %), and took 8 of 8
delivered places on one question and 7 of 8 on the known-positive -- which was
not delivered at all. Identical at 343 and 1651 concepts, so the cause is the
corpus's COMPOSITION, that it holds its own table of contents, and NOT its size.
Splitting the corpus would move the defect, not remove it: any corpus with a
contents list, a project overview or a summary document has the same property.
`--source-quota N` caps how many DELIVERED places one source document may take.
It cuts where the shortlist is cut -- before the pack, never inside the DP,
which maximises a sum over a set it is handed -- so the freed place goes to the
next candidate and `k` is still delivered in full.
DEFAULT 2 SINCE TODAY, and it is the third change here that alters a payload
with NO bundle changing (after `--tie-shared-rank` and `--stem-prefix`).
Opt-out `--no-source-quota` reproduces the previous excerpt order.
Swept over {2, 3, 4, off} on three bundles, with the fasit prefixes validated
against the bundle FIRST (that control caught a defect in the measuring query
itself -- it read the last id segment where the document is the first):
- K2, both bundles: at 2 and 3, hit@8 goes 5 of 6 to 6 of 6 with all five
standing rank-1 rows unmoved. The recovered row had missed on every bundle and
every configuration measured until now. At 4 and off it is 5 of 6.
- The handbook bundle: hit@8 2 of 6 -> 4 of 6, the known-positive from not
delivered to rank 4, and the dominant document's share of delivered places
8 of 8 -> 2 of 8 (7 of 8 -> 2 of 8 on the known-positive).
- 2 rather than 3 on rank alone: the recovered rows come in at 5 and 4 rather
than 7 and 5.
WHAT THE GAIN IS NOT. hit@8 asks whether the gold DOCUMENT appears among the
delivered excerpts, and a document quota directly raises how many distinct
documents a payload holds, so that metric is not neutral with respect to this
rule. The five rows that were already rank 1 are neutral, and they did not move.
THE ADVERSE CASE IS NAMED, not left to a consumer. A bundle built from ONE
document carries the same `source_file` on every concept, so a quota applied
literally would deliver 2 excerpts where `k` were asked for -- a rule against
dominance turned into a rule against small bundles. The shortlist is topped back
up from the best-ranked over-quota candidates, which makes such a bundle
byte-identical to the quota being off, and a test holds it.
`--rarity-weight` was measured against the same defect and does NOT repair it:
it leaves the dominant document at 8 of 8 places on the question it floods,
delivers neither that answer nor the known-positive, and holds 5 of 6 on both K2
bundles. Combined with the quota it is worse than the quota alone (the
known-positive falls back out). It stays off.
The vocabulary stays CLOSED and the new code is published in all three places a
consumer can read it: `WITHHOLDING_RULES` (six -> seven),
`docs/consumption-contract.md` 5.3, and the generated SKILL.md -- verified by
reading the generated file, not the code that writes it. `source_quota_exceeded`
is a DIVERSITY drop and not a relevance one, so folding it into
`no_lexical_match` would tell a consumer the question reached nothing in a
concept the question in fact reached. `okf check --skill --payload` stays
conformant, 0 findings over 15 rules.
Editing the contract moved the 7.4 known-positive, which is the coupling
working as intended: 12 563 -> 13 238 encoded, 12 227 -> 12 893 raw, delta
336 -> 345, updated in the constant, the instantiated skill and the shipped
example payload.
Also adds the O6 guard on the reading side: `build_payload`'s signature defaults
are asserted equal to `okf consume`'s argparse defaults for every same-named
parameter. `okf project` shipped that exact disagreement for two rounds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
732f84df6e
commit
1e9f38b125
9 changed files with 380 additions and 26 deletions
57
CLAUDE.md
57
CLAUDE.md
|
|
@ -22,7 +22,29 @@ one boundary rule:
|
|||
the door's own output reachable as its own input (operator 2026-09-06; the
|
||||
flat listing was not a boundary, it was an absence with no denominator). All
|
||||
file-type→text extraction lives HERE (the guard is text-only). v1 core:
|
||||
`md`, `txt`, `csv`, `json`, `html` (stdlib). `pdf`/`docx`/`xlsx` only via
|
||||
`md`, `txt`, `csv`, `json`, `html` (stdlib). **`html` got a measured
|
||||
segmentation row 2026-09-10.** Until then `_HTMLTextExtractor.text()` was
|
||||
`" ".join("".join(parts).split())`, and `str.split()` with no argument splits
|
||||
on newlines too, so extraction of ANY HTML file returned unconditionally ONE
|
||||
line while every boundary grammar in `propose` is line-anchored -- measured
|
||||
outside this repo, **828 of 828** real sections gave 0 plans and exit 2 at
|
||||
every sample point, and a coarser 145-document cut gave 145 of 145. Block tags
|
||||
now open their own lines and `h1`-`h6` carry the ATX marker for their OWN level
|
||||
(a flat `#` would hand `_ATX` three top-level boundaries where the document
|
||||
declares one section and two subsections). The output grammar is MARKDOWN, the
|
||||
same the office rows reach the proposer through, so **no HTML-only heading
|
||||
grammar exists**; the fix is in the extractor and **never** the converter,
|
||||
because `.html` stays out of `_PANDOC_FORMATS` on CVE-2025-51591. After:
|
||||
**828 of 828 plans, exit 0, 3206 concepts / 6015 md -- the markdown path's
|
||||
count EXACTLY**, and the same at 414 (1651) and 83 (343). Text preservation is
|
||||
an EXACT invariant and not a percentage: strip the added ATX markers and the
|
||||
non-whitespace sequence is identical to the old extractor's, **828 of 828
|
||||
files**, character ratio **1.000000**. `_SKIP_TAGS` stays `{script, style}`.
|
||||
Exposure elsewhere measured rather than argued: **0 of 86** K2 corpus files and
|
||||
**0 of 5** smoke-folder files are HTML, and the smoke bundle is byte-identical
|
||||
before and after. `_EVIDENCE` gains a `.html` row at `measured`, with the limit
|
||||
that travels with it -- one product, one format, one publisher, and a
|
||||
generator's cut, not 828 documents anyone wrote. `pdf`/`docx`/`xlsx` only via
|
||||
the optional `[extract]` extra; without it those types are rejected
|
||||
fail-fast. The extra ships `pdfplumber` for `pdf` (chosen on ONE measured
|
||||
property: it keeps a requirement table's label and value on the same line
|
||||
|
|
@ -598,6 +620,39 @@ and fixtures, never code.
|
|||
here -- a genuine Norwegian morpheme, so that residual is a different answer,
|
||||
never a ceiling. The vocabulary is the BUNDLE's own, so the rule makes a
|
||||
payload corpus-dependent the way `rarity_weights` already is.
|
||||
**A SEVENTH flag, `--source-quota N`, is ON at 2 since 2026-09-10** (opt-out
|
||||
`--no-source-quota`) and is the third change here that alters a payload with
|
||||
NO bundle changing. It caps how many DELIVERED places one `source_file` may
|
||||
take, cutting where `shortlist = candidates[:k]` cuts, so the freed place goes
|
||||
to the next candidate and `k` is still delivered in full. The defect was
|
||||
measured OUTSIDE this repo on a 3206-concept bundle of a published handbook:
|
||||
the code's own process overview is **28 of 3206 concepts (0.87 %)** and **8.0 %
|
||||
of the source characters** yet took **8 of 8** delivered places on one question
|
||||
and **7 of 8** on the known-positive, which was not delivered at all --
|
||||
identical at 343 and 1651 concepts, so it is the corpus's COMPOSITION (it holds
|
||||
its own table of contents) and not its size, and a split would move it rather
|
||||
than remove it. Swept over {2, 3, 4, off} on three bundles with the fasit
|
||||
prefixes validated against the bundle FIRST (that control caught a defect in
|
||||
the measuring query itself): at 2 and 3 hit@8 goes **5 of 6 to 6 of 6 on BOTH
|
||||
K2 bundles** with all five standing rank-1 rows unmoved -- the recovered row
|
||||
had missed on every bundle and every configuration measured until now -- and at
|
||||
4 and off it stays 5 of 6. On the handbook bundle hit@8 goes **2 of 6 to 4 of
|
||||
6** and the dominant document's share **8 of 8 to 2 of 8**. 2 rather than 3 on
|
||||
rank. **What the gain is NOT:** hit@8 asks whether the gold DOCUMENT was
|
||||
delivered and a document quota raises how many distinct documents a payload
|
||||
holds, so that metric is not neutral with respect to this rule; the five rows
|
||||
already at rank 1 are, and did not move. **The adverse case is named:** a
|
||||
one-document bundle has one `source_file` on every concept, so the quota would
|
||||
deliver 2 where `k` were asked -- the shortlist is topped back up from the
|
||||
best-ranked over-quota candidates, making such a bundle byte-identical to the
|
||||
quota being off. The `WITHHOLDING_RULES` vocabulary goes six to seven
|
||||
(`source_quota_exceeded`, a DIVERSITY drop and not a relevance one) and is
|
||||
published in the contract SS 5.3 and in the generated SKILL.md, verified by
|
||||
reading the generated file. Editing the contract moved the SS 7.4
|
||||
known-positive (12 563 -> 13 238 encoded, delta 336 -> 345), which is that
|
||||
coupling working. `--rarity-weight` was measured against the same defect and
|
||||
does NOT repair it -- it leaves the dominant document at 8 of 8 places on the
|
||||
question it floods -- and stays off.
|
||||
`--rarity-weight` is the third: each lexical hit weighs `log(N/df)` over the
|
||||
bundle's own concepts instead of 1, so an identifier is not worth what a
|
||||
common verb is worth. It enters the RANKING and never the GATE — `lexical`
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue