feat(consume): measure the budget lock, add one flag-gated top-rank reservation
The prior measurement (docs/2026-09-08-blindsone-below-k-k2.md SS 3) found that
the budget, not the ranking, is the second lock on a mandate-shaped cost
question -- and that the same mechanism was a REGRESSION on the question that
works: raising `--k` to 16 evicted the gold concept, because the exact knapsack
maximises a SUM of fused scores and has no opinion about rank, so twenty small
excerpts out-value one that costs 56.5 % of the budget.
Measured here on the same 629-concept bundle, with the three known-positive
figures from `4c699fd` reproduced first:
- Corpus distribution, denominator 629: median excerpt 857 B, max 223 391 B,
3 concepts over the limit alone.
- Candidate rule (b), a corpus-derived budget, is FALSIFIED by two numbers: two
defensible derivations are 49x apart on the same corpus, the small one turns
the gold concept into `over_budget_alone` (13 refusals against 2), the large
one changes nothing at the default k. A budget is the consumer's constraint,
not a property of the corpus; `--limit` already belongs to the caller.
- Built instead, behind `--reserve-top-rank` (default OFF): the top-ranked
candidate gets its bytes before the pack runs, AFTER the `over_budget_alone`
pre-exclusion and never before, and the payload declares `budget.reserved`.
- It fixes the eviction: k=16 and k=24 deliver the gold concept at rank 1,
costing one and two excerpts, and 20.4 % / 27.3 % FEWER o200k tokens.
- It changes the delivered list in 2 of 24 measured combinations -- both of them
that eviction. In the other 22 the list, its order and `spent` are identical.
- It does NOT close the mandate-shaped blind spot: that concept ranks 10, not 1.
The one delivering command is `--cost-vocabulary --k 12 --limit 160000`
(62 149 tokens against 58 401), and that is a consumer's decision.
11 new tests (RED first), 7 mutations 7 red with an unmutated negative control
green before and after; two of the seven survived the first test set and the
tests were strengthened. Default payload byte-identical, both goldens unchanged.
Report: docs/2026-09-08-blindsone-laas2-budsjett-k2.md
Suite 1279 green, mypy --strict clean over 28 files, ruff clean.
Co-Authored-By: Claude <claude-opus-5>
This commit is contained in:
parent
4c699fdbb1
commit
6776c37d23
5 changed files with 746 additions and 17 deletions
23
CLAUDE.md
23
CLAUDE.md
|
|
@ -268,15 +268,24 @@ and fixtures, never code.
|
|||
measurement behind it, including the control that FAILED, is
|
||||
`docs/2026-09-07-okf-konsumskill-maaling.md`. **The ranking is this
|
||||
repository's own choice** — the contract binds a payload, not a retrieval
|
||||
algorithm (§ 10) — and it has ONE optional widening, `--cost-vocabulary`,
|
||||
**off by default**: a declared cost/price/quantity vocabulary family that
|
||||
bridges a question and a document naming money with different words, gated on
|
||||
the QUESTION carrying such a term, so a question without one is byte-identical
|
||||
either way. It moves a measured case from candidate rank 249 to 10 and does
|
||||
**not** deliver it: the budget is a second, independent lock, and closing that
|
||||
one is a decision nobody has made. Measured, with the two rules falsified
|
||||
algorithm (§ 10) — and it has TWO optional widenings, both **off by default**
|
||||
and both keeping the default payload byte-identical. `--cost-vocabulary`: a
|
||||
declared cost/price/quantity vocabulary family that bridges a question and a
|
||||
document naming money with different words, gated on the QUESTION carrying
|
||||
such a term, so a question without one is byte-identical either way. It moves
|
||||
a measured case from candidate rank 249 to 10 and does **not** deliver it —
|
||||
the budget is a second, independent lock. Measured, with two rules falsified
|
||||
before building and the `k`-sweep that showed a higher `k` can EVICT a gold
|
||||
concept, in `docs/2026-09-08-blindsone-below-k-k2.md`.
|
||||
`--reserve-top-rank` is that second lock: the pack is an exact knapsack over a
|
||||
SUM, so it has no opinion about rank and out-sums a top-ranked excerpt costing
|
||||
a large share of the budget. The flag gives rank one its bytes first, AFTER
|
||||
the `over_budget_alone` pre-exclusion and never before, and declares
|
||||
`budget.reserved` in the payload. It fixes the eviction and does **not** close
|
||||
the mandate-shaped blind spot (that concept ranks 10, not 1); the budget stays
|
||||
the caller's decision, because deriving a limit from the corpus was measured
|
||||
and falsified — two defensible derivations, 49x apart, one of them breaking
|
||||
the known-positive. `docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
|
||||
|
||||
## Workflow
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue