The consumer report (portfolio-optimiser, S7 SS 2) found that a mandate-shaped
cost question withheld the corpus's one priced table under `below_k`. Measured
here, on a bundle proven byte-identical to a fresh HEAD rebuild:
- The mechanism is a VOCABULARY gap, not a `k` defect: two of three ranking
signals are exactly 0.0 and the concept is candidate 249 of 269.
- The k-sweep buys nothing: k in {8,12,16,24,32,64,128} all withhold it, at
+9.5 % tokens. It also found a regression -- for the question that WORKS,
k >= 16 EVICTS the gold concept, because one 67 838 B excerpt is 56.5 % of
the budget and the knapsack maximises a sum.
- Two proposed rules were falsified BEFORE any code: number/table density ranks
the priced table 178/165/46 of 269 (the form is unfilled, so it is
number-poor), and per-document spread puts its document 30th of 35.
Built instead, behind `--cost-vocabulary` (default OFF, DEFAULT byte-identical):
one declared cost/price/quantity vocabulary family that bridges a question and a
document naming money with different words. It moves the concept from candidate
rank 249 to 10 -- and does NOT close the blind spot: the budget still refuses
it, which is now a separately measured second lock.
Seven RED tests first; six mutations of the rule, six red (two survived the
first version of the tests and the tests were strengthened). Control: a question
with no cost term produces a byte-identical payload with the flag on, at every
k, on the real corpus. Known-positive: 164 987 B / 40 425 o200k tokens, equal to
the published pair.
Report: docs/2026-09-08-blindsone-below-k-k2.md
Suite 1268 green, mypy --strict clean over 28 files, both goldens unchanged.
Co-Authored-By: Claude <claude-opus-5>
392 lines
20 KiB
Markdown
392 lines
20 KiB
Markdown
# The cut's blind spot: a priced table below k, measured on a 629-concept corpus
|
|
|
|
**Date:** 2026-09-08 · **Order:** `20260908T021157Z-6753710732-from-.claude` ·
|
|
**Instrument:** `tools/okf_consume.py` at `5a0c879` plus the one flag this
|
|
document reports · **Upstream finding:** `portfolio-optimiser`
|
|
`docs/2026-09-07-syretest-s7-prepass-k2.md` § 2 and § 4.
|
|
|
|
The consumer report this answers observed that a mandate-shaped question
|
|
("find the cost savings in this tender") delivered 8 of 630 concepts and
|
|
withheld, under the rule `below_k`, the single concept in the corpus that
|
|
carries a price. Free navigation reached that concept in four steps. This
|
|
document measures why, sweeps `k`, and reports one flag-gated rule built after
|
|
the measurement -- including the two candidate rules the measurement killed
|
|
before any code was written.
|
|
|
|
The corpus is external and private to the measurement; no document name, path
|
|
or body from it appears here. Documents are named by shape ("the priced
|
|
table") and the numbers are counts.
|
|
|
|
---
|
|
|
|
## 0. What IS measured, and what is NOT
|
|
|
|
**Measured.** Today's ranking for both questions the upstream report used, with
|
|
each score component that placed the priced table where it sits; the rank of
|
|
that concept among the lexical candidates, with a denominator; a `k` sweep at
|
|
`k` in {8, 12, 16, 24} plus four larger values, with payload bytes and o200k
|
|
tokens at each; the three candidate rules against real numbers; and the one
|
|
rule that was built, on both questions plus a third question that carries no
|
|
cost term at all.
|
|
|
|
**Known-positive, run first.** The payload for the specific question, `k=8`,
|
|
flag off, measures **164 987 B / 40 425 o200k tokens** -- the two figures
|
|
published 2026-09-07 (`portfolio-optimiser`
|
|
`docs/2026-09-07-okf-prepass-i-debatten.md` § 1), byte for byte and token for
|
|
token. The tokenizer is `tiktoken` `o200k_base`, the same counter that produced
|
|
the published number. An instrument that has not reproduced a known figure has
|
|
not been shown to count (consumption contract § 7.4).
|
|
|
|
**NOT measured.** That the rule below helps any corpus other than this one:
|
|
one corpus, two questions and one control question is not a sample, and the
|
|
vocabulary it declares is Norwegian. Not measured either: whether a model
|
|
answers *better* with the priced table in the payload -- that needs a live
|
|
model and is the consumer's measurement, not this one. And not measured: that
|
|
`k=8` is the right default. This document recommends; the default is the
|
|
operator's decision and is unchanged here.
|
|
|
|
**Nothing is decided about the default.** The flag ships OFF. With the flag off
|
|
every payload in this repository is byte-identical to `5a0c879`, and the two
|
|
golden fixtures are unchanged.
|
|
|
|
---
|
|
|
|
## 1. Setup
|
|
|
|
The bundle is the 629-concept build of the corpus produced by `okf build` on
|
|
`5a0c879` with `--ingested-at 2026-09-03T00:00:00Z`, identity
|
|
`sha256-tree:f14872a01104e47474093611b1960c6c541e4701dc40147a00c8e1b337c8a92a`.
|
|
It differs from the bundle the upstream report measured in exactly the two ways
|
|
`5a0c879` fixed: every segmented concept now carries the stamp the flag
|
|
declares, and the run log is no longer walked as a concept. That second fix is
|
|
visible in the denominators below as **629 considered** where the upstream
|
|
report has 630, and as one fewer `no_lexical_match` (358 against 359).
|
|
|
|
**The bundle is controlled, not assumed.** The measurements below ran against a
|
|
bundle produced by an earlier session's working tree. It was rebuilt from the
|
|
raw corpus on committed `5a0c879` while the measurements ran, and `diff -r`
|
|
between the two trees is **exit 0, zero lines** -- so every number here is a
|
|
number about HEAD. That rebuild's own conservation identity holds
|
|
(`merged + coded rejections = 43; N = 43`, 39 substantive, 4 coded rejections,
|
|
780.47 s), its `ref` is the one above, and the contract check on its payload is
|
|
**exit 0** ("conformant: 14 rules over 8 excerpts and 621 withheld entries, 0
|
|
findings").
|
|
|
|
Every command in this document is offline: no model call, no socket, no clock.
|
|
|
|
---
|
|
|
|
## 2. The ranking, and where the priced table sits in it
|
|
|
|
Both questions are the upstream report's, verbatim. The mandate-shaped one is
|
|
in that report § 2; the specific one is quoted in
|
|
`docs/2026-09-07-okf-prepass-i-debatten.md` § 1 -- **not** in the S7a document
|
|
the order named, which mentions neither wording. Stated rather than silently
|
|
corrected.
|
|
|
|
| | mandate-shaped question | specific question |
|
|
|---|---|---|
|
|
| considered | 629 | 629 |
|
|
| delivered | 8 | 8 |
|
|
| withheld | 621 | 621 |
|
|
| — `no_lexical_match` | 358 | 582 |
|
|
| — `below_k` | 261 | 37 |
|
|
| — `over_budget_alone` | 2 | 2 |
|
|
| identity closes | 629 = 621 + 8 | 629 = 621 + 8 |
|
|
| priced table | **withheld, `below_k`** | **delivered, rank 1** |
|
|
|
|
The specific question is the known-positive for the ranker itself: the same
|
|
ranker, the same bundle, the same `k`, and the gold concept comes first.
|
|
|
|
### The priced table's own score, both questions
|
|
|
|
The three signals are the ones `concept_scores` fuses by RRF: (1) the question
|
|
against the concept's title and the segments of its id, (2) the question
|
|
against the body, (3) the stage-one score of the document the concept belongs
|
|
to, which is a **density** over that document's index entries and concept ids.
|
|
|
|
| | mandate-shaped | specific |
|
|
|---|---|---|
|
|
| signal 1 — title + id | **0.0** (rank 616 of 629) | 1.0 (rank 3) |
|
|
| signal 2 — body | 2.0 (rank 193) | 2.0 (rank 7) |
|
|
| signal 3 — document density | **0.0** (rank 616) | 1.0 (rank 1) |
|
|
| fused score | 0.00691115 | 0.04719183 |
|
|
| rank among ALL concepts | 489 of 629 | 1 of 629 |
|
|
| **rank among lexical candidates** | **249 of 269** | **1 of 45** |
|
|
|
|
**The mechanism, in one line: two of three signals are exactly zero.** The
|
|
question normalises to five tokens: a verb, the compound
|
|
`kostnadsbesparelser`, a place name, a building type and the word for the
|
|
tender. The priced table's title, its id and its document's index entries
|
|
contain none of them. Its body earns 2 -- one of them the building type, the
|
|
other the four-character prefix `kost` inside a longer word. The document that IS the
|
|
answer scores 0 at the document level, because a document about `pris` shares
|
|
no four-character prefix with a question about `kostnadsbesparelser`.
|
|
|
|
`MIN_SHARED_PREFIX` is 4, and `tokens_match` is symmetric prefix matching. This
|
|
is not a defect in the matcher; the matcher is doing exactly what it says.
|
|
**The gap is in the vocabulary**, and § 3 shows no value of `k` closes a
|
|
vocabulary gap.
|
|
|
|
---
|
|
|
|
## 3. The k-sweep: what raising k buys, and what it costs
|
|
|
|
`--k` caps the delivered set; the budget (120 000 B, `DEFAULT_LIMIT`) is the
|
|
real gate. Payload bytes are the serialised payload; tokens are o200k over the
|
|
same bytes.
|
|
|
|
| k | delivered | payload B | o200k tok | priced table |
|
|
|---|---|---|---|---|
|
|
| 8 | 8 | 169 583 | 57 289 | `below_k` |
|
|
| 12 | 12 | 172 689 | 58 585 | `below_k` |
|
|
| 16 | 16 | 177 581 | 60 778 | `below_k` |
|
|
| 24 | 23 | 182 715 | 62 723 | `below_k` |
|
|
|
|
Mandate-shaped question, flag off. **Nothing arrives, and 5 434 tokens
|
|
(+9.5 %) are spent discovering that.** Continued past the order's four values,
|
|
on the same run: `k=32` (31 delivered), `k=64` (54), `k=128` (85) -- still
|
|
`below_k`; at `k=249`, the candidate rank itself, the rule finally changes to
|
|
`over_budget_after_knapsack`. So `k` was never the binding constraint for this
|
|
question. **Candidate rule (c) -- "no rule; k=12 alone does the job at a
|
|
measured token price" -- is falsified.**
|
|
|
|
### The sweep also found a regression, on the question that works
|
|
|
|
| k | delivered | payload B | o200k tok | priced table |
|
|
|---|---|---|---|---|
|
|
| 8 | 8 | 164 987 | 40 425 | **delivered, rank 1** |
|
|
| 12 | 11 | 196 550 | 49 571 | **delivered, rank 1** |
|
|
| 16 | 15 | 194 946 | 65 237 | **`over_budget_after_knapsack`** |
|
|
| 24 | 20 | 197 287 | 66 799 | **`over_budget_after_knapsack`** |
|
|
|
|
Specific question, flag off. **Raising `k` EVICTS the gold concept.** The
|
|
knapsack maximises the sum of fused scores under the byte budget; the priced
|
|
table is a 67 838 B spreadsheet render, **56.5 % of the whole budget**, and
|
|
once the pool holds enough small excerpts, twenty of them out-value it. This
|
|
is not a bug in the DP -- it is exact and does what it says -- but it means
|
|
`k` is not a safety dial: raising it can remove the one document a question was
|
|
asked about. Reported here because the sweep the order asked for produced it.
|
|
|
|
---
|
|
|
|
## 4. What scores today, verbatim
|
|
|
|
From `tools/okf_consume.py`, quoted rather than summarised:
|
|
|
|
- `MIN_TOKEN_LENGTH = 3` — "The shortest token this instrument scores."
|
|
- `MIN_SHARED_PREFIX = 4` — "How many leading characters two tokens must share
|
|
to count as a match ... MEASURED 2026-09-07 over a 629-concept corpus".
|
|
- `document_scores` — "One score per top-level document, from the indexes and
|
|
the paths alone ... **The score is a DENSITY, not a sum**".
|
|
- `concept_scores` — "Every concept, ordered best first, fused from three
|
|
signals by RRF ... The third element of each tuple is the concept's OWN
|
|
lexical overlap -- signals 1 and 2 only, with the document prior excluded."
|
|
- `cut` — "**A concept answering nothing in the question is withheld, never
|
|
ranked into the top k as filler.**"
|
|
- `DEFAULT_K = 8` — "`--k` caps the DELIVERED set. The budget is the gate; this
|
|
is a second, cheaper bound".
|
|
|
|
And the contract's own boundary, `docs/consumption-contract.md` § 10: "**No
|
|
engine, ranker or cutter is designed here.** The contract binds a payload and a
|
|
document, not a retrieval algorithm." The ranking is this repository's choice;
|
|
changing it breaks no contract, and it is why the change below is a flag rather
|
|
than a new default.
|
|
|
|
---
|
|
|
|
## 5. Three candidate rules, two killed by measurement before any code
|
|
|
|
**(b) table/number density as a tie-break for cost-vocabulary questions —
|
|
FALSIFIED.** Three density definitions were measured over the 269 lexical
|
|
candidates; the priced table's rank under each: **178/269** (digits over
|
|
alphanumerics), **165/269** (fraction of numeric tokens), **46/269** (fraction
|
|
of lines carrying two or more numeric fields). The documents that rank first
|
|
under all three are room lists and drawing schedules. The reason is in the
|
|
corpus and was already published: the price form is **not filled in** -- one
|
|
priced row in the whole sheet, the rest empty cells the contractor is meant to
|
|
fill. A number-density rule finds the documents full of room numbers and misses
|
|
the one document about money. Building it would have taken a day and produced a
|
|
worse ranking.
|
|
|
|
**(a) spread — at least one delivered concept per top-level document with a
|
|
lexical hit, within the same k — FALSIFIED at the k values the order named.**
|
|
Measured: 269 candidates spread over **35 top-level documents**, and the priced
|
|
table's document ranks **30th of those 35** by its best candidate. One slot per
|
|
document at `k=8` reaches eight documents; the target needs `k>=30`, where § 3
|
|
already shows the knapsack drops a 67 838 B excerpt anyway.
|
|
|
|
**(a') the rule that was built: one declared vocabulary family, behind
|
|
`--cost-vocabulary`, default off.** The measurement in § 2 says the failure is
|
|
that two of three signals are zero because the question and the document use
|
|
different words for money. So: a single list of Norwegian cost/price/quantity
|
|
roots, and within that list any term answers to any other -- in all three
|
|
signals, and only when the QUESTION itself carries such a term.
|
|
|
|
```python
|
|
COST_VOCABULARY = (
|
|
"beløp", "budsjett", "enhet", "honorar", "kost",
|
|
"kroner", "mengde", "pris", "utgift", "vederlag",
|
|
)
|
|
```
|
|
|
|
Three properties, each with a test that goes red without it:
|
|
|
|
- **The gate is the question, not the flag.** A question naming no term in the
|
|
family produces byte-identical bytes with the flag set. Measured on the
|
|
corpus in § 6, not only on the fixture.
|
|
- **The bridge needs a family term on BOTH sides**, and carries only the family
|
|
term: a question's unrelated tokens do not ride along on it. Without this the
|
|
rule would read "everything matches a price document".
|
|
- **Every member is at least `MIN_SHARED_PREFIX` characters.** `sum` is three
|
|
and can never match `Summen`; it was dropped for that reason, and the test
|
|
states the reason.
|
|
|
|
**Honesty about the list, measured leave-one-out on the corpus:** the entire
|
|
effect rests on **two** members, `kost` and `pris`. Removing either returns the
|
|
priced table to rank 249; removing any other member moves it not at all. Three
|
|
members (`budsjett`, and two spellings that cannot match) reach zero concepts
|
|
in this corpus. They are kept because dropping a term for being absent from ONE
|
|
corpus fits the list to that corpus -- but a reader should treat this as a
|
|
**two-word bridge measured on one question**, not as a vocabulary that has been
|
|
shown to generalise.
|
|
|
|
Development order: seven failing tests first, then the implementation. Six
|
|
mutations of the shipped rule were run against the new tests; **all six are
|
|
red** (one-sided bridge; gate stuck open; default flipped on; the load-bearing
|
|
member removed; every question token riding the bridge; a member too short to
|
|
ever match). Two of those six survived the first version of the tests and the
|
|
tests were strengthened until they did not.
|
|
|
|
---
|
|
|
|
## 6. The rule, measured on both questions and on a control
|
|
|
|
`--cost-vocabulary`, same bundle, same budget, `k` swept.
|
|
|
|
| question | flag | k | delivered | payload B | o200k tok | priced table |
|
|
|---|---|---|---|---|---|---|
|
|
| mandate | off | 8 | 8 | 169 583 | 57 289 | `below_k` |
|
|
| mandate | **on** | 8 | 8 | 161 338 | 54 996 | `below_k` |
|
|
| mandate | **on** | 12 | 11 | 172 246 | 58 401 | **`over_budget_after_knapsack`** |
|
|
| mandate | **on** | 16 | 15 | 176 591 | 60 433 | `over_budget_after_knapsack` |
|
|
| mandate | **on** | 24 | 23 | 183 178 | 63 029 | `over_budget_after_knapsack` |
|
|
| specific | off | 8 | 8 | 164 987 | 40 425 | delivered, rank 1 |
|
|
| specific | **on** | 8 | 8 | 164 879 | 40 389 | **delivered, rank 3** |
|
|
| specific | off | 16 | 15 | 194 946 | 65 237 | `over_budget_after_knapsack` |
|
|
| specific | **on** | 16 | 15 | 207 113 | 52 370 | **delivered, rank 3** |
|
|
| control | off | 8 | 7 | — | — | not in this question's answer set |
|
|
| control | **on** | 8 | 7 | — | — | **byte-identical payload** |
|
|
|
|
**What the rule does:** it moves the priced table from candidate rank
|
|
**249 of 269 to 10 of 278** for the mandate-shaped question. The rule
|
|
`below_k` gives way to `over_budget_after_knapsack` from `k=12` on -- the
|
|
ranking objection is gone and a different one takes its place.
|
|
|
|
**What the rule does NOT do: it does not close the blind spot.** At no tested
|
|
`k` does the mandate-shaped question deliver the priced table. Moving a
|
|
document from invisible to visible-but-unaffordable is progress that can be
|
|
measured, and it is not the same as an answer.
|
|
|
|
**Q-good is CHANGED, and that is stated as the order requires.** The specific
|
|
question's delivered SET at `k=8` is the same eight concepts, but the priced
|
|
table moves from rank 1 to rank 3 and the payload is therefore not
|
|
byte-identical (164 987 B against 164 879 B). This is a change to a working
|
|
question and must be read as a cost of the rule. It is not all cost: at `k=16`
|
|
the flag-off run has already evicted the gold concept and the flag-on run
|
|
still delivers it.
|
|
|
|
**The control is the strongest single number here.** A question with no cost
|
|
term produces a **byte-identical payload** with the flag on, at every `k`
|
|
measured, on the real corpus. The widening is confined to the question class it
|
|
names.
|
|
|
|
---
|
|
|
|
## 7. The second lock, isolated
|
|
|
|
With the flag on, `k=12`, only the budget varied:
|
|
|
|
| limit (B) | delivered | spent | priced table |
|
|
|---|---|---|---|
|
|
| 120 000 (default) | 11 | 82 399 | `over_budget_after_knapsack` |
|
|
| 140 000 | 11 | 82 399 | `over_budget_after_knapsack` |
|
|
| **160 000** | 12 | 150 249 | **delivered, rank 10** |
|
|
| 200 000 | 12 | 150 249 | delivered, rank 10 |
|
|
|
|
And the same sweep with the flag OFF: the priced table is `below_k` at every
|
|
limit, because it never reaches the shortlist. **The two locks are independent
|
|
and now separately measured.** Lock 1 is the vocabulary and the flag removes
|
|
it. Lock 2 is that one 67 838 B excerpt is 56.5 % of a 120 000 B budget and the
|
|
knapsack, maximising a sum of scores, prefers twenty small excerpts. Closing
|
|
lock 2 is a second rule -- reserving budget for the top-ranked candidate, or
|
|
sizing the budget to the corpus -- and this order allowed one.
|
|
|
|
Consumption contract § 7.6 asked for exactly this number: "the corpus size at
|
|
which its strategy stops fitting its budget". For this corpus it is not a
|
|
size; it is a single document that costs more than half the budget.
|
|
|
|
---
|
|
|
|
## 8. Honesty limits
|
|
|
|
- **One corpus, two questions, one control.** Generality is NOT demonstrated.
|
|
The vocabulary is Norwegian, and a corpus in another language gets nothing
|
|
from it.
|
|
- **The list was written with both words visible.** `kost` and `pris` are the
|
|
two words in the question and in the document that failed. The same
|
|
disclosure the ranker already carries about `MIN_SHARED_PREFIX` applies here:
|
|
the rule is not blind to the case that motivated it.
|
|
- **The rank improvement is real and the delivery is not.** Every claim that
|
|
the rule "finds" the document should be read against § 6: it ranks it 10th
|
|
and the budget still refuses it.
|
|
- **The eviction finding in § 3 is measured on one question.** That raising `k`
|
|
can evict a gold concept is demonstrated for this pair of question and
|
|
corpus, not proven as a general property of the DP.
|
|
- **`--cost-vocabulary` has no consumer.** Nobody asked for it; it exists so
|
|
the measurement above could be made against real code rather than a
|
|
simulation, and so a decision about it can be made on numbers.
|
|
|
|
---
|
|
|
|
## 9. Recommendation
|
|
|
|
1. **Keep `DEFAULT_K = 8`.** The sweep shows raising `k` buys no answer for the
|
|
mandate-shaped question and can evict the gold concept from the specific
|
|
one. This is the opposite of what the order's option (c) expected, and it is
|
|
measured.
|
|
2. **Keep `--cost-vocabulary` OFF by default.** It is a two-word bridge
|
|
measured on one question; the number that would justify a default is a
|
|
hit-rate over a question set nobody has built yet.
|
|
3. **The blind spot stays open, and it is a BUDGET question now, not a ranking
|
|
question.** If it matters to a consumer, the next order is lock 2: reserve
|
|
budget for the top-ranked candidate, or let a profile size its budget to its
|
|
corpus. That is one rule, it has a clean red test (§ 3's eviction), and it
|
|
is a decision about what a payload is for.
|
|
4. **A mandate is not a query, and no lexical ranker will make it one.** The
|
|
upstream report's own arm reached this document in four navigational steps.
|
|
The honest boundary of a declared cut is that it answers questions, and a
|
|
mandate is a brief. Saying that in the skill's own words costs nothing and
|
|
is more accurate than any `k`.
|
|
|
|
---
|
|
|
|
## 10. Verification log
|
|
|
|
| # | Claim | Command → result |
|
|
|---|---|---|
|
|
| 1 | The token instrument reproduces a published figure | payload for the specific question, `k=8`, flag off → **164 987 B / 40 425 o200k tok**, equal to the 2026-09-07 published pair |
|
|
| 2 | The ranker's known-positive still holds | same question, flag off, `k=8` → priced table delivered at **rank 1** |
|
|
| 3 | The denominators close | 629 = 621 + 8, both questions |
|
|
| 4 | The priced table's rank, mandate-shaped question | **249 of 269** lexical candidates; signals 1 and 3 both 0.0 |
|
|
| 5 | `k` never delivers it | `k` in {8, 12, 16, 24, 32, 64, 128} → `below_k`; at 249 → `over_budget_after_knapsack` |
|
|
| 6 | Raising `k` evicts the gold on the specific question | `k=16` and `k=24` → `over_budget_after_knapsack` |
|
|
| 7 | Number density does not find it | three definitions → rank 178, 165, 46 of 269 |
|
|
| 8 | Spread does not find it at these `k` | document rank **30 of 35** |
|
|
| 9 | The rule moves it | candidate rank **249 → 10** |
|
|
| 10 | The rule does not deliver it | every `k` in {8, 12, 16, 24} → withheld, by two different rules |
|
|
| 11 | The gate is the question | control question, flag on vs off, every `k` → **byte-identical payload** |
|
|
| 12 | The default does not move | flag off ⇒ payload byte-identical to `5a0c879`; both goldens unchanged |
|
|
| 13 | Six mutations, six red | one-sided bridge, gate open, default on, member removed, every token bridges, member too short |
|
|
| 14 | The measured bundle IS a HEAD build | `diff -r` fresh `5a0c879` rebuild against the measured tree → **exit 0, 0 lines**; contract check on its payload → exit 0 |
|
|
| 15 | Suite, types, lint | `pytest -q` **1268 passed**; `mypy --strict src/ tools/` 28 files clean; `ruff check` + `ruff format --check` clean |
|