# The cut's blind spot: a priced table below k, measured on a 629-concept corpus **Date:** 2026-09-08 · **Order:** `20260908T021157Z-6753710732-from-.claude` · **Instrument:** `tools/okf_consume.py` at `5a0c879` plus the one flag this document reports · **Upstream finding:** `portfolio-optimiser` `docs/2026-09-07-syretest-s7-prepass-k2.md` § 2 and § 4. The consumer report this answers observed that a mandate-shaped question ("find the cost savings in this tender") delivered 8 of 630 concepts and withheld, under the rule `below_k`, the single concept in the corpus that carries a price. Free navigation reached that concept in four steps. This document measures why, sweeps `k`, and reports one flag-gated rule built after the measurement -- including the two candidate rules the measurement killed before any code was written. The corpus is external and private to the measurement; no document name, path or body from it appears here. Documents are named by shape ("the priced table") and the numbers are counts. --- ## 0. What IS measured, and what is NOT **Measured.** Today's ranking for both questions the upstream report used, with each score component that placed the priced table where it sits; the rank of that concept among the lexical candidates, with a denominator; a `k` sweep at `k` in {8, 12, 16, 24} plus four larger values, with payload bytes and o200k tokens at each; the three candidate rules against real numbers; and the one rule that was built, on both questions plus a third question that carries no cost term at all. **Known-positive, run first.** The payload for the specific question, `k=8`, flag off, measures **164 987 B / 40 425 o200k tokens** -- the two figures published 2026-09-07 (`portfolio-optimiser` `docs/2026-09-07-okf-prepass-i-debatten.md` § 1), byte for byte and token for token. The tokenizer is `tiktoken` `o200k_base`, the same counter that produced the published number. An instrument that has not reproduced a known figure has not been shown to count (consumption contract § 7.4). **NOT measured.** That the rule below helps any corpus other than this one: one corpus, two questions and one control question is not a sample, and the vocabulary it declares is Norwegian. Not measured either: whether a model answers *better* with the priced table in the payload -- that needs a live model and is the consumer's measurement, not this one. And not measured: that `k=8` is the right default. This document recommends; the default is the operator's decision and is unchanged here. **Nothing is decided about the default.** The flag ships OFF. With the flag off every payload in this repository is byte-identical to `5a0c879`, and the two golden fixtures are unchanged. --- ## 1. Setup The bundle is the 629-concept build of the corpus produced by `okf build` on `5a0c879` with `--ingested-at 2026-09-03T00:00:00Z`, identity `sha256-tree:f14872a01104e47474093611b1960c6c541e4701dc40147a00c8e1b337c8a92a`. It differs from the bundle the upstream report measured in exactly the two ways `5a0c879` fixed: every segmented concept now carries the stamp the flag declares, and the run log is no longer walked as a concept. That second fix is visible in the denominators below as **629 considered** where the upstream report has 630, and as one fewer `no_lexical_match` (358 against 359). **The bundle is controlled, not assumed.** The measurements below ran against a bundle produced by an earlier session's working tree. It was rebuilt from the raw corpus on committed `5a0c879` while the measurements ran, and `diff -r` between the two trees is **exit 0, zero lines** -- so every number here is a number about HEAD. That rebuild's own conservation identity holds (`merged + coded rejections = 43; N = 43`, 39 substantive, 4 coded rejections, 780.47 s), its `ref` is the one above, and the contract check on its payload is **exit 0** ("conformant: 14 rules over 8 excerpts and 621 withheld entries, 0 findings"). Every command in this document is offline: no model call, no socket, no clock. --- ## 2. The ranking, and where the priced table sits in it Both questions are the upstream report's, verbatim. The mandate-shaped one is in that report § 2; the specific one is quoted in `docs/2026-09-07-okf-prepass-i-debatten.md` § 1 -- **not** in the S7a document the order named, which mentions neither wording. Stated rather than silently corrected. | | mandate-shaped question | specific question | |---|---|---| | considered | 629 | 629 | | delivered | 8 | 8 | | withheld | 621 | 621 | | — `no_lexical_match` | 358 | 582 | | — `below_k` | 261 | 37 | | — `over_budget_alone` | 2 | 2 | | identity closes | 629 = 621 + 8 | 629 = 621 + 8 | | priced table | **withheld, `below_k`** | **delivered, rank 1** | The specific question is the known-positive for the ranker itself: the same ranker, the same bundle, the same `k`, and the gold concept comes first. ### The priced table's own score, both questions The three signals are the ones `concept_scores` fuses by RRF: (1) the question against the concept's title and the segments of its id, (2) the question against the body, (3) the stage-one score of the document the concept belongs to, which is a **density** over that document's index entries and concept ids. | | mandate-shaped | specific | |---|---|---| | signal 1 — title + id | **0.0** (rank 616 of 629) | 1.0 (rank 3) | | signal 2 — body | 2.0 (rank 193) | 2.0 (rank 7) | | signal 3 — document density | **0.0** (rank 616) | 1.0 (rank 1) | | fused score | 0.00691115 | 0.04719183 | | rank among ALL concepts | 489 of 629 | 1 of 629 | | **rank among lexical candidates** | **249 of 269** | **1 of 45** | **The mechanism, in one line: two of three signals are exactly zero.** The question normalises to five tokens: a verb, the compound `kostnadsbesparelser`, a place name, a building type and the word for the tender. The priced table's title, its id and its document's index entries contain none of them. Its body earns 2 -- one of them the building type, the other the four-character prefix `kost` inside a longer word. The document that IS the answer scores 0 at the document level, because a document about `pris` shares no four-character prefix with a question about `kostnadsbesparelser`. `MIN_SHARED_PREFIX` is 4, and `tokens_match` is symmetric prefix matching. This is not a defect in the matcher; the matcher is doing exactly what it says. **The gap is in the vocabulary**, and § 3 shows no value of `k` closes a vocabulary gap. --- ## 3. The k-sweep: what raising k buys, and what it costs `--k` caps the delivered set; the budget (120 000 B, `DEFAULT_LIMIT`) is the real gate. Payload bytes are the serialised payload; tokens are o200k over the same bytes. | k | delivered | payload B | o200k tok | priced table | |---|---|---|---|---| | 8 | 8 | 169 583 | 57 289 | `below_k` | | 12 | 12 | 172 689 | 58 585 | `below_k` | | 16 | 16 | 177 581 | 60 778 | `below_k` | | 24 | 23 | 182 715 | 62 723 | `below_k` | Mandate-shaped question, flag off. **Nothing arrives, and 5 434 tokens (+9.5 %) are spent discovering that.** Continued past the order's four values, on the same run: `k=32` (31 delivered), `k=64` (54), `k=128` (85) -- still `below_k`; at `k=249`, the candidate rank itself, the rule finally changes to `over_budget_after_knapsack`. So `k` was never the binding constraint for this question. **Candidate rule (c) -- "no rule; k=12 alone does the job at a measured token price" -- is falsified.** ### The sweep also found a regression, on the question that works | k | delivered | payload B | o200k tok | priced table | |---|---|---|---|---| | 8 | 8 | 164 987 | 40 425 | **delivered, rank 1** | | 12 | 11 | 196 550 | 49 571 | **delivered, rank 1** | | 16 | 15 | 194 946 | 65 237 | **`over_budget_after_knapsack`** | | 24 | 20 | 197 287 | 66 799 | **`over_budget_after_knapsack`** | Specific question, flag off. **Raising `k` EVICTS the gold concept.** The knapsack maximises the sum of fused scores under the byte budget; the priced table is a 67 838 B spreadsheet render, **56.5 % of the whole budget**, and once the pool holds enough small excerpts, twenty of them out-value it. This is not a bug in the DP -- it is exact and does what it says -- but it means `k` is not a safety dial: raising it can remove the one document a question was asked about. Reported here because the sweep the order asked for produced it. --- ## 4. What scores today, verbatim From `tools/okf_consume.py`, quoted rather than summarised: - `MIN_TOKEN_LENGTH = 3` — "The shortest token this instrument scores." - `MIN_SHARED_PREFIX = 4` — "How many leading characters two tokens must share to count as a match ... MEASURED 2026-09-07 over a 629-concept corpus". - `document_scores` — "One score per top-level document, from the indexes and the paths alone ... **The score is a DENSITY, not a sum**". - `concept_scores` — "Every concept, ordered best first, fused from three signals by RRF ... The third element of each tuple is the concept's OWN lexical overlap -- signals 1 and 2 only, with the document prior excluded." - `cut` — "**A concept answering nothing in the question is withheld, never ranked into the top k as filler.**" - `DEFAULT_K = 8` — "`--k` caps the DELIVERED set. The budget is the gate; this is a second, cheaper bound". And the contract's own boundary, `docs/consumption-contract.md` § 10: "**No engine, ranker or cutter is designed here.** The contract binds a payload and a document, not a retrieval algorithm." The ranking is this repository's choice; changing it breaks no contract, and it is why the change below is a flag rather than a new default. --- ## 5. Three candidate rules, two killed by measurement before any code **(b) table/number density as a tie-break for cost-vocabulary questions — FALSIFIED.** Three density definitions were measured over the 269 lexical candidates; the priced table's rank under each: **178/269** (digits over alphanumerics), **165/269** (fraction of numeric tokens), **46/269** (fraction of lines carrying two or more numeric fields). The documents that rank first under all three are room lists and drawing schedules. The reason is in the corpus and was already published: the price form is **not filled in** -- one priced row in the whole sheet, the rest empty cells the contractor is meant to fill. A number-density rule finds the documents full of room numbers and misses the one document about money. Building it would have taken a day and produced a worse ranking. **(a) spread — at least one delivered concept per top-level document with a lexical hit, within the same k — FALSIFIED at the k values the order named.** Measured: 269 candidates spread over **35 top-level documents**, and the priced table's document ranks **30th of those 35** by its best candidate. One slot per document at `k=8` reaches eight documents; the target needs `k>=30`, where § 3 already shows the knapsack drops a 67 838 B excerpt anyway. **(a') the rule that was built: one declared vocabulary family, behind `--cost-vocabulary`, default off.** The measurement in § 2 says the failure is that two of three signals are zero because the question and the document use different words for money. So: a single list of Norwegian cost/price/quantity roots, and within that list any term answers to any other -- in all three signals, and only when the QUESTION itself carries such a term. ```python COST_VOCABULARY = ( "beløp", "budsjett", "enhet", "honorar", "kost", "kroner", "mengde", "pris", "utgift", "vederlag", ) ``` Three properties, each with a test that goes red without it: - **The gate is the question, not the flag.** A question naming no term in the family produces byte-identical bytes with the flag set. Measured on the corpus in § 6, not only on the fixture. - **The bridge needs a family term on BOTH sides**, and carries only the family term: a question's unrelated tokens do not ride along on it. Without this the rule would read "everything matches a price document". - **Every member is at least `MIN_SHARED_PREFIX` characters.** `sum` is three and can never match `Summen`; it was dropped for that reason, and the test states the reason. **Honesty about the list, measured leave-one-out on the corpus:** the entire effect rests on **two** members, `kost` and `pris`. Removing either returns the priced table to rank 249; removing any other member moves it not at all. Three members (`budsjett`, and two spellings that cannot match) reach zero concepts in this corpus. They are kept because dropping a term for being absent from ONE corpus fits the list to that corpus -- but a reader should treat this as a **two-word bridge measured on one question**, not as a vocabulary that has been shown to generalise. Development order: seven failing tests first, then the implementation. Six mutations of the shipped rule were run against the new tests; **all six are red** (one-sided bridge; gate stuck open; default flipped on; the load-bearing member removed; every question token riding the bridge; a member too short to ever match). Two of those six survived the first version of the tests and the tests were strengthened until they did not. --- ## 6. The rule, measured on both questions and on a control `--cost-vocabulary`, same bundle, same budget, `k` swept. | question | flag | k | delivered | payload B | o200k tok | priced table | |---|---|---|---|---|---|---| | mandate | off | 8 | 8 | 169 583 | 57 289 | `below_k` | | mandate | **on** | 8 | 8 | 161 338 | 54 996 | `below_k` | | mandate | **on** | 12 | 11 | 172 246 | 58 401 | **`over_budget_after_knapsack`** | | mandate | **on** | 16 | 15 | 176 591 | 60 433 | `over_budget_after_knapsack` | | mandate | **on** | 24 | 23 | 183 178 | 63 029 | `over_budget_after_knapsack` | | specific | off | 8 | 8 | 164 987 | 40 425 | delivered, rank 1 | | specific | **on** | 8 | 8 | 164 879 | 40 389 | **delivered, rank 3** | | specific | off | 16 | 15 | 194 946 | 65 237 | `over_budget_after_knapsack` | | specific | **on** | 16 | 15 | 207 113 | 52 370 | **delivered, rank 3** | | control | off | 8 | 7 | — | — | not in this question's answer set | | control | **on** | 8 | 7 | — | — | **byte-identical payload** | **What the rule does:** it moves the priced table from candidate rank **249 of 269 to 10 of 278** for the mandate-shaped question. The rule `below_k` gives way to `over_budget_after_knapsack` from `k=12` on -- the ranking objection is gone and a different one takes its place. **What the rule does NOT do: it does not close the blind spot.** At no tested `k` does the mandate-shaped question deliver the priced table. Moving a document from invisible to visible-but-unaffordable is progress that can be measured, and it is not the same as an answer. **Q-good is CHANGED, and that is stated as the order requires.** The specific question's delivered SET at `k=8` is the same eight concepts, but the priced table moves from rank 1 to rank 3 and the payload is therefore not byte-identical (164 987 B against 164 879 B). This is a change to a working question and must be read as a cost of the rule. It is not all cost: at `k=16` the flag-off run has already evicted the gold concept and the flag-on run still delivers it. **The control is the strongest single number here.** A question with no cost term produces a **byte-identical payload** with the flag on, at every `k` measured, on the real corpus. The widening is confined to the question class it names. --- ## 7. The second lock, isolated With the flag on, `k=12`, only the budget varied: | limit (B) | delivered | spent | priced table | |---|---|---|---| | 120 000 (default) | 11 | 82 399 | `over_budget_after_knapsack` | | 140 000 | 11 | 82 399 | `over_budget_after_knapsack` | | **160 000** | 12 | 150 249 | **delivered, rank 10** | | 200 000 | 12 | 150 249 | delivered, rank 10 | And the same sweep with the flag OFF: the priced table is `below_k` at every limit, because it never reaches the shortlist. **The two locks are independent and now separately measured.** Lock 1 is the vocabulary and the flag removes it. Lock 2 is that one 67 838 B excerpt is 56.5 % of a 120 000 B budget and the knapsack, maximising a sum of scores, prefers twenty small excerpts. Closing lock 2 is a second rule -- reserving budget for the top-ranked candidate, or sizing the budget to the corpus -- and this order allowed one. Consumption contract § 7.6 asked for exactly this number: "the corpus size at which its strategy stops fitting its budget". For this corpus it is not a size; it is a single document that costs more than half the budget. --- ## 8. Honesty limits - **One corpus, two questions, one control.** Generality is NOT demonstrated. The vocabulary is Norwegian, and a corpus in another language gets nothing from it. - **The list was written with both words visible.** `kost` and `pris` are the two words in the question and in the document that failed. The same disclosure the ranker already carries about `MIN_SHARED_PREFIX` applies here: the rule is not blind to the case that motivated it. - **The rank improvement is real and the delivery is not.** Every claim that the rule "finds" the document should be read against § 6: it ranks it 10th and the budget still refuses it. - **The eviction finding in § 3 is measured on one question.** That raising `k` can evict a gold concept is demonstrated for this pair of question and corpus, not proven as a general property of the DP. - **`--cost-vocabulary` has no consumer.** Nobody asked for it; it exists so the measurement above could be made against real code rather than a simulation, and so a decision about it can be made on numbers. --- ## 9. Recommendation 1. **Keep `DEFAULT_K = 8`.** The sweep shows raising `k` buys no answer for the mandate-shaped question and can evict the gold concept from the specific one. This is the opposite of what the order's option (c) expected, and it is measured. 2. **Keep `--cost-vocabulary` OFF by default.** It is a two-word bridge measured on one question; the number that would justify a default is a hit-rate over a question set nobody has built yet. 3. **The blind spot stays open, and it is a BUDGET question now, not a ranking question.** If it matters to a consumer, the next order is lock 2: reserve budget for the top-ranked candidate, or let a profile size its budget to its corpus. That is one rule, it has a clean red test (§ 3's eviction), and it is a decision about what a payload is for. 4. **A mandate is not a query, and no lexical ranker will make it one.** The upstream report's own arm reached this document in four navigational steps. The honest boundary of a declared cut is that it answers questions, and a mandate is a brief. Saying that in the skill's own words costs nothing and is more accurate than any `k`. --- ## 10. Verification log | # | Claim | Command → result | |---|---|---| | 1 | The token instrument reproduces a published figure | payload for the specific question, `k=8`, flag off → **164 987 B / 40 425 o200k tok**, equal to the 2026-09-07 published pair | | 2 | The ranker's known-positive still holds | same question, flag off, `k=8` → priced table delivered at **rank 1** | | 3 | The denominators close | 629 = 621 + 8, both questions | | 4 | The priced table's rank, mandate-shaped question | **249 of 269** lexical candidates; signals 1 and 3 both 0.0 | | 5 | `k` never delivers it | `k` in {8, 12, 16, 24, 32, 64, 128} → `below_k`; at 249 → `over_budget_after_knapsack` | | 6 | Raising `k` evicts the gold on the specific question | `k=16` and `k=24` → `over_budget_after_knapsack` | | 7 | Number density does not find it | three definitions → rank 178, 165, 46 of 269 | | 8 | Spread does not find it at these `k` | document rank **30 of 35** | | 9 | The rule moves it | candidate rank **249 → 10** | | 10 | The rule does not deliver it | every `k` in {8, 12, 16, 24} → withheld, by two different rules | | 11 | The gate is the question | control question, flag on vs off, every `k` → **byte-identical payload** | | 12 | The default does not move | flag off ⇒ payload byte-identical to `5a0c879`; both goldens unchanged | | 13 | Six mutations, six red | one-sided bridge, gate open, default on, member removed, every token bridges, member too short | | 14 | The measured bundle IS a HEAD build | `diff -r` fresh `5a0c879` rebuild against the measured tree → **exit 0, 0 lines**; contract check on its payload → exit 0 | | 15 | Suite, types, lint | `pytest -q` **1268 passed**; `mypy --strict src/ tools/` 28 files clean; `ruff check` + `ruff format --check` clean |