K3 round 3, order 20260908T170037Z-3622420612-from-.claude, plus the PM injection of 17:55Z carrying the operator's D3 answer. TWO NEW RULES, both OFF by default. `--sheet-section-rows` cuts an open table block at the rows that label its sections: a RUN of at least three rows whose first cell is a bare numeric label. It is the first rule here that opens a span INSIDE a table, and the opposite direction from Arm E -- that arm decides how far a block extends, this decides where it is cut inside. The spreadsheet is the one file type no arm had ever moved: 1 concept -> 12 on the priced sheet, of which 11 are the reference's 11 cost groups. The run guard buys exactly three candidates on the corpus (the rows stating a computation basis), and each section carries its own `source_sheet` + `source_rows`, verified on the artifact. `--drop-wrapped-outline` declines an outline candidate whose line continues onto the next one. Round 2 falsified the LENGTH axis on this case; this one is not about size. Over the 12-document sample it separates 8 of 34 candidates -- the quoted regulation paragraphs and the risk-table rows -- and none of the 26 the operator kept. Position 4: 4 concepts -> 1, the reference. The cell is `treffer` by count and lossy by bytes, and the report says so. Whole corpus, one cached extraction: 1 of 39 documents changes under the first rule (0 of 32 pdf, 0 of 5 docx), 5 of 39 under the second (all pdf). Arm B and F2 are byte-identical against a frozenc3b645b, both halves by `diff -r`. THE `okf build` DEFAULT MOVED, and this is the operator's decision executed, not a side effect: no flag now means `--outline-run 3 --unit-fold`. Each arm keeps an explicit opt-out (`--outline-run 0`, `--no-unit-fold`) and the pair reproduces the pre-move bytes exactly. The proposer's own defaults do NOT move: they are what the goldens and every published reproduction block are pinned to, so the two layers now disagree on purpose and `cli.py` says where. Two shipped expectations moved with it and both are stated in place. MEASURED AND REPORTED BACK: the configuration made default scores 2 of 12 on the unit worksheet, `docx` 0 of 3 -- Arm F's published 5 of 12 was measured with `--table-grid` ON, and without it the fold has no joined table to fold. Position 1 is declined for the second round on the axis the order named: 3 of 3 recovered chapters have body under them, so "heading-on-heading" separates 0 of 3. The K2 ranking control round 2 could not measure: two bundles from frozenc3b645bdiffering only in `--keep-table-heading`, 2 of 1108 files apart. The priced concept ranks 96 of 629 in BOTH and is delivered at rank 10 in both, so the flag buys 35 payload bytes and zero rank positions. Round 2's prediction is falsified: the concept's NAME was already restored by the orphan-name inheritance; the flag adds the heading LINE. The stale S7 literal is restated where it stood (`c759a657...`, 171 614 B, re-measured on this HEAD), so the next order cannot inherit the superseded one. 1397 -> 1414 tests. Report: docs/2026-09-08-k3-runde3-per-filtype.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
294 lines
15 KiB
Markdown
294 lines
15 KiB
Markdown
# The rarity weight: what it delivered, what it cannot reach, 2026-09-08
|
|
|
|
The previous session made a requirement number VISIBLE to the consumption
|
|
pre-pass and could not make it WORTH more than a common word: `_overlap`
|
|
returns a count, so `10.2-2` scored exactly what `krever` scored, and the gold
|
|
concept stayed `below_k` on all three bundles
|
|
(`docs/2026-09-08-kravnummer-tokenisering.md` § 8). This session weights each
|
|
hit by the token's rarity in the bundle and measures what that buys.
|
|
|
|
It buys one of the three requirement lookups and a large move on K2. It does
|
|
not buy the other two, and the reason is structural rather than a matter of
|
|
tuning. Both halves are below, with the decomposition.
|
|
|
|
---
|
|
|
|
## 0. What IS measured, and what is NOT
|
|
|
|
**Measured.** Nine known-positive figures reproduced before anything was
|
|
changed; the document frequency of every question token on all three bundles,
|
|
under the matcher and under exact equality; the gold's fused rank and its
|
|
per-signal ranks, before and after, on three bundles; the delivered/withheld
|
|
verdict and payload digest per bundle; every K2 control the order named, both
|
|
values; the six hit@8 questions; the `df` pass's cost; eight mutations.
|
|
|
|
**Known-positive, run first.** Every figure this session builds on was
|
|
reproduced on this machine, with the committed tool, before any after-column
|
|
was read:
|
|
|
|
| # | known-positive | source | measured here |
|
|
|---|---|---|---|
|
|
| 1 | gold fused rank per bundle | previous session § 4 | **96, 9, 35** |
|
|
| 2 | concepts sharing the gold's lexical score on N500 | previous session § 8 | **70 of 270** |
|
|
| 3 | considered / withheld / delivered | consumer § 3 | 446/438/8, 1133/1125/8, 270/262/8 |
|
|
| 4 | gold position in `withheld` | consumer § 3.1 | 119 of 438, 204 of 1125, 66 of 262 |
|
|
| 5 | K2 S7 control, two commands | previous session § 5 | `2ae46f68…` 169 573 B, `8aaf8558…` 172 588 B |
|
|
| 6 | priced-sheet candidate rank, four rows, O1 bundle | O1 § 7 | 251 / 19 / 1 / 1 |
|
|
| 7 | the same four rows on the rebuilt bundle | previous session § 7 | 251 / **10** / 1 / 1 |
|
|
| 8 | K2 delivery costs, four commands | previous session § 7 | 169 583 / 183 798 / 190 438 / 108 877 B |
|
|
| 9 | hit@8 over the six published questions | 2026-09-07 § hit@k | **5 of 6** |
|
|
|
|
> **Superseded literal (added 2026-09-08, round 3).** The S7 control figures in
|
|
> this document were correct when it was written and are not the ones to reuse:
|
|
> `a37d5ce`, `17c49fc` and `c95d189` each widened the excerpt form afterwards.
|
|
> The current known-positive for the S7 control question under the default
|
|
> command is **`c759a657dbb41d71`, 171 614 B**, measured on `c3b645b` against
|
|
> the same bundle. `docs/2026-09-08-k3-runde3-per-filtype.md` § 4.
|
|
|
|
**NOT measured.** Two things, both named rather than glossed:
|
|
|
|
- **`o200k_base` token counts.** No `o200k` encoder is installed on this
|
|
machine, and this repository's single-runtime-dependency rule does not admit
|
|
one to produce a caption. Every figure below is therefore in the unit the
|
|
budget gate itself uses — `BUDGET_UNIT`, "utf-8 bytes of emitted JSON" — plus
|
|
the payload `sha256`, which is exact where a byte count is only indicative.
|
|
The previously published o200k figures are not restated as if re-measured.
|
|
- **Whether a live model answers better.** This measures a *ranking*.
|
|
"Delivered" is not "read"; that measurement belongs to the consumer.
|
|
|
|
---
|
|
|
|
## 1. The rule, and why it has no constant in it
|
|
|
|
`rarity_weights(question_tokens, corpus)` returns, per question token,
|
|
|
|
```
|
|
w(t) = log(N / df(t))
|
|
```
|
|
|
|
`N` is the number of concepts in the bundle. `df(t)` is how many of them bear
|
|
the token, counted with the **same** four-character prefix rule a hit is scored
|
|
with. Nothing is set by hand: a word every concept carries weighs `log(1) = 0`
|
|
of itself, and an identifier one concept carries takes the corpus's maximum of
|
|
itself. No token class is declared anywhere.
|
|
|
|
Three decisions inside that sentence, each of them a decision:
|
|
|
|
1. **`df` is counted over the text a hit can be scored on** — title, id and
|
|
body joined, the two lexical signals' union. Counting rarity over one field
|
|
and matching on another would weight a token by how rare it is somewhere it
|
|
is never read.
|
|
2. **The weight enters the RANKING and never the GATE.** `lexical` stays a
|
|
count when weights are supplied, at the price of one more pass over the same
|
|
two fields. This is load-bearing: `krav` weighs exactly 0 on all three
|
|
bundles, so a concept matching only `krav` would have a weighted sum of 0.0
|
|
and would fall to `no_lexical_match` — turning a ranking change into the
|
|
*gate* that `54a0bc2` falsified.
|
|
3. **One `df` table per bundle, reaching every stage that scores the question
|
|
against text**, including the document prior. A weight applied to two of
|
|
three fused signals would be a second ranker rather than one statistic.
|
|
|
|
**Why this is not the sweep `54a0bc2` refuted.** That order asked for smoothed
|
|
IDF as a **port**: a threshold below which a concept is withheld, required to
|
|
zero both known-negative questions. The sweep found the frontier empty — `neg2`
|
|
first reaches 0 at `T = 4.6`, by which point gold-hit reachability has already
|
|
fallen to 0 of 6 — and the order was returned rather than implemented. That
|
|
result stands and is not re-litigated. This is the other use: an ordering
|
|
*inside* the candidate set, with the gate untouched. A ranking cannot withhold
|
|
anything, so the failure mode that refuted the gate has no counterpart here.
|
|
|
|
**The pollution `54a0bc2` § 1 named is inherited deliberately**, and § 3 below
|
|
shows it doing exactly the damage that section predicted.
|
|
|
|
**Cost.** One pass. Measured: **0.062 s** over 446 concepts, **0.241 s** over
|
|
1 133, **0.073 s** over 270 — inside the order's one-second bound on the
|
|
largest bundle by a factor of four.
|
|
|
|
---
|
|
|
|
## 2. The three requirement lookups, before and after
|
|
|
|
Same instrument before and after: the gold's position in the fused `ranked`
|
|
list, and the payload's own verdict from `build_payload`. Command:
|
|
|
|
```
|
|
python3 tools/okf_consume.py <bundle> \
|
|
--question "Hva krever Krav <n> i <normal>? Gjengi det sentrale vilkåret." \
|
|
[--rarity-weight]
|
|
```
|
|
|
|
| | N100:2023 | N200:2024 | N500:2024 |
|
|
|---|---|---|---|
|
|
| concepts | 446 | 1 133 | 270 |
|
|
| gold fused rank, weight **off** | 96 | 9 | 35 |
|
|
| gold fused rank, weight **on** | **103** | **8** | **35** |
|
|
| gold verdict, off | `below_k` | `below_k` | `below_k` |
|
|
| gold verdict, on | `below_k` | **delivered, rank 8** | `below_k` |
|
|
| payload bytes, off → on | 61 566 → 61 452 | 142 861 → **146 641** | 40 626 → 38 869 |
|
|
| payload `sha256`, off | `49e9c718…` | `adc2eac7…` | `21c5bbfe…` |
|
|
| payload `sha256`, on | `ba09a56a…` | `403ee989…` | `1758f6ca…` |
|
|
| `okf_contract_check` | exit 0, 14 rules, 0 findings | idem | idem |
|
|
|
|
**One of three delivers.** N200's gold moves from withheld position 204 of
|
|
1 125 to delivered at rank 8. N500 does not move at all. N100 moves the wrong
|
|
way, by seven positions.
|
|
|
|
---
|
|
|
|
## 3. Why N100 got worse: the matcher makes the identifier common
|
|
|
|
`df` under the matcher, against `df` under exact token equality:
|
|
|
|
| bundle | identifier | `df` under the matcher | `df` exact | `w(identifier)` | `w('sentrale')` |
|
|
|---|---|---|---|---|---|
|
|
| N100 | `3.3.1-13` | **135 of 446** | 1 | 1.195 | **3.798** |
|
|
| N200 | `2.9.2-12` | 17 of 1 133 | 1 | 4.199 | 4.835 |
|
|
| N500 | `10.2-2` | 5 of 270 | 1 | 3.989 | 3.519 |
|
|
|
|
The identifier is unique in all three bundles. It is *common* in N100 only
|
|
because `tokens_match` compares four leading characters, so `3.3.1-13` matches
|
|
every requirement number beginning `3.3.` — 135 of them. The weight then does
|
|
what it was asked to do and ranks a common adjective (`sentrale`, borne by 10
|
|
of 446) as more informative than the exact requirement number.
|
|
|
|
This is `54a0bc2` § 1 — "`df` measured over the colliding matcher measures
|
|
collision breadth, not rarity" — measured on the identifier itself rather than
|
|
on `brann*`. It is not a defect in the weight; it is the weight faithfully
|
|
reporting a `df` the matcher produced. Fixing it means changing the matcher,
|
|
which is a different change with its own measurement, and it is not attempted
|
|
here.
|
|
|
|
---
|
|
|
|
## 4. Why N500 did not move: RRF consumes ranks, and the gold already led
|
|
|
|
The gold's rank on each of the three fused signals, and how many concepts each
|
|
signal reorders:
|
|
|
|
| bundle | signal ranks, off | signal ranks, on | positions moved per signal |
|
|
|---|---|---|---|
|
|
| N100 | 39 / 186 / 127 | 42 / 184 / 127 | 122 / 367 / **0** of 446 |
|
|
| N200 | 8 / 4 / 212 | **4** / 10 / 212 | 37 / 277 / **0** of 1 133 |
|
|
| N500 | **1** / 137 / 74 | **1** / 137 / 74 | **0** / 127 / **0** of 270 |
|
|
|
|
On N500 the gold is **already rank 1** on the title-and-id signal, by counting,
|
|
before any weighting. Weighting raises its score from 3.000 to 3.989 and its
|
|
rank from 1 to 1. RRF was chosen precisely because it **consumes ranks only**
|
|
(`RRF_K`'s comment says so), which makes it invariant to any monotone transform
|
|
of a signal — so the whole increase is invisible to the fusion. What holds the
|
|
gold at 35 is the other two signals, and neither can see the identifier at all:
|
|
the body repeats `krav` and nothing else, and the document prior is a
|
|
two-document density whose order the weight never changes on any of the three
|
|
bundles (0 positions moved, three times).
|
|
|
|
**The general statement, and it is a limit rather than a complaint:** a rarity
|
|
weight can only move a fused rank where it reorders some signal. Where the
|
|
identifier appears in exactly one field and already leads that field, no
|
|
weighting scheme inside the signals can do anything. Closing N500 means
|
|
changing the fusion — giving a signal a weight, or letting a strong single
|
|
signal carry more than 1/(60+1) — and that is a different mechanism from this
|
|
one. It is named here and not built.
|
|
|
|
The same shape is in the test suite, twice, so neither half can quietly stop
|
|
being true: a synthetic corpus where the identifier is in the title alone
|
|
(fused rank 18 → 18, title-signal rank 1 → 1) and one where weighting reorders
|
|
the title signal from last to first (30 → 1) and moves the fused rank only from
|
|
30 to 18.
|
|
|
|
---
|
|
|
|
## 5. K2: every control holds, and the priced sheet moves 10 → 2
|
|
|
|
| control | weight off | weight on | verdict |
|
|
|---|---|---|---|
|
|
| Q-good (`Hvordan skal prisene fylles ut?`), candidate rank | 1 | **1** | holds |
|
|
| Q-good with `--cost-vocabulary` | 1 | **1** | holds |
|
|
| priced sheet, mandate question, vocabulary off | 251 of 271 | **78 of 271** | improves |
|
|
| priced sheet, `--cost-vocabulary --k 12` | 10 of 280 | **2 of 280** | improves |
|
|
| priced sheet, delivered by `--cost-vocabulary --k 12` | rank 10 | **rank 2** | improves |
|
|
| hit@8 over the six published questions | 5 of 6 | **5 of 6**, every rank identical | holds |
|
|
| S7 control question, default command | `2ae46f68…`, 169 573 B | **byte-identical** | holds |
|
|
| S7 control question, `--cost-vocabulary --k 12 --limit 160000` | `8aaf8558…`, 172 588 B | `b205a618…`, 183 462 B | **moves, +10 874 B** |
|
|
|
|
No named K2 control falls. The one figure that moves is a payload the flag
|
|
changes on purpose: with the weight on, that command delivers a different set
|
|
of excerpts and costs 6.3 % more bytes.
|
|
|
|
---
|
|
|
|
## 6. The decision, made by the number
|
|
|
|
The order's rule: default ON if the weight wins on all four corpora with no K2
|
|
control falling; behind a flag otherwise, naming what fell and by how much.
|
|
|
|
**It does not win on all four. N100's gold fused rank moves 96 → 103, a loss of
|
|
seven positions**, on a question where the gold is not delivered either way. So
|
|
the weight ships as `--rarity-weight`, **off by default**.
|
|
|
|
Off means the bytes that were already published, and that is measured rather
|
|
than asserted. Eight payloads — three N-bundle requirement lookups, three
|
|
commands on the rebuilt K2 bundle, two on the O1 bundle — were built from a
|
|
frozen copy of `56c1205` (`git archive` into a scratch tree, `PYTHONPATH` and
|
|
`__file__` both checked) and from this working tree with the flag omitted:
|
|
|
|
**8 of 8 payload digests identical.** No consumer's bytes move because this
|
|
change exists.
|
|
|
|
What the flag buys the caller who sets it, in one line each: N200's requirement
|
|
lookup goes from withheld to delivered; K2's priced sheet goes from candidate
|
|
rank 10 to 2; N100 and N500 are unchanged in verdict and worse and equal in
|
|
rank respectively.
|
|
|
|
---
|
|
|
|
## 7. What did NOT close
|
|
|
|
- **Two of the three requirement lookups are still `below_k`**, and the target
|
|
the order set — rank ≤ 8 on all three — is **not met**. It is met on N200.
|
|
- **N500 cannot be closed by any weighting inside the signals** (§ 4). The next
|
|
mechanism is a change to the fusion, not to the weight, and it has no
|
|
measurement behind it yet.
|
|
- **N100 needs the matcher, not the weight** (§ 3). `MIN_SHARED_PREFIX = 4`
|
|
makes a unique identifier read as 135-of-446 common. That is the same defect
|
|
`54a0bc2` § 1 named a year of sessions ago in another form, and it is now
|
|
measured on the case it costs most.
|
|
- **The 34 % skewed deletion in the orphan gate** is untouched; different axis.
|
|
- **"Delivered" is not "read".** N200's gold is now in the payload. Whether a
|
|
model then answers the engineer's question is the consumer's measurement.
|
|
|
|
---
|
|
|
|
## 8. Tests, mutations, gates
|
|
|
|
Thirteen new tests, red before the rule existed (6 of them failed on
|
|
`AttributeError` alone, and the known-positive — the gold held at rank 18 of 30
|
|
by counting — passed from the start, so the red was the rule's absence and not
|
|
a broken fixture).
|
|
|
|
**Eight mutations, eight red.** Two survived the first pass and neither was
|
|
banded blind:
|
|
|
|
- **M5** (the document prior stops receiving the weights) survived because on
|
|
all three real bundles it is behaviourally equivalent — the two-document
|
|
density is never reordered. The commitment "one statistic reaches every
|
|
stage" is real even where no output distinguishes it, so it is now asserted
|
|
at the call site instead of on an output.
|
|
- **M7** (the title signal stops being weighted) survived because the fixture
|
|
put the identifier in the body as well. The real corpus does not: the gold's
|
|
title answers three question tokens and its body answers one. The fixture was
|
|
wrong, was corrected to the measured shape — and the corrected fixture then
|
|
reproduced the real failure of § 4, which is how § 4 was found at all.
|
|
|
|
| gate | result |
|
|
|---|---|
|
|
| `pytest -q` | **1 308 passed** (1 295 before) |
|
|
| `ruff check src tests tools` | clean |
|
|
| `ruff format --check` | clean, 88 files |
|
|
| `mypy --strict src/` | clean, 17 files |
|
|
| `okf_contract_check` on the three weighted payloads | exit 0, 14 rules, 0 findings |
|
|
| `git diff --stat` | `tools/okf_consume.py`, `tests/test_okf_consume.py` |
|
|
|
|
**No rebuild.** The change is consumption-side only, confirmed by
|
|
`git diff --stat`: no production module is touched, so the K2 bundle ref
|
|
`sha256-tree:2f82fcfea91c3bd3f8ef7147f80cd613227d3ca7975c41d88810233f3f79ab4b`
|
|
is unchanged.
|