docs(measure): K3 Arm D beside a re-rated Arm B -- reach and no threshold

The first rater's row moves 8/4/0/0 to 4/5/0/3; the two blind rows are
identical at 5/1/0/0, and the disagreement sits on exactly the three
positions labelled correct. Both are published, and the verdict is a
conditional on what counts as one unit of knowledge, not a number.

Controls: src/ untouched, flag-absent and --outline-run 0 byte-identical
over the corpus, the consumer bundle unchanged at 1108 files and
9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1,
door-level counts 39/43 unchanged, pre-gate totals 144 and 23/39 as
declared. No threshold is set.
This commit is contained in:
Kjell Tore Guttormsen 2026-09-07 03:04:01 +02:00
commit 190086fc3d

335
docs/2026-09-07-k3-arm-d.md Normal file
View file

@ -0,0 +1,335 @@
# K3 with Arm D beside a re-rated Arm B, 2026-09-07
Two numbers on the same footing, so a threshold can be set afterwards. **No
threshold is set here**, and none is implied: the K3 method
(`docs/2026-09-02-k3-k4-k5-metode.md`) declares none, and inventing one inside
the work that produces a measurement is fitting the bar to the number. This
round ran under order `20260906T213322Z-1044411564-from-.claude`, which refuses
a threshold, refuses a change to the consumer bundle, and permits no model call
in the run path. All three refusals held and each is checked below.
Counts only. The corpus is public procurement material, but nothing here needs a
document body to be checkable.
## Arm D is not defined upstream of this document
`docs/2026-09-02-k3-k4-k5-metode.md` contains **zero** occurrences of the word
"arm" (`grep -c -i "arm"` -> `0`, exit 1). Arm D is a name this repository's
brief gives to one rule, so that a measurement can refer to it:
> read the document's own numbered outline -- the integer chapter headings
> (`N`, `N.`, `N)`) the shipping grammar cannot match, because its `_NUMBERED`
> pattern requires at least one dot -- and admit a boundary only where the
> integers form a **maximal ascending run of length >= 3**, taking the **last**
> such run when the outline repeats, because a contents listing precedes the
> body it lists.
The run length **3 is declared, not swept**. It follows from the corpus's own
distribution of maximal ascending runs (328 of length 1, 37 of length 2, 18 of
length 3 or more), and a sweep over candidate lengths would be choosing the
threshold from the answer.
## The question
K3 asks whether each concept carries **one unit of knowledge** (OKF v0.2 section
2). Four categories, exactly one per document, tie-break coarse before fine
before duplicate: **too coarse / too fine / duplicate / correct**.
## Method
- **Corpus:** `~/corpora/okf-telling-20260829/K2/trinn1`, **N = 43** files, of
which **39/43** are extractable. The other four are `.smc`, `.zip`, a PDF with
no text layer, and a `.doc`.
- **Sample:** **n = 12**, drawn by the method's own rule -- hex SHA-256 of the
**NFC-normalised** filename, stratified by format (8 `pdf`, 3 `docx`,
1 `xlsx`). The draw is now committed in `tools/okf_outline_measure.py` and
re-derived this round rather than copied; it reproduced the twelve published
documents **in order, 12/12**. Under NFD the draw yields a different sample,
so the normalisation is load-bearing.
- **Two arms, one round:** Arm B is the shipping default, **re-rated this
round** rather than carried over. Arm D is the same proposer with
`--outline-run 3`.
- **Raters:** one first-rater identity over both arms, then **two separate blind
raters, one per arm**, `n_blind = 6` each at canonical positions 0, 2, 4, 6,
8, 10 -- **12 blind ratings and two `k/6` figures**. Neither blind rater saw
the other's arm, either first-rater's labels, this report, or the plan.
- **Arm C's `8/4/0/0` is historical context and explicitly not a comparand:** it
was measured in a different round against a different baseline artifact.
## Controls, passed before anything was counted
| control | result |
|---|---|
| `git diff --stat 798f64a..HEAD -- src/` | **empty** -- the library was not touched |
| flag absent vs `--outline-run 0`, whole corpus | **byte-identical**, 28/28 artifacts, exit distribution 28/11/4 both |
| flag-off re-run vs the 28 archived Arm B plans | **byte-identical** (`diff -r`, exit 0) |
| consumer bundle `K2-bundle-20260903`, before and after | **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` -- unchanged |
| door-level counts, Arm D run | extracted **39/43**, gated **39/43**, persisted **39/43** -- unchanged |
| K1b conservation | `merged + coded rejections = 43; N = 43` |
| network imports in either tool | **0** (`grep -cE "anthropic\|openai\|requests\|httpx\|urllib\.request"`) |
| `Claude-Session:` trailers in this round's commits | **0** |
| declared pre-gate totals | **144** boundaries and **23/39** documents, exactly as declared |
The last row is the gate that permitted the rest: a pre-gate total other than
144/23 would mean the implemented rule is not the measured one, and the bundle
build would not have been started.
## The denominator
Every figure below is stated against one of three denominators, and they are not
interchangeable:
- **43** -- corpus files (the door-level denominator);
- **39** -- extractable files (the segmentation denominator);
- **12** -- the K3 sample, of which **at most 8 can move** (below).
## The ceiling: at most 8 of 12
Positions 0, 5, 10 and 11 carry **zero** outline boundaries, so they are the
same proposal in both arms. Measured directly on the plan entries, with
`ingested_at` excluded because the two runs carry different `--proposed-at`:
| pos | document | Arm B | Arm D | entries identical |
|---|---|---|---|---|
| 0 | Bilag 9.1 | no plan | no plan | both absent |
| 5 | Vedlegg 3 | 21 | 21 | **True** |
| 10 | Vedlegg 1 | 15 | 15 | **True** |
| 11 | Dokument for avtaleinngaelse | 2 | 2 | **True** |
*(exploratory -- this identity check is not emitted by a committed instrument.)*
Any reading of the row starts here: a row that moved by four moved four of the
eight it could.
## K3, the two rows side by side
### First rater, n = 12
| arm | too coarse | too fine | duplicate | correct | sum |
|---|---|---|---|---|---|
| Arm B (re-rated this round) | 8 | 4 | 0 | 0 | 12 |
| **Arm D** | **4** | **5** | **0** | **3** | **12** |
| *Arm C, 2026-09-04, historical only* | *8* | *4* | *0* | *0* | *12* |
Four labels changed, all among the eight that could: positions 2, 6 and 8 moved
coarse -> correct, and position 9 moved coarse -> **fine**.
### Blind raters, n = 6 each, one per arm
| arm | rater | too coarse | too fine | duplicate | correct | sum |
|---|---|---|---|---|---|---|
| Arm B | `blind-rater-A` | 5 | 1 | 0 | 0 | 6 |
| Arm D | `blind-rater-D` | 5 | 1 | 0 | 0 | 6 |
**The two blind rows are identical.** Agreement with the first rater, on the
same six positions:
| arm | agreement |
|---|---|
| Arm B | **6/6** |
| Arm D | **3/6** |
## Verdict
**On the first rater's row, Arm D is the first arm to move the number. On the
blind raters' rows, it moved it by zero.** Both statements are measurements of
the same twelve documents, and the report refuses to publish only the first.
The disagreement is not scattered. It sits on **exactly** the three positions
where the first rater wrote `correct` -- 2, 6 and 8 -- and the blind rater wrote
`too coarse` on all three, for one consistent reason: the arm cuts at the
document's **top-level** chapters, and the blind rater judged that the chapters
still fuse their own numbered subsections. Its evidence is concrete rather than
stylistic: at position 8, `Bilag 3.4` lists about 20 second-level sections and
deeper (down to `5.2.1.1.`), and the proposal emits exactly its 8 top-level
chapters, so `Spesielle rom` (6 832 chars) carries `7.1.` through `7.5.` whole.
So the honest form of the finding is a conditional, not a number:
- **If** a top-level chapter counts as one unit of knowledge, Arm D moves K3
from 8/4/0/0 to 4/5/0/3.
- **If** the unit is the numbered subsection, Arm D moves K3 by nothing, and
what it changes is which fusion you get, not whether you get one.
Nothing in the K3 method decides between those two readings, and this round does
not decide it either. That is the operator's call, and it is a **prior**
question to any threshold: a threshold on an undecided unit measures the rater.
The one place both readings agree is criterion 7's case, position 6's
`Planlagt situasjon`: under Arm B that chapter was absorbed into a neighbour and
did not exist as a concept; under Arm D it exists (539 chars, `rule:outline`).
The blind rater still labelled the document `too coarse`, on a *different* span
(`Overvannslosning`, 4 780 chars, fusing two site solutions). The specific
defect the arm was built to fix was fixed; the document did not become correct.
**Criterion 7: PASS**, checked with a command rather than prose --
`any(e['title'] == 'Planlagt situasjon' ...)` over `34.json` -> `True`.
## What did move, with denominators
| figure | Arm B | Arm D | denominator |
|---|---|---|---|
| outline boundaries proposed (pre-gate) | -- | **144** | -- |
| boundaries surviving the orphan gate | -- | **95** | of 144 |
| documents reached (pre-gate) | -- | **23** | of 39 |
| documents reached (post-gate) | -- | **21** | of 39 |
| entries, whole corpus | 618 | **709** | delta **+91** |
| existing Arm B candidates deleted | -- | **4** | all in `Bilag 3.6` |
| documents producing an artifact | 28 | **33** | of 43 |
| documents with nothing to propose | 11 | **6** | of 43 |
| documents with zero entries | 11 | **6** | of 39 |
| unique concept paths | -- | **709** | of 709 entries |
| bundle files | 1108 | **1294** | delta +186 |
| bundle `index.md` files | 478 | **578** | delta +100 |
| proposal wall time | 762 s | **769 s** | 43 documents |
| bundle build wall time | -- | **781.69 s** reported, **1558 s** end to end | 43 documents |
**709 unique paths out of 709 entries**: no collision, so the +91 entries are 91
distinct concepts and not a renaming of existing ones. This was emitted **before**
the bundle was built, which is the point -- a collision found afterwards would be
a fact about the writer, not about the rule.
Span sizes, Arm D: 709 spans, min 10, p50 447, p95 5 848, max 148 051; **185 of
709** are under 200 chars.
**Outline titles carrying no alphabetic word: 11 of 95.** The instrument's own
definition, stated because it is not an upstream term: a word is
`[^\W\d_]{2,}` -- two or more Unicode letters -- so a title made of digits and
single letters (`477 3 025`, `D 1 L`) counts as junk. An ad-hoc count written
during this session with a one-letter threshold gives **4** instead; the
committed instrument's 11 is the figure of record, and the discrepancy is a
difference of definition, not of data.
**Concept paths for unchanged content did not churn.** Of the **569** Arm B
entries whose span survives unchanged into Arm D, **0** received a different
concept path. The plan carried this as a medium risk on the grounds that
`_segment_path`'s `taken` set is order-dependent; on the delivered artifacts the
risk did not fire. *(exploratory -- not emitted by a committed instrument.)*
**0 of 95** post-gate outline titles reduce to the reserved stem `index`.
*(exploratory.)*
## What this does not measure
**The orphan gate deletes 34 % of the arm's own boundaries, and it deletes them
systematically skewed.** 49 of 144 boundaries fall to the parent-span check: a
chapter heading followed immediately by its own `x.y` subsection has an empty
body and is dropped. So the arm keeps `Vedlegg`, `Referanser` and `Innledning`
and loses the chapters that **have** structure beneath them. **What was rated is
therefore the outline rule minus its structurally richest third.** Without this
sentence the row above reads as evidence about "the outline rule" when it is
evidence about a degraded variant of it. The narrower fix (deduplicating
coincident boundaries at insertion) and the wider one (bounding a span to the
next same-or-higher-level heading) were both considered; the wider one is
excluded here under one-change-per-measurement, which is the Arm C lesson. A fix
exists; it is not that none was found.
**An ascending integer run is not the same thing as a chapter outline**, and two
of the twelve show it directly. At position 4 (`Vedlegg 5`) the run the rule
found is the **cited regulation's subsections** -- `1)`, `3)`, `4)`, `5)` -- so
three concepts are 131-394-char statute quotes and the fourth swallows 21 197
chars, 93.2 % of the document, under a subsection's title. At position 9
(`Bilag 5`) the run is a **numbered risk table** whose rows the PDF extractor
flattened into prose, so four rows of one risk assessment became four concepts.
The pre-work control that found "0 of 144 boundaries land on a table row" is not
contradicted by this: it tested markdown table rows (`|`-delimited, 57 of 35 050
lines), and a table geometry flattened into numbered prose is invisible to that
test. The control was right about its own definition and its definition was too
narrow. That is a limit of the control, stated here rather than left implicit.
**Position 0 is unreadable, and no segmentation changes that.** `Bilag 9.1`
extracts as **95.1 %** `(cid:N)` glyph tokens (206 758 of 217 470 chars), because
every embedded font is `/Type3` with no `/ToUnicode` map. Both blind raters
reached that independently. Its `too coarse` label rests on document extent and
the PDF's own bookmark outline -- which names two merged constituent documents --
not on reading the text. It is an **extraction** defect and K3 measures
segmentation, so it did not change a label; but a concept ingested from that file
today would carry almost no readable text however it were cut.
**Every JSON proposal leaves the document's head text uncovered** -- 135 to
3 773 chars of cover page, contents listing, and in two cases the body
`Innledning`. Measured by the Arm D blind rater across all five of its JSON
proposals, recorded here because it is real, and not used as a label: omission
is not one of the four categories.
**All raters are instances of the same model family.** The first rater and both
blind raters are Claude Opus 5. Agreement between them is not independent
confirmation in the sense a human panel would provide; it bounds
self-consistency, not correctness. The first rater had additionally seen the
published Arm C row before rating Arm B, so the Arm B row's reproduction of
`8/4/0/0` is **not** independent confirmation either. What the re-rating does
establish is narrower and sufficient for this comparison: both arms were judged
in the same round, by the same identity, against the same four categories.
**No CHANGELOG entry accompanies this arm.** Measured precedent rather than
preference: `--max-segment-chars` and "Arm C" appear **0** times in
`CHANGELOG.md`, while `--path-prefix` has an entry at `:40-44`. The precedent is
"interface and behaviour changes yes, arm flags no", and `--outline-run` is an
arm flag that defaults to off.
## Reproducing
```
C=~/corpora/okf-telling-20260829
# the flag-off identity half (byte-compare against the archive AND against the
# no-flag run; both were checked)
Z="$C/K2-plans-zero-20260907"; mkdir -p "$Z"; i=0
for f in "$C"/K2/trinn1/*; do
i=$((i+1)); b=$(basename "$f")
.venv/bin/python tools/okf_propose_segments.py "$f" \
--out "$Z/$(printf '%02d' $i).json" \
--path-prefix "${b%.*}" --proposed-at 2026-09-03T00:00:00Z --outline-run 0
done
diff -q -r "$C/K2-plans-baseline-20260907" "$Z" -x '*.err' -x '_index.txt'
# Arm D, into a FRESH dated directory -- never a reused one, because a leftover
# plan matching on source_sha256 would be replayed silently
D="$C/K2-plans-armd-20260907"; mkdir -p "$D"; i=0
for f in "$C"/K2/trinn1/*; do
i=$((i+1)); b=$(basename "$f")
.venv/bin/python tools/okf_propose_segments.py "$f" \
--out "$D/$(printf '%02d' $i).json" \
--path-prefix "${b%.*}" --proposed-at 2026-09-07T00:00:00Z --outline-run 3
done
.venv/bin/python tools/okf_outline_measure.py \
--corpus "$C/K2/trinn1" --report "$C/K2-outline-reach-20260907.md"
.venv/bin/python tools/okf_corpus_run.py \
--corpus "$C/K2/trinn1" \
--report "$C/K2-bundle-armd-20260907-report.md" \
--bundle "$C/K2-bundle-armd-20260907" --plans-dir "$D" \
--bundle-id k2-trinn1-armd-20260907 --okf-version 0.2 \
--ingested-at 2026-09-07T00:00:00Z
cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \
| xargs shasum -a 256 | shasum -a 256
```
The consumer bundle, locale-pinned, before and after this round:
- **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1`
`LC_ALL=C` is not decoration: without it `sort` orders the file list differently
and the aggregate digest changes while the bytes do not.
## Appendix: the twelve first-rater verdicts, both arms
| pos | document | Arm B entries | Arm D entries | Arm B | Arm D |
|---|---|---|---|---|---|
| 0 | Bilag 9.1 | 1 concept | 1 concept | too coarse | too coarse |
| 1 | Bilag 3.2.2 | 20 | 23 | too coarse | too coarse |
| 2 | Bilag 1.1 | 1 concept | 9 | too coarse | **correct** |
| 3 | Bilag 7 | 1 | 3 | too coarse | too coarse |
| 4 | Vedlegg 5 | 1 concept | 4 | too coarse | too coarse |
| 5 | Vedlegg 3 | 21 | 21 | too fine | too fine |
| 6 | Bilag 3.8 | 6 | 7 | too coarse | **correct** |
| 7 | Bilag 1.3 | 45 | 48 | too fine | too fine |
| 8 | Bilag 3.4 | 1 concept | 8 | too coarse | **correct** |
| 9 | Bilag 5 | 5 | 11 | too coarse | **too fine** |
| 10 | Vedlegg 1 | 15 | 15 | too fine | too fine |
| 11 | Dokument for avtaleinngaelse | 2 | 2 | too fine | too fine |
The blind raters covered positions 0, 2, 4, 6, 8, 10 only, and disagreed with
the first rater at 2, 6 and 8 on Arm D -- the three bolded `correct` labels --
and nowhere on Arm B.