# K3 with Arm D beside a re-rated Arm B, 2026-09-07 Two numbers on the same footing, so a threshold can be set afterwards. **No threshold is set here**, and none is implied: the K3 method (`docs/2026-09-02-k3-k4-k5-metode.md`) declares none, and inventing one inside the work that produces a measurement is fitting the bar to the number. This round ran under order `20260906T213322Z-1044411564-from-.claude`, which refuses a threshold, refuses a change to the consumer bundle, and permits no model call in the run path. All three refusals held and each is checked below. Counts only. The corpus is public procurement material, but nothing here needs a document body to be checkable. ## Arm D is not defined upstream of this document `docs/2026-09-02-k3-k4-k5-metode.md` contains **zero** occurrences of the word "arm" (`grep -c -i "arm"` -> `0`, exit 1). Arm D is a name this repository's brief gives to one rule, so that a measurement can refer to it: > read the document's own numbered outline -- the integer chapter headings > (`N`, `N.`, `N)`) the shipping grammar cannot match, because its `_NUMBERED` > pattern requires at least one dot -- and admit a boundary only where the > integers form a **maximal ascending run of length >= 3**, taking the **last** > such run when the outline repeats, because a contents listing precedes the > body it lists. The run length **3 is declared, not swept**. It follows from the corpus's own distribution of maximal ascending runs (328 of length 1, 37 of length 2, 18 of length 3 or more), and a sweep over candidate lengths would be choosing the threshold from the answer. ## The question K3 asks whether each concept carries **one unit of knowledge** (OKF v0.2 section 2). Four categories, exactly one per document, tie-break coarse before fine before duplicate: **too coarse / too fine / duplicate / correct**. ## Method - **Corpus:** `~/corpora/okf-telling-20260829/K2/trinn1`, **N = 43** files, of which **39/43** are extractable. The other four are `.smc`, `.zip`, a PDF with no text layer, and a `.doc`. - **Sample:** **n = 12**, drawn by the method's own rule -- hex SHA-256 of the **NFC-normalised** filename, stratified by format (8 `pdf`, 3 `docx`, 1 `xlsx`). The draw is now committed in `tools/okf_outline_measure.py` and re-derived this round rather than copied; it reproduced the twelve published documents **in order, 12/12**. Under NFD the draw yields a different sample, so the normalisation is load-bearing. - **Two arms, one round:** Arm B is the shipping default, **re-rated this round** rather than carried over. Arm D is the same proposer with `--outline-run 3`. - **Raters:** one first-rater identity over both arms, then **two separate blind raters, one per arm**, `n_blind = 6` each at canonical positions 0, 2, 4, 6, 8, 10 -- **12 blind ratings and two `k/6` figures**. Neither blind rater saw the other's arm, either first-rater's labels, this report, or the plan. - **Arm C's `8/4/0/0` is historical context and explicitly not a comparand:** it was measured in a different round against a different baseline artifact. ## Controls, passed before anything was counted | control | result | |---|---| | `git diff --stat 798f64a..HEAD -- src/` | **empty** -- the library was not touched | | flag absent vs `--outline-run 0`, whole corpus | **byte-identical**, 28/28 artifacts, exit distribution 28/11/4 both | | flag-off re-run vs the 28 archived Arm B plans | **byte-identical** (`diff -r`, exit 0) | | consumer bundle `K2-bundle-20260903`, before and after | **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` -- unchanged | | door-level counts, Arm D run | extracted **39/43**, gated **39/43**, persisted **39/43** -- unchanged | | K1b conservation | `merged + coded rejections = 43; N = 43` | | network imports in either tool | **0** (`grep -cE "anthropic\|openai\|requests\|httpx\|urllib\.request"`) | | `Claude-Session:` trailers in this round's commits | **0** | | declared pre-gate totals | **144** boundaries and **23/39** documents, exactly as declared | The last row is the gate that permitted the rest: a pre-gate total other than 144/23 would mean the implemented rule is not the measured one, and the bundle build would not have been started. ## The denominator Every figure below is stated against one of three denominators, and they are not interchangeable: - **43** -- corpus files (the door-level denominator); - **39** -- extractable files (the segmentation denominator); - **12** -- the K3 sample, of which **at most 8 can move** (below). ## The ceiling: at most 8 of 12 Positions 0, 5, 10 and 11 carry **zero** outline boundaries, so they are the same proposal in both arms. Measured directly on the plan entries, with `ingested_at` excluded because the two runs carry different `--proposed-at`: | pos | document | Arm B | Arm D | entries identical | |---|---|---|---|---| | 0 | Bilag 9.1 | no plan | no plan | both absent | | 5 | Vedlegg 3 | 21 | 21 | **True** | | 10 | Vedlegg 1 | 15 | 15 | **True** | | 11 | Dokument for avtaleinngaelse | 2 | 2 | **True** | *(exploratory -- this identity check is not emitted by a committed instrument.)* Any reading of the row starts here: a row that moved by four moved four of the eight it could. ## K3, the two rows side by side ### First rater, n = 12 | arm | too coarse | too fine | duplicate | correct | sum | |---|---|---|---|---|---| | Arm B (re-rated this round) | 8 | 4 | 0 | 0 | 12 | | **Arm D** | **4** | **5** | **0** | **3** | **12** | | *Arm C, 2026-09-04, historical only* | *8* | *4* | *0* | *0* | *12* | Four labels changed, all among the eight that could: positions 2, 6 and 8 moved coarse -> correct, and position 9 moved coarse -> **fine**. ### Blind raters, n = 6 each, one per arm | arm | rater | too coarse | too fine | duplicate | correct | sum | |---|---|---|---|---|---|---| | Arm B | `blind-rater-A` | 5 | 1 | 0 | 0 | 6 | | Arm D | `blind-rater-D` | 5 | 1 | 0 | 0 | 6 | **The two blind rows are identical.** Agreement with the first rater, on the same six positions: | arm | agreement | |---|---| | Arm B | **6/6** | | Arm D | **3/6** | ## Verdict **On the first rater's row, Arm D is the first arm to move the number. On the blind raters' rows, it moved it by zero.** Both statements are measurements of the same twelve documents, and the report refuses to publish only the first. The disagreement is not scattered. It sits on **exactly** the three positions where the first rater wrote `correct` -- 2, 6 and 8 -- and the blind rater wrote `too coarse` on all three, for one consistent reason: the arm cuts at the document's **top-level** chapters, and the blind rater judged that the chapters still fuse their own numbered subsections. Its evidence is concrete rather than stylistic: at position 8, `Bilag 3.4` lists about 20 second-level sections and deeper (down to `5.2.1.1.`), and the proposal emits exactly its 8 top-level chapters, so `Spesielle rom` (6 832 chars) carries `7.1.` through `7.5.` whole. So the honest form of the finding is a conditional, not a number: - **If** a top-level chapter counts as one unit of knowledge, Arm D moves K3 from 8/4/0/0 to 4/5/0/3. - **If** the unit is the numbered subsection, Arm D moves K3 by nothing, and what it changes is which fusion you get, not whether you get one. Nothing in the K3 method decides between those two readings, and this round does not decide it either. That is the operator's call, and it is a **prior** question to any threshold: a threshold on an undecided unit measures the rater. The one place both readings agree is criterion 7's case, position 6's `Planlagt situasjon`: under Arm B that chapter was absorbed into a neighbour and did not exist as a concept; under Arm D it exists (539 chars, `rule:outline`). The blind rater still labelled the document `too coarse`, on a *different* span (`Overvannslosning`, 4 780 chars, fusing two site solutions). The specific defect the arm was built to fix was fixed; the document did not become correct. **Criterion 7: PASS**, checked with a command rather than prose -- `any(e['title'] == 'Planlagt situasjon' ...)` over `34.json` -> `True`. ## What did move, with denominators | figure | Arm B | Arm D | denominator | |---|---|---|---| | outline boundaries proposed (pre-gate) | -- | **144** | -- | | boundaries surviving the orphan gate | -- | **95** | of 144 | | documents reached (pre-gate) | -- | **23** | of 39 | | documents reached (post-gate) | -- | **21** | of 39 | | entries, whole corpus | 618 | **709** | delta **+91** | | existing Arm B candidates deleted | -- | **4** | all in `Bilag 3.6` | | documents producing an artifact | 28 | **33** | of 43 | | documents with nothing to propose | 11 | **6** | of 43 | | documents with zero entries | 11 | **6** | of 39 | | unique concept paths | -- | **709** | of 709 entries | | bundle files | 1108 | **1294** | delta +186 | | bundle `index.md` files | 478 | **578** | delta +100 | | proposal wall time | 762 s | **769 s** | 43 documents | | bundle build wall time | -- | **781.69 s** reported, **1558 s** end to end | 43 documents | **709 unique paths out of 709 entries**: no collision, so the +91 entries are 91 distinct concepts and not a renaming of existing ones. This was emitted **before** the bundle was built, which is the point -- a collision found afterwards would be a fact about the writer, not about the rule. Span sizes, Arm D: 709 spans, min 10, p50 447, p95 5 848, max 148 051; **185 of 709** are under 200 chars. **Outline titles carrying no alphabetic word: 11 of 95.** The instrument's own definition, stated because it is not an upstream term: a word is `[^\W\d_]{2,}` -- two or more Unicode letters -- so a title made of digits and single letters (`477 3 025`, `D 1 L`) counts as junk. An ad-hoc count written during this session with a one-letter threshold gives **4** instead; the committed instrument's 11 is the figure of record, and the discrepancy is a difference of definition, not of data. **Concept paths for unchanged content did not churn.** Of the **569** Arm B entries whose span survives unchanged into Arm D, **0** received a different concept path. The plan carried this as a medium risk on the grounds that `_segment_path`'s `taken` set is order-dependent; on the delivered artifacts the risk did not fire. *(exploratory -- not emitted by a committed instrument.)* **0 of 95** post-gate outline titles reduce to the reserved stem `index`. *(exploratory.)* ## What this does not measure **The orphan gate deletes 34 % of the arm's own boundaries, and it deletes them systematically skewed.** 49 of 144 boundaries fall to the parent-span check: a chapter heading followed immediately by its own `x.y` subsection has an empty body and is dropped. So the arm keeps `Vedlegg`, `Referanser` and `Innledning` and loses the chapters that **have** structure beneath them. **What was rated is therefore the outline rule minus its structurally richest third.** Without this sentence the row above reads as evidence about "the outline rule" when it is evidence about a degraded variant of it. The narrower fix (deduplicating coincident boundaries at insertion) and the wider one (bounding a span to the next same-or-higher-level heading) were both considered; the wider one is excluded here under one-change-per-measurement, which is the Arm C lesson. A fix exists; it is not that none was found. **An ascending integer run is not the same thing as a chapter outline**, and two of the twelve show it directly. At position 4 (`Vedlegg 5`) the run the rule found is the **cited regulation's subsections** -- `1)`, `3)`, `4)`, `5)` -- so three concepts are 131-394-char statute quotes and the fourth swallows 21 197 chars, 93.2 % of the document, under a subsection's title. At position 9 (`Bilag 5`) the run is a **numbered risk table** whose rows the PDF extractor flattened into prose, so four rows of one risk assessment became four concepts. The pre-work control that found "0 of 144 boundaries land on a table row" is not contradicted by this: it tested markdown table rows (`|`-delimited, 57 of 35 050 lines), and a table geometry flattened into numbered prose is invisible to that test. The control was right about its own definition and its definition was too narrow. That is a limit of the control, stated here rather than left implicit. **Position 0 is unreadable, and no segmentation changes that.** `Bilag 9.1` extracts as **95.1 %** `(cid:N)` glyph tokens (206 758 of 217 470 chars), because every embedded font is `/Type3` with no `/ToUnicode` map. Both blind raters reached that independently. Its `too coarse` label rests on document extent and the PDF's own bookmark outline -- which names two merged constituent documents -- not on reading the text. It is an **extraction** defect and K3 measures segmentation, so it did not change a label; but a concept ingested from that file today would carry almost no readable text however it were cut. **Every JSON proposal leaves the document's head text uncovered** -- 135 to 3 773 chars of cover page, contents listing, and in two cases the body `Innledning`. Measured by the Arm D blind rater across all five of its JSON proposals, recorded here because it is real, and not used as a label: omission is not one of the four categories. **All raters are instances of the same model family.** The first rater and both blind raters are Claude Opus 5. Agreement between them is not independent confirmation in the sense a human panel would provide; it bounds self-consistency, not correctness. The first rater had additionally seen the published Arm C row before rating Arm B, so the Arm B row's reproduction of `8/4/0/0` is **not** independent confirmation either. What the re-rating does establish is narrower and sufficient for this comparison: both arms were judged in the same round, by the same identity, against the same four categories. **No CHANGELOG entry accompanies this arm.** Measured precedent rather than preference: `--max-segment-chars` and "Arm C" appear **0** times in `CHANGELOG.md`, while `--path-prefix` has an entry at `:40-44`. The precedent is "interface and behaviour changes yes, arm flags no", and `--outline-run` is an arm flag that defaults to off. ## Reproducing ``` C=~/corpora/okf-telling-20260829 # the flag-off identity half (byte-compare against the archive AND against the # no-flag run; both were checked) Z="$C/K2-plans-zero-20260907"; mkdir -p "$Z"; i=0 for f in "$C"/K2/trinn1/*; do i=$((i+1)); b=$(basename "$f") .venv/bin/python tools/okf_propose_segments.py "$f" \ --out "$Z/$(printf '%02d' $i).json" \ --path-prefix "${b%.*}" --proposed-at 2026-09-03T00:00:00Z --outline-run 0 done diff -q -r "$C/K2-plans-baseline-20260907" "$Z" -x '*.err' -x '_index.txt' # Arm D, into a FRESH dated directory -- never a reused one, because a leftover # plan matching on source_sha256 would be replayed silently D="$C/K2-plans-armd-20260907"; mkdir -p "$D"; i=0 for f in "$C"/K2/trinn1/*; do i=$((i+1)); b=$(basename "$f") .venv/bin/python tools/okf_propose_segments.py "$f" \ --out "$D/$(printf '%02d' $i).json" \ --path-prefix "${b%.*}" --proposed-at 2026-09-07T00:00:00Z --outline-run 3 done .venv/bin/python tools/okf_outline_measure.py \ --corpus "$C/K2/trinn1" --report "$C/K2-outline-reach-20260907.md" .venv/bin/python tools/okf_corpus_run.py \ --corpus "$C/K2/trinn1" \ --report "$C/K2-bundle-armd-20260907-report.md" \ --bundle "$C/K2-bundle-armd-20260907" --plans-dir "$D" \ --bundle-id k2-trinn1-armd-20260907 --okf-version 0.2 \ --ingested-at 2026-09-07T00:00:00Z cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \ | xargs shasum -a 256 | shasum -a 256 ``` The consumer bundle, locale-pinned, before and after this round: - **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` `LC_ALL=C` is not decoration: without it `sort` orders the file list differently and the aggregate digest changes while the bytes do not. ## Appendix: the twelve first-rater verdicts, both arms | pos | document | Arm B entries | Arm D entries | Arm B | Arm D | |---|---|---|---|---|---| | 0 | Bilag 9.1 | 1 concept | 1 concept | too coarse | too coarse | | 1 | Bilag 3.2.2 | 20 | 23 | too coarse | too coarse | | 2 | Bilag 1.1 | 1 concept | 9 | too coarse | **correct** | | 3 | Bilag 7 | 1 | 3 | too coarse | too coarse | | 4 | Vedlegg 5 | 1 concept | 4 | too coarse | too coarse | | 5 | Vedlegg 3 | 21 | 21 | too fine | too fine | | 6 | Bilag 3.8 | 6 | 7 | too coarse | **correct** | | 7 | Bilag 1.3 | 45 | 48 | too fine | too fine | | 8 | Bilag 3.4 | 1 concept | 8 | too coarse | **correct** | | 9 | Bilag 5 | 5 | 11 | too coarse | **too fine** | | 10 | Vedlegg 1 | 15 | 15 | too fine | too fine | | 11 | Dokument for avtaleinngaelse | 2 | 2 | too fine | too fine | The blind raters covered positions 0, 2, 4, 6, 8, 10 only, and disagreed with the first rater at 2, 6 and 8 on Arm D -- the three bolded `correct` labels -- and nowhere on Arm B.