# K3 with Arm E beside Arm D and a re-rated Arm B, 2026-09-07 Three numbers on the same footing, so a threshold can be set afterwards. **No threshold is set here**, and none is implied: the K3 method (`docs/2026-09-02-k3-k4-k5-metode.md`) declares none, and inventing one inside the work that produces a measurement is fitting the bar to the number. This round ran under order `20260907T075834Z-18584396-from-.claude`, which refuses a threshold, refuses a change to the consumer bundle, permits no model call in the run path, and forbids a push. All four refusals held and each is checked below. One thing IS declared before the row is read, and it is not a threshold: the **direction** that counts as movement -- fewer `too fine` WITHOUT more `too coarse`. It sets no value any count must reach. It is written down in advance precisely so it cannot be chosen after the number is known. Counts only. The corpus is public procurement material, but nothing here needs a document body to be checkable. ## Arm E is not defined upstream of this document `docs/2026-09-02-k3-k4-k5-metode.md` contains **zero** occurrences of the word "arm" (`grep -c -i "arm"` -> `0`, exit 1). Arm E is a name this repository's brief gives to one rule, so that a measurement can refer to it: > a pandoc GRID-table rule line -- `+---+---+`, and `+===+===+` under a header > -- does not close an open table block. A block is marked as JOINED only when > a rule line was actually crossed between two table rows, never merely because > its span contains one, so a single-row grid table stays byte-identical to > Arm D. **The rule has no numeric parameter, so nothing was swept and nothing could be.** What is declared instead is the rule line's character class, `[-=:+]`, and it is measured rather than guessed: across the three grid-bearing documents of this corpus, **38 of 38** lines whose stripped form starts with `+` match the pattern, and those four characters are the complete set occurring on them. The `:` is pandoc's column-alignment marker and it is load bearing -- a first pass with `[-=+]` matched **37 of 38** and, through that single miss, read one document as having two tables where it has one. The 37 is recorded here rather than quietly corrected. ## The question K3 asks whether each concept carries **one unit of knowledge** (OKF v0.2 section 2). Four categories, exactly one per document, tie-break coarse before fine before duplicate: **too coarse / too fine / duplicate / correct**. ## Method - **Corpus:** `K2/trinn1`, **N = 43** files, of which **39/43** are extractable. - **Sample:** **n = 12**, drawn by the method's own rule -- hex SHA-256 of the **NFC-normalised** filename, stratified by format (8 `pdf`, 3 `docx`, 1 `xlsx`), re-derived this round from `tools/okf_outline_measure.py` and reproducing the twelve published documents in order, **12/12**. - **Three arms, one round:** Arm B is the shipping default (`--outline-run 0`), Arm D is `--outline-run 3`, Arm E is `--outline-run 3 --table-grid`. Arm B and Arm D are **re-rated** this round rather than carried over. - **Raters:** one first-rater identity over all three arms, **36 verdicts**; then **two blind raters per arm** at canonical positions 0, 2, 4, 6, 8, 10, `n_blind = 6` each -- **36 blind ratings and six `k/6` figures**. - **The blind protocol is WIDER than Arm D's and the two are not comparable.** `docs/2026-09-07-k3-arm-d.md` used "two separate blind raters, **one per arm**, 12 blind ratings and two `k/6` figures". The order asked for two per arm. Its number governs; the divergence is stated so the rounds' `k/6` figures are not read as like for like. - **Blindness is structural, not promised.** Each blind rater is a separate subagent with its own context, given ONE file: six documents labelled `A`-`F` with their extracted length and, per concept, its length, its title and the first 180 characters of its body. **Rule names were stripped**, because a `rule:table-grid` in the list would have identified the arm. No rater was told which arm it read, that other arms exist, what the first rater said, or that a brief, plan or report exists. - **Arm C's `8/4/0/0` is historical context and explicitly not a comparand.** ## Controls, passed before anything was counted The controls are split by what they do on failure, and that split is the correction this round makes to its own first plan. A **gating** control asks whether the shipped rule is the rule being measured; it halts. A **prediction** is a figure written down in advance from an exploratory replica; it is reported whatever it says, because a replica may not sit in judgement over shipped code. ### Gating -- each one halts the round | control | result | |---|---| | `git diff --stat 54a0bc2..HEAD -- src/` | only `propose.py` and `cli.py`; **136** changed lines in `propose.py`, most of them comments | | consumer bundle `K2-bundle-20260903`, file count | **1108** -- the literal published in the Arm D round | | the same bundle, `LC_ALL=C` aggregate digest | `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` -- the published literal, before and after | | network imports, proposer and both instruments | **0** each | | **Arm B identity**, whole corpus, no flags | **byte-identical** to the archived Arm B plans, `diff -r` exit **0**, `_index.txt` INCLUDED | | **Arm D identity**, whole corpus, `--outline-run 3` | **byte-identical** to the archived Arm D plans, `diff -r` exit **0**, `_index.txt` INCLUDED | | artifact counts asserted BEFORE each diff | Arm B **28** json + **43** `.err` + **43** index lines; Arm D **33** + **43** + **43** | | door-level counts | `.err` files recording `FAILED`: **4** of 43, so extractable **39/43**, unchanged | | declared grid-rule totals | **38** lines on **3** documents, exactly as declared | Two of these rows are corrections to the Arm D round's own published procedure, and both were found by review rather than by failure. The artifact counts are asserted **before** the diff, because `diff -r` over two trees where every document failed would compare nothing and exit 0. And `_index.txt` is **included** in the comparison: the Arm D reproduce block neither generates it nor compares it, which would let `NN.json` name a different document across two runs with nothing saying so. The Arm B identity run is not bookkeeping either. It is the corpus-level half of the promise that `--outline-run 0` with the flag absent is still Arm B, and it is what lets the Arm B row below rest on verified bytes. ### Predictions -- reported, never gating Written into the brief from an exploratory replica **before** the rule was built, and reproduced by the shipped code: | prediction | measured | |---|---| | exactly 3 of 33 plans differ from the Arm D archive | **3** -- and they are the three named | | position 5: 21 -> 6 entries | **21 -> 6** | | position 10: 15 -> 3 entries | **15 -> 3** | | position 11: 2 -> 1 entries | **2 -> 1** | **One prediction was allowed to halt, and only one:** if NO plan had differed, the flag would not have been threaded through to `find_candidates` and the row would have been a wiring bug wearing a null result's clothes -- with the ceiling below standing ready as a plausible wrong explanation. It did not occur. The plan-to-position mapping is **derived, not assumed**: each plan's first entry `path` prefix is matched against the reduced stem of its source filename and that against the committed draw. The `NN` in `NN.json` comes from an unsorted shell glob and names nothing on its own. ## The denominator Every figure below is stated against one of four denominators, and they are not interchangeable: - **43** -- corpus files (the door-level denominator); - **39** -- extractable files (the segmentation denominator); - **12** -- the K3 sample, of which **at most 3 can move** (below); - **6** -- the blind positions, of which **1** is a document Arm E can move. ## The ceiling: at most 3 of 12, and this time it is measured The Arm D round's ceiling was read off entry counts. That is not sound on its own: the orphan check can delete a table candidate before it becomes an entry, so a document could hold table rows that never reach a plan -- and position 0 produces no plan at all, so its zero would be an absence with no denominator. This round measures the ceiling in the text itself, with a committed instrument, over all 43 files and **before any rating began**: | figure | value | denominator | |---|---|---| | documents with at least one table row | **3** | 39 | | table rows in total | **57** | -- | | documents with at least one grid-rule line | **3** | 39 | | grid-rule lines in total | **38** | -- | A document with no table row cannot be moved by this arm, whatever the orphan check later did to its candidates. So **3 of the 12** sample documents can move, and they are positions 5, 10 and 11 -- three of the four the Arm D ceiling excluded, because they carry zero outline boundaries. **The 57 is an independent corroboration and worth stating as one.** The Arm D round's pre-work control counted `|`-delimited table rows over the corpus text it screened and reported **57 of 35 050 lines**. This round's instrument, run against different code for a different purpose, counts **57**. Neither measurement was derived from the other. ## K3, the three rows side by side ### First rater, n = 12 | arm | too coarse | too fine | duplicate | correct | sum | |---|---|---|---|---|---| | Arm B (re-rated this round) | 8 | 4 | 0 | 0 | 12 | | Arm D (re-rated this round) | 4 | 5 | 0 | 3 | 12 | | **Arm E** | **4** | **3** | **0** | **5** | **12** | | *Arm C, 2026-09-04, historical only* | *8* | *4* | *0* | *0* | *12* | Two labels changed, both among the three that could: positions 10 and 11 moved `too fine` -> `correct`. Position 5 did not move. ### Blind raters, n = 6 each, two per arm | arm | rater | too coarse | too fine | duplicate | correct | sum | agreement with first rater | |---|---|---|---|---|---|---|---| | Arm B | `blind-1a` | 4 | 1 | 0 | 1 | 6 | **5/6** | | Arm B | `blind-1b` | 4 | 1 | 0 | 1 | 6 | **5/6** | | Arm D | `blind-2a` | 2 | 1 | 0 | 3 | 6 | **6/6** | | Arm D | `blind-2b` | 3 | 1 | 0 | 2 | 6 | **5/6** | | Arm E | `blind-3a` | 2 | 1 | 0 | 3 | 6 | **5/6** | | Arm E | `blind-3b` | 2 | 1 | 0 | 3 | 6 | **5/6** | Within-arm agreement, which exists for the first time because there are two raters per arm: Arm B **6/6**, Arm D **5/6**, Arm E **6/6**. **`too fine` is 1 of 6 in every arm, Arm E included.** ## Verdict **On the first rater's row, Arm E moves `too fine` from 5 to 3 while `too coarse` stays at 4 -- the direction declared in advance. On the blind raters' rows, `too fine` does not move at all.** Both statements are measurements of the same twelve documents, and the report refuses to publish only the first. The disagreement is one position and it is legible. At position 10 the first rater moved `too fine` -> `correct`; both Arm E blind raters kept `too fine`. Their reason CHANGED rather than persisting. Under Arm B and Arm D they object that thirteen table rows are severed from their header row. Under Arm E, where the table is one concept with its header included, they object that the table is severed from the sentence that introduces it. Arm E fixed the first complaint and does not touch the second. So the honest form of the finding is a conditional, and it has two clauses: - **If** a table is one unit of knowledge, Arm E moves K3 from 4/5/0/3 to 4/3/0/5 and does it without trading a `too fine` for a `too coarse`. - **If** a table is a unit only together with the prose that introduces it, Arm E moves K3 by nothing on the position where both readings were tested, and what it changes is which severance you get, not whether you get one. **And the second clause is measured on ONE document.** The blind positions are fixed at 0, 2, 4, 6, 8, 10, and only position 10 is a document Arm E can move. Positions 5 and 11 -- the other two -- were seen by no blind rater. The blind row is therefore not evidence that Arm E fails on those two; it is evidence that this protocol cannot see them. A round in which the arm's reach and the blind protocol's positions overlap in one document is a round whose blind row carries one document's worth of information about the arm, and no threshold should be set on that. **The unit question Arm D surfaced is still open and is still the operator's.** It fired again here, on identical material: at position 8 one Arm D blind rater called top-level chapters `correct` and the other called them `too coarse` because their numbered subsections are "distinct requirement sets a reader would want separately". Nothing in the K3 method decides between those readings, this round does not decide it either, and it is prior to any threshold -- a threshold on an undecided unit measures the rater. ## What did move, with denominators | figure | Arm B | Arm D | Arm E | denominator | |---|---|---|---|---| | entries, whole corpus | 618 | 709 | **681** | delta **-28** from Arm D | | entries carrying `rule:table-block` | 33 | 33 | **5** | of 681 | | entries carrying `rule:table-grid` | -- | -- | **5** | of 681 | | documents whose entry count changed | -- | -- | **3** | of 39 | | plans differing from the Arm D archive | -- | -- | **3** | of 33 | | blocks joined | -- | -- | **5** | -- | | documents producing an artifact | 28 | 33 | **33** | of 43 | | extractable | 39 | 39 | **39** | of 43 | | position 5 entries | 21 | 21 | **6** | -- | | position 10 entries | 15 | 15 | **3** | -- | | position 11 entries | 2 | 2 | **1** | -- | | largest span Arm E creates | -- | -- | **13 691** chars | position 10 | **Concept paths for unchanged content did not churn, and the denominator is computed rather than declared.** Of the **676** `(plan, span)` pairs present in both Arm D and Arm E, **0** received a different concept path. An earlier draft of this round's plan declared 686 as the expected denominator; that was wrong twice over -- 28 entries are removed, not 33, and a joined block's `end` moves so its pair matches nothing in Arm D. The computed 676 is the figure of record, and the wrong 686 is recorded rather than deleted. ## What this does not measure **The character class was fitted to the same three documents it is measured on.** Arm D's run-length 3 came from a corpus-wide distribution of 328/37/18. Arm E's `[-=:+]` came from 38 lines drawn entirely from the three documents that are 100 % of its movable sample. The whole-corpus screen above tests generalisation outward -- it found no fourth grid-bearing document -- but it cannot break that circularity inward, and no reading of the rows should treat "declared, not swept" as meaning the same thing it meant for Arm D. **`rule:table-grid` is plan-level provenance and does NOT reach the bundle.** Measured on the 1 294-file Arm D bundle: `grep -rl "PROPOSED"` returns **0**, while `derived` appears in 84 files as a frontmatter key with other values. A reviewer looking for the rule name in a built bundle will find nothing, and that absence is a property of materialisation, not evidence that the flag did not fire. **The ceiling is bounded by which table FORM the converter chose, not by how many tables the corpus holds.** Pandoc also emits *simple* and *multiline* tables, whose rows carry no `|` at all. `_TABLE_ROW` never sees those, so they are invisible to the table rule, to Arm E, and to the `|`-row count that measures the ceiling. One document in this sample (position 10) contains such a table in its upper half, and no arm proposes a boundary in it. **Two grid tables separated by a rule line alone would merge into one concept.** Pandoc puts a blank line between adjacent tables, so it does not emit that shape -- but that is a property of the WRITER, not of this code, and `in_table` survives an arbitrary run of rule lines. A unit fixture asserts the merge, so the limit is declared rather than assumed away. The corpus diff found no instance. **Position 5 is the document that shows what Arm E is not.** It has the largest reduction in the round, 21 concepts to 6, and its LABEL DOES NOT CHANGE. Each of its three references is still cut into a 114-character title concept carrying a heading and no body, plus its 1 675-character table. Joining table rows removed most of the fragmentation and left the rest; the remaining cut comes from the heading rule, not from the table rule. **Position 7's cause is diagnosed and deliberately unbuilt.** Its 48 concepts include nine contents-listing lines with dotted leaders, and concepts of 87, 93 and 112 characters. It carries zero table-block entries, so Arm E cannot reach it. One change per measurement is the Arm C lesson; the fix is named and not made. **The orphan gate still deletes 34 % of Arm D's own boundaries, skewed.** Reported in the Arm D round, unfixed, and untouched here. **No bundle was built for Arm E.** The Arm D round's door-level counts came from a bundle run; here they come from the run artifacts themselves -- 4 of 43 `.err` files record `FAILED`, so extractable is 39/43 -- which is the same figure for roughly a twentieth of the wall time. Bundle-level file counts are therefore not reported for Arm E, and that is a gap, not a result. **Position 0 is unreadable and no segmentation changes that.** It extracts as 95.1 % `(cid:N)` glyph tokens. It is an extraction defect, K3 measures segmentation, and it did not move a label in any arm. **All raters are instances of the same model family, and the first rater is not independent.** Agreement between them bounds self-consistency, not correctness. The first rater had read the Arm D report's published rows before rating, so Arm B reproducing `8/4/0/0` and Arm D reproducing `4/5/0/3` is consistency and not confirmation. What the re-rating establishes is narrower and sufficient for this comparison: all three arms were judged in one round, by one identity, against the same four categories. **No CHANGELOG entry accompanies this arm.** Measured precedent, re-checked this round: `grep -c` for "outline-run", "max-segment-chars", "Arm C" and "Arm D" returns 0 in both `README.md` and `CHANGELOG.md`, while `--path-prefix` -- a real interface change -- has a CHANGELOG entry. The rule is "interface and behaviour changes yes, arm flags no", and `--table-grid` is an arm flag that defaults off. ## Reproducing ``` C=~/corpora/okf-telling-20260829 # The corpus loop, in one place. It writes NN.json, NN.err and one # `i|exit|filename` line per document into _index.txt -- which is the format # both archives carry, and which the Arm D round's published block omitted. arm_run() { # $1=outdir $2=proposed-at $3=lo $4=hi then flags DIR=$1; AT=$2; LO=$3; HI=$4; shift 4; mkdir -p "$DIR"; i=0 for f in "$C"/K2/trinn1/*; do i=$((i+1)) [ "$i" -lt "$LO" ] && continue [ "$i" -gt "$HI" ] && continue b=$(basename "$f"); n=$(printf '%02d' "$i") .venv/bin/python tools/okf_propose_segments.py "$f" --out "$DIR/$n.json" \ --path-prefix "${b%.*}" --proposed-at "$AT" "$@" 2> "$DIR/$n.err" echo "$i|$?|$b" >> "$DIR/_index.txt" done } # Run in ascending chunks, or _index.txt line order breaks. Each chunk is a # foreground call under 600 s; documents 13-28 account for most of the time. for lo_hi in "1 12" "13 20" "21 28" "29 36" "37 43"; do set -- $lo_hi arm_run "$C/K2-plans-armB-check-20260907" 2026-09-03T00:00:00Z "$1" "$2" arm_run "$C/K2-plans-armD-check-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3 arm_run "$C/K2-plans-armE-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3 --table-grid done # The identity halves. Assert the counts FIRST: a diff over two trees where # every document failed compares nothing and exits 0. ls "$C"/K2-plans-armB-check-20260907/*.json | wc -l # 28 ls "$C"/K2-plans-armD-check-20260907/*.json | wc -l # 33 diff -r "$C/K2-plans-baseline-20260907" "$C/K2-plans-armB-check-20260907" -x '*.err'; echo $? diff -r "$C/K2-plans-armd-20260907" "$C/K2-plans-armD-check-20260907" -x '*.err'; echo $? # The measurement: exactly three plans differ. diff -rq "$C/K2-plans-armd-20260907" "$C/K2-plans-armE-20260907" -x '*.err' -x '_index.txt' # The ceiling, over all 43 files. .venv/bin/python tools/okf_table_measure.py \ --corpus "$C/K2/trinn1" --report "$C/K2-table-reach-20260907.md" # The door count, without building a bundle. ls "$C"/K2-plans-armE-20260907/*.err | wc -l # 43 grep -l FAILED "$C"/K2-plans-armE-20260907/*.err | wc -l # 4 -> 39/43 cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \ | xargs shasum -a 256 | shasum -a 256 ``` The consumer bundle, locale-pinned, before and after this round: - **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` `LC_ALL=C` is not decoration: without it `sort` orders the file list differently and the aggregate digest changes while the bytes do not. ## Appendix: the twelve first-rater verdicts, three arms | pos | document | B entries | D entries | E entries | Arm B | Arm D | Arm E | |---|---|---|---|---|---|---|---| | 0 | Bilag 9.1 | 1 concept | 1 concept | 1 concept | too coarse | too coarse | too coarse | | 1 | Bilag 3.2.2 | 20 | 23 | 23 | too coarse | too coarse | too coarse | | 2 | Bilag 1.1 | 1 concept | 9 | 9 | too coarse | correct | correct | | 3 | Bilag 7 | 1 | 3 | 3 | too coarse | too coarse | too coarse | | 4 | Vedlegg 5 | 1 concept | 4 | 4 | too coarse | too coarse | too coarse | | 5 | Vedlegg 3 | 21 | 21 | **6** | too fine | too fine | too fine | | 6 | Bilag 3.8 | 6 | 7 | 7 | too coarse | correct | correct | | 7 | Bilag 1.3 | 45 | 48 | 48 | too fine | too fine | too fine | | 8 | Bilag 3.4 | 1 concept | 8 | 8 | too coarse | correct | correct | | 9 | Bilag 5 | 5 | 11 | 11 | too coarse | too fine | too fine | | 10 | Vedlegg 1 | 15 | 15 | **3** | too fine | too fine | **correct** | | 11 | Dokument for avtaleinngaelse | 2 | 2 | **1** | too fine | too fine | **correct** | "1 concept" means no plan was written: the mechanical rules found no boundary and the document lands as one flat concept. The blind raters covered positions 0, 2, 4, 6, 8, 10 only. They disagreed with the first rater at position 6 on Arm B (both raters, `correct` where the first rater says `too coarse` -- his ground is a chapter that Arm B ABSORBS and that therefore does not appear in the material a blind rater sees), at position 8 on Arm D (one rater of two), and at position 10 on Arm E (both raters, `too fine` where the first rater says `correct`). They agreed with the first rater and with each other everywhere else.