llm-ingestion-okf/docs/2026-09-07-k3-arm-e.md
Kjell Tore Guttormsen a7ee942d5f docs(measure): K3 Arm E beside Arm B and Arm D -- three rows, no threshold
Three rows rated in one round by one first-rater identity, then two blind
raters per arm as the order asked -- 36 first-rater verdicts, 36 blind ratings,
six k/6 figures. No threshold is set and none is implied.

First rater: Arm B 8/4/0/0, Arm D 4/5/0/3, Arm E 4/3/0/5. Against the direction
declared before the row was seen -- fewer `too fine` WITHOUT more `too coarse`
-- `too fine` goes 5 -> 3 and `too coarse` stays 4.

Blind raters: `too fine` is 1 of 6 in EVERY arm, Arm E included. The report
publishes both rows and states the verdict as a conditional, because they are
measurements of the same twelve documents and only one of them is favourable.

The disagreement is one position and it is legible. At position 10 both Arm E
blind raters kept `too fine` where the first rater moved to `correct`, and
their REASON changed rather than persisting: under Arm B and D they object that
thirteen table rows are severed from their header; under Arm E, where the table
is one concept with its header, they object that the table is severed from the
sentence that introduces it. Arm E fixed the first complaint and does not touch
the second.

And the denominator that governs how much that can say: the blind positions are
fixed at 0, 2, 4, 6, 8, 10, and only ONE of Arm E's three movable documents
falls in that set. The blind row carries one document's worth of information
about this arm. No threshold should be set on that, and the report says so.

The ceiling is measured this round rather than inferred. A committed instrument
extracted all 43 files and counted table ROWS in the text -- 3 of 39 documents,
57 rows -- because zero ENTRIES is not the same claim: the orphan check can
delete a table candidate before it becomes an entry, and one sample document
produces no plan at all so its zero would be an absence with no denominator.
The 57 reproduces the Arm D round's independent pre-work screen exactly, with
neither measurement derived from the other.

Controls are split by what they do on failure, which is this round's correction
to its own first plan: gating controls halt, predictions are reported whatever
they say. Only one prediction was allowed to halt -- a zero-diff run, which
would be a wiring bug wearing a null result's clothes, with the ceiling standing
ready as a plausible wrong explanation.

Nine limits are stated rather than left implicit, including three that weaken
the arm's own case: the character class was fitted to the same three documents
it is measured on; position 5 has the round's largest reduction (21 concepts to
6) and its label does not change; and pandoc's simple and multiline table forms
carry no pipes at all, so the ceiling is bounded by which form the converter
chose rather than by how many tables the corpus holds.

Pure ASCII (0 non-ASCII bytes), short document labels only (0 corpus filenames).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 11:48:58 +02:00

22 KiB

K3 with Arm E beside Arm D and a re-rated Arm B, 2026-09-07

Three numbers on the same footing, so a threshold can be set afterwards. No threshold is set here, and none is implied: the K3 method (docs/2026-09-02-k3-k4-k5-metode.md) declares none, and inventing one inside the work that produces a measurement is fitting the bar to the number. This round ran under order 20260907T075834Z-18584396-from-.claude, which refuses a threshold, refuses a change to the consumer bundle, permits no model call in the run path, and forbids a push. All four refusals held and each is checked below.

One thing IS declared before the row is read, and it is not a threshold: the direction that counts as movement -- fewer too fine WITHOUT more too coarse. It sets no value any count must reach. It is written down in advance precisely so it cannot be chosen after the number is known.

Counts only. The corpus is public procurement material, but nothing here needs a document body to be checkable.

Arm E is not defined upstream of this document

docs/2026-09-02-k3-k4-k5-metode.md contains zero occurrences of the word "arm" (grep -c -i "arm" -> 0, exit 1). Arm E is a name this repository's brief gives to one rule, so that a measurement can refer to it:

a pandoc GRID-table rule line -- +---+---+, and +===+===+ under a header -- does not close an open table block. A block is marked as JOINED only when a rule line was actually crossed between two table rows, never merely because its span contains one, so a single-row grid table stays byte-identical to Arm D.

The rule has no numeric parameter, so nothing was swept and nothing could be. What is declared instead is the rule line's character class, [-=:+], and it is measured rather than guessed: across the three grid-bearing documents of this corpus, 38 of 38 lines whose stripped form starts with + match the pattern, and those four characters are the complete set occurring on them. The : is pandoc's column-alignment marker and it is load bearing -- a first pass with [-=+] matched 37 of 38 and, through that single miss, read one document as having two tables where it has one. The 37 is recorded here rather than quietly corrected.

The question

K3 asks whether each concept carries one unit of knowledge (OKF v0.2 section 2). Four categories, exactly one per document, tie-break coarse before fine before duplicate: too coarse / too fine / duplicate / correct.

Method

  • Corpus: K2/trinn1, N = 43 files, of which 39/43 are extractable.
  • Sample: n = 12, drawn by the method's own rule -- hex SHA-256 of the NFC-normalised filename, stratified by format (8 pdf, 3 docx, 1 xlsx), re-derived this round from tools/okf_outline_measure.py and reproducing the twelve published documents in order, 12/12.
  • Three arms, one round: Arm B is the shipping default (--outline-run 0), Arm D is --outline-run 3, Arm E is --outline-run 3 --table-grid. Arm B and Arm D are re-rated this round rather than carried over.
  • Raters: one first-rater identity over all three arms, 36 verdicts; then two blind raters per arm at canonical positions 0, 2, 4, 6, 8, 10, n_blind = 6 each -- 36 blind ratings and six k/6 figures.
  • The blind protocol is WIDER than Arm D's and the two are not comparable. docs/2026-09-07-k3-arm-d.md used "two separate blind raters, one per arm, 12 blind ratings and two k/6 figures". The order asked for two per arm. Its number governs; the divergence is stated so the rounds' k/6 figures are not read as like for like.
  • Blindness is structural, not promised. Each blind rater is a separate subagent with its own context, given ONE file: six documents labelled A-F with their extracted length and, per concept, its length, its title and the first 180 characters of its body. Rule names were stripped, because a rule:table-grid in the list would have identified the arm. No rater was told which arm it read, that other arms exist, what the first rater said, or that a brief, plan or report exists.
  • Arm C's 8/4/0/0 is historical context and explicitly not a comparand.

Controls, passed before anything was counted

The controls are split by what they do on failure, and that split is the correction this round makes to its own first plan. A gating control asks whether the shipped rule is the rule being measured; it halts. A prediction is a figure written down in advance from an exploratory replica; it is reported whatever it says, because a replica may not sit in judgement over shipped code.

Gating -- each one halts the round

control result
git diff --stat 54a0bc2..HEAD -- src/ only propose.py and cli.py; 136 changed lines in propose.py, most of them comments
consumer bundle K2-bundle-20260903, file count 1108 -- the literal published in the Arm D round
the same bundle, LC_ALL=C aggregate digest 9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1 -- the published literal, before and after
network imports, proposer and both instruments 0 each
Arm B identity, whole corpus, no flags byte-identical to the archived Arm B plans, diff -r exit 0, _index.txt INCLUDED
Arm D identity, whole corpus, --outline-run 3 byte-identical to the archived Arm D plans, diff -r exit 0, _index.txt INCLUDED
artifact counts asserted BEFORE each diff Arm B 28 json + 43 .err + 43 index lines; Arm D 33 + 43 + 43
door-level counts .err files recording FAILED: 4 of 43, so extractable 39/43, unchanged
declared grid-rule totals 38 lines on 3 documents, exactly as declared

Two of these rows are corrections to the Arm D round's own published procedure, and both were found by review rather than by failure. The artifact counts are asserted before the diff, because diff -r over two trees where every document failed would compare nothing and exit 0. And _index.txt is included in the comparison: the Arm D reproduce block neither generates it nor compares it, which would let NN.json name a different document across two runs with nothing saying so.

The Arm B identity run is not bookkeeping either. It is the corpus-level half of the promise that --outline-run 0 with the flag absent is still Arm B, and it is what lets the Arm B row below rest on verified bytes.

Predictions -- reported, never gating

Written into the brief from an exploratory replica before the rule was built, and reproduced by the shipped code:

prediction measured
exactly 3 of 33 plans differ from the Arm D archive 3 -- and they are the three named
position 5: 21 -> 6 entries 21 -> 6
position 10: 15 -> 3 entries 15 -> 3
position 11: 2 -> 1 entries 2 -> 1

One prediction was allowed to halt, and only one: if NO plan had differed, the flag would not have been threaded through to find_candidates and the row would have been a wiring bug wearing a null result's clothes -- with the ceiling below standing ready as a plausible wrong explanation. It did not occur.

The plan-to-position mapping is derived, not assumed: each plan's first entry path prefix is matched against the reduced stem of its source filename and that against the committed draw. The NN in NN.json comes from an unsorted shell glob and names nothing on its own.

The denominator

Every figure below is stated against one of four denominators, and they are not interchangeable:

  • 43 -- corpus files (the door-level denominator);
  • 39 -- extractable files (the segmentation denominator);
  • 12 -- the K3 sample, of which at most 3 can move (below);
  • 6 -- the blind positions, of which 1 is a document Arm E can move.

The ceiling: at most 3 of 12, and this time it is measured

The Arm D round's ceiling was read off entry counts. That is not sound on its own: the orphan check can delete a table candidate before it becomes an entry, so a document could hold table rows that never reach a plan -- and position 0 produces no plan at all, so its zero would be an absence with no denominator.

This round measures the ceiling in the text itself, with a committed instrument, over all 43 files and before any rating began:

figure value denominator
documents with at least one table row 3 39
table rows in total 57 --
documents with at least one grid-rule line 3 39
grid-rule lines in total 38 --

A document with no table row cannot be moved by this arm, whatever the orphan check later did to its candidates. So 3 of the 12 sample documents can move, and they are positions 5, 10 and 11 -- three of the four the Arm D ceiling excluded, because they carry zero outline boundaries.

The 57 is an independent corroboration and worth stating as one. The Arm D round's pre-work control counted |-delimited table rows over the corpus text it screened and reported 57 of 35 050 lines. This round's instrument, run against different code for a different purpose, counts 57. Neither measurement was derived from the other.

K3, the three rows side by side

First rater, n = 12

arm too coarse too fine duplicate correct sum
Arm B (re-rated this round) 8 4 0 0 12
Arm D (re-rated this round) 4 5 0 3 12
Arm E 4 3 0 5 12
Arm C, 2026-09-04, historical only 8 4 0 0 12

Two labels changed, both among the three that could: positions 10 and 11 moved too fine -> correct. Position 5 did not move.

Blind raters, n = 6 each, two per arm

arm rater too coarse too fine duplicate correct sum agreement with first rater
Arm B blind-1a 4 1 0 1 6 5/6
Arm B blind-1b 4 1 0 1 6 5/6
Arm D blind-2a 2 1 0 3 6 6/6
Arm D blind-2b 3 1 0 2 6 5/6
Arm E blind-3a 2 1 0 3 6 5/6
Arm E blind-3b 2 1 0 3 6 5/6

Within-arm agreement, which exists for the first time because there are two raters per arm: Arm B 6/6, Arm D 5/6, Arm E 6/6.

too fine is 1 of 6 in every arm, Arm E included.

Verdict

On the first rater's row, Arm E moves too fine from 5 to 3 while too coarse stays at 4 -- the direction declared in advance. On the blind raters' rows, too fine does not move at all. Both statements are measurements of the same twelve documents, and the report refuses to publish only the first.

The disagreement is one position and it is legible. At position 10 the first rater moved too fine -> correct; both Arm E blind raters kept too fine. Their reason CHANGED rather than persisting. Under Arm B and Arm D they object that thirteen table rows are severed from their header row. Under Arm E, where the table is one concept with its header included, they object that the table is severed from the sentence that introduces it. Arm E fixed the first complaint and does not touch the second.

So the honest form of the finding is a conditional, and it has two clauses:

  • If a table is one unit of knowledge, Arm E moves K3 from 4/5/0/3 to 4/3/0/5 and does it without trading a too fine for a too coarse.
  • If a table is a unit only together with the prose that introduces it, Arm E moves K3 by nothing on the position where both readings were tested, and what it changes is which severance you get, not whether you get one.

And the second clause is measured on ONE document. The blind positions are fixed at 0, 2, 4, 6, 8, 10, and only position 10 is a document Arm E can move. Positions 5 and 11 -- the other two -- were seen by no blind rater. The blind row is therefore not evidence that Arm E fails on those two; it is evidence that this protocol cannot see them. A round in which the arm's reach and the blind protocol's positions overlap in one document is a round whose blind row carries one document's worth of information about the arm, and no threshold should be set on that.

The unit question Arm D surfaced is still open and is still the operator's. It fired again here, on identical material: at position 8 one Arm D blind rater called top-level chapters correct and the other called them too coarse because their numbered subsections are "distinct requirement sets a reader would want separately". Nothing in the K3 method decides between those readings, this round does not decide it either, and it is prior to any threshold -- a threshold on an undecided unit measures the rater.

What did move, with denominators

figure Arm B Arm D Arm E denominator
entries, whole corpus 618 709 681 delta -28 from Arm D
entries carrying rule:table-block 33 33 5 of 681
entries carrying rule:table-grid -- -- 5 of 681
documents whose entry count changed -- -- 3 of 39
plans differing from the Arm D archive -- -- 3 of 33
blocks joined -- -- 5 --
documents producing an artifact 28 33 33 of 43
extractable 39 39 39 of 43
position 5 entries 21 21 6 --
position 10 entries 15 15 3 --
position 11 entries 2 2 1 --
largest span Arm E creates -- -- 13 691 chars position 10

Concept paths for unchanged content did not churn, and the denominator is computed rather than declared. Of the 676 (plan, span) pairs present in both Arm D and Arm E, 0 received a different concept path. An earlier draft of this round's plan declared 686 as the expected denominator; that was wrong twice over -- 28 entries are removed, not 33, and a joined block's end moves so its pair matches nothing in Arm D. The computed 676 is the figure of record, and the wrong 686 is recorded rather than deleted.

What this does not measure

The character class was fitted to the same three documents it is measured on. Arm D's run-length 3 came from a corpus-wide distribution of 328/37/18. Arm E's [-=:+] came from 38 lines drawn entirely from the three documents that are 100 % of its movable sample. The whole-corpus screen above tests generalisation outward -- it found no fourth grid-bearing document -- but it cannot break that circularity inward, and no reading of the rows should treat "declared, not swept" as meaning the same thing it meant for Arm D.

rule:table-grid is plan-level provenance and does NOT reach the bundle. Measured on the 1 294-file Arm D bundle: grep -rl "PROPOSED" returns 0, while derived appears in 84 files as a frontmatter key with other values. A reviewer looking for the rule name in a built bundle will find nothing, and that absence is a property of materialisation, not evidence that the flag did not fire.

The ceiling is bounded by which table FORM the converter chose, not by how many tables the corpus holds. Pandoc also emits simple and multiline tables, whose rows carry no | at all. _TABLE_ROW never sees those, so they are invisible to the table rule, to Arm E, and to the |-row count that measures the ceiling. One document in this sample (position 10) contains such a table in its upper half, and no arm proposes a boundary in it.

Two grid tables separated by a rule line alone would merge into one concept. Pandoc puts a blank line between adjacent tables, so it does not emit that shape -- but that is a property of the WRITER, not of this code, and in_table survives an arbitrary run of rule lines. A unit fixture asserts the merge, so the limit is declared rather than assumed away. The corpus diff found no instance.

Position 5 is the document that shows what Arm E is not. It has the largest reduction in the round, 21 concepts to 6, and its LABEL DOES NOT CHANGE. Each of its three references is still cut into a 114-character title concept carrying a heading and no body, plus its 1 675-character table. Joining table rows removed most of the fragmentation and left the rest; the remaining cut comes from the heading rule, not from the table rule.

Position 7's cause is diagnosed and deliberately unbuilt. Its 48 concepts include nine contents-listing lines with dotted leaders, and concepts of 87, 93 and 112 characters. It carries zero table-block entries, so Arm E cannot reach it. One change per measurement is the Arm C lesson; the fix is named and not made.

The orphan gate still deletes 34 % of Arm D's own boundaries, skewed. Reported in the Arm D round, unfixed, and untouched here.

No bundle was built for Arm E. The Arm D round's door-level counts came from a bundle run; here they come from the run artifacts themselves -- 4 of 43 .err files record FAILED, so extractable is 39/43 -- which is the same figure for roughly a twentieth of the wall time. Bundle-level file counts are therefore not reported for Arm E, and that is a gap, not a result.

Position 0 is unreadable and no segmentation changes that. It extracts as 95.1 % (cid:N) glyph tokens. It is an extraction defect, K3 measures segmentation, and it did not move a label in any arm.

All raters are instances of the same model family, and the first rater is not independent. Agreement between them bounds self-consistency, not correctness. The first rater had read the Arm D report's published rows before rating, so Arm B reproducing 8/4/0/0 and Arm D reproducing 4/5/0/3 is consistency and not confirmation. What the re-rating establishes is narrower and sufficient for this comparison: all three arms were judged in one round, by one identity, against the same four categories.

No CHANGELOG entry accompanies this arm. Measured precedent, re-checked this round: grep -c for "outline-run", "max-segment-chars", "Arm C" and "Arm D" returns 0 in both README.md and CHANGELOG.md, while --path-prefix -- a real interface change -- has a CHANGELOG entry. The rule is "interface and behaviour changes yes, arm flags no", and --table-grid is an arm flag that defaults off.

Reproducing

C=~/corpora/okf-telling-20260829

# The corpus loop, in one place. It writes NN.json, NN.err and one
# `i|exit|filename` line per document into _index.txt -- which is the format
# both archives carry, and which the Arm D round's published block omitted.
arm_run() {   # $1=outdir $2=proposed-at $3=lo $4=hi then flags
  DIR=$1; AT=$2; LO=$3; HI=$4; shift 4; mkdir -p "$DIR"; i=0
  for f in "$C"/K2/trinn1/*; do
    i=$((i+1))
    [ "$i" -lt "$LO" ] && continue
    [ "$i" -gt "$HI" ] && continue
    b=$(basename "$f"); n=$(printf '%02d' "$i")
    .venv/bin/python tools/okf_propose_segments.py "$f" --out "$DIR/$n.json" \
      --path-prefix "${b%.*}" --proposed-at "$AT" "$@" 2> "$DIR/$n.err"
    echo "$i|$?|$b" >> "$DIR/_index.txt"
  done }

# Run in ascending chunks, or _index.txt line order breaks. Each chunk is a
# foreground call under 600 s; documents 13-28 account for most of the time.
for lo_hi in "1 12" "13 20" "21 28" "29 36" "37 43"; do
  set -- $lo_hi
  arm_run "$C/K2-plans-armB-check-20260907" 2026-09-03T00:00:00Z "$1" "$2"
  arm_run "$C/K2-plans-armD-check-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3
  arm_run "$C/K2-plans-armE-20260907"       2026-09-07T00:00:00Z "$1" "$2" --outline-run 3 --table-grid
done

# The identity halves. Assert the counts FIRST: a diff over two trees where
# every document failed compares nothing and exits 0.
ls "$C"/K2-plans-armB-check-20260907/*.json | wc -l   # 28
ls "$C"/K2-plans-armD-check-20260907/*.json | wc -l   # 33
diff -r "$C/K2-plans-baseline-20260907" "$C/K2-plans-armB-check-20260907" -x '*.err'; echo $?
diff -r "$C/K2-plans-armd-20260907"     "$C/K2-plans-armD-check-20260907" -x '*.err'; echo $?

# The measurement: exactly three plans differ.
diff -rq "$C/K2-plans-armd-20260907" "$C/K2-plans-armE-20260907" -x '*.err' -x '_index.txt'

# The ceiling, over all 43 files.
.venv/bin/python tools/okf_table_measure.py \
  --corpus "$C/K2/trinn1" --report "$C/K2-table-reach-20260907.md"

# The door count, without building a bundle.
ls "$C"/K2-plans-armE-20260907/*.err | wc -l                        # 43
grep -l FAILED "$C"/K2-plans-armE-20260907/*.err | wc -l            # 4  -> 39/43

cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \
  | xargs shasum -a 256 | shasum -a 256

The consumer bundle, locale-pinned, before and after this round:

  • 1108 files, 9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1

LC_ALL=C is not decoration: without it sort orders the file list differently and the aggregate digest changes while the bytes do not.

Appendix: the twelve first-rater verdicts, three arms

pos document B entries D entries E entries Arm B Arm D Arm E
0 Bilag 9.1 1 concept 1 concept 1 concept too coarse too coarse too coarse
1 Bilag 3.2.2 20 23 23 too coarse too coarse too coarse
2 Bilag 1.1 1 concept 9 9 too coarse correct correct
3 Bilag 7 1 3 3 too coarse too coarse too coarse
4 Vedlegg 5 1 concept 4 4 too coarse too coarse too coarse
5 Vedlegg 3 21 21 6 too fine too fine too fine
6 Bilag 3.8 6 7 7 too coarse correct correct
7 Bilag 1.3 45 48 48 too fine too fine too fine
8 Bilag 3.4 1 concept 8 8 too coarse correct correct
9 Bilag 5 5 11 11 too coarse too fine too fine
10 Vedlegg 1 15 15 3 too fine too fine correct
11 Dokument for avtaleinngaelse 2 2 1 too fine too fine correct

"1 concept" means no plan was written: the mechanical rules found no boundary and the document lands as one flat concept.

The blind raters covered positions 0, 2, 4, 6, 8, 10 only. They disagreed with the first rater at position 6 on Arm B (both raters, correct where the first rater says too coarse -- his ground is a chapter that Arm B ABSORBS and that therefore does not appear in the material a blind rater sees), at position 8 on Arm D (one rater of two), and at position 10 on Arm E (both raters, too fine where the first rater says correct). They agreed with the first rater and with each other everywhere else.