Three rows rated in one round by one first-rater identity, then two blind raters per arm as the order asked -- 36 first-rater verdicts, 36 blind ratings, six k/6 figures. No threshold is set and none is implied. First rater: Arm B 8/4/0/0, Arm D 4/5/0/3, Arm E 4/3/0/5. Against the direction declared before the row was seen -- fewer `too fine` WITHOUT more `too coarse` -- `too fine` goes 5 -> 3 and `too coarse` stays 4. Blind raters: `too fine` is 1 of 6 in EVERY arm, Arm E included. The report publishes both rows and states the verdict as a conditional, because they are measurements of the same twelve documents and only one of them is favourable. The disagreement is one position and it is legible. At position 10 both Arm E blind raters kept `too fine` where the first rater moved to `correct`, and their REASON changed rather than persisting: under Arm B and D they object that thirteen table rows are severed from their header; under Arm E, where the table is one concept with its header, they object that the table is severed from the sentence that introduces it. Arm E fixed the first complaint and does not touch the second. And the denominator that governs how much that can say: the blind positions are fixed at 0, 2, 4, 6, 8, 10, and only ONE of Arm E's three movable documents falls in that set. The blind row carries one document's worth of information about this arm. No threshold should be set on that, and the report says so. The ceiling is measured this round rather than inferred. A committed instrument extracted all 43 files and counted table ROWS in the text -- 3 of 39 documents, 57 rows -- because zero ENTRIES is not the same claim: the orphan check can delete a table candidate before it becomes an entry, and one sample document produces no plan at all so its zero would be an absence with no denominator. The 57 reproduces the Arm D round's independent pre-work screen exactly, with neither measurement derived from the other. Controls are split by what they do on failure, which is this round's correction to its own first plan: gating controls halt, predictions are reported whatever they say. Only one prediction was allowed to halt -- a zero-diff run, which would be a wiring bug wearing a null result's clothes, with the ceiling standing ready as a plausible wrong explanation. Nine limits are stated rather than left implicit, including three that weaken the arm's own case: the character class was fitted to the same three documents it is measured on; position 5 has the round's largest reduction (21 concepts to 6) and its label does not change; and pandoc's simple and multiline table forms carry no pipes at all, so the ceiling is bounded by which form the converter chose rather than by how many tables the corpus holds. Pure ASCII (0 non-ASCII bytes), short document labels only (0 corpus filenames). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
22 KiB
K3 with Arm E beside Arm D and a re-rated Arm B, 2026-09-07
Three numbers on the same footing, so a threshold can be set afterwards. No
threshold is set here, and none is implied: the K3 method
(docs/2026-09-02-k3-k4-k5-metode.md) declares none, and inventing one inside
the work that produces a measurement is fitting the bar to the number. This
round ran under order 20260907T075834Z-18584396-from-.claude, which refuses a
threshold, refuses a change to the consumer bundle, permits no model call in the
run path, and forbids a push. All four refusals held and each is checked below.
One thing IS declared before the row is read, and it is not a threshold: the
direction that counts as movement -- fewer too fine WITHOUT more too coarse. It sets no value any count must reach. It is written down in advance
precisely so it cannot be chosen after the number is known.
Counts only. The corpus is public procurement material, but nothing here needs a document body to be checkable.
Arm E is not defined upstream of this document
docs/2026-09-02-k3-k4-k5-metode.md contains zero occurrences of the word
"arm" (grep -c -i "arm" -> 0, exit 1). Arm E is a name this repository's
brief gives to one rule, so that a measurement can refer to it:
a pandoc GRID-table rule line --
+---+---+, and+===+===+under a header -- does not close an open table block. A block is marked as JOINED only when a rule line was actually crossed between two table rows, never merely because its span contains one, so a single-row grid table stays byte-identical to Arm D.
The rule has no numeric parameter, so nothing was swept and nothing could be.
What is declared instead is the rule line's character class, [-=:+], and it is
measured rather than guessed: across the three grid-bearing documents of this
corpus, 38 of 38 lines whose stripped form starts with + match the
pattern, and those four characters are the complete set occurring on them. The
: is pandoc's column-alignment marker and it is load bearing -- a first pass
with [-=+] matched 37 of 38 and, through that single miss, read one
document as having two tables where it has one. The 37 is recorded here rather
than quietly corrected.
The question
K3 asks whether each concept carries one unit of knowledge (OKF v0.2 section 2). Four categories, exactly one per document, tie-break coarse before fine before duplicate: too coarse / too fine / duplicate / correct.
Method
- Corpus:
K2/trinn1, N = 43 files, of which 39/43 are extractable. - Sample: n = 12, drawn by the method's own rule -- hex SHA-256 of the
NFC-normalised filename, stratified by format (8
pdf, 3docx, 1xlsx), re-derived this round fromtools/okf_outline_measure.pyand reproducing the twelve published documents in order, 12/12. - Three arms, one round: Arm B is the shipping default (
--outline-run 0), Arm D is--outline-run 3, Arm E is--outline-run 3 --table-grid. Arm B and Arm D are re-rated this round rather than carried over. - Raters: one first-rater identity over all three arms, 36 verdicts;
then two blind raters per arm at canonical positions 0, 2, 4, 6, 8, 10,
n_blind = 6each -- 36 blind ratings and sixk/6figures. - The blind protocol is WIDER than Arm D's and the two are not comparable.
docs/2026-09-07-k3-arm-d.mdused "two separate blind raters, one per arm, 12 blind ratings and twok/6figures". The order asked for two per arm. Its number governs; the divergence is stated so the rounds'k/6figures are not read as like for like. - Blindness is structural, not promised. Each blind rater is a separate
subagent with its own context, given ONE file: six documents labelled
A-Fwith their extracted length and, per concept, its length, its title and the first 180 characters of its body. Rule names were stripped, because arule:table-gridin the list would have identified the arm. No rater was told which arm it read, that other arms exist, what the first rater said, or that a brief, plan or report exists. - Arm C's
8/4/0/0is historical context and explicitly not a comparand.
Controls, passed before anything was counted
The controls are split by what they do on failure, and that split is the correction this round makes to its own first plan. A gating control asks whether the shipped rule is the rule being measured; it halts. A prediction is a figure written down in advance from an exploratory replica; it is reported whatever it says, because a replica may not sit in judgement over shipped code.
Gating -- each one halts the round
| control | result |
|---|---|
git diff --stat 54a0bc2..HEAD -- src/ |
only propose.py and cli.py; 136 changed lines in propose.py, most of them comments |
consumer bundle K2-bundle-20260903, file count |
1108 -- the literal published in the Arm D round |
the same bundle, LC_ALL=C aggregate digest |
9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1 -- the published literal, before and after |
| network imports, proposer and both instruments | 0 each |
| Arm B identity, whole corpus, no flags | byte-identical to the archived Arm B plans, diff -r exit 0, _index.txt INCLUDED |
Arm D identity, whole corpus, --outline-run 3 |
byte-identical to the archived Arm D plans, diff -r exit 0, _index.txt INCLUDED |
| artifact counts asserted BEFORE each diff | Arm B 28 json + 43 .err + 43 index lines; Arm D 33 + 43 + 43 |
| door-level counts | .err files recording FAILED: 4 of 43, so extractable 39/43, unchanged |
| declared grid-rule totals | 38 lines on 3 documents, exactly as declared |
Two of these rows are corrections to the Arm D round's own published procedure,
and both were found by review rather than by failure. The artifact counts are
asserted before the diff, because diff -r over two trees where every
document failed would compare nothing and exit 0. And _index.txt is
included in the comparison: the Arm D reproduce block neither generates it
nor compares it, which would let NN.json name a different document across two
runs with nothing saying so.
The Arm B identity run is not bookkeeping either. It is the corpus-level half of
the promise that --outline-run 0 with the flag absent is still Arm B, and it
is what lets the Arm B row below rest on verified bytes.
Predictions -- reported, never gating
Written into the brief from an exploratory replica before the rule was built, and reproduced by the shipped code:
| prediction | measured |
|---|---|
| exactly 3 of 33 plans differ from the Arm D archive | 3 -- and they are the three named |
| position 5: 21 -> 6 entries | 21 -> 6 |
| position 10: 15 -> 3 entries | 15 -> 3 |
| position 11: 2 -> 1 entries | 2 -> 1 |
One prediction was allowed to halt, and only one: if NO plan had differed,
the flag would not have been threaded through to find_candidates and the row
would have been a wiring bug wearing a null result's clothes -- with the ceiling
below standing ready as a plausible wrong explanation. It did not occur.
The plan-to-position mapping is derived, not assumed: each plan's first
entry path prefix is matched against the reduced stem of its source filename
and that against the committed draw. The NN in NN.json comes from an
unsorted shell glob and names nothing on its own.
The denominator
Every figure below is stated against one of four denominators, and they are not interchangeable:
- 43 -- corpus files (the door-level denominator);
- 39 -- extractable files (the segmentation denominator);
- 12 -- the K3 sample, of which at most 3 can move (below);
- 6 -- the blind positions, of which 1 is a document Arm E can move.
The ceiling: at most 3 of 12, and this time it is measured
The Arm D round's ceiling was read off entry counts. That is not sound on its own: the orphan check can delete a table candidate before it becomes an entry, so a document could hold table rows that never reach a plan -- and position 0 produces no plan at all, so its zero would be an absence with no denominator.
This round measures the ceiling in the text itself, with a committed instrument, over all 43 files and before any rating began:
| figure | value | denominator |
|---|---|---|
| documents with at least one table row | 3 | 39 |
| table rows in total | 57 | -- |
| documents with at least one grid-rule line | 3 | 39 |
| grid-rule lines in total | 38 | -- |
A document with no table row cannot be moved by this arm, whatever the orphan check later did to its candidates. So 3 of the 12 sample documents can move, and they are positions 5, 10 and 11 -- three of the four the Arm D ceiling excluded, because they carry zero outline boundaries.
The 57 is an independent corroboration and worth stating as one. The Arm D
round's pre-work control counted |-delimited table rows over the corpus text
it screened and reported 57 of 35 050 lines. This round's instrument, run
against different code for a different purpose, counts 57. Neither
measurement was derived from the other.
K3, the three rows side by side
First rater, n = 12
| arm | too coarse | too fine | duplicate | correct | sum |
|---|---|---|---|---|---|
| Arm B (re-rated this round) | 8 | 4 | 0 | 0 | 12 |
| Arm D (re-rated this round) | 4 | 5 | 0 | 3 | 12 |
| Arm E | 4 | 3 | 0 | 5 | 12 |
| Arm C, 2026-09-04, historical only | 8 | 4 | 0 | 0 | 12 |
Two labels changed, both among the three that could: positions 10 and 11 moved
too fine -> correct. Position 5 did not move.
Blind raters, n = 6 each, two per arm
| arm | rater | too coarse | too fine | duplicate | correct | sum | agreement with first rater |
|---|---|---|---|---|---|---|---|
| Arm B | blind-1a |
4 | 1 | 0 | 1 | 6 | 5/6 |
| Arm B | blind-1b |
4 | 1 | 0 | 1 | 6 | 5/6 |
| Arm D | blind-2a |
2 | 1 | 0 | 3 | 6 | 6/6 |
| Arm D | blind-2b |
3 | 1 | 0 | 2 | 6 | 5/6 |
| Arm E | blind-3a |
2 | 1 | 0 | 3 | 6 | 5/6 |
| Arm E | blind-3b |
2 | 1 | 0 | 3 | 6 | 5/6 |
Within-arm agreement, which exists for the first time because there are two raters per arm: Arm B 6/6, Arm D 5/6, Arm E 6/6.
too fine is 1 of 6 in every arm, Arm E included.
Verdict
On the first rater's row, Arm E moves too fine from 5 to 3 while too coarse stays at 4 -- the direction declared in advance. On the blind raters'
rows, too fine does not move at all. Both statements are measurements of the
same twelve documents, and the report refuses to publish only the first.
The disagreement is one position and it is legible. At position 10 the first
rater moved too fine -> correct; both Arm E blind raters kept too fine.
Their reason CHANGED rather than persisting. Under Arm B and Arm D they object
that thirteen table rows are severed from their header row. Under Arm E, where
the table is one concept with its header included, they object that the table is
severed from the sentence that introduces it. Arm E fixed the first complaint
and does not touch the second.
So the honest form of the finding is a conditional, and it has two clauses:
- If a table is one unit of knowledge, Arm E moves K3 from 4/5/0/3 to
4/3/0/5 and does it without trading a
too finefor atoo coarse. - If a table is a unit only together with the prose that introduces it, Arm E moves K3 by nothing on the position where both readings were tested, and what it changes is which severance you get, not whether you get one.
And the second clause is measured on ONE document. The blind positions are fixed at 0, 2, 4, 6, 8, 10, and only position 10 is a document Arm E can move. Positions 5 and 11 -- the other two -- were seen by no blind rater. The blind row is therefore not evidence that Arm E fails on those two; it is evidence that this protocol cannot see them. A round in which the arm's reach and the blind protocol's positions overlap in one document is a round whose blind row carries one document's worth of information about the arm, and no threshold should be set on that.
The unit question Arm D surfaced is still open and is still the operator's.
It fired again here, on identical material: at position 8 one Arm D blind rater
called top-level chapters correct and the other called them too coarse
because their numbered subsections are "distinct requirement sets a reader would
want separately". Nothing in the K3 method decides between those readings, this
round does not decide it either, and it is prior to any threshold -- a threshold
on an undecided unit measures the rater.
What did move, with denominators
| figure | Arm B | Arm D | Arm E | denominator |
|---|---|---|---|---|
| entries, whole corpus | 618 | 709 | 681 | delta -28 from Arm D |
entries carrying rule:table-block |
33 | 33 | 5 | of 681 |
entries carrying rule:table-grid |
-- | -- | 5 | of 681 |
| documents whose entry count changed | -- | -- | 3 | of 39 |
| plans differing from the Arm D archive | -- | -- | 3 | of 33 |
| blocks joined | -- | -- | 5 | -- |
| documents producing an artifact | 28 | 33 | 33 | of 43 |
| extractable | 39 | 39 | 39 | of 43 |
| position 5 entries | 21 | 21 | 6 | -- |
| position 10 entries | 15 | 15 | 3 | -- |
| position 11 entries | 2 | 2 | 1 | -- |
| largest span Arm E creates | -- | -- | 13 691 chars | position 10 |
Concept paths for unchanged content did not churn, and the denominator is
computed rather than declared. Of the 676 (plan, span) pairs present in
both Arm D and Arm E, 0 received a different concept path. An earlier draft
of this round's plan declared 686 as the expected denominator; that was wrong
twice over -- 28 entries are removed, not 33, and a joined block's end moves
so its pair matches nothing in Arm D. The computed 676 is the figure of record,
and the wrong 686 is recorded rather than deleted.
What this does not measure
The character class was fitted to the same three documents it is measured on.
Arm D's run-length 3 came from a corpus-wide distribution of 328/37/18. Arm E's
[-=:+] came from 38 lines drawn entirely from the three documents that are 100
% of its movable sample. The whole-corpus screen above tests generalisation
outward -- it found no fourth grid-bearing document -- but it cannot break that
circularity inward, and no reading of the rows should treat "declared, not
swept" as meaning the same thing it meant for Arm D.
rule:table-grid is plan-level provenance and does NOT reach the bundle.
Measured on the 1 294-file Arm D bundle: grep -rl "PROPOSED" returns 0,
while derived appears in 84 files as a frontmatter key with other values. A
reviewer looking for the rule name in a built bundle will find nothing, and that
absence is a property of materialisation, not evidence that the flag did not
fire.
The ceiling is bounded by which table FORM the converter chose, not by how
many tables the corpus holds. Pandoc also emits simple and multiline
tables, whose rows carry no | at all. _TABLE_ROW never sees those, so they
are invisible to the table rule, to Arm E, and to the |-row count that
measures the ceiling. One document in this sample (position 10) contains such a
table in its upper half, and no arm proposes a boundary in it.
Two grid tables separated by a rule line alone would merge into one concept.
Pandoc puts a blank line between adjacent tables, so it does not emit that
shape -- but that is a property of the WRITER, not of this code, and in_table
survives an arbitrary run of rule lines. A unit fixture asserts the merge, so
the limit is declared rather than assumed away. The corpus diff found no
instance.
Position 5 is the document that shows what Arm E is not. It has the largest reduction in the round, 21 concepts to 6, and its LABEL DOES NOT CHANGE. Each of its three references is still cut into a 114-character title concept carrying a heading and no body, plus its 1 675-character table. Joining table rows removed most of the fragmentation and left the rest; the remaining cut comes from the heading rule, not from the table rule.
Position 7's cause is diagnosed and deliberately unbuilt. Its 48 concepts include nine contents-listing lines with dotted leaders, and concepts of 87, 93 and 112 characters. It carries zero table-block entries, so Arm E cannot reach it. One change per measurement is the Arm C lesson; the fix is named and not made.
The orphan gate still deletes 34 % of Arm D's own boundaries, skewed. Reported in the Arm D round, unfixed, and untouched here.
No bundle was built for Arm E. The Arm D round's door-level counts came from
a bundle run; here they come from the run artifacts themselves -- 4 of 43 .err
files record FAILED, so extractable is 39/43 -- which is the same figure for
roughly a twentieth of the wall time. Bundle-level file counts are therefore not
reported for Arm E, and that is a gap, not a result.
Position 0 is unreadable and no segmentation changes that. It extracts as
95.1 % (cid:N) glyph tokens. It is an extraction defect, K3 measures
segmentation, and it did not move a label in any arm.
All raters are instances of the same model family, and the first rater is not
independent. Agreement between them bounds self-consistency, not correctness.
The first rater had read the Arm D report's published rows before rating, so Arm
B reproducing 8/4/0/0 and Arm D reproducing 4/5/0/3 is consistency and not
confirmation. What the re-rating establishes is narrower and sufficient for this
comparison: all three arms were judged in one round, by one identity, against
the same four categories.
No CHANGELOG entry accompanies this arm. Measured precedent, re-checked this
round: grep -c for "outline-run", "max-segment-chars", "Arm C" and "Arm D"
returns 0 in both README.md and CHANGELOG.md, while --path-prefix -- a real
interface change -- has a CHANGELOG entry. The rule is "interface and behaviour
changes yes, arm flags no", and --table-grid is an arm flag that defaults off.
Reproducing
C=~/corpora/okf-telling-20260829
# The corpus loop, in one place. It writes NN.json, NN.err and one
# `i|exit|filename` line per document into _index.txt -- which is the format
# both archives carry, and which the Arm D round's published block omitted.
arm_run() { # $1=outdir $2=proposed-at $3=lo $4=hi then flags
DIR=$1; AT=$2; LO=$3; HI=$4; shift 4; mkdir -p "$DIR"; i=0
for f in "$C"/K2/trinn1/*; do
i=$((i+1))
[ "$i" -lt "$LO" ] && continue
[ "$i" -gt "$HI" ] && continue
b=$(basename "$f"); n=$(printf '%02d' "$i")
.venv/bin/python tools/okf_propose_segments.py "$f" --out "$DIR/$n.json" \
--path-prefix "${b%.*}" --proposed-at "$AT" "$@" 2> "$DIR/$n.err"
echo "$i|$?|$b" >> "$DIR/_index.txt"
done }
# Run in ascending chunks, or _index.txt line order breaks. Each chunk is a
# foreground call under 600 s; documents 13-28 account for most of the time.
for lo_hi in "1 12" "13 20" "21 28" "29 36" "37 43"; do
set -- $lo_hi
arm_run "$C/K2-plans-armB-check-20260907" 2026-09-03T00:00:00Z "$1" "$2"
arm_run "$C/K2-plans-armD-check-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3
arm_run "$C/K2-plans-armE-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3 --table-grid
done
# The identity halves. Assert the counts FIRST: a diff over two trees where
# every document failed compares nothing and exits 0.
ls "$C"/K2-plans-armB-check-20260907/*.json | wc -l # 28
ls "$C"/K2-plans-armD-check-20260907/*.json | wc -l # 33
diff -r "$C/K2-plans-baseline-20260907" "$C/K2-plans-armB-check-20260907" -x '*.err'; echo $?
diff -r "$C/K2-plans-armd-20260907" "$C/K2-plans-armD-check-20260907" -x '*.err'; echo $?
# The measurement: exactly three plans differ.
diff -rq "$C/K2-plans-armd-20260907" "$C/K2-plans-armE-20260907" -x '*.err' -x '_index.txt'
# The ceiling, over all 43 files.
.venv/bin/python tools/okf_table_measure.py \
--corpus "$C/K2/trinn1" --report "$C/K2-table-reach-20260907.md"
# The door count, without building a bundle.
ls "$C"/K2-plans-armE-20260907/*.err | wc -l # 43
grep -l FAILED "$C"/K2-plans-armE-20260907/*.err | wc -l # 4 -> 39/43
cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \
| xargs shasum -a 256 | shasum -a 256
The consumer bundle, locale-pinned, before and after this round:
- 1108 files,
9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1
LC_ALL=C is not decoration: without it sort orders the file list differently
and the aggregate digest changes while the bytes do not.
Appendix: the twelve first-rater verdicts, three arms
| pos | document | B entries | D entries | E entries | Arm B | Arm D | Arm E |
|---|---|---|---|---|---|---|---|
| 0 | Bilag 9.1 | 1 concept | 1 concept | 1 concept | too coarse | too coarse | too coarse |
| 1 | Bilag 3.2.2 | 20 | 23 | 23 | too coarse | too coarse | too coarse |
| 2 | Bilag 1.1 | 1 concept | 9 | 9 | too coarse | correct | correct |
| 3 | Bilag 7 | 1 | 3 | 3 | too coarse | too coarse | too coarse |
| 4 | Vedlegg 5 | 1 concept | 4 | 4 | too coarse | too coarse | too coarse |
| 5 | Vedlegg 3 | 21 | 21 | 6 | too fine | too fine | too fine |
| 6 | Bilag 3.8 | 6 | 7 | 7 | too coarse | correct | correct |
| 7 | Bilag 1.3 | 45 | 48 | 48 | too fine | too fine | too fine |
| 8 | Bilag 3.4 | 1 concept | 8 | 8 | too coarse | correct | correct |
| 9 | Bilag 5 | 5 | 11 | 11 | too coarse | too fine | too fine |
| 10 | Vedlegg 1 | 15 | 15 | 3 | too fine | too fine | correct |
| 11 | Dokument for avtaleinngaelse | 2 | 2 | 1 | too fine | too fine | correct |
"1 concept" means no plan was written: the mechanical rules found no boundary and the document lands as one flat concept.
The blind raters covered positions 0, 2, 4, 6, 8, 10 only. They disagreed with
the first rater at position 6 on Arm B (both raters, correct where the first
rater says too coarse -- his ground is a chapter that Arm B ABSORBS and that
therefore does not appear in the material a blind rater sees), at position 8 on
Arm D (one rater of two), and at position 10 on Arm E (both raters, too fine
where the first rater says correct). They agreed with the first rater and with
each other everywhere else.