llm-ingestion-okf/docs/2026-09-07-k3-arm-e.md
Kjell Tore Guttormsen a7ee942d5f docs(measure): K3 Arm E beside Arm B and Arm D -- three rows, no threshold
Three rows rated in one round by one first-rater identity, then two blind
raters per arm as the order asked -- 36 first-rater verdicts, 36 blind ratings,
six k/6 figures. No threshold is set and none is implied.

First rater: Arm B 8/4/0/0, Arm D 4/5/0/3, Arm E 4/3/0/5. Against the direction
declared before the row was seen -- fewer `too fine` WITHOUT more `too coarse`
-- `too fine` goes 5 -> 3 and `too coarse` stays 4.

Blind raters: `too fine` is 1 of 6 in EVERY arm, Arm E included. The report
publishes both rows and states the verdict as a conditional, because they are
measurements of the same twelve documents and only one of them is favourable.

The disagreement is one position and it is legible. At position 10 both Arm E
blind raters kept `too fine` where the first rater moved to `correct`, and
their REASON changed rather than persisting: under Arm B and D they object that
thirteen table rows are severed from their header; under Arm E, where the table
is one concept with its header, they object that the table is severed from the
sentence that introduces it. Arm E fixed the first complaint and does not touch
the second.

And the denominator that governs how much that can say: the blind positions are
fixed at 0, 2, 4, 6, 8, 10, and only ONE of Arm E's three movable documents
falls in that set. The blind row carries one document's worth of information
about this arm. No threshold should be set on that, and the report says so.

The ceiling is measured this round rather than inferred. A committed instrument
extracted all 43 files and counted table ROWS in the text -- 3 of 39 documents,
57 rows -- because zero ENTRIES is not the same claim: the orphan check can
delete a table candidate before it becomes an entry, and one sample document
produces no plan at all so its zero would be an absence with no denominator.
The 57 reproduces the Arm D round's independent pre-work screen exactly, with
neither measurement derived from the other.

Controls are split by what they do on failure, which is this round's correction
to its own first plan: gating controls halt, predictions are reported whatever
they say. Only one prediction was allowed to halt -- a zero-diff run, which
would be a wiring bug wearing a null result's clothes, with the ceiling standing
ready as a plausible wrong explanation.

Nine limits are stated rather than left implicit, including three that weaken
the arm's own case: the character class was fitted to the same three documents
it is measured on; position 5 has the round's largest reduction (21 concepts to
6) and its label does not change; and pandoc's simple and multiline table forms
carry no pipes at all, so the ceiling is bounded by which form the converter
chose rather than by how many tables the corpus holds.

Pure ASCII (0 non-ASCII bytes), short document labels only (0 corpus filenames).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 11:48:58 +02:00

418 lines
22 KiB
Markdown

# K3 with Arm E beside Arm D and a re-rated Arm B, 2026-09-07
Three numbers on the same footing, so a threshold can be set afterwards. **No
threshold is set here**, and none is implied: the K3 method
(`docs/2026-09-02-k3-k4-k5-metode.md`) declares none, and inventing one inside
the work that produces a measurement is fitting the bar to the number. This
round ran under order `20260907T075834Z-18584396-from-.claude`, which refuses a
threshold, refuses a change to the consumer bundle, permits no model call in the
run path, and forbids a push. All four refusals held and each is checked below.
One thing IS declared before the row is read, and it is not a threshold: the
**direction** that counts as movement -- fewer `too fine` WITHOUT more `too
coarse`. It sets no value any count must reach. It is written down in advance
precisely so it cannot be chosen after the number is known.
Counts only. The corpus is public procurement material, but nothing here needs a
document body to be checkable.
## Arm E is not defined upstream of this document
`docs/2026-09-02-k3-k4-k5-metode.md` contains **zero** occurrences of the word
"arm" (`grep -c -i "arm"` -> `0`, exit 1). Arm E is a name this repository's
brief gives to one rule, so that a measurement can refer to it:
> a pandoc GRID-table rule line -- `+---+---+`, and `+===+===+` under a header
> -- does not close an open table block. A block is marked as JOINED only when
> a rule line was actually crossed between two table rows, never merely because
> its span contains one, so a single-row grid table stays byte-identical to
> Arm D.
**The rule has no numeric parameter, so nothing was swept and nothing could be.**
What is declared instead is the rule line's character class, `[-=:+]`, and it is
measured rather than guessed: across the three grid-bearing documents of this
corpus, **38 of 38** lines whose stripped form starts with `+` match the
pattern, and those four characters are the complete set occurring on them. The
`:` is pandoc's column-alignment marker and it is load bearing -- a first pass
with `[-=+]` matched **37 of 38** and, through that single miss, read one
document as having two tables where it has one. The 37 is recorded here rather
than quietly corrected.
## The question
K3 asks whether each concept carries **one unit of knowledge** (OKF v0.2 section
2). Four categories, exactly one per document, tie-break coarse before fine
before duplicate: **too coarse / too fine / duplicate / correct**.
## Method
- **Corpus:** `K2/trinn1`, **N = 43** files, of which **39/43** are extractable.
- **Sample:** **n = 12**, drawn by the method's own rule -- hex SHA-256 of the
**NFC-normalised** filename, stratified by format (8 `pdf`, 3 `docx`,
1 `xlsx`), re-derived this round from `tools/okf_outline_measure.py` and
reproducing the twelve published documents in order, **12/12**.
- **Three arms, one round:** Arm B is the shipping default (`--outline-run 0`),
Arm D is `--outline-run 3`, Arm E is `--outline-run 3 --table-grid`. Arm B and
Arm D are **re-rated** this round rather than carried over.
- **Raters:** one first-rater identity over all three arms, **36 verdicts**;
then **two blind raters per arm** at canonical positions 0, 2, 4, 6, 8, 10,
`n_blind = 6` each -- **36 blind ratings and six `k/6` figures**.
- **The blind protocol is WIDER than Arm D's and the two are not comparable.**
`docs/2026-09-07-k3-arm-d.md` used "two separate blind raters, **one per
arm**, 12 blind ratings and two `k/6` figures". The order asked for two per
arm. Its number governs; the divergence is stated so the rounds' `k/6`
figures are not read as like for like.
- **Blindness is structural, not promised.** Each blind rater is a separate
subagent with its own context, given ONE file: six documents labelled `A`-`F`
with their extracted length and, per concept, its length, its title and the
first 180 characters of its body. **Rule names were stripped**, because a
`rule:table-grid` in the list would have identified the arm. No rater was told
which arm it read, that other arms exist, what the first rater said, or that a
brief, plan or report exists.
- **Arm C's `8/4/0/0` is historical context and explicitly not a comparand.**
## Controls, passed before anything was counted
The controls are split by what they do on failure, and that split is the
correction this round makes to its own first plan. A **gating** control asks
whether the shipped rule is the rule being measured; it halts. A **prediction**
is a figure written down in advance from an exploratory replica; it is reported
whatever it says, because a replica may not sit in judgement over shipped code.
### Gating -- each one halts the round
| control | result |
|---|---|
| `git diff --stat 54a0bc2..HEAD -- src/` | only `propose.py` and `cli.py`; **136** changed lines in `propose.py`, most of them comments |
| consumer bundle `K2-bundle-20260903`, file count | **1108** -- the literal published in the Arm D round |
| the same bundle, `LC_ALL=C` aggregate digest | `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` -- the published literal, before and after |
| network imports, proposer and both instruments | **0** each |
| **Arm B identity**, whole corpus, no flags | **byte-identical** to the archived Arm B plans, `diff -r` exit **0**, `_index.txt` INCLUDED |
| **Arm D identity**, whole corpus, `--outline-run 3` | **byte-identical** to the archived Arm D plans, `diff -r` exit **0**, `_index.txt` INCLUDED |
| artifact counts asserted BEFORE each diff | Arm B **28** json + **43** `.err` + **43** index lines; Arm D **33** + **43** + **43** |
| door-level counts | `.err` files recording `FAILED`: **4** of 43, so extractable **39/43**, unchanged |
| declared grid-rule totals | **38** lines on **3** documents, exactly as declared |
Two of these rows are corrections to the Arm D round's own published procedure,
and both were found by review rather than by failure. The artifact counts are
asserted **before** the diff, because `diff -r` over two trees where every
document failed would compare nothing and exit 0. And `_index.txt` is
**included** in the comparison: the Arm D reproduce block neither generates it
nor compares it, which would let `NN.json` name a different document across two
runs with nothing saying so.
The Arm B identity run is not bookkeeping either. It is the corpus-level half of
the promise that `--outline-run 0` with the flag absent is still Arm B, and it
is what lets the Arm B row below rest on verified bytes.
### Predictions -- reported, never gating
Written into the brief from an exploratory replica **before** the rule was
built, and reproduced by the shipped code:
| prediction | measured |
|---|---|
| exactly 3 of 33 plans differ from the Arm D archive | **3** -- and they are the three named |
| position 5: 21 -> 6 entries | **21 -> 6** |
| position 10: 15 -> 3 entries | **15 -> 3** |
| position 11: 2 -> 1 entries | **2 -> 1** |
**One prediction was allowed to halt, and only one:** if NO plan had differed,
the flag would not have been threaded through to `find_candidates` and the row
would have been a wiring bug wearing a null result's clothes -- with the ceiling
below standing ready as a plausible wrong explanation. It did not occur.
The plan-to-position mapping is **derived, not assumed**: each plan's first
entry `path` prefix is matched against the reduced stem of its source filename
and that against the committed draw. The `NN` in `NN.json` comes from an
unsorted shell glob and names nothing on its own.
## The denominator
Every figure below is stated against one of four denominators, and they are not
interchangeable:
- **43** -- corpus files (the door-level denominator);
- **39** -- extractable files (the segmentation denominator);
- **12** -- the K3 sample, of which **at most 3 can move** (below);
- **6** -- the blind positions, of which **1** is a document Arm E can move.
## The ceiling: at most 3 of 12, and this time it is measured
The Arm D round's ceiling was read off entry counts. That is not sound on its
own: the orphan check can delete a table candidate before it becomes an entry,
so a document could hold table rows that never reach a plan -- and position 0
produces no plan at all, so its zero would be an absence with no denominator.
This round measures the ceiling in the text itself, with a committed instrument,
over all 43 files and **before any rating began**:
| figure | value | denominator |
|---|---|---|
| documents with at least one table row | **3** | 39 |
| table rows in total | **57** | -- |
| documents with at least one grid-rule line | **3** | 39 |
| grid-rule lines in total | **38** | -- |
A document with no table row cannot be moved by this arm, whatever the orphan
check later did to its candidates. So **3 of the 12** sample documents can move,
and they are positions 5, 10 and 11 -- three of the four the Arm D ceiling
excluded, because they carry zero outline boundaries.
**The 57 is an independent corroboration and worth stating as one.** The Arm D
round's pre-work control counted `|`-delimited table rows over the corpus text
it screened and reported **57 of 35 050 lines**. This round's instrument, run
against different code for a different purpose, counts **57**. Neither
measurement was derived from the other.
## K3, the three rows side by side
### First rater, n = 12
| arm | too coarse | too fine | duplicate | correct | sum |
|---|---|---|---|---|---|
| Arm B (re-rated this round) | 8 | 4 | 0 | 0 | 12 |
| Arm D (re-rated this round) | 4 | 5 | 0 | 3 | 12 |
| **Arm E** | **4** | **3** | **0** | **5** | **12** |
| *Arm C, 2026-09-04, historical only* | *8* | *4* | *0* | *0* | *12* |
Two labels changed, both among the three that could: positions 10 and 11 moved
`too fine` -> `correct`. Position 5 did not move.
### Blind raters, n = 6 each, two per arm
| arm | rater | too coarse | too fine | duplicate | correct | sum | agreement with first rater |
|---|---|---|---|---|---|---|---|
| Arm B | `blind-1a` | 4 | 1 | 0 | 1 | 6 | **5/6** |
| Arm B | `blind-1b` | 4 | 1 | 0 | 1 | 6 | **5/6** |
| Arm D | `blind-2a` | 2 | 1 | 0 | 3 | 6 | **6/6** |
| Arm D | `blind-2b` | 3 | 1 | 0 | 2 | 6 | **5/6** |
| Arm E | `blind-3a` | 2 | 1 | 0 | 3 | 6 | **5/6** |
| Arm E | `blind-3b` | 2 | 1 | 0 | 3 | 6 | **5/6** |
Within-arm agreement, which exists for the first time because there are two
raters per arm: Arm B **6/6**, Arm D **5/6**, Arm E **6/6**.
**`too fine` is 1 of 6 in every arm, Arm E included.**
## Verdict
**On the first rater's row, Arm E moves `too fine` from 5 to 3 while `too
coarse` stays at 4 -- the direction declared in advance. On the blind raters'
rows, `too fine` does not move at all.** Both statements are measurements of the
same twelve documents, and the report refuses to publish only the first.
The disagreement is one position and it is legible. At position 10 the first
rater moved `too fine` -> `correct`; both Arm E blind raters kept `too fine`.
Their reason CHANGED rather than persisting. Under Arm B and Arm D they object
that thirteen table rows are severed from their header row. Under Arm E, where
the table is one concept with its header included, they object that the table is
severed from the sentence that introduces it. Arm E fixed the first complaint
and does not touch the second.
So the honest form of the finding is a conditional, and it has two clauses:
- **If** a table is one unit of knowledge, Arm E moves K3 from 4/5/0/3 to
4/3/0/5 and does it without trading a `too fine` for a `too coarse`.
- **If** a table is a unit only together with the prose that introduces it, Arm
E moves K3 by nothing on the position where both readings were tested, and
what it changes is which severance you get, not whether you get one.
**And the second clause is measured on ONE document.** The blind positions are
fixed at 0, 2, 4, 6, 8, 10, and only position 10 is a document Arm E can move.
Positions 5 and 11 -- the other two -- were seen by no blind rater. The blind
row is therefore not evidence that Arm E fails on those two; it is evidence that
this protocol cannot see them. A round in which the arm's reach and the blind
protocol's positions overlap in one document is a round whose blind row carries
one document's worth of information about the arm, and no threshold should be
set on that.
**The unit question Arm D surfaced is still open and is still the operator's.**
It fired again here, on identical material: at position 8 one Arm D blind rater
called top-level chapters `correct` and the other called them `too coarse`
because their numbered subsections are "distinct requirement sets a reader would
want separately". Nothing in the K3 method decides between those readings, this
round does not decide it either, and it is prior to any threshold -- a threshold
on an undecided unit measures the rater.
## What did move, with denominators
| figure | Arm B | Arm D | Arm E | denominator |
|---|---|---|---|---|
| entries, whole corpus | 618 | 709 | **681** | delta **-28** from Arm D |
| entries carrying `rule:table-block` | 33 | 33 | **5** | of 681 |
| entries carrying `rule:table-grid` | -- | -- | **5** | of 681 |
| documents whose entry count changed | -- | -- | **3** | of 39 |
| plans differing from the Arm D archive | -- | -- | **3** | of 33 |
| blocks joined | -- | -- | **5** | -- |
| documents producing an artifact | 28 | 33 | **33** | of 43 |
| extractable | 39 | 39 | **39** | of 43 |
| position 5 entries | 21 | 21 | **6** | -- |
| position 10 entries | 15 | 15 | **3** | -- |
| position 11 entries | 2 | 2 | **1** | -- |
| largest span Arm E creates | -- | -- | **13 691** chars | position 10 |
**Concept paths for unchanged content did not churn, and the denominator is
computed rather than declared.** Of the **676** `(plan, span)` pairs present in
both Arm D and Arm E, **0** received a different concept path. An earlier draft
of this round's plan declared 686 as the expected denominator; that was wrong
twice over -- 28 entries are removed, not 33, and a joined block's `end` moves
so its pair matches nothing in Arm D. The computed 676 is the figure of record,
and the wrong 686 is recorded rather than deleted.
## What this does not measure
**The character class was fitted to the same three documents it is measured on.**
Arm D's run-length 3 came from a corpus-wide distribution of 328/37/18. Arm E's
`[-=:+]` came from 38 lines drawn entirely from the three documents that are 100
% of its movable sample. The whole-corpus screen above tests generalisation
outward -- it found no fourth grid-bearing document -- but it cannot break that
circularity inward, and no reading of the rows should treat "declared, not
swept" as meaning the same thing it meant for Arm D.
**`rule:table-grid` is plan-level provenance and does NOT reach the bundle.**
Measured on the 1 294-file Arm D bundle: `grep -rl "PROPOSED"` returns **0**,
while `derived` appears in 84 files as a frontmatter key with other values. A
reviewer looking for the rule name in a built bundle will find nothing, and that
absence is a property of materialisation, not evidence that the flag did not
fire.
**The ceiling is bounded by which table FORM the converter chose, not by how
many tables the corpus holds.** Pandoc also emits *simple* and *multiline*
tables, whose rows carry no `|` at all. `_TABLE_ROW` never sees those, so they
are invisible to the table rule, to Arm E, and to the `|`-row count that
measures the ceiling. One document in this sample (position 10) contains such a
table in its upper half, and no arm proposes a boundary in it.
**Two grid tables separated by a rule line alone would merge into one concept.**
Pandoc puts a blank line between adjacent tables, so it does not emit that
shape -- but that is a property of the WRITER, not of this code, and `in_table`
survives an arbitrary run of rule lines. A unit fixture asserts the merge, so
the limit is declared rather than assumed away. The corpus diff found no
instance.
**Position 5 is the document that shows what Arm E is not.** It has the largest
reduction in the round, 21 concepts to 6, and its LABEL DOES NOT CHANGE. Each of
its three references is still cut into a 114-character title concept carrying a
heading and no body, plus its 1 675-character table. Joining table rows removed
most of the fragmentation and left the rest; the remaining cut comes from the
heading rule, not from the table rule.
**Position 7's cause is diagnosed and deliberately unbuilt.** Its 48 concepts
include nine contents-listing lines with dotted leaders, and concepts of 87, 93
and 112 characters. It carries zero table-block entries, so Arm E cannot reach
it. One change per measurement is the Arm C lesson; the fix is named and not
made.
**The orphan gate still deletes 34 % of Arm D's own boundaries, skewed.**
Reported in the Arm D round, unfixed, and untouched here.
**No bundle was built for Arm E.** The Arm D round's door-level counts came from
a bundle run; here they come from the run artifacts themselves -- 4 of 43 `.err`
files record `FAILED`, so extractable is 39/43 -- which is the same figure for
roughly a twentieth of the wall time. Bundle-level file counts are therefore not
reported for Arm E, and that is a gap, not a result.
**Position 0 is unreadable and no segmentation changes that.** It extracts as
95.1 % `(cid:N)` glyph tokens. It is an extraction defect, K3 measures
segmentation, and it did not move a label in any arm.
**All raters are instances of the same model family, and the first rater is not
independent.** Agreement between them bounds self-consistency, not correctness.
The first rater had read the Arm D report's published rows before rating, so Arm
B reproducing `8/4/0/0` and Arm D reproducing `4/5/0/3` is consistency and not
confirmation. What the re-rating establishes is narrower and sufficient for this
comparison: all three arms were judged in one round, by one identity, against
the same four categories.
**No CHANGELOG entry accompanies this arm.** Measured precedent, re-checked this
round: `grep -c` for "outline-run", "max-segment-chars", "Arm C" and "Arm D"
returns 0 in both `README.md` and `CHANGELOG.md`, while `--path-prefix` -- a real
interface change -- has a CHANGELOG entry. The rule is "interface and behaviour
changes yes, arm flags no", and `--table-grid` is an arm flag that defaults off.
## Reproducing
```
C=~/corpora/okf-telling-20260829
# The corpus loop, in one place. It writes NN.json, NN.err and one
# `i|exit|filename` line per document into _index.txt -- which is the format
# both archives carry, and which the Arm D round's published block omitted.
arm_run() { # $1=outdir $2=proposed-at $3=lo $4=hi then flags
DIR=$1; AT=$2; LO=$3; HI=$4; shift 4; mkdir -p "$DIR"; i=0
for f in "$C"/K2/trinn1/*; do
i=$((i+1))
[ "$i" -lt "$LO" ] && continue
[ "$i" -gt "$HI" ] && continue
b=$(basename "$f"); n=$(printf '%02d' "$i")
.venv/bin/python tools/okf_propose_segments.py "$f" --out "$DIR/$n.json" \
--path-prefix "${b%.*}" --proposed-at "$AT" "$@" 2> "$DIR/$n.err"
echo "$i|$?|$b" >> "$DIR/_index.txt"
done }
# Run in ascending chunks, or _index.txt line order breaks. Each chunk is a
# foreground call under 600 s; documents 13-28 account for most of the time.
for lo_hi in "1 12" "13 20" "21 28" "29 36" "37 43"; do
set -- $lo_hi
arm_run "$C/K2-plans-armB-check-20260907" 2026-09-03T00:00:00Z "$1" "$2"
arm_run "$C/K2-plans-armD-check-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3
arm_run "$C/K2-plans-armE-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3 --table-grid
done
# The identity halves. Assert the counts FIRST: a diff over two trees where
# every document failed compares nothing and exits 0.
ls "$C"/K2-plans-armB-check-20260907/*.json | wc -l # 28
ls "$C"/K2-plans-armD-check-20260907/*.json | wc -l # 33
diff -r "$C/K2-plans-baseline-20260907" "$C/K2-plans-armB-check-20260907" -x '*.err'; echo $?
diff -r "$C/K2-plans-armd-20260907" "$C/K2-plans-armD-check-20260907" -x '*.err'; echo $?
# The measurement: exactly three plans differ.
diff -rq "$C/K2-plans-armd-20260907" "$C/K2-plans-armE-20260907" -x '*.err' -x '_index.txt'
# The ceiling, over all 43 files.
.venv/bin/python tools/okf_table_measure.py \
--corpus "$C/K2/trinn1" --report "$C/K2-table-reach-20260907.md"
# The door count, without building a bundle.
ls "$C"/K2-plans-armE-20260907/*.err | wc -l # 43
grep -l FAILED "$C"/K2-plans-armE-20260907/*.err | wc -l # 4 -> 39/43
cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \
| xargs shasum -a 256 | shasum -a 256
```
The consumer bundle, locale-pinned, before and after this round:
- **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1`
`LC_ALL=C` is not decoration: without it `sort` orders the file list differently
and the aggregate digest changes while the bytes do not.
## Appendix: the twelve first-rater verdicts, three arms
| pos | document | B entries | D entries | E entries | Arm B | Arm D | Arm E |
|---|---|---|---|---|---|---|---|
| 0 | Bilag 9.1 | 1 concept | 1 concept | 1 concept | too coarse | too coarse | too coarse |
| 1 | Bilag 3.2.2 | 20 | 23 | 23 | too coarse | too coarse | too coarse |
| 2 | Bilag 1.1 | 1 concept | 9 | 9 | too coarse | correct | correct |
| 3 | Bilag 7 | 1 | 3 | 3 | too coarse | too coarse | too coarse |
| 4 | Vedlegg 5 | 1 concept | 4 | 4 | too coarse | too coarse | too coarse |
| 5 | Vedlegg 3 | 21 | 21 | **6** | too fine | too fine | too fine |
| 6 | Bilag 3.8 | 6 | 7 | 7 | too coarse | correct | correct |
| 7 | Bilag 1.3 | 45 | 48 | 48 | too fine | too fine | too fine |
| 8 | Bilag 3.4 | 1 concept | 8 | 8 | too coarse | correct | correct |
| 9 | Bilag 5 | 5 | 11 | 11 | too coarse | too fine | too fine |
| 10 | Vedlegg 1 | 15 | 15 | **3** | too fine | too fine | **correct** |
| 11 | Dokument for avtaleinngaelse | 2 | 2 | **1** | too fine | too fine | **correct** |
"1 concept" means no plan was written: the mechanical rules found no boundary
and the document lands as one flat concept.
The blind raters covered positions 0, 2, 4, 6, 8, 10 only. They disagreed with
the first rater at position 6 on Arm B (both raters, `correct` where the first
rater says `too coarse` -- his ground is a chapter that Arm B ABSORBS and that
therefore does not appear in the material a blind rater sees), at position 8 on
Arm D (one rater of two), and at position 10 on Arm E (both raters, `too fine`
where the first rater says `correct`). They agreed with the first rater and with
each other everywhere else.