llm-ingestion-okf/docs/2026-09-04-k3-arm-c.md
Kjell Tore Guttormsen 98be70c144 docs(measure): K3 with Arm C beside the baseline, two numbers, no threshold
Arm C changed the proposal on 6 of the 12 sampled documents and moved the
K3 category counts by ZERO. Both arms: too coarse 8, too fine 4, duplicate
0, correct 0. That is the finding, not a null result -- the size rule makes
concepts smaller, which is what the consumer asked for, and K3 does not
measure size. It asks whether a boundary falls where a unit of knowledge
begins, and an arithmetic cut does not know where one begins.

Blind second rater per arm, n_blind = 6, two separate raters so an arm C
rating cannot anchor on an arm B one: 5/6 in each arm, and the two blind
raters' own six-document counts are identical to each other. The single
disagreement is the same document as the baseline's -- with the roles
reversed, and resolved by a specific offset in one direction and a general
impression in the other, exactly as the baseline described. The `too
coarse` count is still a floor.

Reported rather than smoothed: this first rater records 8/4/0/0 on the same
arm B proposals the 2026-09-02 baseline scored 7/3/1/1. The proposals did
not change; the rater did, and both differences are traced.

What did move, and is not K3: max concept body 217 472 -> 19 988,
concepts over 100 000 characters 4 -> 0, documents proposing nothing
11/39 -> 6/39, concepts 629 -> 680.

Found while measuring, not looked for: Bilag 9.1 -- the largest concept in
the bundle and the one the consumer flagged -- is an EXTRACTION failure.
95.1 % of its 217 470 characters are (cid:N) glyph codes and 98 alphabetic
words of 4+ letters survive in the whole document. Arm C divides unreadable
text into 14 pieces of unreadable text. Its verdict in both arms rests on
document identity, not on readable content.

The bundle a consumer is running against right now is byte-unchanged:
whole tree, 1108 files, sha256 9cd74519... before and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:58:50 +02:00

16 KiB

K3 with Arm C beside the baseline, 2026-09-04

Two numbers on the same footing, so a threshold can be set afterwards. No threshold is set here, and none is implied: the K3 method (docs/2026-09-02-k3-k4-k5-metode.md) declares none, and inventing one inside the work that produces a measurement is fitting the bar to the number.

Counts only. The corpus is public procurement material, but nothing here needs a document body or a full filename list to be checkable.

Arm C is not defined upstream of this document

docs/2026-09-02-k3-k4-k5-metode.md contains zero occurrences of the word "arm" (grep -c -i "arm" -> 0). Neither Arm A nor Arm B nor Arm C is defined there. Where the three names actually occur:

arm where it is named what it says
A docs/2026-08-29-konverteringsarmen-arm-a.md:1 "Conversion arm (arm A): pandoc vs markitdown vs docling"
B tests/test_segmented_inbox.py:509, a code comment "arm B cannot execute without this"
C docs/2026-09-03-k2-bundle-rebuild.md:171 "Reported, not fixed: Arm C is its own decision."

So Arm C had no definition to conform to. The definition measured here is this document's own and is marked as such everywhere it appears, including in the tool's --help:

Arm C = Arm B's mechanical rules, plus one deterministic rule that cuts any proposed span longer than a declared cap at the nearest paragraph boundary at or before it, the whole document counting as one span when the rules find no boundary at all.

One rule and not two, deliberately. The two failure modes docs/2026-09-03-k2-bundle-rebuild.md measured -- a PDF with no outline (Bilag 9.1, 217 472 characters) and a PDF whose headings are its table of contents so the trailing segment absorbs the body (Bilag 3.1, Bilag 1) -- are the same failure of size, and a second rule aimed at each would confound which one moved the number.

What Arm C deliberately does not change: the region before the first candidate is still covered by no segment. That is a real coverage defect -- the baseline's blind rater named it -- and fixing it here would have put two changes behind one measurement.

The cap, declared before any Arm C proposal was judged

max_segment_chars = 20000, fixed from Arm B's own output and not from a K3 outcome:

  • Arm B's segment-body sizes over the whole K2 bundle, n = 629 concepts: p50 441, p75 1 215, p90 2 972, p95 6 099, mean 2 698, max 217 472.
  • 20 000 is ~3.3x that p95, so what Arm B already produces is left alone: measured 15 of 629 concepts (2.4 %) exceed it.
  • At the corpus's measured ~2.20 characters per token, 20 000 characters is ~9 100 tokens -- an order of magnitude below the 100 000-token line that motivated the rebuild.

No sensitivity sweep was run. Sweeping the cap and keeping the value that produced the best K3 number would fit the parameter to the verdict.

The standard profile did not move

--max-segment-chars defaults to 0, which is off. A test writes the artifact with and without the flag at 0 and compares bytes.

The bundle a consumer is running against was hashed before the work started and again after the Arm C bundle was built, whole tree, 1108 files:

find K2-bundle-20260903 -type f | sort | xargs shasum -a 256 | shasum -a 256
9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1   (before)
9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1   (after)

Arm C was built to its own output: K2-bundle-armc-20260904/, plans in K2-plans-armc-20260904/.

K3, the two numbers side by side

Same 12 documents, same canonical hash order, same four categories, same tie-break (coarse before fine before duplicate). First rater claude-opus-5/okf-session-89, labels written to file before any blind rating was read.

category arm B (baseline) arm C
too coarse 8 8
too fine 4 4
duplicate 0 0
correct 0 0

Both sum to n = 12. No replacement was needed; the reconstructed sample is the baseline's own -- a known-positive check, 4 of the 12 have an empty Arm B proposal, which is what docs/2026-09-02-k3-k4-k5-baseline.md reported.

Arm C changed the proposal on 6 of the 12 documents and moved the category counts by zero. That is the finding, not a null result: the size rule makes concepts smaller, which is what a consumer asked for, and K3 does not measure size. K3 asks whether a boundary falls where a unit of knowledge begins, and an arithmetic cut does not know where one begins.

This first rater does not reproduce the 2026-09-02 baseline row

The 2026-09-02 baseline recorded 7 / 3 / 1 / 1 on the same twelve Arm B proposals. This session's first rater records 8 / 4 / 0 / 0. The proposals did not change; the rater did. Two of the differences are traceable:

  • The baseline's single correct was contested by its own blind rater, who named a chapter absorbed into the preceding segment and an introduction covered by no segment. That document is Bilag 3.8 - Overvannsutredning, and this rater checked at that resolution: its chapter 2 (Planlagt situasjon, offset 6325) lies wholly inside the span 4064-6886 of the segment opened by section 1.3. Rated too coarse.
  • The baseline's single duplicate is not reproduced. Where a table-of-contents line and a body heading produce two segments with the same title, this rater read the two as different content -- a contents listing is not the section it lists -- so too fine rather than duplicate.

The baseline said its too coarse count was a floor rather than an estimate. This re-rating is consistent with that and does not overturn it: same-family raters, small n, and a difference in reading rather than in evidence.

Blind second rater, per arm

n_blind = 6 per arm: positions 0, 2, 4, 6, 8, 10 of the sample's canonical hash order. Two separate raters, one per arm, each in its own context, each given only the proposals, the source evidence, the four category definitions and the tie-break -- and neither given the other arm, so an arm C rating cannot anchor on an arm B one.

arm agreement the one disagreement
B (baseline) 5/6 position 6, Bilag 3.8 - Overvannsutredning
C 5/6 position 6, the same document, the same proposal

The two arms agree with their blind raters at the same rate, and the blind raters' own six-document counts are identical to each other: 4 too coarse, 1 too fine, 1 correct, in both arms. Two rater instances that never saw each other's arm reached the same distribution -- which is the same headline as the twelve-document counts, reached independently.

The disagreement is the same document as the baseline's, with the roles reversed. On 2026-09-02 the first rater called Bilag 3.8 correct and the blind rater called it too coarse. Here the first rater calls it too coarse and both blind raters call it correct, on the grounds that "the six segments track the document's own numbered subsections (1.1-1.3, 3.1-3.3), each opening exactly at its heading."

That is true and it is not the whole document. The first rater's evidence is narrower and checkable: the document's chapter 2, Planlagt situasjon, opens at offset 6325, and the segment opened by section 1.3 spans 4064-6886 -- so chapter 2 lies wholly inside it and has no segment of its own. Both arms propose this identical segmentation, because the document is under the cap and Arm C left it untouched.

The asymmetry is the finding, not the ratio. In both the baseline and this run, the disagreement was resolved by a specific offset in one direction and by a general impression in the other, and the specific direction was always too coarse. That is what the baseline meant by calling its too coarse count a floor. It is still a floor.

No verdict is revised. A first rater who edits a label after reading the second rater's is not being second-rated at all.

Cohen's kappa is deliberately not reported, for the reason the method gives: at four categories and n_blind = 6 a chance-corrected statistic would look more rigorous than percent agreement while being less informative.

The independence caveat stands and is not softened. All raters here are instances of the same model family, so agreement overstates independence, probably substantially.

What did move, with denominators

Nothing in this section is K3. It is reported because it is what the consumer's finding was actually about.

figure arm B arm C
documents with something to propose 28/39 33/39
documents proposing nothing 11/39 6/39
proposed entries 618 674
concepts in the bundle 629 680
index.md files 478 483
concepts carrying adjudication 618/629 674/680
concept body characters, median 441 510
concept body characters, max 217 472 19 988
concepts over 100 000 characters 4 0
concepts over 20 000 characters 15 0
wall time, whole run 784.29 s 820.55 s

Door-level counts are unchanged, as they must be -- Arm C touches segmentation, not extraction: N = 43, merged 39/43, coded rejections 4/43 (extractor_unknown 3, extractor_empty_pdf 1), run exited 0.

The 6 documents that still propose nothing under Arm C are the ones already under the cap. Arm C fires on size; where size is not the problem it has nothing to say, and a one-entry plan would only dress a single concept in a plan file.

The three documents the order named

Concept bodies in the two bundles, not plan spans:

document arm B concepts arm B largest arm C concepts arm C largest
Bilag 3.1 34 180 604 43 19 946
Bilag 1 18 142 667 27 19 920
Bilag 9.1 1 217 472 14 19 988

A finding this measurement did not go looking for

Bilag 9.1 -- the largest concept in the bundle, and the one the consumer flagged -- is an extraction failure, not a segmentation failure. Measured on its extracted text: 95.1 % of the 217 470 characters are (cid:N) glyph codes (27 978 of them), and 98 alphabetic words of four or more letters survive in the whole document. The PDF's fonts carry no usable ToUnicode mapping, so the text this repository persisted is very largely undecodable.

This bounds what any arm can claim about that document. Its category in both arms rests on document identity -- an agreement that is transferred, plus attachments -- and not on readable content, so it contributes one count to each arm and no evidence that discriminates between them. Arm C divides unreadable text into 14 pieces of unreadable text.

Reported, not fixed. A CID-mapped PDF is an extraction question, and the extraction warning this library already emits ("figures are not represented") does not cover it: this document's prose did not survive either.

What this does not measure

  • No threshold, on K3 or anything else. Two numbers, so the operator can set one.
  • K4 and K5 are untouched. No adjudication timing was re-run and no SPEC § 5.1 provenance was re-counted; the baseline's figures stand.
  • Arm C's bundle was not re-measured against the consumption contract, and its rebuild-equals-incremental property was not re-run. Both are properties of the harness, unchanged by this flag.
  • pptx, odt, rtf still have denominator 0 in this corpus. Unmeasured, not passing.
  • The cap was not swept. One value, declared first, measured once.

Reproducing

PLANS=~/corpora/okf-telling-20260829/K2-plans-armc-20260904; mkdir -p "$PLANS"
i=0
for f in ~/corpora/okf-telling-20260829/K2/trinn1/*; do
  i=$((i+1)); b=$(basename "$f")
  .venv/bin/python tools/okf_propose_segments.py "$f" \
    --out "$PLANS/$(printf '%02d' $i).json" \
    --path-prefix "${b%.*}" --proposed-at 2026-09-04T00:00:00Z \
    --max-segment-chars 20000
done
.venv/bin/python tools/okf_corpus_run.py \
  --corpus ~/corpora/okf-telling-20260829/K2/trinn1 \
  --report ~/corpora/okf-telling-20260829/K2-bundle-armc-20260904-report.md \
  --bundle ~/corpora/okf-telling-20260829/K2-bundle-armc-20260904 \
  --ingested-at 2026-09-04T00:00:00Z --plans-dir "$PLANS" \
  --bundle-id k2-trinn1-armc-20260904 --okf-version 0.2

Converter as the harness resolved it: the vendored pypandoc binary, version 3.9 -- not the host's 3.10.2.

Appendix: the twelve raw verdicts, both arms

Blind-subset positions are marked *; those six are the ones with a second rater. Documents are named by their short corpus label.

pos document arm B arm C arm B blind arm C blind
0 * Bilag 9.1 (Avtale som tiltransporteres) too coarse too coarse too coarse too coarse
1 Bilag 3.2.2 (RIG-R02 geoteknisk) too coarse too coarse - -
2 * Bilag 1.1 (Generelle tekniske krav) too coarse too coarse too coarse too coarse
3 Bilag 7 (Prisskjema, xlsx) too coarse too coarse - -
4 * Vedlegg 5 (Begrunnelse for krav) too coarse too coarse too coarse too coarse
5 Vedlegg 3 (Erfaringsoversikt) too fine too fine - -
6 * Bilag 3.8 (Overvannsutredning) too coarse too coarse correct correct
7 Bilag 1.3 (BIM-manual) too fine too fine - -
8 * Bilag 3.4 (Premissdokument bygningsfysikk) too coarse too coarse too coarse too coarse
9 Bilag 5 (SHA-plan) too coarse too coarse - -
10 * Vedlegg 1 (Soknadsbrev) too fine too fine too fine too fine
11 Dokument for avtaleinngaelse too fine too fine - -

The six documents whose arm B and arm C proposals are byte-identical are positions 5, 6, 7, 9, 10 and 11: each is under the cap, so Arm C left it untouched and the two verdicts are the same judgement of the same artifact, not two judgements that happened to agree. The six Arm C proposals that are genuinely new are positions 0, 1, 2, 3, 4 and 8 -- and none of them changed category.

The first rater's evidence, per position, in one line each:

  • 0 -- 14 size-cut parts of a document whose text is 95.1 % undecodable; see the finding above. Arm B emitted it as one 217 470-character concept.
  • 1 -- entry 20 fuses section 5.2, chapter 6 and the appendices (22 281 characters); Arm C shortens that entry to 17 756 and adds a 4 525-character part, and the fusion is still there.
  • 2 -- nine numbered chapters, listed in the document's own contents, in one concept; Arm C cuts at a page footer at ~19 164, fusing roughly six chapters then three.
  • 3 -- one entry over F.1 Prissammenstilling, the maintenance-agreement table, F.2 Regningsarbeider and the options table; Arm C's seven parts align with none of those four, and its first entry is a 34-character sliver.
  • 4 -- background, the quoted regulation and the assessment in one concept; Arm C's cut separates the appendix table but leaves the three fused.
  • 5 -- three reference forms shredded into 21 entries.
  • 6 -- chapter 2 absorbed into the span of section 1.3; see the blind-rater disagreement above.
  • 7 -- nine spurious segments cut out of the table of contents, above 36 body segments that do track the document.
  • 8 -- 24 pages in one concept; Arm C's three parts break a sentence across parts 2 and 3.
  • 9 -- entry 5 fuses chapters 2, 3, 4, 5, 6 and the appendix (8 853 characters), and entry 1 is a fragment of the table of contents.
  • 10 -- the checklist table cut into 13 fixed-size slices.
  • 11 -- the agreement's own head (offsets 0-793) covered by no segment, and the signature table emitted as two fragments.