The first rater's row moves 8/4/0/0 to 4/5/0/3; the two blind rows are identical at 5/1/0/0, and the disagreement sits on exactly the three positions labelled correct. Both are published, and the verdict is a conditional on what counts as one unit of knowledge, not a number. Controls: src/ untouched, flag-absent and --outline-run 0 byte-identical over the corpus, the consumer bundle unchanged at 1108 files and 9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1, door-level counts 39/43 unchanged, pre-gate totals 144 and 23/39 as declared. No threshold is set.
335 lines
17 KiB
Markdown
335 lines
17 KiB
Markdown
# K3 with Arm D beside a re-rated Arm B, 2026-09-07
|
|
|
|
Two numbers on the same footing, so a threshold can be set afterwards. **No
|
|
threshold is set here**, and none is implied: the K3 method
|
|
(`docs/2026-09-02-k3-k4-k5-metode.md`) declares none, and inventing one inside
|
|
the work that produces a measurement is fitting the bar to the number. This
|
|
round ran under order `20260906T213322Z-1044411564-from-.claude`, which refuses
|
|
a threshold, refuses a change to the consumer bundle, and permits no model call
|
|
in the run path. All three refusals held and each is checked below.
|
|
|
|
Counts only. The corpus is public procurement material, but nothing here needs a
|
|
document body to be checkable.
|
|
|
|
## Arm D is not defined upstream of this document
|
|
|
|
`docs/2026-09-02-k3-k4-k5-metode.md` contains **zero** occurrences of the word
|
|
"arm" (`grep -c -i "arm"` -> `0`, exit 1). Arm D is a name this repository's
|
|
brief gives to one rule, so that a measurement can refer to it:
|
|
|
|
> read the document's own numbered outline -- the integer chapter headings
|
|
> (`N`, `N.`, `N)`) the shipping grammar cannot match, because its `_NUMBERED`
|
|
> pattern requires at least one dot -- and admit a boundary only where the
|
|
> integers form a **maximal ascending run of length >= 3**, taking the **last**
|
|
> such run when the outline repeats, because a contents listing precedes the
|
|
> body it lists.
|
|
|
|
The run length **3 is declared, not swept**. It follows from the corpus's own
|
|
distribution of maximal ascending runs (328 of length 1, 37 of length 2, 18 of
|
|
length 3 or more), and a sweep over candidate lengths would be choosing the
|
|
threshold from the answer.
|
|
|
|
## The question
|
|
|
|
K3 asks whether each concept carries **one unit of knowledge** (OKF v0.2 section
|
|
2). Four categories, exactly one per document, tie-break coarse before fine
|
|
before duplicate: **too coarse / too fine / duplicate / correct**.
|
|
|
|
## Method
|
|
|
|
- **Corpus:** `~/corpora/okf-telling-20260829/K2/trinn1`, **N = 43** files, of
|
|
which **39/43** are extractable. The other four are `.smc`, `.zip`, a PDF with
|
|
no text layer, and a `.doc`.
|
|
- **Sample:** **n = 12**, drawn by the method's own rule -- hex SHA-256 of the
|
|
**NFC-normalised** filename, stratified by format (8 `pdf`, 3 `docx`,
|
|
1 `xlsx`). The draw is now committed in `tools/okf_outline_measure.py` and
|
|
re-derived this round rather than copied; it reproduced the twelve published
|
|
documents **in order, 12/12**. Under NFD the draw yields a different sample,
|
|
so the normalisation is load-bearing.
|
|
- **Two arms, one round:** Arm B is the shipping default, **re-rated this
|
|
round** rather than carried over. Arm D is the same proposer with
|
|
`--outline-run 3`.
|
|
- **Raters:** one first-rater identity over both arms, then **two separate blind
|
|
raters, one per arm**, `n_blind = 6` each at canonical positions 0, 2, 4, 6,
|
|
8, 10 -- **12 blind ratings and two `k/6` figures**. Neither blind rater saw
|
|
the other's arm, either first-rater's labels, this report, or the plan.
|
|
- **Arm C's `8/4/0/0` is historical context and explicitly not a comparand:** it
|
|
was measured in a different round against a different baseline artifact.
|
|
|
|
## Controls, passed before anything was counted
|
|
|
|
| control | result |
|
|
|---|---|
|
|
| `git diff --stat 798f64a..HEAD -- src/` | **empty** -- the library was not touched |
|
|
| flag absent vs `--outline-run 0`, whole corpus | **byte-identical**, 28/28 artifacts, exit distribution 28/11/4 both |
|
|
| flag-off re-run vs the 28 archived Arm B plans | **byte-identical** (`diff -r`, exit 0) |
|
|
| consumer bundle `K2-bundle-20260903`, before and after | **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` -- unchanged |
|
|
| door-level counts, Arm D run | extracted **39/43**, gated **39/43**, persisted **39/43** -- unchanged |
|
|
| K1b conservation | `merged + coded rejections = 43; N = 43` |
|
|
| network imports in either tool | **0** (`grep -cE "anthropic\|openai\|requests\|httpx\|urllib\.request"`) |
|
|
| `Claude-Session:` trailers in this round's commits | **0** |
|
|
| declared pre-gate totals | **144** boundaries and **23/39** documents, exactly as declared |
|
|
|
|
The last row is the gate that permitted the rest: a pre-gate total other than
|
|
144/23 would mean the implemented rule is not the measured one, and the bundle
|
|
build would not have been started.
|
|
|
|
## The denominator
|
|
|
|
Every figure below is stated against one of three denominators, and they are not
|
|
interchangeable:
|
|
|
|
- **43** -- corpus files (the door-level denominator);
|
|
- **39** -- extractable files (the segmentation denominator);
|
|
- **12** -- the K3 sample, of which **at most 8 can move** (below).
|
|
|
|
## The ceiling: at most 8 of 12
|
|
|
|
Positions 0, 5, 10 and 11 carry **zero** outline boundaries, so they are the
|
|
same proposal in both arms. Measured directly on the plan entries, with
|
|
`ingested_at` excluded because the two runs carry different `--proposed-at`:
|
|
|
|
| pos | document | Arm B | Arm D | entries identical |
|
|
|---|---|---|---|---|
|
|
| 0 | Bilag 9.1 | no plan | no plan | both absent |
|
|
| 5 | Vedlegg 3 | 21 | 21 | **True** |
|
|
| 10 | Vedlegg 1 | 15 | 15 | **True** |
|
|
| 11 | Dokument for avtaleinngaelse | 2 | 2 | **True** |
|
|
|
|
*(exploratory -- this identity check is not emitted by a committed instrument.)*
|
|
|
|
Any reading of the row starts here: a row that moved by four moved four of the
|
|
eight it could.
|
|
|
|
## K3, the two rows side by side
|
|
|
|
### First rater, n = 12
|
|
|
|
| arm | too coarse | too fine | duplicate | correct | sum |
|
|
|---|---|---|---|---|---|
|
|
| Arm B (re-rated this round) | 8 | 4 | 0 | 0 | 12 |
|
|
| **Arm D** | **4** | **5** | **0** | **3** | **12** |
|
|
| *Arm C, 2026-09-04, historical only* | *8* | *4* | *0* | *0* | *12* |
|
|
|
|
Four labels changed, all among the eight that could: positions 2, 6 and 8 moved
|
|
coarse -> correct, and position 9 moved coarse -> **fine**.
|
|
|
|
### Blind raters, n = 6 each, one per arm
|
|
|
|
| arm | rater | too coarse | too fine | duplicate | correct | sum |
|
|
|---|---|---|---|---|---|---|
|
|
| Arm B | `blind-rater-A` | 5 | 1 | 0 | 0 | 6 |
|
|
| Arm D | `blind-rater-D` | 5 | 1 | 0 | 0 | 6 |
|
|
|
|
**The two blind rows are identical.** Agreement with the first rater, on the
|
|
same six positions:
|
|
|
|
| arm | agreement |
|
|
|---|---|
|
|
| Arm B | **6/6** |
|
|
| Arm D | **3/6** |
|
|
|
|
## Verdict
|
|
|
|
**On the first rater's row, Arm D is the first arm to move the number. On the
|
|
blind raters' rows, it moved it by zero.** Both statements are measurements of
|
|
the same twelve documents, and the report refuses to publish only the first.
|
|
|
|
The disagreement is not scattered. It sits on **exactly** the three positions
|
|
where the first rater wrote `correct` -- 2, 6 and 8 -- and the blind rater wrote
|
|
`too coarse` on all three, for one consistent reason: the arm cuts at the
|
|
document's **top-level** chapters, and the blind rater judged that the chapters
|
|
still fuse their own numbered subsections. Its evidence is concrete rather than
|
|
stylistic: at position 8, `Bilag 3.4` lists about 20 second-level sections and
|
|
deeper (down to `5.2.1.1.`), and the proposal emits exactly its 8 top-level
|
|
chapters, so `Spesielle rom` (6 832 chars) carries `7.1.` through `7.5.` whole.
|
|
|
|
So the honest form of the finding is a conditional, not a number:
|
|
|
|
- **If** a top-level chapter counts as one unit of knowledge, Arm D moves K3
|
|
from 8/4/0/0 to 4/5/0/3.
|
|
- **If** the unit is the numbered subsection, Arm D moves K3 by nothing, and
|
|
what it changes is which fusion you get, not whether you get one.
|
|
|
|
Nothing in the K3 method decides between those two readings, and this round does
|
|
not decide it either. That is the operator's call, and it is a **prior**
|
|
question to any threshold: a threshold on an undecided unit measures the rater.
|
|
|
|
The one place both readings agree is criterion 7's case, position 6's
|
|
`Planlagt situasjon`: under Arm B that chapter was absorbed into a neighbour and
|
|
did not exist as a concept; under Arm D it exists (539 chars, `rule:outline`).
|
|
The blind rater still labelled the document `too coarse`, on a *different* span
|
|
(`Overvannslosning`, 4 780 chars, fusing two site solutions). The specific
|
|
defect the arm was built to fix was fixed; the document did not become correct.
|
|
|
|
**Criterion 7: PASS**, checked with a command rather than prose --
|
|
`any(e['title'] == 'Planlagt situasjon' ...)` over `34.json` -> `True`.
|
|
|
|
## What did move, with denominators
|
|
|
|
| figure | Arm B | Arm D | denominator |
|
|
|---|---|---|---|
|
|
| outline boundaries proposed (pre-gate) | -- | **144** | -- |
|
|
| boundaries surviving the orphan gate | -- | **95** | of 144 |
|
|
| documents reached (pre-gate) | -- | **23** | of 39 |
|
|
| documents reached (post-gate) | -- | **21** | of 39 |
|
|
| entries, whole corpus | 618 | **709** | delta **+91** |
|
|
| existing Arm B candidates deleted | -- | **4** | all in `Bilag 3.6` |
|
|
| documents producing an artifact | 28 | **33** | of 43 |
|
|
| documents with nothing to propose | 11 | **6** | of 43 |
|
|
| documents with zero entries | 11 | **6** | of 39 |
|
|
| unique concept paths | -- | **709** | of 709 entries |
|
|
| bundle files | 1108 | **1294** | delta +186 |
|
|
| bundle `index.md` files | 478 | **578** | delta +100 |
|
|
| proposal wall time | 762 s | **769 s** | 43 documents |
|
|
| bundle build wall time | -- | **781.69 s** reported, **1558 s** end to end | 43 documents |
|
|
|
|
**709 unique paths out of 709 entries**: no collision, so the +91 entries are 91
|
|
distinct concepts and not a renaming of existing ones. This was emitted **before**
|
|
the bundle was built, which is the point -- a collision found afterwards would be
|
|
a fact about the writer, not about the rule.
|
|
|
|
Span sizes, Arm D: 709 spans, min 10, p50 447, p95 5 848, max 148 051; **185 of
|
|
709** are under 200 chars.
|
|
|
|
**Outline titles carrying no alphabetic word: 11 of 95.** The instrument's own
|
|
definition, stated because it is not an upstream term: a word is
|
|
`[^\W\d_]{2,}` -- two or more Unicode letters -- so a title made of digits and
|
|
single letters (`477 3 025`, `D 1 L`) counts as junk. An ad-hoc count written
|
|
during this session with a one-letter threshold gives **4** instead; the
|
|
committed instrument's 11 is the figure of record, and the discrepancy is a
|
|
difference of definition, not of data.
|
|
|
|
**Concept paths for unchanged content did not churn.** Of the **569** Arm B
|
|
entries whose span survives unchanged into Arm D, **0** received a different
|
|
concept path. The plan carried this as a medium risk on the grounds that
|
|
`_segment_path`'s `taken` set is order-dependent; on the delivered artifacts the
|
|
risk did not fire. *(exploratory -- not emitted by a committed instrument.)*
|
|
**0 of 95** post-gate outline titles reduce to the reserved stem `index`.
|
|
*(exploratory.)*
|
|
|
|
## What this does not measure
|
|
|
|
**The orphan gate deletes 34 % of the arm's own boundaries, and it deletes them
|
|
systematically skewed.** 49 of 144 boundaries fall to the parent-span check: a
|
|
chapter heading followed immediately by its own `x.y` subsection has an empty
|
|
body and is dropped. So the arm keeps `Vedlegg`, `Referanser` and `Innledning`
|
|
and loses the chapters that **have** structure beneath them. **What was rated is
|
|
therefore the outline rule minus its structurally richest third.** Without this
|
|
sentence the row above reads as evidence about "the outline rule" when it is
|
|
evidence about a degraded variant of it. The narrower fix (deduplicating
|
|
coincident boundaries at insertion) and the wider one (bounding a span to the
|
|
next same-or-higher-level heading) were both considered; the wider one is
|
|
excluded here under one-change-per-measurement, which is the Arm C lesson. A fix
|
|
exists; it is not that none was found.
|
|
|
|
**An ascending integer run is not the same thing as a chapter outline**, and two
|
|
of the twelve show it directly. At position 4 (`Vedlegg 5`) the run the rule
|
|
found is the **cited regulation's subsections** -- `1)`, `3)`, `4)`, `5)` -- so
|
|
three concepts are 131-394-char statute quotes and the fourth swallows 21 197
|
|
chars, 93.2 % of the document, under a subsection's title. At position 9
|
|
(`Bilag 5`) the run is a **numbered risk table** whose rows the PDF extractor
|
|
flattened into prose, so four rows of one risk assessment became four concepts.
|
|
The pre-work control that found "0 of 144 boundaries land on a table row" is not
|
|
contradicted by this: it tested markdown table rows (`|`-delimited, 57 of 35 050
|
|
lines), and a table geometry flattened into numbered prose is invisible to that
|
|
test. The control was right about its own definition and its definition was too
|
|
narrow. That is a limit of the control, stated here rather than left implicit.
|
|
|
|
**Position 0 is unreadable, and no segmentation changes that.** `Bilag 9.1`
|
|
extracts as **95.1 %** `(cid:N)` glyph tokens (206 758 of 217 470 chars), because
|
|
every embedded font is `/Type3` with no `/ToUnicode` map. Both blind raters
|
|
reached that independently. Its `too coarse` label rests on document extent and
|
|
the PDF's own bookmark outline -- which names two merged constituent documents --
|
|
not on reading the text. It is an **extraction** defect and K3 measures
|
|
segmentation, so it did not change a label; but a concept ingested from that file
|
|
today would carry almost no readable text however it were cut.
|
|
|
|
**Every JSON proposal leaves the document's head text uncovered** -- 135 to
|
|
3 773 chars of cover page, contents listing, and in two cases the body
|
|
`Innledning`. Measured by the Arm D blind rater across all five of its JSON
|
|
proposals, recorded here because it is real, and not used as a label: omission
|
|
is not one of the four categories.
|
|
|
|
**All raters are instances of the same model family.** The first rater and both
|
|
blind raters are Claude Opus 5. Agreement between them is not independent
|
|
confirmation in the sense a human panel would provide; it bounds
|
|
self-consistency, not correctness. The first rater had additionally seen the
|
|
published Arm C row before rating Arm B, so the Arm B row's reproduction of
|
|
`8/4/0/0` is **not** independent confirmation either. What the re-rating does
|
|
establish is narrower and sufficient for this comparison: both arms were judged
|
|
in the same round, by the same identity, against the same four categories.
|
|
|
|
**No CHANGELOG entry accompanies this arm.** Measured precedent rather than
|
|
preference: `--max-segment-chars` and "Arm C" appear **0** times in
|
|
`CHANGELOG.md`, while `--path-prefix` has an entry at `:40-44`. The precedent is
|
|
"interface and behaviour changes yes, arm flags no", and `--outline-run` is an
|
|
arm flag that defaults to off.
|
|
|
|
## Reproducing
|
|
|
|
```
|
|
C=~/corpora/okf-telling-20260829
|
|
|
|
# the flag-off identity half (byte-compare against the archive AND against the
|
|
# no-flag run; both were checked)
|
|
Z="$C/K2-plans-zero-20260907"; mkdir -p "$Z"; i=0
|
|
for f in "$C"/K2/trinn1/*; do
|
|
i=$((i+1)); b=$(basename "$f")
|
|
.venv/bin/python tools/okf_propose_segments.py "$f" \
|
|
--out "$Z/$(printf '%02d' $i).json" \
|
|
--path-prefix "${b%.*}" --proposed-at 2026-09-03T00:00:00Z --outline-run 0
|
|
done
|
|
diff -q -r "$C/K2-plans-baseline-20260907" "$Z" -x '*.err' -x '_index.txt'
|
|
|
|
# Arm D, into a FRESH dated directory -- never a reused one, because a leftover
|
|
# plan matching on source_sha256 would be replayed silently
|
|
D="$C/K2-plans-armd-20260907"; mkdir -p "$D"; i=0
|
|
for f in "$C"/K2/trinn1/*; do
|
|
i=$((i+1)); b=$(basename "$f")
|
|
.venv/bin/python tools/okf_propose_segments.py "$f" \
|
|
--out "$D/$(printf '%02d' $i).json" \
|
|
--path-prefix "${b%.*}" --proposed-at 2026-09-07T00:00:00Z --outline-run 3
|
|
done
|
|
|
|
.venv/bin/python tools/okf_outline_measure.py \
|
|
--corpus "$C/K2/trinn1" --report "$C/K2-outline-reach-20260907.md"
|
|
|
|
.venv/bin/python tools/okf_corpus_run.py \
|
|
--corpus "$C/K2/trinn1" \
|
|
--report "$C/K2-bundle-armd-20260907-report.md" \
|
|
--bundle "$C/K2-bundle-armd-20260907" --plans-dir "$D" \
|
|
--bundle-id k2-trinn1-armd-20260907 --okf-version 0.2 \
|
|
--ingested-at 2026-09-07T00:00:00Z
|
|
|
|
cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \
|
|
| xargs shasum -a 256 | shasum -a 256
|
|
```
|
|
|
|
The consumer bundle, locale-pinned, before and after this round:
|
|
|
|
- **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1`
|
|
|
|
`LC_ALL=C` is not decoration: without it `sort` orders the file list differently
|
|
and the aggregate digest changes while the bytes do not.
|
|
|
|
## Appendix: the twelve first-rater verdicts, both arms
|
|
|
|
| pos | document | Arm B entries | Arm D entries | Arm B | Arm D |
|
|
|---|---|---|---|---|---|
|
|
| 0 | Bilag 9.1 | 1 concept | 1 concept | too coarse | too coarse |
|
|
| 1 | Bilag 3.2.2 | 20 | 23 | too coarse | too coarse |
|
|
| 2 | Bilag 1.1 | 1 concept | 9 | too coarse | **correct** |
|
|
| 3 | Bilag 7 | 1 | 3 | too coarse | too coarse |
|
|
| 4 | Vedlegg 5 | 1 concept | 4 | too coarse | too coarse |
|
|
| 5 | Vedlegg 3 | 21 | 21 | too fine | too fine |
|
|
| 6 | Bilag 3.8 | 6 | 7 | too coarse | **correct** |
|
|
| 7 | Bilag 1.3 | 45 | 48 | too fine | too fine |
|
|
| 8 | Bilag 3.4 | 1 concept | 8 | too coarse | **correct** |
|
|
| 9 | Bilag 5 | 5 | 11 | too coarse | **too fine** |
|
|
| 10 | Vedlegg 1 | 15 | 15 | too fine | too fine |
|
|
| 11 | Dokument for avtaleinngaelse | 2 | 2 | too fine | too fine |
|
|
|
|
The blind raters covered positions 0, 2, 4, 6, 8, 10 only, and disagreed with
|
|
the first rater at 2, 6 and 8 on Arm D -- the three bolded `correct` labels --
|
|
and nowhere on Arm B.
|