docs(measure): K3 Arm E beside Arm B and Arm D -- three rows, no threshold
Three rows rated in one round by one first-rater identity, then two blind raters per arm as the order asked -- 36 first-rater verdicts, 36 blind ratings, six k/6 figures. No threshold is set and none is implied. First rater: Arm B 8/4/0/0, Arm D 4/5/0/3, Arm E 4/3/0/5. Against the direction declared before the row was seen -- fewer `too fine` WITHOUT more `too coarse` -- `too fine` goes 5 -> 3 and `too coarse` stays 4. Blind raters: `too fine` is 1 of 6 in EVERY arm, Arm E included. The report publishes both rows and states the verdict as a conditional, because they are measurements of the same twelve documents and only one of them is favourable. The disagreement is one position and it is legible. At position 10 both Arm E blind raters kept `too fine` where the first rater moved to `correct`, and their REASON changed rather than persisting: under Arm B and D they object that thirteen table rows are severed from their header; under Arm E, where the table is one concept with its header, they object that the table is severed from the sentence that introduces it. Arm E fixed the first complaint and does not touch the second. And the denominator that governs how much that can say: the blind positions are fixed at 0, 2, 4, 6, 8, 10, and only ONE of Arm E's three movable documents falls in that set. The blind row carries one document's worth of information about this arm. No threshold should be set on that, and the report says so. The ceiling is measured this round rather than inferred. A committed instrument extracted all 43 files and counted table ROWS in the text -- 3 of 39 documents, 57 rows -- because zero ENTRIES is not the same claim: the orphan check can delete a table candidate before it becomes an entry, and one sample document produces no plan at all so its zero would be an absence with no denominator. The 57 reproduces the Arm D round's independent pre-work screen exactly, with neither measurement derived from the other. Controls are split by what they do on failure, which is this round's correction to its own first plan: gating controls halt, predictions are reported whatever they say. Only one prediction was allowed to halt -- a zero-diff run, which would be a wiring bug wearing a null result's clothes, with the ceiling standing ready as a plausible wrong explanation. Nine limits are stated rather than left implicit, including three that weaken the arm's own case: the character class was fitted to the same three documents it is measured on; position 5 has the round's largest reduction (21 concepts to 6) and its label does not change; and pandoc's simple and multiline table forms carry no pipes at all, so the ceiling is bounded by which form the converter chose rather than by how many tables the corpus holds. Pure ASCII (0 non-ASCII bytes), short document labels only (0 corpus filenames). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
fb46424a05
commit
a7ee942d5f
1 changed files with 418 additions and 0 deletions
418
docs/2026-09-07-k3-arm-e.md
Normal file
418
docs/2026-09-07-k3-arm-e.md
Normal file
|
|
@ -0,0 +1,418 @@
|
|||
# K3 with Arm E beside Arm D and a re-rated Arm B, 2026-09-07
|
||||
|
||||
Three numbers on the same footing, so a threshold can be set afterwards. **No
|
||||
threshold is set here**, and none is implied: the K3 method
|
||||
(`docs/2026-09-02-k3-k4-k5-metode.md`) declares none, and inventing one inside
|
||||
the work that produces a measurement is fitting the bar to the number. This
|
||||
round ran under order `20260907T075834Z-18584396-from-.claude`, which refuses a
|
||||
threshold, refuses a change to the consumer bundle, permits no model call in the
|
||||
run path, and forbids a push. All four refusals held and each is checked below.
|
||||
|
||||
One thing IS declared before the row is read, and it is not a threshold: the
|
||||
**direction** that counts as movement -- fewer `too fine` WITHOUT more `too
|
||||
coarse`. It sets no value any count must reach. It is written down in advance
|
||||
precisely so it cannot be chosen after the number is known.
|
||||
|
||||
Counts only. The corpus is public procurement material, but nothing here needs a
|
||||
document body to be checkable.
|
||||
|
||||
## Arm E is not defined upstream of this document
|
||||
|
||||
`docs/2026-09-02-k3-k4-k5-metode.md` contains **zero** occurrences of the word
|
||||
"arm" (`grep -c -i "arm"` -> `0`, exit 1). Arm E is a name this repository's
|
||||
brief gives to one rule, so that a measurement can refer to it:
|
||||
|
||||
> a pandoc GRID-table rule line -- `+---+---+`, and `+===+===+` under a header
|
||||
> -- does not close an open table block. A block is marked as JOINED only when
|
||||
> a rule line was actually crossed between two table rows, never merely because
|
||||
> its span contains one, so a single-row grid table stays byte-identical to
|
||||
> Arm D.
|
||||
|
||||
**The rule has no numeric parameter, so nothing was swept and nothing could be.**
|
||||
What is declared instead is the rule line's character class, `[-=:+]`, and it is
|
||||
measured rather than guessed: across the three grid-bearing documents of this
|
||||
corpus, **38 of 38** lines whose stripped form starts with `+` match the
|
||||
pattern, and those four characters are the complete set occurring on them. The
|
||||
`:` is pandoc's column-alignment marker and it is load bearing -- a first pass
|
||||
with `[-=+]` matched **37 of 38** and, through that single miss, read one
|
||||
document as having two tables where it has one. The 37 is recorded here rather
|
||||
than quietly corrected.
|
||||
|
||||
## The question
|
||||
|
||||
K3 asks whether each concept carries **one unit of knowledge** (OKF v0.2 section
|
||||
2). Four categories, exactly one per document, tie-break coarse before fine
|
||||
before duplicate: **too coarse / too fine / duplicate / correct**.
|
||||
|
||||
## Method
|
||||
|
||||
- **Corpus:** `K2/trinn1`, **N = 43** files, of which **39/43** are extractable.
|
||||
- **Sample:** **n = 12**, drawn by the method's own rule -- hex SHA-256 of the
|
||||
**NFC-normalised** filename, stratified by format (8 `pdf`, 3 `docx`,
|
||||
1 `xlsx`), re-derived this round from `tools/okf_outline_measure.py` and
|
||||
reproducing the twelve published documents in order, **12/12**.
|
||||
- **Three arms, one round:** Arm B is the shipping default (`--outline-run 0`),
|
||||
Arm D is `--outline-run 3`, Arm E is `--outline-run 3 --table-grid`. Arm B and
|
||||
Arm D are **re-rated** this round rather than carried over.
|
||||
- **Raters:** one first-rater identity over all three arms, **36 verdicts**;
|
||||
then **two blind raters per arm** at canonical positions 0, 2, 4, 6, 8, 10,
|
||||
`n_blind = 6` each -- **36 blind ratings and six `k/6` figures**.
|
||||
- **The blind protocol is WIDER than Arm D's and the two are not comparable.**
|
||||
`docs/2026-09-07-k3-arm-d.md` used "two separate blind raters, **one per
|
||||
arm**, 12 blind ratings and two `k/6` figures". The order asked for two per
|
||||
arm. Its number governs; the divergence is stated so the rounds' `k/6`
|
||||
figures are not read as like for like.
|
||||
- **Blindness is structural, not promised.** Each blind rater is a separate
|
||||
subagent with its own context, given ONE file: six documents labelled `A`-`F`
|
||||
with their extracted length and, per concept, its length, its title and the
|
||||
first 180 characters of its body. **Rule names were stripped**, because a
|
||||
`rule:table-grid` in the list would have identified the arm. No rater was told
|
||||
which arm it read, that other arms exist, what the first rater said, or that a
|
||||
brief, plan or report exists.
|
||||
- **Arm C's `8/4/0/0` is historical context and explicitly not a comparand.**
|
||||
|
||||
## Controls, passed before anything was counted
|
||||
|
||||
The controls are split by what they do on failure, and that split is the
|
||||
correction this round makes to its own first plan. A **gating** control asks
|
||||
whether the shipped rule is the rule being measured; it halts. A **prediction**
|
||||
is a figure written down in advance from an exploratory replica; it is reported
|
||||
whatever it says, because a replica may not sit in judgement over shipped code.
|
||||
|
||||
### Gating -- each one halts the round
|
||||
|
||||
| control | result |
|
||||
|---|---|
|
||||
| `git diff --stat 54a0bc2..HEAD -- src/` | only `propose.py` and `cli.py`; **136** changed lines in `propose.py`, most of them comments |
|
||||
| consumer bundle `K2-bundle-20260903`, file count | **1108** -- the literal published in the Arm D round |
|
||||
| the same bundle, `LC_ALL=C` aggregate digest | `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1` -- the published literal, before and after |
|
||||
| network imports, proposer and both instruments | **0** each |
|
||||
| **Arm B identity**, whole corpus, no flags | **byte-identical** to the archived Arm B plans, `diff -r` exit **0**, `_index.txt` INCLUDED |
|
||||
| **Arm D identity**, whole corpus, `--outline-run 3` | **byte-identical** to the archived Arm D plans, `diff -r` exit **0**, `_index.txt` INCLUDED |
|
||||
| artifact counts asserted BEFORE each diff | Arm B **28** json + **43** `.err` + **43** index lines; Arm D **33** + **43** + **43** |
|
||||
| door-level counts | `.err` files recording `FAILED`: **4** of 43, so extractable **39/43**, unchanged |
|
||||
| declared grid-rule totals | **38** lines on **3** documents, exactly as declared |
|
||||
|
||||
Two of these rows are corrections to the Arm D round's own published procedure,
|
||||
and both were found by review rather than by failure. The artifact counts are
|
||||
asserted **before** the diff, because `diff -r` over two trees where every
|
||||
document failed would compare nothing and exit 0. And `_index.txt` is
|
||||
**included** in the comparison: the Arm D reproduce block neither generates it
|
||||
nor compares it, which would let `NN.json` name a different document across two
|
||||
runs with nothing saying so.
|
||||
|
||||
The Arm B identity run is not bookkeeping either. It is the corpus-level half of
|
||||
the promise that `--outline-run 0` with the flag absent is still Arm B, and it
|
||||
is what lets the Arm B row below rest on verified bytes.
|
||||
|
||||
### Predictions -- reported, never gating
|
||||
|
||||
Written into the brief from an exploratory replica **before** the rule was
|
||||
built, and reproduced by the shipped code:
|
||||
|
||||
| prediction | measured |
|
||||
|---|---|
|
||||
| exactly 3 of 33 plans differ from the Arm D archive | **3** -- and they are the three named |
|
||||
| position 5: 21 -> 6 entries | **21 -> 6** |
|
||||
| position 10: 15 -> 3 entries | **15 -> 3** |
|
||||
| position 11: 2 -> 1 entries | **2 -> 1** |
|
||||
|
||||
**One prediction was allowed to halt, and only one:** if NO plan had differed,
|
||||
the flag would not have been threaded through to `find_candidates` and the row
|
||||
would have been a wiring bug wearing a null result's clothes -- with the ceiling
|
||||
below standing ready as a plausible wrong explanation. It did not occur.
|
||||
|
||||
The plan-to-position mapping is **derived, not assumed**: each plan's first
|
||||
entry `path` prefix is matched against the reduced stem of its source filename
|
||||
and that against the committed draw. The `NN` in `NN.json` comes from an
|
||||
unsorted shell glob and names nothing on its own.
|
||||
|
||||
## The denominator
|
||||
|
||||
Every figure below is stated against one of four denominators, and they are not
|
||||
interchangeable:
|
||||
|
||||
- **43** -- corpus files (the door-level denominator);
|
||||
- **39** -- extractable files (the segmentation denominator);
|
||||
- **12** -- the K3 sample, of which **at most 3 can move** (below);
|
||||
- **6** -- the blind positions, of which **1** is a document Arm E can move.
|
||||
|
||||
## The ceiling: at most 3 of 12, and this time it is measured
|
||||
|
||||
The Arm D round's ceiling was read off entry counts. That is not sound on its
|
||||
own: the orphan check can delete a table candidate before it becomes an entry,
|
||||
so a document could hold table rows that never reach a plan -- and position 0
|
||||
produces no plan at all, so its zero would be an absence with no denominator.
|
||||
|
||||
This round measures the ceiling in the text itself, with a committed instrument,
|
||||
over all 43 files and **before any rating began**:
|
||||
|
||||
| figure | value | denominator |
|
||||
|---|---|---|
|
||||
| documents with at least one table row | **3** | 39 |
|
||||
| table rows in total | **57** | -- |
|
||||
| documents with at least one grid-rule line | **3** | 39 |
|
||||
| grid-rule lines in total | **38** | -- |
|
||||
|
||||
A document with no table row cannot be moved by this arm, whatever the orphan
|
||||
check later did to its candidates. So **3 of the 12** sample documents can move,
|
||||
and they are positions 5, 10 and 11 -- three of the four the Arm D ceiling
|
||||
excluded, because they carry zero outline boundaries.
|
||||
|
||||
**The 57 is an independent corroboration and worth stating as one.** The Arm D
|
||||
round's pre-work control counted `|`-delimited table rows over the corpus text
|
||||
it screened and reported **57 of 35 050 lines**. This round's instrument, run
|
||||
against different code for a different purpose, counts **57**. Neither
|
||||
measurement was derived from the other.
|
||||
|
||||
## K3, the three rows side by side
|
||||
|
||||
### First rater, n = 12
|
||||
|
||||
| arm | too coarse | too fine | duplicate | correct | sum |
|
||||
|---|---|---|---|---|---|
|
||||
| Arm B (re-rated this round) | 8 | 4 | 0 | 0 | 12 |
|
||||
| Arm D (re-rated this round) | 4 | 5 | 0 | 3 | 12 |
|
||||
| **Arm E** | **4** | **3** | **0** | **5** | **12** |
|
||||
| *Arm C, 2026-09-04, historical only* | *8* | *4* | *0* | *0* | *12* |
|
||||
|
||||
Two labels changed, both among the three that could: positions 10 and 11 moved
|
||||
`too fine` -> `correct`. Position 5 did not move.
|
||||
|
||||
### Blind raters, n = 6 each, two per arm
|
||||
|
||||
| arm | rater | too coarse | too fine | duplicate | correct | sum | agreement with first rater |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| Arm B | `blind-1a` | 4 | 1 | 0 | 1 | 6 | **5/6** |
|
||||
| Arm B | `blind-1b` | 4 | 1 | 0 | 1 | 6 | **5/6** |
|
||||
| Arm D | `blind-2a` | 2 | 1 | 0 | 3 | 6 | **6/6** |
|
||||
| Arm D | `blind-2b` | 3 | 1 | 0 | 2 | 6 | **5/6** |
|
||||
| Arm E | `blind-3a` | 2 | 1 | 0 | 3 | 6 | **5/6** |
|
||||
| Arm E | `blind-3b` | 2 | 1 | 0 | 3 | 6 | **5/6** |
|
||||
|
||||
Within-arm agreement, which exists for the first time because there are two
|
||||
raters per arm: Arm B **6/6**, Arm D **5/6**, Arm E **6/6**.
|
||||
|
||||
**`too fine` is 1 of 6 in every arm, Arm E included.**
|
||||
|
||||
## Verdict
|
||||
|
||||
**On the first rater's row, Arm E moves `too fine` from 5 to 3 while `too
|
||||
coarse` stays at 4 -- the direction declared in advance. On the blind raters'
|
||||
rows, `too fine` does not move at all.** Both statements are measurements of the
|
||||
same twelve documents, and the report refuses to publish only the first.
|
||||
|
||||
The disagreement is one position and it is legible. At position 10 the first
|
||||
rater moved `too fine` -> `correct`; both Arm E blind raters kept `too fine`.
|
||||
Their reason CHANGED rather than persisting. Under Arm B and Arm D they object
|
||||
that thirteen table rows are severed from their header row. Under Arm E, where
|
||||
the table is one concept with its header included, they object that the table is
|
||||
severed from the sentence that introduces it. Arm E fixed the first complaint
|
||||
and does not touch the second.
|
||||
|
||||
So the honest form of the finding is a conditional, and it has two clauses:
|
||||
|
||||
- **If** a table is one unit of knowledge, Arm E moves K3 from 4/5/0/3 to
|
||||
4/3/0/5 and does it without trading a `too fine` for a `too coarse`.
|
||||
- **If** a table is a unit only together with the prose that introduces it, Arm
|
||||
E moves K3 by nothing on the position where both readings were tested, and
|
||||
what it changes is which severance you get, not whether you get one.
|
||||
|
||||
**And the second clause is measured on ONE document.** The blind positions are
|
||||
fixed at 0, 2, 4, 6, 8, 10, and only position 10 is a document Arm E can move.
|
||||
Positions 5 and 11 -- the other two -- were seen by no blind rater. The blind
|
||||
row is therefore not evidence that Arm E fails on those two; it is evidence that
|
||||
this protocol cannot see them. A round in which the arm's reach and the blind
|
||||
protocol's positions overlap in one document is a round whose blind row carries
|
||||
one document's worth of information about the arm, and no threshold should be
|
||||
set on that.
|
||||
|
||||
**The unit question Arm D surfaced is still open and is still the operator's.**
|
||||
It fired again here, on identical material: at position 8 one Arm D blind rater
|
||||
called top-level chapters `correct` and the other called them `too coarse`
|
||||
because their numbered subsections are "distinct requirement sets a reader would
|
||||
want separately". Nothing in the K3 method decides between those readings, this
|
||||
round does not decide it either, and it is prior to any threshold -- a threshold
|
||||
on an undecided unit measures the rater.
|
||||
|
||||
## What did move, with denominators
|
||||
|
||||
| figure | Arm B | Arm D | Arm E | denominator |
|
||||
|---|---|---|---|---|
|
||||
| entries, whole corpus | 618 | 709 | **681** | delta **-28** from Arm D |
|
||||
| entries carrying `rule:table-block` | 33 | 33 | **5** | of 681 |
|
||||
| entries carrying `rule:table-grid` | -- | -- | **5** | of 681 |
|
||||
| documents whose entry count changed | -- | -- | **3** | of 39 |
|
||||
| plans differing from the Arm D archive | -- | -- | **3** | of 33 |
|
||||
| blocks joined | -- | -- | **5** | -- |
|
||||
| documents producing an artifact | 28 | 33 | **33** | of 43 |
|
||||
| extractable | 39 | 39 | **39** | of 43 |
|
||||
| position 5 entries | 21 | 21 | **6** | -- |
|
||||
| position 10 entries | 15 | 15 | **3** | -- |
|
||||
| position 11 entries | 2 | 2 | **1** | -- |
|
||||
| largest span Arm E creates | -- | -- | **13 691** chars | position 10 |
|
||||
|
||||
**Concept paths for unchanged content did not churn, and the denominator is
|
||||
computed rather than declared.** Of the **676** `(plan, span)` pairs present in
|
||||
both Arm D and Arm E, **0** received a different concept path. An earlier draft
|
||||
of this round's plan declared 686 as the expected denominator; that was wrong
|
||||
twice over -- 28 entries are removed, not 33, and a joined block's `end` moves
|
||||
so its pair matches nothing in Arm D. The computed 676 is the figure of record,
|
||||
and the wrong 686 is recorded rather than deleted.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
**The character class was fitted to the same three documents it is measured on.**
|
||||
Arm D's run-length 3 came from a corpus-wide distribution of 328/37/18. Arm E's
|
||||
`[-=:+]` came from 38 lines drawn entirely from the three documents that are 100
|
||||
% of its movable sample. The whole-corpus screen above tests generalisation
|
||||
outward -- it found no fourth grid-bearing document -- but it cannot break that
|
||||
circularity inward, and no reading of the rows should treat "declared, not
|
||||
swept" as meaning the same thing it meant for Arm D.
|
||||
|
||||
**`rule:table-grid` is plan-level provenance and does NOT reach the bundle.**
|
||||
Measured on the 1 294-file Arm D bundle: `grep -rl "PROPOSED"` returns **0**,
|
||||
while `derived` appears in 84 files as a frontmatter key with other values. A
|
||||
reviewer looking for the rule name in a built bundle will find nothing, and that
|
||||
absence is a property of materialisation, not evidence that the flag did not
|
||||
fire.
|
||||
|
||||
**The ceiling is bounded by which table FORM the converter chose, not by how
|
||||
many tables the corpus holds.** Pandoc also emits *simple* and *multiline*
|
||||
tables, whose rows carry no `|` at all. `_TABLE_ROW` never sees those, so they
|
||||
are invisible to the table rule, to Arm E, and to the `|`-row count that
|
||||
measures the ceiling. One document in this sample (position 10) contains such a
|
||||
table in its upper half, and no arm proposes a boundary in it.
|
||||
|
||||
**Two grid tables separated by a rule line alone would merge into one concept.**
|
||||
Pandoc puts a blank line between adjacent tables, so it does not emit that
|
||||
shape -- but that is a property of the WRITER, not of this code, and `in_table`
|
||||
survives an arbitrary run of rule lines. A unit fixture asserts the merge, so
|
||||
the limit is declared rather than assumed away. The corpus diff found no
|
||||
instance.
|
||||
|
||||
**Position 5 is the document that shows what Arm E is not.** It has the largest
|
||||
reduction in the round, 21 concepts to 6, and its LABEL DOES NOT CHANGE. Each of
|
||||
its three references is still cut into a 114-character title concept carrying a
|
||||
heading and no body, plus its 1 675-character table. Joining table rows removed
|
||||
most of the fragmentation and left the rest; the remaining cut comes from the
|
||||
heading rule, not from the table rule.
|
||||
|
||||
**Position 7's cause is diagnosed and deliberately unbuilt.** Its 48 concepts
|
||||
include nine contents-listing lines with dotted leaders, and concepts of 87, 93
|
||||
and 112 characters. It carries zero table-block entries, so Arm E cannot reach
|
||||
it. One change per measurement is the Arm C lesson; the fix is named and not
|
||||
made.
|
||||
|
||||
**The orphan gate still deletes 34 % of Arm D's own boundaries, skewed.**
|
||||
Reported in the Arm D round, unfixed, and untouched here.
|
||||
|
||||
**No bundle was built for Arm E.** The Arm D round's door-level counts came from
|
||||
a bundle run; here they come from the run artifacts themselves -- 4 of 43 `.err`
|
||||
files record `FAILED`, so extractable is 39/43 -- which is the same figure for
|
||||
roughly a twentieth of the wall time. Bundle-level file counts are therefore not
|
||||
reported for Arm E, and that is a gap, not a result.
|
||||
|
||||
**Position 0 is unreadable and no segmentation changes that.** It extracts as
|
||||
95.1 % `(cid:N)` glyph tokens. It is an extraction defect, K3 measures
|
||||
segmentation, and it did not move a label in any arm.
|
||||
|
||||
**All raters are instances of the same model family, and the first rater is not
|
||||
independent.** Agreement between them bounds self-consistency, not correctness.
|
||||
The first rater had read the Arm D report's published rows before rating, so Arm
|
||||
B reproducing `8/4/0/0` and Arm D reproducing `4/5/0/3` is consistency and not
|
||||
confirmation. What the re-rating establishes is narrower and sufficient for this
|
||||
comparison: all three arms were judged in one round, by one identity, against
|
||||
the same four categories.
|
||||
|
||||
**No CHANGELOG entry accompanies this arm.** Measured precedent, re-checked this
|
||||
round: `grep -c` for "outline-run", "max-segment-chars", "Arm C" and "Arm D"
|
||||
returns 0 in both `README.md` and `CHANGELOG.md`, while `--path-prefix` -- a real
|
||||
interface change -- has a CHANGELOG entry. The rule is "interface and behaviour
|
||||
changes yes, arm flags no", and `--table-grid` is an arm flag that defaults off.
|
||||
|
||||
## Reproducing
|
||||
|
||||
```
|
||||
C=~/corpora/okf-telling-20260829
|
||||
|
||||
# The corpus loop, in one place. It writes NN.json, NN.err and one
|
||||
# `i|exit|filename` line per document into _index.txt -- which is the format
|
||||
# both archives carry, and which the Arm D round's published block omitted.
|
||||
arm_run() { # $1=outdir $2=proposed-at $3=lo $4=hi then flags
|
||||
DIR=$1; AT=$2; LO=$3; HI=$4; shift 4; mkdir -p "$DIR"; i=0
|
||||
for f in "$C"/K2/trinn1/*; do
|
||||
i=$((i+1))
|
||||
[ "$i" -lt "$LO" ] && continue
|
||||
[ "$i" -gt "$HI" ] && continue
|
||||
b=$(basename "$f"); n=$(printf '%02d' "$i")
|
||||
.venv/bin/python tools/okf_propose_segments.py "$f" --out "$DIR/$n.json" \
|
||||
--path-prefix "${b%.*}" --proposed-at "$AT" "$@" 2> "$DIR/$n.err"
|
||||
echo "$i|$?|$b" >> "$DIR/_index.txt"
|
||||
done }
|
||||
|
||||
# Run in ascending chunks, or _index.txt line order breaks. Each chunk is a
|
||||
# foreground call under 600 s; documents 13-28 account for most of the time.
|
||||
for lo_hi in "1 12" "13 20" "21 28" "29 36" "37 43"; do
|
||||
set -- $lo_hi
|
||||
arm_run "$C/K2-plans-armB-check-20260907" 2026-09-03T00:00:00Z "$1" "$2"
|
||||
arm_run "$C/K2-plans-armD-check-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3
|
||||
arm_run "$C/K2-plans-armE-20260907" 2026-09-07T00:00:00Z "$1" "$2" --outline-run 3 --table-grid
|
||||
done
|
||||
|
||||
# The identity halves. Assert the counts FIRST: a diff over two trees where
|
||||
# every document failed compares nothing and exits 0.
|
||||
ls "$C"/K2-plans-armB-check-20260907/*.json | wc -l # 28
|
||||
ls "$C"/K2-plans-armD-check-20260907/*.json | wc -l # 33
|
||||
diff -r "$C/K2-plans-baseline-20260907" "$C/K2-plans-armB-check-20260907" -x '*.err'; echo $?
|
||||
diff -r "$C/K2-plans-armd-20260907" "$C/K2-plans-armD-check-20260907" -x '*.err'; echo $?
|
||||
|
||||
# The measurement: exactly three plans differ.
|
||||
diff -rq "$C/K2-plans-armd-20260907" "$C/K2-plans-armE-20260907" -x '*.err' -x '_index.txt'
|
||||
|
||||
# The ceiling, over all 43 files.
|
||||
.venv/bin/python tools/okf_table_measure.py \
|
||||
--corpus "$C/K2/trinn1" --report "$C/K2-table-reach-20260907.md"
|
||||
|
||||
# The door count, without building a bundle.
|
||||
ls "$C"/K2-plans-armE-20260907/*.err | wc -l # 43
|
||||
grep -l FAILED "$C"/K2-plans-armE-20260907/*.err | wc -l # 4 -> 39/43
|
||||
|
||||
cd "$C" && LC_ALL=C find K2-bundle-20260903 -type f | LC_ALL=C sort \
|
||||
| xargs shasum -a 256 | shasum -a 256
|
||||
```
|
||||
|
||||
The consumer bundle, locale-pinned, before and after this round:
|
||||
|
||||
- **1108 files**, `9cd745194346cda0c70eab9c7136fa44506203bbe85bc17d7eff2766c6e9b4d1`
|
||||
|
||||
`LC_ALL=C` is not decoration: without it `sort` orders the file list differently
|
||||
and the aggregate digest changes while the bytes do not.
|
||||
|
||||
## Appendix: the twelve first-rater verdicts, three arms
|
||||
|
||||
| pos | document | B entries | D entries | E entries | Arm B | Arm D | Arm E |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 0 | Bilag 9.1 | 1 concept | 1 concept | 1 concept | too coarse | too coarse | too coarse |
|
||||
| 1 | Bilag 3.2.2 | 20 | 23 | 23 | too coarse | too coarse | too coarse |
|
||||
| 2 | Bilag 1.1 | 1 concept | 9 | 9 | too coarse | correct | correct |
|
||||
| 3 | Bilag 7 | 1 | 3 | 3 | too coarse | too coarse | too coarse |
|
||||
| 4 | Vedlegg 5 | 1 concept | 4 | 4 | too coarse | too coarse | too coarse |
|
||||
| 5 | Vedlegg 3 | 21 | 21 | **6** | too fine | too fine | too fine |
|
||||
| 6 | Bilag 3.8 | 6 | 7 | 7 | too coarse | correct | correct |
|
||||
| 7 | Bilag 1.3 | 45 | 48 | 48 | too fine | too fine | too fine |
|
||||
| 8 | Bilag 3.4 | 1 concept | 8 | 8 | too coarse | correct | correct |
|
||||
| 9 | Bilag 5 | 5 | 11 | 11 | too coarse | too fine | too fine |
|
||||
| 10 | Vedlegg 1 | 15 | 15 | **3** | too fine | too fine | **correct** |
|
||||
| 11 | Dokument for avtaleinngaelse | 2 | 2 | **1** | too fine | too fine | **correct** |
|
||||
|
||||
"1 concept" means no plan was written: the mechanical rules found no boundary
|
||||
and the document lands as one flat concept.
|
||||
|
||||
The blind raters covered positions 0, 2, 4, 6, 8, 10 only. They disagreed with
|
||||
the first rater at position 6 on Arm B (both raters, `correct` where the first
|
||||
rater says `too coarse` -- his ground is a chapter that Arm B ABSORBS and that
|
||||
therefore does not appear in the material a blind rater sees), at position 8 on
|
||||
Arm D (one rater of two), and at position 10 on Arm E (both raters, `too fine`
|
||||
where the first rater says `correct`). They agreed with the first rater and with
|
||||
each other everywhere else.
|
||||
Loading…
Add table
Add a link
Reference in a new issue