Three rows rated in one round by one first-rater identity, then two blind
raters per arm as the order asked -- 36 first-rater verdicts, 36 blind ratings,
six k/6 figures. No threshold is set and none is implied.
First rater: Arm B 8/4/0/0, Arm D 4/5/0/3, Arm E 4/3/0/5. Against the direction
declared before the row was seen -- fewer `too fine` WITHOUT more `too coarse`
-- `too fine` goes 5 -> 3 and `too coarse` stays 4.
Blind raters: `too fine` is 1 of 6 in EVERY arm, Arm E included. The report
publishes both rows and states the verdict as a conditional, because they are
measurements of the same twelve documents and only one of them is favourable.
The disagreement is one position and it is legible. At position 10 both Arm E
blind raters kept `too fine` where the first rater moved to `correct`, and
their REASON changed rather than persisting: under Arm B and D they object that
thirteen table rows are severed from their header; under Arm E, where the table
is one concept with its header, they object that the table is severed from the
sentence that introduces it. Arm E fixed the first complaint and does not touch
the second.
And the denominator that governs how much that can say: the blind positions are
fixed at 0, 2, 4, 6, 8, 10, and only ONE of Arm E's three movable documents
falls in that set. The blind row carries one document's worth of information
about this arm. No threshold should be set on that, and the report says so.
The ceiling is measured this round rather than inferred. A committed instrument
extracted all 43 files and counted table ROWS in the text -- 3 of 39 documents,
57 rows -- because zero ENTRIES is not the same claim: the orphan check can
delete a table candidate before it becomes an entry, and one sample document
produces no plan at all so its zero would be an absence with no denominator.
The 57 reproduces the Arm D round's independent pre-work screen exactly, with
neither measurement derived from the other.
Controls are split by what they do on failure, which is this round's correction
to its own first plan: gating controls halt, predictions are reported whatever
they say. Only one prediction was allowed to halt -- a zero-diff run, which
would be a wiring bug wearing a null result's clothes, with the ceiling standing
ready as a plausible wrong explanation.
Nine limits are stated rather than left implicit, including three that weaken
the arm's own case: the character class was fitted to the same three documents
it is measured on; position 5 has the round's largest reduction (21 concepts to
6) and its label does not change; and pandoc's simple and multiline table forms
carry no pipes at all, so the ceiling is bounded by which form the converter
chose rather than by how many tables the corpus holds.
Pure ASCII (0 non-ASCII bytes), short document labels only (0 corpus filenames).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>