llm-ingestion-okf/docs/2026-09-02-k3-k4-k5-baseline.md

6.4 KiB

K3, K4 and K5 — the numbers, on the method committed before them

The method is docs/2026-09-02-k3-k4-k5-metode.md, committed in the preceding commit. It is not revised here. Where adjudication showed something the method did not anticipate, it is stated below as a finding; reopening the method file in this commit would erase the ordering the split exists to prove.

K3 — the four category counts, n = 12

category count
too coarse 7
too fine 3
duplicate 1
correct 1

Sum: 12, which is the n the method committed. No replacement was needed: all twelve drawn documents were in the frame.

The frame's proposals came from tools/okf_propose_segments.py run over the sample; four of the twelve proposals were empty — the tool proposed no split at all — and an empty proposal means the whole document would become one concept. All four fell in the first row.

No threshold is applied. The method declared none, and none is invented here.

Blind second rater and agreement

n_blind = 6, the even positions of the sample's canonical hash order, judged in a separate context by claude-opus-5/blind-rater-1 with the proposals, the source documents, the four category definitions and the tie-break rule, and without sight of the first rater's verdicts. The first rater's labels were written to a file before the second rater's were read.

Agreement: 5/6.

The single disagreement is worth more than the ratio. On one document the first rater judged the six proposed section boundaries to match the document's sections; the second rater judged the same proposal too coarse, and named its evidence — one numbered chapter has no segment of its own and is absorbed into the preceding segment's span, and the document's introduction is covered by no segment at all. That is a specific, checkable claim, and the first rater did not check at that resolution. The verdict is not revised — a first rater who edits their label after reading the second rater's is not a second rater at all — but the asymmetry is the finding: the disagreement was resolved by evidence in one direction only, which suggests the true count of too coarse is a floor rather than an estimate.

The method's independence caveat stands and should be read with the 5/6: both raters are instances of the same model family, so agreement overstates independence.

K4a — determinism, binary

PASS. One document was adjudicated with tools/okf_adjudicate.py, then run through the path twice into two bundles, and the two compared:

diff -r /tmp/k4a/a /tmp/k4a/b     # exit 0, no output

Six concepts and an index were written on each run, byte-identical, under SEGMENTED_OKF_V0_2 with bundle_id supplied by the caller.

K4b — decisions and minutes per document, with the denominator

Denominator: 12 documents, carrying 115 proposed entries between them.

figure value
documents adjudicated 12
entry-level decisions 115
wall clock, whole adjudication 314 s
per document 26.2 s
per entry-level decision 2.73 s

This is not a human throughput figure and must never be quoted as one. The adjudicator is the model identity the method named. A human reading 115 proposed segment boundaries against their source documents would not take 26 seconds per document, and nothing here measures that.

A finding the method did not anticipate

4 of 12 verdicts could not be recorded at all. tools/okf_adjudicate.py parses the proposal before writing, and the parser refuses a plan with zero entries — correctly, because an empty plan replayed on the run path would persist nothing for a document that was dropped. The consequence is that the one judgement most worth recording about those four documents — the proposal is empty, and that is wrong — has no artifact to live in. The K3 counts above include all twelve because a category is a judgement, not a file; the adjudication record covers 8/12.

This is stated, not fixed: fixing it is a change to the adjudication tool or to the plan grammar, and neither belongs in a measurement commit.

K4c — binary, and it is not a share

Yes. A consumer can see the adjudication state and discount on it. The evidence is the test committed in Step 14, which pins the adjudication key, its closed value set and its companion keys: tests/test_segmented_index.py:338 (an unratified segment carries proposed), :346 (a ratified one carries adjudicator, time and dwell), :359 (the state is projected as an index facet) and :410 (the older profile writes no adjudication state at all, so its absence is distinguishable from proposed).

The K4a run confirms the same thing on real output: each of its six concepts carries adjudication: adjudicated, adjudicated_by, adjudicated_at and adjudication_dwell_s in its frontmatter.

A distribution fact, reported as such and not as K4c: in the corpus bundle from the 2026-09-02 run, 0 of 39 concepts carry any adjudication state, because that run used STRUCTURED_V1, which writes none. In the K4a bundle, 6 of 6 do. Neither number answers K4c's question, which is why the two were previously confused.

K5 — SPEC § 5.1 provenance

0 of 39 concepts in the corpus bundle carry any SPEC § 5.1 source entry. The measured keys — sources, sources[].resource, sources[].id, sources[].title, sources[].author, usage_count, last_modified — are all present on 0 concepts. Door B under STRUCTURED_V1 records provenance instead as source_file and source_sha256, present on 39 of 39; those are this repository's keys, not § 5.1's, and a consumer reading for § 5.1 finds nothing.

The known ceiling stands and bounds what a different profile could achieve here: 2 of 5 keys on a single source. A dropped file yields a resource and a title; the credibility keys (author, usage_count, last_modified) have no data behind them on an inbox drop, so no profile can fill them by writing more.

No threshold is applied to K5. The method declared none.

Out of scope, restated

K4 over time is PM-owned and out of scope here. This work produces one baseline. Comparing baselines across runs is a programme-level question the rubric's owner holds, and answering it from inside this repository would be answering it in the wrong place.