8.1 KiB
K1 and K2 over the K2 corpus (trinn 1), 2026-09-02
What this is: the output of two shipped instruments run over a public procurement corpus, written down with the denominators visible. It reports counts, error-code distributions, format classes and per-format fidelity figures. It does not reproduce document bodies or full filename lists — the corpus is public, but this repository's rule about consumer content binds the shape of what a tracked file here carries, and nothing in the measurement needs a document's text to be readable.
Both numbers below can be produced again by anyone with the corpus:
.venv/bin/python tools/okf_corpus_run.py \
--corpus ~/corpora/okf-telling-20260829/K2/trinn1 \
--report /tmp/k1k2.md --bundle /tmp/k2-trinn1-bundle
.venv/bin/python tools/okf_fidelity.py <the corpus's .docx and .xlsx files>
Converter, as the harness resolved and reported it — not as configuration
claims it: .venv/lib/python3.14/site-packages/pypandoc/files/pandoc, version
3.9, the pinned one. The host's own pandoc (3.10.2) was not used; that is
the whole reason _pandoc.py resolves a path and freezes a version.
K1 — what the door did with 43 files
N = 43 is the corpus directory's file count, computed by the harness. It is
the denominator for every number in this section.
Three counts, never one, because the guard sits between extraction and persist and a healthy persisted count can hide a pile of quarantines:
| count | value |
|---|---|
| extracted | 39/43 |
| gated | 39/43 |
| persisted | 39/43 |
The numerator split. A merge is degenerate when the extracted text is zero characters after stripping whitespace — a definition, not a threshold:
| class | value |
|---|---|
| substantive | 39/43 |
| degenerate | 0/43 |
| rejected (coded) | 4/43 |
K1b holds exactly: 39 + 4 = 43 = N. This is not an assertion a reader has
to trust. The harness exits non-zero and names the unaccounted files when the
identity fails, so the run itself is the check; this run exited 0.
Rejection codes, all four accounted for:
| code | value |
|---|---|
extractor_unknown |
3/43 |
extractor_empty_pdf |
1/43 |
The structural ceiling
Three of the four rejections cannot merge under any design in scope here: one
.smc and one .zip (neither is a document this library claims to read), and
one .doc the converter cannot read at all. The fourth is a PDF that extracts
to nothing. Nine further PDFs in the corpus are drawing-dominated: they merge,
and their figures do not survive extraction, which every pdf extraction warns
about. They are counted as substantive merges because they carry text; a reader
should not read that as "the drawing arrived".
Wall time
The harness's own figure for the whole inbox pass: 1175.28 s total, 27.332 s per file over 43 files. This is the only evidence the scale NFR has. It is dominated by PDF extraction (33 of 43 files); the run's second pass, which re-extracts to classify degeneracy, roughly doubles the wall clock and is a property of the instrument, not of the door.
Format classes, and the three rows that exit unmeasured
| format | files in corpus | K2 measured |
|---|---|---|
pdf |
33 | no — out of the K2 instrument's scope |
docx |
5 | yes |
xlsx |
2 | yes |
doc |
1 | no — converter cannot read it |
zip |
1 | n/a |
smc |
1 | n/a |
pptx |
0 | denominator 0 — unmeasured by construction |
odt |
0 | denominator 0 — unmeasured by construction |
rtf |
0 | denominator 0 — unmeasured by construction |
Denominator 0 for pptx, odt and rtf: the corpus contains no file of
those three formats at all. Three of the five office rows therefore leave
this work unmeasured — not unmet, and not "passing". No substitute
corpus was sought.
Comparability with the 2026-08-29 arm A count
The earlier PDF figures were produced with pdftotext; this library extracts
PDFs with pdfplumber. 33 of 43 files are therefore not comparable to the
earlier count, and no attempt is made here to compare them. The docx and
xlsx rows are the ones that carry over, and even they carry over only as
context — see the denominator note under K2.
K2 — product path and standalone converter, side by side
Instrument: tools/okf_fidelity.py, which asks two questions about the source
document's own strings — coverage (of the strings the file stores, how many
appear in the converted text) and pairing (of the source rows carrying a
label, how many keep label and value on one output line). The fasit is read
from the package's own XML, never from a converter's output.
Two legs, both with their denominators: the standalone leg is what the
pinned converter can do at all, the product path is what this library
actually produces through extract_text.
| format | leg | coverage | paired |
|---|---|---|---|
docx |
standalone | 184/189 | 19/33 |
docx |
product path | 184/189 | 19/33 |
xlsx |
standalone | 179/179 | 0/1 |
xlsx |
product path | 179/179 | 0/1 |
The fall from standalone to product is zero, on every one of the seven files individually and in the sum. There is nothing to attribute, because nothing was lost between the converter and the bundle. That is the result the acceptance rule was written to force someone to look for, and this time it came back clean.
The denominators are not arm A's, and that is a named difference
Arm A published docx 193/196 coverage and 51/53 paired, and xlsx
193/193 coverage and 98/98 paired. Those denominators (196, 53, 193, 98)
do not match this instrument's (189, 33, 179, 1). This is not a fall from
standalone to product — both legs above were measured by the same instrument,
and they agree exactly. It is a different fasit: arm A's counting command was
never shipped, which is precisely why one was built, and a denominator that
differs on the source side cannot be a converter or pipeline effect. The
largest gap, xlsx pairing 98 versus 1, is an instrument limitation worth
stating plainly: this instrument approximates "a row expecting a value" by the
colon rule over the workbook's deduplicated shared strings, and an xlsx
shared-string table does not preserve row structure, so almost no xlsx row
qualifies. The xlsx pairing figure here is 0/1 — a denominator of one — and
should be read as not measured on this corpus, not as a failure.
Arm A's figures are recorded here as context. They are not this run's standalone leg, and they are not comparable to it.
What a coverage count cannot see
Coverage asks whether a source string reached the output. Two properties of the corpus price sheet are invisible to that question:
- 207 merged regions in the workbook. The text output flattens the grid; the merged cell's string appears once either way, so coverage scores the same whether the geometry survived or not. It did not survive.
- Number formats. 54 numeric cells carry the accounting/currency format and 2 carry a percent format. Formatting is a style attribute, and the converted text carries the stored value only: a kroner amount arrives as a bare number with nothing marking it as kroner, and coverage counts it as present. A reader of the bundle cannot recover the unit.
This is the same failure class the structured-table work already declared out of scope: the loss is document geometry, not a converter choice.
What this run does not say
- Nothing about
pptx,odtorrtf— denominator 0. - Nothing about
pdffidelity. The K2 instrument reads office packages; PDF fidelity was measured separately on 2026-08-21 with a different tool, against a different extractor, and is not restated here. - Nothing about a threshold. No K2 figure here is compared against a bar, because no bar has been ratified.