Leg 2 of the approved sequence, order 20260829T071322Z-683659895-from-.claude. Arm A of the acid test, never run before because the vegnormal PDFs were rastered 11/11. Measurement only: no parser, no src/ change, no runtime dependency. docling was installed as a measuring tool in an isolated venv. pandoc 3.10.2 ships an xlsx reader. The 08-29 count reached for markitdown on xlsx because that reader was assumed absent; it is not. G2 holds on the Excel half: the 207 merged ranges in Bilag 7 Prisskjema pair 61/61 for all three converters. Merged geometry inflates volume without breaking label-value pairing. pandoc takes docx (193/196 source-string coverage, 51/53 paired, 1/1 footnote) and xlsx (193/193, 98/98, at 9 979 non-whitespace chars against docling's 58 875 for identical coverage - docling replicates a merged cell into every spanned column, 8x on the two measured). markitdown corrupts embedded cell newlines into the literal two-character \n; correcting for that moves its xlsx coverage from 184/193 to 191/193. No converter beats the incumbent pdftotext on pdf: pandoc has no reader (rc 21, 33/33), markitdown pairs 534/9 027 (5.9 %, same pdfminer.six failure this repo measured independently on 2026-08-21), and docling's PDF path is blocked - transformers 5.8.1 needs torch >= 2.4 and no x86-64 macOS wheel exists past 2.2.2. Reported as 3 attempted, 3 failed, 30 not attempted, because the failure precedes any file read. docling is the only converter returning content from the Word 97 .doc. markitdown returns exit 0 with a wrong document class on it, and a stub on .smc/.zip; docling returns exit 0 with no file at all on those two. Denominators come from the source's own OOXML and from raw page geometry, never from a converter under test. Notes have no denominator here - the corpus holds 0 real footnotes or comments - so capability is characterised against an injected known-positive instead, where docling drops the footnote 0/1 while the other two capture it. Stated as a measured gap, not a footnote: K1 carried the revision pairs, K2's two stages are byte-identical, so kravspek B1 4.7 (a) utgavepar is uncovered by this corpus. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011Ge8kWitAAx5bJh7FW9r3u
16 KiB
Conversion arm (arm A): pandoc vs markitdown vs docling on K2
Date: 2026-08-29 · Order: 20260829T071322Z-683659895-from-.claude
(from .claude) · leg 2 of the approved sequence in
.claude/docs/2026-08-28-scenarioanalyse-portfolio-optimiser.md § 6 pt 2.
Measurement only. No parser was written, no seam built, no src/ change, no
runtime dependency added, no bundle, no release, no push. docling was
installed as a measuring tool in an isolated venv, never into
pyproject.toml. No OCR was used or installed — the corpus has 0 raster PDFs,
so docling ran with --no-ocr.
This is the arm that was never run in okf-bundle-prosessen.md § 9.5, because
the vegnormal PDFs were rastered 11/11. This corpus is not.
0. Headline
| Question | Answer |
|---|---|
Best converter for docx |
pandoc — 193/196 cell coverage, 51/53 paired, 1/1 footnote |
Best converter for xlsx |
pandoc — 193/193 coverage, 98/98 paired, least noise |
Best converter for pdf |
none of the three beat the incumbent pdftotext |
Best converter for doc (Word 97) |
docling — the only one that returns content |
| G2 on the 207-merged-range price form | 61/61 paired, by all three — merges do not break G2 |
docling PDF path |
NOT MEASURED — blocked by hardware, see § 2 |
| Revision pair (kravspek B1 § 4.7 (a)) | UNCOVERED by this corpus, see § 7 |
The single most consequential finding is not a ranking: pandoc 3.10.2
ships an xlsx reader (and a pptx one). The 08-29 count reached for
markitdown on xlsx because the reader was assumed absent. It is not.
1. Corpus, verified before use
The order says to verify rather than re-download. ~/corpora/okf-telling-20260829/
is present. integrity.tsv carries sizes, not hashes, so the sha256 claim that
can actually be re-checked is the report's cross-stage identity claim.
| Check | Command | Result |
|---|---|---|
| Files per stage | find K2/trinn1 -type f | wc -l |
43 (and 43 in trinn2) |
| Bytes | stat -f%z summed over trinn1 |
247 107 243 — matches 08-29 exactly |
| Stages byte-identical | shasum -a 256 pairwise, 43 files |
43 SAME, 0 DIFF, 0 MISSING |
| PDF pages | independent re-count via pdftotext -bbox-layout |
844 — matches 08-29 exactly |
Two independently reproduced numbers (bytes, pages) from a different command
path than the original. Nothing was re-downloaded. K1/ is empty, as expected.
2. Availability and exit status, per converter
| Converter | Version | docx |
xlsx |
pdf |
doc |
smc |
zip |
|---|---|---|---|---|---|---|---|
pandoc |
3.10.2 | rc 0 (5/5) | rc 0 (2/2) | rc 21 (33/33) | rc 21 | rc 95 | rc 21 |
markitdown |
0.0.2 | rc 0 (5/5) | rc 0 (2/2) | rc 0 (33/33) | rc 0 † | rc 0 † | rc 0 † |
docling |
2.123.1 | rc 0 (5/5) | rc 0 (2/2) | rc 1 (3/3) | rc 0 | rc 0 ‡ | rc 0 ‡ |
pandoc rc 21 is an honest refusal with a message: Unknown input format 'pdf' — Pandoc can convert to PDF, but not from PDF. --list-input-formats confirms
xlsx and pptx are present and pdf is not.
† markitdown returns exit 0 with false content. On the Word 97 .doc it
emits # Email Message / ## Content — 28 characters, and a wrong document
class. On .smc (86 MB Solibri model) and .zip it emits a one-line stub.
A zero exit code here describes the wrapper, not the work.
‡ docling returns exit 0 with no file on .smc/.zip. It does log
ERROR ... format None does not match any allowed format, so the refusal is
stated — but the exit code contradicts the log. Different defect from
markitdown's, and worth separating: one lies silently, one lies only in $?.
docling's PDF path is blocked, and the block is environmental
Installed successfully (uv venv + uv pip install docling, exit 0). docx
and xlsx convert fine — those use SimplePipeline, which needs no model.
The PDF pipeline needs the layout model, which needs torch:
ImportError: AutoImageProcessor requires the PyTorch library
[transformers] Disabling PyTorch because PyTorch >= 2.4 is required but found 2.2.2
transformers 5.8.1 requires torch >= 2.4. The installed torch is 2.2.2
because that is the last version with an x86-64 macOS wheel. Measured, not
assumed:
uv pip install "torch>=2.4" --dry-run
-> hint: Wheels are available for `torch` (v2.13.0) on the following
platforms: manylinux_2_28_aarch64, manylinux_2_28_x86_64,
macosx_14_0_arm64, win_amd64
This machine is macosx-26.0-x86_64. Downgrading to numpy<2 fixed the ABI
warning and did not change the outcome.
Denominator: 3 PDFs attempted, 3 failed, 30 not attempted. The failure is raised while constructing the pipeline, before any file is opened, so it is file-independent; three files across the size range (2, 7 and 150 pages) establish that. This is stated as a gap, not papered over: on arm64 or Linux this arm would produce a docling PDF number, and it does not here.
3. Denominators, and where they come from
No converter under test defines its own denominator.
docx/xlsxtable rows and notes — read from the source's own OOXML (word/document.xml,word/footnotes.xml,xl/viaopenpyxl3.1.5). A table row is one with ≥ 2 non-empty cells; label = first, value = last. → 161 rows (docx 53, xlsx 108); 151 usable after the key rule below.pdftable rows — frompdftotext -bbox-layout, which emits word bounding boxes. Coordinates are not a rendering decision. Poppler puts each cell in its own<line>, so visual rows are rebuilt by clustering words on vertical centre (tolerance 0.5× median word height). A row qualifies on a column gap ≥ max(3× page median inter-word gap, 1.5 % page width, 8 pt). → 844 pages, 31 095 visual rows, 9 961 row candidates; 9 027 usable.- Pairing rule (disputable, so stated):
label_key= last ≤ 3 words of the label (≥ 4 chars),value_key= first ≤ 3 words of the value. A row is paired if one output line holdslabel_keyand thenvalue_keyafter it. Markdown pipes are flattened first, so| a | b |counts as one line.
Known-positives, per L93. The PDF row query returned 0 on the first
attempt; that was my query, not the data — it assumed a row was one poppler
<line>. After the fix it returns 103 rows on the 3-page room list, from 188
visual rows. The note extractor finds 2/2 injected notes in a docx and 1/1 in
an xlsx from the same directory and file class, and 0 in the real files. The
pairing matcher finds 534 pairs in markitdown's PDF output, so its low score
below is a measurement, not a matcher that cannot see that format.
One trap: pdftotext exits 0 while emitting a raw 0x0E byte inside a
<word>, which is not well-formed XML. One of 33 files. The parser was fixed
to strip XML-invalid control bytes; the converter was not at fault.
4. (a) Text preserved and (d) characters out
Raw characters out, and — because a bigger number is not automatically better — non-whitespace characters, plus coverage of the source's own cell/paragraph strings, which is what "preserved" actually means.
| Format | Converter | (d) chars | non-ws chars | source-string coverage |
|---|---|---|---|---|
docx |
pandoc |
31 710 | 19 895 | 193/196 (98.5 %) |
docx |
docling |
21 228 | 10 825 | 189/196 (96.4 %) |
docx |
markitdown |
11 598 | 9 356 | 187/196 (95.4 %) |
xlsx |
docling |
379 296 | 58 875 | 193/193 (100 %) |
xlsx |
pandoc |
103 829 | 9 979 | 193/193 (100 %) |
xlsx |
markitdown |
14 578 | 10 632 | 191/193 (99.0 %) |
pdf |
markitdown |
1 752 548 | — | see § 5 |
pdf |
pandoc / docling |
no reader / blocked | — | — |
doc |
docling |
2 919 | — | real structured text |
doc |
markitdown |
28 | — | wrong document class |
docling's 6× volume on xlsx is noise with a named mechanism: it
replicates a merged cell's value into every spanned column. Two long cells in
the price form appear 8 times each in docling's output and once each in
pandoc's and markitdown's. For a workbook with 207 merged ranges that is
the whole difference between 58 875 and 9 979 non-whitespace characters, at
identical coverage.
markitdown does not lose the 9 long cells it first appeared to lose — it
corrupts them. An embedded line break inside a cell is emitted as the literal
two-character sequence \n. Correcting for that moves xlsx coverage from
184/193 to 191/193; 2 cells are genuinely absent. The corrupted 9 are the price
form's instruction cells, the ones carrying commercial obligation. Same defect
family as the escape-sequence tokens found in session 73.
On docx, volume does not predict fidelity. pandoc emits 2.7× more
characters than markitdown and gains 6 source strings; all three pair 51/53.
5. (b) Label + value on the same row — the G2 criterion
| Format | Converter | Paired | Usable rows | Share |
|---|---|---|---|---|
docx |
pandoc / markitdown / docling |
51 | 53 | 96.2 % each |
xlsx |
pandoc / markitdown / docling |
98 | 98 | 100 % each |
xlsx — Bilag 7 Prisskjema alone |
all three | 61 | 61 | 100 % each |
pdf |
markitdown |
534 | 9 027 | 5.9 % |
pdf |
pandoc |
— | 9 027 | no reader, rc 21 |
pdf |
docling |
— | 9 027 | blocked, rc 1 |
pdf |
pdftotext -layout (reference, not a contender) |
7 600 | 9 027 | 84.2 % |
The hard test passed. The 207 merged ranges in Bilag 7 Prisskjema were
named as where G2 typically breaks. They do not break it: 61/61 for all three.
Merged geometry inflates volume (§ 4) without breaking pairing. This is the
number that says the Excel half can go through Door A.
markitdown destroys row pairing in PDFs, and the mechanism is visible.
It uses pdfminer.six, which emits each cell on its own line — label and value
never share one. This repo measured the same failure independently on 2026-08-21
(pdfminer.six scored 0/4 on the N200 requirement table). Two corpora, same
result.
The pdftotext row is a reference line, not a fourth candidate, and its
84.2 % is an upper bound. It shares an engine with the ground truth, so it is
graded partly on its own coordinates. It is reported because leg 3 needs to
know whether the incumbent is beaten, and the honest answer is: not by anything
measured here.
6. (c) Notes captured
The corpus has no notes. 0 real footnotes, endnotes or comments in the 5
docx (the footnotes.xml/endnotes.xml parts exist but hold only the
separator stubs, ids -1/0), and no comment part at all in the 2 xlsx.
That is a measured 0, not an unmeasured one: the same extractor finds 2/2
injected notes in a docx and 1/1 in an xlsx built from corpus files.
So dimension (c) has no denominator on this corpus. Rather than report a blank, each converter was run on the injected known-positive, which characterises the capability without claiming corpus evidence:
| Note channel | pandoc |
markitdown |
docling |
|---|---|---|---|
docx footnote (1 injected, pandoc-authored file) |
1/1 | 1/1 | 0/1 |
xlsx cell comment (1 injected via openpyxl) |
reader failed, rc 64 | 0/1 | 0/1 |
docling 2.123.1 silently drops the footnote: its output is the body sentence
with no note and no marker. The other two capture it, so the 0 discriminates.
pandoc's xlsx reader failed on the openpyxl-rewritten workbook —
Failed to parse XLSX: Entry not found: xl//xl/worksheets/sheet1.xml. It reads
both real Excel-authored workbooks without complaint. This is a fragility to
non-Excel writers, on a synthetic file, and it belongs in the recommendation as
a caveat rather than as a corpus result.
7. The known cost of running on K2 alone
The operator chose 2026-08-29 to proceed on K2 after K1 Skram proved undownloadable. The cost of that choice, stated as a measured gap:
K1 carried the revision pairs. K2's two stages are byte-identical — 43/43 sha256 equal, re-verified in § 1 — not revisions. Kravspek B1 § 4.7 (a) "utgavepar" is therefore uncovered by this corpus, and "generell" remains the weakest-covered requirement. Nothing in this report bears on how a converter handles a revised edition of a document it has already seen, because this corpus contains no such pair. No substitute corpus was sought.
Two further holes, both denominator-honest: pptx is 0 of 43 files, so
pandoc's and docling's pptx readers are untested here; and docling's
PDF path is 0 of 33 measured for the hardware reason in § 2.
8. Recommendation for leg 3
| Format | Take | The number that decides it |
|---|---|---|
docx |
pandoc -t plain --wrap=none |
193/196 coverage (best), 51/53 paired (tied), 1/1 footnote where docling gives 0/1 |
xlsx |
pandoc |
193/193 coverage and 98/98 paired at 9 979 non-ws chars, vs docling's 58 875 for the same coverage; markitdown corrupts embedded newlines to literal \n |
pdf |
keep pdftotext; none of the three earns the slot |
pandoc cannot read PDF (rc 21, 33/33); markitdown pairs 5.9 %; docling unmeasured |
doc |
docling |
the only converter returning content (2 919 chars, structured); pandoc rc 21, markitdown emits a wrong document class |
smc, zip |
none — and do not trust the exit code | markitdown rc 0 with a stub; docling rc 0 with no file |
pandoc wins on both office formats, and the reason is the same one both
times: it carries the most source content at the least invented volume. It
is also the only one of the three already installed, with no ML stack, no model
download, and no network at run time.
Two caveats that belong to the recommendation, not to a footnote. First,
pandoc's xlsx reader broke on an openpyxl-written workbook (§ 6) — if leg
3 ever converts a workbook this pipeline itself wrote, that path needs its own
test. Second, docling is recommended for exactly one format on the strength
of one file; that is a thin basis, and it is the only basis this corpus offers.
What would change this recommendation: a docling PDF number. It is the one candidate whose PDF path is designed around layout, and it is the one that could not run here. Producing it needs an arm64 or Linux runner, not a new corpus.
9. Verification log
| Claim | Command | Outcome |
|---|---|---|
| Corpus present, unchanged | find, stat -f%z, shasum -a 256 |
43+43 files, 247 107 243 B, 43 SAME 0 DIFF |
| 844 PDF pages (independent re-count) | pdftotext -bbox-layout page count |
844 — matches 08-29 |
pandoc has no pdf reader |
pandoc --list-input-formats; stderr |
xlsx,pptx present; pdf absent; rc 21 ×33 |
pandoc has an xlsx reader |
conversion + coverage check | 193/193 source strings recovered |
docling installed |
uv pip install docling |
exit 0, v2.123.1 |
docling PDF blocked |
docling convert --no-ocr ×3 |
rc 1 ×3, AutoImageProcessor requires PyTorch |
torch>=2.4 unobtainable here |
uv pip install "torch>=2.4" --dry-run |
unsatisfiable; no x86-64 macOS wheel |
| Ground truth rows (ooxml) | word/document.xml, openpyxl |
161 rows; 151 usable |
| Ground truth rows (pdf) | pdftotext -bbox-layout + y-band clustering |
31 095 visual rows → 9 961; 9 027 usable |
| Row query known-positive | same query, 3-page room list | 103 rows from 188 visual rows |
| Note extractor known-positive | injected footnote + comment | docx 2/2, xlsx 1/1; corpus 0 |
| Matcher can see markitdown pdf output | pairing run | 534 pairs > 0 |
| G2 on 207 merged ranges | pairing run, Bilag 7 |
61/61 for all three |
docling replicates merged cells |
grep -o count per converter |
8 vs 1 vs 1, two cells |
markitdown literal \n |
coverage re-run with \n → space |
184/193 → 191/193 |
docling drops footnotes |
3 converters on the injected docx | 1/1, 1/1, 0/1 |
markitdown false success on .doc |
output inspection | rc 0, 28 chars, # Email Message |
docling rc 0 with no file |
output dir listing + log | ERROR logged, exit 0, no file |
| No revision pair | sha256 across stages | 43 identical, 0 differing |
0 pptx in corpus |
extension count | 0 of 43; known-positive xlsx = 2 |
Tool versions
pandoc 3.10.2 · markitdown 0.0.2 · docling 2.123.1 (docling-core
2.92.0, isolated venv, Python 3.12) · poppler 26.08.0 · openpyxl 3.1.5 ·
transformers 5.8.1 · torch 2.2.2 · numpy 1.26.4. Host:
macosx-26.0-x86_64.
Nothing was installed into this repo's environment. docling and its stack
live only in a scratch venv and are not referenced by pyproject.toml.