Step 1 of the approved sequence (scenarioanalyse SS 6): produce numbers for a
heterogeneous project corpus after conversion. Measurement only - no code, no
parser, no bundle, no dependency; src/ untouched and the corpus lives outside
the repo.
Headline: 1 595 054 characters after conversion across 43 unique files,
844 PDF pages, 0 conversion failures of 40 attempted.
Two order premises moved under measurement:
- SS 9 marks K1 Skram "open, tested". It is not. K1 serves 79 filenames as
plain text with no link and no file id for an anonymous visitor, on all
three URL variants (known-positive: the same parser extracts 43/43 links
from K2). The 142.8 KB PDF that "proved the mechanism" on 28.08 is a K2
file - 146 242 bytes, Del I Vedlegg 5. The tested corpus was K2 all along.
Per the order, K1 is reported blocked rather than substituted.
- SS 9 calls K2's two stages a near-duplicate. All 43 files are byte-identical
by sha256, 0 differing. The corpus therefore contains no revision pair.
SS 9's file counts were exact for both corpora (79 and 43); the access and
duplication claims were not.
Also measured, closing an explicit "not verified" in SS 8: openpyxl
data_only=True returned a cached value for 52 of 52 formula cells, 0 None.
Bounded to the one workbook that has formulas.
Absences carry denominator, exit status and a known-positive throughout:
0 scanned PDFs (0 of 33 zero-font), 0 pptx (0 of 86, exit 0, xlsx control = 4),
0 login walls (0 of 86, control = 1).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V3Ghu6sgMSsycFDrZDGzd6