docs(measure): count the public project corpus - K2 measured, K1 blocked

Step 1 of the approved sequence (scenarioanalyse SS 6): produce numbers for a
heterogeneous project corpus after conversion. Measurement only - no code, no
parser, no bundle, no dependency; src/ untouched and the corpus lives outside
the repo.

Headline: 1 595 054 characters after conversion across 43 unique files,
844 PDF pages, 0 conversion failures of 40 attempted.

Two order premises moved under measurement:

- SS 9 marks K1 Skram "open, tested". It is not. K1 serves 79 filenames as
  plain text with no link and no file id for an anonymous visitor, on all
  three URL variants (known-positive: the same parser extracts 43/43 links
  from K2). The 142.8 KB PDF that "proved the mechanism" on 28.08 is a K2
  file - 146 242 bytes, Del I Vedlegg 5. The tested corpus was K2 all along.
  Per the order, K1 is reported blocked rather than substituted.
- SS 9 calls K2's two stages a near-duplicate. All 43 files are byte-identical
  by sha256, 0 differing. The corpus therefore contains no revision pair.

SS 9's file counts were exact for both corpora (79 and 43); the access and
duplication claims were not.

Also measured, closing an explicit "not verified" in SS 8: openpyxl
data_only=True returned a cached value for 52 of 52 formula cells, 0 None.
Bounded to the one workbook that has formulas.

Absences carry denominator, exit status and a known-positive throughout:
0 scanned PDFs (0 of 33 zero-font), 0 pptx (0 of 86, exit 0, xlsx control = 4),
0 login walls (0 of 86, control = 1).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V3Ghu6sgMSsycFDrZDGzd6
This commit is contained in:
Kjell Tore Guttormsen 2026-08-29 07:03:50 +02:00
commit 6572e9ec4d

View file

@ -0,0 +1,265 @@
# Count: public project corpus (K1 Skram + K2 Stange), 2026-08-29
Measurement only. No code was written, no parser built, no bundle produced,
no dependency added. `src/` is untouched. The corpus lives outside the repo
in `~/corpora/okf-telling-20260829/` and is not tracked by git.
This is step 1 of the approved sequence in
`.claude/docs/2026-08-28-scenarioanalyse-portfolio-optimiser.md` § 6. It
produces numbers, not machinery.
## 0. Headline
| Question | Answer |
|---|---|
| Characters after conversion, whole downloadable corpus | **1 595 054** (unique set) |
| Same number counting both published stages as delivered | 3 190 108 |
| Files measured | 43 unique (86 downloaded) |
| PDF pages | 844 |
| Bytes on disk | 247 107 243 unique (494 214 486 downloaded) |
| Conversion failures | **0 of 40 attempted** |
| K1 Skram | **NOT DOWNLOADABLE — order premise (a) does not hold** |
Two premises the order told me to verify rather than assume both moved:
- **(a) "open without login" was tested against K1.** It was not. The single
142.8 KB PDF from 28.08 is `Del I Vedlegg 5 Begrunnelse for å bruke krav i
stedet for tildelingskriterier.pdf`, **146 242 bytes, and it belongs to K2**,
not K1. § 9 marks K1 "åpen, **testet**" and K2 "samme mekanisme". That is
reversed: K2 is the tested one, and K1 is the one that fails.
- **(b) K1 is blocked**, so per the order's instruction I did not substitute a
corpus. I measured K2 in full and report K1 as blocked. See § 5 for the
decision this leaves to the operator.
## 1. Files per format — K2, unique set (43 files)
Denominator is three numbers per format: found / attempted / failed.
| Format | Found | Attempted | Failed | Pages | Chars out | Bytes | Tool |
|---|---:|---:|---:|---:|---:|---:|---|
| `doc` | 1 | 0 | 0 | — | *not converted* | 63 488 | — |
| `docx` | 5 | 5 | 0 | — | 31 710 | 138 462 | `pandoc -t plain` |
| `pdf` | 33 | 33 | 0 | 844 | 1 548 766 | 155 747 491 | `pdftotext` |
| `smc` | 1 | 0 | 0 | — | *not converted* | 86 779 456 | — |
| `xlsx` | 2 | 2 | 0 | — | 14 578 | 31 382 | `markitdown` |
| `zip` | 1 | 0 | 0 | — | *not converted* | 4 346 964 | — |
| **SUM** | **43** | **40** | **0** | **844** | **1 595 054** | **247 107 243** | |
Three files (`doc`, `smc`, `zip`) are reported as **not converted**, not as 0
characters: no local converter exists for binary Word 97, for a Solibri model,
or for a zip of dwg drawings. Nothing was installed to change that.
`.doc` is the analogue of the order's `.dwg`/`.ifc` note: counted as file and
bytes, conversion absent rather than empty.
### Access test per file class (order premise (a))
Every class was fetched separately. Type confirmed with `file`, not with the
served `Content-Type`.
| Class | HTTP | curl exit | Bytes | `file` says |
|---|---:|---:|---:|---|
| `docx` | 200 | 0 | 30 410 | Microsoft Word 2007+ |
| `pdf` | 200 | 0 | 146 242 | PDF document, version 1.7, 7 pages |
| `xlsx` | 200 | 0 | 10 169 | Microsoft Excel 2007+ |
| `smc` | 200 | 0 | 86 779 456 | Zip archive data |
| `zip` | 200 | 0 | 4 346 964 | Zip archive data |
| `doc` | 200 | 0 | 63 488 | Composite Document File V2 |
All six classes download anonymously. The bulk run then fetched 86/86 files:
0 missing, 0 zero-byte, **0 HTML login walls**. The login-wall query is proven
able to find: run against a saved tender page it returns 1 hit.
## 2. PDF: scanned vs digital — both signals
Per the order, two signals per file: `pdftotext` character count AND
`pdffonts` embedded-font count. Neither alone decides.
**Result: 0 raster/scanned PDFs.** 33 of 33 carry embedded fonts (zero-font
count = 0), so no OCR question arises and none was attempted.
One file is text-poor despite having a font — a vector drawing, which is the
`figures do not survive extraction` case this library already documents:
| PDF | Pages | `pdftotext` chars | `pdffonts` fonts | Chars/page | Reading |
|---|---:|---:|---:|---:|---|
| Bilag 2.4 - VA-tegninger | 2 | **2** | 1 | 1 | **vector drawing** |
| Bilag 2.7 - Eksisterende EL-infrastruktur | 7 | 4 155 | 14 | 594 | digital, drawing-heavy |
| Bilag 3.2.1 - RIG-R01 Datarapport | 93 | 72 914 | 50 | 784 | digital |
| Bilag 2.1 - Arkitekttegninger | 62 | 161 856 | 147 | 2 611 | digital |
| Bilag 3.1 - Miljøteknisk rapport | 150 | 284 854 | 24 | 1 899 | digital |
| Bilag 1 - Kravspesifikasjon | 105 | 251 914 | 10 | 2 399 | digital |
Chars-per-page across the 33 PDFs: min 1, median 1 899, max 5 731.
The `1 chars/page` case is why one signal is not enough: on font count alone
that file looks digital, and on text alone it looks scanned. It is neither —
it is a drawing.
## 3. Excel structure, and the `data_only` test § 8 left unverified
Denominator: 4 xlsx files found, 4 opened, 0 raised.
| Workbook | Sheets | Non-empty cells | Merged ranges | Formula cells |
|---|---:|---:|---:|---:|
| Bilag 0 Dokumentliste del II | 1 | 76 | 0 | 0 |
| Bilag 7 Prisskjema | 1 | 206 | **207** | **52** |
**`data_only=True`, measured on real files rather than assumed:** for all
**52 of 52** formula cells in `Bilag 7 Prisskjema`, `data_only=True` returned
a cached value. **0 cells came back `None`.**
```
Prissammenstilling!G16 formula '=F92' data_only 5647500
Prissammenstilling!G18 formula '=SUM(G7:H17)' data_only 5647500
Prissammenstilling!G19 formula '=G18*0.25' data_only 1411875
```
So on this corpus the feared failure mode — formulas with no cached value —
**did not occur**. § 8 can be updated from "known behaviour, test on a real
file" to "tested, 52/52 cached, on one workbook". The claim is bounded to one
workbook with formulas, because that is how many the corpus contains.
**Known-positive for the merged-cell absence:** claiming "0 merged ranges in
Bilag 0" needs proof the query can find merged ranges. The same query against
`Bilag 7` in the **same directory and same file class** returns **207**. So
the 0 is measured, not a broken query.
More cells are merged (207) than are non-empty (206). Merged geometry is the
dominant structure in the price form — which is the shape this library already
declines to recover, and this corpus does not change that.
## 4. `.pptx`: absence with denominator and exit status
```
find ~/corpora/okf-telling-20260829/K2 -type f -iname '*.pptx'
-> no output, exit 0
```
**0 of 86 files searched.** Known-positive, same command shape and directory:
`-iname '*.xlsx'` returns 4 hits, exit 0 — the query can find.
No pptx was constructed to fill the hole. The hole is stated: this corpus
cannot exercise a pptx path.
## 5. K1 Skram: blocked, with the measurement
The order says to return rather than substitute if download is blocked. K1 is
blocked. K2 is not, so this report delivers K2 in full and leaves K1 open.
**What the page gives an anonymous visitor:** 79 file rows, each a filename as
**plain text**. No `<a>`, no `onclick`, no `data-` attribute, no `<input>`, and
**no file id of any kind** — so no download URL can be constructed without
guessing, and guessing ids was out of scope.
| Page | Rows | `<a href>` in file list | 6+ digit ids |
|---|---:|---:|---:|
| K1 Skram (3 URL variants, all HTTP 200) | 79 | **0** | **0** |
| K2 trinn 1 | 43 | 43 | 43 |
| K2 trinn 2 | 43 | 43 | 43 |
The known-positive is the same parser on the same run: it extracts 43/43 links
from K2. The 0 on K1 is a measured absence, not a parser that failed.
Three URL forms were tried (`/nb-no/anbud/246749360/…`,
`/permalink/246749360.aspx`, `/en/tender/246749360/…`); all resolve 200 and all
show 79 rows with 0 links.
**Correlation, not established cause:** K1 carries `Dato for innlevering er
passert` with a deadline of 20.01.2025; neither K2 page carries that flag.
Mercell plausibly withdraws download links after the deadline. I did not verify
that rule, and do not assert it.
## 6. Deviation against § 9
| § 9 claim | Measured | Verdict |
|---|---|---|
| K1: pdf 57 · docx 8 · xlsx 3 · dwg 4 · ifc 3 · annet 4 = 79 | pdf 57, docx 8, dwg 4, xlsx 3, ifc 3, smi/smc/pln/jpg 4 = 79 | **exact match** |
| K2: pdf 33 · docx 5 · xlsx 2 · doc 1 · annet 2 = 43 | pdf 33, docx 5, xlsx 2, doc 1, zip+smc 2 = 43 | **exact match** |
| K1 "åpen, **testet**" | 0 download links; the 28.08 test file is a K2 file | **false** |
| K2 "åpen, samme mekanisme" | true — and K2 is the one actually tested | **true, mislabelled** |
| K2 "delvis (to trinn, **nær-duplikat**)" | **all 43 files byte-identical (sha256)** across both stages | **false — identical, not near** |
| "≈ 120 filer" for K1+K2 | 43 unique downloadable; 79 listed but unreachable | **not reached** |
The file counts read off the page were right. The access claim and the
duplication claim were not.
**On the revision pair (D5):** § 9 justified K1 partly by "ekte revisjonspar"
and K2 by "to trinn, nær-duplikat". K1 is unreachable, and K2's two stages are
byte-identical — sha256 equal for 43/43 files, 0 files differing, 0 files
present in one stage only. **This corpus contains no revision pair.** Whatever
step 1 was meant to feed about revisions, it does not have the data.
## 7. What this number means for the union arm
The order asks for the character total because it decides whether the union
arm is a context-threshold question at all.
**1 595 054 characters** is the whole downloadable corpus after conversion —
roughly 0.40.5 M tokens at 34 chars/token. That is one number from one
corpus, and it is a *lower* bound on a real project: three of 43 files could
not be converted locally, K1's 79 files are absent entirely, and the 86 MB
Solibri model and the dwg archive carry content no text pipeline reaches.
No decision follows from this report. It supplies the number § 6 point 1 asked
for; the union-arm call is the operator's, and no union-arm work was started.
## 8. Verification log
Every number above, with the command that produced it.
| Claim | Command | Outcome |
|---|---|---|
| K1 resolves, 79 rows | `curl -sSL …/anbud/246749360/…` | 200, 79 `tdDownload` |
| K1 has no download links | python parse of `id="fileslist"` block | 0 `<a`, 0 ids, 3 URL variants |
| K1 parser known-positive | same parser on K2 pages | 43/43 links found |
| K2 trinn 1 = 43 files | `curl …/permalink/217958336.aspx` | 200, 43 rows, 43 links |
| K2 trinn 2 = 43 files | `curl …/permalink/224154780.aspx` | 200, 43 rows, 43 links |
| All 6 classes downloadable | `curl` per class + `file -b` | 6/6 HTTP 200, exit 0, real types |
| 86/86 downloaded | manifest vs disk comparison | 0 missing, 0 zero-byte |
| No login walls | `file`-type scan for HTML/ASCII | 0 of 86; known-positive = 1 |
| 28.08 test file is a K2 file | `pdfinfo` on the fetched PDF + `grep` in K1 | title matches K2 `Vedlegg 5`; `grep 'Begrunnelse' k1.html` exit 1, known-positive `Forpliktelseserkl` = 1 hit exit 0 |
| Stages byte-identical | sha256 of all 86 files | 43 identical, 0 differing |
| PDF pages = 844 | `pdfinfo` per file | 33/33 read |
| PDF chars = 1 548 766 | `pdftotext … -` piped to char count | 33/33 exit 0 |
| 0 scanned PDFs | `pdffonts` per file | zero-font count = 0 of 33 |
| docx chars = 31 710 | `pandoc -t plain --wrap=none` | 5/5 exit 0 |
| xlsx chars = 14 578 | `markitdown` | 2/2 exit 0 |
| Excel structure | `openpyxl` 3.1.5 | 4/4 opened, 0 raised |
| `data_only=True` cached | `openpyxl` both modes, same cells | 52/52 cached, 0 None |
| Merged-cell known-positive | same query on Bilag 7 | 207 ranges found |
| 0 pptx | `find … -iname '*.pptx'` | no output, exit 0, n=86; known-positive xlsx = 4 |
### Tool versions (nothing installed)
`pdfinfo`/`pdftotext`/`pdffonts` poppler 26.08.0 · `pandoc` 3.10.2 ·
`markitdown` 0.0.2 · `openpyxl` 3.1.5 · `pypdf` 6.14.2 (unused, fallback).
Confirmed absent and left absent: `tesseract`, `pdfplumber`, `pymupdf`/`fitz`,
`docling`.
### Two traps this run hit, recorded so they are not repeated
- **`curl` inside a `while read` loop consumes stdin.** The first bulk run
stopped at 12 of 43 files and the wrapper still exited 0. A clean exit code
described the wrapper, not the work.
- **A backgrounded `nohup … &` is reported "completed" when the wrapper
returns, not when the download ends.** Three overlapping `download.sh`
processes wrote to one log, producing 65 rows for a 43-row manifest. The
fix was to verify files against the manifest on disk rather than to trust
the log — disk is ground truth, the log is a claim.
### Cross-check: `markitdown` loses text `pandoc` keeps (docx)
Both tools were run on all 5 docx files, as the order asked.
| File | `pandoc` chars | `markitdown` chars | Ratio |
|---|---:|---:|---:|
| Vedlegg 1 Søknadsbrev | 15 165 | 2 700 | 0.18 |
| Vedlegg 2 Forpliktelseserklæring | 1 800 | 1 211 | 0.67 |
| Vedlegg 3 Erfaringsoversikt | 5 135 | 825 | 0.16 |
| Vedlegg 4 Tilbudsbrev | 7 855 | 5 389 | 0.69 |
| Dokument for avtaleinngåelse | 1 755 | 1 473 | 0.84 |
`markitdown` returns as little as 16 % of `pandoc`'s text on form-shaped
documents. The `docx` total in § 1 uses `pandoc`. This is a measurement of two
tools on this corpus, not a recommendation — no tool choice was made here.