llm-ingestion-okf/docs/2026-09-08-proveniens-k2.md
Kjell Tore Guttormsen 9d1f4b14ed test(fixtures): replace sector-specific example material with generic, fictitious examples — green
Every fixture, test document, tool example and document now uses an invented
kitchen-and-baking handbook series, written in this repository. The package's
behaviour is unchanged; src/ changes are comments and help text only.

- Generated fixtures are regenerated from their generators. Their structural
  counts are identical before and after: elements, images, rows, cells,
  headings, bookmarks and the witness inventory's per-document totals. The
  image-inbox and accounting documents are renamed kapittel-84-*.
- tools/okf_accounting_gate.py: the two options that named one real corpus
  each are replaced by a generic, repeatable --corpus PATH with no default.
  Row 5 compares the PDF pair alone. Gate verdict unchanged: RED rows 2, 3, 6.
- tools/okf_witness.py: the STS JSON reader for one publisher's delivery is
  removed, along with its three twins and five tests. The mutation harness
  loses W09.
- docs/: 13 dated reports that documented runs on a retired reference corpus
  are removed, and 40 are neutralized. Dead links are removed, and no new
  dangling path is introduced.
- The synthetic MCP-gate corpus and the residual probe words are neutral.

Valgt: keep the `okf quality --fasit` bar value (the measured fraction, one corpus) and
rewrite only its provenance, because the verdict stays unchanged and the
number names nothing.

Term check with the local list: 0 of 411 tracked files, 0 file names, 0 of
27 binary fixtures. Suite after git add: 2457 passed, 1 skipped. The base
tree had 2460 passed and 2 skipped; five tests went with the JSON reader and
four were added by the term check. ruff, ruff format and mypy --strict src/
are clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:52:02 +02:00

386 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Provenance to the original: an address and a locator per format, measured on 629 concepts
**Date:** 2026-09-08 · **Order:** `20260908T104812Z-6143923560-from-.claude` ·
**Instrument:** `src/llm_ingestion_okf/{extract,inbox,profiles}.py` plus the
`SEGMENTED_OKF_V0_2` profile · **Corpus:** the same 43-document public tender
folder every K2 measurement in this repository runs on.
A consuming agent is expected to falsify a hypothesis against the ORIGINAL
document. Before this step it could not: a concept named its source file by
basename, digested it, and — when segmented — carried a `source_offset` that
indexes the text THIS LIBRARY extracted, not the document anybody wrote. The
corpus directory, the extractor and its exact version were all required to
follow that pointer, and none of the three is in the bundle.
The corpus is external and no document body from it appears here. Documents are
named by shape, and everything else is a count.
---
## 0. What IS measured, and what is NOT
**Measured.** The provenance keys present on all 629 K2 concepts before and
after, per source format with a denominator; the three questions the order asks
per format ("can a consumer find the file, verify it, open the right place?");
one hand-walked known-positive from `source_offset` back to a page of the
original, counted in steps; the guard's verdict on seven candidate `sources`
forms, known-positive first; a full K2 rebuild with the file-set,
byte-diff and content identity; and the payload a consumer receives.
**Known-positive, run first.** The rebuild reproduces the bundle this
repository has been measuring on all week before anything else is read from it:
`merged + coded rejections = 39 + 4 = 43 = N`, exit 0; **629 concepts**;
**1 108 files, and the file-set is identical name for name** — 0 concept ids
moved. The 479 files that are not concepts (478 `index.md` plus `log.md`) are
byte-identical to the previous bundle.
**NOT measured.**
- **That a locator is CORRECT for any document not in this corpus.** The page
numbers are the extractor's own page numbers and the sheet rows were checked
against the workbook XML; the line numbers for the converted formats are
lines of OUR extraction and are labelled as such (§ 4).
- **That any model reads the new keys.** Whether an agent handed a payload
actually opens the original is `portfolio-optimiser`'s measurement, not this
one — and it cannot be, because the payload does not carry these keys at all
today (§ 6).
- **A `.doc`, `.zip` or `.smc` row.** Those three types are rejected before any
concept exists, so they have no provenance to carry and no denominator here.
---
## 1. Before: what a concept said, per format, with denominators
Measured on the bundle built at `56c1205` (`/tmp/okf-kravnummer-20260908/k2-new`),
which is what HEAD produced — the following commit `116d3e1` changed the
consumption side only.
```
python3 - <<'PY' # frontmatter key presence over every non-index .md
... # the script is /tmp/o3/measure1.py; it parses the line-oriented block
PY
```
| | concepts | `source_file` | `source_sha256` | `source_offset` | `sources` |
|---|---|---|---|---|---|
| `.pdf` (32 of 33 files) | 586 | 586 | 586 | 576 | **0** |
| `.docx` (5 of 5) | 41 | 41 | 41 | 40 | **0** |
| `.xlsx` (2 of 2) | 2 | 2 | 2 | 2 | **0** |
| `.doc` / `.zip` / `.smc` | 0 | — | — | — | — |
| **total** | **629** | 629 | 629 | 618 | **0** |
The three questions the order asks, per format:
| format | (a) find the original file? | (b) verify it? | (c) open the right PLACE? |
|---|---|---|---|
| `.pdf` | **no** — a bare basename, no directory and no URL | **yes** — `source_sha256` over the dropped bytes | **no** — `source_offset` is a character position in the joined page text; `_extract_pdf` joins pages with a blank line and drops empty ones, so the page number is gone |
| `.docx` | **no** — same | **yes** | **no** — offset only; no paragraph, no heading path |
| `.xlsx` | **no** — same | **yes** | **partly** — the sheet name survives inside the concept TITLE (`Prissammenstilling {#sheet-1}`), as prose; no row, and no key a consumer can read |
Column (b) is a yes with a caveat worth stating: the digest verifies a file you
have already found by other means. It cannot help you find it.
### The known-positive: one hand-walk from an offset back to a page
Taken on a real segmented concept — a fire-safety requirement, `source_offset:
[24966, 25869]`, `source_sha256: 973e9490…` — using a script that shares no
code with the bundle builder:
1. Know the corpus directory. **Not in the bundle.** (`log.md` names it once,
for the run, not per concept — and `log.md` is a run log, not a concept.)
2. Read the file's bytes; confirm `sha256` matches. ✔
3. Know that the extractor is `pdfplumber` behind `[extract]`. **Not in the
bundle.**
4. Know its transitive text pin, `pdfminer.six==20260107`. **Not in the
bundle**, and the extracted text is pinned to it by a frozen literal in
`tests/test_extract.py`.
5. Re-extract, slice `[24966:25869]`. The requirement text comes back. ✔
6. To learn the PAGE, re-open the PDF page by page, rebuild the same join, and
bisect the offset. **Pages 12–13.**
**Six steps, four of which need knowledge the bundle does not carry.** The
document has 20 pages; nothing in the concept says which two.
---
## 2. What the spec has, and what it does not
SPEC v0.2 § 5.1:303-306, verbatim:
> `resource`: REQUIRED within an entry. Names either a concrete artifact a
> consumer can follow (an absolute URL, a bundle-relative path, or a path
> into a `references/` subdirectory, §6) or a population or scope descriptor
> it cannot (for example `all queries in BigQuery project X`).
A dropped file's inbox-relative path is precisely "a bundle-relative path", so
the ADDRESS needs no key of ours. An earlier key inventory had already named
that re-alignment for `source_file` (`sources[].resource` as a bundle-relative
or `references/` path, § 5.1:303-305 + § 6.3); this step executes it.
§ 5.1 lists exactly six entry keys — `resource`, `id`, `title`, `author`,
`usage_count`, `last_modified` — and **none of them is a location within the
resource.** A page, a sheet row and a line have no spec key, at any level.
### Why the locator cannot live inside a `sources` entry
Not a preference — a measured gate. Against the pinned guard (1.3.0), with a
known-positive first:
```
.venv/bin/python -c 'from llm_ingestion_guard import okf; okf.parse_frontmatter(DOC)'
```
| candidate | verdict |
|---|---|
| `sources: [{ id: a, resource: fixture }]` (our Door A golden) | **PASS** — known-positive |
| `sources: [{ resource: a.pdf, title: A }]` | **PASS** |
| `sources: [{ resource: a.pdf, sha256: abc }]` | REJECT — `mapping key 'sha256' is not on the OKF mapping allowlist under 'sources'` |
| `sources: [{ resource: a.pdf, pages: [1, 2] }]` | REJECT — `a flow mapping admits scalar leaves only, not '['` |
| `sources: [{ resource: a.xlsx, sheet: Ark1, rows: [1, 9] }]` | REJECT — same |
| `sources: [{ resource: "Del II, Bilag.pdf" }]` | REJECT — `a quoted scalar inside a flow mapping is not a supported form` |
| `sources: [{ resource: Del II Bilag 3.3.1 - Brannkonsept.pdf }]` | **PASS** — spaces are fine unquoted |
Two of those rows decide the design. A locator inside the entry has **no
expressible form** that survives the gate, so it would produce bundles this
project emits and could never read back through Door C. And because quoting is
also refused, a path containing a flow terminator cannot be escaped — it is
refused fail-fast instead (`inbox_source_file_unaddressable`), which is this
library's standing rule of validation over repair.
**The spec deviation, in four parts.**
| | |
|---|---|
| **spec says** | `sources[].resource` is the address (§ 5.1:303-306); there is no field for a place within it |
| **we do** | write the address in spec form, and the locator as top-level keys of ours — `source_pages`, `source_sheet`, `source_rows`, `source_lines` — in the shape `source_offset` already had |
| **measurement** | the guard rejects every route to putting a locator inside the entry (table above), and § 5.1's six keys contain no locator to reuse |
| **re-align** | if upstream ever gives `sources` a location field, these four keys move into the entry and become a rename. Until then they are ours, named as ours, and they say what they index |
---
## 3. The rule
**A unit table is built AT EXTRACTION**, where the extracted text and the
original's structure are the only moment they are known to agree, and it is
stored as `(start offset, original unit number, scope)` triples. A locator is
then a lookup, never a guess about text whose page breaks are already gone.
- `.pdf` → **`source_pages: [from, to]`**, from `pdfplumber`'s own page numbers.
A page that yielded no text leaves the table without renumbering the pages
after it, so the number is the document's, not "the n-th page that produced
text".
- `.xlsx` → **`source_sheet: <name>` + `source_rows: [from, to]`**. The sheet
comes from the heading the converter writes per sheet; the row restarts at
every sheet and skips the separator line **by position, never by content**
(§ 4).
- `.docx` and every text type → **`source_lines: [from, to]`**, lines of the
extracted text (§ 4).
- `source_offset` **stays**, unchanged. It is what an existing consumer joins
on, and a locator that replaced it would break them in order to help them.
`sources` is written whether or not a locator can be computed: the address
answers "which document", the locator answers "where in it", and a consumer is
owed the first even when the second is unavailable.
**One profile moves.** `provenance` is a policy object on `BundleProfile`,
defaulted to `None`, and set on `SEGMENTED_OKF_V0_2` alone. `sources` is a v0.2
key so a v0.1 profile must not name it; `DEFAULT` and `STRICT_V1` state
contracts owned in other repositories; `OKF_V0_2` is Door A's, where `sources`
already comes from the manifest. Five shipped profiles are byte-identical, and
a test asserts the field set so a sixth cannot arrive unnoticed.
---
## 4. Two things the measurement CHANGED before they shipped
Both were found by running the rule against the real corpus, and both would
have shipped a confident wrong number.
**A `paragraphs` key for `.docx` would have named a number the document does
not have.** The order asks for a paragraph or heading index. Measured on the
five K2 `.docx` files — `<w:p>` elements against converted-markdown lines:
| document | `<w:p>` | md lines | headings |
|---|---|---|---|
| 1 | 108 | 75 | 2 |
| 2 | 27 | 33 | 0 |
| 3 | 65 | 67 | 4 |
| 4 | 176 | 144 | 2 |
| 5 | 57 | 63 | 0 |
Not one pair agrees, and the heading-path alternative is empty for 2 of the 5.
So the key is `source_lines` and it says what it indexes: lines of the text
this library extracted. For `md`/`txt` that text IS the dropped file, and the
number is the original's own line.
**An empty spreadsheet row renders exactly like a table separator.** The first
implementation skipped the separator by matching the LINE (`^\|[-|: ]+\|$`).
On the K2 price sheet that rule also ate **8 empty rows**, and reported the
sheet's last row as **92** against a workbook that says **100** — every row
after the first blank one shifted, silently, for the whole sheet. The rule is
now positional: the second table line of a sheet is the separator, and nothing
else is. A new hand-laid fixture (`tomrad.xlsx`, four rows with the third
empty) is what keeps that red.
After the fix, against the workbook XML:
| spreadsheet | locator rows | `<row>` elements | contiguous from 1 |
|---|---|---|---|
| K2 document list | 1–39 | 39 | yes |
| K2 price sheet | 1–100 | 100 | yes |
Exact, both. The honesty limit: this holds as far as one converted line per
`<row>` holds. A sheet whose XML omits a row entirely would be numbered from
the converted table, and nothing here can see that.
---
## 5. After: the rebuild
`okf build` over the same 43-document folder, same arguments
(`--ingested-at 2026-09-03T00:00:00Z --bundle-id k2-trinn1-20260903
--okf-version 0.2`), run from a frozen copy of `src/` on `PYTHONPATH` so no
later edit could reach a measurement in flight — the log's first line names
that path.
| | before | after |
|---|---|---|
| `merged + coded rejections == N` | `39 + 4 = 43` | `39 + 4 = 43`, exit 0 |
| concepts | 629 | **629** |
| files in the bundle | 1 108 | **1 108**, identical name set |
| concept ids that moved | — | **0** |
| files byte-identical | — | **479 of 1 108** (every `index.md`, plus `log.md`) |
| files changed | — | **629 — every concept, and nothing else** |
| lines REMOVED across all 629 | — | **0** |
The diff is additive by measurement, not by intention: every changed file gained
`sources` plus exactly one locator shape, and no file lost a line or changed a
body.
| added keys | concepts | matches |
|---|---|---|
| `sources`, `source_pages` | 586 | every `.pdf` concept |
| `sources`, `source_lines` | 41 | every `.docx` concept |
| `sources`, `source_sheet`, `source_rows` | 2 | every `.xlsx` concept |
**The three questions, after:**
| format | (a) find | (b) verify | (c) open the right place |
|---|---|---|---|
| `.pdf` | **yes** | **yes** | **yes** — `source_pages` |
| `.docx` | **yes** | **yes** | **partly** — `source_lines` locates the extraction, not the document |
| `.xlsx` | **yes** | **yes** | **yes** — `source_sheet` + `source_rows` |
**629 of 629 carry an address and a locator.** For `.pdf` and `.xlsx`, all
three columns are yes; for `.docx` the third is honest rather than yes.
**The known-positive, walked again.** The same fire-safety concept now reads
`source_pages: [12, 13]` — the two pages the six-step hand-walk in § 1 arrived
at, computed by a script that shares no code with the builder. The walk is now
zero steps and needs nothing the bundle does not carry.
### The new content identity
**`sha256-tree:665563a2f74423fcbcc8e4f0b0954ee73b73985ac0418de4f6987bd162a1f7c8`.**
The previous ref, `sha256-tree:2f82fcfe…`, is superseded and stale. It was
recomputed from the previous bundle in the same run, as the known-positive, and
it matched — an instrument that has not reproduced a known figure has not been
shown to count.
The bundle on disk grows **97 011 B, 1 943 406 → 2 040 417 (+4.99 %)**, which is
**154 B per concept** — one address line plus one locator line.
### The payload regression
The order asks whether the pre-pass payload's token count moves now that the
frontmatter is bigger. Measured, same question, same `k`, both bundles through
the same tool:
| | before | after |
|---|---|---|
| payload bytes | 209 092 | **209 092** |
| delivered excerpts | 8 | 8 |
| budget spent (`BUDGET_UNIT`) | 118 409 | 118 409 |
| concept ids delivered | — | identical |
| excerpt texts, `text_sha256`, `withheld`, `denominators` | — | identical |
| changed lines in the whole payload | — | **18** |
The 18 are the bundle `ref` and the eight per-concept `sha256` values. **The
payload does not grow by one byte**, and the reason is § 6's finding stated from
the other side: an excerpt carries the concept's BODY, so frontmatter growth is
invisible to it. The frontmatter got bigger and the thing a consumer is billed
for did not move.
There is no o200k figure here. `tiktoken` is not installed on this machine, so
the byte count and the payload digest are the measurement; repeating a published
token number would be quoting, not measuring.
---
## 6. The consumer side: measured, not fixed
`tools/okf_consume.py` builds one excerpt per delivered concept, and
`excerpt_for` names its members explicitly. Measured on a real payload:
```
python3 -c "import json; d=json.load(open('…/k2-new-default.json')); \
print(sorted({k for x in d['excerpts'] for k in x}))"
# -> ['adjudication', 'bundle_id', 'bundle_id_inherited', 'concept_id',
# 'rank', 'sha256', 'text', 'text_sha256', 'trust_tier']
```
**No. Nine members, and not one of them is a provenance field** — not
`source_file`, not `sources`, not a locator, not even the concept's `title`.
An agent holding this payload cannot say "file X, page 12". It can say
`concept_id`, which is a path INTO THE BUNDLE, so the provenance is one file
read away — but only for a consumer that also has the bundle.
What it would take: `excerpt_for` carries `sources` and the locator keys
through from the concept's frontmatter. That is a payload-shape change, the
consumption contract § 8 enumerates an excerpt's members, and the shape is
`portfolio-optimiser`'s to decide. **Not done here**, deliberately, and this
paragraph is the whole of the fix's cost.
---
## 8. Honesty limits
1. **`source_lines` is not a paragraph number**, and for `.docx`, `.pptx`,
`.odt`, `.rtf`, `.csv`, `.json` and `.html` it indexes our extraction rather
than the document. § 4 has the numbers that make that the honest answer
rather than a shortcut.
2. **A row number holds only while one converted line means one `<row>`.**
Verified exact on both K2 spreadsheets and all three fixtures, all
contiguous from row 1. A sheet with a genuinely missing row would number
from the converted table.
3. **The address is inbox-relative, not absolute.** It resolves against the
folder that was dropped, and the bundle still does not say where that folder
is. Naming an absolute path would leak an operator's filesystem into a
published bundle; naming a URL is available whenever the source HAS one, and
for a dropped file it does not.
4. **The original is still not in the bundle.** § 6.3's `references/`
convention would put it there; copying binaries into a bundle is a decision
about size and licensing that nobody has taken, and it is not this order's.
5. **The rebuild is one corpus.** Every number above is 43 documents of one
public tender.
6. **`sources` carries no `id`.** § 5.1 says it SHOULD be present when the body
cites the source, and no body here cites with a `[^id]` footnote — measured
0 across the corpus by an earlier key inventory. When one does, the key
is available and unwritten.
---
## 9. Verification log
| claim | command |
|---|---|
| 629 concepts, per-format key counts, before and after | `python3 /tmp/o3/measure1.py <bundle>` |
| the file set and the byte diff | a `rglob` set comparison plus `read_bytes()` equality over 1 108 files |
| every diff is additive | `difflib.unified_diff` per concept; 0 removed lines over 629 files |
| the guard's verdict on seven `sources` forms | `okf.parse_frontmatter` on each, known-positive first |
| docx paragraphs ≠ markdown lines | `<w:p>` count from `word/document.xml` against `extract_text(...).split("\n")` |
| spreadsheet rows against the workbook | `<row>` elements from `xl/worksheets/sheet1.xml` against `source_units(...).covering(0, len(text))` |
| the known-positive page walk | an independent script re-extracting and bisecting page offsets: pages 12–13, matching the emitted `source_pages` |
| the payload members | `json.load` over a shipped K2 payload |
| suite, lint, types | `pytest -q`, `ruff check src tests tools`, `ruff format --check`, `mypy --strict src` |