llm-ingestion-okf/docs/2026-09-10-k3-runde15-bundle-mismatch.md
Kjell Tore Guttormsen 9d1f4b14ed test(fixtures): replace sector-specific example material with generic, fictitious examples — green
Every fixture, test document, tool example and document now uses an invented
kitchen-and-baking handbook series, written in this repository. The package's
behaviour is unchanged; src/ changes are comments and help text only.

- Generated fixtures are regenerated from their generators. Their structural
  counts are identical before and after: elements, images, rows, cells,
  headings, bookmarks and the witness inventory's per-document totals. The
  image-inbox and accounting documents are renamed kapittel-84-*.
- tools/okf_accounting_gate.py: the two options that named one real corpus
  each are replaced by a generic, repeatable --corpus PATH with no default.
  Row 5 compares the PDF pair alone. Gate verdict unchanged: RED rows 2, 3, 6.
- tools/okf_witness.py: the STS JSON reader for one publisher's delivery is
  removed, along with its three twins and five tests. The mutation harness
  loses W09.
- docs/: 13 dated reports that documented runs on a retired reference corpus
  are removed, and 40 are neutralized. Dead links are removed, and no new
  dangling path is introduced.
- The synthetic MCP-gate corpus and the residual probe words are neutral.

Valgt: keep the `okf quality --fasit` bar value (the measured fraction, one corpus) and
rewrite only its provenance, because the verdict stays unchanged and the
number names nothing.

Term check with the local list: 0 of 411 tracked files, 0 file names, 0 of
27 binary fixtures. Suite after git add: 2457 passed, 1 skipped. The base
tree had 2460 passed and 2 skipped; five tests went with the JSON reader and
four were added by the term check. ruff, ruff format and mypy --strict src/
are clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:52:02 +02:00

8.7 KiB

K3 round 15: which bundle was this skill generated for?

Date: 2026-09-10 · HEAD at start: 113b3f8 (v0.8.1 tagged locally on 3daf983) · Red test: b5df335 · No push, no tag, no version bump.

1. What was wrong

okf skill --help had carried this sentence about its own tooling since 2026-09-08:

Measured 2026-09-08: the contract checker passes the UNFILLED template against a real payload (exit 0, 15 rules, 0 findings), and passes a skill built for a different bundle against this one's payload. So the checker cannot tell the two forms apart.

Documented, not closed. okf check had fifteen rules and none of them asked whether the skill and the payload were talking about the same bundle.

2. Reproduced first, on this HEAD

Every figure below was re-measured on 113b3f8 in a frozen export (git archive + uv sync --frozen --extra extract), because the order's figures were taken on 958e9bc. Zero divergence: all five rows reproduced byte for byte, including the exact report lines.

Two real artefacts, both read and never written: a SKILL generated by okf skill from the pinned K2 bundle (k2-trinn1-20260903 at sha256-tree:18ae18ab..., 453 concepts) and a payload from an unrelated corpus (a 446-concept bundle at sha256-tree:da6b8204...).

# pair rc before rc after line after
A skill(k2) x payload(unrelated corpus) 0 1 NOT conformant: 16 rules over 8 excerpts and 438 withheld entries, 1 findings
B unfilled template x payload(unrelated corpus) 0 1 same denominators, 1 finding
C skill(k2) x payload(k2 id, foreign ref) 0 1 same denominators, 1 finding, message names ref only
KP skill(k2) x its own payload 0 0 conformant: 16 rules over 8 excerpts and 445 withheld entries, 0 findings
KN skill(k2) x {} 1 1 NOT conformant: 16 rules over 0 excerpts and 0 withheld entries, 9 findings

Before-rows were measured with the frozen export of 113b3f8; after-rows with the working tree. Exit codes were captured directly (out=$(okf check ...); rc=$?) and never through a pipe, which would report tail's status.

3. Why the rule compares the ref and not only the id

Three distinct builds on this machine carry the same bundle_id k2-trinn1-20260903 at three different refs (f14872a0..., 9a4e5561..., 18ae18ab...). So bundle_id does not identify the bytes, and an id-only rule would pass a stale skill at exit 0 -- the case the generated skill warns about in its own words: "If the bundle moves, the ref moves with it and this file is stale -- regenerate". docs/consumption-contract.md SS 3.3 is the ground: "a version is the producer's assertion; a ref is a fact about bytes". SS 3.1 is the ground for the excerpt-level half: "identity across bundles is the tuple, never concept_id alone".

Row C is the arm that separates a whole rule from half of one. It is the only row where an id comparison and an identity comparison disagree.

4. What was built

One new rule, contract_check.rule_bundle_identity, one new code bundle_mismatch, RULES 15 -> 16. No existing rule weakened, no existing code renamed, no rule short-circuiting another.

No new field was needed, and that was verified rather than assumed. The identity is already in the generated skill's prose. It was factored into one authored copy at skill.py:395 (identity_line), which _rewrite uses at the place it was previously inlined, and contract_check.skill_identity reads back. Generated skill bytes are unchanged: both tracked bundles produce byte-identical skills before and after when rendered on the same interpreter (e67827a6... and 890aeac2... both ways). The earlier apparent 9-byte difference was the bundle path's length in two different export trees, not the code -- measured and discarded.

Three cases, in one rule:

  1. The skill declares no readable identity -> finding. This is what refuses the unfilled template, and it is load-bearing: a silent pass here would let row B through and leave the round undone.
  2. The declared identities disagree in bundle_id, in ref, or in both -> one finding naming which halves disagreed.
  3. An excerpt names a bundle the payload does not -> one finding per excerpt. Same question one level down (SS 3.1).

A payload declaring no identity at all is left to rule_bundle_ref: restating it would report one hole twice. Measured: the known-negative {} is unchanged at 9 findings, and bundle_mismatch is not among them.

5. What the rule found in this repository

skills/okf-consume/SKILL.md -- the first instantiated skill, filled in by hand before okf skill existed -- declares no bundle identity a reader can act on, so okf check now refuses it against its own shipped example payload: NOT conformant: 16 rules over 3 excerpts and 0 withheld entries, 1 findings, rc 1. 1 of 1 shipped hand-made instantiated skill. A second fact sits behind it: that skill's prose states K2 numbers while the payload shipped beside it is regenerated from the tracked golden bundle, because K2 cannot ship. Repairing it is a decision about what that artefact is for, and is left to the operator rather than taken here.

Nine tests asserted the old, false conformance by pairing a skill with a payload from another bundle (or with the template standing in for a skill). They now pair correctly -- the checker's control is an INSTANTIATED skill, since the template can no longer be one.

6. Nothing else moved, measured on the bytes

  • Build output identical. ~/okf-test/dokumenter built with the old source and the new source on the same interpreter: diff -r empty, 52 files, 26 concepts. (Comparing across the two venvs shows one line in log.md -- the converter's own path and Python version -- which is the venv, not this change.)
  • okf project is still byte-equal to okf build on that folder: diff -r empty against <out>/.okf/<id>.
  • K2 pin untouched and green. tests/test_default_bundle_pin.py is unmodified and RAN (7 passed, not skipped): 453 concepts, per-row ranks (1, 1, 1, 1, 1, 5).
  • The repository's own generated pair is conformant at the full 16 rules: conformant: 16 rules over 2 excerpts and 1 withheld entries, 0 findings, and the okf project skill likewise, 16 rules over 8 excerpts and 18 withheld entries, 0 findings.
  • Gates, ruff 0.16.6: ruff check src tests tools clean, ruff format --check . 119 files, mypy src clean, pytest -q 1593 passed, 1 skipped (baseline 1582 + 1 on this HEAD; the frozen export additionally fails test_segmented_golden.py::test_the_four_existing_goldens_are_untouched, which is the export artefact and was red before this round too).

7. Honesty limits

  • The rule compares a DECLARED identity against a DECLARED identity and never opens the bundle. A payload that misreports its own ref passes. Proving a ref against bytes is okf consume --ref's job and needs a bundle path this command deliberately does not take. Nothing here makes a ref true.
  • The denominator for the reproduction is two real SKILL/payload pairs plus arms built from this repository's own two tracked bundles -- not a corpus. Every fixture is generated from bundles already in the tree, so no sentence from any document, no corpus path and no gold question reaches a tracked file.
  • The excerpt-level arm is built, not found. Across the two real payloads reproduced this round, 0 of 16 excerpts carried a bundle_id differing from the payload's own. The arm covers a case that has not been observed in the wild; it is cheap, and its absence has a denominator.
  • Conformance is the floor, never the proof. The checker's own docstring says why: the division of labour (SS 2) and the prohibitions (SS 9) are properties of a RUN -- whether an agent went looking for context the pre-pass withheld -- and no static check reaches them. A green okf check is not a quality gate, and this repository measured that directly on 2026-09-10: three arms over one corpus all returned 0 findings from okf check while their hit@k ranged from 6 of 6 to 0 of 6.
  • "Closed" means the three measured forms now fail, not that no fourth form exists. A skill whose identity sentence has been edited by hand to name a bundle it was not generated from would pass. So would a payload and a skill that agree with each other and disagree with the bytes.
  • The rule count is a contract change downstream. Every consumer quoting "15 rules" is quoting a number that has changed, and one pair portfolio-optimiser already runs goes from rc 0 to rc 1. That is the rule working, not a regression there -- but it is their build's exit code, so it was sent to them rather than left to be discovered.