test(fixtures): replace sector-specific example material with generic, fictitious examples — green

Every fixture, test document, tool example and document now uses an invented
kitchen-and-baking handbook series, written in this repository. The package's
behaviour is unchanged; src/ changes are comments and help text only.

- Generated fixtures are regenerated from their generators. Their structural
  counts are identical before and after: elements, images, rows, cells,
  headings, bookmarks and the witness inventory's per-document totals. The
  image-inbox and accounting documents are renamed kapittel-84-*.
- tools/okf_accounting_gate.py: the two options that named one real corpus
  each are replaced by a generic, repeatable --corpus PATH with no default.
  Row 5 compares the PDF pair alone. Gate verdict unchanged: RED rows 2, 3, 6.
- tools/okf_witness.py: the STS JSON reader for one publisher's delivery is
  removed, along with its three twins and five tests. The mutation harness
  loses W09.
- docs/: 13 dated reports that documented runs on a retired reference corpus
  are removed, and 40 are neutralized. Dead links are removed, and no new
  dangling path is introduced.
- The synthetic MCP-gate corpus and the residual probe words are neutral.

Valgt: keep the `okf quality --fasit` bar value (the measured fraction, one corpus) and
rewrite only its provenance, because the verdict stays unchanged and the
number names nothing.

Term check with the local list: 0 of 411 tracked files, 0 file names, 0 of
27 binary fixtures. Suite after git add: 2457 passed, 1 skipped. The base
tree had 2460 passed and 2 skipped; five tests went with the JSON reader and
four were added by the term check. ruff, ruff format and mypy --strict src/
are clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-23 13:54:57 +02:00
commit 9d1f4b14ed
174 changed files with 1889 additions and 6512 deletions

View file

@ -28,9 +28,9 @@ import pytest
from llm_ingestion_okf import propose as okf_propose_segments
from llm_ingestion_okf.segmentation import parse_segmentation_plan
DOCUMENT = """# N500 Vegbygging
DOCUMENT = """# Q500 Kakebaking
Innledende tekst om vegbygging og dens omfang.
Innledende tekst om kakebaking og dens omfang.
## 3.1 Brannkonsept
@ -84,7 +84,7 @@ To uavhengige roemningsveier fra hver branncelle.
GOLDEN_PROPOSED_AT = "2026-09-03T00:00:00Z"
def write(tmp_path: Path, text: str = DOCUMENT, name: str = "n500.md") -> Path:
def write(tmp_path: Path, text: str = DOCUMENT, name: str = "q500.md") -> Path:
path = tmp_path / name
path.write_text(text, encoding="utf-8", newline="")
return path
@ -420,7 +420,7 @@ LONG_PARAGRAPH = ("Krav til seksjonering av bygget over flere etasjer. " * 40).s
UNSTRUCTURED = "\n\n".join(f"{LONG_PARAGRAPH} Avsnitt {i}." for i in range(12)) + "\n"
STRUCTURED_WITH_A_LONG_TAIL = (
"# N500 Vegbygging\n\nInnledende tekst om vegbygging.\n\n"
"# Q500 Kakebaking\n\nInnledende tekst om kakebaking.\n\n"
"## 3.1 Brannkonsept\n\nKort avsnitt om seksjonering.\n\n"
"## 3.2 Roemning\n\n" + UNSTRUCTURED
)