A concept named its source file by basename and, when segmented, carried a
`source_offset` into the text THIS LIBRARY extracted. Following that pointer
needed the corpus directory, the extractor and its exact transitive version --
none of which the bundle carries. Hand-walked on a real K2 concept: six steps,
four of them requiring knowledge from outside the bundle, to learn that a
requirement sits on pages 12-13 of a 20-page document.
The address is spec's: `sources: [{ resource, title }]`, where `resource` is
the dropped file's inbox-relative path (SPEC v0.2 5.1:303-306 -- "an absolute
URL, a bundle-relative path, or a path into a `references/` subdirectory").
The locator is ours, and it has to be: 5.1 has no field for a place within a
resource, and the pinned guard (1.3.0) rejects every route to putting one
inside a `sources` entry -- a non-allowlisted key by name, a nested flow list
as "scalar leaves only", and quoting as an unsupported form. So the locator is
top-level keys shaped like `source_offset`, and a path carrying a flow
terminator is refused fail-fast rather than mangled.
The unit table is built AT EXTRACTION, where the extracted text and the
original's structure are known to agree: pdf -> `source_pages` from
pdfplumber's own page numbers (a page that yielded no text does not renumber
the ones after it), xlsx -> `source_sheet` + `source_rows`, everything else ->
`source_lines`. `source_offset` stays.
Two measurements changed the design before it shipped. A `paragraphs` key for
docx would name a number the document does not have: `<w:p>` counts of
108/27/65/176/57 against converted-markdown lines of 75/33/67/144/63, not one
pair agreeing -- so the key is `source_lines` and says what it indexes. And an
empty spreadsheet row renders exactly like a table separator: the content-based
rule ate 8 empty rows on the K2 price sheet and reported its last row as 92
against a workbook that says 100. The separator is now found by position, and
`tomrad.xlsx` keeps that red.
One profile moves. `provenance` is a policy object, `None` everywhere but
`SEGMENTED_OKF_V0_2`; the other five shipped profiles are byte-identical.
K2 rebuilt from a frozen src copy: 629 concepts, 1108 files, name set identical,
0 ids moved, 479 files byte-identical, 629 changed and 0 lines removed anywhere.
629/629 now carry an address and a locator. New ref
`sha256-tree:665563a2f74423fcbcc8e4f0b0954ee73b73985ac0418de4f6987bd162a1f7c8`;
`2f82fcfe...` is stale. The pre-pass payload does not grow by one byte
(209 092 B before and after, 18 changed lines: the ref and eight per-concept
digests) -- because an excerpt carries the body, not the frontmatter, which is
also why the consumer still cannot cite "file X page 12" from a payload alone.
Report: docs/2026-09-08-proveniens-k2.md. 1339 tests, ruff and mypy clean.
Co-Authored-By: Claude <claude-opus-5>
188 lines
10 KiB
Markdown
188 lines
10 KiB
Markdown
# Test fixtures
|
|
|
|
## The PDF fixtures
|
|
|
|
`two-line-krav.pdf` and `no-text-layer.pdf` are hand-written minimal PDFs,
|
|
regenerated by `make_fixtures.py` in this directory:
|
|
|
|
```
|
|
python3 tests/fixtures/make_fixtures.py
|
|
```
|
|
|
|
They carry no library's output — the objects are laid out by hand and the xref
|
|
offsets computed from the emitted bytes — so they are auditable byte for byte
|
|
and reproducible from that one file.
|
|
|
|
| Fixture | What it is for |
|
|
|---|---|
|
|
| `two-line-krav.pdf` | One heading plus one requirement row with label and value on the **same line**. That pairing is the property `pdfplumber` was chosen for. |
|
|
| `no-text-layer.pdf` | A structurally valid page with no text operators — the shape a scanned or image-only PDF presents. Must fail fast (`extractor_empty_pdf`), never persist as an empty concept. |
|
|
| `three-page-krav.pdf` | Three pages, one line of text each, and **the middle page carries no text operators**. The extractor drops empty pages, so the last page's text belongs to page 3 — which is what separates a page NUMBER from a count of the pages that produced text. Two pages could not tell those apart. |
|
|
|
|
## The office fixtures
|
|
|
|
`two-line-krav.docx`, `no-styles-krav.docx` and `two-line-krav.xlsx` are
|
|
hand-laid OOXML containers, regenerated by the same `make_fixtures.py`. Every
|
|
part is written out by hand and zipped with a fixed `date_time`, so they are
|
|
byte-reproducible and carry no converter's output.
|
|
|
|
**That last point is the whole policy, not a preference.** A `.docx` written by
|
|
the converter and then read by the converter proves only that the converter
|
|
agrees with itself, and would stay green through any conversion defect that is
|
|
symmetric — which is most of them.
|
|
|
|
| Fixture | What it is for |
|
|
|---|---|
|
|
| `two-line-krav.docx` | A heading plus one requirement row with label and value on the **same line** — the docx mirror of `two-line-krav.pdf`. |
|
|
| `no-styles-krav.docx` | The **same document without `word/styles.xml`**. A negative control: the body survives and the heading marker does not, which is what proves the styles part is load-bearing rather than decoration. |
|
|
| `two-line-krav.xlsx` | A sheet name that becomes a heading, plus a label/value pair on one row. |
|
|
| `tomrad.xlsx` | Four rows with the **third one empty**. The converter renders an empty row as a pipe line of nothing but spaces, which is what a pipe table's own separator line also looks like — so a rule that reads the line rather than its position swallows the row and renumbers every row after it. Found on the K2 price sheet (8 empty rows, last row reported as 92 against a workbook that says 100); this fixture is what keeps it red. |
|
|
|
|
Two things were measured while building these, and both are the same shape —
|
|
structurally valid input, silently reduced output, exit code 0 and no warning:
|
|
|
|
- **Without `word/styles.xml`** the docx extracts as flat prose with no heading.
|
|
A fixture lacking that part would pin the body and pin nothing about
|
|
structure, while looking exactly as convincing. Structure is the half the
|
|
segment proposer reads.
|
|
- **With inline strings (`t="inlineStr"`)** rather than a shared string table,
|
|
the xlsx extracts with the sheet name intact and **every cell value gone**.
|
|
The fixture therefore uses a `dimension` element and a shared string table.
|
|
|
|
## The K2 office fixture set (`k2-office/`)
|
|
|
|
`krav-presentasjon.pptx`, `krav-tekstdokument.odt` and
|
|
`krav-rikt-tekstformat.rtf` are the synthetic denominator for the three office
|
|
rows the corpus has none of. `docs/2026-09-04-k2-pptx-odt-rtf.md` measured that
|
|
denominator at **zero** — `K2/trinn1` holds 43 files and not one is a `pptx`,
|
|
an `odt` or an `rtf` — so those rows were `unmeasured` in the sense of never
|
|
having met a document at all. Regenerated by `make_k2_office.py` in this
|
|
directory:
|
|
|
|
```
|
|
python3 tests/fixtures/make_k2_office.py
|
|
```
|
|
|
|
**One document, three containers.** All three carry the same authored content —
|
|
a title, an intro, a 20-row label/value table, a caption and a 4x4 grid — so
|
|
the only variable between the three measurements is the container and the
|
|
reader that opens it. The counts are hand-counted once, in
|
|
`k2-office-fasit.json`, and shared: **56 cells, 20 pairs, 59 distinct strings.**
|
|
|
|
**The generator and the fasit live one level up, and that is not tidiness.**
|
|
Door B walks its drop directory recursively, so anything parked inside
|
|
`k2-office/` would enter the run and N would stop being 3.
|
|
|
|
Same policy as the office fixtures above, for the same reason: every part is
|
|
hand-laid and no converter wrote any of them. The commissioning order offered
|
|
pandoc as a generator option; a file written by the converter and then read by
|
|
the converter would prove only that the converter agrees with itself.
|
|
|
|
Two things were measured while building this set, both against the vendored
|
|
pandoc 3.9, and both are the house shape — structurally plausible input,
|
|
silently wrong output, exit code 0 and no warning:
|
|
|
|
- **RTF cell paragraphs need `\pard\intbl`.** Without it, consecutive
|
|
`\trowd…\row` rows are read as each row NESTED inside the previous one:
|
|
five label/value rows came back as five levels of nested table, 2076
|
|
characters where 117 were expected.
|
|
- **The `\uN?` unicode escape loses the character after it.** Measured
|
|
directly: `A\u248?BC` reads back as `AoC` (ring letter present, `B` gone) and
|
|
`A\u248?xBC` reads back as `AoBC`. The `?` is taken as the control word's
|
|
delimiter and `\uc1` then skips a real character. The fixture writes
|
|
`\uN ?` with an explicit space, which round-trips. **This is the form Word
|
|
emits**, so it is a converter finding rather than a fixture quirk — recorded
|
|
in `docs/2026-09-07-k2-pptx-odt-rtf-fixtures.md`, not worked around anywhere
|
|
in `src/`.
|
|
|
|
**Three synthetic documents in one house style are not a corpus.** The rows stay
|
|
`unmeasured` in `extract._EVIDENCE` and `tests/test_k2_office_fixtures.py`
|
|
asserts that they do.
|
|
|
|
## The proposer's default-profile golden
|
|
|
|
`propose-golden-default.json` is the artifact `tools/okf_propose_segments.py`
|
|
produces for `OUTLINE_DOCUMENT` (defined in `tests/test_propose_segments.py`)
|
|
with **no flags at all**, generated at commit `798f64a` with
|
|
`--proposed-at 2026-09-03T00:00:00Z`. The timestamp is an explicit argument
|
|
because the artifact carries it verbatim; a wall-clock default would make the
|
|
golden unreproducible by construction.
|
|
|
|
**Why the fixture is `OUTLINE_DOCUMENT` and not `DOCUMENT`.** The golden exists
|
|
to go red if any later rule is accidentally defaulted ON. `DOCUMENT` was
|
|
measured to contain **zero** bare-integer lines, so a golden over it would stay
|
|
byte-identical through exactly the regression it was named to catch -- a trap
|
|
written down but unable to fire. `OUTLINE_DOCUMENT` carries a bare-integer
|
|
ascending run of three, which today's rules do not match (measured: bare `1` /
|
|
`1.` / `1)` yield 0 candidates), so the golden pins that absence and breaks the
|
|
moment it stops being true.
|
|
|
|
**It transitively pins `observed_extractor_version`**
|
|
(`src/llm_ingestion_okf/segmentation.py`): the field is written into every
|
|
artifact, so a converter or extractor bump turns this golden red. That red is
|
|
legitimate -- read the diff and decide, exactly as for the frozen PDF literal
|
|
below. Regenerate only after that decision, never to make a red go away.
|
|
|
|
## Why the expected office text is frozen as a literal
|
|
|
|
The same reason as the PDF text below, with one addition: the literals are
|
|
pinned to a **named converter version**. `_pandoc.py` refuses any binary but
|
|
the vendored 3.9, and `tests/test_extract.py` asserts that version beside the
|
|
literals. A frozen literal without a named converter pins nothing — it says
|
|
"these bytes" without saying what produced them.
|
|
|
|
## Why the expected PDF text is frozen as a literal
|
|
|
|
`tests/test_extract.py` asserts the extracted text of `two-line-krav.pdf` as an
|
|
exact string. That is deliberate, and it is the mechanism behind a promise this
|
|
library makes everywhere else:
|
|
|
|
- Extraction is deterministic **within** a parser version. Measured 2026-08-21
|
|
across five configurations, two runs each, compared byte for byte
|
|
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`).
|
|
- Extraction is **not** guaranteed stable **across** parser versions.
|
|
`pdfplumber` pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships
|
|
date-stamped releases with no stability contract. So the real pin on extracted
|
|
text is a transitive one, and it is exact.
|
|
|
|
The consequence is worth stating plainly: **any golden fixture built on
|
|
extracted PDF text is pinned to an exact parser version, and a parser upgrade
|
|
is a fixture migration, not a routine bump.** The frozen literal is what makes
|
|
that upgrade break something visible instead of drifting silently. If it goes
|
|
red after a dependency change, the correct response is to read the diff and
|
|
decide, not to re-record the expectation.
|
|
|
|
The version range that carries this lives in `pyproject.toml`'s
|
|
`[project.optional-dependencies] extract`, with the same reasoning at the
|
|
declaration site.
|
|
|
|
## What these fixtures do not cover
|
|
|
|
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
|
|
detected table objects are clean enough to hand to `render_table` unchanged;
|
|
two independent parsers return the same wrong shape, because the breakage is in
|
|
the documents' ruling geometry rather than in either library. PDFs enter this
|
|
library as **prose**, and structured tables are out of scope until that is
|
|
decided separately.
|
|
|
|
## `propose-golden-grid-default.json`
|
|
|
|
Pins the DEFAULT proposer artifact over a document containing a pandoc **grid
|
|
table**. Generated on unmodified code, before Arm E's rule existed, with
|
|
exactly the command the test runs:
|
|
|
|
```
|
|
write(tmp_path, GRID_GOLDEN_DOCUMENT, "grid.md")
|
|
okf_propose_segments.main([source, "--out", out, "--proposed-at", "2026-09-03T00:00:00Z"])
|
|
```
|
|
|
|
It exists because `propose-golden-default.json` **cannot** pin this. That
|
|
golden is taken over `OUTLINE_DOCUMENT`, which contains no `|` row and no `+`
|
|
rule line, so no table rule -- present or future -- can move its bytes. A guard
|
|
that is structurally incapable of firing is a trap written down but never
|
|
armed. This one is taken over a document that has a grid table, so an
|
|
accidentally default-on table rule turns it red.
|
|
|
|
Both goldens transitively pin `observed_extractor_version` and `PROPOSER_VERSION`.
|
|
A red here after a dependency change is a legitimate red: read the diff and
|
|
decide, do not re-record the expectation.
|