A concept named its source file by basename and, when segmented, carried a
`source_offset` into the text THIS LIBRARY extracted. Following that pointer
needed the corpus directory, the extractor and its exact transitive version --
none of which the bundle carries. Hand-walked on a real K2 concept: six steps,
four of them requiring knowledge from outside the bundle, to learn that a
requirement sits on pages 12-13 of a 20-page document.
The address is spec's: `sources: [{ resource, title }]`, where `resource` is
the dropped file's inbox-relative path (SPEC v0.2 5.1:303-306 -- "an absolute
URL, a bundle-relative path, or a path into a `references/` subdirectory").
The locator is ours, and it has to be: 5.1 has no field for a place within a
resource, and the pinned guard (1.3.0) rejects every route to putting one
inside a `sources` entry -- a non-allowlisted key by name, a nested flow list
as "scalar leaves only", and quoting as an unsupported form. So the locator is
top-level keys shaped like `source_offset`, and a path carrying a flow
terminator is refused fail-fast rather than mangled.
The unit table is built AT EXTRACTION, where the extracted text and the
original's structure are known to agree: pdf -> `source_pages` from
pdfplumber's own page numbers (a page that yielded no text does not renumber
the ones after it), xlsx -> `source_sheet` + `source_rows`, everything else ->
`source_lines`. `source_offset` stays.
Two measurements changed the design before it shipped. A `paragraphs` key for
docx would name a number the document does not have: `<w:p>` counts of
108/27/65/176/57 against converted-markdown lines of 75/33/67/144/63, not one
pair agreeing -- so the key is `source_lines` and says what it indexes. And an
empty spreadsheet row renders exactly like a table separator: the content-based
rule ate 8 empty rows on the K2 price sheet and reported its last row as 92
against a workbook that says 100. The separator is now found by position, and
`tomrad.xlsx` keeps that red.
One profile moves. `provenance` is a policy object, `None` everywhere but
`SEGMENTED_OKF_V0_2`; the other five shipped profiles are byte-identical.
K2 rebuilt from a frozen src copy: 629 concepts, 1108 files, name set identical,
0 ids moved, 479 files byte-identical, 629 changed and 0 lines removed anywhere.
629/629 now carry an address and a locator. New ref
`sha256-tree:665563a2f74423fcbcc8e4f0b0954ee73b73985ac0418de4f6987bd162a1f7c8`;
`2f82fcfe...` is stale. The pre-pass payload does not grow by one byte
(209 092 B before and after, 18 changed lines: the ref and eight per-concept
digests) -- because an excerpt carries the body, not the frontmatter, which is
also why the consumer still cannot cite "file X page 12" from a payload alone.
Report: docs/2026-09-08-proveniens-k2.md. 1339 tests, ruff and mypy clean.
Co-Authored-By: Claude <claude-opus-5>
|
||
|---|---|---|
| .. | ||
| ingest-golden-file | ||
| ingest-golden-http | ||
| ingest-golden-okf-v0-2 | ||
| ingest-golden-segmented | ||
| ingest-golden-segmented-okf-v0-2 | ||
| ingest-golden-sql | ||
| README.md | ||
Golden extractions (ingest-spec §11)
One directory per case, ingest-golden-{source type}/, each containing:
| Entry | Meaning |
|---|---|
manifest.json |
The manifest under test. |
fixture/ |
The source content (CSV catalogue, sqlite database, or mock payloads). |
ingested-at.txt |
The fixed timestamp, one line, §5 format. |
expected-bundle/ |
The expected materialized bundle, compared byte for byte. |
ingest-golden-file and ingest-golden-sql are the conformance MUSTs. The
http source type is an OPTIONAL extension point (spec §1); this repository
implements it, and ingest-golden-http exercises it against the mock payloads
in its fixture/ directory — the conformance suite never opens a socket and
runs without credentials (§11). For the sql case, the test sets the
OKF_GOLDEN_SQL_DB environment variable to the fixture database path; for the
http case, the injected mock transport serves fixture/{query path}.
The suite (tests/test_golden.py) re-materializes each case into a fresh
directory and compares the result against expected-bundle/ file by file,
byte for byte — the load-bearing golden-regression seam (§11).