Operator decision 2026-09-21: nothing from that consumer's collection goes out on the public remote. The NAME stays where it is already published -- it is a consumer of this library, named as such, and removing it would mean rewriting published history, which this repository does not do. What goes is everything that describes their CONTENT. Removed across README, CLAUDE.md, CHANGELOG, four dated reports, the consumption contract, three source modules and three test modules: their corpus's document and page counts, the concept count of a bundle built from it, the byte figures of a payload built from it, the question and fasit counts and recorded score of their evaluation set, a bundle id with two content refs, an order id naming them, and a path into their repository. Kept, because the argument survives without the corpus: RATIOS and percentages. A ratio is the finding -- a withheld list that is 65.5 % of a payload is a defect at any corpus size -- and it discloses nothing about how large anyone's collection is. Where a claim lost its denominator it now SAYS so rather than quietly reading as unmeasured: the gate-refusal limitation in the README states that the corpus and its counts are deliberately withheld and points the reader at their own build, which is the number that binds them anyway. One integrity pin is kept and named here rather than left to be found: the retrieval gate still pins that set by sha256, because the pin is what refuses a self-written file in the right shape, and a checksum discloses nothing about what it checksums. Its recorded SCORE is gone -- that was their figure about their own corpus, and the row now says so instead of restating it. The known-positive constants move with the contract document, as they must. Suite green, 2372 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2265 lines
141 KiB
Markdown
2265 lines
141 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to this project will be documented in this file.
|
||
|
||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||
|
||
## [1.0.0] — 2026-09-20
|
||
|
||
### Added
|
||
|
||
- **A document the gate refuses WHOLE is named in the run's own summary.**
|
||
Measured 2026-09-20 against a real corpus of official documentation built
|
||
with the shipped default gate: 17 sources were refused outright, 16 of them
|
||
among its ordinary reference pages, and the summary said only
|
||
one `fail_secure` line and one `quarantine_review` line. The count of
|
||
documents the gate dropped was not there (`rejected (coded)` sums gate
|
||
refusals and extraction failures, which have different remedies), the names
|
||
were not there, and neither was the way out. `okf build` now prints a
|
||
`Documents the gate refused WHOLE` section carrying all four — the count with
|
||
its denominator, the names (capped at ten, with the rest in the bundle's
|
||
`log.md`), the codes, and `--gate none` for a source you vouch for yourself —
|
||
and repeats it in one line on stderr, where a redirected stdout cannot hide
|
||
it. `log.md` gains a bullet naming every refused document, uncapped. **The
|
||
exit code does not move**: the build is valid, every refusal is coded and the
|
||
bundle is a true record of what the gate allowed; what was wrong was the
|
||
silence. A run the gate refused nothing from is byte-identical, in the
|
||
summary and in `log.md`.
|
||
- **A `Known limitations` section on the front page**, high up and before the
|
||
install detail: the gate's measured refusals and the way out, the absent
|
||
ceiling on what one run pays for images, the three gates of this repository
|
||
that are RED and what each red row means for a user, what the content
|
||
accounting does not count, and the rough edges nothing is planned for. No new
|
||
measurement — every number was already taken.
|
||
- **The payload says what of the question it reached** — a new top-level
|
||
`coverage` member carrying three lists: the terms the pre-pass read the
|
||
question as, the terms no concept in the bundle answers, and the terms no
|
||
delivered excerpt answers. Without it a reader holding eight excerpts cannot
|
||
tell a bundle that ANSWERED its question from one that merely ranked
|
||
something; the two payloads have the same shape. Documented as SS 8 point 7
|
||
of `docs/consumption-contract.md`, and the generated consumption skill is
|
||
told to read it.
|
||
- **Facts, and no verdict, which is a measurement rather than caution.** Two
|
||
readings were built and both falsified over **81 questions** (16 synthetic,
|
||
65 across three real gold sets, 2026-09-20): the share of a question's
|
||
terms a delivered excerpt answers separates the synthetic controls at 0.33
|
||
against 0.50 and then REVERSES on real data, where covered questions run
|
||
down to 0.27 while one genuinely uncovered question sits at 0.71; and the
|
||
share of a bundle tying the best lexical match is ~0.00 for every question
|
||
in a large bundle, covered or not. Question style dominates the first and
|
||
corpus size the second, so a pre-pass emitting a verdict would assert
|
||
across corpora what was measured on one.
|
||
- **Contract change, and the cost to a consumer is a re-run.** Every payload
|
||
grows the member; the checker does not read it, so a third-party pre-pass
|
||
that omits it stays conformant. The SS 7.4 known-positive moves with the
|
||
document it is measured on (14 721 / 375 → 16 389 / 417).
|
||
- **The retrieval gate is measurable where it was assertable**
|
||
(`tools/okf_retrieval_gate.py`, not shipped in the wheel):
|
||
- Row 8 prints the identity of every bundle it measured — path,
|
||
`bundle_id` and content ref — beside the set's sha256. Measured the same
|
||
day: two builds of one consumer's corpus carrying the SAME `bundle_id` at
|
||
different refs score differently on the same pinned set, which is why the
|
||
ref and not the id is what a row is attributed to.
|
||
- `REAL_SET_PINS` states what each of the three real sets IS — questions,
|
||
fasit entries, controls and sha256 — so a self-written file in the right
|
||
shape is refused instead of reading `1 of 1 | 3 of 3 | GREEN`.
|
||
- Row 5 reads the hold-out threshold as a number in [0, 1] and RUNS the
|
||
registered set against the registered bundle, printing
|
||
`answered of asked = share against threshold`. `bool(threshold)` was the
|
||
whole check, so `report-only; any number is acceptable for v1` passed it.
|
||
- Row 4's marking reads the payload's `coverage`: `UNANSWERED_BAR = 2/3`,
|
||
swept and collapsing at both ends (at 0.50 eleven real covered questions
|
||
are marked; at 0.70 the row falls to 5 of 6). The margin is thin — 0.6087
|
||
against 0.6667 — and what it does not catch is published with it.
|
||
|
||
### Changed
|
||
|
||
- **The two `pip install` lines under "Install in detail" install
|
||
`[extract]`.** The first screen installs `llm-ingestion-okf[extract]` and
|
||
those two omitted it, so a reader following the detailed instructions got a
|
||
build that reports `resolved converter path: unresolved
|
||
(extractor_extra_missing)` and reads no binary format. Two recipes, two
|
||
different installations.
|
||
- **Version `1.0.0`.** The scope this tool is finished at. It adds no
|
||
capability over `v0.10.1`; what it adds is that the tool says what it does
|
||
not do. After this tag the library is touched for defects found in its own
|
||
use, and the next round is Google OKF v0.3.
|
||
- **A withheld concept now carries the rule that actually decided it.** The
|
||
source quota filters the WHOLE ranked candidate list rather than the top
|
||
`k`, so every over-quota candidate came back `source_quota_exceeded` —
|
||
including the ones the RANK had already put outside `k`, which the quota
|
||
only reached because it ran first. Measured on 25 real misses 2026-09-17:
|
||
**13 of them** were labelled by the quota and decided by the rank.
|
||
`consume._fates_without_quota` asks the same cut what would have become of
|
||
each candidate with no quota in force, and the drop keeps THAT rule; only a
|
||
candidate the quota-off cut would have delivered is named as the quota's.
|
||
The budget step is lifted into `consume._pack` and used by both, so the
|
||
quota-off fate is decided by the code the run itself uses. This moves the
|
||
`rule` string a consumer reads for some withheld entries; no delivery, no
|
||
rank and no excerpt byte moves, and no committed payload in this repository
|
||
changed. The retrieval gate's row 3 goes **2 of 5 RED to 5 of 5 GREEN**.
|
||
|
||
## [0.10.1] — 2026-09-19
|
||
|
||
### Removed
|
||
|
||
- **`tools/okf_adjudicate.py`.** It shelled out to a model CLI at an absolute
|
||
path on one machine, which is the one thing nothing in this repository does:
|
||
no code here starts another program to judge anything. Its tests go with it.
|
||
The two entries below under earlier versions describe what that tool did
|
||
while it existed and are left standing — a changelog that edits its own past
|
||
is not a record. The K3/K4/K5 reports that used it now say so in the past
|
||
tense.
|
||
|
||
### Added
|
||
|
||
- **An MCP surface over OKF bundles, in two shapes, plus a generic
|
||
consumption skill.** `okf mcp --bundle <dir>` serves exactly one bundle;
|
||
`okf mcp --root <dir>` (repeatable) serves every bundle under the roots and
|
||
knows none of them by name. Four tools — `okf_list`, `okf_describe`,
|
||
`okf_ask`, `okf_fetch` — each with its reason written into the description a
|
||
client reads. A single-bundle server exposes **three**: `okf_list` is absent
|
||
where there is nothing to list, because a tool that always returns the same
|
||
one row invites a client to treat discovery as available when the deployment
|
||
does not have it. The eval was written RED first (`tools/okf_mcp_gate.py`,
|
||
`5f1772e`); the capability follows.
|
||
- **The protocol is written narrowly with stdlib only, and that is the
|
||
packaging invariant kept rather than an aesthetic.** An MCP SDK would be
|
||
this package's second runtime dependency on the DEFAULT install path, for
|
||
four JSON-RPC methods and a newline framing, and
|
||
`test_the_only_runtime_dependency_is_the_security_boundary` pins that list
|
||
literally. `uv.lock` is untouched.
|
||
- **Nothing is cached across calls.** Every call re-walks the roots and
|
||
recomputes the bundle's content identity, so a bundle added, removed or
|
||
rebuilt while the server runs is seen by the next call without a restart, a
|
||
configuration edit or a code change — measured, 9 of 9 discovery checks over
|
||
three bundles written while the process was serving. The cost is paid per
|
||
call: 0.75 s for the identity of a 2 756-concept bundle, 5.6 s for one ask.
|
||
- **Containment is two independent checks**: the bundle's own index must name
|
||
the concept, and the resolved path must be inside the bundle. Removing
|
||
either one alone still refuses — with a different code, which the gate
|
||
asserts by name — and removing both is caught by the gate's row 6.
|
||
- **`okf card <bundle>`** prints one bundle's identity, concept count,
|
||
conditional-field counts and whole-bundle cost as JSON, DERIVED on every run
|
||
and never written into the bundle. **`okf skill --generic`** writes one
|
||
installable consumption skill for ANY bundle, carrying no bundle's identity
|
||
or numbers and pointing its reader at the card. Measured: two per-bundle
|
||
skills are identical on 280 of 312 and 310 lines, and what differs is
|
||
exactly what goes stale on a rebuild.
|
||
- Report: `docs/2026-09-20-mcp-to-varianter.md`. The gate stands RED on row 2
|
||
(83 of 181 anchors of the frozen graded set reached, of which 99 are present
|
||
in the bundles at all and 0 were met by a concept-id lookup), and the
|
||
architecture choice between the two shapes is the operator's.
|
||
|
||
- **Every carried image is now one a model can be SHOWN, and the ones that
|
||
cannot be are refused out loud.** Until this round the asset path carried
|
||
whatever format a publisher shipped. Measured 2026-09-19 over the frozen
|
||
R761 delivery's own `assets/` (denominator 50): 29 JPEG, 2 PNG and **19 "PC
|
||
bitmap, Windows 3.x, 8-bit, compression 1"** — RLE8 BMP. The 19 are
|
||
byte-correct files nothing reads, so 19 of that document's figures were
|
||
present and invisible at once, with `images: N` reporting that they had
|
||
arrived.
|
||
- `assets.VIEWABLE_MEDIA_TYPES` states the set (`image/png`, `image/jpeg`,
|
||
`image/gif`, `image/webp`) and `read_image` tests every asset's SNIFFED
|
||
type against it. It is a property, not a list of formats we happened to
|
||
meet: a format nobody here has seen is refused by the same rule that
|
||
refuses TIFF.
|
||
- **BMP is converted losslessly to PNG** — 8-bit uncompressed, 8-bit RLE8
|
||
and 24-bit uncompressed. The reader is stdlib (`struct` + the existing
|
||
`zlib` PNG writer) and adds NO dependency. Pillow was measured first and
|
||
rejected on two counts: `read_image` is on the CORE path (`.html` and
|
||
`.xml` carry images with no `[extract]` extra), and an asset's name is its
|
||
content digest, so encoding through an installed library would make a
|
||
bundle's identity move with that library's version — the property 0.10.0
|
||
felled page rasterisation over. Pillow is the INDEPENDENT decoder in the
|
||
tests instead.
|
||
- **Lossless, measured on the real files:** all **19 of 19** R761 RLE8
|
||
assets convert with RGB identical to Pillow's decoding of the source,
|
||
**2 366 365 pixels** compared.
|
||
- **Traceability per converted asset**, on the pointer line where the rest
|
||
of the asset metadata already lives: the original media type, the original
|
||
sha256 in full, and the new one. A converted asset is ONE asset — one file
|
||
in `assets/`, one pointer, one row in the accounting.
|
||
- **The ceiling is paid before the pixels exist.** The BMP reader bounds the
|
||
DECLARED size through the same `check_size` the rest of the image path
|
||
uses, before a row is allocated, and an RLE run is written as one clipped
|
||
slice — painting pixel by pixel would leave the memory bounded and the CPU
|
||
unbounded, since a megabyte of `FF` runs is a hundred million paint steps
|
||
against a 32-pixel frame.
|
||
- **Two new codes.** `asset_not_viewable` — a real image in a format no
|
||
model can be shown and with no lossless conversion here (TIFF, JPEG 2000).
|
||
`asset_bmp_unsupported` — a BMP variant this reader does not express
|
||
(RLE4, BITFIELDS, 16/32-bit, BITMAPCOREHEADER, over 256 palette entries).
|
||
Both leave a "not carried" line in the concept and a row in the run log.
|
||
- **The cost, measured with a committed script** (`tools/okf_asset_census.py`,
|
||
one row per image, run from two pinned trees over 18 403 files and 67
|
||
PDFs, **9 714 image rows**): exactly **35 rows moved** — 19 BMP now
|
||
carried as PNG, and **16 JPEG 2000 objects** that stop being carried and
|
||
become `asset_not_viewable`, because no stdlib route decodes JPEG 2000.
|
||
**9 321 of 9 321** JPEG and PNG rows are byte-identical on both sides, so
|
||
not one already-viewable picture changed hands.
|
||
|
||
- **One normalisation door in front of the persist gate: U+00AD is removed and
|
||
COUNTED** (operator decision 2026-09-18). `llm-ingestion-guard` 1.4.0 keeps
|
||
the soft hyphen in `_ZERO_WIDTH_CPS`, and `output:zero-width-present` is an
|
||
any-tier carrier — `fail_secure` at every trust level, no sanitisation, no
|
||
exception. R761 Prosesskoden:2025 carries 71 of them and 0 of the four real
|
||
zero-width characters; all 71 are Norwegian hyphenation points inside words,
|
||
so a 701-page process code was unreadable for the whole chain over
|
||
typography. `extract.normalise_extracted` removes that one character from
|
||
every extracted text; `ExtractedDocument.soft_hyphens`,
|
||
`InboxResult.normalised` and the accounting's `normalised_soft_hyphen` carry
|
||
the number per document and per run, and `log.md` gains a `**Normalisation**`
|
||
bullet. The guard is untouched, the other four characters and U+00A0 NBSP are
|
||
untouched, and a real zero-width character is still `fail_secure`. Reach,
|
||
measured: **0 of the 78** readable documents of the reference corpus carry
|
||
any of the six, so no bundle measured here moves.
|
||
- **`refused` in the accounting: a partial refusal is never silent.** The
|
||
report and `log.md` now say how many of M documents the run persisted nothing
|
||
of. The exit code is unchanged — it belongs to the whole run.
|
||
|
||
- **`okf build --accounting PATH`: content accounting per element.** Before
|
||
extraction, every source document is inventoried in a per-format element
|
||
vocabulary: headings, paragraphs, tables, cells, images, and so on. After
|
||
the run, every element gets exactly one fate: `carried`, `pointer` or a
|
||
coded rejection. The fates are written as JSON to PATH and summarised in
|
||
`log.md`. The build then exits 1 when any element is unaccounted or booked
|
||
twice.
|
||
- **`carried` is checked, not declared.** A persisted document's element is
|
||
carried when all of its text is found in the concept bodies written for
|
||
that document (letters and digits, case-folded). A document the gate
|
||
refused books every element as rejected with the gate's code, and its
|
||
`log.md` line says what the source held.
|
||
- **The judge is `tools/okf_accounting_gate.py`**, written red first
|
||
against an independent witness (`tools/okf_witness.py`, which imports
|
||
nothing from this package). At this change it is green on all six rows,
|
||
including R761 Prosesskoden:2025: 110 of 110 units under both the default
|
||
gate and `--gate none`.
|
||
- **Opt-in, measured.** On the 43-document reference corpus the build took
|
||
+744 s (+19 %) and +0.53 GB peak RSS.
|
||
- **The account is over the element classes the vocabulary knows.** A file
|
||
whose suffix has no reader is accounted at file level only, and a part of
|
||
a document no vocabulary names is not counted — `.docx` headers, footers,
|
||
endnotes and comments, `.pptx` speaker notes, `.xlsx` cell comments and
|
||
formulas, the `.rtf` header/footer groups. Content there can go missing
|
||
under exit 0 and `0 unaccounted`; README states the list.
|
||
- **The reference corpus fails the check, with 24 real losses:** 22 images
|
||
on PDF pages without a text layer, which the reader drops together with
|
||
the page, and 2 docx Title paragraphs, which the converter moves into
|
||
metadata. A default-on door would therefore fail builds that pass today.
|
||
Report: `docs/2026-09-17-innholdsregnskapet-bygget.md`.
|
||
- **Opt-in by operator decision (2026-09-17)**, until the losses it reports
|
||
on the reference corpus are fixed. Of the three exceptions the gate
|
||
proposes, the operator approved the PDF one only; approving it moves no
|
||
number, because no witness counts a heading in a PDF.
|
||
- **Limit, measured:** the check proves that a string is present, not
|
||
where. Short elements such as a section label or a one-word title are
|
||
often found elsewhere in the same document. With R761's concept text cut
|
||
to half, 4 823 paragraphs and 3 621 sections were reported lost, but only
|
||
3 titles and 16 labels.
|
||
|
||
### Changed
|
||
|
||
- **`okf build` exits 1 when it extracted at least one document and
|
||
persisted none.** Until now such a run exited 0, because every refusal was
|
||
coded and the conservation identity held. The bundle was nonetheless empty.
|
||
Measured case: guard 1.4.0 refuses R761 Prosesskoden:2025 whole, because of
|
||
its 71 soft hyphens (U+00AD). Door B's library function
|
||
(`process_inbox`) and `corpus.measure` are unchanged; for a hostile inbox,
|
||
"all rejected" is a correct outcome.
|
||
- In this repository, one test relied on exit 0:
|
||
`tests/test_cli_gate.py::test_build_refuses_a_document_the_real_guard_refuses`.
|
||
- No script here does. `okf project` calls the build as a function and is
|
||
unaffected.
|
||
- **A file carried through a document is no longer also a coded rejection.**
|
||
Since 0.10.0, an image beside a document was carried into `assets/` through
|
||
that document and was ALSO counted as `extractor_unknown`, so one file had
|
||
two fates. On R761 under `--gate none` that was 50 files.
|
||
- The conservation identity is now `merged + files carried through a
|
||
document + coded rejections = N`.
|
||
- `log.md` writes the middle term only when it is non-zero, so a corpus with
|
||
no such files keeps its line byte for byte.
|
||
- The carried files are the references the reader actually resolved and
|
||
carried (`ExtractedDocument.files`), never a byte match. A byte match
|
||
would credit R761's 7 unpointed duplicates.
|
||
- An unpointed file beside a document stays a coded rejection.
|
||
- **`log.md`'s `Images: C carried of F found`**: with `--accounting`, F is
|
||
what the SOURCES declare. A refused document's pictures therefore no longer
|
||
read as "0 of 0 found".
|
||
|
||
### Security
|
||
|
||
- **A document can no longer forge a carry in the content-accounting gate
|
||
(0.10.1).** New in the viewable-asset round: a converted image's own bytes
|
||
are not in `assets/`, so `asset_holds` gained a second route that reads the
|
||
two digests the bundle states on the pointer line. The expression ran over
|
||
the WHOLE bundle text, so a document could simply write the sentence.
|
||
Measured by PM 2026-09-19: a BMP declaring 50 000 x 50 000, refused
|
||
`asset_too_large` and absent from `assets/`, was reported as held — through
|
||
an image's `alt` text, and through ordinary body text. Before that route
|
||
existed, `asset_holds` hashed the source file and nothing a document wrote
|
||
could reach it; the gate's own first sentence is that the fasit never comes
|
||
from the reader it judges, and `claimed and not found` could be silenced by
|
||
a document that asked for it.
|
||
- The claim now counts only inside a POINTER BLOCK this code wrote, and only
|
||
where it names the asset that block points at. That closes body text and a
|
||
table cell.
|
||
- An image's LABEL is document text written inside a pointer block, so
|
||
`assets._inline` disarms a checksum field in anything that came from the
|
||
document: the digits are kept, the colon that makes them a field is not.
|
||
That closes the `alt` route. Neither half is sufficient alone.
|
||
- Three mutants in `tools/okf_gate_mutants.py`, one per check, each felled by
|
||
its own arm; the harness now runs the copy with its own `src/` on
|
||
`PYTHONPATH`, because an editable install made a `src/` mutant resolve to
|
||
the working tree and survive without having been applied.
|
||
|
||
- **A remote image reference is no longer a live markdown image link
|
||
(0.10.1).** New in 0.10.0: before it, no reader read an `<img>` attribute at
|
||
all. A document could put ``
|
||
into a persisted concept, with the address and query string chosen by
|
||
whoever wrote the document. This package opens no socket, but a consumer
|
||
that renders the bundle — or an agent that fetches what it renders — does,
|
||
which turns "this bundle was opened" into a beacon, and a server-side
|
||
consumer into an SSRF. The guard refuses such a line at
|
||
`guard-user-upload` and the build's default tier does not, so the same bytes
|
||
were persisted under the default and refused one tier up. A remote reference
|
||
is now inert text with the address in ONE code span, and a property test over
|
||
the readers asserts that no reference produces a markdown image link outside
|
||
`assets/`. Found by an independent review of 0.10.0 before it was pushed.
|
||
- The first fix wrote the address **twice** — once in a code span and once
|
||
bare — and a GFM/linkify renderer autolinks a bare URL into `<a href>`.
|
||
It takes a click rather than a render, so it is weaker than an image link,
|
||
but "inert" was half true. The address is now written once.
|
||
- The first fix also **dropped the caption**: `label` stayed in the
|
||
signature of the line that says what is missing, and no branch read it, so
|
||
the alt text or figure caption of an image the bundle does not carry was
|
||
lost — a regression against 0.10.0 and against that line's own reason for
|
||
existing. It is written again, in the same `-- <label>` form a carried
|
||
pointer uses.
|
||
- **An image is bounded in three places, and the third is what the run pays
|
||
(0.10.1).** Nothing limited a PDF image's size: a 9.6 KB file declaring
|
||
3 000 x 3 000 grayscale zeros took 83 MB of peak RSS and a 63 KB one
|
||
declaring 8 000 x 8 000 took 276 MB, linear in the pixel count, so one
|
||
document could exhaust memory and take a whole batch build with it — before
|
||
any gate, because the guard never sees image bytes.
|
||
- The size a container **declares** (`/Width` x `/Height`, an IHDR, a
|
||
`data:` payload's encoded length) is checked against `MAX_IMAGE_PIXELS`
|
||
(40 000 000) and `MAX_IMAGE_BYTES` (256 MiB) before anything is decoded.
|
||
- The size a carried **file** has is checked the same way. This package
|
||
never decodes such a file, so it pays nothing for it — but a 7 000 x 7 000
|
||
PNG of 47 705 bytes written into a bundle hands the consumer the same bomb
|
||
with `7000x7000 px` printed beside it.
|
||
- What the **stream** behind a PDF image decompresses to is measured, a
|
||
chunk at a time and discarded, before `get_data()` is called. That is an
|
||
independent number from the declared size: `/Length` is the compressed
|
||
length, and a second independent review measured a 408 516-byte PDF
|
||
declaring a 1x1 picture and carrying 400 MB of deflated zeros being
|
||
CARRIED, with no rejection, at 892 MB of peak RSS. With the bound: 0
|
||
carried, `asset_too_large`, 54 MB — and 62 MB where the old path cost
|
||
2 436 MB, so the cost no longer scales with the bomb.
|
||
- **Every link of the chain is measured, not only the first.** A PDF
|
||
decodes a stream through a list of filters, and the first fix read
|
||
`filters[0]`: `/Filter [/FlateDecode /FlateDecode]` therefore cost
|
||
886 554 624 bytes of peak RSS from 1 636 bytes of file, and three links
|
||
cost the same from 1 070 — about 542 000x, with the picture still refused
|
||
at the end, after the memory had been spent. It also left the 16 corpus
|
||
image objects behind an `[/ASCII85Decode /FlateDecode]` chain unmeasured,
|
||
because `filters[0]` is not `FlateDecode` there. Bounded, measured idle in paired subprocesses: 52 367 360 bytes at two
|
||
links, 61 390 848 at three, and 60 403 712 where the old path cost
|
||
2 567 204 864.
|
||
- **What a link COSTS is bounded, not the size of its output.** Bounding
|
||
every `FlateDecode` was still not a bound: the bomb moved into
|
||
`ASCII85Decode`, which the previous fix had classed as safe "because it
|
||
shrinks". `z` is that encoding's shorthand for four zero bytes, so the
|
||
filter quadruples its input, and `base64.a85decode` appends one 4-byte
|
||
object per group to a list — about a hundred bytes of memory per byte of
|
||
INPUT (measured on CPython 3.14: 101.4x at 1 MiB, 96.1x at 4 MiB, 94.5x at
|
||
16 MiB). Measured in paired subprocesses, idle machine, the document built
|
||
once and read from a file: a 33 475-byte PDF decoding through
|
||
`[/FlateDecode /ASCII85Decode]` cost 3 261 599 744 bytes of peak RSS and
|
||
the picture was CARRIED; bounded, 42 070 016 and `asset_too_large`.
|
||
Doubling the run of `z` takes the old cost to 6 461 558 784 and the
|
||
bounded one to 40 280 064, so the cost no longer follows the bomb. A
|
||
single `[/ASCII85Decode]` link went 933 085 184 → 62 484 480, and
|
||
`[/Fl /A85 /Fl]` 3 519 180 800 → 43 438 080 (and from
|
||
`asset_samples_invalid` to a bound's own code).
|
||
- **Every permitted filter now carries a measured cost ratio**
|
||
(`assets.PDF_FILTER_COST_RATIO`) and a per-link budget
|
||
(`MAX_FILTER_DECODE_BYTES`, 512 MiB). `FlateDecode` is measured a chunk at
|
||
a time as it is paid, under a limit that is the smaller of the picture's
|
||
own bound and what the NEXT link's decoder may be handed, so the budget
|
||
travels down the chain. Every other permitted filter has its cost
|
||
PREDICTED from its input size before its decoder is called, because those
|
||
decoders take a whole string and return a whole string. The cap that falls
|
||
out for `ASCII85Decode` is read off the corpora: of the 9 668 image
|
||
objects of the 77 PDFs measured, 16 decode through such a link and the
|
||
largest input to one is 450 739 bytes, more than ten times under it.
|
||
- **A filter with no measured ratio is refused unread**, with its own code
|
||
`asset_pdf_unbounded`: `LZWDecode`, `RunLengthDecode`, `CCITTFaxDecode`,
|
||
`/Crypt` and anything unknown.
|
||
- **A property test runs every chain of length 1–3** over the ten filters
|
||
pdfminer decodes — 1 110 of them, each with an amplifying payload —
|
||
and requires each to be delivered under the bound or refused with a code
|
||
in the published vocabulary, never paid for on the way. The known-positive
|
||
beside it holds that every chain over the permitted filters still carries
|
||
a small image. Over the 5 142 image objects of
|
||
the 78 PDFs measured, the refused class is 4 `CCITTFaxDecode` objects,
|
||
which are 1-bit stencil masks and were already refused one step later by
|
||
the encoder. Measured by name over the same 78 documents, carried images
|
||
are 9 356 before and 9 356 after: no document loses a picture, and the 8
|
||
objects that move code (4 masks, counted twice) were refused on both sides.
|
||
- **An encrypted stream is deciphered and then measured.** Deciphering does
|
||
not change a stream's length, so this does what pdfminer's own `decode()`
|
||
does; before, `stream.decipher is not None` returned without measuring,
|
||
which made "the document declares encryption" a way past the bound.
|
||
- What remains outside the bound is a stream something else has already
|
||
decoded, where the memory is spent before this package is asked. That one
|
||
is caught by a check on `len(data)` AFTER `get_data()`, which is a counted
|
||
refusal and not a bounded one. The difference is stated in the code rather
|
||
than implied — and, since this change, held by a test: deleting exactly
|
||
that line passed all 2 132 tests before it.
|
||
- **A declared size that is not a size is refused with its own code
|
||
(0.10.1).** `/Width -1 /Height 40000000000` multiplies to a NEGATIVE pixel
|
||
count, under which every bound read as satisfied: the check returned
|
||
silently, 400 MB was decompressed, and the refusal arrived from the PNG
|
||
encoder as `asset_samples_invalid` — a code about a sample buffer for a
|
||
defect in the declaration. A non-positive dimension is now `asset_size_invalid`,
|
||
raised before the stream is read. Its own code because a legitimate
|
||
publisher shipping a picture larger than this package carries and a
|
||
dictionary written to be read wrong are different facts about a document.
|
||
|
||
### Fixed
|
||
|
||
- **The retrieval gate had to resist the work it judges: four of eight
|
||
cheating attacks went through it, and they are closed (2026-09-19).** PM's
|
||
checkpoint on `2c8296b` took rows 3, 5, 7 and 8 GREEN without one label
|
||
becoming true or one concept ranking better. An eval written before the
|
||
capability has one job beyond being red today, so the gate was repaired
|
||
before anything is built against it. `src/` is untouched.
|
||
- **Row 8 requires all three named sets** (`wiki-20`, `r761-sk2`,
|
||
`vegnormal-32`) and is NOT RUN otherwise. It counted whatever `--real`
|
||
gave it, so one set of three read `6 of 6 GREEN` — and this repository's
|
||
own test asserted `(1, 1, GREEN)` for a single set. The numbers the run
|
||
DID measure are still printed: a missing set must not cost the reader the
|
||
set that was measured.
|
||
- **Its headline is at QUESTION granularity**, and was `quoted + concept`
|
||
over `quoted_units + concept_units` on the line above the detail saying
|
||
the two are not summed. The three sets share no unit — a citation, a
|
||
section title and a requirement number — so their sum is a number that is
|
||
none of them.
|
||
- **Rows 2 and 3 take their denominator from the pinned set, not the run.**
|
||
At `k = 32` the fixtures declaring class b are delivered, and they used to
|
||
leave the denominator: rows 1, 2, 3 and 6 all read green at once. A forced
|
||
fixture that stops producing its declared class is a BROKEN PREMISE now,
|
||
printed as one and counted against its row.
|
||
- **Row 3 carries a known-positive.** With `--source-quota` off every
|
||
printed reason is true — not a lie, an empty measurement — so a set may
|
||
declare `source_quota_in_force` and the row is NOT RUN for it when the
|
||
default and quota-off cuts deliver the same concepts. **The control's own
|
||
premise was measured first and was false where it was first put:** over
|
||
the five existing sets the two cuts deliver the SAME concepts (the quota
|
||
is topped back up), 52 labels moving `source_quota_exceeded` → `below_k`
|
||
with 0 deliveries changing. `set-quota.json` is the one set where the
|
||
quota genuinely decides.
|
||
- **Row 5 reads git for the half a registration cannot assert about
|
||
itself.** Two files PM wrote in the moment came back `7 of 7 GREEN`. Three
|
||
of its ten checks now read history: committed and unmodified, its commit
|
||
is not itself a ranking change, and a ranking change landed AFTER it —
|
||
the last being the one that cannot be self-attested. What git cannot prove
|
||
(that nobody read the number first) is stated in the row.
|
||
- **Row 7's roster is pinned apart from the list it names.** The bar is a
|
||
share, so seven duplicate `k = 1` mutants read `18 of 20 GREEN` with the
|
||
same two survivors. `MUTANT_ROSTER` and `MUTANT_COUNT` are separate
|
||
constants, duplicates are refused, and the bar is the roster's length.
|
||
- **The synthetic corpus is pinned like the sets** (`SPECS_SHA256`). A tuned
|
||
corpus was caught by row 2's forced classes and not by a pin.
|
||
- **`M14` closes PM's G9**: `hit = bool(hit_ids) and bool(confirmed)` is
|
||
reached only by a delivery that still carries the citation and is no
|
||
longer the concept file's bytes. It is felled and no production line
|
||
changed — the term was observable and unobserved. The judge's
|
||
independence is measured with it: index warmed BEFORE the patch, every
|
||
unit a miss; index built UNDER it, every unit a hit. The gate never builds
|
||
one under a mutation, and that is in `LIMITS`.
|
||
- **Row 9 takes `--k2 SET SHA256 BUNDLE`** in this gate's own set shape, and
|
||
a set of another size is refused as another set wearing K2's name. Without
|
||
one it stays RED rather than NOT RUN: its denominator is known.
|
||
- **Row 8 ran, against all three real sets**: **44 of 64 questions**, 7 of
|
||
29 at citation granularity and 38 of 50 at concept granularity, 33 of 34
|
||
misses class b. wiki (6 of 20) and r761 (7 of 7) reproduce PM's recorded
|
||
figures exactly; vegnormal measures 31 of 43 citations where PM recorded
|
||
32, a one-citation disagreement between two instruments over the same
|
||
pinned bytes, stated and not resolved here.
|
||
- Rows 1 and 6 go 9 of 9 to 10 of 10 (one added fixture, one added hit).
|
||
Every other row is unchanged and the verdict is unchanged:
|
||
`GATE RED: rows 3, 4, 5, 7, 8, 9`, exit 1, byte-identical over two runs.
|
||
`mypy --strict` on the gate goes 8 errors to 0. Report:
|
||
[`docs/2026-09-19-gjenfinningsgaten-motstand.md`](docs/2026-09-19-gjenfinningsgaten-motstand.md).
|
||
|
||
- **The conversion claim the content-accounting gate believes now comes from
|
||
the RUN, not from the bundle's prose (0.10.1).** The previous round bound
|
||
the claim to a pointer block, which closed the two forgeries PM had
|
||
measured and did not close the class: a pointer block is two lines of
|
||
markdown, and one ordinary HTML file with two `<p>` elements writes them.
|
||
Reproduced through the real `okf build` — a BMP refused `asset_too_large`
|
||
and absent from `assets/` read as CARRIED, from a document naming one
|
||
digest that is public in the bundle and one that is computable in advance.
|
||
- `okf build --accounting` now books every conversion the run performed:
|
||
`assets.conversion` names the `(source digest, asset digest)` pair,
|
||
`DocumentAssets.conversions` carries it out of the run and the accounting
|
||
JSON states it per document as `conversions: [{from, to}]`.
|
||
- The gate reads the pair from there and uses the bundle text only to
|
||
CONFIRM it. The confirmation can be forged and the ledger cannot, which is
|
||
why the ledger decides.
|
||
- **Chosen over neutralising pointer-shaped text at extraction**, because
|
||
that fix changes what every document SAYS in order to defend a tool
|
||
outside the build: a source quoting a bundle listing would come out
|
||
altered and existing bundles would move bytes.
|
||
- A build run with no accounting door has no ledger, so a converted image
|
||
is reported claimed-and-not-found rather than believed. That is the same
|
||
reading the gate had before the conversion route existed.
|
||
- Measured: the three arms PM reproduced go forged → refused, 3 of 3, with
|
||
the known-positive (a BMP the run really does convert) True in all three.
|
||
The text-level regression guard goes 3 arms to 13. R761 rebuilt is
|
||
`diff -r`-identical, 50 assets (29 JPEG + 21 PNG), 19 of 19 conversions
|
||
confirmed against 19 declared, SHY 71, u = 0, d = 0.
|
||
- **An RLE8 stream that stops before the frame is refused (0.10.1).** The
|
||
terminator rule added earlier in this version asks only that a stream SAY it
|
||
is finished, and a stream can say so anywhere: measured 2026-09-19, one
|
||
whose FIRST two bytes are the end-of-bitmap escape was carried with 32 of 32
|
||
pixels never decoded, while an independent decoder refuses the same file.
|
||
`_bmp_rle8_rows` now also requires the cursor to stand at or past the end of
|
||
the last row (`asset_samples_invalid`).
|
||
- **The line is the cursor and not the pixels.** A delta escape and an
|
||
end-of-line escape state their skip, so the pixels they pass over keep
|
||
index 0 and every decoder produces the same picture; a pixel-coverage
|
||
count would refuse both constructions the format defines. The corpus
|
||
cannot choose between the two rules — over the 25 RLE8 BMPs the R761
|
||
delivery ships, 25 of 25 paint every pixel, 25 of 25 reach the end of the
|
||
frame and 0 of 25 use a delta — and an independent decoder can: Pillow
|
||
reads 5 of the 8 streams in the table and refuses the same 3 the new rule
|
||
does, one of them short by a single pixel.
|
||
- Two docstrings this round was sent to correct are rewritten: the test no
|
||
longer claims every pixel is decoded, and `_bmp_rle8_rows` no longer
|
||
frames the delta argument as read off the corpus, which it never was.
|
||
- **An end-of-line escape at column 0 states no skip, and neither does a delta
|
||
out of its row (0.10.1).** The entry above says a delta escape and an
|
||
end-of-line escape both leave pixels every decoder agrees on. Measured by PM
|
||
and reproduced here: that is true of the delta and false of the end-of-line.
|
||
Four end-of-line escapes and an end-of-bitmap carried an 8x4 frame with 32 of
|
||
32 pixels never decoded, and Pillow refuses those same bytes.
|
||
- The class is wider than the one construction, and this round measured it
|
||
rather than patching it: over every opcode sequence of length 1 to 4 on a
|
||
4x3 frame — **22 620 streams**, swept in the suite — this package carried
|
||
**703** streams the independent decoder refuses and drew **1 492** more
|
||
differently. PM's recommendation on its own (refuse a stream that painted
|
||
nothing) leaves **512** and **1 171** of those, so it would have narrowed
|
||
the class for the third round running.
|
||
- `_bmp_rle8_rows` refuses an end-of-line escape at column 0 (it closes no
|
||
row, so the row it passes over is one the stream never wrote) and a delta
|
||
whose horizontal offset would leave the row (the format puts that offset
|
||
inside the line; this reader keeps the cursor past the row end and a flat
|
||
decoder rolls it into the next row). Both with `asset_samples_invalid`.
|
||
After: **0** carried-here-refused-there and **32** drawn differently.
|
||
- **What is not closed is stated.** All 32 residual streams are a run or
|
||
absolute block that OVERRUNS its row. Refusing those gives 0 and 0 — and
|
||
costs **15 of the 25** real RLE8 files, which would drop 15 real figures
|
||
and move a pinned bundle's bytes.
|
||
- **Cost measured on the corpus first:** over **11 441** files scanned across
|
||
the four raw standard deliveries and the K2 reference corpus, the only
|
||
**25** BMPs on this machine use an end-of-line at column 0 in **0 of 25**
|
||
and a delta in **0 of 25**, and **25 of 25** still decode to Pillow's
|
||
pixels exactly (**3 117 220** pixels compared) after the change.
|
||
- `CURSOR_CASES` goes 8 arms to 12: one for the cursor rule's ROW clause
|
||
(PM's `P8`, `height - 1` → `height - 2`, which survived 51 tests) and four
|
||
for the end-of-line class. `P8` and `P13` join the mutant runner.
|
||
- **The published `--accounting` contract names every key the gate reads
|
||
(0.10.1).** The JSON sketch in `tools/okf_accounting_gate.py` is what a
|
||
consumer implements the door from, and it did not name `conversions`, which
|
||
`asset_holds`' conversion route depends on, nor `normalised_soft_hyphen`,
|
||
`unaccounted` or `double_booked`, which the door had written for a round
|
||
longer. A door built from the contract writes a ledger the gate reads as
|
||
"nothing was converted", and every converted image comes out
|
||
claimed-and-not-found — 19 of 50 on R761. Two tests hold the sketch against
|
||
both sides: what the gate LOOKS UP (measured with a ledger that records its
|
||
own lookups, not by grep) and what the door SERIALISES.
|
||
- **A bundle built without the door now says why a converted image cannot be
|
||
proved (0.10.1).** Without `--accounting` there is no ledger, so `asset_holds`
|
||
falls back to its first route and a converted picture is counted
|
||
claimed-and-not-found. The fallback was honest and silent; the count was
|
||
printed and its cause was not. `_tally` names the missing ledger when, and
|
||
only when, something was claimed and not found, and `asset_holds` says it in
|
||
its own docstring. The gate always passes the flag, so no row moves.
|
||
- **The published `tbx:` count is one number, guarded without the delivery
|
||
(0.10.1).** `assert sum(tbx.values()) == 568` sat behind a `skipif` on a file
|
||
only one machine has, so on a fresh clone the sentence five files publish was
|
||
unguarded — the state in which 574 survived in four docstrings. `N101_TBX_TAGS`
|
||
is now the one place it lives and a second test holds all five published
|
||
sentences to it, with no corpus and no clock. What it does not prove is
|
||
stated: five files agreeing is agreement, not a count.
|
||
- **The mutant runner judges a mutant by the suite that owns it, and the
|
||
catalogue goes 39 to 45 (0.10.1).** It could only run one test file, which is
|
||
why PM's three survivors from `43331fc` could not be added. `X3`/`X4` were
|
||
rewritten against the code as it now stands; `X6` is the defeated state
|
||
exactly (a pointer block believed without the run having booked it); `X7`
|
||
cuts the ledger off at its source; `X8` removes the cursor rule; `P6`, `P11`
|
||
and `P12` are PM's three. Two survivors appeared on the first run and both
|
||
were findings — the asset binding had stopped being exercised, and the
|
||
`_inline` disarming survived the WHOLE suite (2134 passed) because the gate
|
||
no longer reads its claim from the bundle. The disarming is KEPT and now
|
||
measured in `tests/test_assets.py`: the property is about the bundle, not
|
||
about one judge. `killed 45 of 45`, exit 0.
|
||
- **`test_the_four_existing_goldens_are_untouched` skips, with its reason, in
|
||
a `git archive` extract (0.10.1).** It called `git status` with `check=True`
|
||
outside a repository and raised. It was the single failure of the whole
|
||
suite run from a clean extract, twice reported as a round's one failure by a
|
||
round that had not touched the file.
|
||
|
||
- **A truncated RLE8 BMP is refused instead of carried as a partly blank PNG
|
||
(0.10.1).** `_bmp_rle8_rows` painted what the stream held and left the rest
|
||
of the frame at palette index 0 — which is what the format says about a
|
||
SKIPPED pixel, so no decoder disagreed and the picture was wrong with no
|
||
code and no row. Measured by PM 2026-09-19 on a real R761 asset (352x548 =
|
||
192 896 pixels): cut to 90 % it was carried with 13 923 pixels wrong, to
|
||
50 % with 95 890, to 10 % with 166 525. The uncompressed path already
|
||
refused the same shape.
|
||
- The decode may now end at an explicit end-of-bitmap escape and nowhere
|
||
else; running out of bytes raises `asset_samples_invalid`, the code the
|
||
uncompressed path uses. NO PIXEL IS GUESSED: either every one is decoded
|
||
from the stream, or the picture is refused with a line in the concept.
|
||
- The rule is the terminator rather than `biSizeImage` (a claim by the same
|
||
untrusted header) or a coverage count (which would refuse the delta escape
|
||
the format defines), and it is read off the corpus: over the 19 real RLE8
|
||
assets of the frozen R761 delivery, **19 of 19** end at an explicit
|
||
end-of-bitmap, on **19 of 19** it is the stream's last two bytes, and on
|
||
**19 of 19** `biSizeImage` equals the available bytes. A whole stream that
|
||
omits the terminator is refused alongside a cut one.
|
||
- **Nothing real changes hands:** the same 19 files still convert losslessly
|
||
after the rule, **2 366 365 pixels** compared — this time with stdlib on
|
||
BOTH sides, an independent BMP reader and an independent `zlib` +
|
||
filter-reversal PNG decoder, with a one-byte control proving the
|
||
comparison can fail.
|
||
|
||
- **The content-accounting gate: a document refused whole is never clean.**
|
||
Its elements are all booked as coded rejections, so u = 0 and d = 0, and
|
||
`refused_whole` asks its question only for a corpus that persisted NOTHING —
|
||
one refused source beside an accepted one read as clean with the content
|
||
gone. `Unit.refused` is that loss with its own column and the document's code
|
||
in the note.
|
||
- **The STS JSON role map reads the publisher's own tags.** `count_sts_json`
|
||
compared the raw tag string where the XML witness has always used `_local`,
|
||
so `mml:math` reached `tag == "math"` on nothing — 74 formulas in N200
|
||
Vegbygging:2024 counted as 0. And the publisher's JSON writes a figure's
|
||
caption as `figcaption` under the `graphic`, not as the `fig/caption`
|
||
NISO-STS writes — 430 of them over the eight deliveries measured. No other
|
||
count moves, measured role by role over those eight and the committed twins.
|
||
- **The mutation harness is a gate.** A surviving mutant now exits 1; the run
|
||
ended `2 if errors else 0`, so `killed 0 of 1` exited 0. PM's X2 mutant — a
|
||
report may declare a document rejected while the bundle holds it — is in the
|
||
set and is killed by a new test driven from both sides.
|
||
- **The skipped-row guard measures the machine, not the argument.** Row 6 is
|
||
SKIPPED exactly when the corpora the arguments name are absent, so asking the
|
||
arguments made the branch unreachable.
|
||
- **Row 6 says when a corpus measures no element class at all.** On N200 the
|
||
build proposes 0 plans and exits 2 before the accounting door, so 16 549
|
||
elements land as unaccounted with no declared fate — a finding about the run
|
||
that read as a finding about the build.
|
||
|
||
- **An inline PDF image gets a stable name (0.10.1).** pdfminer names an
|
||
inline image (`BI … EI`) from `id()` of a Python object, so a pointer line
|
||
changed between two runs of one build and two concept files of the reference
|
||
corpus differed — breaking the bit-exact rebuild invariant. Such an image is
|
||
now named from its position on the page.
|
||
|
||
- **Three sentences this release publishes are now held by tests.** PM's
|
||
checkpoint on `43331fc` found three mutants surviving the entire suite: a
|
||
normalisation door that ALSO removes U+00A0 NBSP — which would have eaten
|
||
all 6 633 of them in R761 while `log.md` went on saying "No other character
|
||
is touched" — and row 3 of the accounting gate losing either its
|
||
`refused={n}` column or its "N element(s) lost with R of D document(s)
|
||
refused whole" clause. Each is killed now by a test that counts its own
|
||
numbers: ten characters the door must leave exactly where they were
|
||
(compared against a filter written in the test, so ORDER is pinned as well
|
||
as multiset), and R, D and the element total counted over the units the test
|
||
builds.
|
||
- **`log.md` says where the soft-hyphen count comes from.** The
|
||
`**Normalisation**` bullet now ends "The count is the door's own, read off
|
||
the run and not recounted from the source." Chosen over adding a second,
|
||
independent counter: the door acts on the EXTRACTED text, so a counter over
|
||
the source bytes would disagree by construction on every type extraction
|
||
does not carry verbatim, and the gate would have to decide which difference
|
||
was a loss.
|
||
- **Two published numbers were wrong and are corrected.** N101 ships **568**
|
||
`tbx:` tags, not 574 — three independent counts agree (raw substring, regex
|
||
over the JSON `tag` field, node traversal), and a test now counts them over
|
||
the delivery instead of repeating the number in a fourth docstring. And the
|
||
reach clause "0 across `tests/fixtures`, `examples`, …" was true only of
|
||
U+00AD: **2 of 230** readable tracked files carry U+200B, this repo's own
|
||
known-negative fixture, which the door is built not to touch.
|
||
|
||
### Documented, not changed
|
||
|
||
- **The accounting gate proves CARRIAGE, not FIDELITY.** Neither of
|
||
`asset_holds`'s routes decodes a pixel: a converter writing a blank PNG is
|
||
accepted, because the bundle is internally consistent. The suite fells that
|
||
mutant by decoding both sides; the judge cannot, and its docstring now says
|
||
so — "claiming a conversion it did not perform" means claiming one whose
|
||
FILE is missing, never one whose pixels are wrong.
|
||
|
||
- **The lossless guard now has an arm that runs on a core install.** It decoded
|
||
through Pillow, an OPTIONAL dependency here (transitive under `pdfplumber` in
|
||
`[extract]`), so 4 of the 13 guards in `tests/test_asset_viewable.py` were
|
||
SKIPPED on a plain `pip install llm-ingestion-okf` — the lossless one among
|
||
them. The new arm decodes the carried PNG with `zlib` and the five PNG SS 9.2
|
||
filters and compares against pixels written out in the test file, and it is
|
||
re-run under a `sys.meta_path` finder that makes every `PIL` import fail.
|
||
|
||
|
||
- `images: N` in a concept counts POINTER BLOCKS, not unique pictures (12
|
||
pointers to 2 files is `images: 12`). Now stated in the README.
|
||
- A concept that is only a pointer block is persisted as substantive, because
|
||
"degenerate" means zero characters after stripping whitespace and a pointer
|
||
block is text.
|
||
|
||
## [0.10.0] — 2026-09-17
|
||
|
||
### Added
|
||
|
||
- **A bundle carries the images its sources declare (0.10.0).** Until now no
|
||
reader in this package fetched, named, described or copied a single image:
|
||
`<img>`'s attributes were never read, a NISO-STS `<graphic>` was walked past,
|
||
a PDF was opened for its text alone, the converter's markdown writer dropped
|
||
every picture, and the only writer into a bundle took `content: str`. The two
|
||
lossiness warnings said so on every run, which made the loss honest and did
|
||
not make it smaller. Measured on R761 Prosesskoden:2025: the process text is
|
||
carried in full while 12 `Tabell N-N` and 9 `Figur N-N` captions stand over
|
||
nothing, so process 84's "toleranseklasse ... er gitt i tabell 84-2" points
|
||
at empty space.
|
||
**Five readers place, one module decides.** `pdf` (embedded image XObjects),
|
||
`docx`/`pptx`/`odt`/`rtf` (the converter's media, through `--extract-media`),
|
||
`html`/`htm` (`<img src alt>`, local paths and inline data URIs) and `xml`
|
||
(`<graphic xlink:href>`, resolved against the href and then against a sibling
|
||
`graphics/`). `llm_ingestion_okf.assets` decides what an image IS, what it is
|
||
called and how it is pointed at, so "carried N of M" means one thing across
|
||
all five. `.xlsx` is deliberately excluded: a two-line block inside its pipe
|
||
tables would break the row locator read back out of them.
|
||
**The bytes go to `assets/`** at the bundle root under
|
||
`<sha256[:12]>-<the source's own base name>`, and the concept carries a
|
||
two-line pointer where the picture stood -- a markdown image, then the
|
||
source's own file name and the size in pixels. A PDF stream that is already a
|
||
file (`DCTDecode`, `JPXDecode`) is carried VERBATIM; raw samples are encoded
|
||
to PNG with `zlib` from the stdlib, so no new dependency and no rasteriser
|
||
version enters an asset's bytes or its content-addressed name. What this
|
||
encoder cannot express exactly -- a stencil mask, a `Decode` array, CMYK,
|
||
anything but 8-bit samples -- is refused with a code and counted, never
|
||
approximated.
|
||
**ON by default, with `--no-assets` reproducing the pre-0.10.0 bytes.**
|
||
Measured over the 43-document reference corpus, two builds of one commit:
|
||
453 -> 454 concepts, 865 -> 867 markdown files, 0 -> 2 964 assets (2 964
|
||
carried of 3 145 found, 4 622 pointers), 4.7 MB -> 115 MB, 2 414 s ->
|
||
3 088 s, peak RSS 6.26 -> 8.74 GB, and 422 of 865 markdown files differ. The
|
||
one new concept has a measured cause: the pointers are body text, so a
|
||
section holding 146 of that document's images grew from 19.0 % to 30.6 % of
|
||
the extracted text and crossed `--outline-gate`'s 0.20 share clause.
|
||
**`log.md` states it either way** -- "N carried of M found", or `NOT CARRIED`
|
||
under `--no-assets`, so a bundle nobody looked for figures in cannot be
|
||
mistaken for a bundle of documents that had none. A concept on this
|
||
repository's own profiles also carries `images: N`, conditional, counted out
|
||
of the concept's own text.
|
||
**The image bytes are NOT screened**, and the log says so: the guard is
|
||
text-only, the pointer block passes the gate as body text, and the picture
|
||
beside it passes nothing.
|
||
**Door C carries them too.** Measured before the repair: importing a bundle
|
||
built with `--assets` merged 6 of 6 concepts and wrote no `assets/` at all,
|
||
so every pointer in the imported bundle named a missing file. Only the assets
|
||
a MERGED concept points at are carried -- an asset belonging to a refused
|
||
concept must not ride in on the back of a cleared one.
|
||
A proposed SPEC section 6.4 for the layout is in
|
||
`docs/plan/okf-assets-section-6-4.md`; `_okf-canonical` is not edited from
|
||
here.
|
||
|
||
### Fixed
|
||
|
||
- **A markdown image is no longer read as a cross-reference.**
|
||
`structure._LINK` reads `[...](target)` and never looked at the character in
|
||
front of the bracket, so an asset pointer would have arrived in the index as
|
||
a `references` edge to a concept that cannot exist -- and the digits in an
|
||
asset's file name would have been read as a document number. The link's span
|
||
is still masked, so the number scan cannot see it either.
|
||
|
||
- **`okf build` now runs a real guard, and the bundle says which one (F1).**
|
||
From the day the command was packaged until 2026-09-15, `corpus.measure`
|
||
wired an unconditional approve-everything stub into `process_inbox` and no
|
||
`add_argument` call anywhere in the package named a gate -- so the only path
|
||
most people use screened nothing, while `pyproject.toml` made the guard a
|
||
MANDATORY runtime dependency and the README recommended a composition the
|
||
command line could not reach. Reported from outside by `claude-code-llm-wiki`
|
||
and reproduced here before anything moved.
|
||
**`--gate` takes `guard-trusted-source` (the new default),
|
||
`guard-user-upload` or `none`**, and the name is written into the bundle's
|
||
section 9 `log.md` either way, so a consumer holding a bundle can tell a
|
||
screened one from an unscreened one. An unknown name is refused rather than
|
||
resolved to the stub: falling back would reproduce the defect with an extra
|
||
step. `okf project` owns no flag that moves a bundle's bytes and takes the
|
||
default; the corpus harness carries the same flag and the same default,
|
||
because a test holds the two paths byte-equal.
|
||
**The default was chosen on a measurement, not on caution.** Over the 453
|
||
concept bodies of the pinned reference bundle, `PRESET_TRUSTED_SOURCE`
|
||
returns the persist disposition on **453 of 453** while `PRESET_USER_UPLOAD`
|
||
holds **1**, taking one of the 39 source documents out. Neither tier waves
|
||
anything through: an invisible carrier and a CRITICAL finding fail secure at
|
||
both, measured against guard 1.4.0. Door B's own library default is
|
||
unchanged at `PRESET_USER_UPLOAD` -- an inbox drop is an untrusted upload,
|
||
an operator pointing `okf build` at their own folder is not. The second tier
|
||
ships as `guard_adapter.inbox_gate_trusted_source`, the three-line adapter
|
||
that module's own docstring describes, rather than as a preset parameter.
|
||
**The composition the README recommends is now tested.** Before this change
|
||
`grep -rl inbox_gate tests/` gave ONE file with 0 occurrences of `segment`,
|
||
while the nine files passing `segmentation=` all injected a local warn-stub:
|
||
no test ran a real guard verdict and a segmentation plan in the same call.
|
||
- **A fenced code block no longer declares structure (F2).** The proposer read
|
||
every line of the extracted text with the same grammars, so `# Use the
|
||
opus[1m] alias` inside a ```bash fence became a level-1 ATX heading. Two
|
||
effects, and the smaller one was the visible one: the document was REFUSED
|
||
entirely when the line carried `[` or `]` (Door B validates a title fail-fast
|
||
and never repairs one) -- 5 of 191 pages of the reporter's corpus -- and the
|
||
concept TITLE was silently taken from somebody's shell session everywhere
|
||
else, on **62 of 191 pages (32.5 %)**.
|
||
No rule in `find_candidates` reads a fenced line now: not ATX, not the
|
||
numbered grammar, not a table row, not a bold title, and not Arm D's outline
|
||
run, which selects from the whole line list and would otherwise let a fenced
|
||
install listing decide which run wins. Backtick and tilde fences, up to three
|
||
leading spaces, a closing fence at least as long as its opener, and
|
||
CommonMark's rule that a backtick fence's info string may not contain a
|
||
backtick -- that last one is what keeps a line holding only `` `okf build` ``
|
||
from silencing the rest of a document.
|
||
**It lands unconditionally rather than behind a flag, and the exposure is
|
||
measured on the bytes**: 0 of 865 concept files in the pinned default bundle
|
||
and 0 of the shipped fixtures and goldens that reach the proposer carry a
|
||
fence of either kind, so a rule that can only fire INSIDE a fence cannot have
|
||
moved anything this repository has measured. It is a defect, not a default
|
||
move.
|
||
|
||
### Added
|
||
|
||
- **`okf quality <bundle> --fasit <json>` -- boundary recall against the
|
||
structure the source itself declares (G37b).** The bundle-only gate returned
|
||
`UNMEASURED` and exit **3** on the very arm it was built for, because no
|
||
bundle-only metric reaches it; `boundary_share` -- declared boundaries that
|
||
became a concept, over declared boundaries -- is the one metric measured that
|
||
orders the arms correctly, and it needs the publisher's own structure, so it
|
||
arrives as an input rather than as a constant. The fasit is a JSON list whose
|
||
rows carry `title` and `norm`, validated at the door: a file that is not a
|
||
list, a row missing either key, or anything that is not JSON exits **2** with
|
||
the reason, never a quiet `UNMEASURED`.
|
||
**The normalisation was measured before the metric was built**: stripping all
|
||
whitespace and lowercasing reproduces the fasit's own `norm` from its own
|
||
`title` on **2 761 of 2 761** rows (keeping only alphanumerics scores 58).
|
||
**A boundary is recovered in either of two forms**, and both are load-bearing:
|
||
a concept whose normalised title equals `norm`, or the pair of the concept's
|
||
own directory and its residual title -- because the numbering token a
|
||
publisher glues into a heading lands in the concept id on one route and in the
|
||
title on another. Measured on the known-good arm, the literal form alone
|
||
reaches **22 of 2 761** where the two together reach **2 759**; on another
|
||
build of the same product the split is the exact opposite (2 727 literal, 0
|
||
paired). One bar, at the value measured on the pinned artifact: **2 759/2 761**,
|
||
`corpora = 1`. It separates the known-bad arm at **1 148 of 2 761 (41.6 %)**,
|
||
which is now a `FAIL` and exit 1 instead of exit 3. **`--fasit` is an
|
||
assertion**, like `okf consume --ref`: a bundle of another product scores 0 of
|
||
2 761 (measured on two of them) and reads `FAIL` -- the assertion being wrong,
|
||
not the bundle. The bar rests on **one product**, and the run says so on every
|
||
boundary row. `docs/2026-09-12-g37-terskler.md` SS 7 carries the premises
|
||
re-measured, the seven bundles, the interval any bar could sit in, and the two
|
||
R761 builds this one fails.
|
||
|
||
### Unchanged
|
||
|
||
- **Without `--fasit` the command is exactly what it was**, held by a test: no
|
||
boundary row, and `860019-mdb-100` still exits 3. No version bump, no tag,
|
||
`okf check` untouched.
|
||
|
||
## [0.9.0] — 2026-09-13
|
||
|
||
### Added
|
||
|
||
- **`okf quality <bundle>` -- a per-file-type verdict, with the denominator
|
||
(G37).** `okf check` is a CONTRACT check, and a green one is not a quality
|
||
gate: measured 2026-09-10 by `vegnormal-okf`, three arms over one corpus all
|
||
returned 0 findings and exit 0 while their hit@k ranged from 6 of 6 to 0 of 6.
|
||
The new command asks the other question. Three verdicts and no fourth --
|
||
`PASS`, `FAIL`, `UNMEASURED` -- and a type with no measured threshold is never
|
||
`PASS`; exit **0** judged and clean, **1** at least one `FAIL`, **2** did not
|
||
run, **3** nothing could be judged, because exit 0 over a table of unmeasured
|
||
rows would be the silent pass the command exists to stop. Two thresholds
|
||
today, both `structure_null_share` (the share of a type's documents that
|
||
yielded exactly one concept), read off the pinned 43-document reference
|
||
bundle: `.pdf` 8/32 and `.docx` 2/5, plus one definitional bar that applies to
|
||
every type (0 concepts with an empty body, measured 0 of 8 602 over four
|
||
bundles). A bar needs at least five documents on BOTH sides -- its own and the
|
||
judged bundle's -- so `.xlsx` (2), `.xml` (1) and every type with no corpus
|
||
class in `extract._EVIDENCE` are `UNMEASURED` and print their numbers without
|
||
a verdict. The gate reads the index tree, never a directory (SS 9.2), and
|
||
prints the bundle's own run log beside its counts, because a document rejected
|
||
at extraction leaves no row in the bundle at all -- the pinned corpus holds 33
|
||
PDFs and the bundle shows 32. `okf check` is untouched.
|
||
`docs/2026-09-12-g37-terskler.md` carries the table, the nine bundles behind
|
||
it, the order's own premises re-measured (three of five moved), and three
|
||
candidate metrics measured and NOT shipped -- two of them ordering a known-bad
|
||
arm and a known-good one the wrong way round. The README publishes the bars
|
||
behind a `<!-- quality-thresholds: ... -->` marker that
|
||
`tests/test_docs_promises.py` pins to the code and to the document.
|
||
|
||
### Changed
|
||
|
||
- **The README states every file type the extractor registry reads (K3-26).**
|
||
It read 13 extensions (`.md`, `.txt`, `.csv`, `.json`, `.html`, `.htm`,
|
||
`.xml`, `.pdf`, `.docx`, `.xlsx`, `.pptx`, `.odt`, `.rtf`) while the opening
|
||
line named five of them and the full list existed only in a hidden
|
||
`<!-- extract-formats: ... -->` comment, which no reader reads. A visible
|
||
`## Supported file types` table now carries one row per extension with its
|
||
reader, its dependency (core or the `[extract]` extra), the evidence class
|
||
the code records for it, and one honest note -- so a `constructed` row at
|
||
N = 1 cannot read as a supported one. A `Not read today` section states what
|
||
is absent (`.doc`, `.epub`, `.eml`/`.msg`, image files, source files,
|
||
`.one`/`.vsd`) as facts rather than as a queue. Four new assertions in
|
||
`tests/test_docs_promises.py` pin the table's row set, its evidence cells and
|
||
its core/extra split to the registry, and the opening to the table. The
|
||
six-row Format/Reader/Evidence table that `### Binary extraction` carried is
|
||
gone with them: it duplicated three evidence classes in prose no test read,
|
||
and that section now points at the pinned table. A fifth assertion holds it
|
||
gone -- the same query finds 8 table lines in that section before the change.
|
||
No change to `extract.py`: nothing about what is read moved, only what the
|
||
README says about it.
|
||
|
||
## [0.8.5] — 2026-09-12
|
||
|
||
### Added
|
||
|
||
- **`consume` can rank a body without the door's own link line (K3-23).** One
|
||
parameter, `link_in_signal`, default `True`, no CLI flag, no payload moved:
|
||
it is the instrument that separates a ranking movement from a budget
|
||
displacement on ONE bundle, and the measurement it was built for is
|
||
`docs/2026-09-12-k3-runde23-stien-i-kroppssignalet.md`. Measured on R761
|
||
(2 761 concepts, 710 heading-only sections, 675 carrying the line): of the
|
||
**39** newly delivered concepts the line ever added a question token to,
|
||
**39** gained it from the bundle-absolute PATH and **0** from the link's
|
||
title, and every token it ever contributed is a segment of the document's
|
||
own directory -- the saturation `shared_id_prefix` takes out of the id
|
||
signal, back in through the body. hit@1/8/50 stays **6/6 · 6/6 · 6/6** and
|
||
the known-positive stays at rank 1 under every reading; delivered sets move
|
||
on 2 of 8 questions at the default `k` and 3 of 8 at `k` 50; the heavier
|
||
excerpts displace **1** concept on **1** of 16 rows, and at the default `k`
|
||
the budget does not bind at all. The recognition is the door's own constant
|
||
and its own place -- last in the body, after a blank line, in the door's link
|
||
form -- so a human line opening with the same two words, and the door's exact
|
||
form anywhere but last, both keep every character they have.
|
||
`--shell-parent` therefore stays OFF at the current link form, and reading
|
||
the body without the line was recommended as `consume`'s default -- carried
|
||
out in the same unreleased block below: it is byte-identical on **5 of 5**
|
||
bundles anyone ships today, none of which carries the line.
|
||
|
||
### Changed
|
||
|
||
- **`consume` no longer scores the door's `Enclosing section:` line (K3-25).**
|
||
`link_in_signal` defaults to `False` (`consume.DEFAULT_LINK_IN_SIGNAL`) on
|
||
all three entry points -- `searchable_text`, `concept_scores` and
|
||
`build_payload` -- so the ranking and the stem vocabulary read a heading-only
|
||
body without the line the door appends under `--shell-parent`. The excerpt
|
||
still carries it and no CLI flag changed, so what moves is ORDER and never an
|
||
excerpt's bytes. Carrying out K3-23's recommendation with its numbers: of the
|
||
newcomers that line ever added a question token to, **39 of 39** gained it
|
||
from the bundle-absolute PATH and **0 of 39** from the link's title, every
|
||
such token being a segment of the document's own directory; under this
|
||
reading a flagged bundle delivers what the unflagged build delivers on **16
|
||
of 16** rows (list, order and `spent`), with hit@1/8/50 **6/6** at both `k`
|
||
and the known-positive at rank **1**. The cost to anyone shipping a bundle
|
||
today is zero bytes: **0 of 5** bundles carry the line, so **5 of 5** payloads
|
||
are byte-identical across the move. A bundle that does carry it -- only
|
||
`okf build --shell-parent` writes it, and that default is unchanged and still
|
||
OFF -- gets a different delivered order; `link_in_signal=True` is still
|
||
reachable for the older reading.
|
||
|
||
- **No bundle bytes move.** A five-document folder built before and after is
|
||
`diff -r`-identical (52 files), so no ranking measurement is owed. The
|
||
emission rule is untouched: this library still writes flow.
|
||
|
||
- **Two docstrings and one README paragraph corrected rather than left
|
||
standing.** `materialize._render_sources` gave three measured reasons for
|
||
refusing to emit the block form; reason 1 (a block list round-trips to an
|
||
empty value) and reason 3 (B6's acceptance test cannot pass) FELL with this
|
||
fix and are struck. Reason 2 STANDS and now carries the rule alone,
|
||
re-measured by reading `portfolio-optimiser` at `6eb58e5`: `read_provenance`
|
||
returns `UnreadableProvenance(reason="block-sequence")`. It is not the
|
||
guard's objection -- guard 1.4.0 reads the block form on 4 609 of 4 609
|
||
files. The README said this library "cannot read the block form" where one
|
||
reader could and one could not; it now separates the two by KEY.
|
||
|
||
- **`okf.parse_frontmatter` returns a flow string for a block `sources:`
|
||
where it returned an EMPTY string (public API).** The fix above changes what
|
||
an outside caller reads: a consumer who read the empty value and concluded
|
||
the address was absent now gets the address, while a consumer who passed the
|
||
return value straight to a YAML reader gets a parse error where they
|
||
previously got something empty that parsed -- a regression for them, and the
|
||
reason it is stated here rather than left inside the fix. Measured
|
||
2026-09-12 over the same four bundles, denominator = concept files carrying
|
||
a block `sources:` (2 756 + 446 + 1 133 + 270): PyYAML 6.0.3 reads the
|
||
returned string back on **0 of 4 605** of them, because the `?` opening a
|
||
query string in the viewer URL ends the flow scalar. The string is a READING
|
||
projection of a value this library does not write in that form; the emitter
|
||
`materialize._render_sources` still writes flow, so no bundle bytes move.
|
||
|
||
### Fixed
|
||
|
||
- **A block `sources:` sequence no longer loses the address in the flat
|
||
readers (K3-24).** `consume.read_sources` has always read both YAML forms;
|
||
the three copies of this library's line-oriented frontmatter grammar read
|
||
only the flow one and returned the key with an EMPTY value for the block
|
||
form -- not a `KeyError` a consumer can catch, an address that disappears.
|
||
Measured 2026-09-12 over four bundles a producer ships, denominator = files
|
||
carrying a frontmatter block: 2 756 of 2 757, 446 of 447, 1 133 of 1 134 and
|
||
270 of 271 concept files lost it, while PyYAML 6.0.3 and the pinned guard
|
||
1.4.0 both read it on 100 % of the same files. After: **0 of each**, and all
|
||
three flat readers return what BOTH reference readers return on 4 609 of
|
||
4 609 files, plus 2 762 of 2 762 in a flow-form bundle that is unchanged.
|
||
`materialize.parse_frontmatter` is public API, so the external consumer is
|
||
the one this repairs.
|
||
|
||
The value type was the choice and it was measured: widening the return type
|
||
from `dict[str, str]` costs 15 `mypy --strict` errors across four modules
|
||
plus a signature every outside caller follows; rendering the entries back
|
||
into the flow form those readers already round-trip costs 0. The rendering
|
||
is a READING projection, not a claim that the value is writable.
|
||
|
||
Narrow on purpose: `profiles.STRUCTURED_BLOCK_KEYS` is `{"sources"}`, the
|
||
key `read_sources` already knows how to read, and one grammar now serves all
|
||
four call sites (`profiles.read_block_mappings`). A shipped fixture carrying
|
||
a block `verified:` still reads as an empty value, pinned by a test so the
|
||
next widening is a decision rather than a side effect. The K3-20 guarantee
|
||
is unmoved and asserted per reader copy: a nested key never enters the
|
||
document's namespace.
|
||
|
||
## [0.8.4] — 2026-09-11
|
||
|
||
### Added
|
||
|
||
- **`parent` reaches the reader (K3-21 A).** `okf consume` resolves a
|
||
concept's `parent:` pointer -- a `segment_id`, unique only inside one
|
||
document's plan -- among the concepts sharing its `source_file`, and an
|
||
excerpt carries `parent: { concept_id, title }`, conditional like
|
||
`req_number`; a pointer that lands nowhere is named `parent_unresolved:
|
||
true`. A heading-only body whose plan entry has a parent gains ONE line,
|
||
`Enclosing section: [<title>](/<bundle-relative path>)` (SPEC § 5.1, § 6.1).
|
||
Measured on the one standard with such sections: 675 of 710 carry exactly
|
||
one link, 0 broken, 72 265 B = 4.49 % of body bytes; hit@1/8/50 6/6 at both
|
||
`k` with the known-positive at rank 1.
|
||
- **`okf consume --follow-parent` (K3-21 B, off).** `parent` also carries the
|
||
enclosing concept's `text` with that concept's own `sha256`, placed after the
|
||
cut from the room it left, in rank order -- the delivered set is the same
|
||
with it as without it (16 of 16 payloads); a text that does not fit is cut
|
||
and marked `truncated`.
|
||
|
||
### Changed
|
||
|
||
- **`okf check` has seventeen rules** (`parent_unfollowable`): a `parent` that
|
||
is not a `concept_id` and `title`, names its own excerpt, or names a concept
|
||
in neither `excerpts` nor `withheld`. Every "16 rules" line a consumer quotes
|
||
is now "17 rules". Contract § 8 gains point 6, and its sentence "additional
|
||
members are permitted and are not read by the checker" now says the checker
|
||
reads only the members § 8 names. A consumer that does not know `parent`
|
||
has nothing to do: it is conditional and absent on every bundle without it.
|
||
- **The § 7.4 known-positive moved** (13 238 / 12 893 / 345 -> 14 721 /
|
||
14 346 / 375), because it IS the contract document: every payload's
|
||
`budget.known_positive` block moves with no bundle changing. Measured on 32
|
||
regression payloads, everything outside that block is byte-identical.
|
||
- **`--shell-parent` stays off, on a measurement** rather than on "`okf
|
||
consume` reads no `parent` key", which A made false: the link's absolute
|
||
path repeats the document's directory in 675 bodies and moved the delivered
|
||
set on 2 of 8 questions at the default `k` (3 of 8 at `k` 50), hit@k
|
||
unchanged.
|
||
- **The guard pin moves from `v1.3.0` to `v1.4.0`** (`[tool.uv.sources]`, and
|
||
the pip fallback in the README). Installing `@v0.8.4` gives a different
|
||
guard than installing `@v0.8.3`: 1.4.0 parses a flow sequence of plain
|
||
scalars -- `source_offset: [1, 24]`, `derived: [number]` -- where 1.3.0
|
||
raised "a flow sequence admits flow mappings only". Measured with both
|
||
guards' `parse_frontmatter` on one bundle built from the five-document
|
||
project: 1.3.0 refuses 26 of its 28 frontmatter blocks, 1.4.0 refuses 0 of
|
||
28. The dependency range `llm-ingestion-guard>=1.2,<2.0` is unchanged.
|
||
|
||
### Fixed
|
||
|
||
- **The index resolves a `parent` naming a segment of its own document
|
||
(K3-21 C).** 675 of 675 such facets rendered `parent: pN?` while the concept
|
||
stood in the bundle; now 0. The two segmented goldens' index files move one
|
||
`?` each (4 lines).
|
||
- **A declared section below markdown's sixth level keeps its level in the
|
||
plan (K3-21 D).** The NISO-STS reader clipped the outline mark to 6 along
|
||
with the heading; the mark now carries the declared depth. On one standard
|
||
the plan moves on exactly 2 entries (`--shell-parent` only), and 710 of 710
|
||
shells point at the ancestor the `<sec>` nesting names (708 before).
|
||
- **Frontmatter this library writes is YAML a YAML reader reads back the same
|
||
(K3-22).** SPEC § 11 point 1 requires "a parseable YAML frontmatter block" in
|
||
every file. Measured with PyYAML 6.0.3 before the change, the pinned K2
|
||
default bundle failed `safe_load` on 41 of 455 blocks (and a 42nd parsed to
|
||
a truncated title), and each R761 build on 1 -- every one a block scalar
|
||
written verbatim: a title with `": "` or `" #"`, a leading `- `, `*` or `**`,
|
||
a trailing `:`.
|
||
- **Block scalars:** a value K3-19's plain-scalar rule refuses is written
|
||
double-quoted, `\` and `"` escaped; every other value keeps its bytes.
|
||
Rebuilt, the five-document project moves 0 files, each R761 build 1 line
|
||
and the K2 default bundle 42 `title` lines, after which all 454 of its
|
||
frontmatters parse and read back the same.
|
||
- **Flow leaves (`sources`, Door A's list, a run-stated flow value):** the
|
||
pinned guard refuses any quote inside a flow mapping, so a leaf PyYAML
|
||
would refuse (`?`, `,[]{}`, `": "`, `" #"`, a quote, a leading indicator)
|
||
has no form both read and is refused with the door's existing code
|
||
(`inbox_source_file_unaddressable`, `inbox_source_title_unaddressable`,
|
||
`source_reference_unquotable`, `run_frontmatter_invalid`).
|
||
- **Behaviour change:** `okf build --frontmatter 'sources=[{ resource: <URL
|
||
with a query string>, … }]'` now exits 2 and writes nothing. K3-19's own
|
||
flagged R761 build used such a URL and wrote 2 761 of 2 761 frontmatters
|
||
PyYAML refuses.
|
||
- **Readers** (`parse_frontmatter`, the index and structure readers, both
|
||
`read_sources` branches) unquote a `"`-wrapped value; `'`-wrapped values
|
||
are untouched. On 25 273 files of existing bundles the readers return
|
||
exactly what they returned before.
|
||
- The generated `SKILL.md` header goes through the same block rule.
|
||
- PyYAML joins the `dev` dependency group only; `src/` imports no yaml.
|
||
- Report: `docs/2026-09-11-k3-runde22-yaml-lesbar-frontmatter.md`.
|
||
|
||
## [0.8.3] — 2026-09-11
|
||
|
||
### Added
|
||
|
||
- **`okf build --shell-parent` (K3-20), off by default.** A concept whose body
|
||
is its heading alone gets `parent:` naming the `segment_id` of the nearest
|
||
ancestor that holds text: the nearest preceding plan entry at a smaller
|
||
level, passing over an ancestor that is empty too. Nothing is copied and no
|
||
boundary moves. It reads the plan's level and order, never the row. Measured
|
||
on one process code, 710 of 2 761 concepts are heading-only; the route names
|
||
the ancestor the document's own nesting names on 708 of them (two sit at
|
||
depth 7, which a markdown heading clips to 6), where reading section numbers
|
||
gets 686.
|
||
- **Off, by measurement:** `okf consume` reads no `parent` key, so no payload
|
||
ranks differently, while the flag moves the bytes of every bundle holding
|
||
a heading-only section.
|
||
- **Known cost:** the index projects `parent` as a document NUMBER, so a
|
||
segment id always renders unresolved there (`parent: p1?`). The same key
|
||
already carries both meanings for an adjudicator's declared parent.
|
||
|
||
### Fixed
|
||
|
||
- **A directory every concept id shares no longer ranks the concepts
|
||
(K3-20).** `okf consume`'s first fusion signal read a concept's title
|
||
together with every segment of its id. On a one-document bundle every id
|
||
starts with the same directory, and since K3-19 an STS document names that
|
||
directory after its own number, so a question naming the document matched
|
||
every concept -- except the one whose title already named it, which gained
|
||
nothing because the overlap counts a question token once. Measured on a
|
||
2 761-concept bundle, the known-positive fell from rank 1 to not delivered
|
||
at the default `k` (13 at `k` = 50). `consume.shared_id_prefix` now keeps
|
||
the leading directories EVERY id shares out of that signal: the
|
||
known-positive is rank 1 at both `k` and S1-S6 stay 6/6.
|
||
- **Consumer cost: payloads move only on a bundle whose ids all share a
|
||
leading directory**, which is what a one-document build produces. Where
|
||
they share none, the signal reads the same string as before, and the
|
||
measured multi-document bundles are byte-identical. The old order is
|
||
reproducible by no flag.
|
||
- **Measured and felled:** dropping each concept's own document directory
|
||
instead took a hit@8 row on the pinned 43-document bundle from rank 5 to
|
||
not delivered.
|
||
|
||
### Changed
|
||
|
||
- **A NISO-STS document's own identity names its directory and titles its
|
||
`sources` entry (K3-19).** `okf build` put every concept of an STS delivery
|
||
under a directory named for the delivery file -- measured, a UUID occurring
|
||
0 times in the document -- while the document's one `<std-ident>` carried a
|
||
`<doc-number>`. `extract.declared_identity` reads exactly one `<std-ident>`
|
||
(`<doc-number>`, `<year>`) and exactly one `<title-wrap>`; a value stated
|
||
more than once is not read. The directory is the `<doc-number>` through the
|
||
id grammar, replacing only the file's stem, and `sources[0].title` is
|
||
`<doc-number>` + `<year>`, then the `<title-wrap>` title, then the file name
|
||
-- the first that can be written into the flow mapping verbatim (the
|
||
measured `<full>` carries commas, so it never is). A declared name two
|
||
documents in one run claim is used by neither, and stderr says so.
|
||
- **Consumer cost: a re-run, and an STS document's concept ids move**
|
||
(`<uuid>/...` -> `<doc-number>/...`). Every other file type is untouched:
|
||
the five-document folder rebuilds identical except the `log.md` line that
|
||
records the venv's converter path.
|
||
- **Measured side effect on ranking, reported rather than repaired:** hit@1
|
||
/ 8 / 50 over S1-S6 stays 6/6 at both `k`, but the known-positive falls
|
||
from rank 1 to not delivered at the default `k` (13 at `k` = 50).
|
||
`consume`'s first signal reads a concept id's segments, and on a
|
||
one-document bundle every id now carries the document's own name; renaming
|
||
only the directory back restores rank 1. `--rarity-weight` delivers it at
|
||
rank 4 with S1-S6 unmoved, and stays off.
|
||
|
||
### Added
|
||
|
||
- **`okf build --frontmatter KEY=VALUE`**, repeatable, stamps a key on every
|
||
concept of a run -- for what an operator knows and a document does not say,
|
||
such as an edition or a publisher's address. SPEC SS 4.1 "Extensions" lets
|
||
a producer add any key; SS 11 forbids a consumer to reject one. Split on the
|
||
FIRST `=`, and the value is written verbatim on ONE line, because this
|
||
package's readers are line-oriented and blind to a block-form `sources`.
|
||
Precedence: a stated value beats what the document declares, which beats the
|
||
file name. A run may add any key and REPLACE only `sources` and
|
||
`description`; every key the door writes itself -- including Door A's
|
||
`ingest_manifest`, which would make that door claim the file -- is refused
|
||
before a proposal is written (`run_frontmatter_invalid`), as is a value that
|
||
would not read back as stated. Reachable as `build(frontmatter=...)` and as a
|
||
keyword-only `concept_frontmatter_values` on `measure`, `process_inbox` and
|
||
`render_inbox_concept`. Without the flag nothing moves. **Note:** a
|
||
`sources` value carrying a URL or `X:Y` in the flow mapping passes `okf
|
||
check` and is refused by PyYAML's `safe_load` -- measured on 2 761 of 2 761
|
||
concepts with such a value -- and it is written verbatim as stated.
|
||
- **`description` for an STS section, from its own first spec point.** The
|
||
first `<p>` of the first direct-child `<sec sec-type="spec">`, whole,
|
||
carried by the plan entry, screened by the gate, and written only where a
|
||
YAML reader reads it verbatim (`inbox._yaml_plain`; over 2 024 measured
|
||
values the rule and PyYAML agree on every one). The spec sets no length, so
|
||
the one-paragraph limit is ours. On the measured document: 2 026 of 2 761
|
||
titled sections carry a spec point, 1 807 descriptions are written (2 points
|
||
have no `<p>`, 217 carry `": "`), none is invented, and none is derived from
|
||
a title.
|
||
|
||
`--ingested-at` alone was confirmed to stamp every concept, on the segmented
|
||
route too, and to date `log.md`: 2 761 of 2 761. Report:
|
||
[`docs/2026-09-11-k3-runde19-dokumentidentitet-og-frontmatter.md`](docs/2026-09-11-k3-runde19-dokumentidentitet-og-frontmatter.md).
|
||
|
||
## [0.8.2] — 2026-09-11
|
||
|
||
### Added
|
||
|
||
- **`okf check` refuses a skill and a payload that name different bundles
|
||
(`bundle_mismatch`).** The checker had published this hole about itself since
|
||
2026-09-08 and not closed it: reproduced on this repository's HEAD, it
|
||
reported `conformant: 15 rules over 8 excerpts and 438 withheld entries, 0
|
||
findings` for a skill generated from one corpus against a payload assembled
|
||
from another; the same line for the UNFILLED template against that payload;
|
||
and the same line again for a payload sharing the skill's `bundle_id` at a
|
||
foreign `ref`. All three now exit **1** with one finding. The right pair is
|
||
untouched at exit 0 with 0 findings, and a payload declaring no identity at
|
||
all stays `ref_missing`'s defect at 9 findings -- no rule restates another.
|
||
- **BOTH halves are compared, and the `ref` half is the load-bearing one.**
|
||
Three distinct builds on one machine were measured carrying the same
|
||
`bundle_id`, so an id comparison alone would pass a stale skill -- the case
|
||
the generated skill warns about in its own words ("if the bundle moves, the
|
||
ref moves with it and this file is stale"). SS 3.3: "a version is the
|
||
producer's assertion; a ref is a fact about bytes".
|
||
- **An identity the rule cannot read is a finding, never a silent pass.**
|
||
That is what refuses the unfilled template, whose `<CORPUS>` and `<REF>`
|
||
are not an identity. It also refuses the repository's own hand-made
|
||
`skills/okf-consume/SKILL.md`, which predates `okf skill` and declares no
|
||
bundle identity a reader can act on -- **1 of 1** shipped hand-made skill,
|
||
a real find and not a fixture.
|
||
- **The rule compares a DECLARED identity against a DECLARED identity and
|
||
never opens the bundle**, so a payload misreporting its own `ref` still
|
||
passes. Proving a ref against bytes is `okf consume --ref`'s job and needs
|
||
a bundle path this command deliberately does not take.
|
||
|
||
### Changed
|
||
|
||
- **The checker's rule count is 16, not 15**, and `Report.rules_evaluated` is
|
||
the denominator every report line quotes -- so `15 rules` becomes `16 rules`
|
||
in every published line. A consumer citing the old number is citing a number
|
||
that has changed. No payload bytes move: this is the checker, not the
|
||
pre-pass.
|
||
- `skill.identity_line` is now the single authored copy of the sentence a
|
||
generated skill declares its bundle in, read back by
|
||
`contract_check.skill_identity` and held to it by a test. Generated skill
|
||
bytes are unchanged -- measured, both tracked bundles byte-identical before
|
||
and after on the same interpreter.
|
||
|
||
### Fixed
|
||
|
||
- **The shipped `skills/okf-consume/` is generated, and passes the check it
|
||
tells its reader to run.** The hand-filled copy predated `okf skill`,
|
||
declared no bundle identity, and was refused against the payload shipped
|
||
beside it: `NOT conformant: 16 rules over 3 excerpts and 0 withheld entries,
|
||
1 findings` (`bundle_mismatch`, exit 1). It is now `okf skill`'s output for
|
||
`examples/ingest-golden-segmented-okf-v0-2/expected-bundle` -- the bundle its
|
||
payload always came from -- and the pair is `conformant: 16 rules over 3
|
||
excerpts and 0 withheld entries, 0 findings`, exit 0. The payload's bytes do
|
||
not move. Regenerate with the command in
|
||
`skills/okf-consume/references/README.md`; two tests hold the pair and the
|
||
generator's bytes.
|
||
- **Its frontmatter `name` changed** from `okf-consume` to
|
||
`b-golden-segmented-okf-v0-2-consume`. Claude Code takes a project or
|
||
personal skill's command from its DIRECTORY, which stays `okf-consume`, so
|
||
a copy at `.claude/skills/okf-consume/` is still `/okf-consume`; only the
|
||
display label moves. Nothing in this repository named the skill
|
||
`okf-consume`.
|
||
- **Its prose no longer states K2 numbers.** The 629-concept figures belonged
|
||
to a corpus that cannot ship; the generated numbers describe the
|
||
three-concept golden bundle and nothing larger.
|
||
- **`uv.lock` records this package at 0.8.1.** The 0.8.1 version bump never
|
||
reached the lockfile, which still said 0.7.0, so `uv lock --check` exited 1
|
||
on a clean checkout and any non-frozen `uv` command rewrote the file. One
|
||
line; nothing else in the lock moved.
|
||
- **`--title-covered` no longer lifts a short title over a title that answers
|
||
more of the question.** 0.8.1's partition read every concept whose WHOLE
|
||
title the question accounts for before everything the fusion ranked above
|
||
it. That is a claim about the covered title's PRECISION, and it overrode the
|
||
fusion even against a title answering MORE of the question: measured on a
|
||
26-concept bundle of five tender documents, a question naming a section by
|
||
three of its title tokens also held a neighbour's whole one-token title, and
|
||
the neighbour took rank 1 from the section the question names. A covered
|
||
concept now RISES through the fusion's order and stops beneath the first
|
||
concept whose title answers more question tokens, by equality, than it
|
||
holds. Same flag, no new parameter, no new constant.
|
||
- **What the rule was built for does not move.** On the 2 761-concept bundle
|
||
of one standard no covered concept had such a title above it, so all 8
|
||
payloads are byte-identical to 0.8.1's at default `k` AND at `--k 50`;
|
||
hit@1/8/50 stays 6/6 · 6/6 · 6/6 with the known-positive at rank 1.
|
||
- **Where the rule never fires nothing moves either**, measured on the bytes:
|
||
the pinned K2 bundle 6 of 6 payloads identical, Arm B 6 of 6, the three N
|
||
bundles 15 of 15.
|
||
- **Four other repairs were measured and not taken**: a minimum title length
|
||
sold hit@1 back to 3 of 6; a share of the question holds only in a band
|
||
set by the question's word count (1/9 < s <= 1/6); an order inside the
|
||
covered group cannot act on a group of one; a closed stop list touches no
|
||
title involved and would be a new vocabulary to maintain.
|
||
- **0.8.1's unbounded order is reproducible by no flag.** It differs from
|
||
this one only where a covered concept has such a title above it -- 1 of
|
||
the 4 questions measured on that bundle, 0 of 8 on the standard, 0 of 27
|
||
elsewhere. `--no-title-covered` still reproduces the pre-0.8.1 order.
|
||
- **One constructed variant still reads the short title first, and no rule
|
||
reading titles alone separates it.** A shortened form of the same question
|
||
shares ONE token with each of the two titles; a form that blocks on any
|
||
token the covered title lacks fixes it and takes one of the standard's
|
||
scored questions, the same shape with the opposite answer, from rank 1 to
|
||
3. Report: `docs/2026-09-11-k3-runde17-dekningen-stopper-ved-en-bredere-tittel.md`.
|
||
|
||
## [0.8.1] — 2026-09-10
|
||
|
||
### Added
|
||
|
||
- **`--title-covered` (ON by default, opt out with `--no-title-covered`): a
|
||
question that accounts for a concept's WHOLE title reads that concept first.**
|
||
On the 2 761-concept bundle of one standard, the answering section was
|
||
delivered at rank 1 on **3 of 6** scored questions and **none of the reading
|
||
side's six flags moved that number** -- the whole sweep sits at 3/6 or worse.
|
||
Measured on that bundle, before and after: hit@1/8/50 **3/6 - 5/6 - 5/6 ->
|
||
6/6 - 6/6 - 6/6** at default `k` and **3/6 - 5/6 - 6/6 -> 6/6 - 6/6 - 6/6** at
|
||
`--k 50`, with the known-positive holding rank 1 at both and the
|
||
known-negative still not a hit. S1 4 -> 1, S5 not delivered -> 1, S6 3 -> 1.
|
||
- **THE DEFECT IS THAT BOTH LEXICAL SIGNALS ARE UNNORMALISED COVERAGE
|
||
COUNTS.** They measure how much of the QUESTION a candidate answers and
|
||
nothing measures how much of the CANDIDATE the question accounts for, so a
|
||
section titled with the question's subject alone scores what a narrower
|
||
section titled with that subject PLUS a qualifier scores, and then loses on
|
||
the body count. Decomposed per miss: S1 turns on `hvordan`, an interrogative
|
||
pronoun; S5 on `hvilke` and `stilles` in a body 7x the gold's, and on
|
||
`betonghvelv ~ betongkonstruksjoner` through the four-character stem
|
||
`betong`; S6 on an exact tie broken by `concept_id`.
|
||
- **A PARTITION, never a fourth RRF signal, and the arithmetic is why.** RRF
|
||
consumes ranks alone, so with shared ranks a rule whose positive group has
|
||
`m` members is worth `1/61 - 1/(61 + m)` -- a rule firing on ONE concept of
|
||
2 761 is worth 0.00026 against a body gap of 0.0029. **A precise rule is
|
||
worth LEAST under this fusion.** Measured as a signal it moves hit@1 not at
|
||
all (3/6, both as a third and as a fourth signal); as a partition it reaches
|
||
6/6. `lookup_hits` is the same shape for the same measured reason, and it
|
||
still wins: the new partition lands below it, with a test and its control.
|
||
- **By EQUALITY, never by shared prefix.** Four shared leading characters take
|
||
the group from 1 to 6 on one question and 9 to 31 on another, with the
|
||
answering section falling to candidate rank 6 and the known-positive to 2.
|
||
- **TWO CANDIDATE REPAIRS WERE MEASURED AND FELLED FIRST.** Pivoted length
|
||
normalisation of the body signal collapses at every value swept
|
||
(b = 0.25/0.5/0.75/1.0 -> hit@8 3/6, 1/6, 1/6, 0/6, and at b = 1.0 the
|
||
known-positive falls to rank 49): the median concept holds 22 tokens against
|
||
a mean of 60, so length normalisation promotes thousands of tiny concepts.
|
||
Title PRECISION as a signal reaches candidate hit@1 5/6 and takes the
|
||
known-positive from 1 to 4 every time it does.
|
||
- **NOTHING ELSE MOVES AND IT IS MEASURED ON THE BYTES.** The pinned K2 bundle
|
||
keeps `(1,1,1,1,1,5)` and its 7 pin tests, Arm B keeps `(1,1,1,1,1,5)`, and
|
||
the payloads on both are **byte-identical on 6 of 6 questions**; n100/n200/
|
||
n500 payloads are byte-identical on 5 questions each; the 828-file HTML
|
||
corpus still gives 828 plans, 0 unreadable and 6 015 md with `diff -rq`
|
||
empty; the five-document folder is `diff -r`-identical at 26 concepts / 52
|
||
md; `okf project` stays byte-equal to `okf build`. hit@k on n100/n200/n500
|
||
is **NOT MEASURED** -- this repository holds no gold set for them, which is
|
||
0 gold sets and not 0 hits.
|
||
- **THIS IS THE FOURTH READING-SIDE CHANGE THAT MOVES A PAYLOAD WITH NO BUNDLE
|
||
CHANGING.** A consumer pinned to the previous excerpt order needs
|
||
`--no-title-covered`. The rule fires on **0 of 21** measured cells outside
|
||
that one bundle, so "no regression" there means it never fires -- not that
|
||
it fires harmlessly. Report:
|
||
`docs/2026-09-10-k3-runde16-hele-tittelen-tar-ruten.md`.
|
||
- **A MEASURED DOWNSIDE, written as a known limitation and not as a fixed
|
||
defect: a SHORT, GENERIC title is covered in full by more questions than a
|
||
long one is.** On a five-document folder (26 concepts) a constructed
|
||
known-negative question demoted the answering section from delivered rank 1
|
||
to rank 2: a neighbouring concept titled with a single common process word
|
||
has its WHOLE title accounted for by that question, while the answering
|
||
section's longer title does not. The other seven delivered places did not
|
||
move. The repair is a later round's; nothing here narrows the rule, and the
|
||
opt-out is `--no-title-covered`.
|
||
- `build_payload`'s signature defaults are now held equal to the consume CLI's
|
||
argparse defaults by a test, for every same-named parameter. This is O6's
|
||
defect in the other command: `cli.build` defaulted two flags `False` in the
|
||
signature and `True` in argparse, and a caller reaching it as a function read
|
||
the signature.
|
||
|
||
## [0.8.0] — 2026-09-10
|
||
|
||
### Added
|
||
|
||
- **A section the SOURCE DECLARES now takes the route declared structure takes,
|
||
at the shipped defaults.** `.xml` gained a reader in the entry above and the
|
||
reader reached its ceiling -- **2 761 of 2 761** heading lines -- while the
|
||
build delivered **23 concepts and 15 of 2 761 boundaries**. Everything after
|
||
the reader ate it, and both steps are measured: the **orphan check** removed
|
||
**710 of 2 761** (710 of 710 removed headings are followed immediately by
|
||
another heading and **0 of 2 051** delivered ones are -- they are container
|
||
sections), and **Arm F** folded **2 066** more, 2 089 -> 23.
|
||
- `extract.xml_outline` reports the marks the reader wrote itself. There is
|
||
**no bridge** and therefore no tolerance constant and no `unresolved`
|
||
bucket: the reader appended the line it names. That is the difference from
|
||
`pdf_outline`, whose naive nearest-line rule was wrong on 1 840 of 2 762.
|
||
- `propose.RULE_XML_SECTION` (`rule:xml-section`) is its own name in
|
||
`RULE_NAMES` and in `_ORPHAN_EXEMPT`, so an artifact still distinguishes an
|
||
element the reader transcribed from a bridged bookmark
|
||
(`rule:pdf-outline`) and from a heading somebody guessed (`rule:heading`).
|
||
- The route is chosen by the ROW (`DECLARED_STRUCTURE_IDS = {"xml"}`), never
|
||
by the text: the same markdown arriving from a `.md` file is still a guess.
|
||
**No other file type changes one byte** -- `diff -r` on the five-document
|
||
reference folder is empty (52 md, 26 concepts, 0 of 5 rejected, 0 `.xml`
|
||
files in it), `okf project` is still byte-equal to `okf build`, the pinned
|
||
K2 bundle is unchanged, and the PDF arm still proposes 2 762 segments.
|
||
- Measured at SHIPPED DEFAULTS, not behind a flag: **2 761 concepts**;
|
||
**2 761 of 2 761** declared sections became a concept with the source's own
|
||
directory and title; **0** concepts match no declaration; `a)`-points
|
||
**0 of 4 954**; table blocks **10 of 10**; hit@1/8/50 **3/6 · 5/6 · 6/6**
|
||
(from 0/6 · 0/6 · 0/6) with the known-positive at rank 1. Cross-arm,
|
||
**2 761 shared concept ids** -- 100 % of the XML bundle and 2 761 of 2 762
|
||
of the PDF arm's, up from round 13's 2 022.
|
||
- Two directories of 2 738 still hold two concepts (`11`, `12`): the
|
||
publisher reuses a section number for two distinct sections, and it is the
|
||
same 2 the PDF arm has. Round 13's 14 such directories were false positives
|
||
of the text route reading the document's own contents listing, and they are
|
||
gone.
|
||
|
||
- **`.xml` is a core file type, NISO-STS aware, with a generic fallback.** A
|
||
publisher's own viewer delivers a zip that holds 0 html, 1 xml and 109
|
||
images; `okf build` on it was **110 of 110 unreadable, 0 plans, exit 2**, and
|
||
the conservation identity `merged + coded rejections == N` was never written
|
||
because the run aborted earlier. The one xml file is the whole product: 7 715
|
||
`<sec>`, **2 761 with a `<title>`**, 4 954 lettered points, 10
|
||
`<table-wrap>`, and a `<sec>`-nesting depth distribution row-for-row
|
||
identical to the publisher's own structure fasit.
|
||
- The output grammar is MARKDOWN, the same the office and HTML rows reach the
|
||
proposer through: `propose.py` is untouched. `<label>` + `<title>` become
|
||
one ATX line at the section's own depth; a `<sec>` with only a `<label>` is
|
||
a body line and never a heading (**0 of 4 954** became concepts);
|
||
`<table-wrap>` becomes one markdown table (**10 of 10**, against 0 of 10 on
|
||
the PDF path).
|
||
- The reader emits **2 761 of 2 761** heading lines and preserves text
|
||
exactly -- 1 283 395 of 1 283 395 non-whitespace characters, ratio
|
||
**1.000000**. The BUILD reaches 2 065 of 2 761 with `--no-unit-fold` and 15
|
||
of 2 761 on the shipped defaults; the whole distance is two proposer rules,
|
||
decomposed with denominators in the report.
|
||
- hit@k over six questions, k=50: **3/6 · 5/6 · 6/6**, matching the PDF arm
|
||
row for row, with the known-positive moving from **rank 13 to rank 1**.
|
||
2 022 concept ids are shared between the two channels -- 96.8 % of the XML
|
||
bundle.
|
||
- **No new dependency:** `xml.etree.ElementTree` is stdlib and `uv.lock` is
|
||
untouched. A `<!DOCTYPE` is REFUSED unparsed with its own code, which is a
|
||
guarantee about this package rather than about the installed libexpat.
|
||
- `.xml` never routes through the converter, and it is measured about 12x
|
||
faster and about 30x smaller in peak memory than the PDF arm on the same
|
||
document and the same machine.
|
||
- **XML that is not STS gives 0 plans and a FAILED build, and that is not an
|
||
`.xml` defect.** The known-positive that decides it: a folder holding one
|
||
`.txt` of prose with no headings gives exactly the same three lines and the
|
||
same exit 2. This is general `okf build` behaviour for any structureless
|
||
document -- extraction works, 0 unreadable, the text is there, and the
|
||
proposer has nothing to propose. The gate that refuses a run with no plans
|
||
stays: a run replaying zero plans would emit a flat bundle and report it as
|
||
a success. Separating "0 plans, 0 unreadable" from "0 plans because nothing
|
||
could be read" would change the outcome on **0 of the 4** reference
|
||
corpora, so it is not separated.
|
||
- Report: `docs/2026-09-11-k3-runde13-xml-sts.md`.
|
||
|
||
### Fixed
|
||
|
||
- **A PDF bookmark sharing a line with another left no trace.** `pdf_outline`
|
||
collected marks in a dict keyed on the destination line index, so a second
|
||
bookmark on a line was discarded by `setdefault` in silence: measured on a
|
||
701-page document, **2 763 nodes in, 2 762 marks out, `unresolved` = 0**.
|
||
`PdfOutline` now carries `collided`, and the identity `nodes in == marks +
|
||
unresolved + collided` holds. Keeping both nodes was measured and felled --
|
||
the two candidates then open at one offset and the first closes with an empty
|
||
span the orphan check deletes.
|
||
|
||
- **`--pdf-outline` (OFF): cut a PDF at the boundaries its own `/Outlines`
|
||
bookmark tree declares.** Measured outside this repository on one 701-page
|
||
process code whose publisher also ships a NISO-STS structure for it: the
|
||
shipped default recovers **1 967 of 2 761** titled sections, **0 of its 28**
|
||
chapters, and **794 of 794** misses have their heading text present in the
|
||
extracted text -- the line was read, the boundary was never opened. The same
|
||
file carries a 2 763-node bookmark tree that matches **2 761 of 2 761** STS
|
||
titles exactly after `re.sub(r"\s+","",s).lower()`. With the arm on:
|
||
**2 759 of 2 761 boundaries (99.9 %)**, depth 1 **28 of 28**, concept titles
|
||
identical to the publisher's own after that normalisation **2 761 of 2 761**,
|
||
false positives **3 of 2 762** (was 163 of 2 182), directories carrying two
|
||
concept files **2** (was 132, of which 65 were a contents copy and a body
|
||
section under one id), front-matter concepts **2 of 2 762** (was 72). Seven
|
||
of seven consumption fasit now exist in the bundle (was four); hit@1/8/50 is
|
||
**3/6 · 5/6 · 6/6** against **1/6 · 2/6 · 4/6**.
|
||
- It is a SEGMENTATION arm, not a reader option: the extracted text is byte
|
||
for byte the same either way, and `--pdf-headings`/`--ocr` stay the only
|
||
two things that change what a PDF says.
|
||
- A PDF with no bookmark tree is **byte-identical with the flag on**;
|
||
`pdfminer`'s `PDFNoOutlines` is "this file has no index", never an error.
|
||
Measured on the five-document smoke folder: `diff -r` empty against both
|
||
the arm off and the pre-change tree.
|
||
- No new dependency and no second parse of the file's pages: the tree is read
|
||
through `pdfminer.six`'s `PDFDocument.get_outlines()`, which
|
||
`pdfplumber` already ships under the existing `[extract]` extra. Cost on
|
||
the 701-page document: 119.22 s -> 183.31 s wall, peak RSS 3 252 -> 3 251
|
||
MiB.
|
||
- An unresolvable `/Dest` is dropped and COUNTED, never fabricated into a
|
||
boundary and never a refusal of the file. That document has 0 of 2 763;
|
||
one of the eight reference PDFs in this repository's own sample has 2 of 2.
|
||
- **The default does not move in this release.** Reach measured: 1 of the 8
|
||
reference PDFs carries a usable tree at all.
|
||
|
||
- **`.pdf` has a row in `extract._EVIDENCE`, as `measured`.** It was the row
|
||
with the most measurement behind it and no entry in the table, which is the
|
||
one way a table like that misleads while every entry in it is true.
|
||
|
||
|
||
## [0.7.0] — 2026-09-09
|
||
|
||
The first screen an agent reads, the three shapes of request the skill answers,
|
||
and one defect that made the command the first screen recommends build a worse
|
||
bundle than the command it claims to be.
|
||
|
||
> **Note for path importers, added 2026-09-09 after a report from a consumer.**
|
||
> `tools/okf_consume.py` ends by replacing its own `sys.modules` entry with the
|
||
> packaged `llm_ingestion_okf.consume`. A caller importing it with
|
||
> `importlib.util.spec_from_file_location` holds the object `module_from_spec`
|
||
> returned, which that line does not reach: read the module back out of
|
||
> `sys.modules[<name>]` after `exec_module`, or import `llm_ingestion_okf.consume`
|
||
> directly. From the commit that adds this note the file also copies the
|
||
> implementation's public names into its own globals, so a path-imported object
|
||
> carries them -- but that restores attribute ACCESS only, never patch-through.
|
||
|
||
### Fixed
|
||
|
||
- **`okf project` built a bundle two rules behind `okf build`.** `cli.build`'s
|
||
Python signature defaulted `keep_table_heading` and `sheet_section_rows` to
|
||
`False` while argparse defaulted both to `True`; `project.create` calls
|
||
`build()` as a function and passes no flag list, so it read the signature.
|
||
Measured on a five-document folder: `okf project` wrote **15 concepts / 30
|
||
files** where `okf build` on the same folder wrote **26 / 52**, and the whole
|
||
difference was in the priced spreadsheet — the document a question about price
|
||
has to reach. The invariant test that was supposed to catch this could not:
|
||
it compared `project.create` against the same `build()` function, so both
|
||
sides carried the same wrong value, and its two fixture documents had neither
|
||
a table nor a sheet. Both gaps are now tests: one compares the signature's
|
||
defaults against argparse's, the other builds a document whose concept count
|
||
actually moves with the two flags. After the fix the two paths are byte-equal
|
||
on that folder (`diff -rq`, 0 differences).
|
||
|
||
### Added
|
||
|
||
- **A first screen for a reader who has not used this before**, agent or human:
|
||
what it is in one sentence, one install line, two commands, and the three
|
||
shapes of request. The phase-status paragraph that used to open the README
|
||
moved down to `## What this library is`; nothing was deleted.
|
||
- **Three modes in the consumption skill**, stated in the template, the
|
||
instantiated skill and the generator:
|
||
- **Question** — as before, the default.
|
||
- **Hypothesis** — decomposed into premises and answered **per premise** as
|
||
`confirmed` / `refuted` / `undecidable-from-bundle`, three literals with no
|
||
fourth value. A premise whose excerpt is real but does not carry the
|
||
conclusion is `[sourced-not-sufficient]` on **that premise**, not on the
|
||
whole answer: four premises and one weak source is three answers and one
|
||
gap, and reporting it as one refusal throws the three away.
|
||
- **Task that produces a document or a paragraph** — every claim in the
|
||
written artefact carries `(bundle_id, concept_id)`, the excerpt's `sha256`,
|
||
its `title` and whichever `source_*` keys it has; an ungrounded paragraph is
|
||
**written and marked**, never dropped; and the cut (`considered`,
|
||
`withheld`, `delivered`) is declared inside the document, because the
|
||
document travels without the chat.
|
||
The five markings are untouched — the modes add no sixth.
|
||
|
||
### Changed
|
||
|
||
- **The generated skill states relative paths where it can.** In the layout
|
||
`okf project` writes, the commands are now `okf consume .okf/<id>` and
|
||
`okf check --skill .claude/skills/<id>-consume/SKILL.md`, runnable from the
|
||
project root — which is where `okf project`'s own closing line tells the
|
||
reader to start `claude`. A path outside the project root stays absolute on
|
||
purpose: `../../..` is not more portable, only harder to read. The two
|
||
absolute paths a generated skill carried are now zero, measured with a query
|
||
shown capable of finding first — O5's published "4 → 0" used `grep -c "^/"`
|
||
against paths indented by two spaces, which could not have matched either way.
|
||
- **One tag is pinned everywhere.** `README.md` pinned `v0.4.0` on its install
|
||
lines and `v0.6.0` further down, and `llms.txt` pinned `v0.4.0`; an agent
|
||
reading from the top installed a tag without `okf project`. All install lines
|
||
now name `v0.7.0`, and the earlier tags are kept as a labelled history
|
||
section rather than as commands. `llms.txt` gained the `okf project` form and
|
||
a pointer to the Claude Code section.
|
||
|
||
## [0.6.0] — 2026-09-08
|
||
|
||
The first tag since `v0.5.0a2`, so everything that had accumulated as
|
||
Unreleased is in it — those sections are kept below, under their own heading,
|
||
rather than folded together. What follows first is what O5 added, and why the
|
||
release needed a minor of its own: the installed command grew from one
|
||
subcommand to five, and one build default moved.
|
||
|
||
### Added
|
||
|
||
- **`okf consume`, `okf check`, `okf skill` and `okf project` are subcommands
|
||
of the installed `okf` command.** Until now the pre-pass, the contract
|
||
checker and the skill generator lived in `tools/` and were reachable only
|
||
from a clone; a consumer who installed this library could build a bundle and
|
||
had no way to read one back. Measured before the move: a consumption skill
|
||
generated from a checkout carried **four** lines naming that checkout by
|
||
absolute path, two of them the commands the skill tells a reader to run, so
|
||
the skill could not be moved, shared, or run by anyone else. The generated
|
||
skill now names `okf consume` and `okf check` — names on PATH — and a test
|
||
asserts the repository appears in it nowhere, with a known-positive so the
|
||
zero is a measurement rather than a search that could not find.
|
||
`tools/okf_consume.py`, `tools/okf_contract_check.py` and `tools/okf_skill.py`
|
||
remain as thin aliases, so every published reproduction block still runs.
|
||
- **`okf project <folder>`: a folder of documents to a bundle you can ask a
|
||
question of, in one command.** It runs `okf build` with this package's
|
||
default into `<out>/.okf/<id>/`, generates the skill into
|
||
`<out>/.claude/skills/<id>-consume/`, and prints what it read, what it wrote,
|
||
which documents are in the folder but not in the bundle, and which landed
|
||
whole as a single concept — the two cases a question can only be answered
|
||
`[sourced-not-sufficient]` in. `<out>` defaults to the current directory and
|
||
`<id>` to the folder's name reduced to `[a-z0-9-]`. It owns no flag that
|
||
changes a bundle's bytes, and a test holds the project bundle byte-equal to
|
||
the `okf build` bundle of the same folder.
|
||
- **`skills/okf-prosjekt/`**, a Claude Code skill (Norwegian) that wraps
|
||
`okf project` and reads its summary back.
|
||
- **The template and the contract document travel in the wheel.** `okf skill`
|
||
instantiates `skills/okf-consume-template/SKILL.md` and `okf consume`
|
||
measures `docs/consumption-contract.md` as its section 7.4 known-positive;
|
||
neither was installable before. Both are force-included from the file they
|
||
are authored in, so there is still exactly one copy of each.
|
||
|
||
### Changed
|
||
|
||
- **`okf build`'s default now includes Arm E (`--table-grid`), with
|
||
`--no-table-grid` as its opt-out.** The default moved to Arm D plus Arm F
|
||
earlier in the same day; measured afterwards, that combination is Arm F with
|
||
nothing to fold. The fold's table clause folds a table back into the heading
|
||
that introduces it, and with Arm E off a grid table is not one block but one
|
||
block per rule line. On the operator's twelve-document reference the shipped
|
||
default scored **2 of 12** and `docx` **0 of 3**, against the **5 of 12** the
|
||
fold was published with — which had been measured with Arm E on. Three arms
|
||
are now on by default, each with an explicit opt-out; `--outline-run 0
|
||
--no-table-grid --no-unit-fold` reproduces the pre-2026-09-08 bytes.
|
||
- **This tag does not make OKF v0.2 generally available.** `OKF_LATEST` is
|
||
unchanged.
|
||
|
||
### Also in this release: everything that had accumulated since v0.5.0a2
|
||
|
||
|
||
- **The corpus harness (`tools/okf_corpus_run.py`) can replay segmentation
|
||
plans, and it writes the bundle's `log.md`.** `--plans-dir` names the
|
||
proposals to replay and the profile follows from it; `--bundle-id` and
|
||
`--okf-version` are arguments, never constants, because a profile names a key
|
||
and the caller owns its value. `log.md` is written in SPEC section 9 form and
|
||
dated from `ingested_at`, so `merged + sum(coded rejections) == N` is
|
||
checkable from the bundle alone rather than only from a report that does not
|
||
travel with it. Without `--plans-dir` a run is unchanged.
|
||
- **The K2 rebuild that measured all of this.** The harness had passed
|
||
`STRUCTURED_V1` and no plans, so a 43-document corpus arrived as 39 flat
|
||
concepts with no `adjudication` key anywhere. Rebuilt, the same corpus yields
|
||
629 concepts, 618 of them `adjudication: proposed`. Record:
|
||
`docs/2026-09-03-k2-bundle-rebuild.md`.
|
||
|
||
- **`tools/okf_propose_segments.py` takes `--path-prefix`.** Section numbering
|
||
is document-local, so across 39 documents 618 proposed entries claimed only
|
||
601 distinct paths -- 17 collisions that Door B's gate refuses per document.
|
||
Scoping each document's entries under a caller-supplied prefix removes all
|
||
17. Without the flag every artifact already produced is byte-identical.
|
||
|
||
### Changed
|
||
|
||
- **A spreadsheet's tables are written as PIPE tables, so a row survives as a
|
||
row.** The converter's default markdown writer emits SIMPLE tables, which pad
|
||
every cell out to the width of the widest cell in its column. Measured on a
|
||
real 43-document corpus: one 594-character prose cell turned the sheet holding
|
||
the tender's prices into a 67 244-character whitespace carpet with runs of up
|
||
to **887 characters between a label and its amount**, and the header row named
|
||
a single column because only the first cell of the source's row 1 is filled.
|
||
The bytes reached a live consumer's model in two of eleven prompts and
|
||
appeared in none of its eleven answers. The same sheet through the pipe writer
|
||
is **11 048 characters with no whitespace run longer than two**, one row per
|
||
line, each source column its own cell. `--columns=1` is part of the fix and
|
||
not cosmetic: the pipe writer pads to a width computed from that setting, so
|
||
at the default a NARROW table gains runs of up to 45. Spreadsheet-only: the
|
||
other four office rows have the same defect available to the same one-line
|
||
change, but a spreadsheet is a grid with no prose fallback, while moving the
|
||
prose rows would move a corpus denominator nothing has measured.
|
||
`tests/test_extract.py` pins that scoping with three digests. Record:
|
||
`docs/2026-09-08-prisform-og-loggen-k2.md`.
|
||
|
||
- **An integral spreadsheet number loses the converter's trailing `.0`, and the
|
||
shared string table is what makes that safe.** The converter renders a numeric
|
||
cell as a double, so `5647500` arrives as `5647500.0` -- and a TEXT cell
|
||
reading `92.0` arrives as `92.0` too, which the output alone cannot tell
|
||
apart. The rewrite is bounded to a table cell whose entire content is such a
|
||
number, by unescaped pipes on both sides, and it is skipped whenever the same
|
||
literal is in the workbook's shared string table. Any workbook this cannot
|
||
read keeps its converter decimals rather than being guessed at.
|
||
|
||
- **The bundle's root `index.md` no longer links `log.md`.** The harness added
|
||
that link (`95eb271`) so a reader entering at `index.md` could reach the one
|
||
file carrying `N`; it was a LOCAL choice and said so. Consumption contract
|
||
SS 9.2 forbids a consumer from enumerating the bundle directory unless the
|
||
named profile says the index is derived, which makes the index tree the entire
|
||
map a consumer may use -- so everything it links is a document. Measured: a
|
||
consumer walking a 629-concept bundle that way returned **630**, and a corpus
|
||
run's own log became readable and citable as content. `5a0c879` excluded
|
||
`log.md` from OUR walk, which fixed the count on one side of a disagreement
|
||
produced on the other. The log is still written to the bundle root, which is
|
||
where SPEC section 9 puts it, and `tools/okf_consume.py` still excludes a
|
||
linked `log.md` -- every bundle built between `95eb271` and this change
|
||
carries the link.
|
||
|
||
- **The pinned `llm-ingestion-guard` moves to `v1.3.0`** (`[tool.uv.sources]`
|
||
and `uv.lock`; the `>=1.2,<2.0` range in `[project.dependencies]` already
|
||
covered it and is unchanged). What this fixes is that the guard could not
|
||
read back what this library WRITES: at `1.2.0` the flow-form `sources` in the
|
||
OKF v0.2 golden was refused outright, and flow is the only form this library
|
||
is able to emit, because its own line-oriented parser cannot round-trip the
|
||
block form. Measured before the bump so the test discriminates rather than
|
||
merely passes, and pinned by
|
||
`tests/test_guard_adapter.py::test_the_guard_parses_the_flow_form_sources_our_goldens_emit`.
|
||
Two rows of the gate table in `docs/okf-nokkelinventar.md` moved, not one:
|
||
the BLOCK form of `sources` now passes too, which retires G30 -- though it
|
||
changes nothing about what we emit, since our own parser is still the binding
|
||
constraint. `resource` is allowlisted only inside a `sources` entry, so
|
||
section 10.2's `executor`/`attester` resource stays rejected through every
|
||
carrier and the Door C boundary is unmoved. No public API of this library
|
||
changes; a consumer's cost is a re-run.
|
||
- **`uv.lock` also picks up `pypandoc-binary==1.17`**, which is a stale lockfile
|
||
being corrected rather than a new dependency: the package was already declared
|
||
in the `[extract]` extra, and `uv lock --check` reports the lockfile as out of
|
||
date on the untouched tree. Core still has exactly one runtime dependency.
|
||
- **`tools/okf_propose_segments.py` writes no artifact when it has nothing to
|
||
propose**, exiting `1` (distinct from `2`, "could not do the job") instead of
|
||
`0` with an empty plan. An empty plan cannot be replayed -- `process_inbox`
|
||
refuses one, because a plan naming no entry would persist nothing for a
|
||
document that was dropped -- so the file's only possible use was to fail a
|
||
run later, and it did. Measured: 11 of 39 documents in the K2 corpus propose
|
||
zero segments.
|
||
|
||
- **A concept can now record more than one source.** `sources` renders a flow
|
||
sequence of N flow mappings on one line, so a v0.2 profile can express
|
||
multi-source provenance instead of the single entry that was the measured
|
||
ceiling on SPEC 5.1 coverage. **A single source is byte-identical to before**,
|
||
so every golden is unmoved and no existing bundle changes.
|
||
|
||
The form is flow, not the block list PM decision B6 prescribed, and the
|
||
reason is measured: this library's frontmatter parser is line-oriented and
|
||
skips indented lines, so a block list round-trips to an EMPTY value with every
|
||
entry silently gone -- and the consumer the decision was written for accepts
|
||
the multi-entry flow sequence while classifying a block sequence as unreadable
|
||
provenance. Emitting block would have produced records neither side can read.
|
||
A negative-control test pins the block form's data loss so the reason stays
|
||
measurable rather than remembered.
|
||
|
||
New error code `sources_empty`: an empty list is refused, because
|
||
`sources: []` reads as a measured absence when it is the absence of a
|
||
measurement.
|
||
|
||
- **The consumption contract, stated normatively** in
|
||
`docs/consumption-contract.md`, with a copyable skill template
|
||
(`skills/okf-consume-template/`) and a checker (`tools/okf_contract_check.py`)
|
||
that reads its mechanically checkable half: payload shape, source marking per
|
||
excerpt, the closed `adjudication` and `trust_tier` state sets, denominator
|
||
identity, and the budget gate with its validated known-positive. The checker
|
||
ships outside `src/`, so no consumer's install surface changes.
|
||
|
||
- **Five office formats through a vendored converter**, behind the same
|
||
optional `[extract]` extra: `docx`, `xlsx`, `pptx`, `odt`, `rtf`. The
|
||
converter binary travels inside the wheel and is resolved by path rather than
|
||
found on `PATH`, with its version asserted against a pin -- `pypandoc`
|
||
searches `PATH` first and takes the highest version it finds, so a vendored
|
||
binary buys nothing until something resolves it explicitly. A host carrying a
|
||
different converter is refused rather than silently used.
|
||
|
||
**Two of the five rows are measured; three are not.** The corpus this work
|
||
was measured on contains zero `pptx`, `odt` and `rtf` files, so those rows
|
||
work by construction and have never met a document anyone wrote. The
|
||
distinction is asserted in the suite, not left in a comment.
|
||
|
||
`.doc` (Word 97) stays out -- the converter does not read it. Drawn content
|
||
does not survive extraction in any format, and every conversion warns about
|
||
it.
|
||
|
||
- **Four error codes** for the converter path: `extractor_binary_missing`,
|
||
`extractor_binary_version`, `extractor_convert_error` and
|
||
`extractor_empty_conversion`.
|
||
|
||
- **`tests/test_docs_promises.py`**, which asserts the README's published
|
||
format list against the registries it describes. This exists because the
|
||
README's previous promise -- that `docx` and `xlsx` always fail fast -- went
|
||
false silently when the converter landed. A published guarantee is a test
|
||
obligation.
|
||
|
||
|
||
- **Door B extracts `pdf` behind the optional `[extract]` extra.** The extra is
|
||
populated for the first time, with one parser: `pdfplumber>=0.11.10,<0.12`
|
||
(MIT). The default install is unchanged — still exactly one runtime
|
||
dependency, still stdlib otherwise — and a packaging test enforces that.
|
||
|
||
**The parser choice was forced by a measurement, not by preference**
|
||
(`docs/2026-08-21-g2-pdf-extraction-measurement.md`). On a real Vegnormalene
|
||
requirement table, `pdfplumber` keeps 4 of 4 rows with label and value on the
|
||
same line; `pypdf`, `pdfminer.six` and `pymupdf` each keep 0 of 4, emitting
|
||
all labels and then all values. A downstream reader can only re-pair those by
|
||
guessing, and in a requirements document a wrong pairing looks right.
|
||
`pymupdf` was additionally excluded on licence (AGPL-3.0 or commercial):
|
||
this package is MIT and an optional extra must not hand a consumer copyleft
|
||
they did not choose.
|
||
|
||
**`docx` and `xlsx` were unchanged at this release**, shipping no parser and
|
||
failing fast with `extractor_extra_missing`. That is no longer true: a later
|
||
release added a vendored converter reaching five office formats. This entry
|
||
is left as written — a changelog records what a release did — and the
|
||
correction is stated here so a reader arriving at this line is not misled.
|
||
|
||
- **`ExtractionWarning`**, exported from the package. Every `pdf` extraction
|
||
emits one. Text extraction recovers text; anything a PDF *draws* — figures,
|
||
diagrams, images — has no text to recover, so only captions survive and a
|
||
bundle built from drawn documents is **incomplete by construction**. That is
|
||
categorically true rather than document-specific, so it is stated rather than
|
||
detected: deciding "is there a figure on this page" is a layout heuristic this
|
||
library does not own. A named class so it can be filtered deliberately.
|
||
|
||
- **Two error codes**, both mirroring existing patterns rather than inventing
|
||
behaviour: `extractor_empty_pdf` (a PDF yielded no text on any page — a
|
||
scanned or image-only document; refused rather than persisted as an empty
|
||
concept, which would be the silent skip this registry exists to prevent) and
|
||
`extractor_pdf_error` (the parser failed on the bytes; the third-party
|
||
exception is wrapped, never leaked).
|
||
|
||
### Changed
|
||
|
||
- **The `[extract]` gate for `pdf` is now an import probe rather than a
|
||
membership test.** The rejection did not change: without the extra installed,
|
||
`pdf` still raises `ExtractionError` with code `extractor_extra_missing` and
|
||
the same message naming the remedy. A consumer that has not installed the
|
||
extra sees no difference at all. That behaviour is asserted unconditionally,
|
||
including on machines where the parser *is* installed, so it cannot rot into
|
||
a skipped test.
|
||
|
||
**Known limitation, stated rather than worked around:** structured table
|
||
recovery is out of scope. `pdfplumber.extract_tables()` and
|
||
`PyMuPDF.find_tables()` — two independent implementations — return the *same*
|
||
wrong shape for the measured requirement table, and across the whole handbook
|
||
only 45 of 196 detected table objects are clean enough to hand to
|
||
`render_table` unchanged. The breakage is in the documents' ruling geometry,
|
||
not in either library. PDFs therefore enter this library as **prose**, with
|
||
table lines correctly paired.
|
||
|
||
**Extracted PDF text is pinned to an exact parser version.** `pdfplumber`
|
||
pins `pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
|
||
releases with no stability contract. Extraction is deterministic within a
|
||
parser version (measured across five configurations) and not guaranteed
|
||
across one, so any golden fixture built on extracted PDF text is a fixture
|
||
migration away from a parser upgrade. `tests/test_extract.py` freezes the
|
||
expected text of a committed fixture so that upgrade breaks something visible
|
||
instead of drifting silently; see `tests/fixtures/README.md`.
|
||
|
||
- **`DEFAULT` now stamps `generated: { by: process:okf-ingest, at: <ingested_at> }`
|
||
instead of `generated: true`.** This is a byte change in every concept file Door A
|
||
writes under `DEFAULT`, so a consumer's own golden fixtures will show one changed
|
||
line per generated file.
|
||
|
||
The trigger is not an upstream Google release. `DEFAULT` states the ingest-spec
|
||
layer owned by `portfolio-optimiser-commons`, and they ratified this shape
|
||
(2026-08-02) and executed it in their spec on 2026-08-09. §7 defines the value:
|
||
the actor is the fixed `process:okf-ingest`, `at` repeats `ingested_at` verbatim,
|
||
and it is unquoted because frontmatter is parsed line-oriented — a quote would be
|
||
a character in the value rather than syntax a reader strips.
|
||
|
||
**What it costs a consumer is a re-run, and nothing else.** Ownership recognition
|
||
is one-way: a profile carrying an actor still owns the older literal stamp, so a
|
||
bundle written by an earlier version re-runs in place rather than tripping the §3
|
||
collision gate. No call signature changed and no key was added or removed.
|
||
|
||
Two things it is NOT. It is not a migration of `DEFAULT` onto OKF v0.2 — the
|
||
profile remains v0.1 on every axis upstream owns, and still emits no `sources`.
|
||
And it does not make `DEFAULT` and `OKF_V0_2` the same profile; they now agree on
|
||
the stamp and continue to differ in index root frontmatter, type-conditional
|
||
requirements, and `sources` derivation.
|
||
|
||
Door B (`process_inbox`) is deliberately unchanged: its marker is `generated`
|
||
plus `source_file`, it is not governed by the ingest-spec, and it stays disjoint
|
||
from Door A's `ingest_manifest`.
|
||
|
||
- **The v0.5.0 pilot set gained a fourth member: `portfolio-optimiser`.** Admitted
|
||
2026-08-09 on their request. `v0.5.0a1`'s entry below says "do not pin this tag
|
||
outside the pilot set" and names three repos; that entry is left as written,
|
||
because it records what was true when the tag shipped. This is the amendment,
|
||
and the sentence still binds — the set is now those three plus
|
||
`portfolio-optimiser`.
|
||
|
||
The reason is the producer axis, not the count. `portfolio-optimiser-claude` is
|
||
parked, and with it parked no original member could *emit* a v0.2 bundle at all:
|
||
`claude-code-llm-wiki` is read-only in the pilot and `catalog` is gate-side. The
|
||
new member consumes the same Door A. Nothing about the provisional surface
|
||
changes: `OKF_LATEST` still points at `DEFAULT`, and the v0.2 surface may still
|
||
move on pilot feedback without a deprecation cycle.
|
||
|
||
**If you pin `v0.5.0a2`, its guard pin is `>=0.2,<0.3`** — that tag resolves
|
||
against guard `v0.2.0`, *not* the `v0.3.4` this repo's `main` now uses. `main`
|
||
moved to `>=0.3,<0.4` after the tag. Also note `tool.uv.sources` is not
|
||
inherited transitively: a consumer supplies the guard's source itself.
|
||
|
||
- **The guard pin moved to `>=0.3,<0.4`, resolved against `v0.3.4`.** The
|
||
window is widened only after measurement, never before: the 19-fixture
|
||
guard-surface suite was run against `v0.3.4` in a scratch venv first, and
|
||
reproduced exactly the three deltas measured against `v0.3.3` — no new ones.
|
||
`v0.3.4`'s own fixes are regex-complexity repairs, one of them
|
||
(`okf._MD_LINK_RE`) on Door C's call path, with no disposition changes.
|
||
|
||
- **Door C now passes `allow_reserved=False` to `okf.import_bundle`.** The
|
||
guard added the keyword in the `0.3.x` line and defaults it `True` for the
|
||
received-bundle path, which would merge a sender's `index.md` / `log.md`
|
||
instead of rejecting them. Door C overrides it, keeping the unconditional
|
||
reserved-name refusal committed to before the keyword existed. The reason is
|
||
structural rather than a second opinion on the guard's scan: Door C generates
|
||
the merged bundle's `index.md` from what it merged and writes every merged
|
||
concept verbatim, so a sender's `index.md` would be a second, irreconcilable
|
||
claim on one path.
|
||
|
||
**This is not a behaviour change for anyone on the previous pin.** Under
|
||
`v0.2.0` the keyword did not exist and reserved names were refused by
|
||
construction; the explicit argument preserves that outcome across the bump.
|
||
A consumer sees the same rejections, with the same reasons, before and after.
|
||
|
||
### Fixed
|
||
|
||
- **`tools/okf_adjudicate.py` exits 2 on a malformed plan instead of raising.**
|
||
A `SegmentationError` from the plan grammar used to escape `main()` as a
|
||
traceback and exit 1, while every other malformed-plan case in the same file
|
||
already returned 2. Exit codes are the interface a caller scripts against,
|
||
and exit 1 with a traceback says "this command broke" where the truth is
|
||
"this file is not a plan" -- the two are the same code an unhandled bug would
|
||
produce. The refusal itself is unchanged: nothing is written either way, and
|
||
the grammar is untouched. Now one line on stderr naming the error code, and
|
||
exit 2. A caller that treated a non-zero exit as failure sees no difference;
|
||
one that distinguished 1 from 2 sees a malformed plan move into the class it
|
||
belonged to. Pre-existing since the empty-verdict branch landed, and pinned
|
||
until now by a test that recorded it as a finding rather than fixing it.
|
||
|
||
- **A verdict `okf_adjudicate` cannot read back is no longer reported as a
|
||
malformed plan.** The command parses the plan before building anything and
|
||
parses its own verdict after, so only the first failure is the operator's
|
||
file. The second is reachable -- an empty `--adjudicator` produces a verdict
|
||
the grammar refuses -- and before the change above it surfaced as a
|
||
traceback. Catching the grammar error at the top would have turned it into a
|
||
clean, confident, wrong diagnosis naming the plan file, which is worse than
|
||
the traceback: it sends the operator to fix the one artifact that was fine.
|
||
The verdict parse now raises `AdjudicationError`, so the message names what
|
||
actually failed. Exit 2 either way; nothing is written either way.
|
||
|
||
- **`okf build --ingested-at` alone now stamps every concept, not just the
|
||
unsegmented ones.** `--proposed-at` defaulted independently to
|
||
`DEFAULT_STAMP`, so a caller who set only `--ingested-at` stamped the 11 of
|
||
629 concepts that read the call's value directly, while the 618 segmented
|
||
concepts -- which read `segment.ingested_at`, the plan's `proposed_at` --
|
||
stayed on `1970-01-01T00:00:00Z`. Measured on K2 rebuilt at `fbaac6d`.
|
||
`proposed_at` now defaults to `ingested_at` when omitted; a caller who wants
|
||
the proposal and the replay dated differently still passes both explicitly.
|
||
Neither flag passed still yields `DEFAULT_STAMP` for both, byte-identical to
|
||
before.
|
||
|
||
- **The consumption pre-pass (`tools/okf_consume.py`) no longer counts a
|
||
linked `log.md` as a concept.** `link_log_in_root_index` (`corpus.py`,
|
||
`95eb271`) links a run's own log from the root index for bundle navigation;
|
||
the index walk that enumerates concepts followed that link like any other
|
||
and counted the log as one, inflating a 629-concept K2 rebuild to 630 and
|
||
letting the log rank and get cut like real content. The link stays --
|
||
`docs/consumption-contract.md` is silent on `log.md`, and `95eb271` already
|
||
named the link a LOCAL choice rather than conformance -- but the walk now
|
||
recognises `LOG_NAME` the same way it recognises the index itself: reachable
|
||
for navigation, never a concept. A bundle whose index does not link its log
|
||
(every bundle built before `95eb271`, including the delivered
|
||
`K2-bundle-20260903`) computes the same `sha256-tree` ref before and after.
|
||
|
||
## [0.5.0a2] — 2026-07-31
|
||
|
||
**This is the pre-release the pilots pin. `v0.5.0a1` was tagged and abandoned
|
||
unused — do not pin it.** It carried a `generated.by` actor that the spec owner
|
||
had already excluded, and it was caught before any pilot was notified.
|
||
|
||
### Fixed
|
||
|
||
- **`OKF_V0_2`'s `generated.by` actor is `process:okf-ingest`**, not
|
||
`process:llm-ingestion-okf`. Commons decided the fixed id on this repo's own
|
||
proposal 2026-07-31, superseding option (d) chosen here 2026-07-27. The
|
||
exclusion is `ingest-spec.md:7-8`, frozen on the spec being framework-neutral:
|
||
normalising *our* repo name into the id would force every other conformant
|
||
implementation to write it into its own output.
|
||
|
||
Nothing in the wild carried the excluded value — `OKF_V0_2` did not exist at
|
||
`v0.4.0`, and no pilot had been notified — so this costs a tag rather than a
|
||
migration. It is recorded rather than quietly folded in because the failure it
|
||
avoids is specific: `actor` is both the stamp written and the value owned back
|
||
(`OwnershipPolicy`), and recognition is one-way, so a pilot holding bundles
|
||
stamped with the excluded id would have hit `collision_unstamped` on its own
|
||
files the moment the id was corrected.
|
||
|
||
## [0.5.0a1] — 2026-07-31 — ABANDONED, do not pin
|
||
|
||
**Pre-release. PROVISIONAL surface.** Shipped to a named pilot set —
|
||
`portfolio-optimiser-claude`, the plugin marketplace catalog, and
|
||
`claude-code-llm-wiki` — so that real data can find what fixtures cannot. The
|
||
v0.2 surface may change on their feedback **without a deprecation cycle**.
|
||
Saying so up front is what buys the freedom to act on the feedback; discovering
|
||
it later is what would make a pilot a de-facto release. Do not pin this tag
|
||
outside the pilot set.
|
||
|
||
Support for a new upstream version is **additive — a new profile, never a
|
||
migration**. `OKF_LATEST` still points at `DEFAULT`; flipping it is the GA
|
||
event, not a side effect of this tag.
|
||
|
||
### Added
|
||
|
||
- **`OKF_V0_2` profile.** Google OKF v0.2 as a profile alongside `DEFAULT` and
|
||
`STRICT_V1`. It closes nothing: §14 forbids a conformant consumer to reject on
|
||
an unknown `type` value or on unknown additional keys, so `type` is the only
|
||
required key. It NAMES `okf_version` but never carries its value — that value
|
||
tracks the upstream Google version and belongs to catalog (decision E1), so
|
||
the caller supplies it via `materialize_bundle(..., root_frontmatter_values=…)`
|
||
and the profile fixes only the key and its position.
|
||
- **`root_frontmatter_values`** on `materialize_bundle`, keyword-only: the
|
||
mechanism behind "a profile names a key, a caller owns its value".
|
||
- **`sources` emitted as an inline flow sequence**, populated from the manifest.
|
||
- **The v0.2 golden fixture**, compared byte-for-byte like the others.
|
||
|
||
### Changed
|
||
|
||
- **`profile` on `materialize_bundle` is keyword-only**, so existing positional
|
||
call sites stay source-compatible.
|
||
- **`_is_ingest_owned` is profile-aware.** Ownership is a policy on the profile,
|
||
not a literal; recognition stays one-way.
|
||
|
||
### Fixed
|
||
|
||
- **`root_frontmatter` permitted a key without requiring it — the two had been
|
||
the same profile field.** Upstream says MAY where the field said MUST, so a
|
||
root index that legally omitted an optional key was rejected. `root_frontmatter`
|
||
now permits and orders; the new `root_frontmatter_required` requires.
|
||
`STRICT_V1` keeps requiring its three (unchanged for its consumer), `OKF_V0_2`
|
||
requires none, `DEFAULT` is untouched. Measured against the eight root indexes
|
||
available locally: 7 of 8 failed before, 0 of 8 after.
|
||
|
||
### Notes
|
||
|
||
- **`DEFAULT` and `STRICT_V1` are byte-stable across this release.** The emit
|
||
path is byte-identical; the golden suite is the check.
|
||
- **Known, deliberately unfixed:** `TypePolicy`'s closed branch does not strip
|
||
quotes from a declared type. All three call sites pass `DEFAULT.types`, and
|
||
both `DEFAULT` and `OKF_V0_2` set `allowed=None`, so the branch is
|
||
unreachable in shipped code — it fires only for a caller constructing
|
||
`STRICT_V1` and calling `rejection` directly. Fixing it would repair a write
|
||
path that never sees quotes, and would fix the meaning of a quote character
|
||
without a value model that can express one, in a line-oriented format.
|
||
|
||
## [0.4.0] — 2026-07-25
|
||
|
||
Phase 2. The bundle inbox (Door B) and external-bundle import (Door C) ship,
|
||
and with them this library's first — and only ever — runtime dependency.
|
||
|
||
### Added
|
||
|
||
- **Door B — bundle inbox (`process_inbox`).** Per dropped file: bytes →
|
||
extraction → persist gate → render → collision gate → write → index link.
|
||
Returns an `InboxResult` whose four buckets (`persisted`, `quarantined`,
|
||
`rejected`, `failed`) are disjoint and complete, so a file that vanished
|
||
mid-run surfaces as a missing entry rather than as nothing. One bad file
|
||
never aborts the run: only an invalid `ingested_at`, a reserved `okf_type`,
|
||
and a missing inbox directory fail the whole run, each being wrong for every
|
||
file at once. Concept files are named `inbox-{slug}.md`, disjoint from
|
||
`index.md`, Door A's `ingest-*`, and `promoted-verdict-*`; `source_sha256`
|
||
is taken over the original dropped bytes, so provenance stays re-verifiable
|
||
against the operator's file.
|
||
- **Door C — external bundle import (`import_bundle`).** Reads an external OKF
|
||
bundle and hands it *whole* to an injected gate over the guard's
|
||
`okf.import_bundle` — a bundle-level call, because it resolves the
|
||
cross-link graph across concepts — merging only concepts that clear the
|
||
non-blocking floor. Two invariants, both load-bearing: a merged concept is
|
||
written **verbatim** (stamping it would require round-tripping frontmatter
|
||
through this library's line-oriented parser, which cannot represent the
|
||
block lists the guard's parser accepts, and would persist bytes the guard
|
||
never screened), and ownership is therefore proven by **content identity** —
|
||
an occupied target name is re-used only when the bytes there are already
|
||
identical, never overwritten. Re-importing an unchanged bundle is a no-op.
|
||
- **`extract_text` — the Door B extraction registry.** All file-type → text
|
||
extraction lives in this library, because the guard is text-only. `md`/`txt`
|
||
pass through, `csv` renders the markdown table, `json` is fenced verbatim,
|
||
and `html`/`htm` reduce to text with `html.parser` — stdlib throughout.
|
||
`pdf`, `docx`, and `xlsx` are gated behind the `[extract]` extra, which
|
||
ships no parser yet: those types fail fast with a typed error naming the
|
||
extra, never a silent skip.
|
||
- **`llm_ingestion_okf.guard_adapter`** — the shipped gate over the real
|
||
guard (`inbox_gate`, `import_gate`), and the only module here that imports
|
||
it. Door B screens the exact bytes it persists: the guard's `prepare_input`
|
||
bookend prepares text for a model call this library never makes, so
|
||
`screen_output` alone is used and the screened string is the written string.
|
||
It follows that the gate **refuses rather than repairs** — a file carrying
|
||
an invisible carrier is rejected, not stripped and persisted. The policy is
|
||
`PRESET_USER_UPLOAD`, so any finding at all is held back.
|
||
- **15 new stable error codes**, each registered in the `errors.py` docstrings
|
||
(the stability contract) and pinned by the error-code conformance suite:
|
||
`extractor_decode_error`, `extractor_empty_csv`, `extractor_extra_missing`,
|
||
`extractor_unknown`, `import_label_invalid`, `import_path_empty`,
|
||
`import_path_too_long`, `import_provenance_invalid`,
|
||
`import_slug_collision`, `inbox_slug_collision`, `inbox_slug_empty`,
|
||
`inbox_slug_too_long`, `inbox_source_file_invalid`, `inbox_title_invalid`,
|
||
`okf_type_reserved`.
|
||
|
||
### Changed
|
||
|
||
- **`llm-ingestion-guard>=0.2,<0.3` is now a mandatory runtime dependency.**
|
||
Installing this package installs the guard. Importing it does not: only
|
||
`guard_adapter` imports the guard, so a Door A consumer keeps working
|
||
whatever state the dependency is in, and the doors themselves still take an
|
||
*injected* gate. Until the guard is published to a package index, install it
|
||
from its tag — see the README. A packaging test enforces that this stays the
|
||
only runtime dependency.
|
||
- **An extraction `title` containing `[` or `]` is rejected at manifest load**
|
||
(ingest-spec §4, ratified D1). Previously only single-line was validated. A
|
||
manifest that loaded before and carries a bracket in a title now fails fast
|
||
with code `manifest_schema`: the title renders verbatim into the index link
|
||
label `- [title](target)`, where a bracket breaks index-link and navigation
|
||
parsing downstream.
|
||
|
||
### Fixed
|
||
|
||
- **Materializing one manifest no longer deletes the files another manifest
|
||
stamped into the same bundle.** The §3 ownership scan classified every
|
||
ingest-stamped file as replaceable, so a second manifest sharing a bundle
|
||
removed the first one's concept files and their index links. Ownership is
|
||
now narrowed to files whose stamp names the running manifest by stem — the
|
||
stem is stable across content edits, so an edited manifest still reclaims
|
||
what a prior run of itself wrote. The operator-copy restriction (a generated
|
||
file copied into curated content) remains documented, not enforced.
|
||
|
||
### Notes
|
||
|
||
- Phase 2's binary extraction is **not** in this release: the `[extract]`
|
||
extra is declared but empty, and `pdf`/`docx`/`xlsx` therefore fail fast.
|
||
That is the one outstanding item from the phase, and it ships separately.
|
||
- **The ownership change above is an extension point, not a spec fix.** Filed
|
||
under "Fixed" because it stops silent data loss, but the spec owner
|
||
(`portfolio-optimiser-commons`) has since pointed out that removing *every*
|
||
ingest-stamped file is verbatim what ingest-spec v1 §5 mandates: v1 assumes
|
||
one manifest per bundle and defers multiple manifests feeding one bundle as
|
||
a named extension point. So this release implements that extension point and
|
||
outruns the frozen text rather than conforming to it. The mechanic itself is
|
||
not at risk — matching by manifest stem follows from the spec's own
|
||
`{stem}@{h}` stamp definition, since `{h}` changes on every content edit —
|
||
but the surrounding prose is queued for amendment in commons and is not
|
||
ratified. Treat multi-manifest bundles as ahead of the spec until it is.
|
||
- Exception `__cause__` preservation is now pinned by a conformance suite, one
|
||
test per fail-fast wrap site, alongside the existing `.code` suite.
|
||
|
||
## [0.3.2] — 2026-07-23
|
||
|
||
### Fixed
|
||
|
||
- **Frontmatter and index-label values are emitted verbatim; only
|
||
`source_query` is whitespace-collapsed.** Earlier releases collapsed every
|
||
whitespace run in every frontmatter value and index link label to a single
|
||
space. §5 of `ingest-spec.md` mandates that collapse for `source_query`
|
||
alone — where a multi-line SQL `SELECT` must render on one line — while
|
||
every other value is validated single-line at manifest load and passed
|
||
through unchanged: validation, not repair. A `title` carrying an internal
|
||
whitespace run now survives byte-for-byte at both the `title` frontmatter
|
||
and the index link label, instead of being silently altered. Output bytes
|
||
change only for values that contained a collapsible whitespace run; the
|
||
shipped golden fixtures and both consumers are unaffected.
|
||
|
||
## [0.3.1] — 2026-07-19
|
||
|
||
### Fixed
|
||
|
||
- **Documentation corrected a security claim that did not hold.** The module
|
||
docstring and README stated that this library "calls the guard at every persist
|
||
gate". That described the intended end state in the present tense. Door A — the
|
||
only door shipped — is ungated: the package has zero runtime dependencies and
|
||
calls no guard function before writing to disk. Both places now say so plainly,
|
||
and state that gating external or untrusted content is the caller's
|
||
responsibility (`okf.import_bundle`, or `prepare_input`/`screen_output`) until
|
||
the persist gates land with Doors B and C.
|
||
|
||
No behavior changed in this release. The correction is published because a
|
||
consumer read the earlier wording as safe-by-default and would have persisted
|
||
ungated content on that basis.
|
||
|
||
## [0.3.0] — 2026-07-17
|
||
|
||
### Added
|
||
|
||
- **Stable machine-readable error codes.** `IngestError` gained a keyword-only
|
||
`code` attribute (default `"unspecified"`). Roughly 24 codes are documented in
|
||
the `errors.py` docstrings, and those docstrings are the stability contract.
|
||
Consumers should assert on `exc.value.code`, not on message text.
|
||
|
||
### Changed
|
||
|
||
- **Exception message text is explicitly declared unstable.** It may change in any
|
||
release. Tests matching on message strings (`pytest.raises(match=...)`) should
|
||
migrate to code comparisons.
|
||
|
||
### Notes
|
||
|
||
- Generic schema shape violations deliberately share the single code
|
||
`manifest_schema`. One code per validation rule would have frozen an
|
||
unnecessarily large surface. Finer resolution is a separate decision, not an
|
||
assumed requirement.
|
||
|
||
## [0.2.0] — 2026-07-17
|
||
|
||
### Added
|
||
|
||
- **PEP 561 support.** The `py.typed` marker ships with the package, so consumers
|
||
get the inline type hints without a mypy override.
|
||
- **`IngestResult.stamp`** exposes the spec §5 provenance stamp on the result
|
||
object.
|
||
|
||
## [0.1.0] — 2026-07-16
|
||
|
||
Phase 1 (Door A) implemented against the normative `ingest-spec.md` owned by
|
||
`portfolio-optimiser-commons`. Never tagged; consumers pinned the commit
|
||
`dae0bd1a` directly.
|
||
|
||
### Added
|
||
|
||
- Fail-fast manifest validation (spec §3–§4).
|
||
- Spec §5 body renderers as pure functions.
|
||
- The `file` connector (CSV, fail-closed path boundary), the `sql` connector
|
||
(read-only sqlite, env-resolved credentials), and the `http` connector behind an
|
||
explicit per-run network opt-in.
|
||
- Spec §5 materialization with an in-memory staging collision gate.
|
||
- Index maintenance on re-materialization (spec §6).
|
||
- The spec §11 golden fixtures, compared byte-for-byte.
|
||
- The Door A public surface: `materialize_bundle` plus the typed error hierarchy
|
||
rooted in `IngestError`.
|
||
|
||
[0.6.0]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/v0.5.0a2...v0.6.0
|
||
[0.5.0a2]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/v0.5.0a1...v0.5.0a2
|
||
[0.5.0a1]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/v0.4.0...v0.5.0a1
|
||
[0.4.0]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/v0.3.2...v0.4.0
|
||
[0.3.2]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/v0.3.1...v0.3.2
|
||
[0.3.1]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/v0.3.0...v0.3.1
|
||
[0.3.0]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/v0.2.0...v0.3.0
|
||
[0.2.0]: https://git.fromaitochitta.com/open/llm-ingestion-okf/compare/dae0bd1a...v0.2.0
|
||
[0.1.0]: https://git.fromaitochitta.com/open/llm-ingestion-okf/src/commit/dae0bd1a
|