llm-ingestion-okf/pyproject.toml
Kjell Tore Guttormsen bc39e8091f feat(assets): a bundle carries the images its sources declare (0.10.0)
Until now no reader in this package fetched, named, described or copied a
single image. `<img>`'s attributes were never read, a NISO-STS `<graphic>`
was walked past, a PDF was opened for its text alone, the converter's
markdown writer dropped every picture, and the only writer into a bundle
took `content: str`. The two lossiness warnings said so on every run, which
made the loss honest and did not make it smaller.

Measured on R761 Prosesskoden:2025, published as a 701-page PDF and as a
NISO-STS delivery: the process text is carried in full while 12 `Tabell N-N`
and 9 `Figur N-N` captions stand over nothing, because that publisher ships
those tables as raster pictures in both. Process 84's "toleranseklasse ...
er gitt i tabell 84-2" points at empty space.

THE GATE WAS WRITTEN FIRST AND RED. `tests/test_asset_gate.py` reads its
denominator out of the source (`page.images`, `word/media/`, `ppt/media/`,
`<img`, `<graphic`), never from a constant here. Measured at 332961a, built
from `git archive` and not from the editable tree: carried 0 of 8 local
images across 5 documents (9 declared), and no `assets/` at all. After: 8 of
8, with the ninth a remote source carried as a pointer without a file.

FIVE READERS PLACE, ONE MODULE DECIDES. `assets.py` owns what an image is
(sniffed from the bytes, never from the claimed extension), what it is
called (`<sha256[:12]>-<the source's own basename>`) and how it is pointed
at (one two-line block, one regex). `.xlsx` is deliberately not a row: a
block inside its pipe tables would break the `source_rows` locator, and 0 of
4 K2 workbooks hold media.

A PDF stream that is already a file is carried VERBATIM (29 of R761's 50
objects are DCTDecode); raw samples are encoded to PNG with stdlib zlib, so
no new dependency. Rendering the page region was the alternative and was
felled on determinism: a rasterised crop's bytes, and therefore the asset's
content-addressed name and the bundle's digest, would depend on the
installed rasteriser. What the encoder cannot express exactly is refused
with a code and counted, never approximated.

NO SIZE FLOOR, and that is a measurement: over the 4 828 image objects of
the K2 corpus the size distribution is a broad spread with no gap, unlike
OCR_CID_SHARE's bimodal one, so a threshold would be a number we chose.

ON BY DEFAULT, AND THE CONTROL IS TWO WHOLE BUILDS. The 43-document
reference corpus at 332961a versus rebuilt at HEAD with `--no-assets`:
865 files on both sides, `diff -rq` reports ONE difference, the added
`Images: NOT CARRIED` line in log.md. Every concept byte-identical.
Against the default: 453 -> 454 concepts, 865 -> 867 md, 0 -> 2 964 assets
(2 964 carried of 3 145 found, 4 622 pointers), 4.7 MB -> 115 MB, 2 414 s ->
3 088 s, peak RSS 6.26 -> 8.74 GB, 422 of 865 md files differ. The one new
concept has a measured cause: the pointers are body text, so a section
holding 146 of that document's images grew from 19.0 % to 30.6 % of the
extracted text and crossed `--outline-gate`'s 0.20 share clause.

THE IMAGE BYTES ARE NOT SCREENED. The guard is text-only, the pointer block
passes the gate as body text, the picture beside it passes nothing, and
log.md says so on every run.

Also fixed, both found by measuring rather than by reading:

- a markdown image is no longer read as a cross-reference. `structure._LINK`
  never looked at the character in front of the bracket, so every pointer
  would have arrived in the index as an edge to a concept that cannot exist.
- Door C carries the assets its merged concepts point at. Before this,
  importing a bundle built with `--assets` merged 6 of 6 concepts and wrote
  no `assets/` at all, so every pointer named a missing file.

Report: docs/2026-09-17-bilder-i-bundlen-trinn1.md
Spec proposal: docs/plan/okf-assets-section-6-4.md
Suite 1 955 passed / 1 skipped (from 1 896), ruff and mypy --strict clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 10:01:31 +02:00

258 lines
14 KiB
TOML

[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[project]
name = "llm-ingestion-okf"
version = "0.10.0"
description = "Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import, with security delegated to llm-ingestion-guard."
readme = "README.md"
license = "MIT"
requires-python = ">=3.10"
authors = [{ name = "Kjell Tore Guttormsen" }]
classifiers = [
"Development Status :: 3 - Alpha",
"Intended Audience :: Developers",
"Operating System :: OS Independent",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3.10",
]
# Exactly one runtime dependency, ever: the security boundary. Everything
# else is stdlib. The version range is the real pin — it resolves normally
# against a package index, and is satisfied today by the git+https tag
# install documented in the README (a direct reference is an install-time
# channel, not a dependency declaration).
#
# Floor 1.2, not the 1.0.0 freeze: this library needs the flow-mapping
# frontmatter support (`generated: { by: x, at: y }`) that landed in the
# guard's 1.2.0, without which Door C fail-secures every concept carrying
# it. Ceiling <2.0, not a narrower minor: the guard's own 1.0.0 release
# promises no exported name is removed, renamed or given a different
# meaning short of a 2.0.0 — calibration (severities, dispositions) is
# explicitly free to move within 1.x under that same promise, so a tighter
# ceiling here would claim a stability guarantee the guard does not need to
# keep and we do not need to demand.
dependencies = ["llm-ingestion-guard>=1.2,<2.0"]
# The installed command. `okf build <folder> --bundle <dir>` is the packaged
# form of a path that was two unpackaged scripts under `tools/` and nine flags
# -- reachable only from a clone, which is not where a consumer stands.
[project.scripts]
okf = "llm_ingestion_okf.cli:main"
[project.optional-dependencies]
# Binary file-type extraction parsers. OPT-IN ONLY: this extra pulls binary
# wheels (pillow, pypdfium2) and a transitive tree that core must never have —
# the "exactly one runtime dependency" rule above covers the default install,
# and this extra is outside it by construction.
#
# The extra names the parsers it actually ships, so a consumer installing it
# gets what the error message promised and nothing else. It ships two: a `pdf`
# reader, and a converter that reaches the office types.
#
# WHY pdfplumber, and why the floor is not free (measured 2026-08-21,
# docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real Vegnormalene
# requirement table pdfplumber keeps 4 of 4 rows with label and value on the
# same line; pypdf, pdfminer.six and pymupdf each keep 0 of 4, emitting all
# labels then all values, which a downstream reader can only re-pair by
# guessing. In a `krav` document that is a wrong answer that looks right.
# pymupdf is additionally out on LICENSE (AGPL-3.0 or commercial) — this
# package is MIT and an extra must not hand a consumer copyleft they did not
# choose.
#
# PARSER VERSION IS PART OF THE OUTPUT CONTRACT. pdfplumber pins
# `pdfminer.six==20260107` exactly, and pdfminer.six ships date-stamped
# releases with no stability contract. Extraction is deterministic WITHIN a
# parser version (measured, 5 configurations) and NOT guaranteed across one.
# `tests/test_extract.py` holds that promise against a committed fixture, so
# widening this range makes a test go red instead of letting extracted text
# drift silently. See tests/fixtures/README.md.
#
# WHY THE CONVERTER BINARY IS VENDORED RATHER THAN FOUND ON PATH. The `xlsx`
# and `pptx` readers exist only from pandoc 3.8.3. Debian 12 ships 2.17.1.1
# and Ubuntu 24.04 ships 3.1.3, so a PATH binary cannot deliver two of the
# five office formats on current stable distributions -- and a library whose
# output depends on which pandoc a host happens to carry is not deterministic
# in the sense the rest of this package means it.
#
# `pypandoc-binary` carries the binary inside the wheel (7 platform wheels at
# 1.17, including macosx x86_64/arm64, manylinux and musllinux x86_64/aarch64,
# and win_amd64 -- measured on the PyPI JSON API 2026-09-02). The pin is
# EXACT, not a range, because the binary's version is part of the output
# contract in the same way pdfminer.six's is: extraction is deterministic
# within a converter version and not across one.
#
# This does not widen the runtime dependency surface. The rule above governs
# `project.dependencies`, which still names the guard alone; the extra is
# outside it by construction, and the test below now pins its contents so a
# third entry cannot arrive unexamined.
extract = ["pdfplumber>=0.11.10,<0.12", "pypandoc-binary==1.17"]
# The OCR engine for `--ocr`, and NEVER a runtime dependency. It is a separate
# group from `extract` rather than three more entries in it, because it buys
# something categorically different: `extract` decides which file types can be
# read at all, while this one only changes how a PDF page is read when the
# page's own text never arrived. A consumer who installs `[extract]` gets every
# file type; a consumer who never meets a scanned document should never carry
# an inference runtime.
#
# WHY rapidocr ON onnxruntime, and why not the obvious alternative. Docling was
# measured first and is OUT on a platform fact, not a preference: it needs
# torch, and torch stopped publishing macOS x86_64 wheels after 2.2.2, with no
# `transformers` version inside Docling's own window that works against that
# one (4 tried, 2026-09-08). rapidocr on onnxruntime installs and runs on this
# machine, and it carries its ONNX models inside its own wheel, so `--ocr`
# needs no network at run time -- which matters here more than usual, since
# this library's network gate is an explicit per-run opt-in and an engine that
# downloaded a model on first use would walk straight through it.
#
# `pypdfium2` is named although `[extract]` already reaches it through
# pdfplumber: the OCR path RENDERS a page before reading it, and the renderer
# is a dependency of that path rather than a happy accident of another one.
#
# The pins are ranges rather than exact versions, and that is a weaker promise
# than `[extract]` makes on purpose: OCR output is a model's reading of an
# image, so it is deterministic within one model version and NOT across one,
# and no range can make it otherwise. A bundle built with `--ocr` is
# reproducible against the versions it was built with, which is stated in the
# report rather than implied by a pin.
ocr = ["rapidocr>=3.9,<4", "onnxruntime>=1.20,<2", "pypdfium2>=4,<6"]
[dependency-groups]
# ruff PINNED TO A RANGE, not floored at an ancient version. `ruff>=0.9` let
# the lockfile decide which ruff ran, and a frozen 0.15.22 is what made this
# tree read green while 0.16.6 found 148 things in it. The floor is now the
# version the acceptance was measured under, and the ceiling is the next minor,
# because 0.16 is itself the release that widened the default rule set.
#
# pyyaml is a TEST reader and nothing else (K3-22): SPEC SS 11 requires "a
# parseable YAML frontmatter block" in every file, and the only way to measure
# that is to ask a YAML reader. How a value is WRITTEN stays decided by a rule
# in `profiles`, never by a parser, so `src/` imports no yaml; the tests
# validate the rule against this reader. The floor is the version it was
# measured under (6.0.3, 2026-09-11).
dev = ["pytest>=8", "mypy>=1.14", "ruff>=0.16.6,<0.17", "pyyaml>=6.0.3,<7"]
[tool.hatch.build.targets.wheel]
packages = ["src/llm_ingestion_okf"]
# Two AUTHORED files the packaged commands cannot run without, carried into the
# wheel from where they are edited rather than committed a second time under
# `src/`. A duplicate would drift, and both of these are checked against
# literals in the code: a template whose blocks are pinned by a test, and a
# known-positive artefact whose byte count is a constant in `consume.py`.
#
# `okf skill` instantiates the template; `okf consume` measures the contract
# document as its section 7.4 known-positive and refuses without it. Before
# 2026-09-08 neither command was installable, so neither file had to travel.
[tool.hatch.build.targets.wheel.force-include]
"skills/okf-consume-template/SKILL.md" = "llm_ingestion_okf/_data/okf-consume-template.md"
"docs/consumption-contract.md" = "llm_ingestion_okf/_data/consumption-contract.md"
[tool.ruff]
line-length = 100
target-version = "py310"
# THE RULE SET IS DECLARED, and that is the whole repair rather than a
# preference. Until 2026-09-09 this table set only `line-length` and
# `target-version`, so the ACCEPTANCE was whatever ruff's default happened to
# be -- and the tree read green only because `uv.lock` froze ruff at 0.15.22.
# Upgrading to 0.16.6 turned up 148 findings in code nobody had touched, all of
# them new rules rather than new defects: 0.16 widened the default set to
# include whole families (YTT, ASYNC, PL, ISC, C4, UP, B, SIM, FURB, ...). An
# undeclared `select` means every ruff release silently redefines what "clean"
# means, which is exactly how a formatter gate went red unseen.
#
# WHAT IS HERE AND WHY. The historical default (`E4`, `E7`, `E9`, `F`), plus
# `I` because this tree already keeps its imports sorted, plus `RUF100` so a
# `noqa` that has stopped meaning anything is caught rather than left as
# decoration.
#
# WHAT IS NOT HERE, MEASURED RATHER THAN ASSUMED. `S` (bandit) reports **2657**
# `S101` on a test suite whose every assertion is an `assert`, and `S603`
# reports **19** subprocess calls of which one was ever marked -- selecting it
# would buy 18 new suppressions and no defect. The remaining families the 0.16
# default adds are a real question and a separate one: they are worth adopting
# deliberately, not inside a version-pin commit, and the number to start from
# is the 148 above.
[tool.ruff.lint]
select = ["E4", "E7", "E9", "F", "I", "RUF100"]
# ONE FILE IS EXEMPT, and it is a fence rather than a judgement about the code.
# `tools/okf_consume_measure.py` is a measurement instrument that published
# figures were produced with, and the order that authorised this cleanup fenced
# it explicitly: it is RUN, not edited, so its bytes stay as the numbers were
# taken. Its three findings are a stale `noqa: E402` twice over and an import
# order -- none of them a defect, all of them the same churn this upgrade
# produced everywhere else, and all of them to be cleaned the next time the
# fence is lifted. Named here rather than left to make the gate red for a
# reason nobody could see.
[tool.ruff.lint.per-file-ignores]
"tools/okf_consume_measure.py" = ["I001", "RUF100"]
# MARKDOWN IS NOT FORMATTED, and this is a decision the 0.16 upgrade forced.
# ruff 0.16 formats fenced Python inside markdown. Two files here would change
# under it, and both are RECORDS rather than source: `README.md`'s call example
# and `docs/2026-09-08-blindsone-below-k-k2.md`'s QUOTATION of `COST_VOCABULARY`
# as it stood when that measurement was taken. Reformatting a quotation makes it
# stop being one, and this repository publishes reproduction blocks that a
# reader is meant to be able to compare against what was run. The formatter's
# job here is Python source; `ruff format --check .` is part of the acceptance
# and stays so, over `.py`.
[tool.ruff.format]
exclude = ["*.md"]
[tool.mypy]
strict = true
python_version = "3.10"
# llm-ingestion-guard ships no py.typed marker, so its symbols arrive as Any.
# The adapter coerces every value it carries across the seam to a concrete
# type, which is what keeps --strict meaningful on this side of it.
[[tool.mypy.overrides]]
module = ["llm_ingestion_guard", "llm_ingestion_guard.*"]
ignore_missing_imports = true
# `pypandoc` ships no py.typed marker either. Only `_pandoc.py` imports it, and
# every value it hands back is coerced to `str`/`Path` there before it reaches
# the rest of the package -- the same discipline as the guard adapter above.
[[tool.mypy.overrides]]
module = ["pypandoc", "pypandoc.*"]
ignore_missing_imports = true
# `rapidocr` ships no py.typed marker either, and it is behind an OPTIONAL
# group -- so on a machine without that group installed the import does not
# resolve at all. Only `_ocr_reader` imports it, and the only value that
# crosses back is coerced to `str` there, the same discipline as the two
# overrides above.
[[tool.mypy.overrides]]
module = ["rapidocr", "rapidocr.*"]
ignore_missing_imports = true
# Install CHANNEL for the guard, which is not on a package index yet. It is
# uv-specific, and it reaches further than a dev-only setting: a consumer
# installing this package from git WITH UV picks the guard up from this tag
# automatically, because uv reads this file when it builds from the source
# tree. Measured against an empty cache 2026-07-25, 2026-08-20, and
# 2026-08-21 on uv 0.9.8. The 08-21 run also measured the TRANSITIVE form: a
# separate consumer project naming only this package still resolves the guard
# from the entry below, because this package reaches it as a git source.
#
# That source is the whole reach. A wheel carries Requires-Dist and nothing
# else, so this entry cannot survive an index install — and while the guard is
# off-index, removing it would break the one-command uv path the README
# documents.
#
# pip does not read it at all: it resolves [project.dependencies] alone and
# fails with "No matching distribution found for llm-ingestion-guard" until
# the guard is installed from its own tag first (README; measured 2026-08-21,
# both the failure and the two-command recovery).
#
# Either way the range above stays the pin, and the pin is per-tree: a wheel
# built from THIS tree carries `Requires-Dist: llm-ingestion-guard<2.0,>=1.2`,
# measured 2026-08-23 against the built wheel. The `<0.4,>=0.3` this comment
# carried before was the `v0.3.4` tag's range — still true of that tag, never
# true of this tree. Reading a range off one and installing it against the
# other is the one combination that fails.
[tool.uv.sources]
llm-ingestion-guard = { git = "https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git", tag = "v1.4.0" }