The prior measurement (docs/2026-09-08-blindsone-below-k-k2.md SS 3) found that
the budget, not the ranking, is the second lock on a mandate-shaped cost
question -- and that the same mechanism was a REGRESSION on the question that
works: raising `--k` to 16 evicted the gold concept, because the exact knapsack
maximises a SUM of fused scores and has no opinion about rank, so twenty small
excerpts out-value one that costs 56.5 % of the budget.
Measured here on the same 629-concept bundle, with the three known-positive
figures from `4c699fd` reproduced first:
- Corpus distribution, denominator 629: median excerpt 857 B, max 223 391 B,
3 concepts over the limit alone.
- Candidate rule (b), a corpus-derived budget, is FALSIFIED by two numbers: two
defensible derivations are 49x apart on the same corpus, the small one turns
the gold concept into `over_budget_alone` (13 refusals against 2), the large
one changes nothing at the default k. A budget is the consumer's constraint,
not a property of the corpus; `--limit` already belongs to the caller.
- Built instead, behind `--reserve-top-rank` (default OFF): the top-ranked
candidate gets its bytes before the pack runs, AFTER the `over_budget_alone`
pre-exclusion and never before, and the payload declares `budget.reserved`.
- It fixes the eviction: k=16 and k=24 deliver the gold concept at rank 1,
costing one and two excerpts, and 20.4 % / 27.3 % FEWER o200k tokens.
- It changes the delivered list in 2 of 24 measured combinations -- both of them
that eviction. In the other 22 the list, its order and `spent` are identical.
- It does NOT close the mandate-shaped blind spot: that concept ranks 10, not 1.
The one delivering command is `--cost-vocabulary --k 12 --limit 160000`
(62 149 tokens against 58 401), and that is a consumer's decision.
11 new tests (RED first), 7 mutations 7 red with an unmutated negative control
green before and after; two of the seven survived the first test set and the
tests were strengthened. Default payload byte-identical, both goldens unchanged.
Report: docs/2026-09-08-blindsone-laas2-budsjett-k2.md
Suite 1279 green, mypy --strict clean over 28 files, ruff clean.
Co-Authored-By: Claude <claude-opus-5>
457 lines
23 KiB
Markdown
457 lines
23 KiB
Markdown
# llm-ingestion-okf
|
||
|
||
Shared OKF (Open Knowledge Format) ingestion library: spec-based connectors, bundle inbox, and external-bundle import. Security delegated to llm-ingestion-guard.
|
||
|
||
Status: phases 1–3 are implemented. Phase 1 (spec-based ingestion) covers
|
||
manifest validation, the `file`/`sql`/`http` connectors, deterministic
|
||
materialization, index generation, and the golden fixture suite under
|
||
`examples/`. Phase 2 adds the bundle inbox (`process_inbox`) and
|
||
external-bundle import (`import_bundle`), both against an **injected** persist
|
||
gate, with `llm_ingestion_okf.guard_adapter` wiring that gate to the real
|
||
guard (see below). Phase 3 makes the bundle contract configurable, so types,
|
||
layers, frontmatter sets, index shape, and reserved-file policy are carried by
|
||
a profile rather than by constants (see [Upstream OKF
|
||
versions](#upstream-okf-versions)). Binary extraction runs behind the
|
||
optional `[extract]` extra: `pdf` through a PDF parser, and five office
|
||
formats through a vendored document converter. Three of those five office
|
||
rows are **unmeasured** — see [Binary extraction](#binary-extraction). Phase 4
|
||
(the Node half) is planned (see `docs/plan/`).
|
||
|
||
## Install
|
||
|
||
Python 3.10+. Neither this package nor the guard it depends on is on a package
|
||
index yet. With uv, one command is enough:
|
||
|
||
```
|
||
uv pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
|
||
```
|
||
|
||
uv resolves the guard on its own, because it reads the `[tool.uv.sources]`
|
||
entry in the `pyproject.toml` **of the tag it is installing**, and `v0.4.0`
|
||
points that entry at the guard tag below. Measured 2026-07-25 and re-measured
|
||
2026-08-20 with an empty `uv` cache; both runs installed
|
||
`llm-ingestion-guard==0.2.0` + `llm-ingestion-okf==0.4.0` and imported clean.
|
||
|
||
With plain pip, the transitive git dependency does not resolve on its own —
|
||
**install the guard first**, or installing this package fails with
|
||
`No matching distribution found for llm-ingestion-guard`:
|
||
|
||
```
|
||
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v0.2.0"
|
||
pip install "llm-ingestion-okf @ git+https://git.fromaitochitta.com/open/llm-ingestion-okf.git@v0.4.0"
|
||
```
|
||
|
||
The guard tag is paired to the okf tag, not to this branch: `v0.4.0` declares
|
||
`llm-ingestion-guard>=0.2,<0.3`, which `v0.2.0` satisfies and later guard tags
|
||
do not. `main` has since moved its own pin to `>=1.2,<2.0` (see
|
||
[Requirements](#requirements)); that pin reaches you in the next stable tag,
|
||
not in the commands above. Reading a pin off this branch and installing it
|
||
against `v0.4.0` is the one combination that fails.
|
||
|
||
`v0.4.0` is the current stable tag. `v0.5.0a2` is a pre-release for the named
|
||
OKF v0.2 pilot set only; pin it only if you are one of them (see
|
||
[Upstream OKF versions](#upstream-okf-versions)).
|
||
|
||
## Build
|
||
|
||
Installing the package installs one command. A folder of documents in, an OKF
|
||
bundle out:
|
||
|
||
```
|
||
okf build ./documents --bundle ./bundle --bundle-id my-bundle --okf-version 0.2
|
||
```
|
||
|
||
It walks the folder recursively, proposes a segmentation for each document with
|
||
the mechanical rules, replays those proposals through the bundle inbox, writes
|
||
the bundle and its `log.md`, and prints the run's numbers. Every proposal is
|
||
marked `PROPOSED` and `adjudicated: false` — the command segments nothing a
|
||
human has approved, and says so in the artifact.
|
||
|
||
The last line that matters is the conservation identity: `merged + coded
|
||
rejections == N`, where `N` is the folder's file count read at run time. **The
|
||
run exits non-zero when it does not hold**, and names the unaccounted files, so
|
||
a pipeline cannot mistake a partial bundle for a complete one.
|
||
|
||
Flags worth knowing: `--segments off` ingests each document as one concept and
|
||
asks for no root values; `--plans-dir` keeps the proposals instead of
|
||
discarding them; `--report` writes the full report to a file as well as stdout.
|
||
`--ingested-at` and `--proposed-at` default to `1970-01-01T00:00:00Z` rather
|
||
than the clock, so two builds of the same folder are byte-identical — a
|
||
wall-clock default would break rebuild-equals-incremental for every caller who
|
||
did not pass them.
|
||
|
||
Measured 2026-09-07 on a 43-file corpus (33 `pdf`, 5 `docx`, 2 `xlsx`, and
|
||
three files no reader accepts), one
|
||
`okf build` invocation replacing the shell loop over `tools/` that produced the
|
||
same corpus's bundle on 2026-09-03:
|
||
|
||
| figure | value |
|
||
|---|---|
|
||
| `N` (folder file count, computed) | 43 |
|
||
| merged | 39/43 |
|
||
| coded rejections | 4/43 (`extractor_unknown` 3, `extractor_empty_pdf` 1) |
|
||
| K1b | `39 + 4 = 43 = N`, exit `0` |
|
||
| files written | 1108 |
|
||
| identical to the 2026-09-03 bundle | 1107/1108 |
|
||
| wall time | 779.43 s total, 18.126 s per file |
|
||
|
||
The one file that differs is the root `index.md`, by exactly one line: the link
|
||
to the bundle's own `log.md`, added by a commit that postdates the stored
|
||
artifact by fifteen hours. Appending that line to the stored `index.md`
|
||
reproduces the new one byte for byte. The command's own byte-identity test
|
||
compares it against the two scripts at the current commit, where the two agree
|
||
over the whole tree.
|
||
|
||
## Consume
|
||
|
||
The other direction: a bundle plus one question in, one bounded, contract-shaped
|
||
payload out.
|
||
|
||
```
|
||
python3 tools/okf_consume.py ./bundle --question "your question" --out payload.json
|
||
```
|
||
|
||
`tools/okf_consume.py` is the **pre-pass** `docs/consumption-contract.md` § 1
|
||
defines — the deterministic program that reads the bundle, ranks its concepts,
|
||
cuts them to a bounded set and emits one payload. It decides nothing about the
|
||
question; the skill that reads the payload does the judgement. It calls no
|
||
model, opens no socket, imports nothing outside the standard library and this
|
||
package, and takes no clock: the same bundle bytes and the same
|
||
`(question, k, limit, cost_vocabulary, reserve_top_rank)` produce
|
||
byte-identical output.
|
||
|
||
`--cost-vocabulary` is off by default and widens one question class: it lets a
|
||
declared list of cost/price/quantity terms bridge a question and a document that
|
||
name money with different words. The gate is the question — one naming no such
|
||
term gets byte-identical bytes either way — and what it does and does not close
|
||
is measured in `docs/2026-09-08-blindsone-below-k-k2.md`.
|
||
|
||
`--reserve-top-rank` is off by default and answers a different objection: the
|
||
budget is packed by an exact knapsack, which maximises a SUM of scores and
|
||
therefore has no opinion about rank, so a top-ranked excerpt costing a large
|
||
share of the budget is out-summed by many small ones. Measured, that made `--k`
|
||
a dial that could EVICT the concept a question was asked about. The flag gives
|
||
rank one its bytes before the pack runs — after the `over_budget_alone`
|
||
pre-exclusion, never before — and the payload then declares
|
||
`budget.reserved`. On a 629-concept corpus it changed the delivered set in 2 of
|
||
24 measured combinations, both of them that eviction:
|
||
`docs/2026-09-08-blindsone-laas2-budsjett-k2.md`.
|
||
|
||
It emits the § 8 shape — `contract`, `bundle` (`bundle_id` plus a
|
||
`sha256-tree:` content identity), `budget` (unit, instrument, limit, spent and a
|
||
validated known-positive), `denominators`, `excerpts` and `withheld` — and every
|
||
withheld concept names the rule that dropped it, from a closed set of six.
|
||
`considered == withheld + delivered` closes by construction, and the payload is
|
||
refused rather than reported when it does not.
|
||
|
||
Three exit codes, not two: **0** a payload was written, **1** the run happened
|
||
and refused (the budget admitted none of the concepts that answered the
|
||
question, or an asserted `--ref` contradicted the bytes), **2** the run did not
|
||
happen. Collapsing 2 into 1 would report an unread bundle as a failed cut.
|
||
`--ref` is an **assertion**, never an override — the identity is always computed
|
||
from the bytes, because labelling a payload with an identity its bytes do not
|
||
have is the one thing § 3.3 exists to prevent.
|
||
|
||
Check any payload against the skill that will read it:
|
||
|
||
```
|
||
python3 tools/okf_contract_check.py --skill skills/okf-consume/SKILL.md --payload payload.json
|
||
```
|
||
|
||
`skills/okf-consume/` is the first instantiated consumption skill: a filled copy
|
||
of `skills/okf-consume-template/` naming this pre-pass, with every per-corpus
|
||
hole replaced by a measured value. Measured 2026-09-07 on a 629-concept bundle,
|
||
hit@8 was **5 of 6** questions at rank 1 against a chance baseline of **1.35 of
|
||
6** — with one control that failed, and both are in
|
||
`docs/2026-09-07-okf-konsumskill-maaling.md` with the honesty limits stated.
|
||
|
||
## Implemented scope (v1)
|
||
|
||
The library provides three entry points for getting content into an OKF
|
||
bundle:
|
||
|
||
1. **Spec-based ingestion.** An implementation of the normative ingest
|
||
specification owned by `portfolio-optimiser-commons`: manifest →
|
||
`file`/`sql`/`http` connector → deterministic materialization of
|
||
`ingest-{id}.md` concept files → index generation. Zero model calls in the
|
||
run path; output is reproducible byte-for-byte against golden fixtures.
|
||
2. **Bundle inbox.** A drop directory where common file types are converted
|
||
to OKF concept files. All file-type→text extraction lives in this library:
|
||
`md`, `txt`, `csv`, `json`, and `html` are handled by the stdlib core;
|
||
`pdf` and the five office formats (`docx`, `xlsx`, `pptx`, `odt`, `rtf`)
|
||
require the optional `[extract]` extra and are rejected fail-fast without
|
||
it. Extracted text passes the security gate before anything is persisted.
|
||
The drop directory is walked **recursively**, in sorted relative-path order:
|
||
a file at any depth is ingested and records its path relative to the inbox
|
||
root as its `source_file`, while dot-directories and a bundle directory
|
||
sitting inside the inbox are skipped with a reported code.
|
||
|
||
<!-- extract-formats: .md, .txt, .csv, .json, .html, .htm, .pdf, .docx, .xlsx, .pptx, .odt, .rtf -->
|
||
3. **External bundle import.** Import and merge of third-party OKF bundles:
|
||
each concept is assessed via the security gate, and only concepts that
|
||
pass are merged, materialized, and linked into the index.
|
||
|
||
## Boundary: security is delegated
|
||
|
||
Security is owned by the sibling package
|
||
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
|
||
(pinned `>=1.2,<2.0`). The division is strict:
|
||
|
||
- **guard** answers "is this content safe to persist?" — scan, sanitize,
|
||
quarantine, fail-secure, provenance stamping.
|
||
- **this library** does the plumbing — connect a source, materialize a
|
||
deterministic OKF bundle, generate the index.
|
||
|
||
No security functionality is reimplemented here.
|
||
|
||
### What is gated today: read this before trusting a door
|
||
|
||
- **Door A (`materialize_bundle`) is ungated.** It calls nothing before
|
||
writing to disk and writes what it is given. A caller materializing
|
||
untrusted content is responsible for gating it.
|
||
- **Doors B and C (`process_inbox`, `import_bundle`) gate through an adapter
|
||
you pass in.** Each takes a `gate` argument; the flow hands it the content
|
||
and obeys the verdict, refusing to persist anything that does not clear the
|
||
guard's non-blocking floor — including a disposition it does not recognise,
|
||
and (at Door C) a concept the gate returned no verdict for. What it cannot
|
||
do is check that your adapter is a real guard: a permissive stub approves
|
||
everything, and the flow will believe it.
|
||
|
||
`llm_ingestion_okf.guard_adapter` is the adapter over the real guard, and the
|
||
only module here that imports it — importing the package itself does not:
|
||
|
||
```python
|
||
from llm_ingestion_okf import process_inbox
|
||
from llm_ingestion_okf.guard_adapter import inbox_gate
|
||
|
||
result = process_inbox(inbox_dir, bundle_dir, "2026-07-25T12:00:00Z",
|
||
okf_type="reference", gate=inbox_gate)
|
||
```
|
||
|
||
Two properties of that adapter are worth knowing before you rely on it.
|
||
It screens the **exact bytes it persists** — the guard's `prepare_input`
|
||
bookend prepares text for a model call, which this library never makes, so
|
||
only `screen_output` is used and the screened string is the written string.
|
||
And it **refuses rather than repairs**: a file carrying an invisible
|
||
zero-width or bidi character is rejected, not silently stripped and written.
|
||
Door B screens under the untrusted-upload policy, so any finding at all is
|
||
held back rather than persisted.
|
||
|
||
This section is stated plainly because earlier wording ("calls the guard at
|
||
every persist gate") described the intended end state in the present tense,
|
||
and a consumer reasonably read it as safe-by-default.
|
||
|
||
## Roadmap
|
||
|
||
The library is built in four phases so that every known OKF surface in the
|
||
ecosystem is eventually covered. Each phase has a detailed plan with
|
||
verification criteria:
|
||
|
||
1. Spec-based ingestion (Python) with byte-exact golden fixtures —
|
||
[plan](docs/plan/phase-1-door-a.md).
|
||
2. Bundle inbox and external-bundle import (Python), guard-gated —
|
||
[plan](docs/plan/phase-2-doors-b-c.md).
|
||
3. Configurable bundle contract (types, layers, frontmatter sets, index
|
||
shape, and reserved-file policy as configuration), enabling stricter
|
||
bundle profiles such as `strict-v1` —
|
||
[plan](docs/plan/phase-3-configurable-contract.md).
|
||
4. A `node/` half: a zero-dependency Node/ESM package (importable and
|
||
CLI-invokable, vendored per consumer) providing bundle checking, index
|
||
generation, inbox processing, and document conversion for the OKF
|
||
second-brain plugin ecosystem. The Python and Node halves share the OKF
|
||
contract and fixture suite, not code —
|
||
[plan](docs/plan/phase-4-node-half.md).
|
||
|
||
## Upstream OKF versions
|
||
|
||
The library targets the current latest version of Google's OKF. Support is
|
||
**additive** — a new upstream version arrives as a new profile, never as a
|
||
migration of an existing one — so an *upstream* release does not change the
|
||
bytes an existing profile emits.
|
||
|
||
That guarantee is about upstream, and one profile tracks a second contract as
|
||
well. `DEFAULT` states the ingest-spec owned by `portfolio-optimiser-commons`,
|
||
so when they change that spec, `DEFAULT` follows them. It happened on
|
||
2026-08-09: `generated` moved from `true` to
|
||
`{ by: process:okf-ingest, at: <ingested_at> }`, one changed line per generated
|
||
file. Upgrading across it costs a re-run and nothing more — a profile still
|
||
recognises bundles stamped by earlier versions, so re-running writes in place
|
||
instead of refusing. `DEFAULT` remains OKF v0.1 on every axis upstream owns.
|
||
|
||
| Profile | Contract | Status |
|
||
|---|---|---|
|
||
| `DEFAULT` | commons' ingest-spec layer (OKF v0.1 semantics) | stable |
|
||
| `STRICT_V1` | a consumer's ratified v0.1 contract | stable |
|
||
| `OKF_V0_2` | OKF v0.2 | **provisional**, pre-release only |
|
||
| `STRUCTURED_V1` | `DEFAULT` plus a faceted, derived index | stable |
|
||
| `OKF_LATEST` | alias for the latest version supported as *stable* | currently `DEFAULT` |
|
||
|
||
`STRUCTURED_V1` is `DEFAULT` in every respect but the index. Under it, Door B
|
||
derives each dropped document's title, number, hierarchy and cross-references,
|
||
writes them into the concept's own frontmatter, and carries them into the index
|
||
entry — so a consumer can reason over the bundle rather than only look things
|
||
up in it. Every inferred field is named in a `derived` list, because an
|
||
unmarked heuristic is worse than no heuristic: the consumer cannot know when to
|
||
doubt it. A pointer to a document not dropped yet is rendered `N200?` rather
|
||
than omitted, since a bundle is built up over several drops and an absence that
|
||
leaves no trace is the dangerous kind. Carrying the metadata costs index
|
||
characters — roughly 3x to 6x the flat index, depending on how many facets the
|
||
profile names — and the facet key set is the dial. Design record and
|
||
measurements: [`docs/plan/structure-derivation.md`](docs/plan/structure-derivation.md).
|
||
|
||
`OKF_V0_2` ships first as a pre-release to a named pilot set and may change on
|
||
their feedback without a deprecation cycle. Pin the versioned constant rather
|
||
than `OKF_LATEST` unless you have explicitly opted into tracking; `OKF_LATEST`
|
||
moves at general availability, which is a deliberate release event rather than
|
||
a side effect of an upgrade.
|
||
|
||
Selecting a profile is keyword-only, so existing call sites are unaffected:
|
||
|
||
```python
|
||
materialize_bundle(manifest, bundle_dir, ingested_at, profile=OKF_V0_2)
|
||
```
|
||
|
||
A bundle may declare the version it targets. OKF v0.2 §12 makes this a MAY, and
|
||
puts the declaration in the bundle-root `index.md`'s frontmatter block. The
|
||
profile names the key; the **caller supplies the value**, because that value
|
||
tracks the upstream version and is not this library's to decide:
|
||
|
||
```python
|
||
materialize_bundle(
|
||
manifest, bundle_dir, ingested_at,
|
||
profile=OKF_V0_2,
|
||
root_frontmatter_values={"okf_version": "0.2"},
|
||
)
|
||
```
|
||
|
||
Omit the argument and no frontmatter block is written. Offering a key the
|
||
profile does not name is refused before anything is written to disk.
|
||
|
||
### Attested computations (v0.2 §10)
|
||
|
||
`OKF_V0_2` supports the `Attested Computation` type as a **format**: its five
|
||
contract fields — `runtime`, `parameters`, `computation`, `executor`,
|
||
`attester` — are emitted in canonical position, judged, and round-tripped.
|
||
`runtime` is required for that type and for no other, which the profile
|
||
expresses through `FrontmatterSchema.required_by_type`; a type the mapping does
|
||
not name carries no extra requirement, because §14 forbids a consumer to reject
|
||
on an unknown `type`.
|
||
|
||
Nothing here executes a computation or checks an attestation. Upstream defers
|
||
the receipt and verdict wire formats, so there is no contract to implement, and
|
||
the question an attestation answers — was this value produced the sanctioned
|
||
way — is not this library's. It re-enters scope when upstream specifies the
|
||
protocol.
|
||
|
||
On the import side, a third-party concept may name an `executor` or `attester`
|
||
resource pointing at executable code. Door C imports the **pointer** and never
|
||
the code — it writes concepts verbatim and skips every non-`.md` file — so such
|
||
a reference may not resolve, or may resolve to a file the destination tree
|
||
already holds under that path. Each one is reported in
|
||
`ImportResult.unverified_references`; the concept still merges, because §14
|
||
forbids rejecting a bundle over a broken cross-link while §10.5 asks a consumer
|
||
to surface rather than silently drop. The report names the pointer key, not the
|
||
resource it points at: recovering the resource needs the structured reader.
|
||
|
||
One limit worth knowing before you write such a concept: §10.2 presents
|
||
`executor` and `attester` as nested block mappings, and this library's
|
||
frontmatter parser is line-oriented. It reads inline **flow** mappings
|
||
(`executor: { resource: …, receipt: [ … ] }`) as opaque values that round-trip
|
||
unchanged, but it cannot read the block form — two block mappings that both
|
||
carry a `resource` collapse into one namespace and the first is lost. Write the
|
||
flow form; both are valid YAML, and a real YAML consumer recovers the same
|
||
structure from either.
|
||
|
||
## Non-goals
|
||
|
||
- Verdict/feedback machinery from the method specification (stays in the
|
||
consuming repositories).
|
||
- Embedding- or retrieval-layer functionality.
|
||
- Security functionality, in either runtime — that is always
|
||
`llm-ingestion-guard`'s domain.
|
||
|
||
## Requirements
|
||
|
||
Python 3.10+, and exactly one runtime dependency — the security boundary,
|
||
`llm-ingestion-guard>=1.2,<2.0`. Everything else is stdlib. The commands are
|
||
under [Install](#install); what follows is why they look the way they do.
|
||
|
||
From a checkout, the test suite runs with:
|
||
|
||
```
|
||
.venv/bin/python -m pytest
|
||
```
|
||
|
||
The suite is the verification surface for everything above: 596 tests, run on
|
||
2026-08-21 against this branch with the `[extract]` extra installed. Without
|
||
the extra the same suite is 589 passed and 7 skipped, measured the same day:
|
||
the seven cover the parser path, and the tests holding the fail-fast rejection
|
||
for an uninstalled extra run in both. It is not shipped in an installed
|
||
distribution — `tests/` lives at the repository root, so this command needs a
|
||
clone rather than a `pip install`.
|
||
|
||
A git URL is a PEP 508 direct reference and pins one exact tag, so it is an
|
||
install-time *channel*, not the pin: the range above stays the declared
|
||
dependency — a wheel built from this branch carries `Requires-Dist:
|
||
llm-ingestion-guard<2.0,>=1.2`, measured 2026-08-23 — and resolves normally
|
||
once the package index exists. A wheel built from a *tag* carries that tag's
|
||
range instead, which is why the install commands pair tag with tag.
|
||
|
||
### Binary extraction
|
||
|
||
The optional `[extract]` extra ships two things: `pdfplumber` (MIT) for `pdf`,
|
||
and `pypandoc-binary` for five office formats. It is opt-in because it pulls
|
||
binary wheels, which the default install must never do — the single runtime
|
||
dependency rule covers the default install and this extra sits outside it.
|
||
|
||
The converter **binary travels inside the wheel** and is resolved by path
|
||
rather than found on `PATH`, with its version asserted against a pin. A host
|
||
carrying a different converter is refused, not silently used: extraction is
|
||
deterministic within a converter version and not across one.
|
||
|
||
| Format | Reader | Evidence |
|
||
|---|---|---|
|
||
| `pdf` | `pdfplumber` | measured |
|
||
| `docx` | converter | measured |
|
||
| `xlsx` | converter | measured |
|
||
| `pptx` | converter | **unmeasured** |
|
||
| `odt` | converter | **unmeasured** |
|
||
| `rtf` | converter | **unmeasured** |
|
||
|
||
**`unmeasured` means what it says.** The corpus this work was measured on
|
||
contains **zero** `pptx`, `odt` and `rtf` files, so those three rows work by
|
||
construction and have never been checked against a document anyone wrote.
|
||
They are not known to be broken; they are not known to be right either, and
|
||
the distinction is the point.
|
||
|
||
**What stays out.** `.doc` (Word 97) is not supported — the converter does not
|
||
read it. Rastered or scanned PDFs are refused rather than persisted as empty
|
||
concepts, because this library does not do OCR. Drawn content — figures,
|
||
diagrams, shapes — does not survive extraction in any format here, and every
|
||
extraction says so with a warning. Structured table recovery is out of scope.
|
||
|
||
Request it by appending `[extract]` to the package name in whichever install
|
||
command from [Install](#install) you are using — this package is not on an
|
||
index, so a bare `pip install 'llm-ingestion-okf[extract]'` does **not** work
|
||
today, and the error message naming that command is written for the day it
|
||
does. The extra is unreleased: it reaches a consumer through a tag that
|
||
contains it, and no such tag exists yet.
|
||
|
||
Two properties of the extra are worth knowing before depending on its output:
|
||
|
||
- **Extracted text is pinned to an exact parser version.** `pdfplumber` pins
|
||
`pdfminer.six==20260107` exactly, and `pdfminer.six` ships date-stamped
|
||
releases with no stability contract. Extraction is deterministic within a
|
||
parser version and not guaranteed across one, so a golden fixture built on
|
||
extracted PDF text is a fixture migration away from any parser upgrade.
|
||
- **Text extraction recovers text, and nothing that is drawn.** Figures,
|
||
diagrams and images have no text to recover — only their captions survive —
|
||
so a bundle built from drawn documents is incomplete by construction. The
|
||
library says so itself: every `pdf` extraction emits an `ExtractionWarning`.
|
||
Structured table recovery is separately out of scope; PDFs enter as prose.
|
||
|
||
The planned Node half targets Node/ESM with zero npm dependencies.
|
||
|
||
## License
|
||
|
||
MIT — see [LICENSE](LICENSE).
|