llm-ingestion-okf/README.md
2026-07-16 11:01:09 +02:00

80 lines
3.4 KiB
Markdown

# llm-ingestion-okf
Shared ingestion library for OKF (Open Knowledge Format) bundles.
Status: pre-implementation. This repository currently contains the project
scaffold, the scope definition, and detailed plans for each phase (under
`docs/plan/`); no pipeline code has been written yet.
## Planned scope (v1)
The library provides three entry points for getting content into an OKF
bundle:
1. **Spec-based ingestion.** An implementation of the normative ingest
specification owned by `portfolio-optimiser-commons`: manifest →
`file`/`sql`/`http` connector → deterministic materialization of
`ingest-{id}.md` concept files → index generation. Zero model calls in the
run path; output is reproducible byte-for-byte against golden fixtures.
2. **Bundle inbox.** A drop directory where common file types are converted
to OKF concept files. All file-type→text extraction lives in this library:
`md`, `txt`, `csv`, `json`, and `html` are handled by the stdlib core;
`pdf`, `docx`, and `xlsx` require the optional `[extract]` extra and are
rejected fail-fast without it. Extracted text passes the security gate
before anything is persisted.
3. **External bundle import.** Import and merge of third-party OKF bundles:
each concept is assessed via the security gate, and only concepts that
pass are merged, materialized, and linked into the index.
## Boundary: security is delegated
Security is owned by the sibling package
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
(pinned `>=0.2,<0.3`). The division is strict:
- **guard** answers "is this content safe to persist?" — scan, sanitize,
quarantine, fail-secure, provenance stamping.
- **this library** does the plumbing — connect a source, materialize a
deterministic OKF bundle, generate the index — and calls the guard at every
persist gate (`prepare_input`/`screen_output` for extracted text,
`okf.import_bundle` for received bundles).
No security functionality is reimplemented here.
## Roadmap
The library is built in four phases so that every known OKF surface in the
ecosystem is eventually covered. Each phase has a detailed plan with
verification criteria:
1. Spec-based ingestion (Python) with byte-exact golden fixtures —
[plan](docs/plan/phase-1-door-a.md).
2. Bundle inbox and external-bundle import (Python), guard-gated —
[plan](docs/plan/phase-2-doors-b-c.md).
3. Configurable bundle contract (types, layers, frontmatter sets, and
reserved-file policy as configuration), enabling stricter bundle
profiles such as `strict-v1`
[plan](docs/plan/phase-3-configurable-contract.md).
4. A `node/` half: a zero-dependency Node/ESM package (importable and
CLI-invokable, vendored per consumer) providing bundle checking, index
generation, inbox processing, and document conversion for the OKF
second-brain plugin ecosystem. The Python and Node halves share the OKF
contract and fixture suite, not code —
[plan](docs/plan/phase-4-node-half.md).
## Non-goals
- Verdict/feedback machinery from the method specification (stays in the
consuming repositories).
- Embedding- or retrieval-layer functionality.
- Security functionality, in either runtime — that is always
`llm-ingestion-guard`'s domain.
## Requirements
Python 3.10+. The core has zero runtime dependencies. The planned Node half
targets Node/ESM with zero npm dependencies.
## License
MIT — see [LICENSE](LICENSE).