Scope settled 2026-07-16: implements portfolio-optimiser-commons ingest-spec (commons keeps spec authorship), Python 3.10+ stdlib-only core, security delegated to llm-ingestion-guard at persist gates. Doors: spec-based ingestion, bundle inbox (md/txt/csv/json/html core, binary formats behind [extract]), external bundle import. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QeqhJpYQyghASjiJo5EhGg
60 lines
2.5 KiB
Markdown
60 lines
2.5 KiB
Markdown
# llm-ingestion-okf
|
|
|
|
Shared ingestion library for OKF (Open Knowledge Format) bundles.
|
|
|
|
Status: pre-implementation. This repository currently contains the project
|
|
scaffold and scope definition only; no pipeline code has been written yet.
|
|
|
|
## Planned scope (v1)
|
|
|
|
The library provides three entry points for getting content into an OKF
|
|
bundle:
|
|
|
|
1. **Spec-based ingestion.** An implementation of the normative ingest
|
|
specification owned by `portfolio-optimiser-commons`: manifest →
|
|
`file`/`sql`/`http` connector → deterministic materialization of
|
|
`ingest-{id}.md` concept files → index generation. Zero model calls in the
|
|
run path; output is reproducible byte-for-byte against golden fixtures.
|
|
2. **Bundle inbox.** A drop directory where common file types are converted
|
|
to OKF concept files. All file-type→text extraction lives in this library:
|
|
`md`, `txt`, `csv`, `json`, and `html` are handled by the stdlib core;
|
|
`pdf`, `docx`, and `xlsx` require the optional `[extract]` extra and are
|
|
rejected fail-fast without it. Extracted text passes the security gate
|
|
before anything is persisted.
|
|
3. **External bundle import.** Import and merge of third-party OKF bundles:
|
|
each concept is assessed via the security gate, and only concepts that
|
|
pass are merged, materialized, and linked into the index.
|
|
|
|
## Boundary: security is delegated
|
|
|
|
Security is owned by the sibling package
|
|
[`llm-ingestion-guard`](https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security)
|
|
(pinned `>=0.2,<0.3`). The division is strict:
|
|
|
|
- **guard** answers "is this content safe to persist?" — scan, sanitize,
|
|
quarantine, fail-secure, provenance stamping.
|
|
- **this library** does the plumbing — connect a source, materialize a
|
|
deterministic OKF bundle, generate the index — and calls the guard at every
|
|
persist gate (`prepare_input`/`screen_output` for extracted text,
|
|
`okf.import_bundle` for received bundles).
|
|
|
|
No security functionality is reimplemented here.
|
|
|
|
## Non-goals (v1)
|
|
|
|
- The `strict-v1` bundle profile used by `claude-code-llm-wiki` (candidate
|
|
for v2; requires a configurable bundle contract).
|
|
- A Node/JavaScript port for the OKF second-brain plugin ecosystem, which is
|
|
deliberately Node/ESM zero-dependency. See CLAUDE.md for the recorded
|
|
decision criteria.
|
|
- Verdict/feedback machinery from the method specification (stays in the
|
|
consuming repositories).
|
|
- Embedding- or retrieval-layer functionality.
|
|
|
|
## Requirements
|
|
|
|
Python 3.10+. The core has zero runtime dependencies.
|
|
|
|
## License
|
|
|
|
MIT — see [LICENSE](LICENSE).
|