llm-ingestion-okf/docs/plan/phase-2-doors-b-c.md
Kjell Tore Guttormsen b0ad2aedfb docs(plan): add detailed phase 1-4 implementation plans
Phase 1 (Door A): ingest-spec v1 implementation + §11 golden fixtures,
pinned to the current spec revision, with TDD order and byte-exact
verification criteria. Phase 2 (Doors B/C): guard-gated inbox and
external bundle import. Phase 3: configurable bundle contract via a
profile object, default profile locked to the golden suite. Phase 4:
zero-dep Node half, coordination-first with owner sign-offs.

Each plan carries explicit key assumptions with tests and a
verification section.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QeqhJpYQyghASjiJo5EhGg
2026-07-16 10:58:08 +02:00

7 KiB

Phase 2 plan — Door B (bundle inbox) + Door C (external bundle import)

Status: approved roadmap phase (see CLAUDE.md); details settled here before code. Depends on: Phase 1 (materialization + index primitives are reused, never duplicated). This phase adds the library's first — and only permitted — runtime dependency: llm-ingestion-guard>=0.2,<0.3.

Goal

Two guard-gated persist paths on top of the Phase 1 plumbing:

  • Door B — bundle inbox: convert operator-dropped files into OKF concept files. All file-type→text extraction lives here (the guard is text-only).
  • Door C — external bundle import: assess a third-party OKF bundle per concept via the guard's okf.import_bundle; merge, materialize, and index only the non-rejected concepts.

The boundary rule is absolute: the guard answers "is this content safe to persist?"; this library only connects, extracts, materializes, and indexes. No scanning, sanitizing, or quarantine logic is implemented here.

Door B — deliverables

  1. Extraction registry — file extension → extractor:
    • Core (stdlib only): md/txt passthrough, csv → markdown table (reuses the Phase 1 renderer), json → verbatim inside a fenced block, html → text via html.parser.
    • Optional ([extract] extra only): pdf, docx, xlsx. Without the extra installed those types are rejected fail-fast with a typed error naming the extra — never a silent skip, never a bundled parser in core.
  2. Inbox flow — an explicit operator command (mirrors the spec §9 HITL rule: never automatic, no watcher, no scheduler): process_inbox(inbox_dir, bundle_dir, ingested_at, ...) -> InboxResult. Per file: read bytes → extract text → guard gate → materialize concept → index link. ingested_at explicit as in Phase 1; no wall-clock anywhere.
  3. Guard gate (persist gate #1): extracted text passes through the guard (prepare_input/screen_output bookends with an upload-grade policy) before any write. Disposition handling: blocking dispositions → the file is NOT persisted and is reported in InboxResult; quarantine-grade dispositions → reported for operator review, not persisted (a quarantine directory is an extension point, not v1); pass/warn → persisted with the report attached to the result.
  4. Inbox provenance layer — machine-generated files carry the honesty marker, analogous to spec §7: type, title, source_file (inbox-relative name), source_sha256 (of the original bytes), ingested_at, generated: true. Filenames namespaced inbox-{slug}.md with the slug reduced to the Phase 1 id grammar — disjoint from index.md, promoted-verdict-*, and ingest-*. The verdict reservation and the pre-mutation collision gate apply unchanged.

Door C — deliverables

  1. Import flow — explicit operator command: import_bundle(source_dir, bundle_dir, ingested_at, *, origin, channel) -> ImportResult. Reads the external bundle as {relative path -> text}, hands it to the guard's okf.import_bundle(bundle, origin=…, channel=…, allow_reserved=False), and merges ONLY concepts whose per-concept result carries no error.
  2. Merge + index: accepted concepts go through the Phase 1 staging/collision gate and index primitives (idempotent linking, curated content preserved). Rejected concepts are reported per concept with the guard's reason — never partially written.
  3. Import report: the guard's log body (BundleResult.log()) is returned to the caller in ImportResult; persisting it into the bundle is NOT done in v1 (reserved-file policy differs per consumer and becomes configurable in Phase 3).

Design decisions fixed by this plan

  • Trust follows origin, never channel (guard doctrine) — callers must supply both; no defaults.
  • Extraction failure (corrupt file, missing extra) is a typed error on that file; the run continues with the remaining files and the result reports per-file outcomes. A guard FAIL-grade disposition on one file likewise never aborts the whole inbox run.
  • Determinism: same inbox bytes + same ingested_at + same guard version ⇒ byte-identical persisted output. Guard verdicts are part of that input surface.

TDD order

  1. Extraction registry: core types (fixtures per type), fail-fast on unknown extensions, fail-fast on optional types without [extract].
  2. Inbox provenance rendering + filename slugging (pure functions).
  3. Door B flow against a stub guard (in tests only — a test double standing in for the pinned guard API, so persist/reject branches are exercised deterministically; the real guard is exercised in integration tests).
  4. Door B integration with the real guard: benign fixture persists; a fixture the guard fails-secure on is not persisted.
  5. Door C flow: mixed fixture bundle (accepted + rejected concepts) → only accepted concepts merged; collision and curated-preservation semantics reused from Phase 1 tests.

Key assumptions (each with its test)

# Assumption Test
B1 Guard 0.2 API matches the pinned surface (prepare_input, screen_output, okf.import_bundle, disposition enum) Import-and-signature smoke test that fails on upgrade drift
B2 Guard is installable where this library's CI runs Resolve the distribution channel with the operator before implementation starts (risk: not yet decided)
B3 html.parser-based extraction is adequate for v1 Golden fixtures for representative HTML; anything richer is explicitly out of scope
B4 Guard verdicts are deterministic for fixed input + version Same fixture run twice → identical InboxResult

Non-goals

  • Any security decision logic (scan, sanitize, lexicons, fencing) — guard only.
  • Quarantine storage/workflow and re-scan over time — extension points; the guard's companion-expectations list is revisited in Phase 3.
  • Model-assisted enrichment of inbox content — never in the run path.
  • Watchers, schedulers, implicit triggers.

Verification

  1. pytest green with the guard installed; mypy --strict src/; ruff check . and ruff format --check . clean.
  2. pdf file without [extract] → typed error naming the extra; with pip install .[extract] the same file extracts (integration test, may be skipped in environments without the extra — but must run in CI).
  3. Persist-gate proof: a fixture the guard fails-secure on results in ZERO new files in the bundle and an explicit rejection entry in the result (named test asserts the bundle directory is byte-identical before/after).
  4. Door C partial-merge proof: mixed fixture → accepted concepts present, rejected absent, index links only for accepted, curated lines untouched.
  5. Phase 1 golden suite still passes byte-for-byte (no regression from reuse).
  6. Grep-gate: grep -rn "sanitize\|quarantine\|lexicon" src/ shows no local security reimplementation (guard imports only).
  7. pyproject.toml runtime dependencies == exactly llm-ingestion-guard>=0.2,<0.3.