# llm-ingestion-guard ![Version](https://img.shields.io/badge/version-0.2.0-blue) ![Status](https://img.shields.io/badge/status-alpha-orange) ![Python](https://img.shields.io/badge/python-3.10%2B-purple) ![Tests](https://img.shields.io/badge/tests-522_passing-green) ![License](https://img.shields.io/badge/license-MIT-lightgrey) **Write-time ingestion is the trust boundary that query-time guardrails structurally cannot see.** When untrusted content passes through an LLM enrichment/summarization/extraction step into a *persisted* artifact — a RAG corpus, a knowledge base, an LLM wiki — the poisoned result is read *later* by a downstream agent as **trusted context**. That agent's guardrail never sees where the content came from. The only place the provenance still exists is the write. This library packages that write-time contract — sanitize → fence → tool-less quarantined transform → per-stage capability isolation → scan-before-commit → fail-secure — as composable, stdlib-first, framework-agnostic code. It is the write-time **sibling** of query-time tools (LLM Guard, NeMo Guardrails, Rebuff, Vigil), not a competitor: those harden material as it enters the model; this hardens it as it is committed for a *later* reader. Where existing OSS tooling is mostly single-stage *detectors* — a risk verdict, with quarantine, capability isolation, scan-before-persist, and fail-secure left to the integrator — this packages the full contract as code. (Neighbours surveyed in [`docs/BRIEF.md`](docs/BRIEF.md) §11.) **Why an LLM wiki (e.g. Google OKF) needs this specifically.** OKF and second-brain formats have no schema registry, no central authority, and no signing — a bundle's claimed origin is not verifiable at the format level. So *your ingestion pipeline is the trust boundary*: provenance must be stamped by you at write time, never assumed from the format. Any pipeline ingesting external data into an agent-read store has this shape; an OKF wiki is its canonical form — which is why the guard ships a first-class OKF adapter (below). **Status:** `v0.2`, alpha. The stdlib-only core — its detector, contract, and OKF-adapter modules plus the top-level wiring — is built and tested, exercised by an end-to-end showcase and adversarial + false-positive corpora. The public API may still change. There are real limitations, stated plainly below; read them. ## Install ```bash pip install llm-ingestion-guard # stdlib-only core, zero dependencies ``` Optional ML/judge detectors live behind extras (`[ml]`, `[judge]`) and are not required — the core is deterministic and dependency-free. ## Quickstart — the two bookends The library never makes the model call itself. It gives you the two library-side halves around your own **tool-less** transform: ```python from llm_ingestion_guard import ( prepare_input, screen_output, Disposition, PRESET_USER_UPLOAD, ) prepared = prepare_input(untrusted_content) # sanitize + fence enriched = your_model(prepared.fenced) # tool-less — YOUR call decision = screen_output(enriched, PRESET_USER_UPLOAD) # scan + dispose if decision.disposition is Disposition.FAIL_SECURE: alert(gate_code=decision.reasons) # minimal payload, no content raise SystemExit # halt — never persist ``` `screen_output` fails **closed**: if the scanner itself errors on crafted input, the disposition is `FAIL_SECURE`, never a silent persist. Pass `transform_failed=True` when your model call raised or fell back — a scan hit together with a transform failure is treated as a probable forced-fallback attack and halts regardless of trust tier. Every primitive is also exported for pipelines that compose the checklist themselves — `sanitize`, `scan_lexicon`, `scan_entropy`, `scan_output`, `scan_active_content`, `neutralize`, the `decide` / `guard` disposition machinery, and the contract asserters `assert_tool_less` / `assert_credential_allowlist` / `scoped_env`. See [the end-to-end showcase](tests/test_showcase.py) for a full worked pipeline. ## OKF / LLM-wiki support (shipped) For a bundle-shaped store, the `okf` submodule sits **on top of** the format-agnostic core: it knows OKF structure (frontmatter, paths, links, `resource`, bundles) and feeds scannable regions into the same `sanitize` / `scan_output` / disposition machinery — no YAML/format awareness leaks into the core. Two ingestion modes: - **(a) Own enrichment output** — your agent writes concepts. Run the two bookends above per concept before commit. - **(b) Received external bundle** — you merge a whole third-party OKF bundle. `import_bundle` iterates concept-by-concept and runs the full per-concept gate: ```python from llm_ingestion_guard.okf import import_bundle, Origin, Channel # bundle: {concept_path -> raw document text}, e.g. {"tables/users.md": "---\n..."} result = import_bundle(bundle, origin=Origin.EXTERNAL, channel=Channel.AUTOMATIC) for c in result.concepts: if c.error: # hard reject: bad path, unsafe frontmatter, non-https resource skip(c.path) # disposition is FAIL_SECURE; do not merge this concept # result.disposition = most-severe across concepts # result.links = the cross-link graph (dangling/rejected/resolved edges) # result.log() = the log.md body (one provenance-stamped line per concept) ``` Per-concept gates: **path / reserved-name** (rejects `..` traversal and reserved `index.md`/`log.md` shadowing); **frontmatter parse-safety** — a strict, reject-by-default loader that refuses anchors, aliases, and explicit tags *by construction*, so a billion-laughs alias expansion or a `!!python/object` coercion cannot occur (it is deliberately **not** a general YAML engine, whose own features are the attack surface); **`resource` https-allowlist** (hard-rejects `data:`/`javascript:`/`file:` before commit — a reject-gate, not defang); **whole-concept scan** (frontmatter *values* + body through `scan_output`); **cross-link graph** (surfaces dangling targets, the dormant-injection signal, and rejects bundle-escaping links); **provenance stamping** (`Origin` × `Channel` → tier + disposition, emitted to `log.md`). ## The reusable contract (adopt-this checklist) The actual product is this checklist, encoded as code you wire in order: 1. **Sanitize before fence.** Strip carrier classes (zero-width, BIDI, Unicode-tag, HTML comment, `data:`) from untrusted input first. 2. **Fence untrusted input.** Spotlight-mark it in a randomized per-call delimiter; strip attacker fence markers from the payload. 3. **Tool-less transform.** Call the model with zero tools. A successful injection then has nothing to act with. 4. **Per-stage capability isolation.** The enrichment stage holds only the model key; the publish stage holds only the publish credential; no stage holds both. 5. **Treat output as data.** Parse to a frozen schema; reject on structural violation. The output never reaches a shell, git, or a filesystem path. 6. **Scan output before persist.** Run the lexicon + entropy + active-content scan over the emitted text. Verbatim-carried payloads, model-emitted instructions, and zero-click exfil carriers (EchoLeak-class markdown images/links, raw active HTML, `data:` URIs) are caught here; `neutralize` additionally defangs them, opt-in. 7. **Fail-secure on compound signals.** Injection hit + transform failure = halt + alert, never a silent verbatim commit. 8. **Minimal alert payloads.** Alert with a gate code + run ID, never content. Steps 1-2 are `prepare_input`; steps 6-7 are `screen_output`; steps 3-5 are yours; the contract asserters harden step 3-4. ## Verify what it stops (coverage matrix) Before wiring the guard into your pipeline, see — in one place — every vulnerability class it stops, and the ones it deliberately does not: ```bash python -m llm_ingestion_guard.coverage # prints the matrix; exit 0 = all as documented ``` Each row drives the real guard with a live payload and prints `class -> OWASP anchor -> expected -> observed -> verdict`. As of `v0.2`: **126 of 126 defended classes demonstrated (recall 100%)** and **4 of 4 documented gaps still hold**. The manifest ([`coverage.py`](src/llm_ingestion_guard/coverage.py)) is the single source of truth for [`tests/test_coverage_matrix.py`](tests/test_coverage_matrix.py), which asserts **total recall** across every defended class, that **every documented gap still holds** (a closed gap fails the test, forcing a doc update), and — the honesty guard — that **every lexicon pattern has a case** so the matrix cannot fall behind the lexicon. The full LLM02 secret-egress set and the container-layer front-end (CSV formula-injection, zip-slip/bomb, symlink) are asserted there too. ## Honest limitations (shipped as a control) Conceding these plainly is itself a control — it prevents the false assurance that a green scan means safe content: - **Structural unsolvability at the text layer.** Pattern/lexicon detection is bypassable in isolation; character-injection and novel phrasings evade it. The *contract* (tool-less transform, capability isolation, fail-secure) carries the security — the lexicon is defense-in-depth, not a wall. - **A lone HIGH finding in trusted prose disposes to WARN, not quarantine.** Under `PRESET_TRUSTED_SOURCE`, trust-scaling downgrades a single HIGH to WARN, and one HIGH is not "compound" (escalation needs ≥2 findings at MEDIUM+). So a HIGH injection reproduced verbatim under a *trusted* policy persists with a WARN. This is by design: if your "trusted" sources can carry attacker-influenced text, run them as untrusted (or add a quarantine floor). - **The upload preset's quarantine floor is currently vacuous.** `PRESET_USER_UPLOAD` sets `quarantine_default`, but detectors emit CRITICAL/HIGH/MEDIUM only, and under untrusted trust a MEDIUM already escalates to QUARANTINE_REVIEW — so the floor changes no outcome today. It is headroom for a future LOW/INFO finding, documented so the preset is not over-read. - **Semantic / factual poisoning is invisible** to lexicon + entropy: a false claim in clean prose carries no suspicious token. **Highest impact for a wiki.** The `grounding` module ships only a `SourceGroundingCheck` *seam* — the deterministic core does not judge semantics; a `[judge]` implementation must be plugged in. - **Adversarial-ML evasion** can survive normalization; **tokenizer mismatch** between scanner and model leaves gaps. **Latent / dormant memory poisoning** is not judgeable at write time. - **Dormant / broken-link injection** in a linked corpus (e.g. an OKF bundle): a link to a not-yet-existing target passes a per-concept write-time scan clean — the payload is planted later, when that target is written. `link_graph` surfaces the *dangling* edge as the signal, but catching the payload needs cross-write re-scan over time (the caller's disposition call). - **OKF reserved files (`index.md` / `log.md`).** In a *received* bundle these are legitimate structure, so mode-b `import_bundle` scans their body and frontmatter (an injection in a directory listing is caught) rather than path-rejecting the conformant bundle. A front-end materialising individual uploads keeps the opposite rule (`allow_reserved=False`): a reserved basename is a listing-shadow and refused. - **A document that *describes* attacks is a false positive.** Content documenting prompt-injection payloads (security notes, this project's own corpus) trips carrier-strip / fail-secure. At the text layer "*about* an attack" and "*carrying* an attack" are indistinguishable; such content needs a deliberate, explicitly escaped path, never a silent allow. - **Bilingual text trips the Cyrillic/Latin homoglyph rule.** `homoglyph:cyrillic-latin-mix` (MEDIUM) flags a Latin letter adjacent to a Cyrillic look-alike, so genuine bilingual prose → MEDIUM → under untrusted → QUARANTINE_REVIEW — a real false positive for an inbox that expects multilingual content. A calibration fix is pending. - **Insider in-place edits** by a trusted author are out of the untrusted-content threat model. - **Text-only.** The core is `text -> findings`: it parses no files (no `pypdf`/`python-docx`/archive deps). Extract text first, then scan it with the high-untrust upload provenance. OCR-embedded instructions and multimodal stego are out of scope beyond the sanitizer's character-layer stripping. - **Uploaded files: only the *extracted text* is scanned.** The dev-scoped OKF inbox showcase (`tests/test_okf_inbox_uploads.py` + `tests/inbox_frontend.py`; parsers in the `[dev]` extra, never core `dependencies`) reads `.txt`/`.md`/`.csv`/`.docx`/`.pptx`/`.xlsx`, folders and `.zip`, materializes an OKF bundle, then guards it. What survives extraction is **out of scope**: macros, OLE/embedded objects, OCR-needing images, font/render stego, encrypted files — the binary layer needs a separate scanner. The front-end owns the container threats it *can* see (zip-slip → path gate, zip-bomb → size cap, symlink refusal, CSV/XLSX formula-lead cells). **`.pdf` is a deliberate concession:** a top-level `.pdf` is *refused as unsupported* rather than half-scanned (a PDF parser is disproportionate for a dev showcase, and the OCR/stego it would smuggle is already out of scope). One known gap: the numeric `-`/`+` CSV false positive (a typed XLSX numeric cell does not trip it). - **Lexicon findings are deduplicated by pattern id** — `count=1` and the first offset are reported, so a class matched across several channels collapses to one finding at its first location: a deliberate readability tradeoff. - **Secret egress: base64-wrapped is caught, hex-wrapped is not.** The output gate decodes base64 blobs and re-scans the plaintext, so a base64-*wrapped* secret surfaces as `decoded:egress:*`. `entropy` exposes decoded plaintext for base64 only, so hex (and other encodings, or nested wraps) is a deliberate boundary — decode the transport layer first if you need it scanned. ## Out-of-scope (documented boundary) Embedding/vector-layer defenses (OWASP LLM08, downstream of persist); multimodal steganography; query-time / runtime guardrails; semantic factuality verification. ## Design & threat model - [Design brief](docs/BRIEF.md) — what this repo contains and why. - [Build plan](docs/PLAN.md) — module build order and the reuse map. - [Adoption brief](docs/ADOPTION-BRIEF.md) — wiring the guard into an OKF second-brain / LLM wiki, and a checklist for *when* to include it. The contract is extracted from a working reference implementation (the `claude-code-llm-wiki` Stage B enrichment pipeline). Threat-model anchors: OWASP LLM Top-10 2025 (LLM01/02/04/05/06 strongest, LLM08 boundary, LLM09/10), PoisonedRAG, guardrail-evasion (arXiv 2504.11168), EchoLeak (CVE-2025-32711). ## License MIT — see [LICENSE](LICENSE).