feat(wiring): §6 bookends (prepare_input/screen_output) + full public API + README v0.1 (TDD)
Module 11 (final) — the top-level wiring. The library never makes the model call (no SDK imported by the core), so the public surface is the toolkit plus two library-side bookends around the caller's tool-less transform (Form 1, chosen with the operator over an export-only toolkit and a full orchestrator — the bookends fit existing pipelines with least friction, encode the two halves the library can stand for, and impose no control flow): - prepare_input(text, source=INPUT) -> PreparedInput(fenced, nonce, report): §6 steps 1-2, sanitize THEN fence (carrier can never smuggle a forged delimiter). Merged report carries both steps' findings; renders no disposition. - screen_output(text, policy, *, provenance, transform_failed) -> DispositionResult: §6 steps 6-7, scan_output under guard() so a scanner error fails CLOSED (FAIL_SECURE, never a silent persist). transform_failed routes the compound forced-fallback rule. - __all__ exports the full framework-agnostic surface: detectors, result types, disposition machinery + presets, contract asserters, the grounding seam. Docs: README refreshed from the stale "brief / pre-implementation" line to a v0.1 alpha status with a Form-1 quickstart, the §6 adopt-this checklist, and an honest -limitations section (structural unsolvability at the text layer; semantic poisoning invisible to lexicon+entropy; text-only, no multimodal). CHANGELOG seeded; CLAUDE.md remote/status lines corrected (remote IS set, no longer brief-stage). 10 wiring tests (public surface, prepare_input compose, screen_output fail-closed + compound). 189 green (showcase + corpora follow). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HyRCQMocjZ6SmSQ6JidJ2k
This commit is contained in:
parent
1e63643157
commit
0af8f68cae
5 changed files with 410 additions and 10 deletions
119
README.md
119
README.md
|
|
@ -1,4 +1,4 @@
|
|||
# llm-ingestion-pipeline-security
|
||||
# llm-ingestion-guard
|
||||
|
||||
A reusable, minimal, dependency-light defensive layer for **LLM ingestion
|
||||
pipelines** — the write-time siblings of query-time chatbot guardrails.
|
||||
|
|
@ -9,11 +9,120 @@ untrusted content flowing through an LLM enrichment/summarization/extraction ste
|
|||
into a **persisted, downstream-consumed artifact** (RAG corpus, knowledge base,
|
||||
wiki). It packages the architectural contract — sanitize → fence → tool-less
|
||||
quarantined transform → per-stage capability isolation → scan output before
|
||||
commit → fail-secure — as composable, framework-agnostic code.
|
||||
commit → fail-secure — as composable, stdlib-first, framework-agnostic code.
|
||||
|
||||
**Status:** brief / pre-implementation. Start with the design brief:
|
||||
The gap it fills is **not** "no one detects injection." It is a small *library*
|
||||
(not a hosted service, not a fine-tuned model) that packages the **write-time
|
||||
ingestion contract** — the part query-time tooling structurally cannot see,
|
||||
because a poisoned artifact committed at write time is read by a *downstream*
|
||||
agent whose guardrail never sees where it came from.
|
||||
|
||||
- [Design brief](docs/BRIEF.md) — what this repo should contain and why.
|
||||
**Status:** `v0.1`, alpha. The stdlib-only core is built and tested — ten
|
||||
detector/contract modules and the top-level wiring, exercised by an end-to-end
|
||||
showcase and adversarial + false-positive corpora. The public API may still
|
||||
change. There are real limitations, stated plainly below; read them.
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
pip install llm-ingestion-guard # stdlib-only core, zero dependencies
|
||||
```
|
||||
|
||||
Optional ML/judge detectors live behind extras (`[ml]`, `[judge]`) and are not
|
||||
required — the core is deterministic and dependency-free.
|
||||
|
||||
## Quickstart — the two bookends
|
||||
|
||||
The library never makes the model call itself. It gives you the two library-side
|
||||
halves around your own **tool-less** transform:
|
||||
|
||||
```python
|
||||
from llm_ingestion_guard import (
|
||||
prepare_input, screen_output, Disposition, PRESET_USER_UPLOAD,
|
||||
)
|
||||
|
||||
prepared = prepare_input(untrusted_content) # §6 1-2: sanitize + fence
|
||||
enriched = your_model(prepared.fenced) # §6 3: tool-less — YOUR call
|
||||
decision = screen_output(enriched, PRESET_USER_UPLOAD) # §6 6-7: scan + dispose
|
||||
|
||||
if decision.disposition is Disposition.FAIL_SECURE:
|
||||
alert(gate_code=decision.reasons) # §6 8: minimal payload, no content
|
||||
raise SystemExit # §6 7: halt — never persist
|
||||
```
|
||||
|
||||
`screen_output` fails **closed**: if the scanner itself errors on crafted input,
|
||||
the disposition is `FAIL_SECURE`, never a silent persist. Pass
|
||||
`transform_failed=True` when your model call raised or fell back — a scan hit
|
||||
together with a transform failure is treated as a probable forced-fallback attack
|
||||
and halts regardless of trust tier.
|
||||
|
||||
Every primitive is also exported for pipelines that compose the checklist
|
||||
themselves — `sanitize`, `scan_lexicon`, `scan_entropy`, `scan_output`,
|
||||
`neutralize`, the `decide` / `guard` disposition machinery, and the contract
|
||||
asserters `assert_tool_less` / `assert_credential_allowlist` / `scoped_env`. See
|
||||
[the end-to-end showcase](tests/test_showcase.py) for a full worked pipeline.
|
||||
|
||||
## The reusable contract (adopt-this checklist)
|
||||
|
||||
The actual product is this checklist, encoded as code you wire in order:
|
||||
|
||||
1. **Sanitize before fence.** Strip carrier classes (zero-width, BIDI,
|
||||
Unicode-tag, HTML comment, `data:`) from untrusted input first.
|
||||
2. **Fence untrusted input.** Spotlight-mark it in a randomized per-call
|
||||
delimiter; strip attacker fence markers from the payload.
|
||||
3. **Tool-less transform.** Call the model with zero tools. A successful
|
||||
injection then has nothing to act with.
|
||||
4. **Per-stage capability isolation.** The enrichment stage holds only the model
|
||||
key; the publish stage holds only the publish credential; no stage holds both.
|
||||
5. **Treat output as data.** Parse to a frozen schema; reject on structural
|
||||
violation. The output never reaches a shell, git, or a filesystem path.
|
||||
6. **Scan output before persist.** Run the lexicon + entropy over the emitted
|
||||
text. Verbatim-carried payloads and model-emitted instructions are caught here.
|
||||
7. **Fail-secure on compound signals.** Injection hit + transform failure = halt
|
||||
+ alert, never a silent verbatim commit.
|
||||
8. **Minimal alert payloads.** Alert with a gate code + run ID, never content.
|
||||
|
||||
Steps 1-2 are `prepare_input`; steps 6-7 are `screen_output`; steps 3-5 are
|
||||
yours; the contract asserters harden step 3-4.
|
||||
|
||||
## Honest limitations (shipped as a control)
|
||||
|
||||
Conceding these plainly is itself a control — it prevents the false assurance
|
||||
that a green scan means safe content:
|
||||
|
||||
- **Structural unsolvability at the text layer.** Pattern/lexicon detection is
|
||||
bypassable in isolation; character-injection and novel phrasings evade it. The
|
||||
*contract* (tool-less transform, capability isolation, fail-secure) is what
|
||||
carries the security — the lexicon is defense-in-depth, not a wall.
|
||||
- **Semantic / factual poisoning is invisible** to lexicon + entropy: a
|
||||
factually false claim in clean prose carries no suspicious token. The
|
||||
`grounding` module ships only a `SourceGroundingCheck` *seam* — the deterministic
|
||||
core does not judge semantics; a `[judge]` implementation must be plugged in.
|
||||
- **Adversarial-ML evasion** can survive normalization; **tokenizer mismatch**
|
||||
between scanner and model leaves gaps.
|
||||
- **Latent / dormant memory poisoning** is not judgeable at write time.
|
||||
- **Insider in-place edits** by a trusted author are out of the untrusted-content
|
||||
threat model.
|
||||
- **Text-only.** The core is `text -> findings`: it parses no files (no
|
||||
`pypdf`/`python-docx`/archive deps). Extract text first, then scan it with the
|
||||
high-untrust upload provenance. OCR-embedded instructions and multimodal stego
|
||||
in images/PDFs are out of scope beyond the sanitizer's character-layer stripping.
|
||||
|
||||
## Out-of-scope (documented boundary)
|
||||
|
||||
Embedding/vector-layer defenses (OWASP LLM08, downstream of persist); multimodal
|
||||
steganography; query-time / runtime guardrails; semantic factuality verification.
|
||||
|
||||
## Design & threat model
|
||||
|
||||
- [Design brief](docs/BRIEF.md) — what this repo contains and why.
|
||||
- [Build plan](docs/PLAN.md) — module build order and the reuse map.
|
||||
|
||||
The contract is extracted from a working reference implementation (the
|
||||
`claude-code-llm-wiki` Stage B enrichment pipeline).
|
||||
`claude-code-llm-wiki` Stage B enrichment pipeline). Threat-model anchors: OWASP
|
||||
LLM Top-10 2025 (LLM01/02/04/05/06 strongest, LLM08 boundary, LLM09/10),
|
||||
PoisonedRAG, guardrail-evasion (arXiv 2504.11168), EchoLeak (CVE-2025-32711).
|
||||
|
||||
## License
|
||||
|
||||
MIT — see [LICENSE](LICENSE).
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue