1
0
Fork 0
llm-ingestion-pipeline-secu.../docs/ADOPTION-BRIEF.md
Kjell Tore Guttormsen a1f3fe1983 docs(adoption): reusable consumer adoption brief for OKF second-brain repos
Self-contained brief a consumer repo can plan an inclusion from: what the
guard is (write-time, not query-time), the two bookends + 8-step contract,
the shipped OKF adapter (import_bundle mode-b, per-concept gates), how to
verify (coverage matrix -> 126 classes), how to depend (stdlib-only core),
and a planning checklist for WHEN/WHERE to wire it (untrusted boundary, not
first-party onboarding). Every claim verified against v0.2 code.
2026-07-15 12:35:09 +02:00

237 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Adoption brief — wiring `llm-ingestion-guard` into an OKF second-brain / LLM wiki
**Audience:** a repo that is building (or planning) an LLM wiki / second-brain —
especially one converging on Google's Open Knowledge Format (OKF v0.1) — and needs
to decide **when** and **where** to add a write-time ingestion guard.
**Status of the guard:** `v0.2` (alpha). Stdlib-only core, framework-agnostic.
Public API may still change. Read the honest-limitations section before you rely
on it.
This brief is self-contained: you can plan an inclusion from it alone. Every
technical claim below is checkable against the guard repo (commands given inline).
---
## 1. What this is — and what it is *not*
`llm-ingestion-guard` is the **write-time** sibling of query-time chatbot
guardrails. It does **not** sit between a user and a model at query time (that is
LLM Guard / NeMo Guardrails / Rebuff / Vigil territory). It hardens the other
shape: **untrusted content flowing through an LLM enrichment/summarization/
extraction step into a *persisted, downstream-consumed* artifact** — a RAG corpus,
a knowledge base, a wiki, an OKF bundle.
Why this matters for a second-brain: a poisoned concept committed at **write**
time is later read by a *downstream* agent as **trusted context**. That agent's
query-time guardrail never sees where the concept came from. The write gate is the
only place the provenance is still known. **Your ingestion pipeline *is* the trust
boundary** — OKF has no schema registry, no central authority, and no signing, so
a received bundle's claimed origin is not verifiable at the format level.
The library never makes the model call itself. It gives you the two library-side
halves around your own **tool-less** transform, plus an OKF adapter for bundles.
## 2. The two bookends (the minimal integration)
```python
from llm_ingestion_guard import (
prepare_input, screen_output, Disposition, PRESET_USER_UPLOAD,
)
prepared = prepare_input(untrusted_content) # sanitize + fence
enriched = your_model(prepared.fenced) # tool-less — YOUR call
decision = screen_output(enriched, PRESET_USER_UPLOAD) # scan + dispose
if decision.disposition is Disposition.FAIL_SECURE:
alert(gate_code=decision.reasons) # minimal payload, no content
raise SystemExit # halt — never persist
```
`screen_output` fails **closed**: if the scanner itself errors on crafted input,
the disposition is `FAIL_SECURE`, never a silent persist. Pass
`transform_failed=True` when your model call raised or fell back — a scan hit
together with a transform failure is treated as a probable forced-fallback attack
and halts regardless of trust tier.
## 3. The reusable contract (the actual product — an adopt-this checklist)
The library is this checklist encoded as composable code you wire in order. Steps
12 are `prepare_input`; steps 67 are `screen_output`; steps 35 are yours (the
contract asserters harden 34):
1. **Sanitize before fence.** Strip carrier classes (zero-width, BIDI,
Unicode-tag, HTML comment, `data:`) from untrusted input first.
2. **Fence untrusted input.** Spotlight-mark it in a randomized per-call delimiter;
strip attacker fence markers from the payload.
3. **Tool-less transform.** Call the model with zero tools. A successful injection
then has nothing to act with.
4. **Per-stage capability isolation.** The enrichment stage holds only the model
key; the publish stage holds only the publish credential; no stage holds both.
5. **Treat output as data.** Parse to a frozen schema; reject on structural
violation. Output never reaches a shell, git, or a filesystem path.
6. **Scan output before persist.** Lexicon + entropy + active-content scan over the
emitted text (catches verbatim-carried payloads, model-emitted instructions,
EchoLeak-class markdown images/links, raw active HTML, `data:` URIs).
7. **Fail-secure on compound signals.** Injection hit + transform failure = halt +
alert, never a silent verbatim commit.
8. **Minimal alert payloads.** Alert with a gate code + run ID, never content.
You do not have to take all eight at once — every primitive is exported
(`sanitize`, `scan_lexicon`, `scan_entropy`, `scan_output`, `scan_active_content`,
`neutralize`, the `decide`/`guard` disposition machinery, and the contract
asserters `assert_tool_less` / `assert_credential_allowlist` / `scoped_env`).
## 4. The OKF adapter (for bundle-shaped ingestion)
If your second-brain is (or is converging on) OKF, the `okf` submodule sits **on
top of** the format-agnostic core: it knows OKF structure (frontmatter, paths,
links, `resource`, bundles) and feeds scannable regions into the same
`sanitize` / `scan_output` / disposition machinery. No YAML/format awareness leaks
into the core.
Two ingestion modes:
- **(a) Own enrichment output** — your agent writes concepts. Run the two bookends
(§2) per concept before commit.
- **(b) Received external bundle** — you merge a whole third-party OKF bundle.
`import_bundle` iterates concept-by-concept and runs the full per-concept gate:
```python
from llm_ingestion_guard.okf import import_bundle, Origin, Channel
# bundle: {concept_path -> raw document text}, e.g. {"tables/users.md": "---\n..."}
result = import_bundle(bundle, origin=Origin.EXTERNAL, channel=Channel.AUTOMATIC)
for c in result.concepts:
if c.error: # hard reject: bad path, unsafe frontmatter, non-https resource
skip(c.path) # disposition is FAIL_SECURE; do not merge this concept
# result.disposition = most-severe across concepts; result.links = cross-link graph
# result.log() = the log.md body (one provenance-stamped line per concept)
```
Per-concept gates the adapter applies (each maps to a named control):
| Gate | What it does |
|---|---|
| **Path / reserved-name** | Rejects `..` traversal, absolute paths, and reserved-name shadowing. `validate_concept_path`. |
| **Frontmatter parse-safety** | `parse_frontmatter` is a strict, reject-by-default loader for the minimal OKF subset — anchors, aliases, and explicit tags are refused *by construction*, so billion-laughs alias DoS and `!!python/object` coercion cannot occur. It is deliberately **not** a general YAML engine (that engine's features *are* the attack surface). |
| **`resource` URL allowlist** | `validate_resource_url` hard-rejects non-`https` (`data:`/`javascript:`/`file:`) before commit — a reject-gate, not defang. |
| **Whole-concept scan** | Frontmatter *values* + body run through the core `scan_output`. |
| **Cross-link graph** | `link_graph` resolves in-bundle links, flags dangling targets (the dormant-injection signal), and rejects dangerous-scheme / bundle-escaping targets. |
| **Provenance stamping** | `Origin` × `Channel` → trust tier + disposition per concept, emitted to `log.md`. |
**Reserved files (`index.md` / `log.md`) — a deliberate mode difference.** In a
*received* bundle these are legitimate structure (directory listing, update log),
so `import_bundle` defaults to `allow_reserved=True`: it **scans their body and
frontmatter** (a directory listing is a high-priority injection surface) rather
than path-rejecting an otherwise-conformant third-party bundle. A front-end that
materialises **individual uploads** should pass `allow_reserved=False` instead —
there a reserved basename is a shadow of the listing and must be refused. Pick the
rule that matches your channel.
## 5. Verify what it stops (before you wire it in)
The guard ships a runnable coverage matrix — every vulnerability class it stops,
and the ones it deliberately does not, each row driving the **real** guard with a
live payload:
```bash
python -m llm_ingestion_guard.coverage # exit 0 = all as documented
```
As of `v0.2`: **126 / 126 defended classes demonstrated (recall 100%)** and **4 /
4 documented gaps still hold** (a *closed* gap fails the test, forcing a doc
update). The matrix is the single source of truth for the test suite (**522
passing**), which also asserts total recall, that every lexicon pattern has a
case (so the matrix cannot fall behind the lexicon), the full LLM02 secret-egress
set, and the container-layer front-end (CSV formula-injection, zip-slip/bomb,
symlink). Run it once; it tells you exactly what assurance you are buying.
## 6. How to depend on it
```bash
pip install llm-ingestion-guard # stdlib-only core, zero dependencies
```
- **Core is stdlib-only** (`dependencies = []`), Python **3.10+**. Nothing to
vet for supply-chain beyond the package itself; it parses no files and makes no
network calls.
- **Optional extras**, none required: `[ml]` / `[judge]` (heavier detectors, e.g.
a semantic-poisoning judge behind the `grounding` seam), `[dev]` (file-extraction
libs used only by the dev-scoped upload showcase — `python-docx`/`python-pptx`/
`openpyxl`/`lxml`/`Pillow`; never core dependencies).
- The core is `text -> findings`. If you ingest files, **extract text first**, then
scan the extracted text with high-untrust upload provenance.
## 7. Planning checklist — *when* to include the guard
Score your ingestion pipeline. The guard earns its place at the **persist gate**
when the untrusted-ingest condition holds:
- [ ] You persist LLM-enriched or externally-received content into a store a
*downstream* agent later reads as trusted (RAG / KB / wiki / OKF bundle).
- [ ] **At least one ingest path takes UNTRUSTED content** — an external URL, an
uploaded file, a received third-party bundle, or auto-fetched web content.
*(This is the decisive one.)*
- [ ] An LLM step (summarize / extract / classify / rewrite) sits between the
untrusted source and the store.
- [ ] You want **fail-secure** (halt before persist) rather than best-effort
detection with a silent commit on error.
**Where it applies vs. where it doesn't.** A second-brain that ingests primarily
the **user's own context** (onboarding writes conformant concepts, the user edits
their own notes) is a *first-party* path — the guard's untrusted-content threat
model does **not** target it, and trusted-author in-place edits are out of scope
by design. Wire the guard specifically at the **untrusted boundary**: a
"react-to-URL" command, an inbox that accepts external drops, a manual-import of a
foreign file, an auto-fetch of web/vendor content, or a received third-party OKF
bundle. Trust follows the data's **origin**, not the insertion channel — a manual
paste of external material is still external.
**When in your roadmap.** It is a persist-gate, so include it *before the first
untrusted ingest path goes live*, wired at the point where enriched content is
committed. If today you only have first-party ingest, note the guard as a
dependency to add when (not if) you open an external/inbox/received-bundle path.
## 8. Honest limitations (read these — a green scan is not "safe")
Conceding these is itself a control. The full list is in the guard's `README.md`
("Honest limitations"); the ones that matter most for a wiki/second-brain:
- **Semantic / factual poisoning is invisible** to lexicon + entropy: a
plausible-but-wrong concept (wrong join-path, wrong metric, wrong runbook step)
carries no suspicious token and passes clean. **Highest impact for a wiki.**
Needs human review or source verification — the deterministic core does not judge
semantics; a `[judge]` implementation plugs into the `grounding` seam.
- **Dormant / broken-link injection**: a link to a not-yet-existing target passes a
per-concept write-time scan; the payload is planted later when that target is
written. `link_graph` surfaces the *dangling* edge as the signal, but whether to
block is your disposition call, and cross-write re-scan over time is your
responsibility.
- **A document that *describes* attacks is a false positive.** Security notes that
legitimately document injection payloads trip carrier-strip / fail-secure. At the
text layer "about an attack" is indistinguishable from "carrying an attack" — such
content needs a deliberate, explicitly-marked escaped path, never a silent allow.
- **Structural unsolvability at the text layer.** Pattern/lexicon detection is
bypassable in isolation; novel phrasings and character-injection evade it. The
*contract* (tool-less transform, capability isolation, fail-secure) carries the
security — the lexicon is defense-in-depth, not a wall.
- **Text-only, extracted-text-only.** No file parsing in the core; what survives
text extraction (macros, OLE objects, OCR-embedded instructions, render/font
stego, encrypted files) is out of scope beyond the sanitizer's character layer.
- **Secret egress: base64-wrapped is caught, hex-wrapped is not** (a documented
boundary — decode the transport layer first if you need it scanned).
## 9. Where to read more (in the guard repo)
- `README.md` — usage, the full contract, and the complete honest-limitations list.
- `docs/BRIEF.md` — design rationale and the nearest-neighbour survey (§11).
- `docs/OKF-INGESTION-BRIEF.md` — the OKF threat-surface analysis (frontmatter,
`resource`, cross-link graph, reserved names, provenance) the adapter implements.
- `python -m llm_ingestion_guard.coverage` — the runnable "verify what it stops".
Threat-model anchors: OWASP LLM Top-10 2025 (LLM01/02/04/05/06 strongest, LLM08
boundary), PoisonedRAG, guardrail-evasion (arXiv 2504.11168), EchoLeak
(CVE-2025-32711). OKF: Google Cloud Open Knowledge Format v0.1 (announced
2026-06-12) — `GoogleCloudPlatform/knowledge-catalog/okf/SPEC.md`.