1
0
Fork 0

docs(adoption): reusable consumer adoption brief for OKF second-brain repos

Self-contained brief a consumer repo can plan an inclusion from: what the
guard is (write-time, not query-time), the two bookends + 8-step contract,
the shipped OKF adapter (import_bundle mode-b, per-concept gates), how to
verify (coverage matrix -> 126 classes), how to depend (stdlib-only core),
and a planning checklist for WHEN/WHERE to wire it (untrusted boundary, not
first-party onboarding). Every claim verified against v0.2 code.
This commit is contained in:
Kjell Tore Guttormsen 2026-07-15 12:35:09 +02:00
commit a1f3fe1983

237
docs/ADOPTION-BRIEF.md Normal file
View file

@ -0,0 +1,237 @@
# Adoption brief — wiring `llm-ingestion-guard` into an OKF second-brain / LLM wiki
**Audience:** a repo that is building (or planning) an LLM wiki / second-brain —
especially one converging on Google's Open Knowledge Format (OKF v0.1) — and needs
to decide **when** and **where** to add a write-time ingestion guard.
**Status of the guard:** `v0.2` (alpha). Stdlib-only core, framework-agnostic.
Public API may still change. Read the honest-limitations section before you rely
on it.
This brief is self-contained: you can plan an inclusion from it alone. Every
technical claim below is checkable against the guard repo (commands given inline).
---
## 1. What this is — and what it is *not*
`llm-ingestion-guard` is the **write-time** sibling of query-time chatbot
guardrails. It does **not** sit between a user and a model at query time (that is
LLM Guard / NeMo Guardrails / Rebuff / Vigil territory). It hardens the other
shape: **untrusted content flowing through an LLM enrichment/summarization/
extraction step into a *persisted, downstream-consumed* artifact** — a RAG corpus,
a knowledge base, a wiki, an OKF bundle.
Why this matters for a second-brain: a poisoned concept committed at **write**
time is later read by a *downstream* agent as **trusted context**. That agent's
query-time guardrail never sees where the concept came from. The write gate is the
only place the provenance is still known. **Your ingestion pipeline *is* the trust
boundary** — OKF has no schema registry, no central authority, and no signing, so
a received bundle's claimed origin is not verifiable at the format level.
The library never makes the model call itself. It gives you the two library-side
halves around your own **tool-less** transform, plus an OKF adapter for bundles.
## 2. The two bookends (the minimal integration)
```python
from llm_ingestion_guard import (
prepare_input, screen_output, Disposition, PRESET_USER_UPLOAD,
)
prepared = prepare_input(untrusted_content) # sanitize + fence
enriched = your_model(prepared.fenced) # tool-less — YOUR call
decision = screen_output(enriched, PRESET_USER_UPLOAD) # scan + dispose
if decision.disposition is Disposition.FAIL_SECURE:
alert(gate_code=decision.reasons) # minimal payload, no content
raise SystemExit # halt — never persist
```
`screen_output` fails **closed**: if the scanner itself errors on crafted input,
the disposition is `FAIL_SECURE`, never a silent persist. Pass
`transform_failed=True` when your model call raised or fell back — a scan hit
together with a transform failure is treated as a probable forced-fallback attack
and halts regardless of trust tier.
## 3. The reusable contract (the actual product — an adopt-this checklist)
The library is this checklist encoded as composable code you wire in order. Steps
12 are `prepare_input`; steps 67 are `screen_output`; steps 35 are yours (the
contract asserters harden 34):
1. **Sanitize before fence.** Strip carrier classes (zero-width, BIDI,
Unicode-tag, HTML comment, `data:`) from untrusted input first.
2. **Fence untrusted input.** Spotlight-mark it in a randomized per-call delimiter;
strip attacker fence markers from the payload.
3. **Tool-less transform.** Call the model with zero tools. A successful injection
then has nothing to act with.
4. **Per-stage capability isolation.** The enrichment stage holds only the model
key; the publish stage holds only the publish credential; no stage holds both.
5. **Treat output as data.** Parse to a frozen schema; reject on structural
violation. Output never reaches a shell, git, or a filesystem path.
6. **Scan output before persist.** Lexicon + entropy + active-content scan over the
emitted text (catches verbatim-carried payloads, model-emitted instructions,
EchoLeak-class markdown images/links, raw active HTML, `data:` URIs).
7. **Fail-secure on compound signals.** Injection hit + transform failure = halt +
alert, never a silent verbatim commit.
8. **Minimal alert payloads.** Alert with a gate code + run ID, never content.
You do not have to take all eight at once — every primitive is exported
(`sanitize`, `scan_lexicon`, `scan_entropy`, `scan_output`, `scan_active_content`,
`neutralize`, the `decide`/`guard` disposition machinery, and the contract
asserters `assert_tool_less` / `assert_credential_allowlist` / `scoped_env`).
## 4. The OKF adapter (for bundle-shaped ingestion)
If your second-brain is (or is converging on) OKF, the `okf` submodule sits **on
top of** the format-agnostic core: it knows OKF structure (frontmatter, paths,
links, `resource`, bundles) and feeds scannable regions into the same
`sanitize` / `scan_output` / disposition machinery. No YAML/format awareness leaks
into the core.
Two ingestion modes:
- **(a) Own enrichment output** — your agent writes concepts. Run the two bookends
(§2) per concept before commit.
- **(b) Received external bundle** — you merge a whole third-party OKF bundle.
`import_bundle` iterates concept-by-concept and runs the full per-concept gate:
```python
from llm_ingestion_guard.okf import import_bundle, Origin, Channel
# bundle: {concept_path -> raw document text}, e.g. {"tables/users.md": "---\n..."}
result = import_bundle(bundle, origin=Origin.EXTERNAL, channel=Channel.AUTOMATIC)
for c in result.concepts:
if c.error: # hard reject: bad path, unsafe frontmatter, non-https resource
skip(c.path) # disposition is FAIL_SECURE; do not merge this concept
# result.disposition = most-severe across concepts; result.links = cross-link graph
# result.log() = the log.md body (one provenance-stamped line per concept)
```
Per-concept gates the adapter applies (each maps to a named control):
| Gate | What it does |
|---|---|
| **Path / reserved-name** | Rejects `..` traversal, absolute paths, and reserved-name shadowing. `validate_concept_path`. |
| **Frontmatter parse-safety** | `parse_frontmatter` is a strict, reject-by-default loader for the minimal OKF subset — anchors, aliases, and explicit tags are refused *by construction*, so billion-laughs alias DoS and `!!python/object` coercion cannot occur. It is deliberately **not** a general YAML engine (that engine's features *are* the attack surface). |
| **`resource` URL allowlist** | `validate_resource_url` hard-rejects non-`https` (`data:`/`javascript:`/`file:`) before commit — a reject-gate, not defang. |
| **Whole-concept scan** | Frontmatter *values* + body run through the core `scan_output`. |
| **Cross-link graph** | `link_graph` resolves in-bundle links, flags dangling targets (the dormant-injection signal), and rejects dangerous-scheme / bundle-escaping targets. |
| **Provenance stamping** | `Origin` × `Channel` → trust tier + disposition per concept, emitted to `log.md`. |
**Reserved files (`index.md` / `log.md`) — a deliberate mode difference.** In a
*received* bundle these are legitimate structure (directory listing, update log),
so `import_bundle` defaults to `allow_reserved=True`: it **scans their body and
frontmatter** (a directory listing is a high-priority injection surface) rather
than path-rejecting an otherwise-conformant third-party bundle. A front-end that
materialises **individual uploads** should pass `allow_reserved=False` instead —
there a reserved basename is a shadow of the listing and must be refused. Pick the
rule that matches your channel.
## 5. Verify what it stops (before you wire it in)
The guard ships a runnable coverage matrix — every vulnerability class it stops,
and the ones it deliberately does not, each row driving the **real** guard with a
live payload:
```bash
python -m llm_ingestion_guard.coverage # exit 0 = all as documented
```
As of `v0.2`: **126 / 126 defended classes demonstrated (recall 100%)** and **4 /
4 documented gaps still hold** (a *closed* gap fails the test, forcing a doc
update). The matrix is the single source of truth for the test suite (**522
passing**), which also asserts total recall, that every lexicon pattern has a
case (so the matrix cannot fall behind the lexicon), the full LLM02 secret-egress
set, and the container-layer front-end (CSV formula-injection, zip-slip/bomb,
symlink). Run it once; it tells you exactly what assurance you are buying.
## 6. How to depend on it
```bash
pip install llm-ingestion-guard # stdlib-only core, zero dependencies
```
- **Core is stdlib-only** (`dependencies = []`), Python **3.10+**. Nothing to
vet for supply-chain beyond the package itself; it parses no files and makes no
network calls.
- **Optional extras**, none required: `[ml]` / `[judge]` (heavier detectors, e.g.
a semantic-poisoning judge behind the `grounding` seam), `[dev]` (file-extraction
libs used only by the dev-scoped upload showcase — `python-docx`/`python-pptx`/
`openpyxl`/`lxml`/`Pillow`; never core dependencies).
- The core is `text -> findings`. If you ingest files, **extract text first**, then
scan the extracted text with high-untrust upload provenance.
## 7. Planning checklist — *when* to include the guard
Score your ingestion pipeline. The guard earns its place at the **persist gate**
when the untrusted-ingest condition holds:
- [ ] You persist LLM-enriched or externally-received content into a store a
*downstream* agent later reads as trusted (RAG / KB / wiki / OKF bundle).
- [ ] **At least one ingest path takes UNTRUSTED content** — an external URL, an
uploaded file, a received third-party bundle, or auto-fetched web content.
*(This is the decisive one.)*
- [ ] An LLM step (summarize / extract / classify / rewrite) sits between the
untrusted source and the store.
- [ ] You want **fail-secure** (halt before persist) rather than best-effort
detection with a silent commit on error.
**Where it applies vs. where it doesn't.** A second-brain that ingests primarily
the **user's own context** (onboarding writes conformant concepts, the user edits
their own notes) is a *first-party* path — the guard's untrusted-content threat
model does **not** target it, and trusted-author in-place edits are out of scope
by design. Wire the guard specifically at the **untrusted boundary**: a
"react-to-URL" command, an inbox that accepts external drops, a manual-import of a
foreign file, an auto-fetch of web/vendor content, or a received third-party OKF
bundle. Trust follows the data's **origin**, not the insertion channel — a manual
paste of external material is still external.
**When in your roadmap.** It is a persist-gate, so include it *before the first
untrusted ingest path goes live*, wired at the point where enriched content is
committed. If today you only have first-party ingest, note the guard as a
dependency to add when (not if) you open an external/inbox/received-bundle path.
## 8. Honest limitations (read these — a green scan is not "safe")
Conceding these is itself a control. The full list is in the guard's `README.md`
("Honest limitations"); the ones that matter most for a wiki/second-brain:
- **Semantic / factual poisoning is invisible** to lexicon + entropy: a
plausible-but-wrong concept (wrong join-path, wrong metric, wrong runbook step)
carries no suspicious token and passes clean. **Highest impact for a wiki.**
Needs human review or source verification — the deterministic core does not judge
semantics; a `[judge]` implementation plugs into the `grounding` seam.
- **Dormant / broken-link injection**: a link to a not-yet-existing target passes a
per-concept write-time scan; the payload is planted later when that target is
written. `link_graph` surfaces the *dangling* edge as the signal, but whether to
block is your disposition call, and cross-write re-scan over time is your
responsibility.
- **A document that *describes* attacks is a false positive.** Security notes that
legitimately document injection payloads trip carrier-strip / fail-secure. At the
text layer "about an attack" is indistinguishable from "carrying an attack" — such
content needs a deliberate, explicitly-marked escaped path, never a silent allow.
- **Structural unsolvability at the text layer.** Pattern/lexicon detection is
bypassable in isolation; novel phrasings and character-injection evade it. The
*contract* (tool-less transform, capability isolation, fail-secure) carries the
security — the lexicon is defense-in-depth, not a wall.
- **Text-only, extracted-text-only.** No file parsing in the core; what survives
text extraction (macros, OLE objects, OCR-embedded instructions, render/font
stego, encrypted files) is out of scope beyond the sanitizer's character layer.
- **Secret egress: base64-wrapped is caught, hex-wrapped is not** (a documented
boundary — decode the transport layer first if you need it scanned).
## 9. Where to read more (in the guard repo)
- `README.md` — usage, the full contract, and the complete honest-limitations list.
- `docs/BRIEF.md` — design rationale and the nearest-neighbour survey (§11).
- `docs/OKF-INGESTION-BRIEF.md` — the OKF threat-surface analysis (frontmatter,
`resource`, cross-link graph, reserved names, provenance) the adapter implements.
- `python -m llm_ingestion_guard.coverage` — the runnable "verify what it stops".
Threat-model anchors: OWASP LLM Top-10 2025 (LLM01/02/04/05/06 strongest, LLM08
boundary), PoisonedRAG, guardrail-evasion (arXiv 2504.11168), EchoLeak
(CVE-2025-32711). OKF: Google Cloud Open Knowledge Format v0.1 (announced
2026-06-12) — `GoogleCloudPlatform/knowledge-catalog/okf/SPEC.md`.