Answers the gap that the README said what it does NOT stop (a long limitations section) but never plainly listed what it DOES. Add a 'What it protects against' section high up: attack classes grouped by OWASP anchor (LLM01 injection + 83 lexicon classes + carriers, LLM02 egress, LLM05 EchoLeak, LLM06 agency, LLM10 fail-secure, OKF T1-T6, container front-end), each driven by a live coverage-matrix payload. Move the full honest-limitations list to docs/LIMITATIONS.md; README keeps a high-impact summary + link. Net: protection and limits read in balance, 261 -> 216 lines. Coverage 126/126 and 522 tests unchanged; every class listed is real.
6 KiB
6 KiB
Honest limitations
Conceding these plainly is itself a control — it prevents the false assurance that a green scan means safe content. The README carries a summary of the highest-impact items; this is the full list, each with the mechanism.
- Structural unsolvability at the text layer. Pattern/lexicon detection is bypassable in isolation; character-injection and novel phrasings evade it. The contract (tool-less transform, capability isolation, fail-secure) carries the security — the lexicon is defense-in-depth, not a wall.
- A lone HIGH finding in trusted prose disposes to WARN, not quarantine.
Under
PRESET_TRUSTED_SOURCE, trust-scaling downgrades a single HIGH to WARN, and one HIGH is not "compound" (escalation needs ≥2 findings at MEDIUM+). So a HIGH injection reproduced verbatim under a trusted policy persists with a WARN. This is by design: if your "trusted" sources can carry attacker-influenced text, run them as untrusted (or add a quarantine floor). - The upload preset's quarantine floor is currently vacuous.
PRESET_USER_UPLOADsetsquarantine_default, but detectors emit CRITICAL/HIGH/MEDIUM only, and under untrusted trust a MEDIUM already escalates to QUARANTINE_REVIEW — so the floor changes no outcome today. It is headroom for a future LOW/INFO finding, documented so the preset is not over-read. - Semantic / factual poisoning is invisible to lexicon + entropy: a false claim
in clean prose carries no suspicious token. Highest impact for a wiki. The
groundingmodule ships only aSourceGroundingCheckseam — the deterministic core does not judge semantics; a[judge]implementation must be plugged in. - Adversarial-ML evasion can survive normalization; tokenizer mismatch between scanner and model leaves gaps. Latent / dormant memory poisoning is not judgeable at write time.
- Dormant / broken-link injection in a linked corpus (e.g. an OKF bundle): a
link to a not-yet-existing target passes a per-concept write-time scan clean — the
payload is planted later, when that target is written.
link_graphsurfaces the dangling edge as the signal, but catching the payload needs cross-write re-scan over time (the caller's disposition call). - OKF reserved files (
index.md/log.md). In a received bundle these are legitimate structure, so mode-bimport_bundlescans their body and frontmatter (an injection in a directory listing is caught) rather than path-rejecting the conformant bundle. A front-end materialising individual uploads keeps the opposite rule (allow_reserved=False): a reserved basename is a listing-shadow and refused. - A document that describes attacks is a false positive. Content documenting prompt-injection payloads (security notes, this project's own corpus) trips carrier-strip / fail-secure. At the text layer "about an attack" and "carrying an attack" are indistinguishable; such content needs a deliberate, explicitly escaped path, never a silent allow.
- Bilingual text trips the Cyrillic/Latin homoglyph rule.
homoglyph:cyrillic-latin-mix(MEDIUM) flags a Latin letter adjacent to a Cyrillic look-alike, so genuine bilingual prose → MEDIUM → under untrusted → QUARANTINE_REVIEW — a real false positive for an inbox that expects multilingual content. A calibration fix is pending. - Insider in-place edits by a trusted author are out of the untrusted-content threat model.
- Text-only. The core is
text -> findings: it parses no files (nopypdf/python-docx/archive deps). Extract text first, then scan it with the high-untrust upload provenance. OCR-embedded instructions and multimodal stego are out of scope beyond the sanitizer's character-layer stripping. - Uploaded files: only the extracted text is scanned. The dev-scoped OKF
inbox showcase (
tests/test_okf_inbox_uploads.py+tests/inbox_frontend.py; parsers in the[dev]extra, never coredependencies) reads.txt/.md/.csv/.docx/.pptx/.xlsx, folders and.zip, materializes an OKF bundle, then guards it. What survives extraction is out of scope: macros, OLE/embedded objects, OCR-needing images, font/render stego, encrypted files — the binary layer needs a separate scanner. The front-end owns the container threats it can see (zip-slip → path gate, zip-bomb → size cap, symlink refusal, CSV/XLSX formula-lead cells)..pdfis a deliberate concession: a top-level.pdfis refused as unsupported rather than half-scanned (a PDF parser is disproportionate for a dev showcase, and the OCR/stego it would smuggle is already out of scope). One known gap: the numeric-/+CSV false positive (a typed XLSX numeric cell does not trip it). - Lexicon findings are deduplicated by pattern id —
count=1and the first offset are reported, so a class matched across several channels collapses to one finding at its first location: a deliberate readability tradeoff. - Secret egress: base64-wrapped is caught, hex-wrapped is not. The output gate
decodes base64 blobs and re-scans the plaintext, so a base64-wrapped secret
surfaces as
decoded:egress:*.entropyexposes decoded plaintext for base64 only, so hex (and other encodings, or nested wraps) is a deliberate boundary — decode the transport layer first if you need it scanned.
The four documented gaps (tracked by the coverage matrix)
These are asserted to still hold by tests/test_coverage_matrix.py — a closed gap
fails the test, forcing this doc to be updated:
- Hex-wrapped secret egress —
entropydecodes base64 only (above). - Semantic / factual poisoning — invisible to token analysis (above).
- A lone HIGH in trusted prose → WARN — the §4.7 trust-scaling design (above).
- Lexicon dedup (
count=1) — first offset only, by design (above).
Out-of-scope (documented boundary)
Embedding/vector-layer defenses (OWASP LLM08, downstream of persist); multimodal steganography; query-time / runtime guardrails; semantic factuality verification.