1
0
Fork 0
llm-ingestion-pipeline-secu.../src/llm_ingestion_guard/neutralize.py
Kjell Tore Guttormsen 2d98d6809d feat(sanitize,fence,neutralize): reject oversize input instead of half-transforming it
The scanners cap by truncating: they return findings, so reading a prefix costs
detection in the tail and nothing else. The three transform surfaces return
*content*, where the same move is not available — a shortened document is silent
data loss, and a transformed prefix followed by an untransformed tail is a
bypass, since the attacker chooses where in the document the payload sits.

So they fail secure instead. Above MAX_INPUT_CHARS (1 000 000) sanitize, fence
and neutralize raise OversizeInputError. sanitize is step 1 of prepare_input and
only ever removes, so that one refusal bounds the whole input path.

OversizeInputError subclasses ContractViolation: a pipeline already bracketing
its quarantined stage keeps failing closed rather than meeting a type it has
never heard of. It inherits the alert-routable property too — sizes in the
message, refusing surface in details, no input in either.

Invariant now pinned across all three: returned text is always fully
transformed, or not returned at all.

Still uncapped and recorded in LIMITATIONS: scan_active_content called directly
(through scan_output it inherits that cap) and the okf link graph. Both are
detection-shaped, so truncate-and-flag transfers unchanged — mechanical, not
policy.

699 tests (+23), coverage 128/128 + 6/6, ReDoS sweep 0 candidates / 150.
2026-08-02 21:13:08 +02:00

174 lines
6.8 KiB
Python

"""neutralize — opt-in, pure defang of active content in model OUTPUT.
Query-time guardrails guard the answer; this guards the *persisted artifact*.
When model output is written to a wiki, doc, or knowledge base and later rendered,
active-content constructs become an exfiltration channel: a markdown image URL is
auto-fetched the moment the page renders, leaking whatever the attacker packed
into it — with no click. This is the EchoLeak class (CVE-2025-32711). Such
carriers are neither injection strings nor high-entropy blobs, so ``lexicon`` and
``entropy`` do not see them; neutralizing them is a distinct control (OWASP
LLM05 — Improper Output Handling).
Defang, don't delete. Each active construct is rewritten to an inert but still
human-auditable form: URLs get a non-resolvable scheme and bracketed dots
(``https://evil.com`` -> ``hxxps://evil[.]com``), and raw active HTML is escaped
so a renderer shows it as literal text instead of executing it. The visible
information survives review; only the machine-actionable affordance dies.
The pattern table this mutator rewrites is shared with the report-only detector
:func:`~llm_ingestion_guard.active_content.scan_active_content` and lives in
``active_content`` — detection feeds the standard gate; defanging stays the
separate, opt-in mutation below.
Two properties are load-bearing and mirror the sanitizer:
1. **Opt-in and separate.** Calling this function *is* the opt-in to mutate.
Detection elsewhere in the library stays report-only (design principles 3 & 4);
the report-only output gate (``output``) never rewrites. A caller that wants
findings without mutation reads ``result.report`` and discards ``result.text``.
2. **Byte-identical on clean input.** Output with no active construct is returned
unchanged with an empty report. Benign inline formatting (``<b>``, ``<em>``)
and prose containing stray ``<``, ``>``, ``[`` are left untouched.
Scope note (conceded, not hidden): this is a targeted defanger, not a full HTML
sanitizer. Text *between* escaped tags (e.g. a ``<script>`` body) is neutralized
as active content by escaping the tags, but bare URLs left in that residual text
stay visible; balanced-parenthesis link URLs are matched conservatively. The goal
is to kill the zero-click auto-fetch/execute affordance, not to rewrite every URL.
"""
from __future__ import annotations
import re
from dataclasses import dataclass
from .active_content import (
AUTOLINK_RE,
DATA_URI_RE,
HTML_TAG_RE,
MD_IMAGE_RE,
MD_LINK_RE,
MD_REFDEF_RE,
URL_IN_TEXT_RE,
defang_url,
is_active_tag,
redact,
)
from .calibration import MAX_INPUT_CHARS
from .contract import assert_within_input_cap
from .report import Finding, Report, Severity, Source
@dataclass(frozen=True)
class NeutralizeResult:
"""The defanged text plus a report of every construct that was neutralized."""
text: str
report: Report
def neutralize(
text: str,
source: Source = Source.OUTPUT,
max_input_chars: int = MAX_INPUT_CHARS,
) -> NeutralizeResult:
"""Defang active-content constructs in ``text`` and report each class.
Rewrites markdown images/links, reference-link definitions, angle-bracket
autolinks, raw active HTML, and ``data:`` URIs into inert forms. Text with no
such construct is returned byte-identical with an empty report.
Raises :class:`~llm_ingestion_guard.contract.OversizeInputError` above
``max_input_chars``. A partially defanged artifact is the worst outcome
available here: it *looks* neutralized, and the live constructs are all in
the tail nobody re-reads.
"""
assert_within_input_cap(text, surface="neutralize", max_input_chars=max_input_chars)
report = Report()
out = text
def _flag(label: str, severity: Severity, count: int, evidence: str) -> None:
report.add(Finding(
label=label, severity=severity, source=source, detector="neutralize",
count=count, evidence=redact(evidence), owasp="LLM05",
))
# 1. Markdown images — the zero-click auto-fetch primitive (EchoLeak). Run
# first so the leading `!` is consumed before the inline-link pass.
img_ev: list[str] = []
def _img(m: re.Match[str]) -> str:
defanged = defang_url(m.group("url"))
img_ev.append(defanged)
return f'![{m.group("alt")}]({defanged}{m.group("title")})'
out, n_img = MD_IMAGE_RE.subn(_img, out)
if n_img:
_flag("neutralize:markdown-image", Severity.HIGH, n_img, img_ev[0])
# 2. Markdown inline links — clickable / prefetchable exfil target.
link_ev: list[str] = []
def _link(m: re.Match[str]) -> str:
defanged = defang_url(m.group("url"))
link_ev.append(defanged)
return f'[{m.group("text")}]({defanged}{m.group("title")})'
out, n_link = MD_LINK_RE.subn(_link, out)
if n_link:
_flag("neutralize:markdown-link", Severity.MEDIUM, n_link, link_ev[0])
# 3. Reference-style link definitions — the documented image-filter bypass.
ref_ev: list[str] = []
def _ref(m: re.Match[str]) -> str:
defanged = defang_url(m.group("url"))
ref_ev.append(defanged)
return m.group("pre") + defanged
out, n_ref = MD_REFDEF_RE.subn(_ref, out)
if n_ref:
_flag("neutralize:reference-link", Severity.MEDIUM, n_ref, ref_ev[0])
# 4. Angle-bracket autolinks.
auto_ev: list[str] = []
def _auto(m: re.Match[str]) -> str:
defanged = defang_url(m.group("url"))
auto_ev.append(defanged)
return f"<{defanged}>"
out, n_auto = AUTOLINK_RE.subn(_auto, out)
if n_auto:
_flag("neutralize:autolink", Severity.MEDIUM, n_auto, auto_ev[0])
# 5. Raw active HTML — escape so a renderer shows it as inert literal text.
html_state = {"count": 0, "ev": ""}
def _html(m: re.Match[str]) -> str:
tag = m.group(0)
if not is_active_tag(m.group("name"), m.group("attrs") or ""):
return tag
html_state["count"] += 1
if not html_state["ev"]:
html_state["ev"] = tag
inert = URL_IN_TEXT_RE.sub(lambda u: defang_url(u.group(0)), tag)
return inert.replace("<", "&lt;").replace(">", "&gt;")
out = HTML_TAG_RE.sub(_html, out)
if html_state["count"]:
_flag("neutralize:raw-html", Severity.HIGH, html_state["count"], html_state["ev"])
# 6. Standalone `data:` URIs left in prose (those inside constructs above are
# already defanged; the literal `data:` colon is gone, so no double count).
data_ev: list[str] = []
def _data(m: re.Match[str]) -> str:
defanged = defang_url(m.group(0))
data_ev.append(defanged)
return defanged
out, n_data = DATA_URI_RE.subn(_data, out)
if n_data:
_flag("neutralize:data-uri", Severity.HIGH, n_data, data_ev[0])
return NeutralizeResult(text=out, report=report)