The scanners cap by truncating: they return findings, so reading a prefix costs detection in the tail and nothing else. The three transform surfaces return *content*, where the same move is not available — a shortened document is silent data loss, and a transformed prefix followed by an untransformed tail is a bypass, since the attacker chooses where in the document the payload sits. So they fail secure instead. Above MAX_INPUT_CHARS (1 000 000) sanitize, fence and neutralize raise OversizeInputError. sanitize is step 1 of prepare_input and only ever removes, so that one refusal bounds the whole input path. OversizeInputError subclasses ContractViolation: a pipeline already bracketing its quarantined stage keeps failing closed rather than meeting a type it has never heard of. It inherits the alert-routable property too — sizes in the message, refusing surface in details, no input in either. Invariant now pinned across all three: returned text is always fully transformed, or not returned at all. Still uncapped and recorded in LIMITATIONS: scan_active_content called directly (through scan_output it inherits that cap) and the okf link graph. Both are detection-shaped, so truncate-and-flag transfers unchanged — mechanical, not policy. 699 tests (+23), coverage 128/128 + 6/6, ReDoS sweep 0 candidates / 150.
174 lines
6.8 KiB
Python
174 lines
6.8 KiB
Python
"""neutralize — opt-in, pure defang of active content in model OUTPUT.
|
|
|
|
Query-time guardrails guard the answer; this guards the *persisted artifact*.
|
|
When model output is written to a wiki, doc, or knowledge base and later rendered,
|
|
active-content constructs become an exfiltration channel: a markdown image URL is
|
|
auto-fetched the moment the page renders, leaking whatever the attacker packed
|
|
into it — with no click. This is the EchoLeak class (CVE-2025-32711). Such
|
|
carriers are neither injection strings nor high-entropy blobs, so ``lexicon`` and
|
|
``entropy`` do not see them; neutralizing them is a distinct control (OWASP
|
|
LLM05 — Improper Output Handling).
|
|
|
|
Defang, don't delete. Each active construct is rewritten to an inert but still
|
|
human-auditable form: URLs get a non-resolvable scheme and bracketed dots
|
|
(``https://evil.com`` -> ``hxxps://evil[.]com``), and raw active HTML is escaped
|
|
so a renderer shows it as literal text instead of executing it. The visible
|
|
information survives review; only the machine-actionable affordance dies.
|
|
|
|
The pattern table this mutator rewrites is shared with the report-only detector
|
|
:func:`~llm_ingestion_guard.active_content.scan_active_content` and lives in
|
|
``active_content`` — detection feeds the standard gate; defanging stays the
|
|
separate, opt-in mutation below.
|
|
|
|
Two properties are load-bearing and mirror the sanitizer:
|
|
|
|
1. **Opt-in and separate.** Calling this function *is* the opt-in to mutate.
|
|
Detection elsewhere in the library stays report-only (design principles 3 & 4);
|
|
the report-only output gate (``output``) never rewrites. A caller that wants
|
|
findings without mutation reads ``result.report`` and discards ``result.text``.
|
|
2. **Byte-identical on clean input.** Output with no active construct is returned
|
|
unchanged with an empty report. Benign inline formatting (``<b>``, ``<em>``)
|
|
and prose containing stray ``<``, ``>``, ``[`` are left untouched.
|
|
|
|
Scope note (conceded, not hidden): this is a targeted defanger, not a full HTML
|
|
sanitizer. Text *between* escaped tags (e.g. a ``<script>`` body) is neutralized
|
|
as active content by escaping the tags, but bare URLs left in that residual text
|
|
stay visible; balanced-parenthesis link URLs are matched conservatively. The goal
|
|
is to kill the zero-click auto-fetch/execute affordance, not to rewrite every URL.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import re
|
|
from dataclasses import dataclass
|
|
|
|
from .active_content import (
|
|
AUTOLINK_RE,
|
|
DATA_URI_RE,
|
|
HTML_TAG_RE,
|
|
MD_IMAGE_RE,
|
|
MD_LINK_RE,
|
|
MD_REFDEF_RE,
|
|
URL_IN_TEXT_RE,
|
|
defang_url,
|
|
is_active_tag,
|
|
redact,
|
|
)
|
|
from .calibration import MAX_INPUT_CHARS
|
|
from .contract import assert_within_input_cap
|
|
from .report import Finding, Report, Severity, Source
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class NeutralizeResult:
|
|
"""The defanged text plus a report of every construct that was neutralized."""
|
|
|
|
text: str
|
|
report: Report
|
|
|
|
|
|
def neutralize(
|
|
text: str,
|
|
source: Source = Source.OUTPUT,
|
|
max_input_chars: int = MAX_INPUT_CHARS,
|
|
) -> NeutralizeResult:
|
|
"""Defang active-content constructs in ``text`` and report each class.
|
|
|
|
Rewrites markdown images/links, reference-link definitions, angle-bracket
|
|
autolinks, raw active HTML, and ``data:`` URIs into inert forms. Text with no
|
|
such construct is returned byte-identical with an empty report.
|
|
|
|
Raises :class:`~llm_ingestion_guard.contract.OversizeInputError` above
|
|
``max_input_chars``. A partially defanged artifact is the worst outcome
|
|
available here: it *looks* neutralized, and the live constructs are all in
|
|
the tail nobody re-reads.
|
|
"""
|
|
assert_within_input_cap(text, surface="neutralize", max_input_chars=max_input_chars)
|
|
report = Report()
|
|
out = text
|
|
|
|
def _flag(label: str, severity: Severity, count: int, evidence: str) -> None:
|
|
report.add(Finding(
|
|
label=label, severity=severity, source=source, detector="neutralize",
|
|
count=count, evidence=redact(evidence), owasp="LLM05",
|
|
))
|
|
|
|
# 1. Markdown images — the zero-click auto-fetch primitive (EchoLeak). Run
|
|
# first so the leading `!` is consumed before the inline-link pass.
|
|
img_ev: list[str] = []
|
|
|
|
def _img(m: re.Match[str]) -> str:
|
|
defanged = defang_url(m.group("url"))
|
|
img_ev.append(defanged)
|
|
return f'})'
|
|
|
|
out, n_img = MD_IMAGE_RE.subn(_img, out)
|
|
if n_img:
|
|
_flag("neutralize:markdown-image", Severity.HIGH, n_img, img_ev[0])
|
|
|
|
# 2. Markdown inline links — clickable / prefetchable exfil target.
|
|
link_ev: list[str] = []
|
|
|
|
def _link(m: re.Match[str]) -> str:
|
|
defanged = defang_url(m.group("url"))
|
|
link_ev.append(defanged)
|
|
return f'[{m.group("text")}]({defanged}{m.group("title")})'
|
|
|
|
out, n_link = MD_LINK_RE.subn(_link, out)
|
|
if n_link:
|
|
_flag("neutralize:markdown-link", Severity.MEDIUM, n_link, link_ev[0])
|
|
|
|
# 3. Reference-style link definitions — the documented image-filter bypass.
|
|
ref_ev: list[str] = []
|
|
|
|
def _ref(m: re.Match[str]) -> str:
|
|
defanged = defang_url(m.group("url"))
|
|
ref_ev.append(defanged)
|
|
return m.group("pre") + defanged
|
|
|
|
out, n_ref = MD_REFDEF_RE.subn(_ref, out)
|
|
if n_ref:
|
|
_flag("neutralize:reference-link", Severity.MEDIUM, n_ref, ref_ev[0])
|
|
|
|
# 4. Angle-bracket autolinks.
|
|
auto_ev: list[str] = []
|
|
|
|
def _auto(m: re.Match[str]) -> str:
|
|
defanged = defang_url(m.group("url"))
|
|
auto_ev.append(defanged)
|
|
return f"<{defanged}>"
|
|
|
|
out, n_auto = AUTOLINK_RE.subn(_auto, out)
|
|
if n_auto:
|
|
_flag("neutralize:autolink", Severity.MEDIUM, n_auto, auto_ev[0])
|
|
|
|
# 5. Raw active HTML — escape so a renderer shows it as inert literal text.
|
|
html_state = {"count": 0, "ev": ""}
|
|
|
|
def _html(m: re.Match[str]) -> str:
|
|
tag = m.group(0)
|
|
if not is_active_tag(m.group("name"), m.group("attrs") or ""):
|
|
return tag
|
|
html_state["count"] += 1
|
|
if not html_state["ev"]:
|
|
html_state["ev"] = tag
|
|
inert = URL_IN_TEXT_RE.sub(lambda u: defang_url(u.group(0)), tag)
|
|
return inert.replace("<", "<").replace(">", ">")
|
|
|
|
out = HTML_TAG_RE.sub(_html, out)
|
|
if html_state["count"]:
|
|
_flag("neutralize:raw-html", Severity.HIGH, html_state["count"], html_state["ev"])
|
|
|
|
# 6. Standalone `data:` URIs left in prose (those inside constructs above are
|
|
# already defanged; the literal `data:` colon is gone, so no double count).
|
|
data_ev: list[str] = []
|
|
|
|
def _data(m: re.Match[str]) -> str:
|
|
defanged = defang_url(m.group(0))
|
|
data_ev.append(defanged)
|
|
return defanged
|
|
|
|
out, n_data = DATA_URI_RE.subn(_data, out)
|
|
if n_data:
|
|
_flag("neutralize:data-uri", Severity.HIGH, n_data, data_ev[0])
|
|
|
|
return NeutralizeResult(text=out, report=report)
|