feat(sanitize,fence,neutralize): reject oversize input instead of half-transforming it
The scanners cap by truncating: they return findings, so reading a prefix costs detection in the tail and nothing else. The three transform surfaces return *content*, where the same move is not available — a shortened document is silent data loss, and a transformed prefix followed by an untransformed tail is a bypass, since the attacker chooses where in the document the payload sits. So they fail secure instead. Above MAX_INPUT_CHARS (1 000 000) sanitize, fence and neutralize raise OversizeInputError. sanitize is step 1 of prepare_input and only ever removes, so that one refusal bounds the whole input path. OversizeInputError subclasses ContractViolation: a pipeline already bracketing its quarantined stage keeps failing closed rather than meeting a type it has never heard of. It inherits the alert-routable property too — sizes in the message, refusing surface in details, no input in either. Invariant now pinned across all three: returned text is always fully transformed, or not returned at all. Still uncapped and recorded in LIMITATIONS: scan_active_content called directly (through scan_output it inherits that cap) and the okf link graph. Both are detection-shaped, so truncate-and-flag transfers unchanged — mechanical, not policy. 699 tests (+23), coverage 128/128 + 6/6, ReDoS sweep 0 candidates / 150.
This commit is contained in:
parent
adf93e47fb
commit
2d98d6809d
10 changed files with 272 additions and 21 deletions
|
|
@ -15,6 +15,8 @@ from __future__ import annotations
|
|||
import re
|
||||
from dataclasses import dataclass
|
||||
|
||||
from .calibration import MAX_INPUT_CHARS
|
||||
from .contract import assert_within_input_cap
|
||||
from .report import Finding, Report, Severity, Source
|
||||
|
||||
# Invisible / steganographic character classes (codepoints).
|
||||
|
|
@ -29,8 +31,10 @@ _TAG_LO, _TAG_HI = 0xE0000, 0xE007F # Unicode Tags block (U+E0000–U+E007F)
|
|||
# was wrong in the same way `output`'s was: a lazy run in front of a REQUIRED
|
||||
# literal costs a full tail rescan at *every* start position when the literal
|
||||
# never arrives, so `<!--` repeated to 100_000 chars measured 20.1s (exponent
|
||||
# 1.96–2.14 over four doublings) — and this module, unlike `scan_lexicon` /
|
||||
# `scan_output`, applies no input cap, so nothing bounds that above.
|
||||
# 1.96–2.14 over four doublings) — and at the time this module, unlike
|
||||
# `scan_lexicon` / `scan_output`, applied no input cap, so nothing bounded that
|
||||
# above. MAX_INPUT_CHARS now does, but as a second line only: the cap bounds a
|
||||
# *future* quadratic pattern's damage, it does not make a quadratic one safe.
|
||||
#
|
||||
# Neither of the two fixes used elsewhere fits here. Excluding the opener (`<`)
|
||||
# from the run would drop every comment containing markup — `<!-- <b>x</b> -->`
|
||||
|
|
@ -94,8 +98,19 @@ def _decode_tags(codepoints: list[int]) -> str:
|
|||
return "".join(out)
|
||||
|
||||
|
||||
def sanitize(text: str, source: Source = Source.INPUT) -> SanitizeResult:
|
||||
"""Strip carrier classes from ``text`` and report per-class counts."""
|
||||
def sanitize(
|
||||
text: str,
|
||||
source: Source = Source.INPUT,
|
||||
max_input_chars: int = MAX_INPUT_CHARS,
|
||||
) -> SanitizeResult:
|
||||
"""Strip carrier classes from ``text`` and report per-class counts.
|
||||
|
||||
Raises :class:`~llm_ingestion_guard.contract.OversizeInputError` above
|
||||
``max_input_chars``: this is step 1 of the input path, so the refusal bounds
|
||||
the whole path, and a *partially* sanitized document is worse than none —
|
||||
the unstripped tail is where a carrier would be placed.
|
||||
"""
|
||||
assert_within_input_cap(text, surface="sanitize", max_input_chars=max_input_chars)
|
||||
report = Report()
|
||||
|
||||
# Character-class carriers: single pass, keep everything else verbatim.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue