fix(sanitize,okf,active_content): three quadratic patterns, two on the input path
The generalised sweep found what 0.3.2's hand-written rows missed. All three are
the documented class -- a run in front of a required literal that never arrives,
so every start position rescans the tail -- and all three are worse than the
0.3.3 findings, because `sanitize`, `neutralize`, `scan_active_content` and the
okf link graph apply NO input cap. `scan_lexicon`/`scan_output` are the only
entry points that do, so there is no ceiling to extrapolate to.
sanitize._HTML_COMMENT_RE `<!--`*100_000 20.1s, exponent 1.96-2.14
active_content.URL_IN_TEXT_RE `<a `+`A`*100_000 12.99s / 14.9s, exponent ~2.0
okf._MD_LINK_RE `[`*100_000 7.1s, exponent 1.99-2.05
Each fix is the one the pattern's own shape allows, not a copied choice:
- The comment stripper drops the regex for `str.find`. Excluding `<` would lose
every comment containing markup; bounding the run would be a carrier bypass
of the exact construct the stripper exists to remove.
- `URL_IN_TEXT_RE` bounds its scheme run to an RFC 3986 scheme (`{0,63}`).
Bounding is safe *here* only because it is a defanger inside a tag already
flagged `active:raw-html`. A lookbehind was measured too and rejected: it
drops `-http://evil.com`, a one-character evasion. Bounded: 0.185s at 1M.
- `_MD_LINK_RE` excludes `[`, matching `active_content.MD_LINK_RE` exactly,
including the nested-label trade already documented there.
`sanitize` claimed "no catastrophic backtracking" in a comment; that claim was
wrong in the same way `output`'s was before 0.3.2, and is corrected in place.
676 tests (+10), coverage 128/128 + 6/6 gaps, sweep clean across 150 patterns.
The okf destination run gets no row: `[^)\s]+` cannot fail, so a row for it
could never go red.
This commit is contained in:
parent
abbfe5f0fd
commit
73fa1b99ae
9 changed files with 223 additions and 7 deletions
|
|
@ -11,6 +11,8 @@ empty report; only active-content constructs are ever rewritten. Mutation lives
|
|||
here, kept separate from the report-only output gate (design principles 3 & 4).
|
||||
The transform is pure ``text -> (defanged_text, report)`` — no I/O, no globals.
|
||||
"""
|
||||
import time
|
||||
|
||||
from llm_ingestion_guard.neutralize import neutralize
|
||||
from llm_ingestion_guard.report import Severity, Source
|
||||
|
||||
|
|
@ -139,3 +141,23 @@ def test_prose_with_lone_brackets_and_angles_is_identical():
|
|||
result = neutralize(text)
|
||||
assert result.text == text
|
||||
assert result.report.found is False
|
||||
|
||||
|
||||
# --- self-safety (OWASP LLM10): the long-attribute arm -----------------------
|
||||
# Second call site of the same defect pinned in test_active_content.py: the
|
||||
# defanger runs `URL_IN_TEXT_RE` over each active tag's body. 14.9s at 100_000
|
||||
# chars, exponent 1.91-2.22. `neutralize` applies no input cap either.
|
||||
_ATTR_REDOS_N = 100_000
|
||||
|
||||
|
||||
def test_crafted_long_attribute_tag_stays_bounded():
|
||||
payload = "<a " + "A" * _ATTR_REDOS_N + ">"
|
||||
start = time.monotonic()
|
||||
neutralize(payload)
|
||||
assert time.monotonic() - start < 2.0
|
||||
|
||||
|
||||
def test_url_defanging_inside_a_tag_survives_the_redos_fix():
|
||||
result = neutralize("<a href=-http://evil.com>x</a>")
|
||||
assert "hxxp" in result.text
|
||||
assert "http://evil.com" not in result.text
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue