fix(sanitize,okf,active_content): three quadratic patterns, two on the input path
The generalised sweep found what 0.3.2's hand-written rows missed. All three are
the documented class -- a run in front of a required literal that never arrives,
so every start position rescans the tail -- and all three are worse than the
0.3.3 findings, because `sanitize`, `neutralize`, `scan_active_content` and the
okf link graph apply NO input cap. `scan_lexicon`/`scan_output` are the only
entry points that do, so there is no ceiling to extrapolate to.
sanitize._HTML_COMMENT_RE `<!--`*100_000 20.1s, exponent 1.96-2.14
active_content.URL_IN_TEXT_RE `<a `+`A`*100_000 12.99s / 14.9s, exponent ~2.0
okf._MD_LINK_RE `[`*100_000 7.1s, exponent 1.99-2.05
Each fix is the one the pattern's own shape allows, not a copied choice:
- The comment stripper drops the regex for `str.find`. Excluding `<` would lose
every comment containing markup; bounding the run would be a carrier bypass
of the exact construct the stripper exists to remove.
- `URL_IN_TEXT_RE` bounds its scheme run to an RFC 3986 scheme (`{0,63}`).
Bounding is safe *here* only because it is a defanger inside a tag already
flagged `active:raw-html`. A lookbehind was measured too and rejected: it
drops `-http://evil.com`, a one-character evasion. Bounded: 0.185s at 1M.
- `_MD_LINK_RE` excludes `[`, matching `active_content.MD_LINK_RE` exactly,
including the nested-label trade already documented there.
`sanitize` claimed "no catastrophic backtracking" in a comment; that claim was
wrong in the same way `output`'s was before 0.3.2, and is corrected in place.
676 tests (+10), coverage 128/128 + 6/6 gaps, sweep clean across 150 patterns.
The okf destination run gets no row: `[^)\s]+` cannot fail, so a row for it
could never go red.
This commit is contained in:
parent
abbfe5f0fd
commit
73fa1b99ae
9 changed files with 223 additions and 7 deletions
|
|
@ -80,7 +80,26 @@ _SCHEME_SUBS = (
|
|||
# Dot-defang that is idempotent: never touches a `.` already inside `[.]`.
|
||||
_DOT_RE = re.compile(r"(?<!\[)\.(?!\])")
|
||||
# A bare http(s)/ftp URL embedded in other text (used inside escaped HTML).
|
||||
URL_IN_TEXT_RE = re.compile(r"[A-Za-z][A-Za-z0-9+.\-]*://[^\s'\"<>]+")
|
||||
#
|
||||
# ReDoS note (OWASP LLM10). The scheme run sits in front of a REQUIRED `://`, so
|
||||
# a long run of scheme characters that never reaches it costs a full rescan at
|
||||
# every start position: `<a ` + `A`*100_000 + `>` measured 12.99s through
|
||||
# `scan_active_content` and 14.9s through `neutralize`, exponent ~2.0 over four
|
||||
# doublings, on entry points that apply no input cap. The 0.3.2 sweep missed it
|
||||
# because its payloads repeat a unit, and this arm needs the tag to CLOSE before
|
||||
# the body is handed on.
|
||||
#
|
||||
# The exclusion trick used by the constructs below does not apply — the attack
|
||||
# repeats a plain scheme character, not this pattern's anchor — so the run is
|
||||
# bounded to an RFC 3986 scheme instead (`ALPHA *( ALPHA / DIGIT / "+" / "-" /
|
||||
# "." )`; the longest registered scheme is far under 64). Unlike the detector
|
||||
# tables, bounding costs nothing here: this is a defanger applied INSIDE a tag
|
||||
# already flagged `active:raw-html`, padding merely shifts where the match
|
||||
# starts, and a 64+ character "scheme" is not resolvable by any renderer. A
|
||||
# lookbehind that killed interior start positions was measured too and rejected:
|
||||
# it drops `-http://evil.com` and `.http://x.com`, a one-character evasion of
|
||||
# the defanger. Bounded: 0.185s at the full 1_000_000-char cap.
|
||||
URL_IN_TEXT_RE = re.compile(r"[A-Za-z][A-Za-z0-9+.\-]{0,63}://[^\s'\"<>]+")
|
||||
|
||||
|
||||
def defang_url(url: str) -> str:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue