fix(output): 19 quadratic regex runs on the output path, worst ~5.7h at the cap
The output gate claimed LLM10 self-safety on the grounds that its patterns have no nested quantifiers. True, and irrelevant: nesting is not what makes these blow up. A run in front of a REQUIRED literal, reachable from a short anchor, is enough -- crafted input repeats the anchor and never supplies the literal, so every start position rescans the tail. Quadratic, not exponential, and the max_scan_chars cap does not help: it bounds the input, and quadratic work on a bounded input is still hours. Measured, not argued. `<a:` x 100_000 took 23.4s in AUTOLINK_RE alone; the composed gate on that payload took 458.7s, extrapolating to ~5.7 hours at the 1_000_000-char input the gate itself accepts. Size-matched ordinary prose runs 0.31s, so the separation is 18x-660x -- unlike the blob in the neighbouring test, which is the *faster* side of prose and never exercised backtracking. Two fixes, chosen per pattern rather than uniformly: - active_content + lexicon JSON (15 runs): exclude the character that opens the pattern's own anchor (`[` for markdown, `<` for tags), so a run cannot reach past the next start position and the per-start costs telescope. Verified to cost no recall: long URLs, long alt text, and `<` inside a quoted attribute all still match. Bounding instead would have been linear too but wrong here -- the content is attacker-controlled, so padding past a bound would be a one-line bypass of the EchoLeak class this table exists to catch. - connstr egress (4 runs): bound the password at MAX_CONNSTR_VALUE. The exclusion fix is unavailable -- the anchor character is `/` and passwords containing `/` are the common case (measured: they match today). The residual miss is a credential over 256 chars; a token that long is still caught by egress:jwt-token. hybrid-xss:script-tag had neither option: its run is the script BODY, which may legitimately contain `<`. It now matches the opening tag and drops the `</script>` requirement. That also closes a fail-open -- `<script>alert(1)` unclosed was silently missed -- at the cost of flagging prose that merely mentions `<script>`, now documented. Found by the composed-gate test staying red after every individual scanner was already linear: the lexicon's six html-obfuscation patterns were the remaining 813x. A per-scanner test alone would have shipped that. 662 passed (was 642), and faster than before the fix.
This commit is contained in:
parent
8deca93ee1
commit
cff043787d
10 changed files with 220 additions and 30 deletions
|
|
@ -40,10 +40,25 @@ ENTROPY_HEX_FLOOR_LEN = 64
|
|||
|
||||
# --- lexicon: self-safety + variant thresholds ------------------------------
|
||||
# Input-size cap (OWASP LLM10): large enough for a real ingested document;
|
||||
# beyond it the scanner reads the prefix and flags, so runtime stays bounded
|
||||
# even on a decompression-bomb-sized input.
|
||||
# beyond it the scanner reads the prefix and flags, so every sub-scanner sees a
|
||||
# bounded input. Note what this cap does NOT buy: bounded input is only bounded
|
||||
# runtime if the patterns are linear in it. A quadratic pattern turns this cap
|
||||
# into hours of work, which is what crafted input against the output path was
|
||||
# measured to do before the ReDoS fix (see active_content's pattern-table note).
|
||||
MAX_SCAN_CHARS = 1_000_000
|
||||
|
||||
# --- output: secret-egress self-safety --------------------------------------
|
||||
# Longest password a connection-string pattern will match. A bound is required
|
||||
# (not merely nice) because the run sits in front of a mandatory `@`: unbounded,
|
||||
# crafted input repeating `redis://:` and never supplying the `@` makes every
|
||||
# start position rescan the tail — quadratic. Excluding the anchor character the
|
||||
# way the active-content table does is not available here, since that character
|
||||
# is `/` and passwords containing `/` are the common case.
|
||||
# 256 is generous for a password and cheap to scan; the residual miss is a
|
||||
# credential longer than this, which for the realistic case (a token used as a
|
||||
# DB password) is still caught by the jwt-token / high-specificity patterns.
|
||||
MAX_CONNSTR_VALUE = 256
|
||||
|
||||
# Minimum length before the rot13 variant is scanned — shorter strings hit
|
||||
# rot13-look-alike false positives.
|
||||
ROT13_MIN_LEN = 40
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue