1
0
Fork 0

fix(output): 19 quadratic regex runs on the output path, worst ~5.7h at the cap

The output gate claimed LLM10 self-safety on the grounds that its patterns have
no nested quantifiers. True, and irrelevant: nesting is not what makes these
blow up. A run in front of a REQUIRED literal, reachable from a short anchor, is
enough -- crafted input repeats the anchor and never supplies the literal, so
every start position rescans the tail. Quadratic, not exponential, and the
max_scan_chars cap does not help: it bounds the input, and quadratic work on a
bounded input is still hours.

Measured, not argued. `<a:` x 100_000 took 23.4s in AUTOLINK_RE alone; the
composed gate on that payload took 458.7s, extrapolating to ~5.7 hours at the
1_000_000-char input the gate itself accepts. Size-matched ordinary prose runs
0.31s, so the separation is 18x-660x -- unlike the blob in the neighbouring
test, which is the *faster* side of prose and never exercised backtracking.

Two fixes, chosen per pattern rather than uniformly:

- active_content + lexicon JSON (15 runs): exclude the character that opens the
  pattern's own anchor (`[` for markdown, `<` for tags), so a run cannot reach
  past the next start position and the per-start costs telescope. Verified to
  cost no recall: long URLs, long alt text, and `<` inside a quoted attribute
  all still match. Bounding instead would have been linear too but wrong here --
  the content is attacker-controlled, so padding past a bound would be a
  one-line bypass of the EchoLeak class this table exists to catch.
- connstr egress (4 runs): bound the password at MAX_CONNSTR_VALUE. The
  exclusion fix is unavailable -- the anchor character is `/` and passwords
  containing `/` are the common case (measured: they match today). The residual
  miss is a credential over 256 chars; a token that long is still caught by
  egress:jwt-token.

hybrid-xss:script-tag had neither option: its run is the script BODY, which may
legitimately contain `<`. It now matches the opening tag and drops the
`</script>` requirement. That also closes a fail-open -- `<script>alert(1)`
unclosed was silently missed -- at the cost of flagging prose that merely
mentions `<script>`, now documented.

Found by the composed-gate test staying red after every individual scanner was
already linear: the lexicon's six html-obfuscation patterns were the remaining
813x. A per-scanner test alone would have shipped that.

662 passed (was 642), and faster than before the fix.
This commit is contained in:
Kjell Tore Guttormsen 2026-07-31 18:31:58 +02:00
commit cff043787d
10 changed files with 220 additions and 30 deletions

View file

@ -46,8 +46,19 @@ length. The report is meant to be logged; it must not become the leak.
**Self-safety (OWASP LLM10).** The output is capped once to ``max_scan_chars``
and a single ``output:oversize-input`` finding is emitted if it was truncated;
every sub-scanner then sees bounded input. The egress patterns are linear
(anchored prefixes / negated character classes no nested quantifiers).
every sub-scanner then sees bounded input.
Bounded input is not by itself bounded runtime, and this module used to claim it
was. The egress patterns have no nested quantifiers that part was true but
absence of nesting does not imply linearity. A run in front of a *required*
literal (here: the password run before ``@``) makes every start position rescan
the tail when the literal never arrives, which is quadratic in the scanned
length. Crafted input repeating ``redis://:`` measured 8.2s at 100_000 chars and
extrapolated to hours at the 1_000_000-char cap this gate itself accepts. The
connection-string runs are therefore bounded to
:data:`~llm_ingestion_guard.calibration.MAX_CONNSTR_VALUE`; the same defect in
the active-content table is fixed there by excluding the anchor character. Both
are pinned by ``tests/test_output.py::test_crafted_redos_payload_stays_bounded``.
"""
from __future__ import annotations
@ -56,6 +67,7 @@ from dataclasses import dataclass, replace
from typing import Optional, Union
from .active_content import scan_active_content
from .calibration import MAX_CONNSTR_VALUE
from .entropy import scan_entropy
from .lexicon import MAX_SCAN_CHARS, scan_lexicon
from .report import Finding, Report, Severity, Source
@ -122,19 +134,25 @@ _SECRET_PATTERNS: list[_SecretPattern] = [
_SecretPattern("pkcs8-private-key", _p(r"-{5}BEGIN PRIVATE KEY-{5}"),
Severity.CRITICAL, "PEM PKCS#8 private key header"),
# --- DB connection strings (suppress placeholder passwords) -------------
# The password run is bounded at MAX_CONNSTR_VALUE per the ReDoS note on
# _SECRET_PATTERNS above. Unlike the active-content table, excluding the
# anchor character is NOT available here: the anchor opens with `/`, and a
# password containing `/` is the common case (a base64-ish secret), so
# excluding it would drop real credentials. The bound is the lesser loss.
_SecretPattern("postgres-connstr",
_p(r"postgres(?:ql)?://[^:@\s]+:(?P<val>[^@\s]+)@[^\s'\"]+"),
_p(r"postgres(?:ql)?://[^:@\s]+:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
Severity.CRITICAL, "PostgreSQL connection string with credentials",
value_group="val"),
_SecretPattern("mongodb-connstr",
_p(r"mongodb(?:\+srv)?://[^:@\s]+:(?P<val>[^@\s]+)@[^\s'\"]+"),
_p(r"mongodb(?:\+srv)?://[^:@\s]+:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
Severity.CRITICAL, "MongoDB connection string with credentials",
value_group="val"),
_SecretPattern("mysql-connstr",
_p(r"mysql(?:2)?://[^:@\s]+:(?P<val>[^@\s]+)@[^\s'\"]+"),
_p(r"mysql(?:2)?://[^:@\s]+:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
Severity.CRITICAL, "MySQL/MariaDB connection string with credentials",
value_group="val"),
_SecretPattern("redis-connstr", _p(r"redis://:(?P<val>[^@\s]+)@[^\s'\"]+"),
_SecretPattern("redis-connstr",
_p(r"redis://:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
Severity.HIGH, "Redis connection string with password",
value_group="val"),
# --- JWT (high false-positive rate -> MEDIUM, flag for review) ----------