1
0
Fork 0

test(redos): the dead row had the wrong payload, and the last wall clock is retired

Both rows that could not go red are decided, each by measurement.

test_lexicon.py::test_redos_pathological_subagent_input_returns_fast is REVIVED,
not retired. The row was not dead because the seed form is safe -- it was dead
because both earlier payloads made the prefix match at ONE start position, and
the cost is per-prefix-match. Repeating `spawn an agent that ` instead makes it
match K times, each driving its own O(N) lazy scan for a keyword never supplied:
K x O(N) against the seed's `(?:.*?\s+)?`, K x O(1) against the shipped
`{0,12}?` bound. Measured through scan_lexicon at 1500/3000/6000/12000 words:
seed 0.091/0.283/1.085/4.091s (exponent 1.92), shipped 0.047/0.051/0.094/0.190s
(exponent 1.01). Verified red with the seed form patched in: 4.21s against the
2.0s bound. The nesting the old comment blamed was a red herring.

test_output.py::test_pathological_input_returns_within_a_bound moves to CPU time
with a 20.0s bound, and the "or a hang" half of its claim is retired. The wall
clock was kept because a blocking hang burns no CPU -- true in general, and
inapplicable to a path with no open(), socket, subprocess, thread, lock or sleep
anywhere on it. Same payload, idle vs ~4x oversubscription: wall 3.30 -> 21.63s
(2x over the old 10.0s bound), cpu 3.30 -> 7.62s. It guarded a mode it could not
have while paying the full false-red premium. No in-repo vulnerable form can
turn this row red, so the bound was proved live against what it actually guards
-- a future pattern quadratic on long runs, `A+\s*EXFILTRATE` -- which failed it
at 64.77s CPU, 3.2x over.

redos_clock.py and the clock's pin test both documented this row as the
deliberate wall-clock holdout; both corrected.

792 passed, 6/6 documented gaps hold.
This commit is contained in:
Kjell Tore Guttormsen 2026-08-13 21:47:43 +02:00
commit 2466d260d3
3 changed files with 78 additions and 47 deletions

View file

@ -332,23 +332,50 @@ def test_no_double_oversize_flag_from_lexicon():
def test_pathological_input_returns_within_a_bound():
# A scanner that hangs on crafted input IS the DoS. This bounds the runtime
# so a hang or a blowup fails loudly; it is NOT a throughput regression test.
# The name overstates what the payload proves: measured against size-matched
# The composed gate terminates on a full-cap payload. It is NOT a ReDoS row
# and NOT a throughput regression test: measured against size-matched
# ordinary prose this blob is the FASTER side (0.93x / 0.96x, order swapped),
# so it does not exercise catastrophic backtracking. That duty is carried by
# test_lexicon.py::test_redos_pathological_subagent_input_returns_fast, which
# crafts against a known-bad nested `.*?` pattern.
# Bound set from measurement, not preference: the slowest legitimate run of
# this size is ordinary prose on a cold process (~4.2s); the blob itself runs
# 4.70s cold / 3.40-3.73s warm. The old 5.0s sat ~6% over that and failed on
# a loaded machine. 10.0s is ~2.4x the slowest observed legitimate run.
# The payload is 1_000_200 chars -- 200 over the max_scan_chars default, so
# this also exercises the truncate-and-flag oversize path. Do not resize it.
# so it exercises no catastrophic backtracking. That duty is carried by the
# crafted table below and by test_lexicon.py. What is unique here is the size:
# 1_000_200 chars, 200 over the max_scan_chars default, so this also drives
# the truncate-and-flag oversize path. Do not resize it.
#
# THE WALL CLOCK IS GONE, and the "or a hang" half of the old claim with it.
# It was kept on `time.monotonic()` on the grounds that a BLOCKING hang burns
# no CPU and only a wall clock catches it. True in general, and inapplicable
# here: `scan_output` is pure `re` over an in-memory `str` -- no open(), no
# socket, no subprocess, no threading, no lock, no sleep anywhere on the path
# (`urllib.parse` is string splitting). There is no way for this code to stop
# without spending cycles, so the wall clock guarded a mode that cannot occur
# while measurably producing false red. Measured on this machine, same
# payload, idle vs 48 busy processes (~4x oversubscription on 16 logical):
#
# wall 3.30 / 3.29 / 3.13s -> 20.71 / 21.63s <- 2x OVER the old bound
# cpu 3.30 / 3.18 / 3.20s -> 7.02 / 7.62s <- bounded by SMT, ~2.4x
#
# Bound derivation on the surviving clock: slowest legitimate content of this
# size is ordinary prose (3.02-3.07s idle CPU, ~4.7s on a cold process), and
# CPU inflation under contention tops out near 2x -- 7.62s measured, flat
# beyond, for the reason `test_the_redos_clock_ignores_time_this_process_did_
# not_spend` derives. 20.0s is ~2.6x the slowest observed legitimate run and
# still catches a blowup by orders of magnitude.
#
# That last claim is measured, not extrapolated, because no in-repo
# vulnerable form can turn this row red: its payload is a blob, not a crafted
# one, so none of the quadratic patterns this suite fixed (`[`, `<a:`,
# long-attribute, the sub-agent lazy run) fire on it. What the row actually
# guards is a FUTURE pattern that is quadratic on long runs -- so that is
# what was patched in to prove the bound live: `A+\s*EXFILTRATE`, one run
# followed by a required literal the payload never supplies, the exact defect
# class 0.3.2 and the input-path sweep both fixed. The row failed at 64.77s
# CPU against the 20.0s bound, 3.2x over. Removed again after.
#
# The cost of dropping the wall clock, stated: an infinite loop in the gate
# would now hang the suite instead of failing it. That is the same trade
# `tests/redos_clock.py` documents and every other bound in this suite already
# takes; this row was the last one paying false-red premiums to opt out of it.
payload = ("A" * 5000 + " ") * 200 # ~1MB of blob-ish text
start = time.monotonic()
scan_output(payload)
assert time.monotonic() - start < 10.0
assert scan_seconds(scan_output, payload) < 20.0
# --- crafted ReDoS payloads against OUR OWN patterns (OWASP LLM10) -----------
@ -458,10 +485,11 @@ def test_the_redos_clock_ignores_time_this_process_did_not_spend():
# What this clock gives up: a scan that BLOCKS forever burns no CPU, so it
# would hang the suite instead of failing it. Acceptable here -- these
# scanners are pure regex over an in-memory string, with no I/O and no locks,
# so the only way they can be slow is by spending cycles. It is also why
# `test_pathological_input_returns_within_a_bound` above keeps a wall clock:
# that test claims to catch "a hang or a blowup", and only a wall clock
# catches the first.
# so the only way they can be slow is by spending cycles. That held for
# `test_pathological_input_returns_within_a_bound` above too, once its path
# was actually checked for something that could block; it kept a wall clock
# on the "or a hang" claim until then, and paid 21.6s against a 10.0s bound
# under load for a mode it could not have.
#
# A sleep is the defect class at its purest: wall-clock seconds this process
# did not spend. 0.4s is 4x the assertion, so this cannot pass by timing luck.