feat(sanitize,fence,neutralize): reject oversize input instead of half-transforming it
The scanners cap by truncating: they return findings, so reading a prefix costs detection in the tail and nothing else. The three transform surfaces return *content*, where the same move is not available — a shortened document is silent data loss, and a transformed prefix followed by an untransformed tail is a bypass, since the attacker chooses where in the document the payload sits. So they fail secure instead. Above MAX_INPUT_CHARS (1 000 000) sanitize, fence and neutralize raise OversizeInputError. sanitize is step 1 of prepare_input and only ever removes, so that one refusal bounds the whole input path. OversizeInputError subclasses ContractViolation: a pipeline already bracketing its quarantined stage keeps failing closed rather than meeting a type it has never heard of. It inherits the alert-routable property too — sizes in the message, refusing surface in details, no input in either. Invariant now pinned across all three: returned text is always fully transformed, or not returned at all. Still uncapped and recorded in LIMITATIONS: scan_active_content called directly (through scan_output it inherits that cap) and the okf link graph. Both are detection-shaped, so truncate-and-flag transfers unchanged — mechanical, not policy. 699 tests (+23), coverage 128/128 + 6/6, ReDoS sweep 0 candidates / 150.
This commit is contained in:
parent
adf93e47fb
commit
2d98d6809d
10 changed files with 272 additions and 21 deletions
|
|
@ -47,6 +47,17 @@ ENTROPY_HEX_FLOOR_LEN = 64
|
|||
# measured to do before the ReDoS fix (see active_content's pattern-table note).
|
||||
MAX_SCAN_CHARS = 1_000_000
|
||||
|
||||
# --- transform surfaces: input cap ------------------------------------------
|
||||
# The same size, and deliberately NOT the same constant, because the two caps
|
||||
# buy different things and may need to move apart. MAX_SCAN_CHARS truncates: a
|
||||
# scanner returns findings, so reading the prefix costs *detection* on the tail
|
||||
# and nothing else. `sanitize` / `fence` / `neutralize` return content, so the
|
||||
# equivalent move would hand back either a shortened document (silent data loss)
|
||||
# or an untransformed tail (a bypass an attacker positions the payload into).
|
||||
# They reject instead — see contract.OversizeInputError. Held at 1M so a
|
||||
# document accepted by the input path is one the scanners can also read whole.
|
||||
MAX_INPUT_CHARS = 1_000_000
|
||||
|
||||
# --- output: secret-egress self-safety --------------------------------------
|
||||
# Longest password a connection-string pattern will match. A bound is required
|
||||
# (not merely nice) because the run sits in front of a mandatory `@`: unbounded,
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue