fix(calibration): grade active content on URL shape, not construct type
v0.3.0 made the untrusted upload path unusable: measured on both doors, an ordinary remote image fail_secure'd and an ordinary link/autolink/refdef quarantined, so only documents without external references persisted. Two independent defects compounded; neither fix works alone: 1. `markdown-image: HIGH` fired on any external image. The exfil primitive is a URL that moves bytes outward, not an image. `is_ordinary_url` now grades on shape - http(s)/protocol-relative, no query, no userinfo, no percent-escape, no opaque host label or path segment -> LOW; anything data-carrying keeps the carrier's severity. raw-html and data: URIs stay HIGH unconditionally. Opacity reuses entropy's primitives; floors calibrated against real doc URLs (worst legit token H=4.08, exfil segments 4.36-4.54) and frozen in calibration. 2. The quarantine_default floor fired on ANY finding, a premise that broke when every ordinary link became a finding. It now fires at MEDIUM+ - a no-op for every detector that shipped before 0.3.0 (no LOW/INFO exists), which is what makes this a patch rather than a minor. The corpus blind spot that let this pass 522 green tests is closed: the FP corpus carries realistic markdown and is asserted on the OUTPUT gate under PRESET_USER_UPLOAD, with a counter-corpus of exfil-shaped URLs that must still block. Beaconing and short opaque segments are conceded in LIMITATIONS and asserted by the coverage matrix rather than papered over. No new public API; no new preset (0.4.0 work); allow_reserved default unchanged.
This commit is contained in:
parent
da7421e6c8
commit
6e9b8168e3
13 changed files with 533 additions and 46 deletions
|
|
@ -183,11 +183,20 @@ def _base_disposition(
|
|||
disposition = Disposition.WARN
|
||||
reasons.append(f"{max_sev.value} -> WARN")
|
||||
|
||||
# quarantine_default floor (upload preset): any finding is held for review.
|
||||
if policy.quarantine_default and report.found:
|
||||
# quarantine_default floor: a finding at MEDIUM+ is held for review.
|
||||
#
|
||||
# Through 0.3.0 this floor fired on *any* finding, on the premise that a
|
||||
# finding is the exception. Adding the active-content detector broke that
|
||||
# premise — every ordinary markdown link became a finding — and the floor
|
||||
# then quarantined documents whose only sin was linking somewhere. Raising it
|
||||
# to MEDIUM+ restores the intent (hold what is actually suspicious) and is a
|
||||
# no-op for every detector that existed before 0.3.0: none of them emit LOW.
|
||||
if policy.quarantine_default and max_sev is not None and (
|
||||
severity_rank(max_sev) >= severity_rank(Severity.MEDIUM)
|
||||
):
|
||||
floored = _more_severe(disposition, Disposition.QUARANTINE_REVIEW)
|
||||
if floored is not disposition:
|
||||
reasons.append("quarantine-floor: untrusted upload, any finding -> QUARANTINE_REVIEW")
|
||||
reasons.append("quarantine-floor: MEDIUM+ finding -> QUARANTINE_REVIEW")
|
||||
disposition = floored
|
||||
|
||||
return disposition
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue