# Changelog All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [Unreleased] Nothing yet. ## [0.6.1] — 2026-08-11 ### Fixed — the zero-width check tested identity, so emoji-composed documents were hard-blocked `_ZERO_WIDTH` (sanitize, input) and `_ZERO_WIDTH_CPS` (output, `_scan_invisible_carriers`) tested U+200D on codepoint membership alone. `disposition._CARRIER_LABELS` grades a carrier as any-tier `fail_secure` with no appeal, so **any** first-party document containing a ZWJ-composed emoji — professions, families, skin tones, flag variants — was blocked permanently, with no preset able to release it. Reported by `ms-ai-architect`, confirmed here against the code. The worse half was not in the report: sanitize *removed* the joiner, silently decomposing one emoji into two unrelated ones. A module whose published contract is "only ever removes carriers" was corrupting content. The fix is the shape our own lexicon row `unicode:zero-width-in-word` (`\w[ZW]\w`) already used — **judge the joiner by context, not identity**. A ZWJ is exempt only when *both* neighbours are emoji-context codepoints. Half-context is not context, so `a😀` stays a carrier and an attacker cannot buy exemption with a single trailing emoji. **Blocks, not an emoji table.** Measured against Unicode 17.0's `emoji-zwj-sequences.txt`: 1614 RGI sequences use 122 distinct codepoints adjacent to a ZWJ, and five block ranges cover 122/122. The measurement earned its keep — the hand-reasoned candidate table missed U+2194, U+2195 and U+2B1B. Shipping the RGI list itself would be exact the day it landed and stale at the next Unicode release, reopening this false positive for every new emoji; whole blocks carry the unassigned headroom (458 `Cn` codepoints) that future emoji are allocated into, so the table does not age. The predicate is defined once in `sanitize` and imported by `output`; a cross-surface test asserts the two halves agree on six inputs. ### Known behaviour change **Documents whose only finding was an emoji-context ZWJ now persist unattended.** On `PRESET_USER_UPLOAD` they move from `fail_secure` to `WARN`. Measured across the three false-positive populations at one corpus state: **1 document of 1126** (reference-corpus 1/389, vendor-harvest 0/187, generated-notes 0/550). This is a loosening of the *upload door*, not of detection — recall is unchanged at 128/128 demonstrated classes with 6/6 documented gaps holding, and a ZWJ anywhere else, including between a word character and an emoji, is graded exactly as before. **Patch, not minor.** 0.6.0 called itself minor for loosening the same door, but that was a policy choice — `` left the active name set by decision. This one restores a contract the module already published, against a class that was never meant to be blocked. A fix whose observable effect is the point of the fix is what the patch level is for. ### Residuals — both documented (`docs/LIMITATIONS.md`, 33 → 34 items) - **A ZWJ between two emoji is now exempt**, so it can carry a narrowband covert channel: one emoji per bit, and it cannot split a word. A deliberate narrowing, stated rather than hidden. - **U+200C (ZWNJ) still has no context test.** It is orthographically *required* in Persian, Arabic and Devanagari, so those documents stay hard-blocked. The criterion has to be script-based rather than pictographic, and no corpus is here to verify one against — parked as a known false-positive class rather than guessed at. 736 tests pass (was 727). ## [0.6.0] — 2026-08-11 ### Changed — `active:raw-html` stops firing on two things that carry no affordance `is_active_tag` had two over-reaching branches, both measured on consumer corpora rather than argued from the code: - **The URL-attribute branch was a presence test.** Any element carrying `href=`, `src=`, `action=` … graded HIGH regardless of where the URL pointed. An MDX `` — an internal doc route — reaches no attacker-controlled host, and neither does Azure APIM policy XML's ``. The branch now requires an **external** target (absolute scheme or protocol-relative), the rule the markdown paths have applied since 0.3.1. - **`` left the active name set.** HTML's `` has its entire affordance in its `href`, which the URL-attribute branch still catches. APIM's attribute-less `` means "run the inherited policy" and is inert in every renderer. **Measured before and after in one session, against one corpus state** — the two wiki corpora are living, so a before/after split across sessions would mix this change with re-harvest drift: | population | n | before | after | ceiling (raw-HTML off) | |---|---|---|---|---| | reference-corpus | 389 | 133 | **108** | 107 | | vendor-harvest | 187 | 100 | **98** | 62 | | generated-notes | 550 | 90 | **88** | 49 | 96% of the achievable reduction in reference-corpus, 5% in the two wiki corpora: the over-reach was nearly the whole raw-html cost in APIM policy XML and nearly none of it in vendor documentation, where what remains is real HTML — `` 298, `` 94, `` 63 — caught correctly by the name branch. **The two classes had to be measured together.** Alone they free 3 and 13 documents in reference-corpus; together, 25. A document carrying one usually carries the other, so closing either alone leaves it blocked by its twin. `docs/rawhtml-census.py` now carries a `PRODUCTION` row that re-measures the shipped predicate instead of a hypothetical, so a doc number and the code cannot drift apart unnoticed. ### Known behaviour change **Documents whose only finding was one of these two classes now persist unattended.** On `PRESET_USER_UPLOAD` they move from `fail_secure` / `quarantine_review` to `WARN` — 25 documents in the reference corpus, 2 in each wiki corpus. This is a deliberate loosening of the *upload door*, not of detection: recall is unchanged at 128/128 demonstrated classes with 6/6 documented gaps holding, and a tag that is active by name, carries an `on*=` handler, or points anywhere external is graded exactly as before. An element outside the active name set whose only URL attribute is doc-relative is the whole of what changed. ### Fixed — the scanner and the mutator no longer share one predicate `neutralize` imported `is_active_tag` from `active_content` by name, so narrowing the scanner would have silently narrowed the opt-in mutator as well — and **no test in the suite discriminated the two halves**: every `neutralize:raw-html` payload stayed active under each narrowing considered. The predicates are now separate symbols, `is_active_tag` (scanner, external-target rule) and `is_defangable_tag` (mutator, unchanged broad behaviour), and the mutator half is pinned by its own test. Over- defanging costs nothing there — `neutralize` is opt-in and blocks no disposition — while under-defanging would hand a human a live construct. ### Self-safety (OWASP LLM10) Reading a URL attribute's *value* needs a pattern the presence test does not provide. It reuses the same literal alternation with the value attached, so no new run shape enters the table, and `_REDOS_PAYLOADS` gains a row (`active-url-attr-value`) whose unit **denies** the `=` the pattern requires — a unit supplying it matches at once and never exercises the run. Measured at 100_000 chars: 0.031–0.046s across five attack shapes, against the suite's 2.0s bound. A gap between the two patterns fails secure: an attribute seen by the presence test but unreadable by the value parser counts as external, so it over-blocks rather than under-blocks. 727 tests pass (was 717). ## [0.5.0] — 2026-08-11 ### Added — the axis separation: assessment (`Risk`) vs action (`Disposition`) > **Additive, and measured to be so.** Every disposition 0.4.0 rendered is > rendered identically: the full suite went 703 → 715 with no test changed, the > coverage matrix holds at 128/128 recall with 6/6 documented gaps, and the > `PRESET_USER_UPLOAD` grading table locked in 0.3.1 was re-measured row by row > and is unchanged. A caller that never reads the new field sees no difference. `decide` and `guard` returned a `Disposition` — `WARN` / `QUARANTINE_REVIEW` / `FAIL_SECURE` — which names an **action**. But BRIEF design principle 4 says the library reports and the *pipeline* decides, and `disposition.py` admitted the gap in its own docstring: *"It imposes no blocking of its own."* So the library returned an action it cannot enforce, while discarding the judgement that produced it. A consumer wanting different behaviour had to reinterpret the action itself, which is why a consumer ends up pinning our *grading* — the action was all they got. - **`Risk`** — the new assessment axis: `NONE` / `LOW` / `ELEVATED` / `SEVERE`. It answers *how dangerous is this artifact given its source context*, and is trust-aware exactly as BRIEF §4.7 describes the domain: the same finding genuinely is a different judgement in authored prose than in a code fence. - **`DispositionResult.assessment`** — carries that judgement alongside the action. The field is **required, with no default**: `Risk.NONE` would be the natural-looking default and is the wrong one, since a construction site that forgot it would report *clean* and the axis would fail open. - **`Policy.action_map`** — an optional `Risk -> Disposition` override, so "hold for review where you would block" is a policy statement rather than a reason to pin our grading. Defaults to `None`, which means `DEFAULT_ACTION_MAP` and keeps an untouched `Policy` hashable as before. A partial map falls back per-level instead of raising. - Both overlays — compound escalation and the quarantine floor — now move the **assessment**, so a custom action map cannot silently drop them. - The fail-closed path in `guard` pins both axes to their most severe value and deliberately does **not** route through the action map: a policy that downgrades `SEVERE` means "I accept this class of finding", never "I accept a scanner that crashed on crafted input" (§4.6). `NONE` and `LOW` both map to `WARN`, which is the point rather than an oversight: a clean document and one carrying only low-severity findings were a single indistinguishable value through 0.4.0. ### Known limitation recorded (32, was 31) `Severity` still carries disposition intent on the *detection* side — the separation above is caller-side only. Two places say so outright: `ACTIVE_CONTENT_ORDINARY_SEVERITY = LOW` exists because grading an ordinary external image `HIGH` fail-secured ordinary uploads, and the quarantine floor was raised to `MEDIUM+` to repair the same regression from the other end. Closing it changes the grading and so fires a consumer-notification promise; deferred deliberately. See `docs/LIMITATIONS.md`. ### Measured — what the upload door costs on benign documents (33rd limitation) Every field measurement this project had published was **per URL**. None of them answered the question a consumer actually asks: *how often does an ordinary document cost me a human?* Three benign populations were run through `screen_output` under `PRESET_USER_UPLOAD` and counted at document granularity — 98 of 185 vendor-published doc pages (53.0%), 88 of 547 model-written notes (16.1%), and 133 of 389 first-party reference documents (34.2%) disposed to something other than WARN. The number is bad and is published as measured. `docs/PLAN-v1.md` committed to that in advance — *"et rødt FP-resultat er like verdifullt"* — and the response here is a documented limitation, not a recalibration: moving the grading would fire a locked consumer-notification promise, and the drivers are residuals this document already concedes rather than anything newly discovered. - **`docs/fp-sweep.py`** — the method, re-runnable, corpus roots as arguments. It refuses to print a pooled total (the populations have different provenance and different denominators) and it aborts if the default action map stops sending exactly `NONE` and `LOW` to WARN, since the published count is a statement about *assessed risk* and only equals one while that holds. - **`tests/test_corpus.py::test_the_published_fp_metric_is_a_risk_statement`** — the same equivalence, pinned in the suite. An `action_map` override is a supported feature as of the axis separation above, so without this pin a consumer-facing number could change meaning with nothing failing. ### Corrected — 0.3.3 listed two behaviour-change classes and there were three Sweeping the same population against the **v0.3.1 tag a consumer actually pins** (scratch venv, `git+file://…@v0.3.1`, resolved version asserted) returned 99 of 185 (53.5%) where the current tree returns 98 (53.0%). One document moved, and chasing it corrects a claim rather than confirming one. `docs-en-fullscreen.txt` disposed QUARANTINE_REVIEW at 0.3.1 and WARN now, because `markdown:link-anchor-injection` no longer fires on it. Under 0.3.1 that pattern matched **300 characters of ordinary prose**: it opened at a `[`, ran across intervening text containing the word *execute*, and closed at a distant `](…)` belonging to a different construct. The 0.3.3 ReDoS fix excluded `[` from the anchor class and `(` from the target class, which telescopes the runaway — and, as a side effect nobody measured at the time, deletes this false-positive class too. 0.3.3's *Known behaviour changes* said **"None measured"** and then named two exceptions: URLs with a literal `(` in the target, and comment bodies with a literal `(` before the keyword. It missed the third: an anchor can no longer span a `[`, so a match that used to bridge two separate markdown constructs no longer forms. The correction is in our favour — one fewer false positive per 185 documents of vendor documentation — but it was a behaviour change presented as none, and it took a field sweep to find it. ### Fixed - A **retracted** number was still living in a test comment. `tests/test_wiring.py` credited a consumer's capture store with 35 of 35 query-carrying URLs. That consumer retracted it the next day and re-measured 28 of 28 on the same 81-URL corpus; `docs/LIMITATIONS.md` was corrected then and the comment was not. Corrected, with the retraction written into the comment so it cannot read as a second, disagreeing measurement. - **Five current-state version claims had never been updated by any release.** The 0.4.0 release commit touched three files — `CHANGELOG.md`, `pyproject.toml`, `src/llm_ingestion_guard/__init__.py` — and deferred the README deliberately, so that the install block would not point at a tag before a clean-venv install had proven it resolved. That proof step never ran, so tag `v0.4.0` permanently carries a README advertising `v0.3.4`. The tag is not moved; the ordering is. Sweeping *every* tracked file for a version claim, rather than the four surfaces the release checklist named, found four more that no release had ever touched — plus a stale test count: - `SECURITY.md` — "The project is pre-1.0 (`0.2.x`, alpha). Only the latest published version receives fixes." The only one with a consequence for an outsider: it named a support window two minor lines behind the code. - `README.md` — `**Status:** v0.3`, stale since 0.4.0. - `docs/BRIEF.md` and `CLAUDE.md` — "v0.2 (alpha)", stale since 0.3.0. The latter also claimed 12 modules where `src/` has 15. - `docs/ADOPTION-BRIEF.md` — "**703 passing**", where the suite is at 717. Every one of them is a *current-state* claim. Measurement provenance — "New in `v0.4.0`", "verified identical on 0.2.0 and 0.3.1", "measured against the v0.3.1 tag" — is left exactly as written, because bumping those would falsify the record rather than update it. From here all current-state surfaces move in the release commit itself and are verified by `git show ` *before* the tag exists, since that is the only check the previous ordering could not perform. Found because `llm-ingestion-okf` took our report of this defect class as a hypothesis about their own repo, measured it, found a worse instance, and sent back the generalization: writing down a trap is not the same as applying it. ## [0.4.0] — 2026-08-10 > **Behaviour change, not a pure fix — and that is why this is 0.4.0 and not > 0.3.5.** The three transform surfaces gain a refusal path they did not have. A > caller that passes a document larger than 1 000 000 characters now gets an > exception where it previously got a result. Adding a raise to a function that > was previously total is breaking under SemVer whatever the measured blast > radius turns out to be, so the number follows the change, not the survey. > > **The measured blast radius, for the record: zero.** `linkedin-studio` pins an > exact tag, so nothing reaches it until it re-pins. `llm-ingestion-okf` moved to > the range `>=0.3,<0.4` (their `f536e13`), so a 0.3.5 would have landed on them > at their next resolve without an action on their part — but they answered our > query (`20260802T193351Z`) with a measured **no**: zero call sites for > `sanitize` / `fence` / `neutralize` / `prepare_input` anywhere in their `src/`. > Their `screen_output` path reaches only `scan_output`, which truncates and does > not raise. Releasing as 0.4.0 puts this outside their ceiling regardless, so > they cross it deliberately rather than by resolving. ### Added — input-size cap on the transform surfaces (OWASP LLM10) `sanitize`, `fence` and `neutralize` now raise `OversizeInputError` above `MAX_INPUT_CHARS` (1 000 000) instead of accepting text of any length. Since `sanitize` is step 1 of `prepare_input` and only ever *removes*, that single refusal bounds the whole input path. They **reject** where the scanners **truncate**, and the asymmetry is the point: - `scan_lexicon` / `scan_output` return findings. Reading a prefix costs detection in the tail and nothing else — a lossy answer, but an answer. - `sanitize` / `fence` / `neutralize` return *content*. Truncating would return a shortened document (silent data loss for anything that persists the result) or a transformed prefix followed by an untransformed tail — a bypass, since an attacker chooses where in the document the payload sits. The invariant the three now keep: **returned text is always fully transformed, or not returned at all.** `OversizeInputError` subclasses `ContractViolation`, so a pipeline already bracketing its quarantined stage in `except ContractViolation` keeps failing closed. Like its parent it is alert-routable: the message carries the size and the cap, `details` names the refusing surface, and neither carries input. `max_input_chars` is a per-call parameter, defaulting to the single calibrated constant. ### Added — the last two detection surfaces bound their input too `scan_active_content` **called directly** and `okf.link_graph` were the two surfaces still reading attacker-supplied text with no cap. Both truncate and record, the way the other scanners do: - `scan_active_content(text, source, max_scan_chars=MAX_SCAN_CHARS)` emits one `active:oversize-input` finding (MEDIUM, LLM10) and scans the prefix. Reached through `scan_output` the text is already under that surface's cap, so the flag is raised once, there — `max_scan_chars` is now passed down. - `link_graph(bundle, max_scan_chars=MAX_SCAN_CHARS)` caps each body and records `(from_id, body_length)` in the new `LinkGraphResult.truncated` field. The field is additive with a default, so existing positional construction and attribute access are unaffected. **What truncation costs is named rather than implied:** past the cap, "no finding" means "not looked at". That is precisely what a silent truncation would hide, and why `truncated` exists as a field instead of a log line — it is what separates "no links past here" from "no links *read* past here". Recorded in `docs/LIMITATIONS.md`. ## [0.3.4] — 2026-08-01 > **Denial-of-service fix on the INPUT path. Upgrade from 0.3.3.** 0.3.3 swept > the 83 lexicon patterns arm by arm and left every other table on 0.3.2's > hand-written rows. Generalising the sweep over all eleven regex-bearing modules > found three more quadratic patterns — two of them on the input path, one in > `sanitize`, the first thing every ingested document touches. No disposition > changes: recall was measured case by case and nothing was lost. Earlier tags > are not moved. ### Fixed — three quadratic patterns, two on the input path Same class as everything 0.3.2 and 0.3.3 fixed: a run in front of a **required** literal, so crafted input that never supplies the literal makes every start position rescan the tail. Each exponent is read across four doublings, not from a two-point ratio. | Pattern | Crafted payload | Measured @ 100 000 | Exponent | |---|---|---|---| | `sanitize._HTML_COMMENT_RE` | `` it replaces. Excluding `<` from the run would lose every comment containing markup (`` is the ordinary case); bounding the run would be a one-line carrier bypass of the exact construct the stripper exists to remove. The module's own "no catastrophic backtracking" comment was wrong in the same way `output`'s was before 0.3.2, and is corrected in place. - **`URL_IN_TEXT_RE`** bounds its scheme run to an RFC 3986 scheme (`{0,63}`). Bounding is safe *here* only because this is a defanger applied inside a tag already flagged `active:raw-html`, so padding shifts where the match starts rather than evading detection. A lookbehind killing interior start positions was measured too and **rejected**: it drops `-http://evil.com` and `.http://x.com`, a one-character evasion of the defanger. Bounded, the pattern runs in 0.185 s at the full 1 000 000-char cap. - **`okf._MD_LINK_RE`** excludes `[`, matching `active_content.MD_LINK_RE` exactly, including the nested-label trade already documented there. ### Changed — the sweep covers every regex surface, not one table `docs/redos-sweep.py` now sweeps **150 patterns across 11 tables** (0.3.3 covered 83 in one). The collector is mechanical on both axes so no one has to remember to list anything: it walks each module's namespace for compiled patterns, and it derives each pattern's call mode from the module source, because `.sub()` and `.finditer()` visit every start position where `.match()` cannot. A pattern reachable only through a helper parameter gets the worst mode, marked `*` — the fallback can over-measure but never miss. Two arm shapes the generator cannot express are pinned by hand as a result: a tag that *closes* around a long body (repeating-unit payloads never close it), and a run of plain characters carrying no anchor at all. The `okf` destination run gets no row on purpose: `[^)\s]+` cannot fail, so a pin for it could never go red. A lexicon candidate flagged at ×2.8 measured **linear** across four doublings (exponent 0.96–1.03) — the near-noise-floor false flag the script's own docstring warns about, confirmed a second time. 676 tests (+10). Coverage matrix unchanged at 128/128 caught, 6/6 gaps holding. ## [0.3.3] — 2026-07-31 > **Denial-of-service fix, and a correction to 0.3.2. Upgrade from 0.3.2.** The > sweep 0.3.2 shipped was incomplete, and it said otherwise. Two lexicon patterns > were still quadratic — reachable through `scan_output`, not only on the input > path. No disposition changes: recall was measured case by case and nothing was > lost. The v0.3.2 tag is not moved. ### Fixed — two quadratic patterns in the lexicon table `8deca93` scoped the remaining ReDoS duty to the lexicon path, and this is that work: all 83 patterns measured, arm by arm. Two are quadratic, same shape as everything 0.3.2 fixed — a run in front of a **required** literal, where the run may cross the pattern's own opening anchor. | Pattern (arm) | Crafted unit | Measured | At the 1 000 000-char cap | |---|---|---|---| | `markdown:link-anchor-injection` (anchor text) | `[` | 1.91 s @ 8 000 | **~8.3 hours** | | `markdown:link-anchor-injection` (URL run) | `[system](` | 0.006 s @ 8 000 | ~89 seconds | | `markdown:link-ref-comment` (`.*` run) | `[//]: # (` | 0.22 s @ 8 000 | **~1.0 hour** | Exponent measured over five points (1 000 → 16 000): **1.98** — quadratic, not exponential. Legitimate content of the same size is unaffected: 0.316 s at N=100 000 (prose 0.316 / html 0.315 / markdown 0.297 / connection-string 0.296). **These were not input-path-only, and that is the correction.** `scan_lexicon` runs on the output path too, so 0.3.2's *"the last quadratic-backtracking site on the output path"* was false when written. Measured through the public gate before this fix: `scan_output("[" * 100_000)` took **334.7 s**. The claim was too broad because the sweep behind it drove the `[` payload only through `scan_active_content` — no row ever drove it through the lexicon. The statement is corrected in `docs/LIMITATIONS.md`. The fix is anchor exclusion, per the rule `active_content` already documents — bounding attacker-controlled content would be a one-line detection bypass. The excluded character is `(`, not `[`: ``` markdown:link-anchor-injection \[[^\]\[]*(?:system|…)[^\]\[]*\]\([^)(]+\) markdown:link-ref-comment \[//\]:\s*#\s*\([^(\n]*(?:ignore|…) ``` `[` was the obvious choice and it was measurably worse. Excluding `[` from the URL run drops `[override your rules](https://[::1]/x)` — still covered, three other patterns fire on it — but excluding `[` from the link-ref comment run drops `[//]: # (see [x] then ignore this)`, which **no other pattern catches**. The anchors contain `(` as well, so excluding `(` telescopes just as effectively at zero measured recall cost. Both forms verified linear (×1.99–2.02 on doubling). ### Known behaviour changes - **None measured.** Every case that matched before still matches, except URLs containing a literal `(` inside a markdown link target and comment bodies containing a literal `(` before the keyword. No corpus, showcase, or coverage row moved; 666 tests pass. ### Tests Four rows added. Three name the guilty pattern per arm (`test_crafted_redos_payload_stays_bounded_in_the_lexicon`), one covers the composed gate (`test_gate_is_bounded_on_the_payload_the_first_sweep_missed`). Pre-fix they failed at 297 s, 8.1 s, 55 s and 334.7 s. `N` is per row deliberately. The URL arm is quadratic with a small constant and ran 0.9 s **unfixed** at N=100 000 — under the 2.0 s bound, so that row would have passed whether or not the pattern was fixed. It is measured at N=300 000 instead, where crafted (8.10 s) and legitimate (0.926 s) separate 8.8×. ### Residual The sweep flags on timing and ignores measurements below a 1.5 ms noise floor at N=8 000. An arm hiding just under it could still cost **~23 s** at the cap, so what this supports is *"no arm worse than ~23 s"*, not *"no quadratic arm remains"*. The blind spot is not hypothetical: a generic-payload pass found only one of the two patterns. The second appeared only once payloads were synthesised per run from each pattern's own skeleton. Recorded in `docs/LIMITATIONS.md`. ## [0.3.2] — 2026-07-31 > **Denial-of-service fix. Upgrade from 0.3.1.** The output gate could be made to > spend hours on a single call by crafted input it accepts by design. No > disposition changes for ordinary documents — the one measured exception is > listed under *Known behaviour changes* below. The v0.3.1 tag is not moved. ### Fixed — 19 quadratic regex runs on the output path `scan_output` claimed LLM10 self-safety on the grounds that its patterns contain no nested quantifiers. That is true and it is not the property that matters. A run in front of a **required** literal, reachable from a short anchor, is enough: crafted input repeats the anchor and never supplies the literal, so every start position rescans the tail. Quadratic, not exponential — and quadratic is sufficient here. Measured, not argued (Python 3.14, this machine): | Input | Time through `scan_output` | |---|---| | ``. ### Known behaviour changes Two, both measured against the v0.3.1 tag rather than reasoned about: - **A JWT used as a DB password, over 256 chars, is no longer CRITICAL.** The remaining detections (`entropy:base64-blob` HIGH, `egress:jwt-token` MEDIUM) top out below CRITICAL, so the any-tier block is lost: under `PRESET_TRUSTED_SOURCE` such a document moves from `fail_secure` to `quarantine_review`. Under `PRESET_USER_UPLOAD` it still `fail_secure`s, and a *generic* long password still trips `entropy:base64-blob` at CRITICAL with no change at all. The credential is never silently missed; on one preset it is held for review instead of halted. - **`hybrid-xss:script-tag` now fires on prose that merely mentions `