Version synced across the four locked points (pyproject, __version__, README badge, README install pin) + CHANGELOG. Corrections this release carries, both measured rather than reasoned: - docs/LIMITATIONS.md said the script-tag change removed "the last" quadratic-backtracking site on the output path. It did not. Corrected, with the 334.7s gate measurement that falsifies it. - README's coverage line claimed 126/126 classes; the matrix reports 128/128. Stale since before v0.3.2. Test badge was 642, actual 666. New residual recorded (LIMITATIONS, now 30 items, README synced): the sweep flags on timing above a 1.5ms noise floor at N=8000, so an arm hiding under it could still cost ~23s at the cap. What this supports is "no arm worse than ~23s", not "no quadratic arm remains" -- and the blind spot is demonstrated, since a generic-payload pass found only one of the two patterns. 666 tests green, coverage matrix 128/128 with 6/6 gaps holding, exit 0.
380 lines
21 KiB
Markdown
380 lines
21 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to this project will be documented in this file.
|
||
|
||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||
|
||
## [Unreleased]
|
||
|
||
## [0.3.3] — 2026-07-31
|
||
|
||
> **Denial-of-service fix, and a correction to 0.3.2. Upgrade from 0.3.2.** The
|
||
> sweep 0.3.2 shipped was incomplete, and it said otherwise. Two lexicon patterns
|
||
> were still quadratic — reachable through `scan_output`, not only on the input
|
||
> path. No disposition changes: recall was measured case by case and nothing was
|
||
> lost. The v0.3.2 tag is not moved.
|
||
|
||
### Fixed — two quadratic patterns in the lexicon table
|
||
|
||
`8deca93` scoped the remaining ReDoS duty to the lexicon path, and this is that
|
||
work: all 83 patterns measured, arm by arm. Two are quadratic, same shape as
|
||
everything 0.3.2 fixed — a run in front of a **required** literal, where the run
|
||
may cross the pattern's own opening anchor.
|
||
|
||
| Pattern (arm) | Crafted unit | Measured | At the 1 000 000-char cap |
|
||
|---|---|---|---|
|
||
| `markdown:link-anchor-injection` (anchor text) | `[` | 1.91 s @ 8 000 | **~8.3 hours** |
|
||
| `markdown:link-anchor-injection` (URL run) | `[system](` | 0.006 s @ 8 000 | ~89 seconds |
|
||
| `markdown:link-ref-comment` (`.*` run) | `[//]: # (` | 0.22 s @ 8 000 | **~1.0 hour** |
|
||
|
||
Exponent measured over five points (1 000 → 16 000): **1.98** — quadratic, not
|
||
exponential. Legitimate content of the same size is unaffected: 0.316 s at
|
||
N=100 000 (prose 0.316 / html 0.315 / markdown 0.297 / connection-string 0.296).
|
||
|
||
**These were not input-path-only, and that is the correction.** `scan_lexicon`
|
||
runs on the output path too, so 0.3.2's *"the last quadratic-backtracking site on
|
||
the output path"* was false when written. Measured through the public gate before
|
||
this fix: `scan_output("[" * 100_000)` took **334.7 s**. The claim was too broad
|
||
because the sweep behind it drove the `[` payload only through
|
||
`scan_active_content` — no row ever drove it through the lexicon. The statement is
|
||
corrected in `docs/LIMITATIONS.md`.
|
||
|
||
The fix is anchor exclusion, per the rule `active_content` already documents —
|
||
bounding attacker-controlled content would be a one-line detection bypass. The
|
||
excluded character is `(`, not `[`:
|
||
|
||
```
|
||
markdown:link-anchor-injection
|
||
\[[^\]\[]*(?:system|…)[^\]\[]*\]\([^)(]+\)
|
||
markdown:link-ref-comment
|
||
\[//\]:\s*#\s*\([^(\n]*(?:ignore|…)
|
||
```
|
||
|
||
`[` was the obvious choice and it was measurably worse. Excluding `[` from the
|
||
URL run drops `[override your rules](https://[::1]/x)` — still covered, three
|
||
other patterns fire on it — but excluding `[` from the link-ref comment run drops
|
||
`[//]: # (see [x] then ignore this)`, which **no other pattern catches**. The
|
||
anchors contain `(` as well, so excluding `(` telescopes just as effectively at
|
||
zero measured recall cost. Both forms verified linear (×1.99–2.02 on doubling).
|
||
|
||
### Known behaviour changes
|
||
|
||
- **None measured.** Every case that matched before still matches, except URLs
|
||
containing a literal `(` inside a markdown link target and comment bodies
|
||
containing a literal `(` before the keyword. No corpus, showcase, or coverage
|
||
row moved; 666 tests pass.
|
||
|
||
### Tests
|
||
|
||
Four rows added. Three name the guilty pattern per arm
|
||
(`test_crafted_redos_payload_stays_bounded_in_the_lexicon`), one covers the
|
||
composed gate (`test_gate_is_bounded_on_the_payload_the_first_sweep_missed`).
|
||
Pre-fix they failed at 297 s, 8.1 s, 55 s and 334.7 s.
|
||
|
||
`N` is per row deliberately. The URL arm is quadratic with a small constant and
|
||
ran 0.9 s **unfixed** at N=100 000 — under the 2.0 s bound, so that row would have
|
||
passed whether or not the pattern was fixed. It is measured at N=300 000 instead,
|
||
where crafted (8.10 s) and legitimate (0.926 s) separate 8.8×.
|
||
|
||
### Residual
|
||
|
||
The sweep flags on timing and ignores measurements below a 1.5 ms noise floor at
|
||
N=8 000. An arm hiding just under it could still cost **~23 s** at the cap, so what
|
||
this supports is *"no arm worse than ~23 s"*, not *"no quadratic arm remains"*.
|
||
The blind spot is not hypothetical: a generic-payload pass found only one of the
|
||
two patterns. The second appeared only once payloads were synthesised per run
|
||
from each pattern's own skeleton. Recorded in `docs/LIMITATIONS.md`.
|
||
|
||
## [0.3.2] — 2026-07-31
|
||
|
||
> **Denial-of-service fix. Upgrade from 0.3.1.** The output gate could be made to
|
||
> spend hours on a single call by crafted input it accepts by design. No
|
||
> disposition changes for ordinary documents — the one measured exception is
|
||
> listed under *Known behaviour changes* below. The v0.3.1 tag is not moved.
|
||
|
||
### Fixed — 19 quadratic regex runs on the output path
|
||
|
||
`scan_output` claimed LLM10 self-safety on the grounds that its patterns contain no
|
||
nested quantifiers. That is true and it is not the property that matters. A run in
|
||
front of a **required** literal, reachable from a short anchor, is enough: crafted
|
||
input repeats the anchor and never supplies the literal, so every start position
|
||
rescans the tail. Quadratic, not exponential — and quadratic is sufficient here.
|
||
|
||
Measured, not argued (Python 3.14, this machine):
|
||
|
||
| Input | Time through `scan_output` |
|
||
|---|---|
|
||
| `<a:` × 100 000 (300 KB) | **458.7 s** |
|
||
| size-matched ordinary prose | 0.31 s |
|
||
| same payload extrapolated to the 1 000 000-char input the gate itself accepts | **~5.7 hours for one call** |
|
||
|
||
`max_scan_chars` does not mitigate this. It bounds the *input*; quadratic work on a
|
||
bounded input is still hours. That claim was stated in both `output.py` and
|
||
`calibration.py` and is corrected in both.
|
||
|
||
The fix is per pattern, not uniform:
|
||
|
||
- **`active_content` + the lexicon table (15 runs)** — exclude the character that
|
||
opens the pattern's own anchor (`[` for markdown, `<` for tags), so a run cannot
|
||
reach past the next start position and the per-start costs telescope. Verified to
|
||
cost no recall: long URLs, long alt text, and `<` inside a quoted attribute all
|
||
still match. Bounding instead would have been linear too, but wrong here — the
|
||
content is attacker-controlled, so padding past a bound would be a one-line bypass
|
||
of the EchoLeak class this table exists to catch.
|
||
- **`*-connstr` secret egress (4 runs)** — bound the password at the new
|
||
`MAX_CONNSTR_VALUE` (256) in `calibration`. The exclusion fix is unavailable: the
|
||
anchor character is `/`, and passwords containing `/` are the common case
|
||
(measured — they match today).
|
||
- **`hybrid-xss:script-tag`** — had neither option, since its run is the script
|
||
*body*, which may legitimately contain `<`. It now matches the opening tag and no
|
||
longer requires `</script>`.
|
||
|
||
### Known behaviour changes
|
||
|
||
Two, both measured against the v0.3.1 tag rather than reasoned about:
|
||
|
||
- **A JWT used as a DB password, over 256 chars, is no longer CRITICAL.** The
|
||
remaining detections (`entropy:base64-blob` HIGH, `egress:jwt-token` MEDIUM) top
|
||
out below CRITICAL, so the any-tier block is lost: under `PRESET_TRUSTED_SOURCE`
|
||
such a document moves from `fail_secure` to `quarantine_review`. Under
|
||
`PRESET_USER_UPLOAD` it still `fail_secure`s, and a *generic* long password still
|
||
trips `entropy:base64-blob` at CRITICAL with no change at all. The credential is
|
||
never silently missed; on one preset it is held for review instead of halted.
|
||
- **`hybrid-xss:script-tag` now fires on prose that merely mentions `<script>`** —
|
||
and this costs no consumer a disposition. Any text containing a literal
|
||
`<script>` already produced `active:raw-html` at HIGH on 0.3.1, so the same
|
||
document disposed identically before and after. The label is new; the outcome is
|
||
not. By the same measurement, the fail-open this closed (an unclosed
|
||
`<script>alert(1)`) was confined to `scan_lexicon` called on its own — through
|
||
either composed gate, `active:raw-html` already caught it.
|
||
|
||
Both are recorded in `docs/LIMITATIONS.md` (still 29 items — these replace nothing).
|
||
|
||
### Method note
|
||
|
||
The defect was found by a composed-gate DoS test that stayed red after every
|
||
individual scanner had been made linear; the remaining 813× was the lexicon's six
|
||
`html-obfuscation` patterns. A per-scanner test alone would have shipped it. The
|
||
static shape analysis used to find candidates also missed `[\s\S]*?` in
|
||
`script-tag` — the sweep that matters is measurement, not a regex over regexes.
|
||
|
||
662 tests pass (was 642), and the suite is faster than before the fix.
|
||
|
||
## [0.3.1] — 2026-07-25
|
||
|
||
> **Regression fix. Upgrade from 0.3.0.** v0.3.0 made the high-untrust upload path
|
||
> unusable for ordinary documents — measured, not projected. `llm-ingestion-okf`
|
||
> projected the consequence from the 0.3.0 changelog text *before* the tag was cut;
|
||
> the release went out without the inbox being read. The v0.3.0 tag is not moved.
|
||
|
||
### Fixed — the upload path is usable again without losing EchoLeak detection
|
||
|
||
Measured on v0.3.0, both doors (`screen_output` under `PRESET_USER_UPLOAD` and
|
||
`okf.import_bundle` with `origin=EXTERNAL`): a document with **one ordinary remote
|
||
image** disposed `fail_secure`; one ordinary link, autolink or reference definition
|
||
disposed `quarantine_review`. Only documents with no external references persisted.
|
||
Two independent defects compounded, and both had to be fixed — either alone leaves
|
||
the path blocked:
|
||
|
||
- **Severity graded on construct type instead of URL shape.** `markdown-image` was
|
||
HIGH for *any* external image, but the exfiltration primitive is not "an image" —
|
||
it is a URL that moves bytes to a host the attacker controls.
|
||
`` carries nothing. Severity now grades on
|
||
shape: a URL that only *names* a remote document (http(s) or protocol-relative,
|
||
no query, no userinfo, no percent-escapes, no opaque host label or path segment)
|
||
is **LOW**; anything that can carry a value keeps the carrier's full severity.
|
||
`raw-html` and `data:` URIs have no ordinary form and stay HIGH unconditionally.
|
||
Opacity reuses `entropy`'s primitives — decodable base64 (≥20 chars), hex id
|
||
(≥32), or Shannon entropy ≥4.4 at ≥24 chars — calibrated 2026-07-25 against real
|
||
documentation URLs (worst legitimate token H=4.08; exfil payload segments
|
||
4.36-4.54). New constants live in `calibration` with the rest.
|
||
- **The `quarantine_default` floor fired on *any* finding.** It rested on the premise
|
||
that a finding is the exception; adding the active-content detector in 0.3.0 made
|
||
every ordinary markdown link a finding, and the floor then held ordinary documents
|
||
for review. The floor now fires at **MEDIUM+**. This is a no-op for every detector
|
||
that shipped before 0.3.0 — the lexicon holds no LOW/INFO pattern and no other
|
||
detector emits LOW (asserted in `tests/test_calibration.py`) — which is why this is
|
||
a patch and not a minor.
|
||
|
||
**Unchanged, deliberately:** no new public API and no new preset (a middle tier is
|
||
0.4.0 work); the `allow_reserved=True` mode-b default stands — two independent
|
||
consumers document it as load-bearing; the gate still never rewrites content.
|
||
|
||
### Added
|
||
|
||
- **False-positive corpus covers ordinary markdown.** The 0.3.0 corpus had zero
|
||
markdown links or images, asserted only under `PRESET_TRUSTED_SOURCE` (where every
|
||
non-CRITICAL finding WARNs anyway), and drove the *input* path — so `scan_output`
|
||
step 6, where active content actually lives, was never reached. That is how a
|
||
regression this size passed 522 green tests. The corpus now carries realistic
|
||
documents and asserts them on the **output gate under the upload preset**, plus a
|
||
counter-corpus of exfil-shaped URLs (query, base64/hex path segment,
|
||
percent-encoded payload, opaque subdomain, userinfo) that must still block.
|
||
- **Two new documented gaps** in `docs/LIMITATIONS.md`, both asserted by the coverage
|
||
matrix: **pure beaconing** (a bare-path image on a hostile host still fetches, and
|
||
the fetch is not graded) and **short opaque URL segments** (<24 chars, below what
|
||
entropy can resolve). Percent-escapes counting as data-carrying is recorded there
|
||
as a known false positive.
|
||
|
||
## [0.3.0] — 2026-07-25
|
||
|
||
> **A minor bump, not a patch — deliberately.** The changes under *Changed* alter what
|
||
> an existing caller observes with no code change on their side, so a `>=0.2,<0.3` pin
|
||
> stops here rather than absorbing them silently. Re-test that branch before widening
|
||
> the pin. Still alpha: the public API may change again before 1.0.
|
||
|
||
### Changed — observable gate behaviour (re-test before upgrading)
|
||
|
||
Three commits since v0.2.0 change dispositions for an unchanged caller. Two tighten
|
||
the gate; one loosens it.
|
||
|
||
- **`okf.import_bundle` no longer path-rejects reserved basenames.** At v0.2.0,
|
||
`index.md` / `log.md` anywhere in a received bundle was an unconditional
|
||
per-concept hard reject (FAIL_SECURE), and `import_bundle` took no keyword for it.
|
||
The new `allow_reserved` keyword **defaults to `True`** on this mode-b
|
||
*received-bundle* path, so those files are scanned — their body is the
|
||
highest-priority injection surface — rather than refused, and may clear the floor
|
||
and become mergeable. **This is the one loosening change:** content a v0.2.0
|
||
consumer never saw can now reach it, so a consumer whose tests pin the v0.2.0
|
||
reject must re-check, not just bump. A front-end materialising individual
|
||
*uploads* must pass `allow_reserved=False` to keep the shadow-reject there;
|
||
`validate_concept_path` still defaults to `False`.
|
||
- **Active content now reaches the disposition engine.** `scan_output` step 6 runs
|
||
`scan_active_content`, so markdown images/links, reference definitions, autolinks,
|
||
raw active HTML and `data:` URIs surface as `active:*` findings (OWASP LLM05 — the
|
||
EchoLeak / CVE-2025-32711 class) instead of being admitted with `findings=[]`.
|
||
These carry real severities (zero-click auto-fetch/execute HIGH, click-required
|
||
MEDIUM), and two MEDIUM+ findings compound-escalate one tier, so a document that
|
||
passed clean at v0.2.0 can now WARN, quarantine, or fail secure in both
|
||
`screen_output` and `okf.import_bundle`.
|
||
- **Base64-wrapped secrets are now caught as egress.** The output gate's
|
||
decode-and-rescan feeds decoded base64 plaintext through both the lexicon and the
|
||
LLM02 secret-egress detector, so a base64-wrapped credential surfaces as
|
||
`decoded:egress:*` instead of disappearing. Hex-wrapped remains a documented gap
|
||
(`docs/LIMITATIONS.md`).
|
||
|
||
### Added — runnable threat-coverage matrix
|
||
|
||
A single declarative manifest (`llm_ingestion_guard.coverage`) that proves, in one
|
||
place, every vulnerability class the guard stops — and the documented gaps it does
|
||
not. Two consumers of the same source of truth:
|
||
|
||
- `python -m llm_ingestion_guard.coverage` — a narrated matrix
|
||
(`class -> OWASP -> expected -> observed -> verdict`); exit 0 iff every caught
|
||
class is caught and every documented gap holds. Stdlib-only, CI-usable.
|
||
- `tests/test_coverage_matrix.py` — asserts total recall over the core matrix
|
||
(carriers, all 83 lexicon patterns, entropy/decode-rescan, active content, the
|
||
contract asserters, the disposition engine, OKF T1–T7), asserts every documented
|
||
gap still holds, and guards completeness (every lexicon pattern id, and every
|
||
OWASP anchor claimed, has a case). Adds the full 25-pattern LLM02 secret-egress
|
||
set and the container-layer front-end classes (CSV formula-injection, zip-slip,
|
||
zip-bomb, symlink). +165 tests (357 → 522), no core dependency added.
|
||
|
||
This is the real-case validation gate ahead of a v1.0 freeze.
|
||
|
||
### Documentation — consumer adoption + README value proposition
|
||
|
||
- `docs/ADOPTION-BRIEF.md` — a self-contained brief a consumer repo (OKF
|
||
second-brain / LLM wiki) can plan an inclusion from: the write-time trust-boundary
|
||
argument, the two bookends + 8-step contract, the shipped OKF adapter
|
||
(`import_bundle` mode-b), how to verify (coverage matrix), how to depend
|
||
(stdlib-only core), and a checklist for *when/where* to wire it.
|
||
- README rewritten to lead with the write-time trust-boundary framing, add a
|
||
first-class **OKF / LLM-wiki support (shipped)** section for `import_bundle`, a
|
||
concrete **What it protects against** catalogue (attack classes grouped by OWASP
|
||
anchor, driven by the coverage matrix), and correct the test badge (357 → 522).
|
||
Every claim verified against the code.
|
||
- `docs/LIMITATIONS.md` — the full honest-limitations list (15 items + the four
|
||
documented gaps + out-of-scope) moved out of the README, which now carries a
|
||
high-impact summary + link, so protection and limits read in balance.
|
||
|
||
## [0.2.0] — 2026-07-06
|
||
|
||
### Added — OKF adapter (stream 1)
|
||
|
||
An OKF (Google Open Knowledge Format v0.1) adapter *on top of* the
|
||
format-agnostic core (`llm_ingestion_guard.okf`). The core stays `text ->
|
||
findings`; the adapter knows OKF structure and routes scannable regions into the
|
||
existing machinery. All TDD (failing test first), +61 tests. Verified against the
|
||
OKF `SPEC.md` (2026-07-06). See `docs/OKF-INGESTION-BRIEF.md` §8.
|
||
|
||
- `parse_frontmatter` — strict, reject-by-default frontmatter loader; refuses
|
||
anchors, aliases, explicit tags, merge keys, block scalars and flow collections
|
||
by construction, so YAML anchor/alias DoS and `!!python/object` coercion cannot
|
||
occur (not a general YAML parser, by design). (T2)
|
||
- `scan_concept` — whole-concept scan surface: frontmatter values (incl.
|
||
`description`, read first under progressive disclosure), `resource` and body all
|
||
go through `scan_output`. (T1)
|
||
- `validate_concept_path` — path / reserved-name gate: rejects `..` traversal,
|
||
absolute paths and `index.md` / `log.md` shadowing; returns the concept-ID. (T4)
|
||
- `validate_resource_url` — `resource` https allowlist: rejects non-https before
|
||
commit (reject, not defang — the format imposes no scheme constraint itself). (T3)
|
||
- `stamp_concept` / `format_log_entry` — provenance stamping: origin × channel →
|
||
trust × disposition per concept, emitted as `log.md` lines. Trust follows the
|
||
origin, never the insertion channel. (T6)
|
||
- `import_bundle` — received-bundle iterator (mode b): validates each concept
|
||
(path, frontmatter, resource, scan, stamp) independently; one bad concept is
|
||
rejected fail-secure while the rest are still checked; the aggregate disposition
|
||
is the most severe. (T7)
|
||
- `link_graph` / `resolve_link` / `extract_link_targets` — in-import cross-link
|
||
graph: resolves `.md` links (bundle-absolute or relative) to concept-IDs, flags
|
||
dangling links (the §7.2 dormant-injection signal) and rejects dangerous-scheme
|
||
or bundle-escaping targets. (T5a)
|
||
|
||
### Deferred
|
||
|
||
- Cross-run persisted link graph (T5b) — catching a link planted in one run whose
|
||
poisoned target is written in a *later* run (§7.2) needs durable graph state
|
||
whose storage/ownership depends on the consuming pipeline. Deferred to the
|
||
consumer-wiring stream; cross-run dormant links remain a documented residual
|
||
risk (README honest-limitations).
|
||
|
||
## [0.1.0] — 2026-07-06 (alpha)
|
||
|
||
The stdlib-only core, built test-first (TDD) per `docs/PLAN.md`. Tagged `v0.1.0`.
|
||
|
||
### Added
|
||
|
||
- `report` — shared `Finding` / `Report` / `Severity` / `Source` types.
|
||
- `sanitize` — carrier stripping (zero-width, BIDI, Unicode-tag, HTML comment,
|
||
`data:`); byte-identical on clean input.
|
||
- `entropy` — Shannon / base64-like / hex-blob detection; base64 decode-and-rescan.
|
||
- `lexicon` — JSON pattern data + loader; raw/normalized/homoglyph/rot13 variants;
|
||
ReDoS-bounded, size-capped.
|
||
- `fence` — randomized per-call spotlight delimiter; attacker marker-strip.
|
||
- `neutralize` — opt-in defang of active-content output (byte-identical when clean).
|
||
- `output` — compose lexicon + entropy + decode-rescan over emitted text; secret
|
||
egress patterns (OWASP LLM02); report-only, never mutates.
|
||
- `disposition` — WARN | QUARANTINE_REVIEW | FAIL_SECURE under a source-trust
|
||
policy; compound-signal escalation; fail-**closed** when the scanner errors.
|
||
- `contract` — write-time asserters that raise: `assert_tool_less`,
|
||
`assert_credential_allowlist`, `scoped_env`.
|
||
- `grounding` — the `SourceGroundingCheck` seam for semantic poisoning (interface
|
||
only; `[judge]` implementation plugs in behind an extra).
|
||
- Top-level wiring — the `prepare_input` / `screen_output` §6 bookends plus the
|
||
full public surface; end-to-end showcase and adversarial + false-positive corpora.
|
||
|
||
### Security
|
||
|
||
Pre-release hardening from an independent adversarial review (all TDD, failing
|
||
test first):
|
||
|
||
- `entropy` — decode-and-rescan now runs **before** false-positive suppression,
|
||
so an injection blob prefixed with an SRI/media marker (to dodge the entropy
|
||
finding) is still decoded and rescanned by the lexicon. Suppression gates only
|
||
the entropy finding, never the decode.
|
||
- `output` — the invisible-carrier invariant now holds on the persist gate:
|
||
`scan_output` flags zero-width / BIDI presence (`output:zero-width-present`,
|
||
`output:bidi-present`) and `disposition` treats those plus
|
||
`lexicon:unicode-tags-present` as any-tier carriers, so a carrier in model
|
||
output fails secure even under a trusted policy.
|
||
- `contract` — `assert_credential_allowlist` catches a bare `<PROVIDER>_KEY`
|
||
(e.g. `STRIPE_KEY`); the previous regex silently missed it (fail-open). The
|
||
rule is deliberately broad (also flags `PARTITION_KEY`/`SORT_KEY` as loud,
|
||
allowlistable false positives) — fail-loud beats fail-silent for isolation.
|
||
- `disposition` — `guard` runs `decide` inside its guarded block, so a malformed
|
||
report can no longer escape the fail-closed guarantee.
|
||
- `output` — secret-egress placeholder suppression anchors word markers
|
||
(`example`, `todo`, …) to a word boundary, so a real secret that merely
|
||
*contains* such a word is no longer suppressed (fail-open egress miss closed).
|