fix(security): harden 5 adversarial-review findings (M1/M2/M3 + m4/m6) via TDD
Pre-release hardening from an independent adversarial review; each fixed test-first (failing test -> fix -> green). 214 tests pass. - entropy (M1): decode-and-rescan now runs BEFORE false-positive suppression, so an SRI/media-prefixed injection blob is still decoded and lexicon-rescanned. Suppression gates only the entropy finding, never the decode. - output/disposition (M3): the invisible-carrier invariant now holds on the persist gate. scan_output flags zero-width/BIDI presence and disposition treats those + lexicon:unicode-tags-present as any-tier carriers, so a carrier in model output fails secure even under a trusted policy. - contract (M2): assert_credential_allowlist catches a bare <PROVIDER>_KEY (e.g. STRIPE_KEY) that the old regex silently missed (fail-open). Deliberately broad: also flags PARTITION_KEY/SORT_KEY as loud, allowlistable FPs -- fail-loud beats fail-silent for an isolation control. - disposition (m6): guard runs decide inside its guarded block -> total fail-closed even on a malformed report. - output (m4): egress placeholder suppression anchors word markers (example, todo, ...) to a word boundary, closing a fail-open where a real secret merely containing such a word was suppressed. Docs: CHANGELOG Security subsection; README honest-limit for lexicon dedup (m5, documented tradeoff, not fixed). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HyRCQMocjZ6SmSQ6JidJ2k
This commit is contained in:
parent
86726ed109
commit
5397ba15a1
10 changed files with 233 additions and 18 deletions
24
CHANGELOG.md
24
CHANGELOG.md
|
|
@ -29,3 +29,27 @@ The stdlib-only core, built test-first (TDD) per `docs/PLAN.md`.
|
|||
only; `[judge]` implementation plugs in behind an extra).
|
||||
- Top-level wiring — the `prepare_input` / `screen_output` §6 bookends plus the
|
||||
full public surface; end-to-end showcase and adversarial + false-positive corpora.
|
||||
|
||||
### Security
|
||||
|
||||
Pre-release hardening from an independent adversarial review (all TDD, failing
|
||||
test first):
|
||||
|
||||
- `entropy` — decode-and-rescan now runs **before** false-positive suppression,
|
||||
so an injection blob prefixed with an SRI/media marker (to dodge the entropy
|
||||
finding) is still decoded and rescanned by the lexicon. Suppression gates only
|
||||
the entropy finding, never the decode.
|
||||
- `output` — the invisible-carrier invariant now holds on the persist gate:
|
||||
`scan_output` flags zero-width / BIDI presence (`output:zero-width-present`,
|
||||
`output:bidi-present`) and `disposition` treats those plus
|
||||
`lexicon:unicode-tags-present` as any-tier carriers, so a carrier in model
|
||||
output fails secure even under a trusted policy.
|
||||
- `contract` — `assert_credential_allowlist` catches a bare `<PROVIDER>_KEY`
|
||||
(e.g. `STRIPE_KEY`); the previous regex silently missed it (fail-open). The
|
||||
rule is deliberately broad (also flags `PARTITION_KEY`/`SORT_KEY` as loud,
|
||||
allowlistable false positives) — fail-loud beats fail-silent for isolation.
|
||||
- `disposition` — `guard` runs `decide` inside its guarded block, so a malformed
|
||||
report can no longer escape the fail-closed guarantee.
|
||||
- `output` — secret-egress placeholder suppression anchors word markers
|
||||
(`example`, `todo`, …) to a word boundary, so a real secret that merely
|
||||
*contains* such a word is no longer suppressed (fail-open egress miss closed).
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue