The 19 fixed credential shapes in pre-edit-secrets.mjs were regex literals; they now come from signatures/secret-egress.json in the vendored commons via a new scanners/lib/secret-egress.mjs. Policy-injected custom patterns (entries 20+) are unchanged and still appended by the hook. Measured before the swap, not assumed: all 19 positions compared for order, name, regex source and flags, plus recompilation identity, against the literal table sliced out of the module text. Zero divergences. Commons had reported the same result; that was their measurement, so this one was run anyway. STATE's expectation that the golden gate would go red on both table records and file sha256 was wrong: pre-edit-secrets.mjs is in neither PINNED_FILES nor WALKED_MODULES, so the table had no golden coverage at all and the swap moved nothing. Rather than leave the vendored data with only behavioural coverage, secret-egress.mjs joins WALKED_MODULES — walked, not pinned, since it inlines no regex of its own. Golden diff was 19 ADDED, 0 CHANGED, 0 REMOVED, each source byte-identical to the pre-swap literal; re-blessed. suite-counts.json untouched. Tests: coverage is derived from the loaded table, so an entry commons adds cannot arrive without an end-to-end probe. All 19 now block through the real hook and are asserted by label, which also pins the ordering contract (a Bearer-wrapped JWT must report as the header). Mutating the vendored JSON fires in both directions plus reorder: under-match (AKIA quantifier) reddens 3 hook tests + golden; over-match (Anthropic key truncated to its prefix) reddens the false-positive probe + golden; moving the JWT entry ahead of the Bearer entry reddens the ordering test. Suite 2231 tests / 2223 pass / 6 skipped. The two parallel-run failures (pre-compact size-cap, benchmark) pass alone — the known timing flakes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MGMv5ZTUhVzZtCCwRrNZG5 |
||
|---|---|---|
| .. | ||
| patterns.json | ||
| README.md | ||
| reference-run.json | ||
| suite-counts.json | ||
Golden baseline — v8 Phase 5
Reference artifacts recorded before any commons extraction, so that each table-by-table swap in Phase 5 can be proven behaviour-preserving (or rolled back). Generated and checked by one code path:
node scripts/golden-baseline.mjs # check only, exits 1 on drift
node scripts/golden-baseline.mjs --write # (re-)bless
node scripts/golden-baseline.mjs --write --suite # also refresh suite counts (slow)
The gate is tests/lib/golden-baseline.test.mjs.
The artifacts
| File | What it pins | Why that layer exists |
|---|---|---|
patterns.json |
.source + .flags of every RegExp reachable from the walked modules' exports |
After a swap a pattern is new RegExp(jsonString, flags). The plan's named hazard — JSON backslash-doubling — is visible only on the compiled object. |
| ↳ table records | key/value digest of HOMOGLYPH_MAP, TYPOSQUAT_SUSPICIOUS_TOKENS, SEVERITY, the four OWASP maps |
Most of what Phase 4 moves is not a regex. A regex-only dump is blind to a broken homoglyph or OWASP-map swap, i.e. blind to the bulk of the payload. |
| ↳ file records | sha256 of the five source files in the moving set | The completeness layer, complete by construction: it covers regexes inlined in function bodies (NAMED at string-utils.mjs:291, the BIDI/tag/PUA ranges at 357–404) that no export walk can reach. |
reference-run.json |
the 61 showcase payloads through the real hook entry points, plus a static reachability probe | Byte identity over a corpus that trips a handful of patterns would prove almost nothing, so reachability is recorded and unreachable patterns are listed by name. |
suite-counts.json |
per-file pass/fail, each file run alone | A total is unattributable, and npm test runs files concurrently where three timing-sensitive files flake. Running each alone is the only reproducible recording. |
Why there is no regex enumerator
The obvious design — parse the sources and enumerate every regex literal — was
rejected. It needs a JS parser this zero-dependency repo does not have, and a
lexical approximation is contaminated by comments and division (severity.mjs
scores 4 "regexes" that way and exports none). The two-layer split — export
walk for what the scanners actually use, file digest for everything else —
answers the same question without a parser.
During a swap
A diff here means the swap changed observable behaviour. Roll the swap back.
--write is for deliberately re-blessing a change you have already decided is
correct, not for making the gate quiet.
Known coverage gaps — read before quoting a number
coverage is static reachability, NOT observed coverage. It probes the 61
payload strings against the 83 injection patterns in-process: "if you threw
every payload string at every regex, how many would match?" It does not
measure what the 61 hook invocations evaluated — the pre-bash-destructive
payloads never reach injection-patterns at all, yet their strings are in the
probe set and can mark a pattern reachable. 47/83 is therefore an upper
bound on what the corpus could protect, not a measurement of what it did.
The block is named kind: "static-reachability" and carries that caveat in a
note field; the gate asserts both, so the honest label cannot be dropped
quietly.
Two gaps follow:
- 36 unreachable patterns, listed by key under
coverage.unreachablePatterns. A swap that breaks one of those is caught by the pattern dump only. Closing it means growing the conformance corpus (Phase 5 step 5). - The four OWASP maps have no behavioural coverage at all. They are scanner-side and no hook in this corpus reaches them. They were the reason the dump was widened to table records, and the digest is their only protection — do not read the reference run as backing them.
corpusContains says a payload contains a homoglyph or is altered by
normalizeForScan. It does not say the run folded one.