llm-security/tests/golden/README.md
Kjell Tore Guttormsen 2fe29152b3 fix(llm-security): golden gate - coverage block measured something else
The `coverage` block never looked at the 61 hook invocations. It probed every
payload STRING against every injection regex in-process, so `47/83` meant
"if you threw all 61 strings at all 83 patterns, 47 would match" - not "the
reference run exercised 47". The pre-bash-destructive payloads never reach
injection-patterns at all, yet their strings were in the probe set and could
mark a pattern exercised.

That number was asserted as reference-run coverage in three places that
instruct a future session: tests/golden/README.md, the STATE golden-gate
section, and the generator's summary line. In a repo whose v7.8.2 lesson was
"the check reported success without running", a figure that measures one
thing while labelled another is the same defect wearing a different hat.

Relabelled rather than re-measured - the probe still honestly bounds the gate
(an unreachable pattern is one the corpus cannot protect under ANY
attribution), it just has to say what it is:

  coverage.kind = 'static-reachability', with the caveat in a `note` field.
  patternsExercised  -> patternsReachable
  uncoveredPatterns  -> unreachablePatterns
  tablesExercised    -> corpusContains  (a payload CONTAINS a homoglyph; it
                        does not say the run folded one)

The gate now pins both the `kind` and the note, so the honest label cannot be
dropped quietly by a later edit.

Also surfaced the gap the old wording hid: the four OWASP maps have NO
behavioural coverage - they are scanner-side and no hook in this corpus
reaches them. They are precisely the tables the dump was widened to cover, so
the table digest is their only protection. Stated in the README next to the
47/83 line, where a reader was previously left to infer the reference run
backed them.

Flakiness check for the new gate (61 sequential spawns, ~12s, the most
process-heavy file in the suite, added to a suite npm test runs concurrently):
three consecutive full runs, 2053/2053, 0 fail. The three known
timing-sensitive files did not destabilise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BB4vXvwvtW4dxbPRd6vsez
2026-08-09 13:06:45 +02:00

66 lines
4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Golden baseline — v8 Phase 5
Reference artifacts recorded **before** any commons extraction, so that each
table-by-table swap in Phase 5 can be proven behaviour-preserving (or rolled
back). Generated and checked by one code path:
```bash
node scripts/golden-baseline.mjs # check only, exits 1 on drift
node scripts/golden-baseline.mjs --write # (re-)bless
node scripts/golden-baseline.mjs --write --suite # also refresh suite counts (slow)
```
The gate is `tests/lib/golden-baseline.test.mjs`.
## The artifacts
| File | What it pins | Why that layer exists |
|------|--------------|-----------------------|
| `patterns.json` | `.source` + `.flags` of every RegExp reachable from the walked modules' exports | After a swap a pattern is `new RegExp(jsonString, flags)`. The plan's named hazard — JSON backslash-doubling — is visible **only** on the compiled object. |
| ↳ table records | key/value digest of `HOMOGLYPH_MAP`, `TYPOSQUAT_SUSPICIOUS_TOKENS`, `SEVERITY`, the four OWASP maps | Most of what Phase 4 moves is not a regex. A regex-only dump is blind to a broken homoglyph or OWASP-map swap, i.e. blind to the bulk of the payload. |
| ↳ file records | sha256 of the five source files in the moving set | The completeness layer, complete **by construction**: it covers regexes inlined in function bodies (`NAMED` at `string-utils.mjs:291`, the BIDI/tag/PUA ranges at 357404) that no export walk can reach. |
| `reference-run.json` | the 61 showcase payloads through the real hook entry points, plus a static reachability probe | Byte identity over a corpus that trips a handful of patterns would prove almost nothing, so reachability is recorded and unreachable patterns are listed by name. |
| `suite-counts.json` | per-file pass/fail, each file run alone | A total is unattributable, and `npm test` runs files concurrently where three timing-sensitive files flake. Running each alone is the only reproducible recording. |
## Why there is no regex enumerator
The obvious design — parse the sources and enumerate every regex literal — was
rejected. It needs a JS parser this zero-dependency repo does not have, and a
lexical approximation is contaminated by comments and division (`severity.mjs`
scores 4 "regexes" that way and exports none). The two-layer split — export
walk for what the scanners actually use, file digest for everything else —
answers the same question without a parser.
## During a swap
A diff here means the swap changed observable behaviour. **Roll the swap back.**
`--write` is for deliberately re-blessing a change you have already decided is
correct, not for making the gate quiet.
## Known coverage gaps — read before quoting a number
**`coverage` is static reachability, NOT observed coverage.** It probes the 61
payload strings against the 83 injection patterns in-process: "if you threw
every payload string at every regex, how many would match?" It does **not**
measure what the 61 hook invocations evaluated — the `pre-bash-destructive`
payloads never reach injection-patterns at all, yet their strings are in the
probe set and can mark a pattern reachable. `47/83` is therefore an upper
bound on what the corpus could protect, not a measurement of what it did.
The block is named `kind: "static-reachability"` and carries that caveat in a
`note` field; the gate asserts both, so the honest label cannot be dropped
quietly.
Two gaps follow:
- **36 unreachable patterns**, listed by key under
`coverage.unreachablePatterns`. A swap that breaks one of those is caught by
the pattern dump only. Closing it means growing the conformance corpus
(Phase 5 step 5).
- **The four OWASP maps have no behavioural coverage at all.** They are
scanner-side and no hook in this corpus reaches them. They were the reason
the dump was widened to table records, and the digest is their *only*
protection — do not read the reference run as backing them.
`corpusContains` says a payload *contains* a homoglyph or is altered by
`normalizeForScan`. It does not say the run folded one.