feat(signatures): add secret-egress.json from verified llm-security dump

The 18 fixed credential and token shapes a pre-write guard matches before
content is persisted: cloud keys, vendor tokens, PEM blocks, connection
strings, JWTs.

Proven, not transcribed: the table was rebuilt from the commons JSON alone
and diffed against the imported dump module — 18/18 identical on name, source
and flags, and all 18 byte-identical, so no normalisation was needed. All 18
compile in Node bare, Node under `u`, and Python `re`.

Array order is normative and is tested as such, not merely asserted: a Bearer
header containing a JWT must be labelled "Authorization header with token"
rather than "JWT (three-part token)", which is why the source puts the bare
JWT entry last. Reproduced from the commons order, and shown to change under a
reversed table. Every entry carries an explicit `order` field so a JSON
round-trip cannot reorder the contract silently.

Corrects the extraction plan's seed-source row in the same commit: it named
knowledge/secrets-patterns.md, but the dump named hooks/scripts/
pre-edit-secrets.mjs and stated the two are different tables. Recording a
source file that was never delivered is the defect class this repository
already caught once in finding.schema.json.

Neither severity nor disposition was supplied, so neither is invented — the
source table carries a name and a pattern and nothing else. Verification log
in docs/extraction-plan.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FaYqid3mejFmd9ZHsiHgp3
This commit is contained in:
Kjell Tore Guttormsen 2026-08-09 21:10:01 +02:00
commit c3852c9148
2 changed files with 177 additions and 1 deletions

View file

@ -49,7 +49,7 @@ data, never from recollection or inference.
|---|---|
| `lexicon/injection-lexicon.json` | `scanners/lib/injection-patterns.mjs` |
| `codepoints/carriers.json` | `scanners/unicode-scanner.mjs` + `scanners/lib/string-utils.mjs` (incl. `HOMOGLYPH_MAP`) |
| `signatures/secret-egress.json` | `knowledge/secrets-patterns.md` — the **18-entry hook table**, NOT the PCRE-flavored agent-consumed variant |
| `signatures/secret-egress.json` | `hooks/scripts/pre-edit-secrets.mjs` — `SECRET_PATTERNS`, the **18-entry hook table**, NOT the PCRE-flavored agent-consumed variant in `knowledge/secrets-patterns.md`. *(Corrected 2026-08-09: this row originally named `knowledge/secrets-patterns.md` as the source file. The delivered dump named `pre-edit-secrets.mjs` and stated explicitly that the two are different files. The row now names the file that was actually delivered.)* |
| `signatures/malware-signatures.json` | `knowledge/signatures.json` (the SIG scanner) |
| `signatures/active-content.json` | currently only in a guard repo's `active_content.py`. If unavailable: stub with a version field and a TODO naming the source |
| `calibration/calibration.json` | `scanners/lib/severity.mjs` — thresholds + scanner caps |
@ -164,6 +164,33 @@ have no entry in the fold map. The dump states the presence set and the fold map
deliberately distinct. The U+0456 / U+0445 divergence is reported to `llm-security` rather
than fixed here.
### `signatures/secret-egress.json` — extracted 2026-08-09
Source: `llm-security/hooks/scripts/pre-edit-secrets.mjs` (`SECRET_PATTERNS`), supplied as
operator dump 2/2 through the local coord mailbox. No commit hash accompanied it.
**The seed-source row above was wrong and has been corrected.** It named
`knowledge/secrets-patterns.md`; the dump named `hooks/scripts/pre-edit-secrets.mjs` and
stated that the two are different tables — the second is PCRE-flavoured and agent-consumed
and stays where it is. Recording a source file that was never delivered is the same defect
class as the lossiness claim corrected in `finding.schema.json`, so it is corrected here in
the same commit as the file it describes.
| Check | Method | Result |
| --- | --- | --- |
| JSON well-formed, `version` present, LF, trailing newline | `python3 -m json.tool` + byte scan | pass |
| Pattern text and flags reproduce the source | rebuilt the table **from the commons JSON alone**, sorted by the declared `order`, and diffed name, `.source` and `.flags` against the imported dump module | 18/18, 0 differences, **18/18 byte-identical** — no normalisation needed |
| Every pattern compiles in both runtimes | `new RegExp` bare and under `u` in Node; `re.compile` with `re.I` where declared in Python | 18/18 in all three modes, 0 failures |
| The ordering contract holds, and is not decorative | reproduced first-match labelling from the commons order for a Bearer header containing a JWT and for a bare JWT, against the source table | both labels identical to source: header case → `Authorization header with token`, bare case → `JWT (three-part token)` |
| Reordering is detectable, not silent | ran the same Bearer input through a reversed table | label changes to `JWT (three-part token)` — order is load-bearing, which is why every entry carries an explicit `order` field |
| `order` is contiguous | compared to `range(18)` | 017, no gaps |
Not verified, and not claimed: that the dump matches the module. Not supplied, and therefore
not invented: any severity or per-entry disposition — the source table carries a name and a
pattern and nothing else. Out of scope by the dump's own statement: the runtime
policy-injected custom patterns (entries 19+). A consumer matching only this table matches
**less** than the seed hook does when a policy is loaded.
## Definition of done for v0.1.0
1. Repository initialized, Forgejo remote `open/llm-security-commons`, MIT, `STATE.md`