The EchoLeak class (CVE-2025-32711): markdown image/link/refdef/autolink, data: URIs and active HTML, plus the URL-shape analysis that separates a URL naming a remote document from one carrying bytes outward. Seeded from llm-ingestion-pipeline-security v0.3.4 @ 0bf0729 (read-only) — active_content.py's pattern table and calibration.py's severities/floors. That module documents itself as the canonical home of this table, which is why the guard is the source here rather than llm-security. One deliberate deviation from byte-identity, recorded in the file: five patterns carried `\"` from Python raw strings, which is a SyntaxError under ECMAScript's u/v flags. Normalised to `"` and proven equivalent by differential match-set comparison (5 patterns x 30 adversarial inputs, 0 differences). Verification log in docs/extraction-plan.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0191AKc2qW6tmXDFSx1xn53q
106 lines
6.3 KiB
Markdown
106 lines
6.3 KiB
Markdown
# Extraction plan — v0.1.0
|
||
|
||
**Status: informative.** This is the plan of record for how this repository came to exist,
|
||
copied verbatim (structure preserved, lightly reformatted) from the operator brief that
|
||
opened it. It is **not** normative: nothing here constrains a consumer. When it disagrees
|
||
with `spec/` or `schema/`, those win.
|
||
|
||
Origin: **Phase 4 of the `llm-security` v8 plan**, which lives in the sibling repository
|
||
`llm-security`. That repository is context only — no session in this repository reads from
|
||
or writes to it.
|
||
|
||
## Charter
|
||
|
||
No engine code. Only: JSON data, normative specs, and a conformance corpus that several
|
||
runtimes (Node in `llm-security`, Python in a guard repo, a wiki) can run against and get
|
||
an identical verdict from. The pattern is copied from the sibling repository
|
||
`portfolio-optimiser-commons` (hard charter: "nothing here may import/depend on a
|
||
framework").
|
||
|
||
## Layout
|
||
|
||
```
|
||
llm-security-commons/
|
||
README.md # charter: data+contract+fixtures only, no engine code
|
||
lexicon/injection-lexicon.json
|
||
codepoints/carriers.json # zero-width, BIDI, Unicode-Tag ranges, homoglyph map
|
||
signatures/secret-egress.json
|
||
signatures/malware-signatures.json
|
||
signatures/active-content.json # EchoLeak: MD image/link/refdef/autolink, data:, active HTML
|
||
calibration/calibration.json # entropy floors, scan caps, disposition ranks
|
||
mapping/owasp-map.json # prefix -> LLM/ASI/AST/MCP
|
||
schema/finding.schema.json # + SARIF & JSONL profiles. Status: normative
|
||
spec/decode-pipeline.md # normative RFC-2119 decode order
|
||
conformance/ # {case}/input.txt + {case}/expected.json
|
||
STATE.md # LOCAL-ONLY / gitignored (mirror commons convention)
|
||
```
|
||
|
||
Every JSON file carries a top-level `"version"` field. Every spec carries a
|
||
`Status: normative` marker.
|
||
|
||
## v0.1.0 seed sources
|
||
|
||
`llm-security` is the canonical and richest source. **This repository's sessions have no
|
||
read access to it** — content arrives only as an operator-supplied dump. Security-critical
|
||
tables (homoglyph map, secret patterns, malware signatures) MUST come from real source
|
||
data, never from recollection or inference.
|
||
|
||
| Target | Seed source in `llm-security` (unless noted) |
|
||
|---|---|
|
||
| `lexicon/injection-lexicon.json` | `scanners/lib/injection-patterns.mjs` |
|
||
| `codepoints/carriers.json` | `scanners/unicode-scanner.mjs` + `scanners/lib/string-utils.mjs` (incl. `HOMOGLYPH_MAP`) |
|
||
| `signatures/secret-egress.json` | `knowledge/secrets-patterns.md` — the **18-entry hook table**, NOT the PCRE-flavored agent-consumed variant |
|
||
| `signatures/malware-signatures.json` | `knowledge/signatures.json` (the SIG scanner) |
|
||
| `signatures/active-content.json` | currently only in a guard repo's `active_content.py`. If unavailable: stub with a version field and a TODO naming the source |
|
||
| `calibration/calibration.json` | `scanners/lib/severity.mjs` — thresholds + scanner caps |
|
||
| `mapping/owasp-map.json` | `scanners/lib/severity.mjs` — `OWASP_MAP` (+ 3 sibling maps in the same file) |
|
||
| `schema/finding.schema.json` | modelled on `scanners/lib/sarif-formatter.mjs`'s SARIF shape |
|
||
| `conformance/` | union of the guard repo's `coverage.py` matrix (126 classes + 4 gaps-must-hold) and `llm-security/examples/` |
|
||
|
||
## Constraints
|
||
|
||
- Offline / deterministic only — no network, no model calls inside the data itself.
|
||
- Forgejo `open/` — never GitHub.
|
||
- MIT license, fork-and-own.
|
||
- `STATE.md` is LOCAL-ONLY (gitignored) — same convention as the rest of the polyrepo.
|
||
- **Behaviour preservation is the point:** this must not change a single finding in
|
||
`llm-security` when it is later consumed from here. That consumption happens in
|
||
`llm-security`'s own Phase 5 steps 3–4 — **not here.**
|
||
|
||
## Verification log
|
||
|
||
Every claim of fidelity below was produced by a command, not by reading. The check scripts
|
||
themselves deliberately do **not** live in this repository — executable code here would
|
||
breach the charter. They are reproducible from the description given.
|
||
|
||
### `signatures/active-content.json` — extracted 2026-08-09
|
||
|
||
Source: `llm-ingestion-pipeline-security` v0.3.4, commit `0bf0729` (2026-08-03),
|
||
`src/llm_ingestion_guard/active_content.py` + `calibration.py`. Read-only; nothing in that
|
||
repository was modified.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed | `python3 -m json.tool` | pass |
|
||
| Patterns compile as Python `re` | translate `(?<` → `(?P<`, compile all 17 with declared flags | 17/17, 0 failures |
|
||
| Patterns compile as ECMAScript | `new RegExp(pattern, flags)` on all 17 | 17/17, 0 failures |
|
||
| Pattern text matches source | compare against the live `re.Pattern.pattern` of each source object, inline flags stripped | 12/17 byte-identical; 5 differ only by the documented `redundant-quote-escape` normalisation |
|
||
| The 5 normalised patterns behave identically | differential match-set comparison (offsets + captured text) against the source objects over a 30-input adversarial corpus: bare quotes, escaped quotes, markdown titles containing quotes, quoted/unquoted HTML attributes, quote runs of length 1–5 | 150 comparisons, 0 differences |
|
||
| The normalisation is necessary | `new RegExp('\\"', 'u')` and `'v'` in Node | both throw `Invalid escape`; the bare form compiles under `""`, `"u"` and `"v"` |
|
||
| Severities, ordinary severity, opacity floors, active-tag set, pass order | compare against `calibration.ACTIVE_CONTENT_SEVERITY`, `ACTIVE_CONTENT_ORDINARY_SEVERITY`, `URL_OPAQUE_*`, `active_content._ACTIVE_TAGS`, and the scan-call order in `scan_active_content` | all identical (23/23 tags, 6/6 severities, 4/4 floors) |
|
||
|
||
Not verified, and not claimed: that the Node consumer's active-content behaviour matches
|
||
this table. The source module states the Node port shares its severities; that is the
|
||
module's claim, and confirming it needs the Node file.
|
||
|
||
## Definition of done for v0.1.0
|
||
|
||
1. Repository initialized, Forgejo remote `open/llm-security-commons`, MIT, `STATE.md`
|
||
gitignored.
|
||
2. Every file in the layout above present and populated from verified seed data — or
|
||
explicitly and visibly stubbed where the source was unavailable.
|
||
3. All JSON well-formed, every data file carrying `"version"`, every spec carrying
|
||
`Status: normative`.
|
||
4. Tagged `v0.1.0` and pushed.
|
||
5. A `coord` message sent to `llm-security` announcing that the repository and `v0.1.0`
|
||
exist, so Phase 5 step 3 (vendoring) can start from there.
|