The EchoLeak class (CVE-2025-32711): markdown image/link/refdef/autolink, data: URIs and active HTML, plus the URL-shape analysis that separates a URL naming a remote document from one carrying bytes outward. Seeded from llm-ingestion-pipeline-security v0.3.4 @ 0bf0729 (read-only) — active_content.py's pattern table and calibration.py's severities/floors. That module documents itself as the canonical home of this table, which is why the guard is the source here rather than llm-security. One deliberate deviation from byte-identity, recorded in the file: five patterns carried `\"` from Python raw strings, which is a SyntaxError under ECMAScript's u/v flags. Normalised to `"` and proven equivalent by differential match-set comparison (5 patterns x 30 adversarial inputs, 0 differences). Verification log in docs/extraction-plan.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0191AKc2qW6tmXDFSx1xn53q
6.3 KiB
Extraction plan — v0.1.0
Status: informative. This is the plan of record for how this repository came to exist,
copied verbatim (structure preserved, lightly reformatted) from the operator brief that
opened it. It is not normative: nothing here constrains a consumer. When it disagrees
with spec/ or schema/, those win.
Origin: Phase 4 of the llm-security v8 plan, which lives in the sibling repository
llm-security. That repository is context only — no session in this repository reads from
or writes to it.
Charter
No engine code. Only: JSON data, normative specs, and a conformance corpus that several
runtimes (Node in llm-security, Python in a guard repo, a wiki) can run against and get
an identical verdict from. The pattern is copied from the sibling repository
portfolio-optimiser-commons (hard charter: "nothing here may import/depend on a
framework").
Layout
llm-security-commons/
README.md # charter: data+contract+fixtures only, no engine code
lexicon/injection-lexicon.json
codepoints/carriers.json # zero-width, BIDI, Unicode-Tag ranges, homoglyph map
signatures/secret-egress.json
signatures/malware-signatures.json
signatures/active-content.json # EchoLeak: MD image/link/refdef/autolink, data:, active HTML
calibration/calibration.json # entropy floors, scan caps, disposition ranks
mapping/owasp-map.json # prefix -> LLM/ASI/AST/MCP
schema/finding.schema.json # + SARIF & JSONL profiles. Status: normative
spec/decode-pipeline.md # normative RFC-2119 decode order
conformance/ # {case}/input.txt + {case}/expected.json
STATE.md # LOCAL-ONLY / gitignored (mirror commons convention)
Every JSON file carries a top-level "version" field. Every spec carries a
Status: normative marker.
v0.1.0 seed sources
llm-security is the canonical and richest source. This repository's sessions have no
read access to it — content arrives only as an operator-supplied dump. Security-critical
tables (homoglyph map, secret patterns, malware signatures) MUST come from real source
data, never from recollection or inference.
| Target | Seed source in llm-security (unless noted) |
|---|---|
lexicon/injection-lexicon.json |
scanners/lib/injection-patterns.mjs |
codepoints/carriers.json |
scanners/unicode-scanner.mjs + scanners/lib/string-utils.mjs (incl. HOMOGLYPH_MAP) |
signatures/secret-egress.json |
knowledge/secrets-patterns.md — the 18-entry hook table, NOT the PCRE-flavored agent-consumed variant |
signatures/malware-signatures.json |
knowledge/signatures.json (the SIG scanner) |
signatures/active-content.json |
currently only in a guard repo's active_content.py. If unavailable: stub with a version field and a TODO naming the source |
calibration/calibration.json |
scanners/lib/severity.mjs — thresholds + scanner caps |
mapping/owasp-map.json |
scanners/lib/severity.mjs — OWASP_MAP (+ 3 sibling maps in the same file) |
schema/finding.schema.json |
modelled on scanners/lib/sarif-formatter.mjs's SARIF shape |
conformance/ |
union of the guard repo's coverage.py matrix (126 classes + 4 gaps-must-hold) and llm-security/examples/ |
Constraints
- Offline / deterministic only — no network, no model calls inside the data itself.
- Forgejo
open/— never GitHub. - MIT license, fork-and-own.
STATE.mdis LOCAL-ONLY (gitignored) — same convention as the rest of the polyrepo.- Behaviour preservation is the point: this must not change a single finding in
llm-securitywhen it is later consumed from here. That consumption happens inllm-security's own Phase 5 steps 3–4 — not here.
Verification log
Every claim of fidelity below was produced by a command, not by reading. The check scripts themselves deliberately do not live in this repository — executable code here would breach the charter. They are reproducible from the description given.
signatures/active-content.json — extracted 2026-08-09
Source: llm-ingestion-pipeline-security v0.3.4, commit 0bf0729 (2026-08-03),
src/llm_ingestion_guard/active_content.py + calibration.py. Read-only; nothing in that
repository was modified.
| Check | Method | Result |
|---|---|---|
| JSON well-formed | python3 -m json.tool |
pass |
Patterns compile as Python re |
translate (?< → (?P<, compile all 17 with declared flags |
17/17, 0 failures |
| Patterns compile as ECMAScript | new RegExp(pattern, flags) on all 17 |
17/17, 0 failures |
| Pattern text matches source | compare against the live re.Pattern.pattern of each source object, inline flags stripped |
12/17 byte-identical; 5 differ only by the documented redundant-quote-escape normalisation |
| The 5 normalised patterns behave identically | differential match-set comparison (offsets + captured text) against the source objects over a 30-input adversarial corpus: bare quotes, escaped quotes, markdown titles containing quotes, quoted/unquoted HTML attributes, quote runs of length 1–5 | 150 comparisons, 0 differences |
| The normalisation is necessary | new RegExp('\\"', 'u') and 'v' in Node |
both throw Invalid escape; the bare form compiles under "", "u" and "v" |
| Severities, ordinary severity, opacity floors, active-tag set, pass order | compare against calibration.ACTIVE_CONTENT_SEVERITY, ACTIVE_CONTENT_ORDINARY_SEVERITY, URL_OPAQUE_*, active_content._ACTIVE_TAGS, and the scan-call order in scan_active_content |
all identical (23/23 tags, 6/6 severities, 4/4 floors) |
Not verified, and not claimed: that the Node consumer's active-content behaviour matches this table. The source module states the Node port shares its severities; that is the module's claim, and confirming it needs the Node file.
Definition of done for v0.1.0
- Repository initialized, Forgejo remote
open/llm-security-commons, MIT,STATE.mdgitignored. - Every file in the layout above present and populated from verified seed data — or explicitly and visibly stubbed where the source was unavailable.
- All JSON well-formed, every data file carrying
"version", every spec carryingStatus: normative. - Tagged
v0.1.0and pushed. - A
coordmessage sent tollm-securityannouncing that the repository andv0.1.0exist, so Phase 5 step 3 (vendoring) can start from there.