llm-security-commons/docs/extraction-plan.md
Kjell Tore Guttormsen f4aa8b66f4 feat(codepoints): add carriers.json from verified llm-security dump
Six independent carrier tables: zero-width characters (5), the Unicode Tags
block with its subtraction decode rule, the two Supplementary Private Use
Areas, BIDI controls (9, the Trojan Source class CVE-2021-42574), the
Cyrillic presence set (13), and the fold-to-Latin homoglyph map (28).

Proven, not transcribed: five of the six tables were rebuilt from the commons
JSON alone and diffed against the imported dump module — every constant
identical, and the homoglyph map identical down to insertion order. Folding a
corpus through the rebuilt table and the source table gives identical results.

The tables overlap but are NOT merged, and cross_table_notes states each
divergence as fact: the zero-width carrier set includes U+00AD while the
lexicon's pattern class does not; the lexicon class holds U+0456 which
CYRILLIC_CONFUSABLES lacks, and CYRILLIC_CONFUSABLES holds U+0445 which the
class lacks. Reporting that to llm-security, not fixing it here.

Two honest limits, marked in the file rather than smoothed over: the private
use ranges arrived as a source comment with no constant behind them and carry
verified: false, and the dump's own "~25 entries" estimate for the homoglyph
map is wrong — counted mechanically it is 28.

Character names resolved through Python unicodedata against the Unicode
character database, not written from recollection. Verification log in
docs/extraction-plan.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FaYqid3mejFmd9ZHsiHgp3
2026-08-09 21:08:01 +02:00

177 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Extraction plan — v0.1.0
**Status: informative.** This is the plan of record for how this repository came to exist,
copied verbatim (structure preserved, lightly reformatted) from the operator brief that
opened it. It is **not** normative: nothing here constrains a consumer. When it disagrees
with `spec/` or `schema/`, those win.
Origin: **Phase 4 of the `llm-security` v8 plan**, which lives in the sibling repository
`llm-security`. That repository is context only — no session in this repository reads from
or writes to it.
## Charter
No engine code. Only: JSON data, normative specs, and a conformance corpus that several
runtimes (Node in `llm-security`, Python in a guard repo, a wiki) can run against and get
an identical verdict from. The pattern is copied from the sibling repository
`portfolio-optimiser-commons` (hard charter: "nothing here may import/depend on a
framework").
## Layout
```
llm-security-commons/
README.md # charter: data+contract+fixtures only, no engine code
lexicon/injection-lexicon.json
codepoints/carriers.json # zero-width, BIDI, Unicode-Tag ranges, homoglyph map
signatures/secret-egress.json
signatures/malware-signatures.json
signatures/active-content.json # EchoLeak: MD image/link/refdef/autolink, data:, active HTML
calibration/calibration.json # entropy floors, scan caps, disposition ranks
mapping/owasp-map.json # prefix -> LLM/ASI/AST/MCP
schema/finding.schema.json # + SARIF & JSONL profiles. Status: normative
spec/decode-pipeline.md # normative RFC-2119 decode order
conformance/ # {case}/input.txt + {case}/expected.json
STATE.md # LOCAL-ONLY / gitignored (mirror commons convention)
```
Every JSON file carries a top-level `"version"` field. Every spec carries a
`Status: normative` marker.
## v0.1.0 seed sources
`llm-security` is the canonical and richest source. **This repository's sessions have no
read access to it** — content arrives only as an operator-supplied dump. Security-critical
tables (homoglyph map, secret patterns, malware signatures) MUST come from real source
data, never from recollection or inference.
| Target | Seed source in `llm-security` (unless noted) |
|---|---|
| `lexicon/injection-lexicon.json` | `scanners/lib/injection-patterns.mjs` |
| `codepoints/carriers.json` | `scanners/unicode-scanner.mjs` + `scanners/lib/string-utils.mjs` (incl. `HOMOGLYPH_MAP`) |
| `signatures/secret-egress.json` | `knowledge/secrets-patterns.md` — the **18-entry hook table**, NOT the PCRE-flavored agent-consumed variant |
| `signatures/malware-signatures.json` | `knowledge/signatures.json` (the SIG scanner) |
| `signatures/active-content.json` | currently only in a guard repo's `active_content.py`. If unavailable: stub with a version field and a TODO naming the source |
| `calibration/calibration.json` | `scanners/lib/severity.mjs` — thresholds + scanner caps |
| `mapping/owasp-map.json` | `scanners/lib/severity.mjs``OWASP_MAP` (+ 3 sibling maps in the same file) |
| `schema/finding.schema.json` | modelled on `scanners/lib/sarif-formatter.mjs`'s SARIF shape |
| `conformance/` | union of the guard repo's `coverage.py` matrix (126 classes + 4 gaps-must-hold) and `llm-security/examples/` |
## Constraints
- Offline / deterministic only — no network, no model calls inside the data itself.
- Forgejo `open/` — never GitHub.
- MIT license, fork-and-own.
- `STATE.md` is LOCAL-ONLY (gitignored) — same convention as the rest of the polyrepo.
- **Behaviour preservation is the point:** this must not change a single finding in
`llm-security` when it is later consumed from here. That consumption happens in
`llm-security`'s own Phase 5 steps 34 — **not here.**
## Verification log
Every claim of fidelity below was produced by a command, not by reading. The check scripts
themselves deliberately do **not** live in this repository — executable code here would
breach the charter. They are reproducible from the description given.
### `signatures/active-content.json` — extracted 2026-08-09
Source: `llm-ingestion-pipeline-security` v0.3.4, commit `0bf0729` (2026-08-03),
`src/llm_ingestion_guard/active_content.py` + `calibration.py`. Read-only; nothing in that
repository was modified.
| Check | Method | Result |
| --- | --- | --- |
| JSON well-formed | `python3 -m json.tool` | pass |
| Patterns compile as Python `re` | translate `(?<``(?P<`, compile all 17 with declared flags | 17/17, 0 failures |
| Patterns compile as ECMAScript | `new RegExp(pattern, flags)` on all 17 | 17/17, 0 failures |
| Pattern text matches source | compare against the live `re.Pattern.pattern` of each source object, inline flags stripped | 12/17 byte-identical; 5 differ only by the documented `redundant-quote-escape` normalisation |
| The 5 normalised patterns behave identically | differential match-set comparison (offsets + captured text) against the source objects over a 30-input adversarial corpus: bare quotes, escaped quotes, markdown titles containing quotes, quoted/unquoted HTML attributes, quote runs of length 15 | 150 comparisons, 0 differences |
| The normalisation is necessary | `new RegExp('\\"', 'u')` and `'v'` in Node | both throw `Invalid escape`; the bare form compiles under `""`, `"u"` and `"v"` |
| Severities, ordinary severity, opacity floors, active-tag set, pass order | compare against `calibration.ACTIVE_CONTENT_SEVERITY`, `ACTIVE_CONTENT_ORDINARY_SEVERITY`, `URL_OPAQUE_*`, `active_content._ACTIVE_TAGS`, and the scan-call order in `scan_active_content` | all identical (23/23 tags, 6/6 severities, 4/4 floors) |
Not verified, and not claimed: that the Node consumer's active-content behaviour matches
this table. The source module states the Node port shares its severities; that is the
module's claim, and confirming it needs the Node file.
### `schema/finding.schema.json` — extracted 2026-08-09
Source: `llm-security/scanners/lib/sarif-formatter.mjs`, supplied as an operator dump. No
commit hash accompanied it, so provenance is recorded as `unknown` rather than guessed.
| Check | Method | Result |
| --- | --- | --- |
| JSON well-formed | `python3 -m json.tool` | pass |
| Valid JSON Schema | `jsonschema` `check_schema` against draft 2020-12 | pass |
| Accepts/rejects findings correctly | 2 valid + 3 invalid findings (missing `scanner`, unknown severity, `line: 0`) | 5/5 as intended |
| SARIF profile reproduces the source | re-implemented the mapping **from the commons JSON alone** and diffed `JSON.stringify` against the real `toSARIF` over 10 envelope shapes: empty, missing `scanners`, empty `scanners`, scanner with no findings, all five severities plus an unknown and an `undefined` one, five slug edge cases (double space, tab, newline, leading/trailing space, mixed case), a rule-id collision, all seven optional-field combinations, two scanners, and an explicit `version` argument | 10/10 identical, 0 differences |
| The three `known_lossiness` claims are true | executed each against the real formatter | all three confirmed, **and one earlier claim corrected**: punctuation does *not* collapse — the slug lowercases and collapses whitespace only, so `Zero-width carrier` and `Zero-width carrier!` remain distinct ids. The wrong claim was published in the first draft of this file and fixed before commit. |
Not verified, and recorded in the file as open: the finding **producer** was not supplied, so
the property list is a lower bound; `scanner` and `severity` are required by design rather
than by evidence; and the JSONL profile is left explicitly `unspecified` rather than
invented, because "one finding per line" is inference.
### `lexicon/injection-lexicon.json` — extracted 2026-08-09
Source: `llm-security/scanners/lib/injection-patterns.mjs`, supplied as operator dump 2/2
through the local coord mailbox. No commit hash accompanied it, so provenance is recorded
as `unknown` rather than guessed.
| Check | Method | Result |
| --- | --- | --- |
| JSON well-formed, `version` present, LF, trailing newline, no raw invisible code points | `python3 -m json.tool` + a byte scan for U+200B/200C/200D/FEFF/00AD and the Tag block | pass, 0 raw invisible code points |
| Pattern text and flags reproduce the source | rebuilt all four arrays **from the commons JSON alone** (`new RegExp(p.pattern, p.flags ?? '')`) and diffed label, `.source` and `.flags` against the imported dump module | 83/83 compared, 0 differences; 81/83 byte-identical, 2 declared-normalised |
| Flags were read mechanically, not by eye | extracted from each literal via `.flags` | critical 15×`i` / 3×`m` / 3 none, high 32×`i`, medium 20×`i` / 2 none, hybrid 8×`i` |
| Every pattern compiles in both runtimes | `new RegExp(src, flags)` and again with `u` in Node; `re.compile` with the equivalent `re.I`/`re.M` in Python | 83/83 in all three modes, 0 failures |
| The 2 normalised patterns behave identically | differential match-set comparison (offsets + matched text) against the source objects, bare and under `u`, over a 208-input adversarial corpus: every class member, the near-misses excluded from each class (U+00AD, U+2060, U+180E, Cyrillic х, the uppercase set, Greek look-alikes), run boundaries, repeats, empty input | 832 comparisons, 0 differences |
| Class membership was counted, not assumed | enumerated the code points inside each character class directly from the dump bytes | zero-width class = 4 (U+200B, U+200C, U+200D, U+FEFF — **not** U+00AD); Cyrillic class = 7 (U+0430, U+0435, U+043E, U+0440, U+0441, U+0456, U+0443) |
| `\/` is portable, not a defect | 9 patterns carry the redundant escape a JS regex literal requires; compiled in Node bare, Node `u`, and Python `re` | accepted by all three — kept byte-identical, recorded as a translation note for engines that reject unknown escapes |
Not verified, and not claimed: that the dump matches the module it was transcribed from.
Every check above proves this JSON agrees with **the dump**; dump-to-module fidelity is
`llm-security`'s assertion, reproducible only in a session with read access to that
repository. The severity the engine assigns to `HYBRID_PATTERNS` was not supplied and is
left `null` rather than inferred from its three sibling arrays.
### `codepoints/carriers.json` — extracted 2026-08-09
Source: `llm-security/scanners/unicode-scanner.mjs` (charset constants) and
`llm-security/scanners/lib/string-utils.mjs` (`HOMOGLYPH_MAP`), supplied as operator dump
2/2 through the local coord mailbox. No commit hash accompanied it.
| Check | Method | Result |
| --- | --- | --- |
| JSON well-formed, `version` present, no raw invisible code points | `python3 -m json.tool` + byte scan for zero-width, BIDI and Tag-block characters | pass, 0 raw invisible code points |
| Five of the six tables reproduce the source constants | rebuilt each **from the commons JSON alone** (`parseInt(codepoint.slice(2), 16)`) and diffed against the imported dump module | `ZERO_WIDTH_CHARS` 5/5, `BIDI_CHARS` 9/9, `CYRILLIC_CONFUSABLES` 13/13, tag start/end — 0 differences |
| The homoglyph map reproduces the source, including order | rebuilt the object from the entries array and compared keys, values and the whole object | 28/28 keys, values and insertion order identical |
| The map folds identically | applied NFKC + lookup with both the rebuilt and the source table over 12 inputs (Cyrillic and Greek injection spellings, Norwegian and German orthography, empty) | 0 differences |
| The exclusion rationale in the source comment is true | checked whether any of `帿ŨÆäöüßéèêñç` is a key | 0 touched — ordinary Norwegian and German orthography is not folded |
| Convenience `char` fields agree with their own `codepoint` field | `String.fromCodePoint` round-trip on every entry | 41/41, 0 mismatches |
| Character names are not from recollection | resolved every name through Python `unicodedata` against the Unicode character database | all resolved |
| Table sizes were counted, not quoted | counted from the imported constants | homoglyph map holds **28** entries, not the "~25" the dump's own comment estimates; the counted number is the one recorded |
Not verified, and marked `verified: false` **in the file itself**: the two Supplementary
Private Use Area ranges. They arrived as a source comment with no constant behind them, so
unlike the other five tables there was nothing to import and diff. That asymmetry is
recorded per-table rather than averaged into a single file-level verdict.
Recorded and deliberately **not** reconciled, in `cross_table_notes`: the repository now
holds three overlapping Cyrillic sets and two overlapping zero-width sets, and none agree
exactly. The zero-width carrier table includes U+00AD while the lexicon's pattern class does
not; the lexicon's class contains U+0456 which `CYRILLIC_CONFUSABLES` lacks, and
`CYRILLIC_CONFUSABLES` contains U+0445 which the class lacks; and six of the confusables
have no entry in the fold map. The dump states the presence set and the fold map are
deliberately distinct. The U+0456 / U+0445 divergence is reported to `llm-security` rather
than fixed here.
## Definition of done for v0.1.0
1. Repository initialized, Forgejo remote `open/llm-security-commons`, MIT, `STATE.md`
gitignored.
2. Every file in the layout above present and populated from verified seed data — or
explicitly and visibly stubbed where the source was unavailable.
3. All JSON well-formed, every data file carrying `"version"`, every spec carrying
`Status: normative`.
4. Tagged `v0.1.0` and pushed.
5. A `coord` message sent to `llm-security` announcing that the repository and `v0.1.0`
exist, so Phase 5 step 3 (vendoring) can start from there.