432 lines
33 KiB
Markdown
432 lines
33 KiB
Markdown
# Extraction plan — v0.1.0
|
||
|
||
**Status: informative.** This is the plan of record for how this repository came to exist,
|
||
copied verbatim (structure preserved, lightly reformatted) from the operator brief that
|
||
opened it. It is **not** normative: nothing here constrains a consumer. When it disagrees
|
||
with `spec/` or `schema/`, those win.
|
||
|
||
Origin: **Phase 4 of the `llm-security` v8 plan**, which lives in the sibling repository
|
||
`llm-security`. That repository is context only — no session in this repository reads from
|
||
or writes to it.
|
||
|
||
## Charter
|
||
|
||
No engine code. Only: JSON data, normative specs, and a conformance corpus that several
|
||
runtimes (Node in `llm-security`, Python in a guard repo, a wiki) can run against and get
|
||
an identical verdict from. The pattern is copied from the sibling repository
|
||
`portfolio-optimiser-commons` (hard charter: "nothing here may import/depend on a
|
||
framework").
|
||
|
||
## Layout
|
||
|
||
```
|
||
llm-security-commons/
|
||
README.md # charter: data+contract+fixtures only, no engine code
|
||
lexicon/injection-lexicon.json
|
||
codepoints/carriers.json # zero-width, BIDI, Unicode-Tag ranges, homoglyph map
|
||
signatures/secret-egress.json
|
||
signatures/malware-signatures.json
|
||
signatures/active-content.json # EchoLeak: MD image/link/refdef/autolink, data:, active HTML
|
||
calibration/calibration.json # risk-score tiers, verdict/band cutoffs, grade thresholds
|
||
mapping/owasp-map.json # prefix -> LLM/ASI/AST/MCP
|
||
schema/finding.schema.json # + SARIF & JSONL profiles. Status: normative
|
||
spec/decode-pipeline.md # normative RFC-2119 decode order
|
||
conformance/ # {case}/input.txt + {case}/expected.json
|
||
STATE.md # LOCAL-ONLY / gitignored (mirror commons convention)
|
||
```
|
||
|
||
Every JSON file carries a top-level `"version"` field. Every spec carries a
|
||
`Status: normative` marker.
|
||
|
||
## v0.1.0 seed sources
|
||
|
||
`llm-security` is the canonical and richest source. Security-critical tables (homoglyph map,
|
||
secret patterns, malware signatures) MUST come from real source data, never from recollection
|
||
or inference.
|
||
|
||
**The read boundary changed on 2026-08-09 (operator decision), and it matters to how the
|
||
rows below should be read.** This section previously stated that sessions here have no read
|
||
access and that content arrives only as an operator-supplied dump. Sessions may now **read
|
||
`llm-security` read-only from its public remote at a pinned commit** —
|
||
`ssh://git@git.fromaitochitta.com/open/llm-security.git` @ `b0de0ca`, verified as
|
||
`origin/main` — by cloning into a scratchpad. The sibling working copy on this machine
|
||
remains **absolutely off limits**, as does any form of writing.
|
||
|
||
The consequence is a method change, not just a convenience: files seeded from a dump are
|
||
**differentially verified against the module** rather than trusted as transcriptions. That
|
||
change has already caught four attribution defects — a re-escaping story that never happened
|
||
in the lexicon, an inverted escape in one pattern, a severity citation pointing at a file with
|
||
no severities in it, and two family names misremembered in a working note. Every one was a
|
||
transcription artefact. Where a row below still says a dump was the delivery, the
|
||
`provenance` block in the file itself records whether that dump has since been superseded by
|
||
a direct read.
|
||
|
||
| Target | Seed source in `llm-security` (unless noted) |
|
||
|---|---|
|
||
| `lexicon/injection-lexicon.json` | `scanners/lib/injection-patterns.mjs` |
|
||
| `codepoints/carriers.json` | `scanners/unicode-scanner.mjs` + `scanners/lib/string-utils.mjs` (incl. `HOMOGLYPH_MAP`) |
|
||
| `signatures/secret-egress.json` | `hooks/scripts/pre-edit-secrets.mjs` — `SECRET_PATTERNS`, the **18-entry hook table**, NOT the PCRE-flavored agent-consumed variant in `knowledge/secrets-patterns.md`. *(Corrected 2026-08-09: this row originally named `knowledge/secrets-patterns.md` as the source file. The delivered dump named `pre-edit-secrets.mjs` and stated explicitly that the two are different files. The row now names the file that was actually delivered.)* |
|
||
| `signatures/malware-signatures.json` | `knowledge/signatures.json` (the SIG scanner) |
|
||
| `signatures/active-content.json` | currently only in a guard repo's `active_content.py`. If unavailable: stub with a version field and a TODO naming the source |
|
||
| `calibration/calibration.json` | `scanners/lib/severity.mjs` — risk-score tier constants, verdict thresholds, risk-band cutoffs, posture grade thresholds. *(Corrected 2026-08-09: this row said "thresholds + scanner caps" and the README said "entropy floors, scan caps, disposition ranks". No scan cap, entropy floor or disposition rank was delivered — searched across the whole dump: 0 occurrences each. Both rows now describe what arrived.)* |
|
||
| `mapping/owasp-map.json` | `scanners/lib/severity.mjs` — `OWASP_MAP` (+ 3 sibling maps in the same file) |
|
||
| `schema/finding.schema.json` | modelled on `scanners/lib/sarif-formatter.mjs`'s SARIF shape |
|
||
| `conformance/` | union of the guard repo's `coverage.py` matrix and `llm-security/examples/`. *(Corrected 2026-08-09: this row said "126 classes + 4 gaps-must-hold". Counted from the imported `CORE_CASES` manifest, the matrix holds **134** cases — 128 `caught` and **6** `gap`. See the conformance-source breakdown below.)* |
|
||
|
||
## Constraints
|
||
|
||
- Offline / deterministic only — no network, no model calls inside the data itself.
|
||
- Forgejo `open/` — never GitHub.
|
||
- MIT license, fork-and-own.
|
||
- `STATE.md` is LOCAL-ONLY (gitignored) — same convention as the rest of the polyrepo.
|
||
- **Behaviour preservation is the point:** this must not change a single finding in
|
||
`llm-security` when it is later consumed from here. That consumption happens in
|
||
`llm-security`'s own Phase 5 steps 3–4 — **not here.**
|
||
|
||
## Verification log
|
||
|
||
Every claim of fidelity below was produced by a command, not by reading. The check scripts
|
||
themselves deliberately do **not** live in this repository — executable code here would
|
||
breach the charter. They are reproducible from the description given.
|
||
|
||
### `signatures/active-content.json` — extracted 2026-08-09
|
||
|
||
Source: `llm-ingestion-pipeline-security` v0.3.4, commit `0bf0729` (2026-08-03),
|
||
`src/llm_ingestion_guard/active_content.py` + `calibration.py`. Read-only; nothing in that
|
||
repository was modified.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed | `python3 -m json.tool` | pass |
|
||
| Patterns compile as Python `re` | translate `(?<` → `(?P<`, compile all 17 with declared flags | 17/17, 0 failures |
|
||
| Patterns compile as ECMAScript | `new RegExp(pattern, flags)` on all 17 | 17/17, 0 failures |
|
||
| Pattern text matches source | compare against the live `re.Pattern.pattern` of each source object, inline flags stripped | 12/17 byte-identical; 5 differ only by the documented `redundant-quote-escape` normalisation |
|
||
| The 5 normalised patterns behave identically | differential match-set comparison (offsets + captured text) against the source objects over a 30-input adversarial corpus: bare quotes, escaped quotes, markdown titles containing quotes, quoted/unquoted HTML attributes, quote runs of length 1–5 | 150 comparisons, 0 differences |
|
||
| The normalisation is necessary | `new RegExp('\\"', 'u')` and `'v'` in Node | both throw `Invalid escape`; the bare form compiles under `""`, `"u"` and `"v"` |
|
||
| Severities, ordinary severity, opacity floors, active-tag set, pass order | compare against `calibration.ACTIVE_CONTENT_SEVERITY`, `ACTIVE_CONTENT_ORDINARY_SEVERITY`, `URL_OPAQUE_*`, `active_content._ACTIVE_TAGS`, and the scan-call order in `scan_active_content` | all identical (23/23 tags, 6/6 severities, 4/4 floors) |
|
||
|
||
Not verified, and not claimed: that the Node consumer's active-content behaviour matches
|
||
this table. The source module states the Node port shares its severities; that is the
|
||
module's claim, and confirming it needs the Node file.
|
||
|
||
### `schema/finding.schema.json` — extracted 2026-08-09
|
||
|
||
Source: `llm-security/scanners/lib/sarif-formatter.mjs`, supplied as an operator dump. No
|
||
commit hash accompanied it, so provenance is recorded as `unknown` rather than guessed.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed | `python3 -m json.tool` | pass |
|
||
| Valid JSON Schema | `jsonschema` `check_schema` against draft 2020-12 | pass |
|
||
| Accepts/rejects findings correctly | 2 valid + 3 invalid findings (missing `scanner`, unknown severity, `line: 0`) | 5/5 as intended |
|
||
| SARIF profile reproduces the source | re-implemented the mapping **from the commons JSON alone** and diffed `JSON.stringify` against the real `toSARIF` over 10 envelope shapes: empty, missing `scanners`, empty `scanners`, scanner with no findings, all five severities plus an unknown and an `undefined` one, five slug edge cases (double space, tab, newline, leading/trailing space, mixed case), a rule-id collision, all seven optional-field combinations, two scanners, and an explicit `version` argument | 10/10 identical, 0 differences |
|
||
| The three `known_lossiness` claims are true | executed each against the real formatter | all three confirmed, **and one earlier claim corrected**: punctuation does *not* collapse — the slug lowercases and collapses whitespace only, so `Zero-width carrier` and `Zero-width carrier!` remain distinct ids. The wrong claim was published in the first draft of this file and fixed before commit. |
|
||
|
||
Not verified, and recorded in the file as open: the finding **producer** was not supplied, so
|
||
the property list is a lower bound; `scanner` and `severity` are required by design rather
|
||
than by evidence; and the JSONL profile is left explicitly `unspecified` rather than
|
||
invented, because "one finding per line" is inference.
|
||
|
||
### `lexicon/injection-lexicon.json` — extracted 2026-08-09
|
||
|
||
Source: `llm-security/scanners/lib/injection-patterns.mjs`, supplied as operator dump 2/2
|
||
through the local coord mailbox. No commit hash accompanied it, so provenance is recorded
|
||
as `unknown` rather than guessed.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed, `version` present, LF, trailing newline, no raw invisible code points | `python3 -m json.tool` + a byte scan for U+200B/200C/200D/FEFF/00AD and the Tag block | pass, 0 raw invisible code points |
|
||
| Pattern text and flags reproduce the source | rebuilt all four arrays **from the commons JSON alone** (`new RegExp(p.pattern, p.flags ?? '')`) and diffed label, `.source` and `.flags` against the imported dump module | 83/83 compared, 0 differences; 81/83 byte-identical, 2 declared-normalised |
|
||
| Flags were read mechanically, not by eye | extracted from each literal via `.flags` | critical 15×`i` / 3×`m` / 3 none, high 32×`i`, medium 20×`i` / 2 none, hybrid 8×`i` |
|
||
| Every pattern compiles in both runtimes | `new RegExp(src, flags)` and again with `u` in Node; `re.compile` with the equivalent `re.I`/`re.M` in Python | 83/83 in all three modes, 0 failures |
|
||
| The 2 normalised patterns behave identically | differential match-set comparison (offsets + matched text) against the source objects, bare and under `u`, over a 208-input adversarial corpus: every class member, the near-misses excluded from each class (U+00AD, U+2060, U+180E, Cyrillic х, the uppercase set, Greek look-alikes), run boundaries, repeats, empty input | 832 comparisons, 0 differences |
|
||
| Class membership was counted, not assumed | enumerated the code points inside each character class directly from the dump bytes | zero-width class = 4 (U+200B, U+200C, U+200D, U+FEFF — **not** U+00AD); Cyrillic class = 7 (U+0430, U+0435, U+043E, U+0440, U+0441, U+0456, U+0443) |
|
||
| `\/` is portable, not a defect | 9 patterns carry the redundant escape a JS regex literal requires; compiled in Node bare, Node `u`, and Python `re` | accepted by all three — kept byte-identical, recorded as a translation note for engines that reject unknown escapes |
|
||
|
||
Not verified, and not claimed: that the dump matches the module it was transcribed from.
|
||
Every check above proves this JSON agrees with **the dump**; dump-to-module fidelity is
|
||
`llm-security`'s assertion, reproducible only in a session with read access to that
|
||
repository. The severity the engine assigns to `HYBRID_PATTERNS` was not supplied and is
|
||
left `null` rather than inferred from its three sibling arrays.
|
||
|
||
### `codepoints/carriers.json` — extracted 2026-08-09
|
||
|
||
Source: `llm-security/scanners/unicode-scanner.mjs` (charset constants) and
|
||
`llm-security/scanners/lib/string-utils.mjs` (`HOMOGLYPH_MAP`), supplied as operator dump
|
||
2/2 through the local coord mailbox. No commit hash accompanied it.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed, `version` present, no raw invisible code points | `python3 -m json.tool` + byte scan for zero-width, BIDI and Tag-block characters | pass, 0 raw invisible code points |
|
||
| Five of the six tables reproduce the source constants | rebuilt each **from the commons JSON alone** (`parseInt(codepoint.slice(2), 16)`) and diffed against the imported dump module | `ZERO_WIDTH_CHARS` 5/5, `BIDI_CHARS` 9/9, `CYRILLIC_CONFUSABLES` 13/13, tag start/end — 0 differences |
|
||
| The homoglyph map reproduces the source, including order | rebuilt the object from the entries array and compared keys, values and the whole object | 28/28 keys, values and insertion order identical |
|
||
| The map folds identically | applied NFKC + lookup with both the rebuilt and the source table over 12 inputs (Cyrillic and Greek injection spellings, Norwegian and German orthography, empty) | 0 differences |
|
||
| The exclusion rationale in the source comment is true | checked whether any of `帿ŨÆäöüßéèêñç` is a key | 0 touched — ordinary Norwegian and German orthography is not folded |
|
||
| Convenience `char` fields agree with their own `codepoint` field | `String.fromCodePoint` round-trip on every entry | 41/41, 0 mismatches |
|
||
| Character names are not from recollection | resolved every name through Python `unicodedata` against the Unicode character database | all resolved |
|
||
| Table sizes were counted, not quoted | counted from the imported constants | homoglyph map holds **28** entries, not the "~25" the dump's own comment estimates; the counted number is the one recorded |
|
||
|
||
Not verified, and marked `verified: false` **in the file itself**: the two Supplementary
|
||
Private Use Area ranges. They arrived as a source comment with no constant behind them, so
|
||
unlike the other five tables there was nothing to import and diff. That asymmetry is
|
||
recorded per-table rather than averaged into a single file-level verdict.
|
||
|
||
Recorded and deliberately **not** reconciled, in `cross_table_notes`: the repository now
|
||
holds three overlapping Cyrillic sets and two overlapping zero-width sets, and none agree
|
||
exactly. The zero-width carrier table includes U+00AD while the lexicon's pattern class does
|
||
not; the lexicon's class contains U+0456 which `CYRILLIC_CONFUSABLES` lacks, and
|
||
`CYRILLIC_CONFUSABLES` contains U+0445 which the class lacks.
|
||
|
||
**Two claims that stood here until 2026-08-09 were wrong, and both are corrected rather
|
||
than quietly dropped.**
|
||
|
||
The first was a count: this paragraph said *six of the confusables have no entry in the fold
|
||
map*. It is **zero**. All 13 members of `CYRILLIC_CONFUSABLES` are keys in the 28-entry fold
|
||
map, and `codepoints/carriers.json` has said so correctly since it was written —
|
||
`cross_table_notes.cyrillic.in_confusables_but_not_foldable` is an empty array. The prose,
|
||
not the data, carried the error. The number six most likely came from the comparison run the
|
||
other way: the fold map holds **five** Cyrillic keys that are not confusables (U+0423,
|
||
U+0455, U+0456, U+0458, U+04CF), which `llm-security` confirms independently. A count
|
||
asserted in prose beside a data file that contradicts it is the cheapest kind of defect to
|
||
catch and the easiest to propagate, so the check is now: read the array, do not restate it.
|
||
|
||
The second was a characterisation. This paragraph said the sets are *deliberately distinct*.
|
||
`llm-security` states plainly that the U+0456 / U+0445 divergence **is drift, not design**.
|
||
They are not harmonising it now — that would be a behaviour change inside a
|
||
behaviour-preservation release — and have logged it as a v8.x task on their side. So the
|
||
divergence stays in the data exactly as it is, but it is recorded here as a known defect
|
||
awaiting a fix upstream, not as an intentional distinction. Those two readings imply opposite
|
||
things about whether a future consumer should expect the sets to converge.
|
||
|
||
### `signatures/secret-egress.json` — extracted 2026-08-09
|
||
|
||
Source: `llm-security/hooks/scripts/pre-edit-secrets.mjs` (`SECRET_PATTERNS`), supplied as
|
||
operator dump 2/2 through the local coord mailbox. No commit hash accompanied it.
|
||
|
||
**The seed-source row above was wrong and has been corrected.** It named
|
||
`knowledge/secrets-patterns.md`; the dump named `hooks/scripts/pre-edit-secrets.mjs` and
|
||
stated that the two are different tables — the second is PCRE-flavoured and agent-consumed
|
||
and stays where it is. Recording a source file that was never delivered is the same defect
|
||
class as the lossiness claim corrected in `finding.schema.json`, so it is corrected here in
|
||
the same commit as the file it describes.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed, `version` present, LF, trailing newline | `python3 -m json.tool` + byte scan | pass |
|
||
| Pattern text and flags reproduce the source | rebuilt the table **from the commons JSON alone**, sorted by the declared `order`, and diffed name, `.source` and `.flags` against the imported dump module | 18/18, 0 differences, **18/18 byte-identical** — no normalisation needed |
|
||
| Every pattern compiles in both runtimes | `new RegExp` bare and under `u` in Node; `re.compile` with `re.I` where declared in Python | 18/18 in all three modes, 0 failures |
|
||
| The ordering contract holds, and is not decorative | reproduced first-match labelling from the commons order for a Bearer header containing a JWT and for a bare JWT, against the source table | both labels identical to source: header case → `Authorization header with token`, bare case → `JWT (three-part token)` |
|
||
| Reordering is detectable, not silent | ran the same Bearer input through a reversed table | label changes to `JWT (three-part token)` — order is load-bearing, which is why every entry carries an explicit `order` field |
|
||
| `order` is contiguous | compared to `range(18)` | 0–17, no gaps |
|
||
|
||
Not verified, and not claimed: that the dump matches the module. Not supplied, and therefore
|
||
not invented: any severity or per-entry disposition — the source table carries a name and a
|
||
pattern and nothing else. Out of scope by the dump's own statement: the runtime
|
||
policy-injected custom patterns (entries 19+). A consumer matching only this table matches
|
||
**less** than the seed hook does when a policy is loaded.
|
||
|
||
### `mapping/owasp-map.json` — extracted 2026-08-09
|
||
|
||
Source: `llm-security/scanners/lib/severity.mjs` (`OWASP_MAP`, `OWASP_AGENTIC_MAP`,
|
||
`OWASP_SKILLS_MAP`, `OWASP_MCP_MAP`), supplied as operator dump 2/2 through the local coord
|
||
mailbox. No commit hash accompanied it.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed, `version` present | `python3 -m json.tool` | pass |
|
||
| All four maps reproduce the source | rebuilt each exported object **from the commons JSON alone**, iterating the declared prefix list, and compared `JSON.stringify` against the imported dump module | 4/4 identical — keys, order, values and empty arrays |
|
||
| Key sets are identical across taxonomies, and counted | compared key order across all four in both the commons file and the source | one shared order, **16** prefixes, in all eight objects |
|
||
| Empty arrays survive as arrays | type- and length-checked every prefix the source maps to nothing | `agentic` TRG/AST, `skills` WFL/SIG, `mcp` WFL/TRG/SIG/AST — all still `[]`, none dropped or nulled |
|
||
| Code prefixes are homogeneous per taxonomy | extracted the alphabetic prefix of every code with a regex and asserted one per map | LLM, ASI, AST, MCP — no mixed map |
|
||
|
||
Recorded as an **open question in the file**, because it is the one thing a consumer can get
|
||
wrong while matching this map exactly: the dump does not state which *edition* of each
|
||
taxonomy the codes belong to. That matters and is not pedantry — OWASP's Top 10 for LLM
|
||
Applications was renumbered between editions, and in the 2025 edition `LLM06` is Excessive
|
||
Agency, with earlier standalone entries consolidated into other numbers and System Prompt
|
||
Leakage and Vector and Embedding Weaknesses added as `LLM07` and `LLM08`
|
||
([OWASP project page](https://owasp.org/www-project-top-10-for-large-language-model-applications/),
|
||
[2025 edition summary](https://www.gravitee.io/blog/owasp-top-10-for-llm-applications-2025-a-practical-guide)).
|
||
A bare `LLM06` therefore does not identify a risk. Two runtimes can reproduce this map
|
||
perfectly and still publish reports that disagree about what a finding means. `taxonomy_name`
|
||
is left `null` rather than guessed, and the question is reported to `llm-security`.
|
||
|
||
Not supplied, and therefore not invented: what each scanner prefix means. `UNI`, `ENT`, `PRM`
|
||
and the rest are reproduced as opaque keys. Deliberately not inherited: the dump notes that a
|
||
fallback map inside one of `llm-security`'s own agent definitions covers only part of the
|
||
prefix set — that is drift in the consumer, and the complete 16-prefix set is what moves here.
|
||
|
||
### `calibration/calibration.json` — transcribed 2026-08-09
|
||
|
||
Source: `llm-security/scanners/lib/severity.mjs`, delivered as a **prose summary** inside
|
||
operator dump 2/2 — not as source code.
|
||
|
||
**This file is the exception in this repository, and it is marked as such in its own
|
||
`verification` block.** Every other data file here was rebuilt from its commons JSON and
|
||
diffed against an imported module. There was nothing to import here: the constants arrived
|
||
as human-written prose describing the module. The differential check was therefore not run
|
||
— not "passed with caveats", not run — and the file records `verified: false` with the
|
||
specific checks that were skipped.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| JSON well-formed, `version` present | `python3 -m json.tool` | pass |
|
||
| Risk bands are contiguous and non-overlapping | compared each band's `max + 1` to the next band's `min` across 0–100 | contiguous, no gap, no shared value |
|
||
| Band boundaries agree with the verdict thresholds | compared the BLOCK and WARNING score triggers to the band lower bounds | BLOCK 65 = Critical band min; WARNING 15 = Medium band min |
|
||
| `reachable_minimum` is arithmetic, not a new claim | recomputed `base + min(cap, log2(2) * mult)` independently for all four tiers | 80 / 48 / 20 / 4 — exact, because `log2(2)` is exactly 1 |
|
||
| Rebuild-from-commons and diff against source | **not run** — no importable source | — |
|
||
| Differential scoring over a corpus | **not run** — the formulas are engine and were not supplied in runnable form | — |
|
||
|
||
Recorded in the file as `not_supplied`, so the absence is visible rather than inferred: the
|
||
README, this plan's seed-source table and this plan's layout block all described *this file*
|
||
as holding **entropy floors, scan caps and disposition ranks**. None of the three arrived in
|
||
any dump — searched across the entire dump message: `entropy` 0 occurrences, `disposition` 0,
|
||
`rank` 0, `floor` 0. All three descriptions have been corrected to name what arrived rather
|
||
than what was expected. (`signatures/active-content.json` does carry a Shannon-entropy floor,
|
||
extracted from the guard repository; that is a different file with a different source, and is
|
||
not what the calibration rows promised.)
|
||
|
||
The consequence is stated in the file and is worth repeating here, because it inverts this
|
||
repository's central rule: for every other file, a consumer that disagrees is wrong. For this
|
||
one, until the module is supplied in executable form, a disagreement is not automatically the
|
||
consumer's bug.
|
||
|
||
## Conformance sources — counted 2026-08-09
|
||
|
||
The guard's `coverage.py` is readable from a session here (the read boundary covers
|
||
`llm-security` only, not the guard repository), so the matrix was imported and counted rather
|
||
than estimated. `CORE_CASES` holds **134** cases: 128 `caught`, 6 `gap`.
|
||
|
||
Only part of it can become a static `input.txt` / `expected.json` pair, because only part of
|
||
it is text-in / findings-out:
|
||
|
||
| Probe kind | Count | Convertible? |
|
||
| --- | --- | --- |
|
||
| `_scan_case` — scan a payload, assert a label | 103 | yes |
|
||
| `_scan_case_okf` — scan an OKF document body | 2 | yes, if the fixture input may be a structured document rather than plain text |
|
||
| `_predicate_case` — assert an engine-internal predicate (disposition compounds, documented gaps) | 13 | no |
|
||
| `_raise_case` — assert the guard raises a typed error (contract asserters, OKF rejecters) | 16 | no |
|
||
|
||
The 29 non-convertible cases are not a coverage hole in this repository. They exercise a
|
||
runtime's **API surface** — that a Python call raises `OKFPathError`, that a disposition
|
||
engine composes two findings a particular way. This repository publishes data and fixtures and
|
||
explicitly does not own an API, so those rows belong to the guard's own suite, the same way
|
||
the `Planned` rows in the README are named rather than faked.
|
||
|
||
**Convertibility has a second condition, and it is the one easy to miss.** A case is a
|
||
candidate only if the probe is text-in / findings-out **and** the label it asserts maps to data
|
||
this repository actually publishes. A fixture asserting a label commons has no table for would
|
||
quietly extend the contract to cover engine behaviour the charter excludes — the corpus would
|
||
then be testing something this repository does not define.
|
||
|
||
| Group | Cases | Asserted label maps to commons data? |
|
||
| --- | --- | --- |
|
||
| `lexicon` | 83 | Yes — `lexicon/injection-lexicon.json`, now via the `id` field |
|
||
| `active` | 7 | Yes — `signatures/active-content.json` |
|
||
| `carrier` | 8 | Partly. `output:*-present` and `lexicon:unicode-tags-present` rest on the code-point tables in `codepoints/carriers.json`. The five `sanitize:*` rows assert **stripper** behaviour, which is engine. |
|
||
| `egress` | 1 | Yes — `signatures/secret-egress.json` |
|
||
| `entropy` | 2 | **No.** Shannon-entropy detection is the guard's `entropy.py`; commons publishes no entropy table. (`active-content.json` carries an opacity floor — a different thing, for a different pass.) |
|
||
| `decoded` | 2 | **No.** These assert a finding *after* a decode pass. The decode order is exactly what `spec/decode-pipeline.md` would fix, and it is Planned because it is engine. |
|
||
| `okf` (`_scan_case_okf`) | 2 | No — the OKF document model is not published here. |
|
||
|
||
So the buildable set is smaller than 105: roughly 83 + 7 + 3 + 1 ≈ 94. The rest are named here
|
||
so a later session does not rediscover them as a gap. *(Revised 2026-08-10: this sentence
|
||
ended "minus the 13 lexicon patterns whose expected behaviour is still disputed". Measurement
|
||
showed the dispute does not reach these inputs, so the 13 are in. What kept the built corpus
|
||
at 83 rather than 94 turned out to be a different constraint entirely — the 11 non-lexicon
|
||
cases have no ratified cross-runtime finding id. See* Conformance corpus *below.)*
|
||
|
||
**Blocking issue found while counting.** The guard and the Node table are two ports of one
|
||
source file and they do not agree: 64 of 83 patterns byte-identical, 6 differing only by
|
||
escaping, and **13 behaviourally divergent with concrete witness inputs** — in both
|
||
directions. They also disagree about pattern identity itself (`override:ignore-previous` vs
|
||
`override: ignore previous instructions`). Full measurement, method and limits:
|
||
[`lexicon-port-divergence.md`](lexicon-port-divergence.md). Nothing was changed in any data
|
||
file as a result — behaviour preservation holds, and the divergence is reported to the two
|
||
owning repositories instead.
|
||
|
||
**Half of it is now cleared (operator decision, 2026-08-09).**
|
||
`lexicon/injection-lexicon.json` 0.2.0 publishes a commons-owned `id` per pattern with an
|
||
`aliases` object per runtime, adopted verbatim from the guard's port rather than invented.
|
||
Findings can therefore be named. ~~What still blocks the corpus is the other half: the 13
|
||
divergent patterns have no agreed expected behaviour, and a fixture for them would have to
|
||
pick a winner.~~ **The other half cleared on 2026-08-10 without anyone picking a winner** —
|
||
the divergence is not reachable from any input the corpus contains. See *Conformance corpus*
|
||
below.
|
||
|
||
## Definition of done for v0.1.0
|
||
|
||
1. Repository initialized, Forgejo remote `open/llm-security-commons`, MIT, `STATE.md`
|
||
gitignored.
|
||
2. Every file in the layout above present and populated from verified seed data — or
|
||
explicitly and visibly stubbed where the source was unavailable.
|
||
|
||
**Status 2026-08-09, second revision: still unmet, and still why no tag has been cut —
|
||
but the remaining gap is two items, not three.** All **eight** JSON artefacts in the
|
||
layout now exist. `signatures/malware-signatures.json` was the last, added once
|
||
`knowledge/signatures.json` became readable at `b0de0ca` and reproduced from the file
|
||
rather than from a dump.
|
||
|
||
Seven of the eight are differentially verified against their source; three of those
|
||
(`lexicon`, `mapping/owasp-map`, `signatures/malware-signatures`) are verified against the
|
||
source **module** at a pinned commit rather than against a transcription, and
|
||
`schema/finding.schema.json` is closed against its producer. The eighth,
|
||
`calibration/calibration.json`, still carries `verified: false` — its source arrived as a
|
||
prose summary with no code behind it, so no differential check was possible, and the file
|
||
names the checks that were not run rather than attaching a caveat to a pass.
|
||
|
||
Two items remain absent, and only one of them still blocks:
|
||
|
||
- `spec/decode-pipeline.md` — needs the decode implementation. A normative spec inferred
|
||
from a data dump would carry a `Status: normative` marker and pass every convention
|
||
check in this document while asserting something nobody verified.
|
||
- ~~`conformance/`~~ — **built 2026-08-10, 83 cases.** See *Conformance corpus* below.
|
||
|
||
`spec/decode-pipeline.md` is absent rather than stubbed on purpose: a contentless
|
||
normative spec would *pass* every mechanical check in this document while making the
|
||
repository look more finished than it is. It is named as **Planned** in the README
|
||
instead, and is not linked, so nothing points at a file that does not exist.
|
||
|
||
### Conformance corpus — built 2026-08-10
|
||
|
||
The remaining blocker was the 13 divergent patterns, and it **dissolved under measurement
|
||
rather than under a decision**. The divergence was measured on witness inputs — padding past
|
||
256 characters, an interior `<`, an unclosed `<script>` — and the corpus payloads contain
|
||
none of those shapes. Run through both runtimes' public entry points, all 83 produce
|
||
identical lexicon finding sets, 13 of 13 divergent patterns included.
|
||
|
||
| Check | Method | Result |
|
||
| --- | --- | --- |
|
||
| Payload id set equals the commons id set | compared the 83 `_LEX_PAYLOADS` keys to the 83 lexicon ids | identical, both directions empty |
|
||
| Asserted id fires in both runtimes | `scanForInjection()` at `b0de0ca` and `scan_output(source=OUTPUT)` at `0bf0729`, labels mapped to ids through the lexicon's own `aliases` | **83/83** |
|
||
| Lexicon finding **sets** identical between runtimes | set equality per case, not membership | **83/83** |
|
||
| The 13 known-divergent patterns, on their own case input | same comparison, restricted to those 13 | **13/13 agree** |
|
||
| Severity agrees with what commons publishes | commons family `severity` vs the severity the guard emitted per finding | **83/83** |
|
||
| The corpus verifies from the corpus alone | a second harness that reads only the case directories, re-runs both runtimes on the bytes it finds, and checks every `expected.json` field including the digests | **83 cases, 0 failures** |
|
||
| The fixtures discriminate | deleted the middle third of each input and re-scanned; a fixture that still passes is not testing anything | 76/83 broke. The 7 survivors are the shortest payloads, where the mutation leaves the trigger intact — a weak mutation, not a weak fixture |
|
||
| Case ids are collision-free after the filesystem transform | `:` → `__` across all 83 | 83 unique; `__` occurs nowhere in the id space, so the transform is one-to-one |
|
||
|
||
Measured at the **public entry point of each runtime**, never at a rebuilt regex table. That
|
||
is the same discipline the retraction in [`lexicon-port-divergence.md`](lexicon-port-divergence.md)
|
||
was forced to adopt: a measurement taken at the wrong layer produces a number that describes
|
||
neither runtime.
|
||
|
||
**Scope, and why it is 83 and not 94.** The guard's matrix holds 105 convertible cases, of
|
||
which ~94 assert a label that maps to data this repository publishes. Only 83 shipped. The
|
||
lexicon is the sole table whose finding id space is **ratified by both runtimes**; the
|
||
carrier, active-content and secret-egress tables have no agreed cross-runtime finding names,
|
||
so writing fixtures for their 11 cases would mint a contract in the same stroke as the tag,
|
||
unilaterally. They are named in `conformance/manifest.json` under `scope_planned`.
|
||
|
||
The inputs were authored by one runtime — the guard's `coverage.py` — and are reproduced
|
||
verbatim. What makes them usable as a *cross-runtime* corpus is not their origin but the
|
||
measurement above, which ran them through the other runtime too. The asymmetry is recorded
|
||
in the manifest rather than averaged away.
|
||
|
||
Normative semantics live in [`spec/conformance-corpus.md`](../spec/conformance-corpus.md),
|
||
not here.
|
||
3. All JSON well-formed, every data file carrying `"version"`, every spec carrying
|
||
`Status: normative`.
|
||
4. Tagged `v0.1.0` and pushed.
|
||
5. A `coord` message sent to `llm-security` announcing that the repository and `v0.1.0`
|
||
exist, so Phase 5 step 3 (vendoring) can start from there.
|