docs: measure the lexicon port divergence, correct the case count
The guard's coverage.py turned out to be readable from a session here — the
read boundary covers llm-security only. Counting it instead of estimating
corrected two published numbers and surfaced the real conformance blocker.
Counted, not estimated: CORE_CASES holds 134 cases (128 caught, 6 gap), not
the "126 classes + 4 gaps-must-hold" the extraction plan claimed. 105 are
convertible to static input/expected; 29 assert a runtime API surface this
repository does not own, and are named as out of scope rather than faked.
Measured while counting, and the reason the corpus is still blocked: the
commons lexicon and the guard's are two ports of one source file, and they
disagree. 83/83 patterns correspond, 64 byte-identical, 6 differ only by
escaping and are proven equivalent by string identity, and 13 behave
differently — with a witness input for each and misses in both directions.
Cause is two different ReDoS mitigations of the same table: the Node side
bounds the run ({0,256}), the guard excludes the anchor character ([^><]).
Each has a recall cost the other does not.
No data file changed. Behaviour preservation is a v0.1.0 invariant, so the
divergence is reported to the owning repositories, not fixed here. The 8
hybrid severities the guard's port assigns are deliberately NOT copied in —
a second-hand port is not the producing module.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FaYqid3mejFmd9ZHsiHgp3
This commit is contained in:
parent
93f7bc024f
commit
22d3a78591
4 changed files with 193 additions and 6 deletions
139
docs/lexicon-port-divergence.md
Normal file
139
docs/lexicon-port-divergence.md
Normal file
|
|
@ -0,0 +1,139 @@
|
|||
# Lexicon port divergence — commons vs the Python guard
|
||||
|
||||
**Status: informative.** Nothing here is normative and nothing here changes a data file. It
|
||||
records a measured disagreement between two ports of one source table, so the decision can be
|
||||
taken where each table is tested. Under this repository's behaviour-preservation invariant,
|
||||
a divergence found here is **reported, not fixed**.
|
||||
|
||||
Produced 2026-08-09. Every number below came from a command; the scripts live in the session
|
||||
scratchpad rather than in this repository, because executable code here would breach the
|
||||
charter. They are reproducible from the method column.
|
||||
|
||||
## What was compared
|
||||
|
||||
| Side | Artefact | Version |
|
||||
| --- | --- | --- |
|
||||
| commons | [`lexicon/injection-lexicon.json`](../lexicon/injection-lexicon.json) | file `version` as committed |
|
||||
| guard | `llm-ingestion-pipeline-security` `src/llm_ingestion_guard/injection_lexicon.json` | lexicon `version` 1.0, repo v0.3.4, commit `0bf0729` |
|
||||
|
||||
Both are **ports of the same file**: `llm-security/scanners/lib/injection-patterns.mjs`. The
|
||||
guard's JSON says so in its own `note` field — *"Injection lexicon ported from llm-security
|
||||
injection-patterns.mjs. Single source of truth."* Commons extracted the same table from an
|
||||
operator dump of that module.
|
||||
|
||||
That is what makes the comparison worth running. These are not two different detectors that
|
||||
happen to overlap; they are two transcriptions of one table, and where they disagree, they
|
||||
disagree about what the same source says.
|
||||
|
||||
## Result
|
||||
|
||||
| Measure | Method | Result |
|
||||
| --- | --- | --- |
|
||||
| Pattern count, both sides | count entries | 83 and 83 |
|
||||
| Label correspondence | match commons `label` to guard `desc`, em-dash normalised to hyphen | **83/83** |
|
||||
| Regex source byte-identical | string compare | **64/83** |
|
||||
| Differing regex text | string compare | 19 |
|
||||
| — of those, provably equivalent | unescape commons' JS-isms (`\/` → `/`, `\uXXXX` → literal) and compare for string identity | **6/6 identical** |
|
||||
| — of those, behaviourally divergent | differential match-set comparison (offsets + matched text), targeted corpus per pattern family | **13**, each with a concrete witness input |
|
||||
| Total input comparisons | count | 401 |
|
||||
| Flags | compare declared flags | **0 differences** |
|
||||
| Severity / family | commons family vs guard `severity` | **8 differences** (all `hybrid`) |
|
||||
|
||||
64 identical + 6 escaping-only + 13 divergent = 83.
|
||||
|
||||
The 6 escaping-only pairs are equivalent by **string identity after unescaping**, not merely
|
||||
by absence of a counter-example: `<\/system>` vs `</system>`, `text\/html` vs `text/html`,
|
||||
`(?:https?:\/\/…)` vs `(?:https?://…)`, `(?:~\/)?` vs `(?:~/)?`, and two `\uXXXX`-vs-literal
|
||||
character classes. Commons writes code points as escapes by convention; the guard's `note`
|
||||
records the same choice for non-ASCII.
|
||||
|
||||
## The 13, and which side misses what
|
||||
|
||||
Divergence is **not one-directional**. Each side has inputs the other's pattern does not match.
|
||||
|
||||
| Pattern | Direction | Witness input |
|
||||
| --- | --- | --- |
|
||||
| `html-obfuscation: display:none element with content` | commons misses; span differs | attribute run padded past 256 chars |
|
||||
| `html-obfuscation: visibility:hidden element` | span only | `<div data-x="<" style="visibility:hidden">x</div>` |
|
||||
| `html-obfuscation: off-screen positioned element` | span only | same shape, `position:absolute` |
|
||||
| `html-obfuscation: zero font-size element` | span only | same shape, `font-size:0` |
|
||||
| `html-obfuscation: zero opacity element` | span only | same shape, `opacity:0` |
|
||||
| `html-obfuscation: zero-size overflow-hidden element` | span only | same shape, `height:0;overflow:hidden` |
|
||||
| `html-obfuscation: injection in aria-label attribute` | **commons misses** | `<a aria-label="` + 300 × `a` + `ignore">x</a>` |
|
||||
| `markdown: suspicious instruction in link-reference comment` | **guard misses** | `[//]: # (a (b) ignore)` |
|
||||
| `markdown: injection payload in link anchor text` | **guard misses** | `[click ignore this](http://e.com/a(b))` |
|
||||
| `sub-agent: spawn instruction with dangerous capability keywords` | **guard misses** | `create an agent that` + 20 filler words + `execute` |
|
||||
| `sub-agent: delegation with safety bypass instruction` | **guard misses** | `delegate to a new agent` + 200 chars + `bypass` |
|
||||
| `hybrid-xss: <script> tag in content (agent context XSS)` | **commons misses**; span differs | `<script>alert(1)` (unclosed), `<script src=x.js>` |
|
||||
| `hybrid-xss: iframe with executable src (agent context XSS)` | **guard misses** | `<iframe data-x="<" src="javascript:alert(1)">` |
|
||||
|
||||
"Span only" means both sides produce a match on the same input but over different extents —
|
||||
the guard's match starts at an interior `<`. Whether that matters depends on whether a
|
||||
consumer reports offsets or evidence text; it does not change whether a finding is raised.
|
||||
|
||||
The commons-side misses were confirmed in a real JS engine (Node v25.8.2, `RegExp` built from
|
||||
the committed JSON), not only in the Python harness used for the differential.
|
||||
|
||||
## Why they diverge: two different ReDoS mitigations of one table
|
||||
|
||||
This is not drift, and framing it as a bug in either repository would be wrong.
|
||||
|
||||
Both ports have been hardened against catastrophic backtracking, by **different strategies**:
|
||||
|
||||
- **The Node side bounds the run.** `[^"]{0,256}`, `[^>]{1,256}`. Cost: an attacker who pads
|
||||
the attribute past 256 characters falls out of the pattern.
|
||||
- **The guard excludes the anchor character.** `[^><]`, `[^\]\[]`, `[^)(]`. Cost: content that
|
||||
legitimately contains that character stops matching.
|
||||
|
||||
The guard's campaign is documented in its own CHANGELOG (v0.3.2 and v0.3.4: *"exclude the
|
||||
character that opens the pattern's own anchor"*, a sweep over *"150 patterns across 11
|
||||
tables"*), including its measured recall cost — it names exactly the two exceptions this
|
||||
comparison rediscovered: *"URLs containing a literal `(` inside a markdown link target and
|
||||
comment bodies containing a literal `(` before the keyword."* The bounded `{0,12}` / `{0,120}`
|
||||
sub-agent quantifiers produce misses that the guard's exception list does not mention.
|
||||
|
||||
Neither strategy is free, and neither is obviously right. That is the decision the two owning
|
||||
repositories have to take, and it is not commons' to take for them.
|
||||
|
||||
## Severity: the 8 hybrid patterns
|
||||
|
||||
Commons records the `hybrid` family with `severity: null` and a note that the seed dump did
|
||||
not supply it, so a consumer **MUST NOT** assume one. The guard's port assigns all eight
|
||||
`high`.
|
||||
|
||||
**This has deliberately not been copied into commons.** The guard's port is a second-hand
|
||||
transcription, not the producing module; adopting its value would convert a documented gap
|
||||
into an unverified claim, which is the defect class this repository's changelog already
|
||||
records twice. It is reported instead, as a strong hint that the Node source assigns `high`,
|
||||
for `llm-security` to confirm from the module.
|
||||
|
||||
## What this does not show
|
||||
|
||||
- **Not that either runtime misses an attack.** All witness payloads were re-tested against
|
||||
every pattern table commons holds (111 compiled rules: 83 lexicon, 28 from
|
||||
`active-content.json` and `secret-egress.json`). All of them still produce a finding, via
|
||||
`active-content: constructs.raw-html`. The divergence is in the **finding set** — which
|
||||
labels are raised — not in whether anything is raised at all. That still matters, because a
|
||||
finding set is exactly what a `conformance/expected.json` would encode.
|
||||
- **Not that `llm-security` behaves as described here.** Only its extracted pattern table was
|
||||
available. Whether the Node engine runs an active-content table that compensates cannot be
|
||||
determined from this repository.
|
||||
- **Not dump-to-module fidelity.** Every check above proves the two *ports* agree or disagree.
|
||||
That either matches `injection-patterns.mjs` remains `llm-security`'s assertion,
|
||||
reproducible only in a session with read access to it.
|
||||
- **Not exhaustive.** The corpus is targeted per pattern family, 401 comparisons. Absence of a
|
||||
witness proves nothing except for the 6 escaping-only pairs, which are settled by string
|
||||
identity rather than by the corpus.
|
||||
|
||||
## Consequence for `conformance/`
|
||||
|
||||
An `expected.json` names findings. Naming a finding needs a stable id, and the two runtimes do
|
||||
not have one: the same pattern is `override:ignore-previous` in the guard and
|
||||
`override: ignore previous instructions` in the Node table. The guard's JSON happens to carry
|
||||
both — `id` and `desc` — which is evidence that a commons-owned id is achievable rather than
|
||||
speculative.
|
||||
|
||||
So the id question is a **prerequisite** for the corpus, not a parallel task: until commons
|
||||
owns a pattern id both ports map to, no fixture can be written, including for the 64 patterns
|
||||
that are byte-identical. And for the 13 divergent patterns a fixture cannot be authored at all
|
||||
without first deciding whose recall cost is the contract.
|
||||
Loading…
Add table
Add a link
Reference in a new issue