llm-security/docs/lexicon-port-divergence.md
Kjell Tore Guttormsen a640f43d73 Squashed 'scanners/commons/' content from commit 0ffee85
git-subtree-dir: scanners/commons
git-subtree-split: 0ffee85a4b83b3661185488c06ed9a9994c11412
2026-08-10 20:40:16 +02:00

271 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Lexicon port divergence — commons vs the Python guard
**Status: informative.** Nothing here is normative and nothing here changes a data file. It
records a measured disagreement between two ports of one source table, so the decision can be
taken where each table is tested. Under this repository's behaviour-preservation invariant,
a divergence found here is **reported, not fixed**.
Produced 2026-08-09. Every number below came from a command; the scripts live in the session
scratchpad rather than in this repository, because executable code here would breach the
charter. They are reproducible from the method column.
**Revised the same day, after `llm-security`'s source became readable and both runtimes
replied.** Four things changed, and three of them are corrections to this file rather than
new results:
1. The claim that **neither runtime misses an attack** is **retracted**. It does. See
*What this does not show* — the measurement behind that claim unioned pattern tables
belonging to two different runtimes and read the result as a statement about each.
2. One of the 13 divergences does not reach report level, so **12** is the number that
changes what a report says. The 13 still blocks `conformance/`.
3. The `hybrid` **severity is resolved** to `high` — the reported hint was correct, and the
citation behind it was not.
4. The **pattern id space is ratified** by both runtimes.
Corrections are marked in place rather than edited away, because a reader who saw the first
version needs to know which sentence moved.
## What was compared
| Side | Artefact | Version |
| --- | --- | --- |
| commons | [`lexicon/injection-lexicon.json`](../lexicon/injection-lexicon.json) | file `version` as committed |
| guard | `llm-ingestion-pipeline-security` `src/llm_ingestion_guard/injection_lexicon.json` | lexicon `version` 1.0, repo v0.3.4, commit `0bf0729` |
Both are **ports of the same file**: `llm-security/scanners/lib/injection-patterns.mjs`. The
guard's JSON says so in its own `note` field — *"Injection lexicon ported from llm-security
injection-patterns.mjs. Single source of truth."* Commons extracted the same table from an
operator dump of that module.
That is what makes the comparison worth running. These are not two different detectors that
happen to overlap; they are two transcriptions of one table, and where they disagree, they
disagree about what the same source says.
## Result
| Measure | Method | Result |
| --- | --- | --- |
| Pattern count, both sides | count entries | 83 and 83 |
| Label correspondence | match commons `label` to guard `desc`, em-dash normalised to hyphen | **83/83** |
| Regex source byte-identical | string compare | **64/83** |
| Differing regex text | string compare | 19 |
| — of those, provably equivalent | unescape commons' JS-isms (`\/``/`, `\uXXXX` → literal) and compare for string identity | **6/6 identical** |
| — of those, behaviourally divergent | differential match-set comparison (offsets + matched text), targeted corpus per pattern family | **13**, each with a concrete witness input |
| — of those, divergent at REPORT level | re-check whether a sibling pattern raises an equivalent finding on the same witness | **12** — one of the 13 is a label-set difference only |
| Total input comparisons | count | 401 |
| Flags | compare declared flags | **0 differences** |
| Severity / family | commons family vs guard `severity` | **8 differences** (all `hybrid`) — resolved, see *Severity* below |
64 identical + 6 escaping-only + 13 divergent = 83.
**Read the 13 and the 12 as answering different questions.** Thirteen patterns produce
different label sets. Twelve of those change what a report would say. The gap is the
`hybrid-xss: iframe with executable src` row: the guard's version of that one pattern misses
the witness, but its `hybrid-xss: javascript: URI scheme` pattern fires on the same input at
the same severity and the same OWASP anchor, so a reader of the guard's report still sees a
`high` / `LLM01` finding on that payload. Measured, not reasoned: the guard's table matched
`hybrid-xss:javascript-uri` (high, LLM01) and nothing else; the Node engine matched both
`hybrid-xss: javascript: URI scheme` and `hybrid-xss: iframe with executable src`. A
`conformance/expected.json` would have to encode the difference, because a fixture names
labels — which is why the 13 stood as the number that blocked the corpus. *(Superseded
2026-08-10: it blocks a fixture written over a **witness** input, and the corpus contains
none. See* Consequence for `conformance/` *below.)*
The 6 escaping-only pairs are equivalent by **string identity after unescaping**, not merely
by absence of a counter-example: `<\/system>` vs `</system>`, `text\/html` vs `text/html`,
`(?:https?:\/\/…)` vs `(?:https?://…)`, `(?:~\/)?` vs `(?:~/)?`, and two `\uXXXX`-vs-literal
character classes. Commons writes code points as escapes by convention; the guard's `note`
records the same choice for non-ASCII.
## The 13, and which side misses what
Divergence is **not one-directional**. Each side has inputs the other's pattern does not match.
| Pattern | Direction | Witness input |
| --- | --- | --- |
| `html-obfuscation: display:none element with content` | commons misses; span differs | attribute run padded past 256 chars |
| `html-obfuscation: visibility:hidden element` | span only | `<div data-x="<" style="visibility:hidden">x</div>` |
| `html-obfuscation: off-screen positioned element` | span only | same shape, `position:absolute` |
| `html-obfuscation: zero font-size element` | span only | same shape, `font-size:0` |
| `html-obfuscation: zero opacity element` | span only | same shape, `opacity:0` |
| `html-obfuscation: zero-size overflow-hidden element` | span only | same shape, `height:0;overflow:hidden` |
| `html-obfuscation: injection in aria-label attribute` | **commons misses** | `<a aria-label="` + 300 × `a` + `ignore">x</a>` |
| `markdown: suspicious instruction in link-reference comment` | **guard misses** | `[//]: # (a (b) ignore)` |
| `markdown: injection payload in link anchor text` | **guard misses** | `[click ignore this](http://e.com/a(b))` |
| `sub-agent: spawn instruction with dangerous capability keywords` | **guard misses** | `create an agent that` + 20 filler words + `execute` |
| `sub-agent: delegation with safety bypass instruction` | **guard misses** | `delegate to a new agent` + 200 chars + `bypass` |
| `hybrid-xss: <script> tag in content (agent context XSS)` | **commons misses**; span differs | `<script>alert(1)` (unclosed), `<script src=x.js>` |
| `hybrid-xss: iframe with executable src (agent context XSS)` | **guard misses the label, not the attack** | `<iframe data-x="<" src="javascript:alert(1)">` |
The last row is the one of the 13 that does not reach report level. The guard's
`hybrid-xss:javascript-uri` (`javascript\s*:`, high, LLM01) matches that witness, so the
payload is still flagged at the same severity and anchor; only the label set differs — one
finding instead of two. The remaining 12 rows change what a report says.
"Span only" means both sides produce a match on the same input but over different extents —
the guard's match starts at an interior `<`. Whether that matters depends on whether a
consumer reports offsets or evidence text; it does not change whether a finding is raised.
The commons-side misses were confirmed in a real JS engine (Node v25.8.2, `RegExp` built from
the committed JSON), not only in the Python harness used for the differential.
## Why they diverge: two different ReDoS mitigations of one table
This is not drift, and framing it as a bug in either repository would be wrong.
Both ports have been hardened against catastrophic backtracking, by **different strategies**:
- **The Node side bounds the run.** `[^"]{0,256}`, `[^>]{1,256}`. Cost: an attacker who pads
the attribute past 256 characters falls out of the pattern.
- **The guard excludes the anchor character.** `[^><]`, `[^\]\[]`, `[^)(]`. Cost: content that
legitimately contains that character stops matching.
**Every divergence on the guard's side is documented at source, and traceable to the commit
that introduced it.** An earlier draft of this file claimed the two sub-agent bounds were not;
that was wrong, and it was wrong because the search behind it never looked outside the
CHANGELOG. Both mechanisms are named in `lexicon.py`'s own module docstring:
> **Bounded token gaps** — the two sub-agent patterns whose seed form nested `.*?` are ported
> with `(?:\S+\s+){0,N}?`.
>
> **Anchor exclusion** — […] Measured across all 83 patterns arm by arm, two markdown patterns
> had this defect; both now exclude the anchor character from the run.
`git log -S` separates the two: the eight `[^><]` patterns (six html-obfuscation, two
hybrid-xss) arrived with `cff0437`, *"fix(output): 19 quadratic regex runs on the output path"*
— the v0.3.2 campaign, whose CHANGELOG describes exactly this remedy (*"exclude the character
that opens the pattern's own anchor (`[` for markdown, `<` for tags)"*) across a sweep of
*"150 patterns across 11 tables"*. The `{0,12}` / `{0,120}` sub-agent bounds are older still:
they arrived with the original port commit `f397cd9`, so they were never a divergence
introduced later — they are how that table was transcribed in the first place.
The v0.3.4 entry also states the measured recall cost, naming precisely the two exceptions this
comparison rediscovered: *"URLs containing a literal `(` inside a markdown link target and
comment bodies containing a literal `(` before the keyword."*
Worth recording, because it anticipates the criticism the Node side invites: the guard
considered bounding those runs and **rejected it**, on the grounds that *"the content is
attacker-controlled, so padding past a bound would be a one-line bypass."* That is the same
objection the `{0,256}` witness above demonstrates against the Node table. The two projects
reached opposite conclusions from the same reasoning, which is the substance of the
disagreement — not an oversight on either side.
Neither strategy is free, and neither is obviously right. That is the decision the two owning
repositories have to take, and it is not commons' to take for them.
## Severity: the 8 hybrid patterns
**Resolved 2026-08-09. The two sides never disagreed; only the evidence did.**
Commons recorded the `hybrid` family with `severity: null` and a note that the seed dump did
not supply it, so a consumer **MUST NOT** assume one. The guard's port assigned all eight
`high`. Copying the guard's value would have converted a documented gap into an unverified
claim, so it was reported instead — and the report was right: the value **is** `high`,
confirmed at the module, and `lexicon/injection-lexicon.json` 0.5.0 now carries it. The eight
differences in the table above are closed.
The part worth keeping is where the value lives. It is not a field. The engine assigns it by
pushing `HYBRID_PATTERNS` matches straight into the `high` bucket at
`injection-patterns.mjs:274-281`. The guard's port cites `severity.mjs` — a file that
contains **no injection-family severity at all**. So the guard held the right value behind a
citation that leads nowhere, and a reviewer following that citation to check the number would
have found nothing and drawn no conclusion. Refusing to copy it was the right call for a
reason better than the one given at the time: not merely that a port is second-hand, but that
this particular port could not have read what it claimed to.
## What this does not show
- ~~**Not that either runtime misses an attack.**~~ **Retracted 2026-08-09. It does.** This
bullet claimed that every witness payload still produced a finding via
`active-content: constructs.raw-html`, so no attack went unflagged. The measurement behind
it was wrong in method, not in arithmetic: the payloads were run against the **union of
every pattern table this repository holds** — 111 rules across the lexicon,
`active-content.json` and `secret-egress.json` — and the rescuing hit came from
`active-content.json`. That table is the **Python guard's**. `llm-security` has no
active-content table at all. A union of commons tables is not any single runtime's
coverage, and treating it as one turned two runtimes' combined reach into a claim about
each of them.
Measured properly, through `llm-security`'s own entry point `scanForInjection()` — the
whole engine, with normalisation, homoglyph folding, the rot13 variant and all four pattern
arrays, at `b0de0ca`:
| Witness | `scanForInjection()` result |
| --- | --- |
| `<script>alert(1)` (unclosed) | `found: false` — no finding at all |
| `<script src=x.js>` | `found: false` — no finding at all |
| `<a aria-label="` + 300 × `a` + `ignore">` | `found: false` — no finding at all |
Controls in the same run behave as expected: `<script>alert(1)</script>` returns `high`
(hybrid-xss), and the short aria-label variant returns `critical`. So the `{0,256}` window
is a real evasion window and the `<script>` pattern really does require a closing tag.
`llm-security` reached the same three results independently and attributes the cause to
their own v7.8.3 #24 ReDoS hardening, which traded recall for boundedness without seeing
the window. Three confirmed recall holes, logged there as a v8.x task.
What survives from the original bullet is only this: the divergence is *also* in the
finding set, which is what a `conformance/expected.json` encodes.
- **Not that the guard misses an attack.** The guard's side of the 13 was re-checked the same
way, and the one row that looked like a miss (`hybrid-xss: iframe with executable src`) is
covered by a sibling pattern at the same severity and anchor. Its remaining divergences are
label-set and span differences.
- **Not dump-to-module fidelity.** *Superseded 2026-08-09.* Every check above proves the two
*ports* agree or disagree. Commons' side is now settled separately: the lexicon is verified
byte-identical to `injection-patterns.mjs` at `b0de0ca`, 83/83, which is recorded in the
data file rather than here. The guard's fidelity to the module remains its own to
establish.
- **Not exhaustive.** The corpus is targeted per pattern family, 401 comparisons. Absence of a
witness proves nothing except for the 6 escaping-only pairs, which are settled by string
identity rather than by the corpus.
## Consequence for `conformance/`
An `expected.json` names findings. Naming a finding needs a stable id, and the two runtimes do
not have one: the same pattern is `override:ignore-previous` in the guard and
`override: ignore previous instructions` in the Node table. The guard's JSON happens to carry
both — `id` and `desc` — which is evidence that a commons-owned id is achievable rather than
speculative.
So the id question is a **prerequisite** for the corpus, not a parallel task: until commons
owns a pattern id both ports map to, no fixture can be written, including for the 64 patterns
that are byte-identical. And for the 13 divergent patterns a fixture cannot be authored at all
without first deciding whose recall cost is the contract.
**Resolved for the first half, 2026-08-09 (operator decision).** `lexicon/injection-lexicon.json`
0.2.0 now carries a commons-owned `id` per pattern, plus an `aliases` object naming what each
seeding runtime calls it. The id was **adopted verbatim from the guard's port**, not invented
here — that port already carried both names, so the mapping came from source data. The
detection data is provably unmoved: stripping `id`, `aliases` and the new `pattern_id_space`
block reproduces the previous committed file **byte for byte** (23 566 bytes, identical).
**Ratified by both runtimes, 2026-08-09.** `llm-security` accepted the id space as-is,
including the 0.2.0 proposal, and treats an id change as breaking on the same terms; the
guard confirmed the space its own port supplied. `lexicon/injection-lexicon.json` 0.5.0
records both. The id is a cross-runtime contract now, not a proposal.
~~**The second half of the blocker stands, and it did not get smaller.**~~ **Dissolved
2026-08-10 by measurement, not by a decision.** This paragraph said the 13 divergent
patterns had no agreed expected behaviour, so their fixtures could not be authored, and that
someone would have to pick whose recall cost was the contract.
Nobody had to. The question was never asked of the right inputs. Every divergence in the
table above was found on a **witness** input — an attribute run padded past 256 characters,
an interior `<`, an unclosed `<script>`. The corpus is built from the seed suite's payloads,
which are short, unpadded and contain none of those shapes. Run through both runtimes'
public entry points, all 83 patterns produce **identical lexicon finding sets**, and that
includes 13 of 13 of the divergent ones on their own case input. Method, commits and counts:
[`conformance/manifest.json`](../conformance/manifest.json).
So the 13 carry no marker in the corpus and no caveat. Marking them would assert a doubt the
measurement disproves for these inputs, which is a different defect from the one it would
appear to prevent.
**What still stands is everything above this heading.** The divergence is real, it is
unresolved, and it will reappear the moment a fixture is written over a witness input.
`llm-security` has decided **not** to adopt the guard's regex strategy at this point: v0.1.0
is a behaviour-preservation release on their side too, and swapping strategies mid-vendoring
would void their own golden gate. Both behaviours stay registered as known divergence per
pattern. Their three confirmed recall holes are logged as a v8.x task; when it lands they
will say so, and those rows can close then. Until then the correct description of each is
*"known divergence, `llm-security` side has an open recall hole, measured 2026-08-09; not
reachable from any input in the v0.1.0 corpus"* — not *"undecided"*, and not *"resolved"*.