# Lexicon port divergence — commons vs the Python guard **Status: informative.** Nothing here is normative and nothing here changes a data file. It records a measured disagreement between two ports of one source table, so the decision can be taken where each table is tested. Under this repository's behaviour-preservation invariant, a divergence found here is **reported, not fixed**. Produced 2026-08-09. Every number below came from a command; the scripts live in the session scratchpad rather than in this repository, because executable code here would breach the charter. They are reproducible from the method column. **Revised the same day, after `llm-security`'s source became readable and both runtimes replied.** Four things changed, and three of them are corrections to this file rather than new results: 1. The claim that **neither runtime misses an attack** is **retracted**. It does. See *What this does not show* — the measurement behind that claim unioned pattern tables belonging to two different runtimes and read the result as a statement about each. 2. One of the 13 divergences does not reach report level, so **12** is the number that changes what a report says. The 13 still blocks `conformance/`. 3. The `hybrid` **severity is resolved** to `high` — the reported hint was correct, and the citation behind it was not. 4. The **pattern id space is ratified** by both runtimes. Corrections are marked in place rather than edited away, because a reader who saw the first version needs to know which sentence moved. ## What was compared | Side | Artefact | Version | | --- | --- | --- | | commons | [`lexicon/injection-lexicon.json`](../lexicon/injection-lexicon.json) | file `version` as committed | | guard | `llm-ingestion-pipeline-security` `src/llm_ingestion_guard/injection_lexicon.json` | lexicon `version` 1.0, repo v0.3.4, commit `0bf0729` | Both are **ports of the same file**: `llm-security/scanners/lib/injection-patterns.mjs`. The guard's JSON says so in its own `note` field — *"Injection lexicon ported from llm-security injection-patterns.mjs. Single source of truth."* Commons extracted the same table from an operator dump of that module. That is what makes the comparison worth running. These are not two different detectors that happen to overlap; they are two transcriptions of one table, and where they disagree, they disagree about what the same source says. ## Result | Measure | Method | Result | | --- | --- | --- | | Pattern count, both sides | count entries | 83 and 83 | | Label correspondence | match commons `label` to guard `desc`, em-dash normalised to hyphen | **83/83** | | Regex source byte-identical | string compare | **64/83** | | Differing regex text | string compare | 19 | | — of those, provably equivalent | unescape commons' JS-isms (`\/` → `/`, `\uXXXX` → literal) and compare for string identity | **6/6 identical** | | — of those, behaviourally divergent | differential match-set comparison (offsets + matched text), targeted corpus per pattern family | **13**, each with a concrete witness input | | — of those, divergent at REPORT level | re-check whether a sibling pattern raises an equivalent finding on the same witness | **12** — one of the 13 is a label-set difference only | | Total input comparisons | count | 401 | | Flags | compare declared flags | **0 differences** | | Severity / family | commons family vs guard `severity` | **8 differences** (all `hybrid`) — resolved, see *Severity* below | 64 identical + 6 escaping-only + 13 divergent = 83. **Read the 13 and the 12 as answering different questions.** Thirteen patterns produce different label sets. Twelve of those change what a report would say. The gap is the `hybrid-xss: iframe with executable src` row: the guard's version of that one pattern misses the witness, but its `hybrid-xss: javascript: URI scheme` pattern fires on the same input at the same severity and the same OWASP anchor, so a reader of the guard's report still sees a `high` / `LLM01` finding on that payload. Measured, not reasoned: the guard's table matched `hybrid-xss:javascript-uri` (high, LLM01) and nothing else; the Node engine matched both `hybrid-xss: javascript: URI scheme` and `hybrid-xss: iframe with executable src`. A `conformance/expected.json` would have to encode the difference, because a fixture names labels — which is why the 13 stood as the number that blocked the corpus. *(Superseded 2026-08-10: it blocks a fixture written over a **witness** input, and the corpus contains none. See* Consequence for `conformance/` *below.)* The 6 escaping-only pairs are equivalent by **string identity after unescaping**, not merely by absence of a counter-example: `<\/system>` vs ``, `text\/html` vs `text/html`, `(?:https?:\/\/…)` vs `(?:https?://…)`, `(?:~\/)?` vs `(?:~/)?`, and two `\uXXXX`-vs-literal character classes. Commons writes code points as escapes by convention; the guard's `note` records the same choice for non-ASCII. ## The 13, and which side misses what Divergence is **not one-directional**. Each side has inputs the other's pattern does not match. | Pattern | Direction | Witness input | | --- | --- | --- | | `html-obfuscation: display:none element with content` | commons misses; span differs | attribute run padded past 256 chars | | `html-obfuscation: visibility:hidden element` | span only | `
x
` | | `html-obfuscation: off-screen positioned element` | span only | same shape, `position:absolute` | | `html-obfuscation: zero font-size element` | span only | same shape, `font-size:0` | | `html-obfuscation: zero opacity element` | span only | same shape, `opacity:0` | | `html-obfuscation: zero-size overflow-hidden element` | span only | same shape, `height:0;overflow:hidden` | | `html-obfuscation: injection in aria-label attribute` | **commons misses** | `x` | | `markdown: suspicious instruction in link-reference comment` | **guard misses** | `[//]: # (a (b) ignore)` | | `markdown: injection payload in link anchor text` | **guard misses** | `[click ignore this](http://e.com/a(b))` | | `sub-agent: spawn instruction with dangerous capability keywords` | **guard misses** | `create an agent that` + 20 filler words + `execute` | | `sub-agent: delegation with safety bypass instruction` | **guard misses** | `delegate to a new agent` + 200 chars + `bypass` | | `hybrid-xss: `. ## Why they diverge: two different ReDoS mitigations of one table This is not drift, and framing it as a bug in either repository would be wrong. Both ports have been hardened against catastrophic backtracking, by **different strategies**: - **The Node side bounds the run.** `[^"]{0,256}`, `[^>]{1,256}`. Cost: an attacker who pads the attribute past 256 characters falls out of the pattern. - **The guard excludes the anchor character.** `[^><]`, `[^\]\[]`, `[^)(]`. Cost: content that legitimately contains that character stops matching. **Every divergence on the guard's side is documented at source, and traceable to the commit that introduced it.** An earlier draft of this file claimed the two sub-agent bounds were not; that was wrong, and it was wrong because the search behind it never looked outside the CHANGELOG. Both mechanisms are named in `lexicon.py`'s own module docstring: > **Bounded token gaps** — the two sub-agent patterns whose seed form nested `.*?` are ported > with `(?:\S+\s+){0,N}?`. > > **Anchor exclusion** — […] Measured across all 83 patterns arm by arm, two markdown patterns > had this defect; both now exclude the anchor character from the run. `git log -S` separates the two: the eight `[^><]` patterns (six html-obfuscation, two hybrid-xss) arrived with `cff0437`, *"fix(output): 19 quadratic regex runs on the output path"* — the v0.3.2 campaign, whose CHANGELOG describes exactly this remedy (*"exclude the character that opens the pattern's own anchor (`[` for markdown, `<` for tags)"*) across a sweep of *"150 patterns across 11 tables"*. The `{0,12}` / `{0,120}` sub-agent bounds are older still: they arrived with the original port commit `f397cd9`, so they were never a divergence introduced later — they are how that table was transcribed in the first place. The v0.3.4 entry also states the measured recall cost, naming precisely the two exceptions this comparison rediscovered: *"URLs containing a literal `(` inside a markdown link target and comment bodies containing a literal `(` before the keyword."* Worth recording, because it anticipates the criticism the Node side invites: the guard considered bounding those runs and **rejected it**, on the grounds that *"the content is attacker-controlled, so padding past a bound would be a one-line bypass."* That is the same objection the `{0,256}` witness above demonstrates against the Node table. The two projects reached opposite conclusions from the same reasoning, which is the substance of the disagreement — not an oversight on either side. Neither strategy is free, and neither is obviously right. That is the decision the two owning repositories have to take, and it is not commons' to take for them. ## Severity: the 8 hybrid patterns **Resolved 2026-08-09. The two sides never disagreed; only the evidence did.** Commons recorded the `hybrid` family with `severity: null` and a note that the seed dump did not supply it, so a consumer **MUST NOT** assume one. The guard's port assigned all eight `high`. Copying the guard's value would have converted a documented gap into an unverified claim, so it was reported instead — and the report was right: the value **is** `high`, confirmed at the module, and `lexicon/injection-lexicon.json` 0.5.0 now carries it. The eight differences in the table above are closed. The part worth keeping is where the value lives. It is not a field. The engine assigns it by pushing `HYBRID_PATTERNS` matches straight into the `high` bucket at `injection-patterns.mjs:274-281`. `severity.mjs` contains **no injection-family severity at all** — re-measured 2026-08-10 at `b0de0ca`: `CRITICAL_PATTERNS`, `HIGH_PATTERNS`, `MEDIUM_PATTERNS` and `HYBRID_PATTERNS` appear there zero times. ~~**The guard's port cites `severity.mjs`.** So the guard held the right value behind a citation that leads nowhere, and refusing to copy it was right for a better reason than the one given at the time: this particular port could not have read what it claimed to.~~ **Retracted 2026-08-10. The guard's port cites the right file.** This paragraph was never measured here; it restated an assertion received from `llm-security` (`20260809T201048Z`: *"Guardens port satte riktig verdi, men kunne ikke ha lest den fra fila den oppgir"*) as a commons finding. Measured against the guard's own tree: `severity.mjs` has **never** appeared in `src/llm_ingestion_guard/injection_lexicon.json` at any point in that file's history (`git log -S` returns no commits), and at `0bf0729` — the commit `conformance/manifest.json` pins — the only tree-wide occurrence is `docs/PLAN.md:114`, correctly attributing the *report* module to `output.mjs` + `severity.mjs`. The guard's only source statement for the lexicon is the `note` at `injection_lexicon.json:3`, and it names `injection-patterns.mjs`. Refusing to copy the value was still the right call — but for the plain reason, that a port is second-hand evidence, not for the sharper one claimed above. The sharper reason was itself a wrong citation to a right value, which is the defect this section was written to warn about. It survived here because it arrived from a repository that had measured the *other* half of the claim correctly, and the correct half carried the incorrect half past review. ## What this does not show - ~~**Not that either runtime misses an attack.**~~ **Retracted 2026-08-09. It does.** This bullet claimed that every witness payload still produced a finding via `active-content: constructs.raw-html`, so no attack went unflagged. The measurement behind it was wrong in method, not in arithmetic: the payloads were run against the **union of every pattern table this repository holds** — 111 rules across the lexicon, `active-content.json` and `secret-egress.json` — and the rescuing hit came from `active-content.json`. That table is the **Python guard's**. `llm-security` has no active-content table at all. A union of commons tables is not any single runtime's coverage, and treating it as one turned two runtimes' combined reach into a claim about each of them. Measured properly, through `llm-security`'s own entry point `scanForInjection()` — the whole engine, with normalisation, homoglyph folding, the rot13 variant and all four pattern arrays, at `b0de0ca`: | Witness | `scanForInjection()` result | | --- | --- | | `` returns `high` (hybrid-xss), and the short aria-label variant returns `critical`. So the `{0,256}` window is a real evasion window and the `