docs: retract the "neither runtime misses an attack" claim, fix three counts
Four corrections across the two informative documents, all of them to
claims this repository made and none of them to a data file.
The retraction is the important one. lexicon-port-divergence.md said
neither runtime misses an attack: every witness payload still produced a
finding, so the divergence was only in which labels were raised. That was
wrong in method rather than in arithmetic. The payloads had been run
against the UNION of every pattern table this repository holds - 111 rules
across the lexicon, active-content.json and secret-egress.json - and the
hit that rescued them came from active-content.json, which is the Python
guard's table. llm-security has no active-content table at all. Unioning
two runtimes' tables and reading the result as a statement about each of
them turns combined reach into individual coverage.
Re-measured through llm-security's own entry point scanForInjection() at
b0de0ca - the whole engine, normalisation, homoglyph folding, rot13 and all
four arrays - all three witness payloads return found: false. Not a
different label set: no finding at all. Controls in the same run behave
normally, so the {0,256} window is a real evasion window and the <script>
pattern really does require a closing tag. llm-security reached the same
three results independently and has logged them as a v8.x task.
Second: one of the 13 divergences does not reach report level. On the
iframe witness the guard misses hybrid-xss:iframe-src but its
hybrid-xss:javascript-uri fires at the same severity and the same OWASP
anchor, so a reader of its report still sees the payload flagged. Both
sides were re-measured to confirm. The report-level number is 12; the 13
stands as the number that blocks conformance/, because a fixture names
labels.
Third: extraction-plan.md said six of the confusables have no entry in the
fold map. It is zero - all 13 are keys in the 28-entry map, and
carriers.json has said so correctly since it was written. The prose, not
the data, carried the error, and the number six most likely came from
running the comparison the other way (five fold-map Cyrillic keys are not
confusables). It also called the diverging sets deliberately distinct;
llm-security states the U+0456 / U+0445 divergence is drift, not design,
and has logged it for v8.x. Those two readings imply opposite things about
whether the sets should be expected to converge.
Fourth: the seed-sources section still asserted that sessions here have no
read access to llm-security and that content arrives only as a dump. The
operator changed that on 2026-08-09. It now records the read-only pinned-
commit boundary, the still-absolute prohibition on the sibling working copy
and on writing, and the reason the change matters - four attribution
defects caught so far, every one a transcription artefact.
Status blocks updated: all eight JSON artefacts now exist, seven verified
differentially and calibration.json still honestly verified: false; the
conformance blocker is half-cleared, with naming ratified and the 13
divergent patterns still unresolved.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SNMcqrfNyoLRQ7qXUFZnb9
This commit is contained in:
parent
47760d2264
commit
ff18768b71
3 changed files with 192 additions and 43 deletions
|
|
@ -9,6 +9,22 @@ Produced 2026-08-09. Every number below came from a command; the scripts live in
|
|||
scratchpad rather than in this repository, because executable code here would breach the
|
||||
charter. They are reproducible from the method column.
|
||||
|
||||
**Revised the same day, after `llm-security`'s source became readable and both runtimes
|
||||
replied.** Four things changed, and three of them are corrections to this file rather than
|
||||
new results:
|
||||
|
||||
1. The claim that **neither runtime misses an attack** is **retracted**. It does. See
|
||||
*What this does not show* — the measurement behind that claim unioned pattern tables
|
||||
belonging to two different runtimes and read the result as a statement about each.
|
||||
2. One of the 13 divergences does not reach report level, so **12** is the number that
|
||||
changes what a report says. The 13 still blocks `conformance/`.
|
||||
3. The `hybrid` **severity is resolved** to `high` — the reported hint was correct, and the
|
||||
citation behind it was not.
|
||||
4. The **pattern id space is ratified** by both runtimes.
|
||||
|
||||
Corrections are marked in place rather than edited away, because a reader who saw the first
|
||||
version needs to know which sentence moved.
|
||||
|
||||
## What was compared
|
||||
|
||||
| Side | Artefact | Version |
|
||||
|
|
@ -35,12 +51,24 @@ disagree about what the same source says.
|
|||
| Differing regex text | string compare | 19 |
|
||||
| — of those, provably equivalent | unescape commons' JS-isms (`\/` → `/`, `\uXXXX` → literal) and compare for string identity | **6/6 identical** |
|
||||
| — of those, behaviourally divergent | differential match-set comparison (offsets + matched text), targeted corpus per pattern family | **13**, each with a concrete witness input |
|
||||
| — of those, divergent at REPORT level | re-check whether a sibling pattern raises an equivalent finding on the same witness | **12** — one of the 13 is a label-set difference only |
|
||||
| Total input comparisons | count | 401 |
|
||||
| Flags | compare declared flags | **0 differences** |
|
||||
| Severity / family | commons family vs guard `severity` | **8 differences** (all `hybrid`) |
|
||||
| Severity / family | commons family vs guard `severity` | **8 differences** (all `hybrid`) — resolved, see *Severity* below |
|
||||
|
||||
64 identical + 6 escaping-only + 13 divergent = 83.
|
||||
|
||||
**Read the 13 and the 12 as answering different questions.** Thirteen patterns produce
|
||||
different label sets. Twelve of those change what a report would say. The gap is the
|
||||
`hybrid-xss: iframe with executable src` row: the guard's version of that one pattern misses
|
||||
the witness, but its `hybrid-xss: javascript: URI scheme` pattern fires on the same input at
|
||||
the same severity and the same OWASP anchor, so a reader of the guard's report still sees a
|
||||
`high` / `LLM01` finding on that payload. Measured, not reasoned: the guard's table matched
|
||||
`hybrid-xss:javascript-uri` (high, LLM01) and nothing else; the Node engine matched both
|
||||
`hybrid-xss: javascript: URI scheme` and `hybrid-xss: iframe with executable src`. A
|
||||
`conformance/expected.json` still has to encode the difference, because a fixture names
|
||||
labels — which is why the 13 stands as the number that blocks the corpus.
|
||||
|
||||
The 6 escaping-only pairs are equivalent by **string identity after unescaping**, not merely
|
||||
by absence of a counter-example: `<\/system>` vs `</system>`, `text\/html` vs `text/html`,
|
||||
`(?:https?:\/\/…)` vs `(?:https?://…)`, `(?:~\/)?` vs `(?:~/)?`, and two `\uXXXX`-vs-literal
|
||||
|
|
@ -65,7 +93,12 @@ Divergence is **not one-directional**. Each side has inputs the other's pattern
|
|||
| `sub-agent: spawn instruction with dangerous capability keywords` | **guard misses** | `create an agent that` + 20 filler words + `execute` |
|
||||
| `sub-agent: delegation with safety bypass instruction` | **guard misses** | `delegate to a new agent` + 200 chars + `bypass` |
|
||||
| `hybrid-xss: <script> tag in content (agent context XSS)` | **commons misses**; span differs | `<script>alert(1)` (unclosed), `<script src=x.js>` |
|
||||
| `hybrid-xss: iframe with executable src (agent context XSS)` | **guard misses** | `<iframe data-x="<" src="javascript:alert(1)">` |
|
||||
| `hybrid-xss: iframe with executable src (agent context XSS)` | **guard misses the label, not the attack** | `<iframe data-x="<" src="javascript:alert(1)">` |
|
||||
|
||||
The last row is the one of the 13 that does not reach report level. The guard's
|
||||
`hybrid-xss:javascript-uri` (`javascript\s*:`, high, LLM01) matches that witness, so the
|
||||
payload is still flagged at the same severity and anchor; only the label set differs — one
|
||||
finding instead of two. The remaining 12 rows change what a report says.
|
||||
|
||||
"Span only" means both sides produce a match on the same input but over different extents —
|
||||
the guard's match starts at an interior `<`. Whether that matters depends on whether a
|
||||
|
|
@ -120,30 +153,65 @@ repositories have to take, and it is not commons' to take for them.
|
|||
|
||||
## Severity: the 8 hybrid patterns
|
||||
|
||||
Commons records the `hybrid` family with `severity: null` and a note that the seed dump did
|
||||
not supply it, so a consumer **MUST NOT** assume one. The guard's port assigns all eight
|
||||
`high`.
|
||||
**Resolved 2026-08-09. The two sides never disagreed; only the evidence did.**
|
||||
|
||||
**This has deliberately not been copied into commons.** The guard's port is a second-hand
|
||||
transcription, not the producing module; adopting its value would convert a documented gap
|
||||
into an unverified claim, which is the defect class this repository's changelog already
|
||||
records twice. It is reported instead, as a strong hint that the Node source assigns `high`,
|
||||
for `llm-security` to confirm from the module.
|
||||
Commons recorded the `hybrid` family with `severity: null` and a note that the seed dump did
|
||||
not supply it, so a consumer **MUST NOT** assume one. The guard's port assigned all eight
|
||||
`high`. Copying the guard's value would have converted a documented gap into an unverified
|
||||
claim, so it was reported instead — and the report was right: the value **is** `high`,
|
||||
confirmed at the module, and `lexicon/injection-lexicon.json` 0.5.0 now carries it. The eight
|
||||
differences in the table above are closed.
|
||||
|
||||
The part worth keeping is where the value lives. It is not a field. The engine assigns it by
|
||||
pushing `HYBRID_PATTERNS` matches straight into the `high` bucket at
|
||||
`injection-patterns.mjs:274-281`. The guard's port cites `severity.mjs` — a file that
|
||||
contains **no injection-family severity at all**. So the guard held the right value behind a
|
||||
citation that leads nowhere, and a reviewer following that citation to check the number would
|
||||
have found nothing and drawn no conclusion. Refusing to copy it was the right call for a
|
||||
reason better than the one given at the time: not merely that a port is second-hand, but that
|
||||
this particular port could not have read what it claimed to.
|
||||
|
||||
## What this does not show
|
||||
|
||||
- **Not that either runtime misses an attack.** All witness payloads were re-tested against
|
||||
every pattern table commons holds (111 compiled rules: 83 lexicon, 28 from
|
||||
`active-content.json` and `secret-egress.json`). All of them still produce a finding, via
|
||||
`active-content: constructs.raw-html`. The divergence is in the **finding set** — which
|
||||
labels are raised — not in whether anything is raised at all. That still matters, because a
|
||||
finding set is exactly what a `conformance/expected.json` would encode.
|
||||
- **Not that `llm-security` behaves as described here.** Only its extracted pattern table was
|
||||
available. Whether the Node engine runs an active-content table that compensates cannot be
|
||||
determined from this repository.
|
||||
- **Not dump-to-module fidelity.** Every check above proves the two *ports* agree or disagree.
|
||||
That either matches `injection-patterns.mjs` remains `llm-security`'s assertion,
|
||||
reproducible only in a session with read access to it.
|
||||
- ~~**Not that either runtime misses an attack.**~~ **Retracted 2026-08-09. It does.** This
|
||||
bullet claimed that every witness payload still produced a finding via
|
||||
`active-content: constructs.raw-html`, so no attack went unflagged. The measurement behind
|
||||
it was wrong in method, not in arithmetic: the payloads were run against the **union of
|
||||
every pattern table this repository holds** — 111 rules across the lexicon,
|
||||
`active-content.json` and `secret-egress.json` — and the rescuing hit came from
|
||||
`active-content.json`. That table is the **Python guard's**. `llm-security` has no
|
||||
active-content table at all. A union of commons tables is not any single runtime's
|
||||
coverage, and treating it as one turned two runtimes' combined reach into a claim about
|
||||
each of them.
|
||||
|
||||
Measured properly, through `llm-security`'s own entry point `scanForInjection()` — the
|
||||
whole engine, with normalisation, homoglyph folding, the rot13 variant and all four pattern
|
||||
arrays, at `b0de0ca`:
|
||||
|
||||
| Witness | `scanForInjection()` result |
|
||||
| --- | --- |
|
||||
| `<script>alert(1)` (unclosed) | `found: false` — no finding at all |
|
||||
| `<script src=x.js>` | `found: false` — no finding at all |
|
||||
| `<a aria-label="` + 300 × `a` + `ignore">` | `found: false` — no finding at all |
|
||||
|
||||
Controls in the same run behave as expected: `<script>alert(1)</script>` returns `high`
|
||||
(hybrid-xss), and the short aria-label variant returns `critical`. So the `{0,256}` window
|
||||
is a real evasion window and the `<script>` pattern really does require a closing tag.
|
||||
`llm-security` reached the same three results independently and attributes the cause to
|
||||
their own v7.8.3 #24 ReDoS hardening, which traded recall for boundedness without seeing
|
||||
the window. Three confirmed recall holes, logged there as a v8.x task.
|
||||
|
||||
What survives from the original bullet is only this: the divergence is *also* in the
|
||||
finding set, which is what a `conformance/expected.json` encodes.
|
||||
- **Not that the guard misses an attack.** The guard's side of the 13 was re-checked the same
|
||||
way, and the one row that looked like a miss (`hybrid-xss: iframe with executable src`) is
|
||||
covered by a sibling pattern at the same severity and anchor. Its remaining divergences are
|
||||
label-set and span differences.
|
||||
- **Not dump-to-module fidelity.** *Superseded 2026-08-09.* Every check above proves the two
|
||||
*ports* agree or disagree. Commons' side is now settled separately: the lexicon is verified
|
||||
byte-identical to `injection-patterns.mjs` at `b0de0ca`, 83/83, which is recorded in the
|
||||
data file rather than here. The guard's fidelity to the module remains its own to
|
||||
establish.
|
||||
- **Not exhaustive.** The corpus is targeted per pattern family, 401 comparisons. Absence of a
|
||||
witness proves nothing except for the 6 escaping-only pairs, which are settled by string
|
||||
identity rather than by the corpus.
|
||||
|
|
@ -168,7 +236,17 @@ here — that port already carried both names, so the mapping came from source d
|
|||
detection data is provably unmoved: stripping `id`, `aliases` and the new `pattern_id_space`
|
||||
block reproduces the previous committed file **byte for byte** (23 566 bytes, identical).
|
||||
|
||||
Neither runtime has agreed to it yet — both were asked by coord the same day — so the file
|
||||
records it as this repository's proposal rather than a ratified contract. The second half of
|
||||
the blocker stands: the 13 divergent patterns still have no agreed expected behaviour, so
|
||||
their fixtures cannot be authored even now that they can be named.
|
||||
**Ratified by both runtimes, 2026-08-09.** `llm-security` accepted the id space as-is,
|
||||
including the 0.2.0 proposal, and treats an id change as breaking on the same terms; the
|
||||
guard confirmed the space its own port supplied. `lexicon/injection-lexicon.json` 0.5.0
|
||||
records both. The id is a cross-runtime contract now, not a proposal.
|
||||
|
||||
**The second half of the blocker stands, and it did not get smaller.** The 13 divergent
|
||||
patterns still have no agreed expected behaviour, so their fixtures cannot be authored even
|
||||
though they can now be named. `llm-security` has decided **not** to adopt the guard's regex
|
||||
strategy at this point: v0.1.0 is a behaviour-preservation release on their side too, and
|
||||
swapping strategies mid-vendoring would void their own golden gate. Both behaviours are
|
||||
therefore registered as known divergence per pattern. Their three confirmed recall holes are
|
||||
logged as a v8.x task; when it lands they will say so, and those rows can close then. Until
|
||||
then the correct description of each is *"known divergence, `llm-security` side has an open
|
||||
recall hole, measured 2026-08-09"* — not *"undecided"*.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue