git-subtree-dir: scanners/commons git-subtree-split: 0ffee85a4b83b3661185488c06ed9a9994c11412
313 lines
21 KiB
Markdown
313 lines
21 KiB
Markdown
# Changelog
|
|
|
|
All notable changes to this project will be documented in this file.
|
|
|
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this
|
|
project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
|
|
Versioning note: the repository tag versions **the contract** (file set, key names,
|
|
case ids, disposition semantics). Each JSON file additionally carries its own
|
|
`"version"` field, bumped when that file changes.
|
|
|
|
## [0.1.0] — 2026-08-10
|
|
|
|
Initial extraction. Runtime-neutral detection data, the finding contract, and a conformance
|
|
corpus, extracted from the `llm-security` Node implementation and a Python guard **without
|
|
behaviour change** — that invariant is the release, not a caveat on it.
|
|
|
|
What the tag is worth resting on: seven of the eight JSON artefacts were rebuilt from the
|
|
commons file alone and diffed against their source implementation, three of them against the
|
|
source module at a pinned commit. The eighth says `verified: false` about itself. The corpus
|
|
holds 83 cases on which both seeding runtimes were measured agreeing exactly.
|
|
|
|
What it is not: `spec/decode-pipeline.md` does not exist, and the corpus constrains one of
|
|
the seven data files. Both absences are named in *Not included* rather than papered over.
|
|
|
|
### Added
|
|
|
|
- `conformance/` — **83 cases, one per injection-lexicon pattern**, plus `manifest.json`.
|
|
Each case is a directory holding `input.txt` (the exact bytes, no trailing newline) and
|
|
`expected.json` (the findings, named by commons pattern `id`).
|
|
|
|
Both seeding runtimes were measured producing the **same lexicon finding set on all 83**,
|
|
through their public entry points — `scanForInjection()` at `b0de0ca` and
|
|
`scan_output(source=OUTPUT)` at `0bf0729` — with labels mapped to commons ids through the
|
|
lexicon's own `aliases` block. Not through rebuilt regex tables: a table-level comparison
|
|
yields a number that describes neither runtime, which is the mistake the divergence
|
|
document had to retract.
|
|
|
|
**The 13 divergent patterns are in, unmarked, and that is the substantive result.** Their
|
|
divergence was measured on witness inputs — an attribute run padded past 256 characters,
|
|
an interior `<`, an unclosed `<script>` — and none of those shapes occurs in a corpus
|
|
payload. All 13 agree on their own case input. Nobody had to pick whose recall cost
|
|
becomes the contract, because the question was never reachable from these inputs. A
|
|
per-case caveat would have asserted a doubt the measurement disproves.
|
|
|
|
Inputs are the guard's `coverage.py` payloads, reproduced verbatim. One runtime authored
|
|
them; what makes them a cross-runtime corpus is the measurement through the other, and the
|
|
manifest records the asymmetry rather than averaging it away.
|
|
|
|
- `spec/conformance-corpus.md` — **normative.** How a case is read: `input.txt` is bytes and
|
|
is not to be trimmed or re-encoded, `expected.json` names findings by `pattern_id` only
|
|
(severity and OWASP anchor are looked up in the lexicon, never restated), and
|
|
`exact-within-scope` requires equality **restricted to the data files the case names**.
|
|
|
|
The field is `pattern_id`, not `id`, because this repository already publishes an unrelated
|
|
finding `id`: `schema/finding.schema.json` defines it as `DS-<scanner>-<counter>` from a
|
|
process-global counter — stable across neither runs nor processes. Two normative documents
|
|
using one word for a stable rule identity and a volatile per-emission sequence number would
|
|
have produced runtimes failing every case for reasons unrelated to detection. §3.1 states
|
|
the distinction and publishes the bridge a runtime actually needs: its own label maps to a
|
|
`pattern_id` through the lexicon's `aliases` object, and a runtime absent from that object
|
|
has no published way to be compared at all.
|
|
|
|
Scoping is what makes exactness safe — the two runtimes do not implement the same set of
|
|
tables, so a whole-report comparison would fail for reasons unrelated to the pattern under
|
|
test. Exactness is what makes the corpus worth running — a contains-only corpus is passed
|
|
by a runtime that flags everything. `observed_out_of_scope` is evidence, never expectation,
|
|
and an absent runtime key means **unmeasured**, not measured-empty.
|
|
|
|
The document also states the one place this repository's "every JSON file carries a
|
|
top-level `version`" convention does not apply: fixtures are versioned as a corpus, in
|
|
`conformance/manifest.json`. Stated rather than left to be discovered.
|
|
|
|
- `schema/finding.schema.json` — the finding contract plus the SARIF output profile.
|
|
Normative. Closed against the producer in 0.2.0; the JSONL profile is `not applicable`.
|
|
- `signatures/active-content.json` — the EchoLeak class (CVE-2025-32711): 17 patterns,
|
|
severities, opacity floors and pass order, from the Python guard.
|
|
- `lexicon/injection-lexicon.json` — 83 prompt-injection patterns in four families
|
|
(21 critical, 32 high, 22 medium, 8 hybrid).
|
|
- `codepoints/carriers.json` — six carrier tables: zero-width characters, the Unicode Tags
|
|
block, the Supplementary Private Use Areas, BIDI controls, the Cyrillic presence set and
|
|
the 28-entry fold-to-Latin homoglyph map.
|
|
- `signatures/secret-egress.json` — the 18 fixed credential and token shapes. Array order
|
|
is normative.
|
|
- `mapping/owasp-map.json` — four taxonomy maps (LLM, ASI, AST, MCP) over one shared
|
|
16-prefix key set.
|
|
- `calibration/calibration.json` — risk-score tier constants, verdict thresholds, risk-band
|
|
cutoffs, posture grade thresholds.
|
|
- `signatures/malware-signatures.json` — the known-bad-identity table for the `SIG` class:
|
|
seven signatures over four families (`webshell`, `reverse_shell`, `cryptominer`,
|
|
`hacktool`), reproduced verbatim from `knowledge/signatures.json` at `b0de0ca`, key order
|
|
included, with the source file's byte length and SHA-256 pinned in `provenance`.
|
|
|
|
The rules were the easy half. The file's substance is the line between the table and the
|
|
engine, drawn in `engine_behaviour_not_data`: **no rule carries a `flags` field**, because
|
|
the engine compiles every pattern with `i` unconditionally — so a consumer that compiles
|
|
these case-sensitively silently under-matches all seven. Each pattern is also run against
|
|
five decode variants, not just raw bytes; rules are filtered by an enabled-families policy;
|
|
a rule fires once per file; operator rules are merged at scan time; and the loader defaults
|
|
four missing fields rather than rejecting a rule. None of that travels with the data, and
|
|
all of it changes what a consumer sees.
|
|
|
|
Two honesty notes are in `evidence_limits` rather than in prose. Seven signatures are not
|
|
malware coverage — a clean `SIG` result is not "no malware", and the seed runtime's own
|
|
header calls the table "deliberately tight". And three of the seven match on **names**
|
|
(`xmrig`, `mimikatz`, `meterpreter`), so a document *discussing* those tools matches; the
|
|
seed runtime papers over this by excluding `knowledge/`, `tests/`, `docs/` and
|
|
`node_modules/` from the scan, which is engine behaviour and does not come with the table.
|
|
|
|
Verified: 7/7 rule objects field-identical to source including key order, no non-ASCII
|
|
bytes, and all seven compile in Node bare, `i` and `iu` (21/21) and in Python `re` (7/7).
|
|
Note the exact family spellings — `reverse_shell`, not `reverse-shell`, and `cryptominer`,
|
|
not `miner`; they are policy keys, and the working note that seeded this file had both wrong.
|
|
|
|
### Verification
|
|
|
|
Every file above except `calibration.json` was proven rather than transcribed: the data was
|
|
rebuilt **from the commons JSON alone** and diffed against the source implementation. Each
|
|
file records its own result and its own limits.
|
|
|
|
`calibration/calibration.json` carries `verified: false`. Its source arrived as a prose
|
|
summary rather than as code, so no differential check was possible, and the file names the
|
|
checks that were not run instead of attaching a caveat to a pass.
|
|
|
|
The corpus was verified the same way the data was — by a harness that does **not** share the
|
|
generator's knowledge. It reads only the case directories, re-runs both runtimes on the bytes
|
|
it finds there, and checks every field of every `expected.json`, digests included: **83
|
|
cases, 0 failures**. Two further checks, because a corpus that cannot fail is not evidence:
|
|
commons' family severity matches the severity the guard emits per finding, **83/83**; and
|
|
deleting the middle third of each input breaks **76 of 83** expectations. The 7 survivors are
|
|
the shortest payloads, where the mutation leaves the trigger intact — that is a weak
|
|
mutation, not a weak fixture, and it is recorded as such rather than rounded up.
|
|
|
|
- `docs/lexicon-port-divergence.md` — informative. A differential comparison of the two
|
|
ports of `injection-patterns.mjs` (this repository's and the Python guard's): 83/83
|
|
patterns correspond, 64 are byte-identical, 6 differ only by escaping and are proven
|
|
equivalent, and **13 behave differently**, with a witness input for each and misses on
|
|
both sides. The cause is two different ReDoS mitigations of one table. **No data file was
|
|
changed** — behaviour preservation holds and the finding is reported to the owning
|
|
repositories.
|
|
|
|
**Revised 2026-08-09 with one retraction.** The document claimed that *neither runtime
|
|
misses an attack*, on the grounds that every witness payload still produced a finding. It
|
|
does miss. That measurement ran the payloads against the **union of every pattern table
|
|
this repository holds**, and the rescuing hit came from `active-content.json` — the Python
|
|
guard's table. `llm-security` has no active-content table at all, so a union of commons
|
|
tables was read as a statement about each runtime separately. Re-measured through
|
|
`llm-security`'s own `scanForInjection()` at `b0de0ca`, all three witness payloads return
|
|
**`found: false`** — no finding whatsoever — while controls in the same run behave
|
|
normally. Three confirmed recall holes, which `llm-security` attributes to its v7.8.3 #24
|
|
ReDoS hardening and has logged as a v8.x task.
|
|
|
|
Also corrected: one of the 13 divergences does not reach report level, because the guard's
|
|
`hybrid-xss:javascript-uri` fires on the same witness at the same severity and anchor. The
|
|
report-level number is **12**. And the `hybrid` severity question that the document reported
|
|
rather than resolved is now closed — the reported hint was right, the citation behind it was
|
|
not.
|
|
|
|
**Revised again 2026-08-10.** The document said 13 was the number blocking `conformance/`,
|
|
since a fixture names labels. It blocks a fixture written over a **witness** input, and the
|
|
corpus contains none — all 13 agree on their own case input. The divergence itself stands
|
|
unresolved and unchanged; what was wrong was the claim about what it blocked.
|
|
|
|
Corrections are marked in place rather than edited away.
|
|
|
|
### Changed
|
|
|
|
- `schema/finding.schema.json` **0.1.0 → 0.2.0** — the schema is **closed**. It was seeded
|
|
from `sarif-formatter.mjs`, which *consumes* findings, so its property list could only ever
|
|
be a lower bound and `additionalProperties` had to stay open. The producer is now known —
|
|
`finding()` in `scanners/lib/output.mjs`, line 32 — and it returns an object literal with
|
|
**exactly ten keys and no spread**: `id`, `scanner`, `severity`, `title`, `description`,
|
|
`file`, `line`, `evidence`, `owasp`, `recommendation`. `additionalProperties` is `false`,
|
|
and the two keys the old schema never knew about (`id`, `evidence`) are added.
|
|
|
|
`id` gets its own definition: `DS-<prefix>-<counter>`, pattern `^DS-[A-Za-z]+-[0-9]{3,}$`.
|
|
The `{3,}` is deliberate — `padStart(3, '0')` is a minimum, so a run emitting more than 999
|
|
findings produces four digits. The id comes from a process-global counter, so it is stable
|
|
neither across runs nor across processes, and the definition says so before someone keys on it.
|
|
|
|
Nullability is now evidence rather than convention. Five keys are emitted as `null` rather
|
|
than omitted (`opts.x || null`), so a serialised finding always carries all ten. The
|
|
exception is the four assigned straight from `opts`: omit `description` and the key is
|
|
`undefined` and vanishes from the JSON. Verified by calling the real producer — ten keys in
|
|
memory, nine after serialisation.
|
|
|
|
**`owasp` is a string, not an array, and not one code.** Multiple codes are joined with
|
|
`, `. Measured across the seed runtime: 31 distinct values over 157 emission sites, 13 of
|
|
them multi-code, and **four mix taxonomies inside a single value** (`LLM06, ASI02` and
|
|
friends) with no discriminator saying which is which. That sharpens the edition problem
|
|
`mapping/owasp-map.json` already records, and it has a consequence nobody had written down:
|
|
`sarif-formatter.mjs` builds `tags: [f.owasp]`, so a finding anchored to two taxonomies
|
|
produces **one** SARIF tag with a comma in it. Nothing filtering on `LLM06` will match.
|
|
Reproduced end to end through the real `finding()` and `toSARIF()`, and logged as
|
|
`known_lossiness.owasp-tag-not-split` — consumer behaviour in `llm-security`, not data, so
|
|
it is reported rather than fixed here.
|
|
|
|
The **JSONL profile is `not applicable`, not `unspecified`** — the distinction is the point.
|
|
`unspecified` would claim a profile exists and merely has not been written down. No
|
|
finding-JSONL exists: findings are emitted only inside a single JSON envelope
|
|
(`output.mjs:140`). The one module that does write JSONL, `audit-trail.mjs`, writes *audit
|
|
events* under a different schema — where `owasp` is an **array**. Same field name, different
|
|
type, same repository. A consumer reading both through one code path will be wrong about one
|
|
of them, so the profile records the trap instead of leaving a TODO.
|
|
|
|
Verified: the schema is valid Draft 2020-12, every finding built by the real producer
|
|
validates against it, and four negative controls (extra property, missing `id`, malformed
|
|
`id`, unknown severity) are all rejected.
|
|
|
|
One new open question, unpatched by design: the producer's JSDoc lists **seventeen** scanner
|
|
prefixes including `IDE`, while all four maps in `mapping/owasp-map.json` are keyed on
|
|
**sixteen** without it. An `IDE` finding has no taxonomy mapping in any map. Adding the key
|
|
would be inventing detection data.
|
|
|
|
- `lexicon/injection-lexicon.json` **0.4.0 → 0.5.0** — the last null in the file is filled and
|
|
the id space is ratified. Two blockers close, no detection data moves.
|
|
|
|
`families[hybrid].severity` was `null`, deliberately, because the seed dump did not supply
|
|
it. It is **`high`** — and the interesting part is where that is written. The hybrid family
|
|
has no severity field anywhere; the engine assigns one by pushing `HYBRID_PATTERNS` matches
|
|
straight into the `high` bucket at `injection-patterns.mjs:274-281`. Both this repository
|
|
and the Python guard had first looked in `severity.mjs`, which contains no injection-family
|
|
severity at all. The guard's port holds the right value behind that wrong citation, so
|
|
`severity_provenance.not_from` records the miss explicitly: a wrong citation to a right
|
|
value is the harder defect to catch later.
|
|
|
|
`pattern_id_space.not_yet_confirmed` is replaced by `ratification`. Both seeding runtimes
|
|
agreed on 2026-08-09 — `llm-security` ratified the 0.2.0 proposal as-is and treats an id
|
|
change as breaking on the same terms, and the guard confirmed the space its own port
|
|
supplied. `id` is now a cross-runtime contract, which is what `conformance/` was waiting
|
|
on to be able to name a finding.
|
|
|
|
`alias_evidence.llm_security` is sharpened rather than upgraded. All 83 alias strings were
|
|
confirmed equal to the module's `label` field, in order — so the alias is certainly the
|
|
pattern's name **in the table**. It is still not established that a finding carries it: the
|
|
producer is `output.mjs:finding()`, which emits `title` and has no `label` key at all.
|
|
Verified at table level, one level short of where it would matter. Match on `id`.
|
|
|
|
- `lexicon/injection-lexicon.json` **0.3.0 → 0.4.0** — verified against the source module
|
|
instead of against the dump it was transcribed from, and **two false provenance claims
|
|
retracted**. The source is now pinned: `b0de0ca` on the public remote, imported in Node
|
|
and compared entry by entry on `source`, `flags` and `label`.
|
|
|
|
The result is **83/83 byte-identical to source**, which is not what the file previously
|
|
claimed. It said two patterns had been rewritten from raw code points into `\uXXXX`
|
|
escapes; the module already writes them escaped, so nothing had been rewritten. The stored
|
|
pattern text was right the whole time — only the account of where it came from was wrong.
|
|
The dump had rendered the module's escapes as the characters they denote, and this
|
|
repository re-escaped them, arriving at the correct bytes by way of an incorrect story.
|
|
|
|
The same inversion ran the other way in `multi-lang:french`, which carried the class
|
|
spelled with a raw accented Latin `e` where the module writes it as the escape
|
|
`\u00e9` inside the same character class.
|
|
That was the one pattern of 83 not byte-identical to source,
|
|
and it is corrected. The two spellings are the same regular expression — verified in Node
|
|
bare and under `u`, and in Python `re`, over accented, unaccented, uppercase and
|
|
non-matching French input, with identical match offsets — so **no behaviour moved**. No
|
|
pattern in the file contains a non-ASCII byte now, matching the module, whose regex
|
|
literals are pure ASCII throughout.
|
|
|
|
Structurally: `normalisations` is now `[]` with a `normalisations_note`, matching the
|
|
convention already used in `signatures/secret-egress.json`, and a new `source_fidelity`
|
|
block carries the counts, the method, the verified class membership, and both retractions
|
|
in full. Retracted claims are recorded rather than deleted — the earlier equivalence
|
|
evidence (692 Node comparisons, 236 Python) remains true, it is simply no longer
|
|
load-bearing.
|
|
|
|
- `lexicon/injection-lexicon.json` **0.2.0 → 0.3.0** — the two aliases are no longer presented
|
|
as equally backed. `pattern_id_space.alias_evidence` now records each one separately:
|
|
`llm_ingestion_guard` is **verified** (the guard's coverage matrix asserts on that exact
|
|
string, so it is demonstrably what a guard finding carries), while `llm_security` is
|
|
**not** — it is the pattern table's own name, and the finding producer was never supplied,
|
|
with the known Node finding shape using `title` rather than `label`. Averaging the two into
|
|
one file-level claim would have repeated the defect this repository corrects per-table
|
|
elsewhere.
|
|
|
|
Also: `normalisations[].affects` now keys on `id` with the prose names kept beside it as
|
|
`affects_labels`. An internal cross-reference on label was a second identity space inside
|
|
the file the id was added to unify.
|
|
|
|
- `lexicon/injection-lexicon.json` **0.1.0 → 0.2.0** — every pattern gains a commons-owned
|
|
`id` and an `aliases` object naming what each seeding runtime calls it, plus a top-level
|
|
`pattern_id_space` block explaining the field. This exists because a `conformance/`
|
|
fixture has to name a finding and the two runtimes do not name the same pattern the same
|
|
way.
|
|
|
|
The id was **adopted verbatim from the guard's port**, which already carried both names,
|
|
rather than invented here. Matching was by `label` ↔ `desc` with em-dash normalised to
|
|
hyphen: 83/83, one-to-one, ids unique.
|
|
|
|
**No detection data moved.** Labels, patterns and flags are byte-identical in sequence,
|
|
no `flags` key was invented (78 before, 78 after), and stripping the three new fields
|
|
reproduces the previous committed file byte for byte — 23 566 bytes, identical. All 83
|
|
patterns still compile in Node bare and under `u` (166/166) and in Python `re` (83/83).
|
|
|
|
Neither `llm-security` nor the guard has ratified this id space yet; both were asked by
|
|
coord on 2026-08-09, and the file says so rather than implying agreement.
|
|
|
|
### Not included
|
|
|
|
- `spec/decode-pipeline.md` — needs the decode implementation. A normative spec inferred
|
|
from a data dump would be worse than an absent one.
|
|
- **Conformance for the other four tables.** The corpus covers the injection lexicon only.
|
|
The carrier, active-content and secret-egress tables have 11 convertible cases waiting in
|
|
the guard's matrix, and no ratified cross-runtime finding id between them — writing those
|
|
fixtures would mint a contract unilaterally, in the same stroke as the tag. Named in
|
|
`conformance/manifest.json` under `scope_planned`.
|
|
- The 29 non-convertible cases of the guard's 134 assert a runtime's **API surface** — that
|
|
a Python call raises `OKFPathError`, that a disposition engine composes two findings a
|
|
particular way. This repository does not own an API, so those belong to the guard's suite.
|
|
|
|
`spec/decode-pipeline.md` is named in the README as planned rather than linked, so nothing
|
|
in the repository points at a file that does not exist.
|