# Changelog All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). Versioning note: the repository tag versions **the contract** (file set, key names, case ids, disposition semantics). Each JSON file additionally carries its own `"version"` field, bumped when that file changes. ## [Unreleased] Initial extraction, in progress. Runtime-neutral detection data and the finding contract, extracted from the `llm-security` Node implementation and a Python guard without behaviour change. **Not yet tagged** — see *Not included* below. ### Added - `schema/finding.schema.json` — the finding contract plus the SARIF output profile. Normative. The JSONL profile is deliberately left `unspecified`. - `signatures/active-content.json` — the EchoLeak class (CVE-2025-32711): 17 patterns, severities, opacity floors and pass order, from the Python guard. - `lexicon/injection-lexicon.json` — 83 prompt-injection patterns in four families (21 critical, 32 high, 22 medium, 8 hybrid). - `codepoints/carriers.json` — six carrier tables: zero-width characters, the Unicode Tags block, the Supplementary Private Use Areas, BIDI controls, the Cyrillic presence set and the 28-entry fold-to-Latin homoglyph map. - `signatures/secret-egress.json` — the 18 fixed credential and token shapes. Array order is normative. - `mapping/owasp-map.json` — four taxonomy maps (LLM, ASI, AST, MCP) over one shared 16-prefix key set. - `calibration/calibration.json` — risk-score tier constants, verdict thresholds, risk-band cutoffs, posture grade thresholds. ### Verification Every file above except `calibration.json` was proven rather than transcribed: the data was rebuilt **from the commons JSON alone** and diffed against the source implementation. Each file records its own result and its own limits. `calibration/calibration.json` carries `verified: false`. Its source arrived as a prose summary rather than as code, so no differential check was possible, and the file names the checks that were not run instead of attaching a caveat to a pass. - `docs/lexicon-port-divergence.md` — informative. A differential comparison of the two ports of `injection-patterns.mjs` (this repository's and the Python guard's): 83/83 patterns correspond, 64 are byte-identical, 6 differ only by escaping and are proven equivalent, and **13 behave differently**, with a witness input for each and misses on both sides. The cause is two different ReDoS mitigations of one table. **No data file was changed** — behaviour preservation holds and the finding is reported to the owning repositories. ### Changed - `lexicon/injection-lexicon.json` **0.3.0 → 0.4.0** — verified against the source module instead of against the dump it was transcribed from, and **two false provenance claims retracted**. The source is now pinned: `b0de0ca` on the public remote, imported in Node and compared entry by entry on `source`, `flags` and `label`. The result is **83/83 byte-identical to source**, which is not what the file previously claimed. It said two patterns had been rewritten from raw code points into `\uXXXX` escapes; the module already writes them escaped, so nothing had been rewritten. The stored pattern text was right the whole time — only the account of where it came from was wrong. The dump had rendered the module's escapes as the characters they denote, and this repository re-escaped them, arriving at the correct bytes by way of an incorrect story. The same inversion ran the other way in `multi-lang:french`, which carried the class spelled with a raw accented Latin `e` where the module writes it as the escape `\u00e9` inside the same character class. That was the one pattern of 83 not byte-identical to source, and it is corrected. The two spellings are the same regular expression — verified in Node bare and under `u`, and in Python `re`, over accented, unaccented, uppercase and non-matching French input, with identical match offsets — so **no behaviour moved**. No pattern in the file contains a non-ASCII byte now, matching the module, whose regex literals are pure ASCII throughout. Structurally: `normalisations` is now `[]` with a `normalisations_note`, matching the convention already used in `signatures/secret-egress.json`, and a new `source_fidelity` block carries the counts, the method, the verified class membership, and both retractions in full. Retracted claims are recorded rather than deleted — the earlier equivalence evidence (692 Node comparisons, 236 Python) remains true, it is simply no longer load-bearing. - `lexicon/injection-lexicon.json` **0.2.0 → 0.3.0** — the two aliases are no longer presented as equally backed. `pattern_id_space.alias_evidence` now records each one separately: `llm_ingestion_guard` is **verified** (the guard's coverage matrix asserts on that exact string, so it is demonstrably what a guard finding carries), while `llm_security` is **not** — it is the pattern table's own name, and the finding producer was never supplied, with the known Node finding shape using `title` rather than `label`. Averaging the two into one file-level claim would have repeated the defect this repository corrects per-table elsewhere. Also: `normalisations[].affects` now keys on `id` with the prose names kept beside it as `affects_labels`. An internal cross-reference on label was a second identity space inside the file the id was added to unify. - `lexicon/injection-lexicon.json` **0.1.0 → 0.2.0** — every pattern gains a commons-owned `id` and an `aliases` object naming what each seeding runtime calls it, plus a top-level `pattern_id_space` block explaining the field. This exists because a `conformance/` fixture has to name a finding and the two runtimes do not name the same pattern the same way. The id was **adopted verbatim from the guard's port**, which already carried both names, rather than invented here. Matching was by `label` ↔ `desc` with em-dash normalised to hyphen: 83/83, one-to-one, ids unique. **No detection data moved.** Labels, patterns and flags are byte-identical in sequence, no `flags` key was invented (78 before, 78 after), and stripping the three new fields reproduces the previous committed file byte for byte — 23 566 bytes, identical. All 83 patterns still compile in Node bare and under `u` (166/166) and in Python `re` (83/83). Neither `llm-security` nor the guard has ratified this id space yet; both were asked by coord on 2026-08-09, and the file says so rather than implying agreement. ### Not included - `signatures/malware-signatures.json` — seed data not yet delivered. - `spec/decode-pipeline.md` — needs the decode implementation. A normative spec inferred from a data dump would be worse than an absent one. - `conformance/` — still absent, with half the blocker cleared. 105 of the guard's 134 coverage cases are convertible to static `input.txt`/`expected.json`; the other 29 assert a runtime's API surface, which this repository does not own. Findings can now be **named** (see `pattern_id_space` above), but 13 patterns still have no agreed **expected behaviour** — the two ports genuinely differ on them — so those fixtures cannot be authored until the owning repositories answer. These are named in the README as planned rather than linked, so nothing in the repository points at a file that does not exist.