{ "version": "0.3.4", "id": "llm-security-commons/conformance", "description": "Enumeration and measurement header for the conformance corpus. Every case directory holds input.txt (the exact bytes to scan) and expected.json (the findings a conforming runtime must produce). The normative reading of those files is spec/conformance-corpus.md; this file records where the cases came from and what was measured.", "$comment": "Fixture files carry no individual version field. The corpus is versioned as a whole, here — a case is added, removed or corrected by bumping this version, and a case-id change is a MAJOR bump because consumers name cases.", "case_id_derivation": { "rule": "case_id = pattern_id with ':' replaced by '__', optionally followed by '--' and a variant slug", "reverse": "pattern_id = case_id truncated at the first '--' if present, then '__' replaced by ':'", "why": "':' is not a legal filename character on Windows, and this repository is fork-and-own. '__' does not occur in the ratified id space, so the transform is one-to-one — verified collision-free across all 90.", "stability": "A case id is a stable identifier. Changing one is a BREAKING change.", "variant_suffix": { "added_in": "0.3.0", "syntax": "'--' + a slug of [a-z0-9-]", "purpose": "To hold a SECOND case for one pattern id when the two inputs answer different questions about the same rule. The first such case is hybrid-xss__script-tag--src-no-close: the original case's input matches the pattern under both its pre-0.7.0 and post-0.7.0 forms, so it cannot gate the difference between them, and the variant input can.", "separator_is_unambiguous": "Measured, not assumed: '--' occurs in none of the 83 ratified pattern ids and in none of the 89 case ids that predate this version. Ids use single hyphens throughout. So the reverse transform stays purely LEXICAL — split at the first '--', no lookup against the id list required — which is the property the original one-to-one rule was protecting.", "constraint": "A variant case MUST be scoped and matched exactly like the base case, and MUST expect the same pattern id. The suffix distinguishes INPUTS, never findings. It is not a licence to record a second, different verdict for one rule.", "supersedes": "This replaces the `one_case_per_pattern_id` note carried through 0.2.0, which read: 'The derivation takes a pattern id and nothing else, so a single-finding scope holds at most one case per pattern id — there is nowhere in the name to put a second.' That was an accurate description of the rule and it cost a real case: see `omitted_payloads`, where the guard's seventh active-content payload was dropped for exactly this reason. The rule is now extended rather than worked around. NOTE that the omitted payload has NOT been added back — extending the derivation makes it expressible, but it was omitted on a second ground as well (its in-scope finding set is identical to the case built from coverage.py:476), and that ground still stands unexamined against the new rule. Adding it is a separate decision, not a consequence of this one." } }, "match_semantics": "exact-within-scope", "scope_covered": [ "lexicon/injection-lexicon.json", "signatures/active-content.json" ], "scope_covered_note": "The two entries do NOT have equal standing, and averaging them would misreport both. `lexicon/injection-lexicon.json` is implemented by both seeding runtimes and its id space is ratified by both. `signatures/active-content.json` is implemented by one; its cases are `not-applicable` for the other under spec/conformance-corpus.md section 1.1, not failures. See that file's `pattern_id_space.single_runtime`.", "scope_planned": { "$comment": "Named rather than faked. Four convertible cases remain in the guard's coverage matrix (3 carrier, 1 secret-egress), and neither table is blocked on effort — each is blocked on a distinct unresolved question, measured 2026-08-10 and recorded below rather than left as 'no agreement yet'.", "codepoints/carriers.json": 3, "signatures/secret-egress.json": 1, "blockers": { "codepoints/carriers.json": "No adoptable id space, and until 0.3.3 the shape of the problem was mis-recorded here. CORRECTED IN 0.3.3: through 0.3.2 this text opened 'The guard emits TWO stage-coupled labels for the same carrier depending on pipeline position — `sanitize:zero-width` (input) versus `output:zero-width-present` (output), and the same split for bidi and unicode-tag'. The split is NOT the same for unicode-tag, and it was not the same on the day that sentence was written: the artifact-side label for tags is `lexicon:unicode-tags-present`, emitted from `lexicon.py`, not an `output:`-prefixed one. Checked at guard commit `e671edb` — the commit this blocker's sibling was measured against — where `coverage.py` already asserted that label, and re-measured at `a59184b`. So the earlier text was wrong when written, not stale. `output.py` states the delegation in a comment: Unicode-tag / PUA stego 'is already surfaced by scan_lexicon (`lexicon:unicode-tags-present`), so it is not repeated'. The labels are therefore not stage-coupled at all. The guard's `Finding` carries a `detector` field alongside `label`, and the prefix is that field's value: `lexicon:unicode-tags-present` carries `detector=\"lexicon\"`, `output:zero-width-present` carries `detector=\"output\"`, `sanitize:*` is emitted by `sanitize.py`. THE PREFIX NAMES THE DETECTOR, not the pipeline stage — and for tags one detector serves both entry points, which is why there is no sixth `output:` label to find. Measured 2026-08-11 at `a59184b`. What remains open is smaller than 'invent a stage-neutral id', and none of it is a naming question. Six labels exist to adopt verbatim, the way the lexicon's 83 ids were adopted from this same runtime: `sanitize:zero-width` / `output:zero-width-present`, `sanitize:bidi-override` / `output:bidi-present`, `sanitize:unicode-tag` / `lexicon:unicode-tags-present`. Three objections stand against adopting them, all measured at llm-security `47905da`: (a) `sanitize:` asserts an ACTION — the carrier was stripped — and llm-security strips nothing. `scanners/unicode-scanner.mjs` exports a single entry point, `scan(targetPath, discovery)`, which reports presence with `scanner: 'UNI'` and a prose `title` and no id at all. Adopting verbatim would hand one runtime an id whose name asserts something it does not do. (b) Three of the six name a persist gate llm-security does not have. The corpus already has the verdict for that — `not-applicable`, spec section 1.1, which attaches to a declared TABLE — and it applies cleanly today, because that runtime's `DECLARED_TABLES` names the lexicon and nothing else (measured at 47905da). WHAT REMOVES IT IS COMMONS PUBLISHING THE ALIAS: their suite derives its registered set by walking each vendored file for any node carrying an `aliases.llm_security` key, then asserts every registered table appears in the declared set — the section 1.1 anti-narrowing floor. The granularity is the FILE, not the entry. So a single carrier id carrying that alias forces `codepoints/carriers.json` into their declared set, obliges them to run all six carrier cases, and converts the three artifact-side ones from not-applicable into failures. Section 1.1 forbids the per-case exit by name, calling it the silent skip section 1 forbids. The choice is therefore between minting only the input-side three, accepting three standing failures that describe an architectural difference rather than a defect, and publishing carrier ids with no llm_security alias at all — a guard-only id space, worth choosing out loud rather than arriving at by omission. (c) The entry point pinned for llm-security in `measurement.runtimes` (`scanForInjection`, the injection scanner) does not reach carriers at all; carrier findings there come from the path-based unicode scanner instead. A carrier case therefore needs a per-scope entry point per runtime, which this manifest expresses nowhere — it pins one entry point per runtime for the whole corpus. Note that both runtimes already consume this file's tables: llm-security builds its zero-width, tag-range and BIDI sets from `codepoints/carriers.json` via `scanners/lib/codepoints.mjs`. The divergence is in what the finding is CALLED and where it can be observed, never in which code points are carriers. STATUS: put to both runtimes as a decision request on 2026-08-11 over coord, carrying exactly the three objections above and asking each the question only it can answer - the guard whether `lexicon:` on the tag path is intentional and whether `sanitize:` names a detector or an action, llm-security whether it adopts commons ids for carriers at all and which of not-implemented-per-runtime or mint-only-the-input-side it wants. Unanswered as of this version. A correction followed the same day, on this repository's own error: the request asserted that the corpus had no third verdict for an unreachable case, when section 1.1's `not-applicable` is exactly that and the registration mechanism above is what would take it away. Nothing is minted until both have ruled, and a further reason not to pre-mint is that publishing a single alias is itself the irreversible act — it moves another runtime's declared set by force of its own test suite.", "signatures/secret-egress.json": "Not an id-naming question at all. The two runtimes carry DIFFERENT TABLES, not two namings of one: this file holds 19 entries from llm-security, the guard's `_SECRET_PATTERNS` (output.py) holds 25 at different cut points — this file's single `GitHub Token` is four ids there, `Private Key PEM Block` is three, `Database connection string` is four. `aws-access-key-id` is the one clean one-to-one, which is why exactly one egress case was ever offered. A shared id space presupposes a table reconciliation that has not happened. Membership diverges both ways, and `Slack/Discord Webhook URL` (order 13) and `Azure AI Services Key` (order 4) are here and absent from the guard. CORRECTED IN 0.3.1: through 0.3.0 this text ended '(the guard has `gcp-service-account-json` and `openai-api-key-legacy`, which are absent here; ...)'. Two entries were named as one kind of fact, and they are two different kinds. Measured here on 2026-08-11 by running this file's own patterns, in `order`, over a service-account document, against the guard at commit `e671edb` — 18 patterns then, 19 since `secret-egress.json` 0.3.0. (1) `gcp-service-account-json` is NOT a coverage hole here. A complete service-account key file matches this table at order 11, `Private Key PEM Block`: that pattern's prefix group `(?:RSA |EC |DSA |OPENSSH )?` is optional, so the bare PKCS#8 header `-----BEGIN PRIVATE KEY-----` such a file carries is matched. What differs is the CUT POINT: the same document with its `private_key` field removed matches nothing here, while the guard's `\"type\"\\s*:\\s*\"service_account\"` still fires — the guard detects the document MARKER, this table detects the KEY MATERIAL. Both fire on a real key file; only the guard fires on a stripped one. That is exactly the cut-point divergence this blocker is about, and it is not a missing entry. (2) `openai-api-key-legacy` WAS a real hole here and is CLOSED IN 0.3.2. Through `secret-egress.json` 0.2.0 no pattern in this file matched a legacy `sk-…T3BlbkFJ…` key, measured the same way, and llm-security's 18 -> 19 report stood recorded as their unreproduced claim because the commit carrying it was not yet on their public remote. It is now: `refs/heads/main` reads `47905da`, `088e458` is an ancestor of it (checked with `git merge-base --is-ancestor`, not read off their log), and `secret-egress.json` 0.3.0 carries the entry re-extracted from the module text at that commit, with all 19 positions compared for name, source, flags and order. So the two tables now agree on this one shape. What has NOT changed is the reason this blocker exists: 19 against 25 at different cut points is still a table reconciliation nobody has performed, and one closed hole is not that reconciliation." } }, "omitted_payloads": [ { "source": "llm-ingestion-pipeline-security src/llm_ingestion_guard/coverage.py:484, at commit de09711 (line numbers are commit-relative; the structure is the fifth `_scan_case` of `_build_cases()`'s `active` group)", "described_as": "opaque (base64) path segment", "expected_label": "active:markdown-image", "reason": "Its in-scope finding set is `[active:markdown-image]` — identical, measured, to the case built from coverage.py:476. The only thing that distinguishes it is `entropy:base64-blob`, and this repository publishes no entropy table, so the difference falls outside every declared scope. A second case could not have failed in any way the first does not, and the case-id derivation had no room for it.", "$comment": "Recorded so that 6 built from 7 offered reads as a decision rather than as a miscount.", "derivation_ground_withdrawn_in_0_3_0": "The second half of `reason` — that the derivation has no room for a second case per pattern id — stopped being true in 0.3.0, when `case_id_derivation.variant_suffix` was added. It is left in the text above with this note rather than silently deleted, because the omission decision was taken on TWO grounds and only this one lapsed. The first ground stands: the payload's in-scope finding set is identical to the case already built, so it still could not fail in any way that case does not. The payload therefore remains omitted, on one ground instead of two. Re-examining it is a separate decision." } ], "authored_payloads": [ { "case_id": "hybrid-xss__script-tag--src-no-close", "added_in": "0.3.0", "added_date": "2026-08-11", "input": "