{ "version": "0.2.0", "id": "llm-security-commons/conformance", "description": "Enumeration and measurement header for the conformance corpus. Every case directory holds input.txt (the exact bytes to scan) and expected.json (the findings a conforming runtime must produce). The normative reading of those files is spec/conformance-corpus.md; this file records where the cases came from and what was measured.", "$comment": "Fixture files carry no individual version field. The corpus is versioned as a whole, here — a case is added, removed or corrected by bumping this version, and a case-id change is a MAJOR bump because consumers name cases.", "case_id_derivation": { "rule": "case_id = pattern_id with ':' replaced by '__'", "reverse": "pattern_id = case_id with '__' replaced by ':'", "why": "':' is not a legal filename character on Windows, and this repository is fork-and-own. '__' does not occur in the ratified id space, so the transform is one-to-one — verified collision-free across all 89.", "stability": "A case id is a stable identifier. Changing one is a BREAKING change.", "one_case_per_pattern_id": "The derivation takes a pattern id and nothing else, so a single-finding scope holds at most one case per pattern id — there is nowhere in the name to put a second. See spec/conformance-corpus.md section 6, and `omitted_payloads` below for the one payload this actually cost." }, "match_semantics": "exact-within-scope", "scope_covered": [ "lexicon/injection-lexicon.json", "signatures/active-content.json" ], "scope_covered_note": "The two entries do NOT have equal standing, and averaging them would misreport both. `lexicon/injection-lexicon.json` is implemented by both seeding runtimes and its id space is ratified by both. `signatures/active-content.json` is implemented by one; its cases are `not-applicable` for the other under spec/conformance-corpus.md section 1.1, not failures. See that file's `pattern_id_space.single_runtime`.", "scope_planned": { "$comment": "Named rather than faked. Four convertible cases remain in the guard's coverage matrix (3 carrier, 1 secret-egress), and neither table is blocked on effort — each is blocked on a distinct unresolved question, measured 2026-08-10 and recorded below rather than left as 'no agreement yet'.", "codepoints/carriers.json": 3, "signatures/secret-egress.json": 1, "blockers": { "codepoints/carriers.json": "No adoptable id space, and a second problem underneath it. The guard emits TWO stage-coupled labels for the same carrier depending on pipeline position — `sanitize:zero-width` (input) versus `output:zero-width-present` (output), and the same split for bidi and unicode-tag; llm-security emits prose titles (unicode-scanner.mjs:191,236). A commons id would therefore have to be invented stage-neutral, which no other id space here required. And because `exact-within-scope` compares a finding SET, a commons id aliasing both guard labels would make the verdict depend on which entry point the runtime was measured through — an entry-point dependence the lexicon cases do not have, since this manifest pins entry point as a measurement fact rather than as contract.", "signatures/secret-egress.json": "Not an id-naming question at all. The two runtimes carry DIFFERENT TABLES, not two namings of one: this file holds 18 entries from llm-security, the guard's `_SECRET_PATTERNS` (output.py) holds 25 at different cut points — this file's single `GitHub Token` is four ids there, `Private Key PEM Block` is three, `Database connection string` is four — and membership diverges both ways (the guard has `gcp-service-account-json` and `openai-api-key-legacy`, which are absent here; this file has `Slack/Discord Webhook URL` and `Azure AI Services Key`, which are absent there). `aws-access-key-id` is the one clean one-to-one, which is why exactly one egress case was ever offered. A shared id space presupposes a table reconciliation that has not happened." } }, "omitted_payloads": [ { "source": "llm-ingestion-pipeline-security src/llm_ingestion_guard/coverage.py:484, at commit de09711 (line numbers are commit-relative; the structure is the fifth `_scan_case` of `_build_cases()`'s `active` group)", "described_as": "opaque (base64) path segment", "expected_label": "active:markdown-image", "reason": "Its in-scope finding set is `[active:markdown-image]` — identical, measured, to the case built from coverage.py:476. The only thing that distinguishes it is `entropy:base64-blob`, and this repository publishes no entropy table, so the difference falls outside every declared scope. A second case could not have failed in any way the first does not, and the case-id derivation has no room for it (see `case_id_derivation.one_case_per_pattern_id`).", "$comment": "Recorded so that 6 built from 7 offered reads as a decision rather than as a miscount." } ], "payload_provenance": { "$scope": "The 83 lexicon cases. The 6 active-content cases have their own provenance in `active_content_provenance` below — they come from a different structure in the same file, at a different commit, and folding them in here would let one pin stand for two measurements.", "source_repo": "llm-ingestion-pipeline-security", "source_file": "src/llm_ingestion_guard/coverage.py", "source_export": "_LEX_PAYLOADS", "source_commit": "0bf07295c2191d5061537834abf22929f7d50826", "source_version": "0.3.4", "$comment": "The inputs were authored by one of the two runtimes, as one payload per pattern id, and are reproduced verbatim. That asymmetry is stated rather than averaged away: what makes them usable as a cross-runtime corpus is not their origin but the measurement below, which ran them through the other runtime as well and found the same lexicon verdict on every one.", "id_set_check": "The 83 payload keys and the 83 commons pattern ids are the same set — compared, not assumed.", "stability_note": "llm-ingestion-pipeline-security states (coord message 2026-08-10T12:42:31Z) that _LEX_PAYLOADS and the 83 pattern ids are an internal surface on their side, with no README/CHANGELOG/docs statement promising id or payload stability - their own test suite enforces id coverage as their gate, not as a promise to this repository. Their stated position: if a payload changes upstream, this manifest's pin diverges and should be re-pinned; divergence is a re-pin signal, not a breach of a contract they never granted." }, "measurement": { "$scope": "The 83 lexicon cases. The 6 active-content cases are measured in `active_content_measurement`.", "date": "2026-08-10", "method": "Each payload was run through both runtimes' PUBLIC entry point — not through a rebuilt regex table — and the resulting finding labels were mapped to commons pattern ids through the lexicon's own aliases block. Comparing at the entry point is deliberate: a table-level comparison produces a number that describes neither runtime.", "runtimes": [ { "name": "llm_security", "repo": "ssh://git@git.fromaitochitta.com/open/llm-security.git", "commit": "b0de0ca6d86ce697f39669d177c2c2654c280128", "entry_point": "scanForInjection(text) — scanners/lib/injection-patterns.mjs", "engine": "Node v25.8.2", "covers": "normalisation, homoglyph folding, the rot13 variant and all four pattern arrays", "measurement_limit": "This entry point is the injection scanner. Whether this runtime raises findings from other commons tables on these inputs was NOT measured, so observed_out_of_scope carries no entry for it — absence of a key means unmeasured, not measured-empty." }, { "name": "llm_ingestion_guard", "repo": "llm-ingestion-pipeline-security", "commit": "0bf07295c2191d5061537834abf22929f7d50826", "version": "0.3.4", "entry_point": "scan_output(text, source=Source.OUTPUT)", "engine": "CPython 3.14.0", "covers": "the whole output gate, which is more than the lexicon", "measurement_limit": "Non-lexicon findings this gate raised are recorded per case in observed_out_of_scope as informative evidence, never as expectation." } ], "results": { "cases": 83, "asserted_id_present_in_both_runtimes": 83, "lexicon_id_sets_identical_between_runtimes": 83, "cases_with_non_lexicon_residue_in_the_guard": 7 }, "known_divergence": { "$comment": "13 of the 83 patterns are recorded in docs/lexicon-port-divergence.md as behaviourally divergent between the two ports. That divergence is real and unresolved, and it is NOT visible here: it was measured on witness inputs — an attribute run padded past 256 characters, an interior '<', an unclosed