Section 1 required every case to be run and every unrunnable one reported as an
error. Scoping a case to signatures/active-content.json would therefore have made
llm-security permanently fail seven cases for having no such table - reporting an
architectural difference as a defect, and telling a reader nothing.
Section 1.1: a runtime declares which commons data files it implements, and a case
scoped outside that set is `not-applicable` - a third verdict, distinct from section
1's error. Section 1's error means the runtime tried and could not; this means the
question was never addressed to it.
Fenced so it cannot become an exit. It attaches to a TABLE, never to a case, since
per-case opt-out is exactly the silent skip section 1 forbids. A declared set may not
be narrowed to convert failures into not-applicable ones. Such cases stay in the
denominator: `76/83` and `76 passed, 7 not-applicable` describe different runtimes,
and only the second can be checked.
Three consequences, written where they are read:
- Section 4 no longer claims scoping "asks a question both can answer". That held
only while every case was scoped to the one table both runtimes implement. Scope
narrows what is compared; it does not make every runtime a valid addressee. The
superseded sentence is named in place rather than edited away.
- Section 4 now states that "belongs to a data file" means published there, never
"shares its prefix". Live witness: the guard emits `active:oversize-input`, a flag
about its own scan cap, which carries the prefix but is no construct in the table.
A prefix-matching runtime would fail a case over a finding the corpus never claimed.
- Section 6 states the derivation's cost: case_id derives from pattern_id alone, so a
single-finding scope holds at most one case per pattern id. When a source runtime
drives two payloads at one pattern, they must be compared within scope before a
second case is minted - and a discriminated case id is forbidden, since it would
break the reverse transform.
Section 8: a pass count is unreadable without the declared set beside it, and a
not-applicable verdict proves nothing about detection in either direction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVouC9nsfrfV5jRSejxbvQ
schema/finding.schema.json defines a finding `id` as DS-<scanner>-<counter>,
built from a process-global counter: stable across neither runs nor
processes, and the schema says so itself. The corpus keys its comparison on
the lexicon's stable rule identity. Two normative documents in one
repository using one word for both would produce runtimes failing every case
for a reason unrelated to detection.
Also adds spec section 3.1, which publishes the bridge a consumer actually
needs and which neither normative document named: a runtime's own label
reaches a pattern_id through the lexicon's `aliases` object, and a runtime
absent from that object has no published way to be compared -- a mapping
kept privately in a consumer is the drift this repository exists to prevent.
Regenerated all 83 fixtures; re-verified from the corpus alone against both
runtimes, 83 cases, 0 failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WhXDL82FRrQWedEmUg12Pj
The corpus was blocked on a decision nobody had to take. The 13 divergent
lexicon patterns were measured on witness inputs -- an attribute run padded
past 256 characters, an interior '<', an unclosed <script> -- and the corpus
payloads contain none of those shapes. Run through both runtimes' public
entry points, all 83 patterns produce identical lexicon finding sets, the 13
included. No winner picked, because the question was never reachable from
these inputs.
Measured at each runtime's entry point (scanForInjection() at b0de0ca,
scan_output(source=OUTPUT) at 0bf0729), never at a rebuilt regex table --
the layer mistake the divergence document already had to retract once.
- conformance/<case>/{input.txt,expected.json} x83, plus manifest.json
- spec/conformance-corpus.md, normative: bytes not text, id-only findings,
exact-within-scope, and observed_out_of_scope as evidence not expectation
Verified by a harness that does not share the generator's knowledge: reads
only the case directories, re-runs both runtimes, checks every field
including digests -- 83 cases, 0 failures. Severity agrees with what commons
publishes 83/83. Deleting the middle third of each input breaks 76 of 83
expectations; the 7 survivors are the shortest payloads, a weak mutation
rather than a weak fixture, and are recorded as such.
Scope is 83 and not 94 for a different reason than expected: the 11
non-lexicon convertible cases have no ratified cross-runtime finding id, and
writing them would mint a contract unilaterally in the same stroke as the
tag. Named in manifest.json under scope_planned.
Docs corrected in place rather than edited away: extraction-plan and
lexicon-port-divergence both claimed the 13 blocked the corpus.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WhXDL82FRrQWedEmUg12Pj