Commit graph

16 commits

Author SHA1 Message Date
f0a511369d fix(spec): section 7 stated its own premise and then applied itself beyond it
Section 7 justified the fixture-is-ground-truth ordering with "Two
implementations that return different verdicts" and then stated the rule
with no scope. For signatures/active-content.json there is no second
implementation, and the seed runtime has stated the classification behind
it is calibration it does not freeze. As written, section 7 turned a change
they reserved into a bug on their side.

New section 7.1 keys the scope on a structural property, never on a table
name: a scope only one runtime implements, whose payload that runtime
authored. It creates no fourth verdict - the declaration schema closes
result with additionalProperties:false over four counts, so a fifth would
break every consumer's parser. The case still fails and is still named in
failed_cases; what changes is what the failure licenses concluding.

Two limits are stated rather than left to inference: it does not reach a
third-party implementer of the same table, and it is not a licence for a
runtime to self-declare its own divergence as calibration.

manifest.json 0.6.1 -> 0.6.2 retires the open-question sentence, quoted
rather than dropped. The retirement is partial: "section 7 is NOT amended
by this block" stays true, because the spec was amended by its own release.

Neighbours measured over the whole repository, widened past "ground truth"
to the second paragraph's own wording. CONVENTIONS.md and CLAUDE.md carried
the premise and are changed; SECURITY.md gets a cross-reference only, since
its claim is about a fixture expecting too little and 7.1 does not narrow
that direction; README.md and docs/extraction-plan.md are named as
deliberately untouched.

Breaking in category, minor in number - 0.x, per the reading [0.3.0]
recorded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012pZ2FLQ6xkWvj2VcwgwnQv
2026-08-13 23:31:28 +02:00
c75c546614 docs(manifest): the precision field stated its exposure as a hand-derived count
0.8.0 added a field whose whole purpose is precision, and bounded the exposure
with "three of the four dimensions ... cannot move one of these cases at all".
The total was derived by hand over a taxonomy the field had itself
recategorized. Lexicon entries ARE one of the seed runtime's four calibration
dimensions, and they are not absent from these fixtures: two of the seven carry
a lexicon id in observed_out_of_scope. And "the fourth" substituted the
classification for lexicon entries as the fourth item of their sentence.

manifest 0.6.0 -> 0.6.1: the field enumerates three named things and totals
none of them. A lexicon change ages the residue as evidence without moving a
verdict (spec section 5); the classification is the one that can move one.
Retired sentence quoted in an AMENDED IN clause, not dropped.

No measurement changed, no verdict moved, no case or data file touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016yZobrgUiRtpLSWx8i7u2Z
2026-08-13 23:15:10 +02:00
c23aea9062 docs(manifest): the active-content fixtures pin a version, and nothing said so
The seed runtime's v1.0.0 freezes its exported Python surface and explicitly
not its detection behaviour. The manifest pinned commit and version per
measurement block but never recorded that the thing pinned is a version rather
than a frozen classification.

manifest 0.5.2 -> 0.6.0, one new field next to active_content_provenance.
asymmetry, bounded by what the fixtures actually assert: all seven carry
pattern_id only, so three of the four calibration dimensions cannot move them.
Both pins named, not one. Their statement is attributed, not restated as ours.

Also closes the omission 0.7.3 named: the same misquote in
docs/secret-egress-divergence.md:75-76. Not the fix the note implied - those
lines are one single-backtick span across a line break, so the outer delimiter
is promoted to double backticks instead.

spec section 7 deliberately untouched and named in CHANGELOG.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016yZobrgUiRtpLSWx8i7u2Z
2026-08-13 23:09:33 +02:00
757570dd49 fix(readme): the paragraph still argued the premise our own commit retired
README.md opened the egress gap with "It is not an id question at all -- the two
runtimes carry different tables, cut at different granularities", and sent the
reader to scope_planned.blockers as the authority. Since 4356caa that blocker
opens reason (1) with "NO ID SPACE ON THE COMMONS SIDE. This is the hard
blocker." Before 0.7.2 the README was out of date; after it, two files on a
public remote disagreed, and it was our commit that made them.

The paragraph now carries the three measured, independent reasons from
docs/secret-egress-divergence.md. The counts survived the falsification, so
"different tables" and 19-against-25 are kept; only "cut at different
granularities" and the pending-reconciliation claim are gone. It deliberately
omits the outgoing question's status (true on the day written, untested by
anything, and dated in the manifest) and the standing entry_points_by_scope
requirement (a requirement, not a fourth reason).

conformance/manifest.json 0.5.1 -> 0.5.2 in the same release: the blocker
misquoted the contract it cites, rendering the value without the backticks the
field carries around `order`. The data file is unchanged and was never wrong.
Round-trip proved byte-neutrality before the edit; the read-back matched the
decoded string against the data file that owns it, since json.tool passes on
wrong escaping.

No case minted, no data file touched, no id proposed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SZ5vrpu2kxcRktiW59b7s3
2026-08-13 22:58:29 +02:00
4356caa689 fix(manifest): the egress blocker named a premise the measurement falsified
`scope_planned.blockers` for secret egress read "19 entries … 25 at different cut
points" — one table cut at two granularities, waiting on a reconciliation of two
ports. `docs/secret-egress-divergence.md` (2d9ee9c, corrected in 4a6f6ff) measured
otherwise: they are ports of two different source tables in the same source
repository, so reconciling the ports was never going to close it.

Replaced with the three measured, independent reasons: no id space on the commons
side (seed A carries name+pattern only, and a fixture names labels); match
semantics disagree (first-match-wins against finditer over all 25 — one label
against two on the Bearer+JWT witness); membership diverges both ways and is
inherited from two different seeds, so re-measuring either port cannot close it.
Only the first is the outgoing question; the other two stand whatever the answer.

Two hand-carried numbers in the retired text corrected in place: aws-access-key-id
was not the one clean 1:1 (2/19 byte-identical, AWS not among them — the guard
anchors with \b), and `GitHub Token` maps to three guard ids, not four, leaving
ghu_ and ghr_ uncovered. `scope_planned.$comment` said "a distinct unresolved
question" — singular — and is amended alongside.

Measured here, not transcribed: `entry_points_by_scope.scopes` carries no entry for
signatures/secret-egress.json at all. Recorded as a standing requirement, not as a
fourth reason. No id string is proposed; checked against both outgoing coord
messages of 2026-08-13. No case, no signature table and no other data file touched.

conformance/manifest.json 0.5.0 -> 0.5.1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016inh17NCrQpWN3mrghfJgT
2026-08-13 22:46:15 +02:00
e6ca5ae5ee feat(active-content): the seventh case, and the classifier it needed came with it
Adopting `active:raw-html-link` was one id. Publishing it honestly was the whole
of `active_tag_class` — one function, three branches, no way to state the split
without the no-URL narrowing and the 0.6.0 external-target rule. On the old
predicate a bare `</a>` is active by name, so a consumer implementing from the
hybrid would emit the new label where the seed runtime emits nothing.

The file is now two pins, stated as two: v0.3.4/0bf0729 everywhere except the
raw-HTML classifier, v0.7.0/be9759b there. The drift between them was measured
field by field against the imported module rather than assumed, after stripping
inline-flag rendering and applying the file's own declared quote normalisation
so a spelling difference could not masquerade as drift. Exactly one published
field had moved, and not the one this release was about: `html.active_tags`
carried the MUTATOR's 23-name set where the gate means the SCANNER's 22. Correct
at the 0.3.4 pin, wrong from 0.6.0 on. Kept as `html.mutator_tags`.

The sweep covered 93 cases, not the 6 obvious ones. The narrowing can silence an
`active:` finding inside the `observed_out_of_scope` evidence of a LEXICON case,
and that field is guarded by no test anywhere — stale entries there survive
forever. One case moved: html-obfuscation__aria-label, whose `<a aria-label=…>`
carries no URL attribute. Its fixture is deliberately not rewritten; the residue
is true at the commit `measurement` pins, and rewriting one of 83 would leave two
commits under a header naming one. Recorded, dated and pinned in the manifest.

The strongest check is not the digest: the checker rebuilds the published
classifier from the JSON alone, importing nothing from the runtime, and
differential-tests it against `active_tag_class` over 42 probe tags. 0
disagreements. That is what licenses shipping a classifier as data.

No `aliases.llm_security` published, on this file or on carriers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LTTaT4quwNPwBYVqmgAt8t
2026-08-13 21:43:27 +02:00
8f08e9ac73 feat(carriers): three cases minted, and the id is named rather than adopted
Both runtimes answered the 2026-08-11 decision request, and they did not ask
for the same name. The guard ruled `sanitize:` names its `Finding.detector`
and offered all three labels for verbatim adoption. llm-security, asked
directly, answered that the name would make its conformance result read as a
claim about neutralisation it does not perform.

Two things decided it. The guard's own unprompted correction: prefix ==
detector holds for those six labels and is no general law in its runtime
(`egress:*` carries detector="output"; decode-and-rescan yields two-part
`decoded:lexicon:*`). A prefix whose meaning is recoverable only by reading
one implementation cannot carry a shared id space. And a measurement taken
here at be9759b: on the surface the guard's own ruling pinned, `sanitize()`
returns changed text on all three carriers, so the counterargument's decisive
case -- that `scan_output` mutates nothing -- does not reach this surface.

Not a mediation. Neither runtime claimed the shared id must equal its label,
and `override:ignore-previous` already carries two different alias strings.

- carriers.json 0.1.0 -> 0.2.0: carrier:zero-width / :bidi-override /
  :unicode-tag, aliased to the guard's labels. No aliases.llm_security --
  that runtime's carrier findings carry no id yet, and publishing the alias
  is the irreversible act that forces the table into its declared set.
- manifest 0.3.4 -> 0.4.0: entry_points_by_scope, carrying findings accessor
  and fixture presentation per scope per runtime. This was objection (c), and
  it blocked minting harder than the name did.
- Corpus 90 -> 93. Measured through sanitize(text, source=Source.INPUT) at
  guard v0.7.0; verified by a separate checker that re-derives everything from
  disk -- a generator agreeing with itself proves nothing.
- CLAUDE.md gains the two rules that are not derivable from the data: a shared
  id space cannot rest on a one-runtime prefix, and publishing an alias -- not
  minting the case -- is the irreversible act.

Not minted on purpose: no artifact-side id (the other runtime would only fail
them), and no ZWJ-exemption case (U+200D between two emoji is exempt on both
guard surfaces since v0.6.1; the fixture avoids it rather than trips it).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U3o4zSQ2kwLsgJNU7apK3Z
2026-08-13 21:17:44 +02:00
7ce0ba706c docs(carriers): the third verdict exists, and publishing an alias is what takes it away
The carrier decision request went to both runtimes today. Then re-reading our own
spec turned up an error in it: the request said the corpus has no verdict for a
case a runtime cannot reach. It has one, and this repository wrote it — §1.1
`not-applicable`, attaching to a declared table. A correction went to both.

What the correction found is sharper than the mistake. llm-security declares the
lexicon table alone (`DECLARED_TABLES`, measured at 47905da), so carrier cases
are not-applicable there today. Their suite derives the registered set by walking
each vendored file for any node carrying `aliases.llm_security`, then asserts
every registered table is declared — the §1.1 anti-narrowing floor. Granularity
is the file. So one aliased carrier id would force `codepoints/carriers.json`
into their declared set, oblige them to run all six carrier cases, and convert
the three artifact-side ones into failures. Publishing the alias is the
irreversible act, not minting the case.

manifest.json 0.3.3 -> 0.3.4; README's carrier paragraph carried the same
superseded "labels differ by pipeline stage" claim and now carries the measured
one. No case, no id and no expected.json moved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F1HPHHP1zDNk1tMKDPuLCC
2026-08-11 22:34:00 +02:00
302625ead5 fix(conformance): the tag carrier has no output: label, and our blocker claimed it did
The carriers blocker described the guard as emitting two stage-coupled labels
per carrier, "the same split for bidi and unicode-tag". Measured at guard
`a59184b`: the artifact-side label for tags is `lexicon:unicode-tags-present`,
emitted from lexicon.py, and `output.py` says in a comment that it deliberately
does not repeat it. Checked at `e671edb` too — coverage.py asserted that label
there as well, so the sentence was wrong when written, not stale.

The correction moves the blocker rather than shrinking it. The guard's `Finding`
carries a `detector` field, and the label prefix is that field's value, so the
prefix names the detector and there was never a stage to be neutral about. Six
labels exist to adopt verbatim. What blocks adoption is measured and named
instead: `sanitize:` asserts a strip that llm-security does not perform, three
of the six name a persist gate it does not have, and the entry point pinned for
it in `measurement.runtimes` does not reach carriers at all.

manifest.json 0.3.2 -> 0.3.3. No case, no id and no expected.json moved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F1HPHHP1zDNk1tMKDPuLCC
2026-08-11 22:25:27 +02:00
d467324380 feat(signatures): the staleness we disclosed is closed by reading the module, not the message
secret-egress.json 0.2.0 -> 0.3.0. `OpenAI Legacy API Key` enters at order 17,
second to last; JWT stays last because ordering.last_entry_is_load_bearing says
it must. 18 -> 19.

The regex was in the coord message that reported it. That is the path
evidence_limits explicitly ruled out, so it was read out of the module text at
a pinned public commit instead: refs/heads/main = 47905da, and 088e458 (which
carries the entry) confirmed an ancestor with `git merge-base --is-ancestor`
rather than accepted from their log.

The entry is the smaller half. All 19 positions were compared against the
module - name, source, flags, order - with 0 divergences, so positions 0-16 are
no longer resting on a 2026-08-09 transcription whose module fidelity stood
recorded as llm-security's assertion. It is reproduced now, and both the
fidelity bullet and the staleness bullet retire.

manifest.json 0.3.2: the blocker prose promised its note would stand until this
landed. Item (2) is marked closed and the count moves 18 -> 19. The blocker
itself does NOT close - 19 against the guard's 25 at different cut points is a
table reconciliation nobody has performed, and one closed hole is not that.

Verified: JSON well-formed, orders contiguous 0..18, count matches array length,
new pattern compiles in Node bare, Node `u` and Python `re`, and does not match
sk-ant-/sk-proj- shapes. Charter clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JLEZ4XCSnSrQUFA8SzkQB4
2026-08-11 21:52:24 +02:00
0e765a02eb docs(security): the attack surface here is data, so the report route had to say where a wrong entry gets fixed
org-ops recorded SECURITY.md as missing against the org standard (coord,
2026-08-11) and this repository owed it for a sharper reason than "given what
the repo is about": nothing here runs, so a report is never a crash — it is a
detection entry that looks like it works and is not looking.

SECURITY.md therefore answers what an ordinary policy does not have to: how to
report that a detection-table entry is WRONG, and why a confirmed defect in
extracted data is decided in the runtime it was extracted from before it is
changed here. Correcting it here would make the copy disagree with the
implementation it was taken from — two runtimes, two answers on one input, the
exact failure this repository exists to prevent. Two classes skip that routing:
a real secret in the history, and data authored here rather than extracted.
Fix latency is stated plainly as bounded by the owning runtime's schedule and
the consumer's pull, not by ours.

secret-egress 0.1.0 -> 0.2.0 is a staleness DISCLOSURE, not a data change: all
18 patterns byte-identical, one evidence_limits entry added. llm-security
reports the source table at 19 entries now; recorded as their report and not
reproduced, because the commit carrying it is not on their public remote —
measured at b1ba1fb today. What was measured here: none of the 18 patterns
matches a legacy sk-...T3BlbkFJ... shape. A consumer vendoring this file
under-matches the seed hook by one entry, and now reads that in the file.

manifest 0.3.0 -> 0.3.1 corrects the secret-egress blocker. Through 0.3.0 it
named gcp-service-account-json and openai-api-key-legacy together as ids
"absent here". Measured against the guard at e671edb by running this file's own
18 patterns over a service-account document: a COMPLETE service-account key
file is matched here at order 11, since the PEM entry's prefix group is
optional and the bare PKCS#8 header matches; the same document with private_key
removed matches nothing here while the guard's marker still fires. That is a
cut-point difference, which is what the blocker is about, not a missing entry.
openai-api-key-legacy IS a real hole and is now recorded as one. Folded into
the existing blocker string rather than a sibling key, because blockers is a
map from table path to text.

Verified: all JSON well-formed; every non-fixture JSON has a top-level version;
charter clean (no executable code); patterns[] and count byte-identical to HEAD
for secret-egress; manifest key set unchanged and count still 90; 90 case
directories untouched; every spec still carries its normative marker.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T85QqeiWBEoBMjWaMiBnxD
2026-08-11 14:03:24 +02:00
25a2cf9643 feat(conformance): the witness case, and the derivation rule that had no room for it
conformance/hybrid-xss__script-tag--src-no-close/ — input `<script src=x.js>`,
17 bytes, expecting hybrid-xss:script-tag. This is the regression gate for the
convergence in the previous commit, and the reason the corpus could not see that
change coming: the existing hybrid-xss__script-tag input `<script>steal()</script>`
matches the pattern under BOTH forms, so it passes either way.

Mutation-verified in both directions across all 90 cases: reverting the pattern to
its 0.6.0 form fails this case and only this case.

The case-id derivation blocked it, and the fix is an extension rather than a
workaround. case_id_derivation gains an optional `--<variant>` suffix; the reverse
transform truncates at the first `--` then maps `__` to `:`. `--` was measured
absent from all 83 ratified pattern ids and all 89 pre-existing case ids, so the
reverse transform stays purely lexical — no lookup against the id list — which is
the property the original one-to-one rule was protecting. No existing case id
moves, so this is additive.

one_case_per_pattern_id is removed, superseded by variant_suffix.supersedes, which
quotes its text. It documented the constraint rather than carrying data a consumer
matches on, but a removed key is normally breaking here, so it is called out.

That same rule cost a real case: omitted_payloads gains
derivation_ground_withdrawn_in_0_3_0. The guard's seventh active-content payload
was omitted on TWO grounds and this change retires one. The other stands — its
in-scope finding set is identical to a case already built — so the payload stays
omitted, on one ground instead of two. It is NOT added back; that is a separate
decision, not a consequence of this one.

First case input authored in this repository rather than reproduced verbatim from
a runtime's payload set, so it goes in a new authored_payloads block instead of
being folded into payload_provenance, whose value is exactly the claim that its
inputs are verbatim upstream. That claim stays as strong as it was: 83 of 83.

Both witnesses for this axis were named by llm-security on 2026-08-10; this is the
first of the two. It was declined that day on the ground that a fixture encoding a
DISAGREEMENT is worse than an absent one — it then contradicted commons' own
published lexicon. Lexicon 0.7.0 removed the contradiction. The stated order was
"settle the row, then the case is trivial to add".

Findings measured through the guard's public API (scan_lexicon,
scan_active_content) at 0dce50f / 0.5.0 — not read off the regex. The same harness
reproduced the existing case's committed bytes and sha256 in the same run as a
control, which is what licenses trusting its output for the new one. Digest
independently recomputed with shasum over the file on disk: agrees.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuudQLGMyMenuyeZk8fKs5
2026-08-11 13:39:19 +02:00
946f51d35e fix(active-content,conformance): cite line numbers per commit — they do not resolve at the pin
pattern_id_space.verified cited active_content.py:337-369 and :309. Those resolve at
de09711, where the check was run; this file's provenance pins 0bf0729, where the same
six call sites are at 316-348 and the emitter at 288. The 23-line scan-cap insert
shifts everything below it by 21, so a reader following the pin landed on the wrong
lines - and on lines that look plausible rather than obviously wrong.

Same defect class as the at_commit_note corrected before the first commit, one layer
deeper: a measured fact stated without the coordinate it is true in. Both commits'
numbers are now given, plus the symbol names, which are stable across the diff and
are what a reader should actually match on.

omitted_payloads[0].source gets the same treatment - coverage.py:484 is de09711-
relative, and the structural description now carries the load instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVouC9nsfrfV5jRSejxbvQ
2026-08-10 21:16:55 +02:00
bdcb1f1080 feat(conformance): ship the six active-content cases; the id space already existed
The corpus goes 83 -> 89 and scope_covered gains signatures/active-content.json.

The blocker in STATE dissolved under measurement, the same way last session's
13-pattern one did. "An id space for carriers/active-content/secret-egress" was one
question in name only; the three tables have three unrelated problems:

- active-content needed NO id space invented. label_format ("active:{class}") and the
  constructs keys were already extracted verbatim from the seed runtime, and their
  concatenation IS what it emits - verified by comparing the six keys to the six class
  strings at its _flag call sites. What blocked these cases was never naming; it was
  spec section 1, fixed in the parent commit.
- carriers has no adoptable id space AND an entry-point dependence underneath it.
- secret-egress is not an id question at all: the two runtimes carry DIFFERENT tables.
  18 entries here against the guard's 25, cut at different granularities - this file's
  single `GitHub Token` is four ids there, `Private Key PEM Block` three, `Database
  connection string` four - with membership diverging both ways. `aws-access-key-id`
  is the one clean 1:1, which is why exactly one egress case was ever offered. That
  number was a symptom, not modesty.

Both blockers are now recorded under scope_planned.blockers, measured, replacing a
blanket "no runtime has agreed to an id space" that was wrong for both.

Generated from measurement, not written. Payloads were extracted from the seed
runtime's coverage.py by AST - evaluating each _scan_case argument in that module's
own namespace rather than retyping detection data - then run through its public
output gate, the same entry point the 83 lexicon cases used. The fixtures were then
re-read from disk by a separate checker that re-computed every digest, re-scanned the
bytes and applied exact-within-scope independently of the generator: 6 cases, 0
failed checks.

Six built from seven offered. The runtime's matrix drives two payloads at
`active:markdown-image`; measured, their in-scope finding sets are identical, and the
second's only distinguishing signal (entropy:base64-blob) falls outside every table
this repository publishes. Dropped rather than given a discriminated case id, and
named under omitted_payloads so the count reads as a decision.

These six prove LESS than the 83, and the manifest says so: their payloads come from
the only runtime implementing the table, so no second implementation's agreement
could be measured. They pin one runtime's behaviour as a contract a future
implementer can be held to - less than cross-runtime agreement, more than nothing.

llm-security's absence of the table is measured at b0de0ca, not assumed: a tree-wide
search finds no implementation, and `git log -S` over --all returns zero commits,
closing the "it was there once" reading. Absent table is not absent capability -
their entropy scanner reaches markdown-image URLs by another route - and the manifest
says that too.

Provenance and measurement for the active-content half are kept in their own blocks:
a different source structure at a different commit, and one pin must not stand for
two measurements. The guard's HEAD moved twice during the work (3c56d50 -> de09711 ->
398eb74); measurement ran at de09711 and the drift is recorded, including that
active_content.py is NOT identical to the 0bf0729 the data file pins - the change
adds a scan-cap self-safety finding and touches no construct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVouC9nsfrfV5jRSejxbvQ
2026-08-10 21:15:25 +02:00
a1578e6f3f fix(conformance): record the guard's internal-surface position on _LEX_PAYLOADS
llm-ingestion-pipeline-security states (coord message 2026-08-10T12:42:31Z)
that _LEX_PAYLOADS and the 83 pattern ids are an internal surface on their
side, with no README/CHANGELOG/docs statement promising id or payload
stability. Their gate is their own test suite, not a promise to this
repository. Their stated position: a future payload change diverges the
pin and should be re-pinned, not treated as a broken contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJTZgfnjsaN14ti5HXiKhQ
2026-08-10 20:47:44 +02:00
49e1e79807 feat(conformance): build the 83-case corpus; the 13-pattern blocker dissolves under measurement
The corpus was blocked on a decision nobody had to take. The 13 divergent
lexicon patterns were measured on witness inputs -- an attribute run padded
past 256 characters, an interior '<', an unclosed <script> -- and the corpus
payloads contain none of those shapes. Run through both runtimes' public
entry points, all 83 patterns produce identical lexicon finding sets, the 13
included. No winner picked, because the question was never reachable from
these inputs.

Measured at each runtime's entry point (scanForInjection() at b0de0ca,
scan_output(source=OUTPUT) at 0bf0729), never at a rebuilt regex table --
the layer mistake the divergence document already had to retract once.

- conformance/<case>/{input.txt,expected.json} x83, plus manifest.json
- spec/conformance-corpus.md, normative: bytes not text, id-only findings,
  exact-within-scope, and observed_out_of_scope as evidence not expectation

Verified by a harness that does not share the generator's knowledge: reads
only the case directories, re-runs both runtimes, checks every field
including digests -- 83 cases, 0 failures. Severity agrees with what commons
publishes 83/83. Deleting the middle third of each input breaks 76 of 83
expectations; the 7 survivors are the shortest payloads, a weak mutation
rather than a weak fixture, and are recorded as such.

Scope is 83 and not 94 for a different reason than expected: the 11
non-lexicon convertible cases have no ratified cross-runtime finding id, and
writing them would mint a contract unilaterally in the same stroke as the
tag. Named in manifest.json under scope_planned.

Docs corrected in place rather than edited away: extraction-plan and
lexicon-port-divergence both claimed the 13 blocked the corpus.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WhXDL82FRrQWedEmUg12Pj
2026-08-10 04:40:56 +02:00