Compare commits

...

11 commits

Author SHA1 Message Date
c362717818 fix(security): use security@ as the reporting contact, not hello@
hello@ works, but two different addresses across sibling org repos
force a reporter finding a vulnerability to guess which one is the
security channel. security@fromaitochitta.com is the designated one.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fm8ErxAacrm5s8ZWWubgMP
2026-08-21 11:22:58 +02:00
e56812eb39 docs(readme): add table of contents
The README crossed 200 lines with eight H2 sections and no navigation
aid, forcing readers to scroll to find whether it solves their
problem before they've decided anything. AAA+ B-axis order 32, round 2.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XZzcSbaiw9nR686HDrF8KN
2026-08-16 16:17:45 +02:00
f0a511369d fix(spec): section 7 stated its own premise and then applied itself beyond it
Section 7 justified the fixture-is-ground-truth ordering with "Two
implementations that return different verdicts" and then stated the rule
with no scope. For signatures/active-content.json there is no second
implementation, and the seed runtime has stated the classification behind
it is calibration it does not freeze. As written, section 7 turned a change
they reserved into a bug on their side.

New section 7.1 keys the scope on a structural property, never on a table
name: a scope only one runtime implements, whose payload that runtime
authored. It creates no fourth verdict - the declaration schema closes
result with additionalProperties:false over four counts, so a fifth would
break every consumer's parser. The case still fails and is still named in
failed_cases; what changes is what the failure licenses concluding.

Two limits are stated rather than left to inference: it does not reach a
third-party implementer of the same table, and it is not a licence for a
runtime to self-declare its own divergence as calibration.

manifest.json 0.6.1 -> 0.6.2 retires the open-question sentence, quoted
rather than dropped. The retirement is partial: "section 7 is NOT amended
by this block" stays true, because the spec was amended by its own release.

Neighbours measured over the whole repository, widened past "ground truth"
to the second paragraph's own wording. CONVENTIONS.md and CLAUDE.md carried
the premise and are changed; SECURITY.md gets a cross-reference only, since
its claim is about a fixture expecting too little and 7.1 does not narrow
that direction; README.md and docs/extraction-plan.md are named as
deliberately untouched.

Breaking in category, minor in number - 0.x, per the reading [0.3.0]
recorded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012pZ2FLQ6xkWvj2VcwgwnQv
2026-08-13 23:31:28 +02:00
c75c546614 docs(manifest): the precision field stated its exposure as a hand-derived count
0.8.0 added a field whose whole purpose is precision, and bounded the exposure
with "three of the four dimensions ... cannot move one of these cases at all".
The total was derived by hand over a taxonomy the field had itself
recategorized. Lexicon entries ARE one of the seed runtime's four calibration
dimensions, and they are not absent from these fixtures: two of the seven carry
a lexicon id in observed_out_of_scope. And "the fourth" substituted the
classification for lexicon entries as the fourth item of their sentence.

manifest 0.6.0 -> 0.6.1: the field enumerates three named things and totals
none of them. A lexicon change ages the residue as evidence without moving a
verdict (spec section 5); the classification is the one that can move one.
Retired sentence quoted in an AMENDED IN clause, not dropped.

No measurement changed, no verdict moved, no case or data file touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016yZobrgUiRtpLSWx8i7u2Z
2026-08-13 23:15:10 +02:00
c23aea9062 docs(manifest): the active-content fixtures pin a version, and nothing said so
The seed runtime's v1.0.0 freezes its exported Python surface and explicitly
not its detection behaviour. The manifest pinned commit and version per
measurement block but never recorded that the thing pinned is a version rather
than a frozen classification.

manifest 0.5.2 -> 0.6.0, one new field next to active_content_provenance.
asymmetry, bounded by what the fixtures actually assert: all seven carry
pattern_id only, so three of the four calibration dimensions cannot move them.
Both pins named, not one. Their statement is attributed, not restated as ours.

Also closes the omission 0.7.3 named: the same misquote in
docs/secret-egress-divergence.md:75-76. Not the fix the note implied - those
lines are one single-backtick span across a line break, so the outer delimiter
is promoted to double backticks instead.

spec section 7 deliberately untouched and named in CHANGELOG.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016yZobrgUiRtpLSWx8i7u2Z
2026-08-13 23:09:33 +02:00
757570dd49 fix(readme): the paragraph still argued the premise our own commit retired
README.md opened the egress gap with "It is not an id question at all -- the two
runtimes carry different tables, cut at different granularities", and sent the
reader to scope_planned.blockers as the authority. Since 4356caa that blocker
opens reason (1) with "NO ID SPACE ON THE COMMONS SIDE. This is the hard
blocker." Before 0.7.2 the README was out of date; after it, two files on a
public remote disagreed, and it was our commit that made them.

The paragraph now carries the three measured, independent reasons from
docs/secret-egress-divergence.md. The counts survived the falsification, so
"different tables" and 19-against-25 are kept; only "cut at different
granularities" and the pending-reconciliation claim are gone. It deliberately
omits the outgoing question's status (true on the day written, untested by
anything, and dated in the manifest) and the standing entry_points_by_scope
requirement (a requirement, not a fourth reason).

conformance/manifest.json 0.5.1 -> 0.5.2 in the same release: the blocker
misquoted the contract it cites, rendering the value without the backticks the
field carries around `order`. The data file is unchanged and was never wrong.
Round-trip proved byte-neutrality before the edit; the read-back matched the
decoded string against the data file that owns it, since json.tool passes on
wrong escaping.

No case minted, no data file touched, no id proposed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SZ5vrpu2kxcRktiW59b7s3
2026-08-13 22:58:29 +02:00
4356caa689 fix(manifest): the egress blocker named a premise the measurement falsified
`scope_planned.blockers` for secret egress read "19 entries … 25 at different cut
points" — one table cut at two granularities, waiting on a reconciliation of two
ports. `docs/secret-egress-divergence.md` (2d9ee9c, corrected in 4a6f6ff) measured
otherwise: they are ports of two different source tables in the same source
repository, so reconciling the ports was never going to close it.

Replaced with the three measured, independent reasons: no id space on the commons
side (seed A carries name+pattern only, and a fixture names labels); match
semantics disagree (first-match-wins against finditer over all 25 — one label
against two on the Bearer+JWT witness); membership diverges both ways and is
inherited from two different seeds, so re-measuring either port cannot close it.
Only the first is the outgoing question; the other two stand whatever the answer.

Two hand-carried numbers in the retired text corrected in place: aws-access-key-id
was not the one clean 1:1 (2/19 byte-identical, AWS not among them — the guard
anchors with \b), and `GitHub Token` maps to three guard ids, not four, leaving
ghu_ and ghr_ uncovered. `scope_planned.$comment` said "a distinct unresolved
question" — singular — and is amended alongside.

Measured here, not transcribed: `entry_points_by_scope.scopes` carries no entry for
signatures/secret-egress.json at all. Recorded as a standing requirement, not as a
fourth reason. No id string is proposed; checked against both outgoing coord
messages of 2026-08-13. No case, no signature table and no other data file touched.

conformance/manifest.json 0.5.0 -> 0.5.1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016inh17NCrQpWN3mrghfJgT
2026-08-13 22:46:15 +02:00
4a6f6ffc16 docs(egress): the summary row was hand-totalled, and the caveat was measurable
Three corrections to a document whose measurements were right.

The "shapes commons reports and the guard is silent on" row said 6 across 4
commons entries. Re-derived from the differential rather than from the table:
7 witnesses across 5 entries. The old figure collapsed the two webhook hosts
into one row while keeping the two GitHub prefixes as two, so it was
inconsistent with itself, and it dropped GitHub Token from the entry count
even though two of its five prefixes are exactly what the guard misses. The
table now runs one row per witness and says why counting either way alone
misleads.

The mapping row for connection strings said "minus one scheme". The section
below it already said the right thing: commons catches mongodb:// and misses
the SRV form. Aligned.

The "seed B was read at a commit the guard did not port from" caveat is gone,
replaced by the measurement that dissolves it. knowledge/secrets-patterns.md
has one commit at or before 47905da - f153f96, 2026-04-08 - and the guard's
output.py was first committed 2026-07-04. The seed had been still for three
months when the port was written. The 5 real drifts are guard-side by
measurement now, not by inference.

Neither coord message carried the bad count, so nothing needs re-sending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UFforUbBA7GnUYijg78kpK
2026-08-13 22:22:29 +02:00
2d9ee9c434 docs(egress): the premise was one table cut two ways, and it is two tables
The open question was recorded as "19 entries here against the guard's 25,
cut at different granularity". Measured at a pinned tag, that premise does
not hold: commons ported llm-security's `hooks/scripts/pre-edit-secrets.mjs`
(19 entries, name + pattern only), the guard ported its
`knowledge/secrets-patterns.md` (33 entries, ids + severity + FP notes) and
took 25 of them. Two seeds, not two cuts.

The guard's docstring asserting the second seed was not relayed - the file
was read at the pinned commit and all 25 guard ids are in it verbatim.

Beyond the count, three things a membership table would have hidden:

- Match semantics disagree. Commons declares first-match-wins with normative
  ordering and a load-bearing last entry; the guard reports every match. One
  Bearer-plus-JWT witness: one label vs two.
- Only 2 of 19 patterns are byte-identical across the two sides. The pair
  earlier called the clean 1:1 (AWS access key) is not one of them.
- The guard suppresses placeholder and variable-reference values; commons has
  no field that could carry that, because its seed carries none.

Membership diverges both ways and neither port is at fault - the seeds
disagree. Reported, not fixed, per the behaviour-preservation invariant.
No id is proposed: that is the carrier rule, and both runtimes are asked
first. manifest.json is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UFforUbBA7GnUYijg78kpK
2026-08-13 22:16:06 +02:00
27b31701e0 docs(readme): the variant rule is normative in the spec now, and the README still pointed at the manifest
Two loose ends from v0.7.0.

The README sent a reader to case_id_derivation.variant_suffix in the manifest
for the rule about variant cases. That was correct until v0.7.0 made the rule
normative in spec section 6 and left the manifest block as the MEASUREMENT
behind it. Left alone it reproduces in one line the defect v0.7.0 closed: a
reader sent to the wrong authority. Both are named now, each for what it is.

And the release that made a point of re-measuring '--' rather than copying the
manifest's 0.3.0 numbers forward had inherited the adjacent '__' claim
untested. Measured now across all five published id spaces - 83 lexicon ids, 7
active: ids, 3 carrier: ids, 7 malware rule ids, 19 secret-egress entry names -
neither '__' nor '--' occurs in any of them. The one-to-one transform holds, so
no text changed; the sentence was true and is now measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M4xrxV3EXbSALqvB23kpeY
2026-08-13 22:04:16 +02:00
cb784fea6f fix(spec): section 6 forbade a case the corpus ships, and the predicate under it was wrong too
The normative spec read "Such a payload MUST NOT be given a discriminated case
id" while manifest.json defined case_id_derivation.variant_suffix and
conformance/hybrid-xss__script-tag--src-no-close/ sat on disk under it. The
manifest was the correct party: the derivation was extended in corpus 0.3.0 and
the spec was never updated. No data moves here, only the text describing it.

The derivation block now carries the optional '--' suffix and the truncating
reverse, with the '--'-absence measurement stated as the reason the reverse
stays LEXICAL - re-measured at this commit rather than copied from 0.3.0's
numbers, and scoped to id spaces because '--' does occur inside pattern values.

The part that would have passed review while still being wrong: fixing only the
permission. Section 6 also reasoned that equal in-scope finding sets mean the
second case "cannot fail in any way the first does not" - and the shipped
variant falsifies exactly that. It expects the same single finding, same scope,
same match, and still gates what the base cannot, because the base input matches
the pattern under both its published and its superseded stricter form. The
discriminating signal is INSIDE the scope, in the form of the scoped rule. So
the MUST NOT is replaced by a predicate about failure surface rather than
finding sets, checked in both directions: it admits the shipped variant and
still excludes the omitted markdown-image payload, whose only distinguisher
lives in a table this repository does not publish.

manifest.json is untouched and stays at 0.5.0; no case directory moved. The
spec has no version of its own - "Through version 0.1.1" in section 4 is the
CORPUS version, verified against CHANGELOG [0.2.0] before acting, because the
session brief said otherwise. Section 6's stable-id and BREAKING sentences were
read, not edited.

Verified in scratchpad, never in the repo: the amended derivation transcribed
into a checker that reads all 94 cases back from disk, derives each pattern id,
round-trips it forward, and asserts the case expects it. All 94 reproduce.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M4xrxV3EXbSALqvB23kpeY
2026-08-13 22:00:24 +02:00
8 changed files with 841 additions and 29 deletions

View file

@ -9,6 +9,399 @@ Versioning note: the repository tag versions **the contract** (file set, key nam
case ids, disposition semantics). Each JSON file additionally carries its own case ids, disposition semantics). Each JSON file additionally carries its own
`"version"` field, bumped when that file changes. `"version"` field, bumped when that file changes.
## [0.9.0] — 2026-08-13
**A normative rule stated its own premise and then applied itself beyond it.**
`spec/conformance-corpus.md` §7 justified the fixture-is-ground-truth ordering with *"**Two
implementations** that return different verdicts…"* and then stated the rule with no scope at
all. For `signatures/active-content.json` there is no second implementation — the seed runtime
authored both the payloads and the table — and that runtime has stated that the classification
behind it is calibration it does not freeze. §7 as written made a reserved change on their side
into a bug on their side.
**Breaking in category, minor in number.** This changes disposition semantics, which the
versioning note at the top of this file counts as contract. The repository is in 0.x, where a
breaking change is a minor bump by the rules — the same reading `[0.3.0]` recorded: *read the
entry, not the version number*.
### Changed
- **`spec/conformance-corpus.md` — new §7.1, *Where the second paragraph does not hold*.** The
scope is keyed on a **structural property**, never on a table name: a case whose scope is a
table only one runtime implements, whose payload that runtime authored. A rule naming
`active-content` would rot the day a second runtime implements it. §7's own second paragraph
already carried the premise; §7.1 makes it explicit and states the disposition for the case
the premise excludes — the fixture is not rewritten on the divergence alone, the divergence is
recorded against the version pinned, and re-pinning is a separate release. That is the
disposition §5 already applies to a stale `observed_out_of_scope` entry, extended to the one
place where it can reach a verdict.
**It creates no fourth verdict, and that constraint shaped the wording.**
`schema/conformance-declaration.schema.json` closes `result` with `additionalProperties: false`
over four counts plus two arithmetic invariants; a fifth verdict would have broken every
consumer's parser, which is a worse break than the one intended. A case whose expected findings
are not produced still **fails** and is still named in `failed_cases`. What §7.1 changes is what
the failure licenses concluding, not what is reported.
Two limits stated in the section rather than left to be inferred: it does **not** reach a
third-party implementer of the same table — against them the fixture is the contract, exactly
as §7 says, and that is the only thing these cases can prove while one runtime is all there is
— and it is **not** a licence for a runtime to self-declare, since the exemption is carried by
the corpus's provenance record for the scope and not asserted per case by whoever failed.
Superseded text is named rather than edited away, following §6's own pattern: *"Through corpus
version 0.8.1 this section stated the rule above with no scope at all."*
**The competing reading was tested and disposed of**, because it is the one that would have
avoided this release: that §7's existing hatch (*"unless the fixture itself is proven wrong"*)
already covered it. It does not. The hatch's consequence is that **the fixture changes**, and
the manifest field asserts the opposite — pinned, not rewritten, re-pinning a separate
decision. And a runtime recalibrating does not prove the earlier classification wrong: the
fixture measured `de09711` / `0.4.0` correctly, and a later release does not reach back and
falsify an earlier measurement. The case fits neither of §7's two dispositions, which is the
defect.
- **`conformance/manifest.json` `0.6.1``0.6.2`
`active_content_provenance.pins_a_version_not_a_frozen_classification` no longer records an
open question.** The retirement is **partial and it is quoted, not dropped**, per the house
style this field established one release ago (*"a correction that does not say what it corrects
cannot be audited"*). What falls is only the open-question status; the clause *"section 7 …
is NOT amended by this block"* **stays true and is kept**, because §7 was amended by its own
release and not by a data file. Value change only — read back from disk against `HEAD` with a
flattened key diff: `added: 0, removed: 0, changed: 2` (the field and `version`), and the new
string printed and read rather than inferred from the count, since a value edit reports
`changed: 1` whatever it wrote.
Six prose dashes in the new text were written `--` and promoted to `—` before commit: `--` is
the variant-suffix separator token of §6's case-id grammar, and every other occurrence of it in
this file is that token, a real case id, or a CLI flag.
### Neighbours — measured, and the ones left alone are named
A sweep for the retired premise was run over the whole repository, widened past *"ground truth"*
to the second paragraph's own wording (*"one of them has a bug"*, *"two implementations"*), since
a restatement in that phrasing would have survived the first search.
- **`CONVENTIONS.md` — changed.** Carried the rule unscoped and called the proven-wrong hatch
*"the one way that reverses"*. There are now two, and both are listed.
- **`CLAUDE.md` — changed.** The Norwegian restatement that governs sessions in this repository
carried the same unscoped rule; left alone, the next session here would have acted on it.
- **`SECURITY.md` §2 — minimal cross-reference only.** Its claim is about a fixture that expects
**too little**, and §7.1 narrows *who the rule reaches*, not that direction. The conclusion
survives intact, so it was not rewritten.
- **`SECURITY.md` "Why a confirmed defect is usually not fixed here first" — untouched.** Its
*"two implementations answering differently"* is about extracted **data** diverging from its
source, not about fixtures.
- **`README.md` — untouched.** Its conformance row says *"Ground truth"* as a descriptor and does
not restate the disagreement rule, and it already names the asymmetry it would otherwise hide:
the seven active-content cases are *"measured against the one runtime that implements that
table"*. Nothing there became false.
- **`docs/extraction-plan.md` — untouched, and it is supporting evidence rather than a stale
neighbour.** It already records that the calibration file *"inverts this repository's central
rule"* — so this is the second place the unscoped rule was known not to hold, and the first was
documented before this release.
### Not in this release
Whether the seven active-content cases still pass at the seed runtime's `v1.1.0` is **unmeasured**,
and §7.1 is silent on it. No case was minted, no data file touched, no id string proposed.
## [0.8.1] — 2026-08-13
**The field 0.8.0 added to make the exposure precise stated it with a hand-derived count, and the
count was wrong.** Caught in the same session, before any consumer read it, and corrected inside
the field rather than by rewriting it. No measurement changed and no verdict moved.
### Fixed
- **`conformance/manifest.json` 0.6.0 → 0.6.1 —
`active_content_provenance.pins_a_version_not_a_frozen_classification` now enumerates instead of
totalling.** As published it read *"three of the four dimensions they name as calibration cannot
move one of these cases at all. The fourth can: which `active:` ids a payload yields IS the
classification"*. Two defects in one sentence. First, the total was derived by hand over a
taxonomy the field had itself recategorized: the seed runtime's four calibration dimensions are
severities, thresholds, **lexicon entries** and dispositions, and lexicon entries are *not* absent
from these fixtures — `active__data-uri` carries `data-uri:executable` and `active__raw-html`
carries `hybrid-xss:script-tag` in `observed_out_of_scope`, both verified as members of
`lexicon/injection-lexicon.json` and non-members of `signatures/active-content.json`. Second,
*"the fourth"* silently substituted the classification for lexicon entries as the fourth item of
their sentence, which it is not — the classification is what they addressed separately.
- The replacement names three things and totals none of them: severities/thresholds/dispositions
are absent and move no verdict; lexicon entries move no verdict either — spec section 5 forbids
failing a runtime over `observed_out_of_scope` — but a lexicon calibration change **ages** those
two entries as evidence, which is the exposure
`active_content_measurement_0_7_0.movement_sweep.residue_is_the_field_no_test_protects` already
names as a class, and this corpus pins a stale residue entry rather than rewriting it; and the
active-content classification is the one thing that can move a verdict. The retired sentence is
**quoted** in the field's `AMENDED IN 0.6.1` clause, not merely dropped, for the same reason
`scope_planned.$comment` quotes what it retired: a correction that does not say what it corrects
cannot be audited.
## [0.8.0] — 2026-08-13
**The seven active-content fixtures pin a VERSION of the seed runtime, and nothing said so.**
That runtime tagged `v1.0.0` on 2026-08-13 and stated that the freeze covers its exported Python
surface only, excluding detection behaviour: severities, thresholds, lexicon entries and
dispositions are calibration there and move in minor and patch releases. The manifest already
pinned commit and version per measurement block, but nowhere recorded that the thing pinned is a
version rather than a frozen classification. No case is minted, no data file is touched, no id is
proposed.
### Added
- **`conformance/manifest.json` 0.5.2 → 0.6.0 —
`active_content_provenance.pins_a_version_not_a_frozen_classification`.** One field, scoping the
neighbouring `asymmetry` rather than replacing it, and deliberately narrower than the runtime's
own statement. The exposure is bounded by what the fixtures assert, which was read from all seven
rather than assumed: every finding carries `pattern_id` and nothing else — no severity, no
threshold, no disposition — so three of the four dimensions that runtime names as calibration
cannot move one of these cases at all. The fourth can, because which `active:` ids a payload
yields *is* the classification. The field names both pins (six at 0.4.0 / `de09711`, the seventh
at 0.7.0 / `be9759b`) rather than one, since a single version would flatten two measurements into
one header — the defect `superseded_for_one_case` exists to prevent. Their v1.0.0 statement is
**attributed** to their coord message of 2026-08-13T20:40:31Z, not restated as a fact measured
from this side.
- The field also names the disposition of a future divergence, so it is not left to be inferred: a
later 1.x that classifies one of these payloads differently is not a breach by them and does not
make the fixture wrong. The fixture stays ground truth at its pinned version, the divergence is
measured and recorded, and re-pinning is a separate decision — the same disposition this corpus
already applies to a stale `observed_out_of_scope` entry.
### Fixed
- **`docs/secret-egress-divergence.md:75-76` carried the same misquote `conformance/manifest.json`
had corrected in 0.7.3**, named there as a deliberate omission and closed here. The field's value
ends ``ascending `order` `` — the backticks are the field's own. The fix is *not* the one the
omission note implied: those two lines are a single code span delimited by **single** backticks
across a line break, so inserting the field's backticks inside it would have terminated the span
at the first one and rendered the quote broken. The outer delimiter is promoted to double
backticks instead, which is what lets the inner singles survive. The manifest's correction ported
as a literal string because JSON has no backtick semantics; markdown does. Verified by extracting
the span from the file on disk, unfolding the line break, and comparing to the decoded value in
`signatures/secret-egress.json` — equal — and by confirming no backtick run of length ≥ 2 sits
inside the span.
### Not done, and named rather than left silent
- **`spec/conformance-corpus.md` section 7 is untouched.** It states the disagreement rule without
scope: *"The fixture is ground truth. A runtime that disagrees is wrong."* Read against the
active-content scope, whose only implementing runtime has now said in writing that its
classification may legitimately move, that rule would call a calibration change there a bug. The
manifest field records the interaction and explicitly does not amend the spec. Whether the
normative rule needs a scope is a decision for its own release.
## [0.7.3] — 2026-08-13
**The README still argued the premise 0.7.2 retired, and the two files sat on a public remote
disagreeing.** `README.md` opened the egress gap with "It is not an id question at all"; the
blocker it sends the reader to for authority now opens reason (1) with "NO ID SPACE ON THE
COMMONS SIDE. This is the hard blocker." Before 0.7.2 the README was merely out of date. After
it, our own commit had made it contradictory — the same defect class 0.7.1 existed to close. No
data moves, no case is minted, no id is proposed.
### Fixed
- **The README now carries the three measured reasons instead of the retired one.** (1) No id
space on the commons side — the hard blocker, and the only one an answer can resolve; the
answer belongs to the runtimes that own the seeds. (2) Match semantics disagree, and an id
space would not close it. (3) Membership diverges in both directions and the divergence is
inherited: the two sides hold 19 entries and 25, and they are ports of two *different* source
tables in one source repository. The counts survived the falsification; only the causal claim
fell, so `different tables` is kept and "cut at different granularities" is gone. The
paragraph deliberately does **not** restate the outgoing question's status: that is true on
the day it is written, nothing tests README prose, and `conformance/manifest.json` already
carries the date. The standing `entry_points_by_scope` requirement is likewise left out rather
than printed as a fourth reason.
- **`conformance/manifest.json` 0.5.1 → 0.5.2: the blocker misquoted the contract it cites.** It
rendered the field as `match_semantics: "… evaluated in ascending order"`; the value in
`signatures/secret-egress.json` ends ``ascending `order` `` — the backticks are the field's
own. A blocker that misquotes the semantics it is blocking on invites a consumer to implement
the wrong one. The data file is
unchanged and was never wrong — only the quotation of it was, which `scope_planned.$comment`
now records. Verified by reading the edited file back from disk and matching the decoded
string against the data file that owns it; `json.tool` passes on wrong escaping.
Known and deliberately left: `docs/secret-egress-divergence.md` renders the same value without
its backticks. That document is `Status: informative` and was outside this release's scope.
## [0.7.2] — 2026-08-13
**`scope_planned.blockers` named the premise that `docs/secret-egress-divergence.md`
falsified.** The blocker read "19 entries … 25 at different cut points" — one table cut at two
granularities, waiting on a reconciliation of two ports. Measured 2026-08-13: they are ports of
**two different source tables** in the same source repository, so no reconciliation of the ports
was ever going to close it. No data moves in this release, and no case is minted — only the
recorded reason a case cannot be.
### Fixed
- **The egress blocker now carries the three measured reasons, kept independent.** (1) Commons
has no id space for this table: seed A (`hooks/scripts/pre-edit-secrets.mjs`) carries a name
and a pattern per entry and nothing else, so entries are keyed by human-readable `name` while
the guard emits `egress:<id>`, and a fixture names labels. This is the only one of the three
an answer can resolve, and it is the outgoing question. (2) Match semantics disagree:
`first match wins` with `ordering.normative: true` here, against `finditer` over all 25
patterns there — one witness, an `Authorization` header holding a three-part JWT, produces
**one** label under commons' declared contract and **two** from the guard. (3) Membership
diverges both ways and is inherited from two different seeds (seed A 19 entries, seed B 33,
the guard ported 25, 8 unported), so re-measuring either port cannot close it. The blocker
points to `docs/secret-egress-divergence.md` for the method behind every number.
- **Two hand-carried numbers in the retired text are corrected in the same string.**
`aws-access-key-id` was called "the one clean one-to-one": measured, only **2 of 19** commons
patterns are byte-identical to a guard pattern after unescaping, and AWS is not among them —
the guard anchors the same run as `\bAKIA[0-9A-Z]{16}\b`. `GitHub Token` was called "four ids
there": measured on witnesses it maps to **three**, and leaves `ghu_` and `ghr_` covered by no
guard id. Both were transcription, not measurement. What is retracted is quoted in place; the
full retired text stands in git at `conformance/manifest.json` 0.5.0.
- **`scope_planned.$comment` said "a distinct unresolved question" — singular.** Left alone it
would tell a reader the case becomes mintable when an answer arrives, which is true of one
reason in three. Amended alongside the blocker rather than after it, since the two are read
together.
### Measured
- **The standing requirement was measured here, not transcribed from the document.**
`entry_points_by_scope.scopes` carries **no entry at all** for `signatures/secret-egress.json`
— the three declared scopes are the lexicon, active-content and carriers. Entry point,
findings accessor and fixture presentation must be filled for both runtimes before a first
egress case, independently of the three reasons. It is recorded as a requirement, not as a
fourth reason: it would stand even if all three were resolved tomorrow.
- **No id string is proposed, in this file or anywhere else.** Checked against the two outgoing
coord messages of 2026-08-13 rather than assumed: both state in as many words that no id is
being proposed. Naming an id in a shared space is the exception `carrier:*` established, it
requires both runtimes asked first, and both are unanswered.
`conformance/manifest.json` 0.5.0 → 0.5.1. No case directory, no `expected.json` and no
signature table changed; `git status` shows one file besides this changelog.
## [0.7.1] — 2026-08-13
Two loose ends from `0.7.0`, neither of which changes a contract.
### Fixed
- **The README still named the manifest as the authority for the variant rule.** It read
"see `case_id_derivation.variant_suffix` in the manifest" — true until `0.7.0`, when the
rule became normative in `spec/conformance-corpus.md` §6 and the manifest's block became
the *measurement* behind it rather than the contract. Left alone it would have reproduced
in one line the same defect `0.7.0` closed: a reader sent to the wrong authority.
### Measured
- **The `__` half of the derivation was re-measured too, not just the `--` half.** `0.7.0`
made a point of re-measuring `--` rather than copying the manifest's `0.3.0` numbers
forward, while the adjacent sentence asserting that `__` "does not occur anywhere in the
ratified id space" was inherited untested — a claim about this release's own soundness that
the release did not check. Measured now across all five published id spaces: the 83 lexicon
ids, the 7 `active:` ids, the 3 `carrier:` ids, the 7 malware rule ids and the 19
secret-egress entry names carry **neither** `__` nor `--`. The one-to-one transform holds.
No text changed; the sentence was true. It is now true *and* measured.
## [0.7.0] — 2026-08-13
**The normative spec forbade, in as many words, a case the corpus has shipped since
`0.3.0`.** `spec/conformance-corpus.md` §6 read *"Such a payload MUST NOT be given a
discriminated case id; the derivation rule is the contract, and a suffix would break the
reverse transform"* while `conformance/manifest.json` defined `case_id_derivation.
variant_suffix` and `conformance/hybrid-xss__script-tag--src-no-close/` sat on disk under it.
The manifest was the correct party; the spec was simply never updated when the derivation was
extended. **No data moves in this release — only the normative text that describes it.**
### Fixed
- **§6's derivation block now states the rule the corpus actually uses.**
```
before case_id = pattern_id with ":" replaced by "__"
pattern_id = case_id with "__" replaced by ":"
after case_id = pattern_id with ":" replaced by "__",
optionally followed by "--" and a variant slug of [a-z0-9-]
pattern_id = case_id truncated at the first "--" if present,
then "__" replaced by ":"
```
The `--`-absence measurement moves into the spec as the reason the reverse transform stays
**lexical**: a runtime MUST be able to recover a `pattern_id` by splitting the string, and
MUST NOT need a lookup against the published id list to find where the id ends and the
variant begins. Re-measured at this commit rather than copied from the manifest's 0.3.0
numbers: `--` occurs in none of the 83 lexicon ids, none of the three `carrier:` ids, none
of the `active:` construct ids, none of the seven malware rule ids and none of the 19
secret-egress entry names — and in exactly one of the 94 case ids, the variant itself. It
does occur inside *pattern* values (`<!--\s*(?:AGENT|AI|…)`, `-----BEGIN … PRIVATE KEY-----`),
which is why the claim is scoped to id spaces and not to the data files as a whole. The
spec states the property, not the counts, which is what keeps it from going stale the way
"verified collision-free across all 90" in the manifest did.
- **The predicate for minting a variant was wrong, and fixing only the permission would have
legalised the shipped case under a rule that still forbids it.** §6 reasoned that if two
payloads' in-scope finding sets are equal, the second "cannot fail in any way the first does
not." The shipped variant falsifies that: it expects the *same* single finding, in the same
scope, under the same `match`, and still gates something its base cannot — the base input
matches `hybrid-xss:script-tag` under both the published form and the stricter form that
preceded it, so reinstating the stricter form leaves it passing, while the variant input
matches only the published form and fails. The distinguishing signal is *inside* the scope,
in the form of the scoped rule itself, which is exactly what a finding-set comparison cannot
see.
The `MUST NOT` is replaced by a predicate that admits the shipped case and still excludes
the omitted one:
> A variant case MAY be minted when the second input can fail, **within the case's scope**,
> under a change to a scoped data file that the first input would pass. Where no edit to a
> published table separates the two inputs, the second case cannot fail in any way the first
> does not, and it MUST NOT be minted.
Checked against `omitted_payloads`: the guard's seventh active-content payload is
distinguished from the case already built only by `entropy:base64-blob`, and this repository
publishes no entropy table, so no edit to any scoped file separates the two inputs. It stays
omitted, on the one ground the manifest already records as standing. **That verdict is
unchanged by this release** — the manifest's own note that the derivation ground lapsed in
`0.3.0` remains the only part of it that has moved.
- **The manifest's `constraint` is now normative rather than metadata.** A variant case MUST
be scoped and matched exactly like its base case and MUST expect the same `pattern_id`; the
suffix distinguishes inputs, never findings. Two cases at one pattern id expecting different
findings within the same scope are not a variant pair.
### Unchanged, deliberately
- **`conformance/manifest.json` stays at `0.5.0` and no case directory was touched.** The
corpus version tracks the corpus; no case, no id, no expectation and no measurement changed
here. Bumping it would date 94 fixtures to a commit that only edited prose.
- **The spec carries no version of its own, and none was added.** The "Through version 0.1.1"
reference in §4 is the *corpus* version (`conformance/manifest.json` went `0.1.1``0.2.0`
in the `0.2.0` release), not a spec version — verified before acting, because the session
brief said otherwise. Normative specs in this repository are versioned by the repository
tag, exactly as in `0.2.0`, which rewrote §4 and §6 prose under the same mechanism.
- **§6's stable-id paragraph is untouched.** "A case id is a stable identifier. Changing one
is a BREAKING change" is a separate rule that sits in the same section; it was read, not
edited.
- **Minor, not major, and the reason is uncomfortable enough to state:** a consumer whose
reverse transform is `--`-naive has been broken since `v0.3.0`, when the case shipped. This
release documents that break; it does not create it. Nothing here changes a key, a case id
or a disposition.
### Verification
Mechanical, in scratchpad, never in the repository (charter). The amended derivation was
transcribed out of the prose into a checker that reads all 94 cases back from disk and, for
each: asserts `case_id` equals the directory name, applies the reverse transform, round-trips
it forward, and asserts the derived `pattern_id` is one the case expects. All 94 reproduce.
The variant's scope, `match` and findings were asserted identical to its base case — the new
MUST, executed rather than eyeballed — and the checker also asserts that the withdrawn
sentences are gone and that the three stable-id sentences are still present verbatim (the
diff carries them as context lines, not as edits).
## [0.6.0] — 2026-08-13 ## [0.6.0] — 2026-08-13
**A seventh active-content case, and the whole raw-HTML classifier moves forward with it. **A seventh active-content case, and the whole raw-HTML classifier moves forward with it.

View file

@ -65,6 +65,13 @@ Ingen. Data + prosa. Filformater: JSON (data + schema), Markdown (spec), rå tek
- `expected.json` er ground truth. Er en runtime uenig med `expected.json`, er runtimen - `expected.json` er ground truth. Er en runtime uenig med `expected.json`, er runtimen
feil — med mindre fixturen selv bevises feil, og da endres fixturen i eget commit med feil — med mindre fixturen selv bevises feil, og da endres fixturen i eget commit med
begrunnelse. begrunnelse.
- **Regelen over er skopet, og skopet er bærende.** Er casens scope en tabell bare ÉN runtime
implementerer, og den runtimen skrev payloaden, finnes ikke den andre implementasjonen
regelen dømmer mellom. Da er en divergens fra *den* runtimen verken en bevist feil fixture
eller nødvendigvis deres bug: fixturen skrives ikke om på divergensen alene, den føres mot
versjonen som er pinnet, og re-pinning er en egen release. Mot en TREDJEPARTS-implementasjon
av samme tabell gjelder §7 uendret. Til og med `v0.8.1` sto regelen uskopet. Se
`spec/conformance-corpus.md` §7.1.
- **En case er ikke mintbar uten inngangspunkt for sitt scope.** Korpuset pinner ikke - **En case er ikke mintbar uten inngangspunkt for sitt scope.** Korpuset pinner ikke
lenger ett inngangspunkt per runtime for alt — `manifest.json` lenger ett inngangspunkt per runtime for alt — `manifest.json`
`entry_points_by_scope` bærer inngangspunkt, **findings-accessor** og `entry_points_by_scope` bærer inngangspunkt, **findings-accessor** og
@ -86,6 +93,13 @@ Ingen. Data + prosa. Filformater: JSON (data + schema), Markdown (spec), rå tek
sann ved commiten `measurement` pinner. Retter du én av 83, står 82 målinger ved én commit sann ved commiten `measurement` pinner. Retter du én av 83, står 82 målinger ved én commit
og én ved en annen, under en header som navngir én. Før avviket i manifestet med dato og og én ved en annen, under en header som navngir én. Før avviket i manifestet med dato og
commit i stedet. Å re-pinne hele korpuset er en egen beslutning. commit i stedet. Å re-pinne hele korpuset er en egen beslutning.
- **Like funn-sett betyr ikke lik feilflate.** Spørsmålet som avgjør om en variant-case skal
mintes er ikke om de to inputene gir ulike funn innenfor scope — det er om den andre
inputen kan FEILE, innenfor scope, under en endring i den scopede datafila som den første
ville bestå. Korpusets første variant forventer nøyaktig samme funn som base-casen og
gater likevel noe base-casen ikke ser: base-inputen matcher mønsteret både i publisert og
i tidligere, strengere form. En payload hvis skille ligger i en tabell vi ikke publiserer
består ikke terskelen og føres som navngitt utelatelse. Spec §6 bærer regelen.
### Id-rom: adoptert vs. navngitt ### Id-rom: adoptert vs. navngitt

View file

@ -107,10 +107,16 @@ this document:
- `<case-id>` is stable and descriptive. **Changing a case id is a breaking change** — a - `<case-id>` is stable and descriptive. **Changing a case id is a breaking change** — a
published conformance result names it. published conformance result names it.
- `expected.json` is **ground truth**. If a runtime disagrees with it, the runtime is wrong. - `expected.json` is **ground truth**. If a runtime disagrees with it, the runtime is wrong.
- The one way that reverses: the fixture is proven wrong. Then the fixture changes **in its own - One way that reverses: the fixture is proven wrong. Then the fixture changes **in its own
commit, with the reason written down** — never folded into a change that does something else, commit, with the reason written down** — never folded into a change that does something else,
because a fixture edit is the one edit that can make every conforming runtime wrong because a fixture edit is the one edit that can make every conforming runtime wrong
identically. identically.
- The other, added in `v0.9.0`: where a case's scope is a table only one runtime implements and
that runtime authored the payload, a divergence by **that** runtime is neither a proven-wrong
fixture nor necessarily its bug. The fixture is not rewritten on the divergence alone — it is
recorded against the version pinned, and re-pinning is a separate release. Through `v0.8.1`
this list carried only the first way. See
[`spec/conformance-corpus.md` §7.1](spec/conformance-corpus.md).
- A case declares the data files it is `scope`d to. A runtime that does not implement a scoped - A case declares the data files it is `scope`d to. A runtime that does not implement a scoped
table reports the case `not-applicable` — a third verdict beside pass and fail, and one that table reports the case `not-applicable` — a third verdict beside pass and fail, and one that
must be reported rather than dropped from the denominator. See must be reported rather than dropped from the denominator. See

View file

@ -16,6 +16,18 @@ unicode-carrier smuggling or active content in untrusted text, on any runtime.
**It holds no runnable code.** Data, specifications and fixtures only. **It holds no runnable code.** Data, specifications and fixtures only.
## Table of Contents
- [Install](#install)
- [Requirements](#requirements)
- [What it does](#what-it-does)
- [Non-goals](#non-goals)
- [Known limitations](#known-limitations)
- [Contributing](#contributing)
- [Reporting a wrong entry](#reporting-a-wrong-entry)
- [Changelog](#changelog)
- [License](#license)
## Install ## Install
Nothing to install — this repository is **vendored into consumers**, not installed. Nothing to install — this repository is **vendored into consumers**, not installed.
@ -93,8 +105,9 @@ The corpus covers three tables, and they do not carry equal weight — treating
number would misreport all three: number would misreport all three:
- `lexicon/injection-lexicon.json` — 84 cases over 83 patterns. Both seeding runtimes - `lexicon/injection-lexicon.json` — 84 cases over 83 patterns. Both seeding runtimes
implement it and both ratified its id space. One pattern carries a second, variant case; implement it and both ratified its id space. One pattern carries a second, variant case:
see `case_id_derivation.variant_suffix` in the manifest. the rule for when that is legal is normative in [§6](spec/conformance-corpus.md), and
`case_id_derivation.variant_suffix` in the manifest carries the measurement behind it.
- `signatures/active-content.json` — 7 cases, one per published id, the seventh added in - `signatures/active-content.json` — 7 cases, one per published id, the seventh added in
v0.6.0 when the seed runtime split raw HTML into two carrier classes. One runtime v0.6.0 when the seed runtime split raw HTML into two carrier classes. One runtime
implements it. For a runtime that implements it. For a runtime that
@ -108,11 +121,22 @@ number would misreport all three:
missing **name** rather than a missing capability. It lapses the moment that runtime names missing **name** rather than a missing capability. It lapses the moment that runtime names
its label and the alias is added. its label and the alias is added.
One case remains unshipped, for the secret-egress table, and it is not blocked on effort. One case remains unshipped, for the secret-egress table, and it is not blocked on effort. The
It is not an id question at all — the two runtimes carry *different tables*, 19 entries reasons are three, they were measured, and they are independent — none of them dissolves under
against 25, cut at different granularities, and a shared id space presupposes a reconciliation anything this repository can run alone. **(1) There is no id space on the commons side.** The
nobody has performed. `conformance/manifest.json` records that blocker under seed this table was ported from carries a name and a pattern per entry and nothing else, so its
`scope_planned.blockers`, measured, so the gap is visible rather than inferred. entries are keyed by human-readable name while the other runtime emits `egress:<id>` labels —
and a fixture names labels. This is the hard blocker, and the only one of the three that an
answer can resolve; the answer belongs to the runtimes that own the seeds, not to a name coined
here. **(2) Match semantics disagree**, and an id space would not close it: this table declares
first-match-wins with `ordering.normative: true`, the other runtime reports every match, and one
witness — an `Authorization` header holding a three-part JWT — produces one label here and two
there. That difference is exactly what an `expected.json` encodes. **(3) Membership diverges in
both directions, and the divergence is inherited rather than introduced.** The two sides hold 19
entries and 25, but they are ports of two *different* source tables in one source repository, so
re-measuring either port cannot close it. `conformance/manifest.json` records all three under
`scope_planned.blockers`, and the method behind every number is in
[the divergence measurement](docs/secret-egress-divergence.md).
The carrier blocker closed in v0.5.0 and is kept, with its retired text, under The carrier blocker closed in v0.5.0 and is kept, with its retired text, under
`scope_planned.blockers_resolved` — including the correction one runtime volunteered against `scope_planned.blockers_resolved` — including the correction one runtime volunteered against

View file

@ -18,7 +18,7 @@ detector, and that is a working bypass against every consumer until it is closed
Report privately by email: Report privately by email:
- **hello@fromaitochitta.com**, with `SECURITY` at the start of the subject. - **security@fromaitochitta.com**, with `SECURITY` at the start of the subject.
Pull requests are not the channel either — they are switched off on the canonical Pull requests are not the channel either — they are switched off on the canonical
repository, and not as an oversight. This repository is vendored into independent runtimes repository, and not as an oversight. This repository is vendored into independent runtimes
@ -45,8 +45,9 @@ In scope — all of these are real reports:
escaping is wrong for the declared dialect, missing or wrong flags, a pattern that fails escaping is wrong for the declared dialect, missing or wrong flags, a pattern that fails
to compile in a documented engine and gets skipped rather than reported. to compile in a documented engine and gets skipped rather than reported.
2. **A conformance fixture that sanctions a miss.** `expected.json` is ground truth: a 2. **A conformance fixture that sanctions a miss.** `expected.json` is ground truth: a
runtime that disagrees with it is deemed wrong. A fixture that expects too little makes runtime that disagrees with it is deemed wrong (as scoped by `spec/conformance-corpus.md`
every conforming runtime wrong identically, and the corpus will not catch it. §7.1, which narrows who that reaches and not this direction). A fixture that expects too
little makes every conforming runtime wrong identically, and the corpus will not catch it.
3. **A normative clause that mandates unsafe behaviour.** The `spec/` files bind the 3. **A normative clause that mandates unsafe behaviour.** The `spec/` files bind the
implementations that consume them, so a weak rule propagates to all of them. implementations that consume them, so a weak rule propagates to all of them.
4. **A real secret or personal data in the repository or its history.** The history is 4. **A real secret or personal data in the repository or its history.** The history is

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,295 @@
# Secret-egress divergence — commons vs the Python guard
**Status: informative.** Nothing here is normative and nothing here changes a data file. It
records a measured disagreement between two tables that were believed to be two cuts of one
source, and turns out not to be. Under this repository's behaviour-preservation invariant,
a divergence found here is **reported, not fixed**.
Produced 2026-08-13. Every number below came from a command; the scripts live in the session
scratchpad rather than in this repository, because executable code here would breach the
charter. They are reproducible from the method column. The guard was read via
`git archive v0.7.0`, never from its working copy.
**The premise this document was opened to test does not survive it.** The open question was
recorded as "19 entries here against the guard's 25, cut at different granularity" — one
table, two granularities. That is not what the two files are. They are ports of **two
different source tables in the same source repository**, and the granularity difference is a
consequence of that, not the cause. Everything below follows from correcting that premise.
## What was compared
| Side | Artefact | Version / coordinate |
| --- | --- | --- |
| commons | [`signatures/secret-egress.json`](../signatures/secret-egress.json) | file `version` 0.3.0, 19 entries |
| guard | `llm-ingestion-pipeline-security` `src/llm_ingestion_guard/output.py` `_SECRET_PATTERNS` | tag `v0.7.0` = commit `be9759b`, 25 entries |
| seed A | `llm-security` `hooks/scripts/pre-edit-secrets.mjs` `SECRET_PATTERNS` | commit `47905da`, 19 entries — what commons ported |
| seed B | `llm-security` `knowledge/secrets-patterns.md` | commit `47905da`, blob `a7ed469`, 33 entries — what the guard ported |
**The guard's v0.7.0 is the guard's current behaviour.** `git diff v0.7.0..aff3511 -- src/`
is empty, where `aff3511` was the guard's head when this was measured. Pinning at the tag
therefore costs no currency; it is not a waypoint measurement.
**Seed B was read, not accepted.** The guard's module docstring asserts *"Ported from the
`llm-security` `knowledge/secrets-patterns.md` seed"*. That assertion is a claim about a
third repository and would be an attribution, not a finding, if it were relayed. It was
measured instead: the file exists at the pinned commit on the public remote, and all 25 of
the guard's ids appear in it verbatim — `0` guard ids are absent from seed B. The docstring
is correct.
**Both seeds are named in commons' own file.** `signatures/secret-egress.json`'s `$comment`
already says which of the two it took and that the other *"is a separate PCRE-flavoured
agent-consumed variant that stays where it is"*. What was not known until now is that the
guard ported the other one.
## Result
| Measure | Method | Result |
| --- | --- | --- |
| Entry count, both sides | count entries | 19 and 25 |
| Seed B entry count | parse the `.md` at the pinned blob | 33 |
| Guard ids present in seed B | set membership on `id` | **25/25** |
| Seed B ids the guard did not port | set difference | **8** |
| Guard vs seed B, field-identical | compare regex (after stripping seed B's `(?i)` inline rendering), flags and severity | **16/25** |
| — of the 9 remaining, escaping-only | unescape the guard's `\"` (a Python raw-string artefact) and compare for string identity | **4/4 identical** |
| — of the 9 remaining, behaviourally real | differential match comparison | **5** — 4 connection strings, 1 capture-group change |
| commons vs guard, byte-identical patterns | unescape both sides' `\/` and `\"`, compare source + flags | **2/19** |
| Differential probe corpus | one witness per guard id, plus each side's exclusive shapes and the semantics witness | 36 probes |
| Shapes commons reports and the guard is silent on | differential | **7 witnesses, across 5 commons entries** |
| Shapes the guard reports and commons is silent on | differential | **3** |
| Match-semantics divergence | one witness matching two entries on both sides | **1 label vs 2 labels** |
16 field-identical + 4 escaping-only + 5 real = 25.
**The `2/19` is the number that says these are not two cuts of one table.** Only
`GitHub Fine-Grained PAT``github-pat-fine-grained` and `OpenAI Legacy API Key`
`openai-api-key-legacy` are byte-identical after unescaping. Even `AWS Access Key ID` is not:
commons has `AKIA[0-9A-Z]{16}` and the guard has the same run anchored, `\bAKIA[0-9A-Z]{16}\b`.
The earlier note calling that pair the one clean 1:1 was wrong, and was wrong by transcription
rather than by measurement.
## Match semantics: the divergence that is not about membership
This is the finding a membership table would hide, and it is the one a consumer implementing
from commons will get wrong first.
`signatures/secret-egress.json` declares ``match_semantics: "first match wins; patterns are
evaluated in ascending `order`"``, marks `ordering.normative: true`, and names
`last_entry_is_load_bearing: "JWT (three-part token)"` — the JWT entry is placed last
precisely so a token inside an `Authorization` header is reported as the header, not as a
bare JWT.
The guard's `scan_secret_egress` runs `finditer` over all 25 patterns and adds a finding for
every match. Order carries **no** semantics there, and there is no first-match-wins layer.
Measured on one witness — an `Authorization` header whose value is a three-part JWT:
| Side | Finding set |
| --- | --- |
| commons, under its own declared contract | `Authorization header with token` — one label |
| guard, `scan_secret_egress` at `v0.7.0` | `egress:bearer-token`, `egress:jwt-token` — two labels |
Both detect the credential. They disagree about what a report says, which is what a
`conformance/expected.json` encodes. Two runtimes that both "pass" here would still produce
different fixture files.
The commons side of this was not hand-rewritten: the evaluator compiles the patterns out of
the JSON, in `order`, applying `re.I` exactly where the file's own `dialect.translation_notes`
say to, and stops at the first hit. The guard side is the imported module. Neither table was
transcribed.
## Membership, measured
Every row below comes from running a witness input through both sides, not from reading the
two regexes side by side.
| commons `order` / name | guard ids observed | Relation |
| --- | --- | --- |
| 0 `AWS Access Key ID` | `aws-access-key-id` | 1:1, guard anchored |
| 1 `AWS Secret Access Key` | — | **guard silent** |
| 2 `Azure Connection String (AccountKey/SharedAccessKey/sig)` | `azure-storage-key` | overlap; see below |
| 3 `Azure AD ClientSecret` | `azure-client-secret` | 1:1 |
| 4 `Azure AI Services Key` | — | **guard silent** |
| 5 `GitHub Token` | `github-pat-classic`, `github-oauth-token`, `github-server-token` | 1:3, **plus 2 prefixes neither guard id covers** |
| 6 `npm Token` | `npm-token` | 1:1 |
| 7 `Anthropic API Key` | `anthropic-api-key` | 1:1 |
| 8 `OpenAI Project Key` | `openai-project-key` | 1:1 |
| 9 `GitHub Fine-Grained PAT` | `github-pat-fine-grained` | 1:1, **byte-identical** |
| 10 `Google API Key` | `gcp-api-key` | 1:1 |
| 11 `Private Key PEM Block` | `rsa-private-key`, `ec-private-key`, `pkcs8-private-key` | 1:3, **minus one PEM label** |
| 12 `JWT Secret` | — | **guard silent** |
| 13 `Slack/Discord Webhook URL` | — | **guard silent** |
| 14 `Generic credential assignment` | `generic-api-key`, `config-password`, `config-secret` | 1:3 |
| 15 `Authorization header with token` | `bearer-token` (+ `jwt-token`, see semantics) | 1:1 |
| 16 `Database connection string` | `postgres-connstr`, `mysql-connstr`, `redis-connstr` | 1:3, **minus the MongoDB SRV form** |
| 17 `OpenAI Legacy API Key` | `openai-api-key-legacy` | 1:1, **byte-identical** |
| 18 `JWT (three-part token)` | `jwt-token` | 1:1 |
Guard ids with no commons entry firing on their own witness: `gcp-service-account-json`,
`mongodb-connstr` — and `ec-private-key` on the `ENCRYPTED` header.
`Azure Connection String` is listed as *overlap* rather than 1:1 deliberately. Commons'
entry is `(?:AccountKey|SharedAccessKey|sig)=[A-Za-z0-9+/=]{20,}` — three alternatives, no
length pin. The guard's `azure-storage-key` is `AccountKey=([A-Za-z0-9+/]{86}==)` — one
alternative, exact length. The corpus witnessed only the `AccountKey` shape, where both fire.
`SharedAccessKey=` and `sig=` were not witnessed; seed B carries them under separate ids
(`azure-servicebus-connstr`, `azure-sas-token`) that the guard did not port. Read this row as
"one witnessed overlap", not as a coverage claim.
## What each side misses that the other catches
**Commons reports, guard silent — 7 witnesses across 5 commons entries:**
| Witness shape | commons entry | Why the guard is silent |
| --- | --- | --- |
| `ghu_` prefixed token | `GitHub Token` | guard ported `ghp`/`gho`/`ghs`; no id for `ghu` |
| `ghr_` prefixed token | `GitHub Token` | same |
| `aws_secret_access_key = <40 chars>` | `AWS Secret Access Key` | seed B has `aws-secret-access-key`; guard did not port it |
| `Ocp-Apim-Subscription-Key` assignment | `Azure AI Services Key` | absent from seed B entirely |
| `JWT_SECRET` assignment | `JWT Secret` | absent from seed B entirely |
| Slack webhook URL | `Slack/Discord Webhook URL` | absent from seed B entirely |
| Discord webhook URL | `Slack/Discord Webhook URL` | same |
One row is one witness, so two commons entries appear twice: `GitHub Token` covers five
prefixes behind one name, and `Slack/Discord Webhook URL` covers two hosts. Counting rows
rather than entries would overstate how much of commons the guard is missing, and counting
entries rather than rows would hide that `GitHub Token` is only *partly* uncovered — its
`ghp`/`gho`/`ghs` prefixes map onto three guard ids just fine.
Three of the seven are the sharper finding: the `Ocp-Apim-Subscription-Key`, `JWT_SECRET` and
webhook shapes are not entries the guard declined to port, they are entries **seed B does not
have**. Seed A carries three shapes seed B never did.
**Guard reports, commons silent — 3 shapes:**
| Shape | guard id | Why commons is silent |
| --- | --- | --- |
| `"type": "service_account"` | `gcp-service-account-json` | seed A has no GCP service-account marker |
| `-{5}BEGIN ENCRYPTED PRIVATE KEY-{5}` | `ec-private-key` | commons' PEM alternation is `(?:RSA \| EC \| DSA \| OPENSSH )?`; `ENCRYPTED` is not in it |
| `mongodb+srv://user:pw@host` | `mongodb-connstr` | commons' scheme run is `(?:postgres\|mysql\|mongodb\|redis)://` — the `+srv` suffix breaks the literal |
The `mongodb+srv` miss is worth naming precisely: commons is not missing MongoDB, it is
missing the **SRV** form, which is the form Atlas hands out. Plain `mongodb://` is caught.
**This asymmetry is not a scoreboard.** Each side is faithful to its own seed. Every shape in
the left table is present in seed A and absent from seed B; every shape in the right table is
the reverse. Neither port is wrong about its source. The seeds disagree.
## False-positive suppression: a layer commons has no field for
The guard applies value-based suppression to the five entries that capture a value
(`_is_fp_value`): structural placeholders (`your-`, `<`, `>`, `***`), word-boundary
placeholder words (`example`, `changeme`, `todo`, …), variable references (`${`, `$(`,
`os.environ`, `process.env`, …), all-same-character values, and values under 8 characters.
Measured:
| Witness | commons (first match) | guard |
| --- | --- | --- |
| `password: 'your-password-here'` | `Generic credential assignment` | — suppressed |
| `api_key: '${MY_API_KEY_VALUE}'` | `Generic credential assignment` | — suppressed |
Commons has no field that could carry this. `dialect.translation_notes` warns in prose that
the generic entries are *"shape matches, not proofs of a live credential"* and assigns the
trade-off to the consumer's policy — which is a correct statement of ownership and is also
why two consumers reading commons will produce different reports on the same placeholder.
Seed B carries the suppression semantics per entry in a `false_positive_notes` field; seed A
carries name and pattern only, so commons had nothing to extract. This is a gap in the seed,
not an omission in the extraction.
## The connection-string bound
The guard bounds the password run in all four connection-string patterns at
`MAX_CONNSTR_VALUE = 256`, and its module explains why in full: an unbounded run in front of
a required literal makes every start position rescan the tail when the literal never arrives.
They measured 8.2 s at 100 000 characters on crafted `redis://:` input and extrapolated to
hours at their own 1 000 000-character cap. Seed B's connection-string patterns are unbounded;
this is one of the 5 real guard-vs-seed-B drifts, and it is a deliberate, documented one.
Commons' `Database connection string` is `(?:postgres|mysql|mongodb|redis):\/\/[^\s]+@[^\s]+`
**shape-analogous** to what the guard bounded. Measured against the exact boundary:
| Password length | commons | guard |
| --- | --- | --- |
| 12 | matches | `egress:postgres-connstr` |
| 256 | matches | `egress:postgres-connstr` |
| 257 | matches | — |
| 300 | matches | — |
Read this as two facts, not one verdict. Commons has recall the guard traded away above 256
characters. Commons also carries the runtime shape the guard's measurement was about — and
carries it in an *unanchored* form (`[^\s]+@[^\s]+` rather than the guard's
`[^:@\s]+:…@[^\s'"]+`), so the two are not the same pattern under load and no timing claim
about commons is made here. **Nothing is changed on that basis.** The entry is faithful to
seed A, the file that owns it is `llm-security`'s, and the behaviour-preservation invariant
puts the decision there. It is reported, and the guard's measurement is cited so the owner
does not have to redo it.
## Severity and ids: what commons does not carry
Seed B carries `id` and `severity` per entry; the guard preserved both, and all 25 severities
are field-identical to the seed. Seed A carries neither, so commons carries neither, and its
`evidence_limits` says so explicitly: *"No severity, and no per-entry disposition, was
supplied … so neither is invented here."*
That restraint was right and it has a consequence: **commons has no id space for this table.**
Its entries are keyed by human-readable `name` (`"GitHub Token"`), while the guard emits
`egress:<id>` labels. A `conformance/expected.json` scoped to secret egress cannot be written
against commons today, because a fixture names labels and commons has none to name.
The 25 guard ids are **not guard-internal labels**. They are seed B's ids, adopted verbatim,
which was measured above (25/25 present in the seed). That makes the id space question a
question for `llm-security` first — they own both seeds and the id space in one of them — and
for the guard second. Per this repository's naming rule, **no id is proposed here.** The rule
that `carrier:*` established applies exactly: naming an id in a shared space is the exception,
it requires both runtimes asked first, and publishing `aliases.<runtime>` is irreversible at
file granularity.
## What this does not show
- **It does not show that either table is wrong.** Both are faithful ports. The disagreement
is between seed A and seed B, inside `llm-security`, and only that repository can say
whether two tables is intentional (one engine-consumed, one agent-consumed) or whether one
supersedes the other.
- **It does not measure seed A's current state.** Commons' fidelity to seed A was verified at
commit `47905da` and this document adds nothing to that.
*(Seed B was read at the same commit, which is a shared coordinate and not the commit the
guard ported from. That was going to be a caveat — a seed-B entry that moved between the
guard's port and `47905da` would show up here as guard drift. It is dissolved by measurement
instead: `git log -- knowledge/secrets-patterns.md` in a deepened mirror returns exactly one
commit at or before `47905da`, `f153f96`, dated 2026-04-08, and the guard's `output.py` was
first committed 2026-07-04. The seed had been still for three months when the port was
written and has not moved since. Reading it at `47905da` reads what the guard ported from,
so the 5 real drifts are guard-side by measurement rather than by inference.)*
- **It does not compare coverage.** The probe corpus has one witness per guard id plus each
side's exclusive shapes — 36 inputs. It is built to expose membership and semantics, not to
estimate recall. `SharedAccessKey=` and `sig=` Azure shapes, and seed B's 8 unported ids,
have no witness here.
- **It does not measure the runtimes' entry points.** Both sides were driven at table level:
commons through an evaluator compiled from its own JSON under its own declared contract, the
guard through `scan_secret_egress` directly. What `scan_output` composes around it —
decode-and-rescan re-labelling findings as `decoded:egress:*`, the oversize cap — is not in
scope and would change the finding sets.
- **It does not touch `manifest.json`.** `scope_planned.blockers` still names this divergence
as the blocker for egress cases. Whether this document dissolves that blocker or merely
describes it is a separate decision, and it depends on answers this document does not have.
## Consequence for `conformance/`
An egress case is not mintable today, and the reason has changed. It was recorded as "the two
tables are cut at different granularity". The measured reasons are three, and they are
independent:
1. **No id space on the commons side.** A fixture names labels. Commons has names, not ids.
This is the hard blocker and it is the subject of the outgoing question to both runtimes.
2. **Match semantics disagree.** Even with an id space, the Bearer-plus-JWT witness produces a
one-label expectation under commons' declared contract and a two-label one from the guard.
A fixture would have to encode one of them.
3. **Membership disagrees in both directions**, and the disagreement is inherited from two
different seeds rather than from a porting error — so it cannot be closed by re-measuring
either port.
None of the three is dissolved by a measurement this repository can run alone. Per
`conformance/manifest.json``entry_points_by_scope`, a new scope also needs an entry point,
a findings accessor and a fixture presentation for every runtime before its first case, and
those three slots are empty for egress on both runtimes. That requirement stands independently
of everything above.

View file

@ -211,8 +211,10 @@ Reading that absence as "this runtime emits nothing else" would be a claim nobod
## 6. Case ids ## 6. Case ids
``` ```
case_id = pattern_id with ":" replaced by "__" case_id = pattern_id with ":" replaced by "__",
pattern_id = case_id with "__" replaced by ":" optionally followed by "--" and a variant slug of [a-z0-9-]
pattern_id = case_id truncated at the first "--" if present,
then "__" replaced by ":"
``` ```
`:` is not a legal filename character on Windows, and fork-and-own is a supported use of `:` is not a legal filename character on Windows, and fork-and-own is a supported use of
@ -220,18 +222,53 @@ this repository, so the id space cannot reach the filesystem unchanged. `__` doe
anywhere in the ratified id space, so the transform is one-to-one — verified collision-free anywhere in the ratified id space, so the transform is one-to-one — verified collision-free
across all cases rather than assumed. across all cases rather than assumed.
The transform carries a consequence that is easy to miss: **a case id is derived from a `--` does not occur there either: the ratified ids use single hyphens throughout, measured
pattern id alone, so the corpus holds at most one case per `pattern_id` in single-finding across every id space this repository publishes and every case id in the corpus. That
scopes.** There is nowhere in the name to put a second one. That is a real constraint, not a measurement is what keeps the reverse transform **lexical**. A runtime recovers a
formality — a source runtime's own test matrix may well drive two payloads at the same `pattern_id` by splitting the string, and MUST NOT need a lookup against the published id
pattern, as one of the seeding runtimes does for `active:markdown-image`. When it does, the list to find where the id ends and the variant begins — a reverse transform that has to ask
two payloads MUST be compared *within the case's scope* before a second case is minted: if which of two readings is real is a different rule from the one written above, and it fails
their in-scope finding sets are equal, the second case cannot fail in any way the first does on the first id space that is vendored without its lookup table.
not, and its distinguishing signal lies outside the scope where this corpus makes no claim.
Such a payload MUST NOT be given a discriminated case id; the derivation rule is the **A `pattern_id` may carry more than one case.** Through corpus version 0.2.0 this section
contract, and a suffix would break the reverse transform. It SHOULD instead be recorded as said the opposite: that a case id derives from a pattern id alone, so a single-finding scope
a named omission in `conformance/manifest.json`, so the drop is visible rather than holds at most one case per pattern id, with nowhere in the name to put a second. Corpus
inferred from a count. version 0.3.0 extended the derivation with the optional suffix above, and the corpus has
shipped a case under it since (`hybrid-xss__script-tag--src-no-close`, whose own `$comment`
carries the rationale for that one). The superseded sentence is named here rather than edited
away, because it was the stated ground on which a real payload was dropped —
`omitted_payloads` in [`conformance/manifest.json`](../conformance/manifest.json) records
that ground as withdrawn and a second, independent ground as still standing.
The bar for minting a second case is neither that the two inputs differ, nor that their
in-scope finding sets differ:
> A variant case MAY be minted when the second input can fail, **within the case's scope**,
> under a change to a scoped data file that the first input would pass. Where no edit to a
> published table separates the two inputs, the second case cannot fail in any way the first
> does not, and it MUST NOT be minted.
Equal in-scope finding sets do not settle that question, and reading them as if they did is
the error this paragraph replaces. The corpus's first variant case expects exactly the
finding set its base case expects — one `pattern_id`, one scope, one `match` — and still
gates something the base cannot: the base input matches the scoped pattern both in its
published form and in the stricter form that preceded it, so reinstating the stricter form
leaves it passing, while the variant input matches only the published form and fails. The
distinguishing signal is *inside* the scope, in the form of the scoped rule itself, which is
exactly what a finding-set comparison cannot see.
The payload that stays out is the mirror image. A source runtime's own test matrix may drive
two payloads at the same pattern, as one of the seeding runtimes does for
`active:markdown-image`, whose only difference is a signal from a table this repository does
not publish. No edit to any scoped file separates them, so the second case could not fail
where the first passes. Such a payload SHOULD be recorded as a named omission in
`conformance/manifest.json`, so the drop is visible rather than inferred from a count.
A variant case MUST be scoped and matched exactly like its base case, and MUST expect the
same `pattern_id`. **The suffix distinguishes inputs, never findings.** It is not a licence
to record a second, different verdict for one rule: two cases at one pattern id expecting
different findings within the same scope are not a variant pair, they are the corpus
contradicting itself.
**A case id is a stable identifier. Changing one is a BREAKING change** and requires a major **A case id is a stable identifier. Changing one is a BREAKING change** and requires a major
bump of the corpus version, exactly like changing a pattern id. Consumers name cases in bump of the corpus version, exactly like changing a pattern id. Consumers name cases in
@ -247,6 +284,47 @@ This ordering is the whole point of the repository. Two implementations that ret
different verdicts on the same input are not holding different opinions; one of them has a different verdicts on the same input are not holding different opinions; one of them has a
bug. bug.
### 7.1 Where the second paragraph does not hold
**Through corpus version 0.8.1 this section stated the rule above with no scope at all**, and
the scope was load-bearing: the justification names *two* implementations. Where a case's
scope is a table only one runtime implements, and that runtime authored the payload the case
was extracted from, there is no second implementation whose disagreement the paragraph could
adjudicate. Which cases those are is recorded in the corpus, not asserted per run — see
`active_content_provenance.asymmetry` in
[`conformance/manifest.json`](../conformance/manifest.json).
For such a case, a disagreement by the **seed runtime itself** is a third thing, and it is
neither of the two the paragraph offers:
- The fixture is not proven wrong. It recorded that runtime's behaviour correctly at the
commit and version its own measurement block pins, and a later classification does not
reach back and falsify an earlier measurement.
- The runtime does not necessarily have a bug. Where the seed runtime has stated that the
classification behind such a table is calibration it does not freeze, a release that
classifies the payload differently is a change it reserved, not a defect.
So: the fixture MUST NOT be rewritten on the strength of the divergence alone; the divergence
SHOULD be recorded against the version pinned; and re-pinning the case to a later version of
the seed runtime is a separate decision, taken deliberately and released on its own. This is
the disposition §5 already applies to a stale `observed_out_of_scope` entry, extended to the
one place where it can reach a verdict — and a divergence recorded here is the signal that
the re-pinning decision is due, not a reason to leave it open.
Three things this does **not** do.
- **It creates no fourth verdict.** The counts of §1.1 and
[`schema/conformance-declaration.schema.json`](../schema/conformance-declaration.schema.json)
are unchanged: a case whose expected findings are not produced still **fails**, and is still
named in `failed_cases`. What changes is what the failure licenses concluding, not what is
reported.
- **It does not reach a third-party implementer** of the same table. Against them the fixture
is the contract, exactly as §7 states — which is what these cases were minted to provide,
and the only thing they can prove while one runtime is all there is.
- **It is not a licence to self-declare.** The exemption is carried by the corpus's own
provenance record for the scope. A runtime MUST NOT claim it for a case by asserting that
its own divergence is calibration.
## 8. What conformance does and does not prove ## 8. What conformance does and does not prove
Passing this corpus proves that a runtime agrees with the other runtimes that pass it, on Passing this corpus proves that a runtime agrees with the other runtimes that pass it, on