git-subtree-dir: scanners/commons git-subtree-split: 0ffee85a4b83b3661185488c06ed9a9994c11412
178 lines
8.8 KiB
Markdown
178 lines
8.8 KiB
Markdown
# Conformance corpus
|
|
|
|
**Status: normative.** This document defines how the files under [`conformance/`](../conformance/)
|
|
are to be read. Where it disagrees with any informative document in `docs/`, this wins.
|
|
|
|
RFC 2119 language is used deliberately: MUST, MUST NOT, SHOULD, MAY.
|
|
|
|
## 1. What a case is
|
|
|
|
A case is a directory under `conformance/` holding exactly two files:
|
|
|
|
```
|
|
conformance/<case-id>/input.txt the exact bytes to scan
|
|
conformance/<case-id>/expected.json the findings a conforming runtime produces
|
|
```
|
|
|
|
A runtime claiming conformance MUST run every case. It MUST NOT skip a case because its
|
|
engine rejects the pattern behind it — a silently skipped case is an invisible false
|
|
negative, which is the failure mode this corpus exists to catch. A case that cannot be run
|
|
MUST be reported as an error, not as a pass.
|
|
|
|
## 2. `input.txt` is bytes, not text
|
|
|
|
`input.txt` holds the input **verbatim**, UTF-8 encoded. A runtime MUST read it as bytes
|
|
and MUST NOT strip, trim, normalise or re-encode anything before scanning.
|
|
|
|
There is **no trailing newline**. A trailing newline would be part of the input, and no
|
|
case wants one. Two of the current cases carry characters that are invisible on screen — a
|
|
zero-width space and a Cyrillic homoglyph — so an editor that "cleans up" a fixture can
|
|
change what it tests without changing how it looks. For that reason every `expected.json`
|
|
records the input's byte length and SHA-256, and a runtime SHOULD verify both before
|
|
scanning. A digest mismatch means the fixture has been altered; the result of scanning it
|
|
means nothing.
|
|
|
|
## 3. `expected.json`
|
|
|
|
```json
|
|
{
|
|
"case_id": "hybrid-xss__script-tag",
|
|
"input": { "file": "input.txt", "bytes": 24, "sha256": "…" },
|
|
"scope": ["lexicon/injection-lexicon.json"],
|
|
"match": "exact-within-scope",
|
|
"findings": [{ "pattern_id": "hybrid-xss:script-tag" }],
|
|
"observed_out_of_scope": { "llm_ingestion_guard": ["active:raw-html"] }
|
|
}
|
|
```
|
|
|
|
| Field | Normative? | Meaning |
|
|
| --- | --- | --- |
|
|
| `$comment` | no | Provenance prose. Carries no contract. |
|
|
| `case_id` | yes | Equals the directory name. |
|
|
| `input` | yes | Integrity of the input file. See §2. |
|
|
| `scope` | yes | Which commons data files this case constrains. See §4. |
|
|
| `match` | yes | How `findings` is to be compared. See §4. |
|
|
| `findings` | yes | The expected findings, named by commons `pattern_id`. See §3.1. |
|
|
| `observed_out_of_scope` | **no — informative** | What a named runtime was measured emitting outside `scope`. Evidence, never expectation. See §5. |
|
|
|
|
A finding carries a `pattern_id` and nothing else. Severity, OWASP anchor and human label
|
|
are properties of the pattern, published once in [`lexicon/injection-lexicon.json`](../lexicon/injection-lexicon.json)
|
|
and looked up there by that id. They MUST NOT be restated in a fixture: a value written in
|
|
two places is a value that will eventually disagree with itself.
|
|
|
|
### 3.1 `pattern_id` is not the finding `id`
|
|
|
|
The field is called `pattern_id` and not `id` because this repository already publishes a
|
|
different `id`, and confusing the two produces a runtime that fails every case for a reason
|
|
unrelated to detection.
|
|
|
|
| | Identifies | Stable? |
|
|
| --- | --- | --- |
|
|
| `pattern_id` — here; `families[].patterns[].id` in the lexicon | the **detection rule** | Yes. A stable cross-runtime contract, ratified by both seeding runtimes. |
|
|
| `id` — in [`schema/finding.schema.json`](../schema/finding.schema.json) | one **emitted finding** in one run | No. `DS-<scanner>-<counter>`, from a process-global counter: not stable across runs, not unique across processes. The schema says so itself. |
|
|
|
|
A runtime therefore MUST NOT compare a serialized finding's `id` to a fixture's
|
|
`pattern_id`. They answer different questions.
|
|
|
|
**Getting from a runtime's own label to a `pattern_id`.** Nothing requires a runtime to name
|
|
its patterns the way this repository does, and the two seeding runtimes do not: the same rule
|
|
is `override: ignore previous instructions` in one and `override:ignore-previous` in the
|
|
other. The mapping is published, per pattern, in the lexicon's `aliases` object:
|
|
|
|
```json
|
|
{
|
|
"id": "override:ignore-previous",
|
|
"aliases": {
|
|
"llm_security": "override: ignore previous instructions",
|
|
"llm_ingestion_guard": "override:ignore-previous"
|
|
}
|
|
}
|
|
```
|
|
|
|
(The lexicon's own field is spelled `id`, because that name is ratified and renaming it
|
|
would break both consumers. `pattern_id` here refers to exactly that value.)
|
|
|
|
A runtime SHOULD carry that id directly on its own pattern table, which makes the
|
|
comparison a lookup rather than a translation. A runtime that instead maps through `aliases`
|
|
MUST register its own alias there — a runtime absent from that object has no published way to
|
|
be compared, and a mapping maintained privately in a consumer is the drift this repository
|
|
exists to prevent.
|
|
|
|
Fixture files carry **no individual `version` field**, which is the one place this
|
|
repository's "every JSON file has a top-level `version`" convention does not apply. The
|
|
corpus is versioned as a whole in [`conformance/manifest.json`](../conformance/manifest.json).
|
|
The exception is deliberate and stated here rather than left to be discovered.
|
|
|
|
## 4. `exact-within-scope`
|
|
|
|
This is the only `match` value defined in this version, and it means:
|
|
|
|
> Restricted to findings whose id belongs to a data file named in `scope`, the runtime's
|
|
> finding set MUST equal `findings` exactly. Findings outside that scope MUST be ignored
|
|
> by the comparison — neither required nor forbidden.
|
|
|
|
So a runtime MUST NOT raise a lexicon finding this case does not list, and MUST raise every
|
|
one it does. Whether the same runtime also flags the input as active content, as a carrier,
|
|
or as anything else is **not** this case's business.
|
|
|
|
The scoping is what makes the exactness safe. The two seeding runtimes do not implement the
|
|
same set of tables — one has no active-content table at all — so a whole-report comparison
|
|
would fail for reasons that have nothing to do with the pattern under test. Scoping the
|
|
comparison to one published table asks a question both can answer.
|
|
|
|
The exactness is what makes the corpus worth running. A corpus that only checks that the
|
|
expected finding is *present* is passed by a runtime that flags everything.
|
|
|
|
A future case MAY declare a different `scope` or a different `match`. A runtime MUST reject
|
|
a `match` value it does not implement rather than degrade to a weaker comparison.
|
|
|
|
## 5. `observed_out_of_scope` is evidence, not expectation
|
|
|
|
This block records what a **named** runtime, at a **named** commit, was measured emitting
|
|
outside the case's scope. It exists so that a reader can see what was left out of the
|
|
contract instead of having to trust that nothing was.
|
|
|
|
A runtime MUST NOT be failed for disagreeing with it, and MUST NOT be required to reproduce
|
|
it.
|
|
|
|
**Absence of a runtime's key means unmeasured, not measured-empty.** The current corpus
|
|
carries no `llm_security` key anywhere, because that runtime was measured through its
|
|
injection-scanner entry point, which cannot emit an out-of-scope finding by construction.
|
|
Reading that absence as "this runtime emits nothing else" would be a claim nobody made.
|
|
|
|
## 6. Case ids
|
|
|
|
```
|
|
case_id = pattern_id with ":" replaced by "__"
|
|
pattern_id = case_id with "__" replaced by ":"
|
|
```
|
|
|
|
`:` is not a legal filename character on Windows, and fork-and-own is a supported use of
|
|
this repository, so the id space cannot reach the filesystem unchanged. `__` does not occur
|
|
anywhere in the ratified id space, so the transform is one-to-one — verified collision-free
|
|
across all cases rather than assumed.
|
|
|
|
**A case id is a stable identifier. Changing one is a BREAKING change** and requires a major
|
|
bump of the corpus version, exactly like changing a pattern id. Consumers name cases in
|
|
their own test suites; a renamed case is a silently disappeared test.
|
|
|
|
## 7. When a runtime and a fixture disagree
|
|
|
|
The fixture is ground truth. A runtime that disagrees is wrong — **unless the fixture itself
|
|
is proven wrong**, in which case the fixture changes in its own commit, with the reason
|
|
written down, and the corpus version is bumped.
|
|
|
|
This ordering is the whole point of the repository. Two implementations that return
|
|
different verdicts on the same input are not holding different opinions; one of them has a
|
|
bug.
|
|
|
|
## 8. What conformance does and does not prove
|
|
|
|
Passing this corpus proves that a runtime agrees with the other runtimes that pass it, on
|
|
these inputs. It does not prove the fixtures are right. A wrong `expected.json` makes every
|
|
passing runtime wrong in the same way, which is a real failure mode and not a rhetorical
|
|
one — it is why each case records how it was produced.
|
|
|
|
Nor is the corpus a coverage claim. It covers the pattern tables named in
|
|
`conformance/manifest.json` under `scope_covered`, and the tables under `scope_planned` are
|
|
named there precisely so that their absence is visible rather than inferred.
|