Advisor review found six gaps the session's own checks did not cover. The one
that mattered was outward-facing: the divergence report told the guard repo
that two sub-agent bounds were undocumented. That was wrong. lexicon.py's
module docstring documents them explicitly under "Bounded token gaps", and
git log -S dates {0,12} to the original port commit f397cd9 and [^><] to the
ReDoS fix cff0437. Every divergence on the guard's side is documented and
traceable. The claim rested on two sed slices of one file; an absence claim
needs a search over the whole repository. Corrected here and by coord.
lexicon/injection-lexicon.json 0.2.0 -> 0.3.0:
- pattern_id_space.alias_evidence records the two aliases separately instead
of averaging them. llm_ingestion_guard is verified — coverage.py asserts on
that exact string, so it is demonstrably what a guard finding carries.
llm_security is not: it is the pattern table's name, the finding producer
was never supplied, and the known Node finding shape uses title, not label.
- normalisations[].affects now keys on id, with the prose names kept beside it
as affects_labels. An internal cross-reference on label was a second
identity space inside the file the id exists to unify.
Detection data unmoved again: labels, patterns, flags and the ids and aliases
added in 7b70f5b are all byte-identical in sequence; 166/166 Node compiles.
Also: the README lexicon row described thematic families the file does not
have (they are severity families; the theme is the id prefix), and the
conformance convertibility table gained its missing second condition — a case
is buildable only if the label it asserts maps to data this repository
publishes. Eleven cases fail that test (entropy, decoded, the sanitize rows,
the OKF scans), so the buildable set is ~94, not 105.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FaYqid3mejFmd9ZHsiHgp3
120 lines
7.2 KiB
Markdown
120 lines
7.2 KiB
Markdown
# llm-security-commons
|
|
|
|
Runtime-neutral core for LLM and agent security detection: detector data, normative contracts and a conformance corpus that several runtimes can share.
|
|
|
|
[](LICENSE)
|
|
|
|
Detection logic gets reimplemented every time it crosses a language boundary, and the
|
|
copies drift: the Node scanner flags a zero-width carrier the Python guard misses, and
|
|
nobody notices until an incident. This repository holds the part that should never have
|
|
been copied — the pattern tables, the code-point carriers, the calibration thresholds, the
|
|
finding contract, and a fixture corpus with expected verdicts — so that two independent
|
|
implementations can be held to the same answer on the same input.
|
|
|
|
It is for anyone building or maintaining a detector for prompt injection, secret egress,
|
|
unicode-carrier smuggling or active content in untrusted text, on any runtime.
|
|
|
|
**It holds no runnable code.** Data, specifications and fixtures only.
|
|
|
|
## Install
|
|
|
|
Nothing to install — this repository is **vendored into consumers**, not installed.
|
|
|
|
As a `git subtree` (recommended: history is preserved and upgrades are a single command):
|
|
|
|
```bash
|
|
git subtree add --prefix vendor/commons \
|
|
https://git.fromaitochitta.com/open/llm-security-commons.git v0.1.0 --squash
|
|
|
|
# later, to move to a newer tag
|
|
git subtree pull --prefix vendor/commons \
|
|
https://git.fromaitochitta.com/open/llm-security-commons.git v0.2.0 --squash
|
|
```
|
|
|
|
Or pin a tag and copy — `fork-and-own` is an explicitly supported path:
|
|
|
|
```bash
|
|
git clone --depth 1 --branch v0.1.0 \
|
|
https://git.fromaitochitta.com/open/llm-security-commons.git
|
|
```
|
|
|
|
Always vendor **a tag**, never `main`. The tag is what a conformance result can be
|
|
attributed to.
|
|
|
|
## Requirements
|
|
|
|
A JSON parser and the ability to read a text file. That is the entire dependency surface,
|
|
and keeping it that small is the point.
|
|
|
|
## What it does
|
|
|
|
| Path | Contents |
|
|
| --- | --- |
|
|
| [`lexicon/injection-lexicon.json`](lexicon/injection-lexicon.json) | Prompt-injection pattern lexicon: 83 patterns in four **severity** families (`critical`, `high`, `medium`, `hybrid`), each with a stable `id` and per-runtime aliases. The thematic class (`override:`, `evasion:`, `hitl-trap:`, …) is the id prefix, not the family. |
|
|
| [`codepoints/carriers.json`](codepoints/carriers.json) | Invisible and deceptive carriers: zero-width characters, BIDI controls, Unicode Tag block ranges, and the homoglyph map. |
|
|
| [`signatures/secret-egress.json`](signatures/secret-egress.json) | Credential and token shapes that must never leave a machine, in a portable regex dialect. |
|
|
| `signatures/malware-signatures.json` | **Planned, not in v0.1.0.** Signature set for the malicious-code class (`SIG`). The seed data has not been delivered yet. |
|
|
| [`signatures/active-content.json`](signatures/active-content.json) | Active content that renders or fetches on its own — Markdown images, links, reference definitions and autolinks, `data:` URIs, active HTML. The EchoLeak class. |
|
|
| [`calibration/calibration.json`](calibration/calibration.json) | The numbers a detector must not invent: risk-score tier constants, verdict thresholds, risk-band cutoffs, posture grade thresholds. Transcribed from a prose summary, not differentially verified — the file says so itself. |
|
|
| [`mapping/owasp-map.json`](mapping/owasp-map.json) | Finding-id prefix → OWASP taxonomy entry (LLM / ASI / AST / MCP). |
|
|
| [`schema/finding.schema.json`](schema/finding.schema.json) | **Normative.** The finding contract, plus the SARIF and JSONL output profiles. |
|
|
| `spec/decode-pipeline.md` | **Planned, not in v0.1.0.** The decode order, in RFC 2119 language. Two runtimes that decode in different orders will disagree on identical input. Writing it needs the decode implementation, which is engine code and has not been supplied — and a normative spec guessed from a data dump would be worse than an absent one. |
|
|
| `conformance/` | **Planned, not in v0.1.0.** One directory per case: `input.txt` in, `expected.json` out. Ground truth. |
|
|
| [`docs/extraction-plan.md`](docs/extraction-plan.md) | Informative: where each file was seeded from, and what v0.1.0 promised. |
|
|
| [`docs/lexicon-port-divergence.md`](docs/lexicon-port-divergence.md) | Informative: a measured disagreement between two ports of the injection lexicon — 13 patterns that behave differently, in both directions, and why no data file was changed because of it. |
|
|
|
|
Every JSON file carries a top-level `version`. Every normative specification carries a
|
|
`Status: normative` marker. Rows marked **Planned** are named here because the layout is
|
|
part of the contract, but the file does not exist yet — they are not links, and nothing in
|
|
v0.1.0 depends on them.
|
|
|
|
Each data file records its own provenance and, in `verified`, how strongly it is backed.
|
|
`calibration/calibration.json` is currently the one file that says `false`: it was
|
|
transcribed from a prose summary rather than diffed against a running implementation.
|
|
|
|
### How a consumer proves it conforms
|
|
|
|
Run every `conformance/<case>/input.txt` through your detector, serialize the result per
|
|
[`schema/finding.schema.json`](schema/finding.schema.json), and compare to
|
|
`expected.json`. Disagreement means your runtime is wrong, or the fixture is — and the
|
|
fixture only changes in its own commit, with the reason written down.
|
|
|
|
There is **no CI in this organisation** and nothing runs that comparison automatically. It
|
|
runs in each consumer's own test suite, against a pinned tag.
|
|
|
|
## Non-goals
|
|
|
|
- **Not a scanner.** There is no engine here, and there will not be one. If you are looking
|
|
for something to run, you want a consumer — `llm-security` for Claude Code.
|
|
- **Not a framework or a library.** No package manifest, no dependencies, no build.
|
|
- **Not a general-purpose Unicode or regex toolkit.** The tables cover what the detection
|
|
classes need, not the standard.
|
|
- **Not a vulnerability feed.** No CVEs, no advisories, nothing time-sensitive. Everything
|
|
here is offline and deterministic.
|
|
- **Not a policy engine.** `calibration.json` publishes the thresholds; deciding what to do
|
|
when one is crossed belongs to the consumer.
|
|
- **Not the place to fix a consumer's behaviour.** Data extracted from an implementation is
|
|
kept behaviour-identical on purpose. A disagreement is reported to that implementation
|
|
and decided there, where it is tested.
|
|
|
|
## Known limitations
|
|
|
|
- **Coverage is the union of what the seed implementations detected**, not of what exists.
|
|
A class absent from `conformance/` has not been shown to work anywhere.
|
|
- **Regex portability is a real risk.** Pattern data is written for a common subset, but
|
|
engines differ (lookbehind, named groups, Unicode property escapes). A consumer whose
|
|
engine rejects a pattern must report it rather than silently skip it — a skipped pattern
|
|
is an invisible false negative.
|
|
- **Fixtures prove agreement, not correctness.** Two runtimes passing the same corpus agree
|
|
with each other and with the fixture author. A wrong `expected.json` makes both wrong
|
|
identically.
|
|
- **The homoglyph map is finite.** Confusable coverage is a long tail; absence from the map
|
|
is not evidence a character is safe.
|
|
|
|
## Changelog
|
|
|
|
See [CHANGELOG.md](CHANGELOG.md).
|
|
|
|
## License
|
|
|
|
MIT — see [LICENSE](LICENSE). Fork-and-own is an intended use, not a tolerated one.
|