llm-security-commons/README.md
Kjell Tore Guttormsen 47760d2264 feat(signatures): add the malware identity table, drawn from source not memory
The last missing data file. Seven known-bad-identity signatures over four
families - webshell, reverse_shell, cryptominer, hacktool - reproduced
verbatim from llm-security/knowledge/signatures.json at b0de0ca, key order
included. The file was generated from the parsed source rather than typed,
and provenance pins the source's byte length (2494) and SHA-256 so the
claim is checkable rather than asserted.

Note the family spellings: reverse_shell with an underscore, and
cryptominer rather than miner. The working note this file was planned from
had both wrong. They are policy keys - the engine filters on them and
interpolates them into every finding title - so a rename is a breaking
change, which is exactly why the table was read instead of recalled.

The rules were the easy half. The substance is engine_behaviour_not_data,
which draws the line between the table and the runtime around it. No rule
carries a flags field, because the engine compiles every pattern with `i`
unconditionally at signature-scanner.mjs:48 - so a consumer compiling these
case-sensitively silently under-matches all seven, and the dialect block
records that where a reader will hit it. Also engine, not data: matching
against five decode variants rather than raw bytes, the enabled-families
policy filter, per-file rule dedup, custom-rule merging, and a loader that
defaults four missing fields instead of rejecting a rule.

Two limits are stated as evidence limits rather than left implied. Seven
signatures are not malware coverage; a clean SIG result is not "no
malware", and the seed runtime's own header calls the table deliberately
tight. And three of the seven match on names - xmrig, mimikatz,
meterpreter - so a document discussing those tools matches. The seed
runtime hides that by excluding knowledge/, tests/, docs/ and
node_modules/, which is scan scoping and does not travel with the table.

Verified: 7/7 rule objects field-identical to source including key order,
no non-ASCII bytes, all seven compile in Node bare, i and iu (21/21) and in
Python re (7/7). Charter guard clean - no executable code in the repository.

README and CHANGELOG updated: the file moves out of "planned, not in
v0.1.0" and out of "not included".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SNMcqrfNyoLRQ7qXUFZnb9
2026-08-09 22:49:21 +02:00

120 lines
7.5 KiB
Markdown

# llm-security-commons
Runtime-neutral core for LLM and agent security detection: detector data, normative contracts and a conformance corpus that several runtimes can share.
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
Detection logic gets reimplemented every time it crosses a language boundary, and the
copies drift: the Node scanner flags a zero-width carrier the Python guard misses, and
nobody notices until an incident. This repository holds the part that should never have
been copied — the pattern tables, the code-point carriers, the calibration thresholds, the
finding contract, and a fixture corpus with expected verdicts — so that two independent
implementations can be held to the same answer on the same input.
It is for anyone building or maintaining a detector for prompt injection, secret egress,
unicode-carrier smuggling or active content in untrusted text, on any runtime.
**It holds no runnable code.** Data, specifications and fixtures only.
## Install
Nothing to install — this repository is **vendored into consumers**, not installed.
As a `git subtree` (recommended: history is preserved and upgrades are a single command):
```bash
git subtree add --prefix vendor/commons \
https://git.fromaitochitta.com/open/llm-security-commons.git v0.1.0 --squash
# later, to move to a newer tag
git subtree pull --prefix vendor/commons \
https://git.fromaitochitta.com/open/llm-security-commons.git v0.2.0 --squash
```
Or pin a tag and copy — `fork-and-own` is an explicitly supported path:
```bash
git clone --depth 1 --branch v0.1.0 \
https://git.fromaitochitta.com/open/llm-security-commons.git
```
Always vendor **a tag**, never `main`. The tag is what a conformance result can be
attributed to.
## Requirements
A JSON parser and the ability to read a text file. That is the entire dependency surface,
and keeping it that small is the point.
## What it does
| Path | Contents |
| --- | --- |
| [`lexicon/injection-lexicon.json`](lexicon/injection-lexicon.json) | Prompt-injection pattern lexicon: 83 patterns in four **severity** families (`critical`, `high`, `medium`, `hybrid`), each with a stable `id` and per-runtime aliases. The thematic class (`override:`, `evasion:`, `hitl-trap:`, …) is the id prefix, not the family. |
| [`codepoints/carriers.json`](codepoints/carriers.json) | Invisible and deceptive carriers: zero-width characters, BIDI controls, Unicode Tag block ranges, and the homoglyph map. |
| [`signatures/secret-egress.json`](signatures/secret-egress.json) | Credential and token shapes that must never leave a machine, in a portable regex dialect. |
| [`signatures/malware-signatures.json`](signatures/malware-signatures.json) | Known-bad **identity** for the malicious-code class (`SIG`): seven tight signatures over four families — PHP webshells, reverse shells, cryptominers, offensive tooling. Seven signatures are not malware coverage, and the file says so. |
| [`signatures/active-content.json`](signatures/active-content.json) | Active content that renders or fetches on its own — Markdown images, links, reference definitions and autolinks, `data:` URIs, active HTML. The EchoLeak class. |
| [`calibration/calibration.json`](calibration/calibration.json) | The numbers a detector must not invent: risk-score tier constants, verdict thresholds, risk-band cutoffs, posture grade thresholds. Transcribed from a prose summary, not differentially verified — the file says so itself. |
| [`mapping/owasp-map.json`](mapping/owasp-map.json) | Finding-id prefix → OWASP taxonomy entry (LLM / ASI / AST / MCP). |
| [`schema/finding.schema.json`](schema/finding.schema.json) | **Normative.** The finding contract — closed against its producer, ten properties — plus the SARIF output profile. The JSONL profile is recorded as `not applicable`, with the reason. |
| `spec/decode-pipeline.md` | **Planned, not in v0.1.0.** The decode order, in RFC 2119 language. Two runtimes that decode in different orders will disagree on identical input. Writing it needs the decode implementation, which is engine code and has not been supplied — and a normative spec guessed from a data dump would be worse than an absent one. |
| `conformance/` | **Planned, not in v0.1.0.** One directory per case: `input.txt` in, `expected.json` out. Ground truth. |
| [`docs/extraction-plan.md`](docs/extraction-plan.md) | Informative: where each file was seeded from, and what v0.1.0 promised. |
| [`docs/lexicon-port-divergence.md`](docs/lexicon-port-divergence.md) | Informative: a measured disagreement between two ports of the injection lexicon — 13 patterns that behave differently, in both directions, and why no data file was changed because of it. |
Every JSON file carries a top-level `version`. Every normative specification carries a
`Status: normative` marker. Rows marked **Planned** are named here because the layout is
part of the contract, but the file does not exist yet — they are not links, and nothing in
v0.1.0 depends on them.
Each data file records its own provenance and, in `verified`, how strongly it is backed.
`calibration/calibration.json` is currently the one file that says `false`: it was
transcribed from a prose summary rather than diffed against a running implementation.
### How a consumer proves it conforms
Run every `conformance/<case>/input.txt` through your detector, serialize the result per
[`schema/finding.schema.json`](schema/finding.schema.json), and compare to
`expected.json`. Disagreement means your runtime is wrong, or the fixture is — and the
fixture only changes in its own commit, with the reason written down.
There is **no CI in this organisation** and nothing runs that comparison automatically. It
runs in each consumer's own test suite, against a pinned tag.
## Non-goals
- **Not a scanner.** There is no engine here, and there will not be one. If you are looking
for something to run, you want a consumer — `llm-security` for Claude Code.
- **Not a framework or a library.** No package manifest, no dependencies, no build.
- **Not a general-purpose Unicode or regex toolkit.** The tables cover what the detection
classes need, not the standard.
- **Not a vulnerability feed.** No CVEs, no advisories, nothing time-sensitive. Everything
here is offline and deterministic.
- **Not a policy engine.** `calibration.json` publishes the thresholds; deciding what to do
when one is crossed belongs to the consumer.
- **Not the place to fix a consumer's behaviour.** Data extracted from an implementation is
kept behaviour-identical on purpose. A disagreement is reported to that implementation
and decided there, where it is tested.
## Known limitations
- **Coverage is the union of what the seed implementations detected**, not of what exists.
A class absent from `conformance/` has not been shown to work anywhere.
- **Regex portability is a real risk.** Pattern data is written for a common subset, but
engines differ (lookbehind, named groups, Unicode property escapes). A consumer whose
engine rejects a pattern must report it rather than silently skip it — a skipped pattern
is an invisible false negative.
- **Fixtures prove agreement, not correctness.** Two runtimes passing the same corpus agree
with each other and with the fixture author. A wrong `expected.json` makes both wrong
identically.
- **The homoglyph map is finite.** Confusable coverage is a long tail; absence from the map
is not evidence a character is safe.
## Changelog
See [CHANGELOG.md](CHANGELOG.md).
## License
MIT — see [LICENSE](LICENSE). Fork-and-own is an intended use, not a tolerated one.