Runtime-neutral core for LLM and agent security detection: detector data, normative contracts and a conformance corpus that several runtimes can share.
Find a file
Kjell Tore Guttormsen e6ca5ae5ee feat(active-content): the seventh case, and the classifier it needed came with it
Adopting `active:raw-html-link` was one id. Publishing it honestly was the whole
of `active_tag_class` — one function, three branches, no way to state the split
without the no-URL narrowing and the 0.6.0 external-target rule. On the old
predicate a bare `</a>` is active by name, so a consumer implementing from the
hybrid would emit the new label where the seed runtime emits nothing.

The file is now two pins, stated as two: v0.3.4/0bf0729 everywhere except the
raw-HTML classifier, v0.7.0/be9759b there. The drift between them was measured
field by field against the imported module rather than assumed, after stripping
inline-flag rendering and applying the file's own declared quote normalisation
so a spelling difference could not masquerade as drift. Exactly one published
field had moved, and not the one this release was about: `html.active_tags`
carried the MUTATOR's 23-name set where the gate means the SCANNER's 22. Correct
at the 0.3.4 pin, wrong from 0.6.0 on. Kept as `html.mutator_tags`.

The sweep covered 93 cases, not the 6 obvious ones. The narrowing can silence an
`active:` finding inside the `observed_out_of_scope` evidence of a LEXICON case,
and that field is guarded by no test anywhere — stale entries there survive
forever. One case moved: html-obfuscation__aria-label, whose `<a aria-label=…>`
carries no URL attribute. Its fixture is deliberately not rewritten; the residue
is true at the commit `measurement` pins, and rewriting one of 83 would leave two
commits under a header naming one. Recorded, dated and pinned in the manifest.

The strongest check is not the digest: the checker rebuilds the published
classifier from the JSON alone, importing nothing from the runtime, and
differential-tests it against `active_tag_class` over 42 probe tags. 0
disagreements. That is what licenses shipping a classifier as data.

No `aliases.llm_security` published, on this file or on carriers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LTTaT4quwNPwBYVqmgAt8t
2026-08-13 21:43:27 +02:00
calibration feat(calibration): add calibration.json, marked transcribed-only 2026-08-09 21:14:13 +02:00
codepoints feat(carriers): three cases minted, and the id is named rather than adopted 2026-08-13 21:17:44 +02:00
conformance feat(active-content): the seventh case, and the classifier it needed came with it 2026-08-13 21:43:27 +02:00
docs fix(divergence): our own iframe number read 3x low, and the reported cause was not the cause 2026-08-11 22:02:15 +02:00
lexicon feat(lexicon): both unbounded rows narrow to [^><]*, and the mechanism is new here 2026-08-11 21:52:43 +02:00
mapping fix(mapping): state that three of four OWASP maps have no production consumer 2026-08-10 20:47:37 +02:00
schema feat(schema,spec): give the §1.1 MUST a shape, since v0.2.0 shipped it without one 2026-08-11 13:39:37 +02:00
signatures feat(active-content): the seventh case, and the classifier it needed came with it 2026-08-13 21:43:27 +02:00
spec feat(schema,spec): give the §1.1 MUST a shape, since v0.2.0 shipped it without one 2026-08-11 13:39:37 +02:00
.gitignore feat: initialize llm-security-commons (charter, license, extraction plan) 2026-08-09 14:29:16 +02:00
CHANGELOG.md feat(active-content): the seventh case, and the classifier it needed came with it 2026-08-13 21:43:27 +02:00
CLAUDE.md feat(active-content): the seventh case, and the classifier it needed came with it 2026-08-13 21:43:27 +02:00
CONVENTIONS.md docs(conventions): the merge button is off for a reason, and the reason now lives in the repo 2026-08-11 22:12:37 +02:00
LICENSE feat: initialize llm-security-commons (charter, license, extraction plan) 2026-08-09 14:29:16 +02:00
README.md feat(active-content): the seventh case, and the classifier it needed came with it 2026-08-13 21:43:27 +02:00
SECURITY.md docs(security): the attack surface here is data, so the report route had to say where a wrong entry gets fixed 2026-08-11 14:03:24 +02:00

llm-security-commons

Runtime-neutral core for LLM and agent security detection: detector data, normative contracts and a conformance corpus that several runtimes can share.

License: MIT

Detection logic gets reimplemented every time it crosses a language boundary, and the copies drift: the Node scanner flags a zero-width carrier the Python guard misses, and nobody notices until an incident. This repository holds the part that should never have been copied — the pattern tables, the code-point carriers, the calibration thresholds, the finding contract, and a fixture corpus with expected verdicts — so that two independent implementations can be held to the same answer on the same input.

It is for anyone building or maintaining a detector for prompt injection, secret egress, unicode-carrier smuggling or active content in untrusted text, on any runtime.

It holds no runnable code. Data, specifications and fixtures only.

Install

Nothing to install — this repository is vendored into consumers, not installed.

As a git subtree (recommended: history is preserved and upgrades are a single command):

git subtree add --prefix vendor/commons \
  https://git.fromaitochitta.com/open/llm-security-commons.git v0.6.0 --squash

# later, to move to a newer tag
git subtree pull --prefix vendor/commons \
  https://git.fromaitochitta.com/open/llm-security-commons.git <newer-tag> --squash

Or pin a tag and copy — fork-and-own is an explicitly supported path:

git clone --depth 1 --branch v0.6.0 \
  https://git.fromaitochitta.com/open/llm-security-commons.git

Always vendor a tag, never main. The tag is what a conformance result can be attributed to.

Requirements

A JSON parser and the ability to read a text file. That is the entire dependency surface, and keeping it that small is the point.

What it does

Path Contents
lexicon/injection-lexicon.json Prompt-injection pattern lexicon: 83 patterns in four severity families (critical, high, medium, hybrid), each with a stable id and per-runtime aliases. The thematic class (override:, evasion:, hitl-trap:, …) is the id prefix, not the family.
codepoints/carriers.json Invisible and deceptive carriers: zero-width characters, BIDI controls, Unicode Tag block ranges, and the homoglyph map. Carries three commons-owned ids (carrier:zero-width, carrier:bidi-override, carrier:unicode-tag) for the carriers observable on an input surface — the only id space here that was named rather than adopted verbatim from a runtime, and the file records why.
signatures/secret-egress.json Credential and token shapes that must never leave a machine, in a portable regex dialect.
signatures/malware-signatures.json Known-bad identity for the malicious-code class (SIG): seven tight signatures over four families — PHP webshells, reverse shells, cryptominers, offensive tooling. Seven signatures are not malware coverage, and the file says so.
signatures/active-content.json Active content that renders or fetches on its own — Markdown images, links, reference definitions and autolinks, data: URIs, active HTML. The EchoLeak class. Raw HTML carries two classes: active:raw-html for what a renderer acts on unattended, active:raw-html-link for anchors, which need a human. One pattern, one scan, two buckets — the file spells that out, because giving the second class its own pass would double-count.
calibration/calibration.json The numbers a detector must not invent: risk-score tier constants, verdict thresholds, risk-band cutoffs, posture grade thresholds. Transcribed from a prose summary, not differentially verified — the file says so itself.
mapping/owasp-map.json Finding-id prefix → OWASP taxonomy entry (LLM / ASI / AST / MCP).
schema/finding.schema.json Normative. The finding contract — closed against its producer, ten properties — plus the SARIF output profile. The JSONL profile is recorded as not applicable, with the reason.
schema/conformance-declaration.schema.json Normative. The shape a runtime publishes alongside a conformance result: which commons tables it implements, the commons commit it measured, and the four verdict counts. Required by the corpus spec §1.1; not validated by anything here, because nothing here runs.
spec/conformance-corpus.md Normative. How to read the corpus: what a case is, why input.txt is bytes rather than text, what exact-within-scope requires of a runtime, and how a runtime declares its table set so a case scoped outside it reads as not-applicable rather than as a failure.
conformance/ 94 cases. One directory per case: input.txt in, expected.json out. Ground truth. 84 cover the injection lexicon — 83 one per pattern, both seeding runtimes measured producing the same verdict on all 83, plus one variant case gating a pattern form against its predecessor. Seven cover active content and are measured against the one runtime that implements that table — not-applicable for the other, not failing. Three cover the input-side carriers, added in v0.5.0. See conformance/manifest.json.
spec/decode-pipeline.md Planned, still not shipped as of v0.6.0. The decode order, in RFC 2119 language. Two runtimes that decode in different orders will disagree on identical input. Writing it needs the decode implementation, which is engine code and has not been supplied — and a normative spec guessed from a data dump would be worse than an absent one.
docs/extraction-plan.md Informative: where each file was seeded from, and what v0.1.0 promised.
docs/lexicon-port-divergence.md Informative: a measured disagreement between two ports of the injection lexicon — 13 patterns that behave differently, in both directions. Most of it is still open, and the two rows that closed in v0.4.0 closed because the runtime that owns the value decided, not because this document found them wrong.

Every JSON file carries a top-level version. Every normative specification carries a Status: normative marker. Rows marked Planned are named here because the layout is part of the contract, but the file does not exist yet — they are not links, and nothing in v0.6.0 depends on them.

Each data file records its own provenance and, in verified, how strongly it is backed. calibration/calibration.json is currently the one file that says false: it was transcribed from a prose summary rather than diffed against a running implementation.

How a consumer proves it conforms

Run every conformance/<case>/input.txt through your detector and compare the finding ids to expected.json — exactly, but only within the data files the case names in scope. spec/conformance-corpus.md is the normative reading; the short version is that a runtime must raise every listed finding and no other finding from the same table, and that what it does with tables outside the case's scope is not compared.

Disagreement means your runtime is wrong, or the fixture is — and the fixture only changes in its own commit, with the reason written down.

There is no CI in this organisation and nothing runs that comparison automatically. It runs in each consumer's own test suite, against a pinned tag.

The corpus covers three tables, and they do not carry equal weight — treating them as one number would misreport all three:

  • lexicon/injection-lexicon.json — 84 cases over 83 patterns. Both seeding runtimes implement it and both ratified its id space. One pattern carries a second, variant case; see case_id_derivation.variant_suffix in the manifest.
  • signatures/active-content.json — 7 cases, one per published id, the seventh added in v0.6.0 when the seed runtime split raw HTML into two carrier classes. One runtime implements it. For a runtime that does not, these cases are not-applicable, a third verdict beside pass and fail: a runtime declares which commons data files it implements, and a case scoped outside that set was never addressed to it. See §1.1 — and note that not-applicable says the corpus did not ask, never that the runtime is blind.
  • codepoints/carriers.json — 3 cases, added in v0.5.0. Both runtimes implement these tables; only one has published a label for the finding. So these three are not-applicable for the other today, and this is the one place in the corpus where that verdict records a missing name rather than a missing capability. It lapses the moment that runtime names its label and the alias is added.

One case remains unshipped, for the secret-egress table, and it is not blocked on effort. It is not an id question at all — the two runtimes carry different tables, 19 entries against 25, cut at different granularities, and a shared id space presupposes a reconciliation nobody has performed. conformance/manifest.json records that blocker under scope_planned.blockers, measured, so the gap is visible rather than inferred.

The carrier blocker closed in v0.5.0 and is kept, with its retired text, under scope_planned.blockers_resolved — including the correction one runtime volunteered against a general rule this repository had written down and should not have.

Non-goals

  • Not a scanner. There is no engine here, and there will not be one. If you are looking for something to run, you want a consumer — llm-security for Claude Code.
  • Not a framework or a library. No package manifest, no dependencies, no build.
  • Not a general-purpose Unicode or regex toolkit. The tables cover what the detection classes need, not the standard.
  • Not a vulnerability feed. No CVEs, no advisories, nothing time-sensitive. Everything here is offline and deterministic.
  • Not a policy engine. calibration.json publishes the thresholds; deciding what to do when one is crossed belongs to the consumer.
  • Not the place to fix a consumer's behaviour. Data extracted from an implementation is kept behaviour-identical on purpose. A disagreement is reported to that implementation and decided there, where it is tested.

Known limitations

  • Coverage is the union of what the seed implementations detected, not of what exists. A class absent from the tables above has not been shown to work anywhere.
  • The corpus is narrower than the data. conformance/ constrains two of the seven data files. The other five are published, provenance-checked and unfixtured: a runtime can pass every case and still read calibration.json wrongly. Passing the corpus is evidence about the injection lexicon and about active content, and about nothing else.
  • A pass count is unreadable without the declared table set. A runtime implementing one table and a runtime implementing four can print the same number. not-applicable cases must be reported, not dropped from the denominator — 76/83 and 76 passed, 6 not-applicable describe different runtimes.
  • The seven active-content cases prove less than the 83. Their payloads come from the only runtime that implements the table, so no second implementation's agreement could be measured. They pin one runtime's behaviour as a contract a future implementer can be held to; they are not cross-runtime agreement.
  • Regex portability is a real risk. Pattern data is written for a common subset, but engines differ (lookbehind, named groups, Unicode property escapes). A consumer whose engine rejects a pattern must report it rather than silently skip it — a skipped pattern is an invisible false negative.
  • Fixtures prove agreement, not correctness. Two runtimes passing the same corpus agree with each other and with the fixture author. A wrong expected.json makes both wrong identically.
  • The homoglyph map is finite. Confusable coverage is a long tail; absence from the map is not evidence a character is safe.

Contributing

CONVENTIONS.md is the whole rule set a change here is held to: the charter (nothing runs, and why that is load-bearing rather than fussy), the file conventions, when a detection value is allowed to move, how the two version numbers work, and the four offline checks that stand in for the CI this organisation does not have.

It also answers the question the forge surface raises on its own: pull requests are switched off, deliberately. This repository is vendored into independent runtimes that pin a tag, so a change to detection data changes what they find — that has to be coordinated with each consumer before it exists, which a merge button cannot do. Fork-and-own is the supported path; a wrong entry is reported privately.

Reporting a wrong entry

A wrong code point or a mis-escaped regex here is a silent false negative in every runtime that reads it, so it is a security report even though nothing runs. Send it privately — see SECURITY.md, which also explains why a confirmed defect in extracted data is decided in the runtime it came from before it is changed here.

Changelog

See CHANGELOG.md.

License

MIT — see LICENSE. Fork-and-own is an intended use, not a tolerated one.