Two changes that had to ship together, because they co-occur. `active:raw-html-link` (MEDIUM) splits the click-required carriers out of `active:raw-html`. The same URL was LOW as `[t](url)` and HIGH as `<a href="url">` — an asymmetry produced by syntax, not by affordance, on a carrier the markdown path has graded MEDIUM since 0.3.1. The event-handler test runs first, so `<a onclick=...>` stays HIGH. The url-attribute branch stays HIGH too: a name outside the active set has unknown rendering, and grading `<Card src=...>` as a link would be reasoning rather than measurement. The no-URL narrowing makes `</a>`, `<Frame>`, `<video />` and `<img alt=...>` without `src` inert — `<base />`'s argument from 0.6.0 applied to the rest of the name branch. It tests for the URL attribute's PRESENCE, not for a readable value, so the fail-secure gap `_url_attr_is_external` leaves open is not reopened here. WHY TOGETHER: the narrowing strips a document's `</a>`/`<Frame>` and what remains is the `<a href=...>` the split grades down, so each alone leaves the document blocked by the other's residue. `active_tag_class` is now the classification point and `is_active_tag` wraps it. The census patches the former: a boolean could only express a narrowing, never a regrade, so every carrier candidate would have measured equal to PRODUCTION — silently, and in the direction that reads as "no change helps". TWO COSTS, BOTH RECORDED RATHER THAN GLOSSED: - The split TIGHTENS the trusted tier. One finding becomes two, and >=2 findings at MEDIUM+ trip the compound overlay, so a document carrying both an `<img src>` and an `<a href>` goes WARN -> quarantine_review on PRESET_TRUSTED_SOURCE. On that preset it is the only direction the split can move anything. The census now reports a TIGHTENS column on both trust tiers against the previously shipped row — "frees N" without "tightens M" is a one-sided number. - `count` drops on documents containing `</a>`, a published field moving under a meaning that did not change. MEASURED: reference-corpus (389) 54 -> 53 fail_secure, tightens 0/0, and the census `PRODUCTION` row equals its `C1 + D` candidate row for row. The census also reproduces 133/3/13/108/25 exactly, so it is calibrated against every published historical number. The two wiki corpora are NOT yet re-measured; the tree says so explicitly in the docstring, LIMITATIONS and CHANGELOG rather than carrying probe numbers as fact. 791 tests (was 759), coverage 129/129, 6/6 documented gaps holding. Version bumped to 0.7.0 across every surface; no tag is set until the measurement lands.
2.5 KiB
Security policy
llm-ingestion-guard is a defensive library for LLM ingestion pipelines. Its own
security posture matters: a flaw here can silently admit a poisoned artifact into a
downstream corpus. Reports are welcome.
Supported versions
The project is pre-1.0 (0.7.x, alpha). Only the latest published version receives
fixes; there are no back-ported security branches yet. Pin a version and watch the
CHANGELOG.md ### Security entries.
Reporting a vulnerability
Do not open a public issue for a vulnerability. Public disclosure before a fix gives an attacker a window against every downstream consumer.
Instead, report it privately to the maintainer via the canonical repository on Forgejo:
- Repository:
git.fromaitochitta.com/open/llm-ingestion-pipeline-security - Contact the maintainer directly through that Forgejo instance (private message /
maintainer contact) and mark the subject
SECURITY.
Please include:
- affected version / commit,
- a minimal reproduction (input → observed disposition/finding vs. expected),
- the impact you see (e.g. a poisoned artifact that disposes
WARNinstead ofFAIL_SECURE).
Obfuscate any real payloads the same way the test corpus does — build attack strings
from chr(0x…) fragments so the report itself does not ship a live carrier.
What counts as a vulnerability
In scope (a real finding):
- a bypass of a stated control — e.g. an invisible carrier that reaches the
persist gate without failing secure, a credential that egresses without a
decoded:egress:*/egress:*label, aguard()path that fails open; - a
prepare_input/screen_outputcode path that raises instead of failing closed; - a ReDoS or unbounded-resource input against the scanner.
Out of scope (documented boundaries — see the Known limitations section of
README.md, not vulnerabilities):
- semantic / factual poisoning invisible to lexicon + entropy;
- a HIGH finding in trusted prose disposing to
WARN(§4.7 trust-scaling); - hex-wrapped (non-base64) secret egress;
- multimodal / binary-layer carriers (OCR, font stego, VBA/macros, encrypted files);
- the multilingual homoglyph-mix false positive.
If you are unsure whether something is in scope, report it privately anyway.
Disclosure
This is a small project without a formal embargo SLA. The maintainer will
acknowledge a report, agree a fix + disclosure timeline with the reporter, and
credit the reporter in the CHANGELOG.md ### Security entry unless they prefer to
remain anonymous.