`_ZERO_WIDTH` (sanitize, input) and `_ZERO_WIDTH_CPS` (output,
`_scan_invisible_carriers`) tested U+200D on codepoint membership alone.
`disposition._CARRIER_LABELS` grades both as any-tier FAIL_SECURE with no
appeal, so any first-party document containing a ZWJ-composed emoji --
professions, families, skin tones, flag variants -- was hard-blocked forever.
Reported by ms-ai-architect; confirmed here against the code.
The strip was the worse half and was not in the report: sanitize *removed* the
joiner, silently decomposing the emoji into two unrelated ones. A module whose
contract is "only ever removes carriers" was corrupting content.
The fix is the one our own lexicon row `unicode:zero-width-in-word` (`\w[ZW]\w`)
already used: judge the joiner by CONTEXT, not identity. A ZWJ is exempt only
when BOTH neighbours are emoji-context codepoints. Half-context is not context,
so `a<ZWJ>{emoji}` stays a carrier and an attacker cannot buy exemption with a
single emoji.
Blocks, not an emoji table. Measured against Unicode 17.0's
`emoji-zwj-sequences.txt`: 1614 RGI sequences use 122 distinct codepoints
adjacent to a ZWJ, and the five ranges cover 122/122. The measurement earned
its keep -- the hand-reasoned candidate table missed U+2194, U+2195 and U+2B1B.
Shipping the RGI list itself would be exact on the day it landed and stale at
the next Unicode release, reopening this same false positive for every new
emoji; whole blocks carry the unassigned headroom (458 Cn codepoints) that
future emoji are allocated into, so the table does not age.
The predicate is defined once in sanitize and imported by output. A second copy
is how the input side stops flagging while the output side keeps blocking; the
cross-surface test asserts the two agree on six inputs.
Two residuals, both in LIMITATIONS (33 -> 34, README bumped): a ZWJ between two
emoji is now exempt and could carry a covert channel (one emoji per bit, cannot
split a word); and U+200C (ZWNJ) still has no context test, so Persian, Arabic
and Devanagari documents -- where it is orthographically required -- stay
blocked. That needs a script-based criterion and no corpus is here to verify it
against, so it is parked as a known FP class rather than guessed at.
736 green (was 727), coverage 128/128, 6/6 documented gaps hold.
350 lines
18 KiB
Python
350 lines
18 KiB
Python
"""output — the report-only OUTPUT gate (compose + secret egress).
|
|
|
|
Query-time guardrails guard the answer; this guards the *persisted artifact*.
|
|
``output`` is the last gate before model output is written to a wiki, doc, or
|
|
knowledge base: it scans the emitted text and reports what must not be persisted.
|
|
It is **report-only** (design principles 3 & 4) — it never mutates the text.
|
|
Mutation is ``neutralize``'s separate, opt-in job; disposition (WARN /
|
|
QUARANTINE_REVIEW / FAIL_SECURE) is the caller's, decided from this ``Report``.
|
|
|
|
It composes the existing detectors over the output and adds the egress layer the
|
|
input-side scanners do not cover:
|
|
|
|
1. :func:`~llm_ingestion_guard.lexicon.scan_lexicon` over the output — injection
|
|
strings the model reproduced into the artifact (RAG poisoning).
|
|
2. :func:`~llm_ingestion_guard.entropy.scan_entropy` over the output — encoded /
|
|
high-entropy carrier blobs.
|
|
3. **Decode-and-rescan** — every base64 blob ``entropy`` decoded to printable
|
|
text is fed back through ``scan_lexicon`` **and** ``scan_secret_egress``.
|
|
This is what turns "a blob is present" into "an injection — or a wrapped
|
|
credential — is hidden *inside* this blob". Findings from the decoded
|
|
plaintext are re-labelled ``decoded:<label>`` (e.g.
|
|
``decoded:egress:aws-access-key-id``) and carry the blob's offset in the
|
|
original text. (Scope: base64 only — ``entropy`` exposes decoded plaintext
|
|
for base64, not hex; a hex-*wrapped* secret stays a documented honest-limit.)
|
|
4. **Secret / credential egress** (:func:`scan_secret_egress`, OWASP LLM02 —
|
|
Sensitive Information Disclosure) — cloud/provider API keys, PEM private-key
|
|
headers, DB connection strings, JWTs, and labelled password/secret/api-key
|
|
assignments, with false-positive suppression for placeholders and variable
|
|
references. Ported from the ``llm-security`` ``knowledge/secrets-patterns.md``
|
|
seed. Classic PII (email / national-ID / card numbers) is intentionally out
|
|
of scope for v1 — high false-positive risk, not in the seed.
|
|
5. **Invisible carrier presence** (:func:`_scan_invisible_carriers`) — zero-width
|
|
and BIDI override/isolate characters in the emitted artifact. Output is
|
|
report-only and never sanitized, so this is the persist-gate analogue of
|
|
``sanitize``'s input-side stripping; disposition treats the labels as
|
|
any-tier carriers. Unicode-tag / PUA stego is already surfaced by step 1.
|
|
6. **Active content** (:func:`~llm_ingestion_guard.active_content.scan_active_content`,
|
|
OWASP LLM05 — Improper Output Handling) — markdown images/links, reference
|
|
definitions, autolinks, raw active HTML and ``data:`` URIs with an external
|
|
target: the zero-click EchoLeak exfil class (CVE-2025-32711). Report-only;
|
|
``neutralize`` remains the separate, opt-in defanger of the same constructs.
|
|
|
|
**Security property (this module specifically).** A finding's ``evidence`` never
|
|
contains the secret value it matched — only a human description and the match
|
|
length. The report is meant to be logged; it must not become the leak.
|
|
|
|
**Self-safety (OWASP LLM10).** The output is capped once to ``max_scan_chars``
|
|
and a single ``output:oversize-input`` finding is emitted if it was truncated;
|
|
every sub-scanner then sees bounded input.
|
|
|
|
Bounded input is not by itself bounded runtime, and this module used to claim it
|
|
was. The egress patterns have no nested quantifiers — that part was true — but
|
|
absence of nesting does not imply linearity. A run in front of a *required*
|
|
literal (here: the password run before ``@``) makes every start position rescan
|
|
the tail when the literal never arrives, which is quadratic in the scanned
|
|
length. Crafted input repeating ``redis://:`` measured 8.2s at 100_000 chars and
|
|
extrapolated to hours at the 1_000_000-char cap this gate itself accepts. The
|
|
connection-string runs are therefore bounded to
|
|
:data:`~llm_ingestion_guard.calibration.MAX_CONNSTR_VALUE`; the same defect in
|
|
the active-content table is fixed there by excluding the anchor character. Both
|
|
are pinned by ``tests/test_output.py::test_crafted_redos_payload_stays_bounded``.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import re
|
|
from dataclasses import dataclass, replace
|
|
from typing import Optional, Union
|
|
|
|
from .active_content import scan_active_content
|
|
from .calibration import MAX_CONNSTR_VALUE
|
|
from .entropy import scan_entropy
|
|
from .lexicon import MAX_SCAN_CHARS, scan_lexicon
|
|
from .report import Finding, Report, Severity, Source
|
|
from .sanitize import _is_joiner_in_emoji_sequence
|
|
|
|
# --- secret / credential egress patterns (OWASP LLM02) ----------------------
|
|
# Ported from knowledge/secrets-patterns.md. ``value_group`` names the capturing
|
|
# group holding the matched secret *value*: when set, value-based false-positive
|
|
# suppression (placeholder / variable-ref / all-same-char / short) is applied to
|
|
# it — these are the low-specificity "labelled" patterns the seed warns about.
|
|
# When ``value_group`` is None the pattern is a high-specificity token (a unique
|
|
# prefix + length) with, per the seed, no known false positives; it is always
|
|
# reported. PEM headers use ``-{5}`` rather than five literal dashes so neither
|
|
# this source nor a scanned example trips a secret scanner on the module itself.
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class _SecretPattern:
|
|
id: str
|
|
regex: re.Pattern[str]
|
|
severity: Severity
|
|
desc: str
|
|
value_group: Optional[Union[int, str]] = None
|
|
|
|
|
|
def _p(pattern: str, flags: int = 0) -> re.Pattern[str]:
|
|
return re.compile(pattern, flags)
|
|
|
|
|
|
_I = re.IGNORECASE
|
|
|
|
_SECRET_PATTERNS: list[_SecretPattern] = [
|
|
# --- cloud / provider keys (high-specificity: always reported) ----------
|
|
_SecretPattern("aws-access-key-id", _p(r"\bAKIA[0-9A-Z]{16}\b"),
|
|
Severity.CRITICAL, "AWS access key ID"),
|
|
_SecretPattern("gcp-api-key", _p(r"\bAIza[0-9A-Za-z_\-]{35}\b"),
|
|
Severity.HIGH, "Google Cloud / Firebase API key"),
|
|
_SecretPattern("gcp-service-account-json", _p(r'"type"\s*:\s*"service_account"'),
|
|
Severity.CRITICAL, "GCP service-account credential marker"),
|
|
_SecretPattern("github-pat-classic", _p(r"\bghp_[A-Za-z0-9]{36}\b"),
|
|
Severity.CRITICAL, "GitHub classic personal access token"),
|
|
_SecretPattern("github-pat-fine-grained", _p(r"\bgithub_pat_[A-Za-z0-9_]{82}\b"),
|
|
Severity.CRITICAL, "GitHub fine-grained personal access token"),
|
|
_SecretPattern("github-oauth-token", _p(r"\bgho_[A-Za-z0-9]{36}\b"),
|
|
Severity.CRITICAL, "GitHub OAuth access token"),
|
|
_SecretPattern("github-server-token", _p(r"\bghs_[A-Za-z0-9]{36}\b"),
|
|
Severity.HIGH, "GitHub App / Actions token"),
|
|
_SecretPattern("npm-token", _p(r"\bnpm_[A-Za-z0-9]{36}\b"),
|
|
Severity.CRITICAL, "npm automation / publish token"),
|
|
_SecretPattern("openai-api-key-legacy",
|
|
_p(r"\bsk-[A-Za-z0-9]{20}T3BlbkFJ[A-Za-z0-9]{20}\b"),
|
|
Severity.CRITICAL, "OpenAI API key (legacy format)"),
|
|
_SecretPattern("openai-project-key", _p(r"\bsk-proj-[A-Za-z0-9\-_]{40,}\b"),
|
|
Severity.CRITICAL, "OpenAI project-scoped API key"),
|
|
_SecretPattern("anthropic-api-key", _p(r"\bsk-ant-api03-[A-Za-z0-9\-_]{93}\b"),
|
|
Severity.CRITICAL, "Anthropic Claude API key"),
|
|
_SecretPattern("azure-storage-key", _p(r"AccountKey=([A-Za-z0-9+/]{86}==)", _I),
|
|
Severity.CRITICAL, "Azure Storage account key"),
|
|
# --- PEM private-key headers (header alone is sufficient to flag) --------
|
|
_SecretPattern("rsa-private-key", _p(r"-{5}BEGIN RSA PRIVATE KEY-{5}"),
|
|
Severity.CRITICAL, "PEM RSA private key header"),
|
|
_SecretPattern("ec-private-key",
|
|
_p(r"-{5}BEGIN (?:EC|DSA|OPENSSH|ENCRYPTED) PRIVATE KEY-{5}"),
|
|
Severity.CRITICAL, "PEM EC/DSA/OpenSSH private key header"),
|
|
_SecretPattern("pkcs8-private-key", _p(r"-{5}BEGIN PRIVATE KEY-{5}"),
|
|
Severity.CRITICAL, "PEM PKCS#8 private key header"),
|
|
# --- DB connection strings (suppress placeholder passwords) -------------
|
|
# The password run is bounded at MAX_CONNSTR_VALUE per the ReDoS note on
|
|
# _SECRET_PATTERNS above. Unlike the active-content table, excluding the
|
|
# anchor character is NOT available here: the anchor opens with `/`, and a
|
|
# password containing `/` is the common case (a base64-ish secret), so
|
|
# excluding it would drop real credentials. The bound is the lesser loss.
|
|
_SecretPattern("postgres-connstr",
|
|
_p(r"postgres(?:ql)?://[^:@\s]+:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
|
|
Severity.CRITICAL, "PostgreSQL connection string with credentials",
|
|
value_group="val"),
|
|
_SecretPattern("mongodb-connstr",
|
|
_p(r"mongodb(?:\+srv)?://[^:@\s]+:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
|
|
Severity.CRITICAL, "MongoDB connection string with credentials",
|
|
value_group="val"),
|
|
_SecretPattern("mysql-connstr",
|
|
_p(r"mysql(?:2)?://[^:@\s]+:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
|
|
Severity.CRITICAL, "MySQL/MariaDB connection string with credentials",
|
|
value_group="val"),
|
|
_SecretPattern("redis-connstr",
|
|
_p(r"redis://:(?P<val>[^@\s]{1,%d})@[^\s'\"]+" % MAX_CONNSTR_VALUE),
|
|
Severity.HIGH, "Redis connection string with password",
|
|
value_group="val"),
|
|
# --- JWT (high false-positive rate -> MEDIUM, flag for review) ----------
|
|
_SecretPattern("jwt-token",
|
|
_p(r"\beyJ[A-Za-z0-9\-_]+\.[A-Za-z0-9\-_]+\.[A-Za-z0-9\-_]+\b"),
|
|
Severity.MEDIUM, "JSON Web Token"),
|
|
# --- labelled / generic (value-based FP suppression applied) ------------
|
|
_SecretPattern("generic-api-key",
|
|
_p(r"\bapi[_\-]?key\s*[:=]\s*[\"']([A-Za-z0-9\-._]{16,64})[\"']", _I),
|
|
Severity.HIGH, "generic api_key assignment", value_group=1),
|
|
_SecretPattern("bearer-token",
|
|
_p(r"Authorization\s*[:=]\s*[\"']?Bearer\s+([A-Za-z0-9\-._~+/]+=*)", _I),
|
|
Severity.HIGH, "Bearer token in Authorization header", value_group=1),
|
|
_SecretPattern("azure-client-secret",
|
|
_p(r"client[_\-]?secret[\"'\s]*[:=][\"'\s]*([A-Za-z0-9~._\-]{34,40})", _I),
|
|
Severity.CRITICAL, "Azure AD client secret", value_group=1),
|
|
_SecretPattern("config-password",
|
|
_p(r"(?:^|[\s,;{(])\bpass(?:word|wd)?\s*[:=]\s*[\"']([^\"'$<>{}\s]{6,})[\"']", _I),
|
|
Severity.HIGH, "password assignment", value_group=1),
|
|
_SecretPattern("config-secret",
|
|
_p(r"(?:^|[\s,;{(])\bsecret\b\s*[:=]\s*[\"']([^\"'$<>{}\s]{8,})[\"']", _I),
|
|
Severity.HIGH, "secret assignment", value_group=1),
|
|
]
|
|
|
|
# Value fragments that mark a placeholder / template rather than a real secret.
|
|
# Structural markers are legitimate anywhere in the value (a template fragment
|
|
# like ``<your-key>`` or ``your-api-key-here``), so they match as a substring.
|
|
_PLACEHOLDER_STRUCTURAL = ("your-", "your_", "<", ">", "***")
|
|
# Word markers are matched on a word boundary, NOT as a bare substring: a real
|
|
# secret that merely *contains* one ("todoAppSecretKey12" contains "todo") must
|
|
# not be suppressed — that would be a fail-open egress miss. Genuine placeholders
|
|
# ("example-secret", "changeme") still match at their boundaries.
|
|
_PLACEHOLDER_WORDS = (
|
|
"example", "placeholder", "replace", "changeme",
|
|
"xxx", "todo", "fixme", "dummy", "sample",
|
|
)
|
|
_PLACEHOLDER_WORD_RE = re.compile(
|
|
r"\b(?:" + "|".join(_PLACEHOLDER_WORDS) + r")\b", re.IGNORECASE
|
|
)
|
|
_VARREF_TOKENS = ("${", "$(", "%{", "env[", "os.environ", "process.env")
|
|
|
|
|
|
def _is_fp_value(value: str) -> bool:
|
|
"""True if ``value`` is a placeholder / variable ref / trivial, not a secret."""
|
|
low = value.lower()
|
|
if any(token in low for token in _PLACEHOLDER_STRUCTURAL):
|
|
return True
|
|
if _PLACEHOLDER_WORD_RE.search(value):
|
|
return True
|
|
if any(token in low for token in _VARREF_TOKENS):
|
|
return True
|
|
if len(set(value)) == 1: # all-same-character (e.g. "xxxxxxxx")
|
|
return True
|
|
if len(value) < 8: # too short for a generic-pattern secret
|
|
return True
|
|
return False
|
|
|
|
|
|
def scan_secret_egress(text: str, source: Source = Source.OUTPUT) -> Report:
|
|
"""Scan ``text`` for secret/credential egress; report-only, evidence-safe.
|
|
|
|
Each finding is labelled ``egress:<id>`` with OWASP ``LLM02``. Evidence is a
|
|
human description plus the match length — never the secret value itself.
|
|
"""
|
|
report = Report()
|
|
for pattern in _SECRET_PATTERNS:
|
|
for match in pattern.regex.finditer(text):
|
|
if pattern.value_group is not None:
|
|
value = match.group(pattern.value_group)
|
|
if value is None or _is_fp_value(value):
|
|
continue
|
|
report.add(
|
|
Finding(
|
|
label=f"egress:{pattern.id}",
|
|
severity=pattern.severity,
|
|
source=source,
|
|
detector="output",
|
|
offset=match.start(),
|
|
owasp="LLM02",
|
|
# Length only — the evidence must never carry the secret.
|
|
evidence=f"{pattern.desc} (match len {len(match.group(0))})",
|
|
)
|
|
)
|
|
return report
|
|
|
|
|
|
# --- invisible carrier presence (BRIEF §4.7) --------------------------------
|
|
# The output-gate analogue of sanitize's input-side stripping: model output is
|
|
# report-only and never sanitized, so an invisible carrier reaching the persist
|
|
# gate must be flagged here. Zero-width / soft-hyphen and BIDI override/isolate
|
|
# code points; Unicode-tag / PUA stego is already surfaced by scan_lexicon
|
|
# (``lexicon:unicode-tags-present``), so it is not repeated. Disposition treats
|
|
# these labels as any-tier carriers (FAIL_SECURE regardless of trust).
|
|
_ZERO_WIDTH_CPS = frozenset({0x200B, 0x200C, 0x200D, 0xFEFF, 0x00AD})
|
|
_BIDI_CPS = frozenset({0x202A, 0x202B, 0x202C, 0x202D, 0x202E, 0x2066, 0x2067, 0x2068, 0x2069})
|
|
|
|
|
|
def _scan_invisible_carriers(text: str, source: Source) -> Report:
|
|
"""Flag invisible zero-width / BIDI carriers present in ``text`` (report-only)."""
|
|
report = Report()
|
|
# The ZWJ exemption is imported from `sanitize`, never re-stated here: the
|
|
# two surfaces are one decision, and a second copy of the rule is how the
|
|
# input side stops flagging while the output side keeps hard-blocking.
|
|
zero_width = sum(
|
|
1 for i, ch in enumerate(text)
|
|
if ord(ch) in _ZERO_WIDTH_CPS
|
|
and not (ord(ch) == 0x200D and _is_joiner_in_emoji_sequence(text, i))
|
|
)
|
|
bidi = sum(1 for ch in text if ord(ch) in _BIDI_CPS)
|
|
if zero_width:
|
|
report.add(Finding(
|
|
label="output:zero-width-present", severity=Severity.HIGH,
|
|
source=source, detector="output", count=zero_width, owasp="LLM01",
|
|
evidence="invisible zero-width/soft-hyphen characters present",
|
|
))
|
|
if bidi:
|
|
report.add(Finding(
|
|
label="output:bidi-present", severity=Severity.HIGH,
|
|
source=source, detector="output", count=bidi, owasp="LLM01",
|
|
evidence="BIDI override/isolate characters present",
|
|
))
|
|
return report
|
|
|
|
|
|
def scan_output(
|
|
text: str,
|
|
source: Source = Source.OUTPUT,
|
|
max_scan_chars: int = MAX_SCAN_CHARS,
|
|
) -> Report:
|
|
"""Scan model OUTPUT before persist; return the merged, report-only findings.
|
|
|
|
Composes lexicon + entropy + decode-and-rescan + secret egress over ``text``.
|
|
Never mutates ``text``. Bounds runtime by capping the scanned length once.
|
|
"""
|
|
report = Report()
|
|
|
|
truncated = len(text) > max_scan_chars
|
|
scan_text = text[:max_scan_chars] if truncated else text
|
|
if truncated:
|
|
report.add(
|
|
Finding(
|
|
label="output:oversize-input",
|
|
severity=Severity.MEDIUM,
|
|
source=source,
|
|
detector="output",
|
|
count=len(text),
|
|
owasp="LLM10",
|
|
evidence=f"output {len(text)} chars exceeds cap {max_scan_chars}; scanned prefix only",
|
|
)
|
|
)
|
|
|
|
# 1. Injection strings the model reproduced into the artifact. scan_text is
|
|
# already <= cap, so the lexicon will not emit a second oversize finding.
|
|
report.extend(scan_lexicon(scan_text, source, max_scan_chars).findings)
|
|
|
|
# 2. Encoded / high-entropy carrier blobs, plus the decoded plaintext blobs.
|
|
entropy_result = scan_entropy(scan_text, source)
|
|
report.extend(entropy_result.report.findings)
|
|
|
|
# 3. Decode-and-rescan: run the lexicon AND the egress scanner over each
|
|
# decoded blob's plaintext, re-labelled so the finding is attributable to
|
|
# the hiding blob. Feeding the egress set here (not only the lexicon) is
|
|
# what catches a base64-*wrapped* secret: the plaintext credential reaches
|
|
# scan_secret_egress as a decoded:egress:* finding instead of vanishing.
|
|
# (Scope: base64 only — entropy exposes decoded plaintext for base64, not
|
|
# hex; a hex-wrapped secret stays a documented honest-limit.)
|
|
for blob in entropy_result.decoded:
|
|
hidden = scan_lexicon(blob.decoded, source, max_scan_chars).findings
|
|
leaked = scan_secret_egress(blob.decoded, source).findings
|
|
for finding in [*hidden, *leaked]:
|
|
report.add(
|
|
replace(
|
|
finding,
|
|
label=f"decoded:{finding.label}",
|
|
offset=blob.offset,
|
|
evidence=f"{finding.evidence} (in base64 blob @ {blob.offset})",
|
|
)
|
|
)
|
|
|
|
# 4. Secret / credential egress (OWASP LLM02).
|
|
report.extend(scan_secret_egress(scan_text, source).findings)
|
|
|
|
# 5. Invisible carriers present in the artifact (zero-width / BIDI). Output
|
|
# is never sanitized, so this is the persist-gate carrier signal; the
|
|
# unicode-tag case is already covered by the lexicon scan in step 1.
|
|
report.extend(_scan_invisible_carriers(scan_text, source).findings)
|
|
|
|
# 6. Active-content constructs with an external target (the EchoLeak class,
|
|
# OWASP LLM05) — reported here so disposition sees them; defanging stays
|
|
# neutralize's separate, opt-in job.
|
|
# scan_text is already <= cap, so no second oversize finding is emitted.
|
|
report.extend(scan_active_content(scan_text, source, max_scan_chars).findings)
|
|
|
|
return report
|