1
0
Fork 0
llm-ingestion-pipeline-secu.../docs/LIMITATIONS.md
Kjell Tore Guttormsen 58704834b6 feat(okf): the mapping class gets one expressible form, typed and allowlisted
OKF v0.2 writes its whole trust and provenance layer as mappings, and T2 gave
the mapping class no expressible form. A consumer measured 0 of 53 upstream
concepts through the gate on 0.3.4, 1.0.0 and 1.1.0. That was a contract
collision, not a calibration setting: SPEC.md @ 62432a09 uses flow mappings in
its own 5.1/5.2 examples, and 11 carries a hard MUST for consumers ("MUST treat
a bare `verified` mapping as a one-element list") that presupposes they parse.

Admitted: a flow mapping, as a value or as a block-list item, whose every key is
on a nine-name allowlist and whose every leaf is a plain scalar run through the
unchanged dangerous-value and mapping-construct predicates. The form is safe
because the allowlist inspects every key -- the blanket refusal was the
enforcement, not the point.

Refused, each on its own rule and ground-truthed against PyYAML 6.0.3: a key off
the allowlist, a nested collection, a quoted leaf, a duplicate key, an empty or
unclosed mapping, trailing junk, and {a:b} (which PyYAML reads as the KEY a:b).
A refused mapping still raises rather than degrading into a string, so the 1.1.0
type-confusion defect is not reopened, and the block, dotted and inline-colon
routes still raise.

`resource` is deliberately off the allowlist though SPEC.md 5.1 names it inside
a `sources` entry: it is a pointer rather than a label and the only key T3
exists for, so admitting it would let `executor: { resource: skills/run.md }`
carry an executable-code pointer through in typed clothes -- the door-C route
closed in 1.1.0. It costs nothing today, because the conformant carrier for
sources[].resource is a block sequence of block mappings, which stays refused.

Mapping leaves are scanned like every other frontmatter value (T1). Coverage
matrix 130/130 (new row: the off-allowlist key). No exported surface, detector
behaviour or calibration changed.
2026-08-21 21:04:51 +02:00

56 KiB
Raw Blame History

Honest limitations

Conceding these plainly is itself a control — it prevents the false assurance that a green scan means safe content. The README carries a summary of the highest-impact items; this is the full list, each with the mechanism.

  • Structural unsolvability at the text layer. Pattern/lexicon detection is bypassable in isolation; character-injection and novel phrasings evade it. The contract (tool-less transform, capability isolation, fail-secure) carries the security — the lexicon is defense-in-depth, not a wall.

  • A lone HIGH finding in trusted prose disposes to WARN, not quarantine. Under PRESET_TRUSTED_SOURCE, trust-scaling downgrades a single HIGH to WARN, and one HIGH is not "compound" (escalation needs ≥2 findings at MEDIUM+). So a HIGH injection reproduced verbatim under a trusted policy persists with a WARN. This is by design: if your "trusted" sources can carry attacker-influenced text, run them as untrusted (or add a quarantine floor).

  • The quarantine floor fires at MEDIUM+, and is a no-op under the shipped upload preset. Through 0.3.0 quarantine_default floored any finding to QUARANTINE_REVIEW. That premise ("a finding is the exception") broke when the active-content detector made every ordinary markdown link a finding, so 0.3.1 raised the floor to MEDIUM+. Under PRESET_USER_UPLOAD (untrusted) a MEDIUM already escalates on trust alone, so the floor still changes no outcome there; it is live only for a caller-defined trusted policy that opts into quarantine_default. Documented so the preset is not over-read: a LOW finding on an upload now disposes WARN.

  • Semantic / factual poisoning is invisible to lexicon + entropy: a false claim in clean prose carries no suspicious token. Highest impact for a wiki. The grounding module ships only a SourceGroundingCheck seam — the deterministic core does not judge semantics; a [judge] implementation must be plugged in.

  • Adversarial-ML evasion can survive normalization; tokenizer mismatch between scanner and model leaves gaps. Latent / dormant memory poisoning is not judgeable at write time.

  • Dormant / broken-link injection in a linked corpus (e.g. an OKF bundle): a link to a not-yet-existing target passes a per-concept write-time scan clean — the payload is planted later, when that target is written. link_graph surfaces the dangling edge as the signal, but catching the payload needs cross-write re-scan over time (the caller's disposition call).

  • OKF reserved files (index.md / log.md). In a received bundle these are legitimate structure, so mode-b import_bundle scans their body and frontmatter (an injection in a directory listing is caught) rather than path-rejecting the conformant bundle. A front-end materialising individual uploads keeps the opposite rule (allow_reserved=False): a reserved basename is a listing-shadow and refused.

  • OKF frontmatter is a restricted grammar: the mapping class has exactly one expressible form. Gate T2 accepts a line-oriented subset deliberately — full YAML is a larger parse-attack surface than a write-time gate needs. Flow sequences ([a, b]) and nested mappings are rejected outright, which fails secure. Three of the four routes to a mapping fail, each on a different rule — block (k:\n sub: v) on the nested-mapping check, dotted keys (k.sub: v) on the key pattern, and the inline second colon (k: sub: v) on the mapping-construct check. The fourth, the flow form, is admitted only when every key is on an allowlist (by, at, from, to, id, title, author, usage_count, last_modified — the keys SPEC.md @ 62432a09 §5.1/§5.2 names inside a mapping) and every leaf is a plain scalar, itself run through the same value predicates as a top-level scalar. Nested collections, quoted leaves, duplicate keys, an empty or unclosed mapping, and {a:b} (which PyYAML 6.0.3 reads as the key a:b, not as a scalar) all raise. The form is expressible, never trusted: the allowlist inspects every key, which is the property that carried the security when the blanket refusal was doing the enforcing. resource is deliberately off the allowlist although §5.1 names it inside a sources entry — it is a pointer rather than a label and the only key T3 exists for, so admitting it would let executor: { resource: skills/run.md } carry an executable-code pointer through a key the https allowlist never inspects. What else survives is scalars and flat lists of strings. Two routes used to degrade into a string instead of failing, and that defect is closed in 1.1.0: a block-sequence item carrying exactly one key (sources:\n - uri: https://e.com/a yielded the string 'uri: https://e.com/a') and the inline second colon (attester: resource: attesters/sql_equality.py, which a real YAML parser refuses outright). Both parsed "successfully" into the wrong type, and a pointer parked in one rode through in a key the resource allowlist never inspects — mode-b import_bundle returned WARN and wrote the merged concept verbatim. Both now FAIL_SECURE at T2, before the allowlist is reached. What closed is the type confusion, not pointer-smuggling as a class: T3 still inspects resource and nothing else, so an honest string under another key rides through exactly as before — attester: attesters/sql_equality.py is WARN, while the same path under resource: FAIL_SECUREs. That is by design (the string is scanned like any other frontmatter value under T1) and it is not what 1.1.0 changed. The boundary is where YAML puts it, ground-truthed against PyYAML 6.0.3: ": " and a trailing ":" open a mapping and are refused; a colon carrying neither a space nor a line end (domain:security, https://e.com:8443/a) does not and still parses, as does a quoted scalar (- "uri: x"). Quotes are retained rather than stripped — a divergence from YAML that remains, pinned in tests/test_okf.py.

  • T2 constrains import, not emission. The frontmatter grammar runs on okf.import_bundle (door C) only — parse_frontmatter is referenced nowhere in the door A/B persist path, so frontmatter that fails secure on import passes screen_output unremarked. The grammar therefore bounds what a consumer can receive, never what a producer can emit. Verified identical on 0.2.0 and 0.3.1.

  • An OKF v0.2 concept traverses the external-import path only if its sources are flat. The wall used to be total: both of v0.2's backward-breaking migration targets are mappings — timestampgenerated.at, and body # Citations → a sources block list of mappings — and a consumer measured 0 of 53 upstream concepts through the gate. The trust and provenance layer now passes in its spec form (generated, verified bare or listed, usage_window), so generated.at is no longer a wall. sources still is: SPEC.md writes each entry as a block mapping under a block sequence (- id: …\n resource: …), and that carrier stays refused — it is the shape whose one-key degradation smuggled a pointer before 1.1.0, and reopening it is a separate parse-safety decision, not a corollary of the flow form. A concept whose sources are flat strings, or absent, imports. The dangling-or-substituted executor/attester pointer question stays out of reach for the same reason: both are mappings whose payload key is resource.

  • A persist gate cannot cover execution risk. OKF v0.2 introduces concepts whose purpose is to name code to be run (runtime, executor.resource, attester.resource). This library answers "is this safe to store"; executable code carries its risk at run. A file that is harmless to persist can be harmful to point at. Upstream defers the attester ABI and sandboxing to a future revision, so there is no runtime contract to gate against — the execution boundary is unowned across the stack rather than covered by anyone's roadmap, and no tightening of a write-time scanner would change that.

  • A document that describes attacks is a false positive. Content documenting prompt-injection payloads (security notes, this project's own corpus) trips carrier-strip / fail-secure. At the text layer "about an attack" and "carrying an attack" are indistinguishable; such content needs a deliberate, explicitly escaped path, never a silent allow.

  • Bilingual text trips the Cyrillic/Latin homoglyph rule. homoglyph:cyrillic-latin-mix (MEDIUM) flags a Latin letter adjacent to a Cyrillic look-alike, so genuine bilingual prose → MEDIUM → under untrusted → QUARANTINE_REVIEW — a real false positive for an inbox that expects multilingual content. Conceded for 1.x: no fix is promised. The rule fires on codepoint adjacency, which genuine bilingual prose produces as readily as a substitution attack does. Narrowing it is a calibration question, not an API one, so a fix may land in any 1.x release without breaking the contract — but none is scheduled, and a caller that ingests multilingual prose should raise its untrusted-tier threshold rather than wait for one. SECURITY.md lists this as a documented boundary, not a vulnerability.

  • Insider in-place edits by a trusted author are out of the untrusted-content threat model.

  • Text-only. The core is text -> findings: it parses no files (no pypdf/python-docx/archive deps). Extract text first, then scan it with the high-untrust upload provenance. OCR-embedded instructions and multimodal stego are out of scope beyond the sanitizer's character-layer stripping.

  • Uploaded files: only the extracted text is scanned. The dev-scoped OKF inbox showcase (tests/test_okf_inbox_uploads.py + tests/inbox_frontend.py; parsers in the [dev] extra, never core dependencies) reads .txt/.md/.csv/.docx/.pptx/.xlsx, folders and .zip, materializes an OKF bundle, then guards it. What survives extraction is out of scope: macros, OLE/embedded objects, OCR-needing images, font/render stego, encrypted files — the binary layer needs a separate scanner. The front-end owns the container threats it can see (zip-slip → path gate, zip-bomb → size cap, symlink refusal, CSV/XLSX formula-lead cells). .pdf is a deliberate concession: a top-level .pdf is refused as unsupported rather than half-scanned (a PDF parser is disproportionate for a dev showcase, and the OCR/stego it would smuggle is already out of scope). One known gap: the numeric -/+ CSV false positive (a typed XLSX numeric cell does not trip it).

  • Lexicon findings are deduplicated by pattern idcount=1 and the first offset are reported, so a class matched across several channels collapses to one finding at its first location: a deliberate readability tradeoff.

  • Active-content severity grades on URL shape, so a pure beacon is only LOW. Since 0.3.1 a URL that merely names a remote document (bare path, no query, no userinfo, no percent-escapes, no opaque segment) is LOW, and only a URL that can move bytes outward keeps HIGH/MEDIUM. The deliberate hole: ![x](https:// evil.test/pixel.png) on an attacker-controlled host still fetches when a renderer touches it, leaking reader IP, user-agent and timing. Grading the fetch itself would re-block every ordinary document, which is precisely the 0.3.0 regression this replaced — so beaconing is conceded, not covered.

  • Short opaque URL segments slip through the same grading. Opacity is decided by entropy's primitives: base64 that decodes to text (≥20 chars), a hex id (≥32 chars), or Shannon entropy ≥4.4 at ≥24 chars. A shorter payload segment — https://evil.test/aGVsbG8gd29ybGQ — cannot be told from a name, because entropy is bounded by log2(length) at short lengths. Mitigation in depth, not in this detector: a literal credential in a URL is still caught by the LLM02 egress patterns in the same scan_output pass, whatever severity the carrier gets.

  • Percent-escapes count as data-carrying — a %20 in a path is a false positive. An ordinary link with an encoded space grades as carrying and reaches QUARANTINE_REVIEW / FAIL_SECURE on an untrusted upload. Obfuscated encoding is a core exfil primitive and the ambiguous case is put on the review side deliberately; it is listed here because it is the same class of over-block that 0.3.1 fixed, in a rarer shape. It is a non-ASCII-language tax, and that is the finding. Three consumer corpora measured it (2026-07-25/26). The two English ones found zero — 0 of 347 external URLs in a 527-document vendor-docs corpus, 0 of 81 in a capture store. The third, a 389-file Norwegian/Microsoft reference corpus, found 10 distinct real escapes and every one of them Norwegian: %C3%B8 and %C3%A5 are simply ø and å in UTF-8, and legal/government sources turn titles into paths (lovdata.no/…/kap2/%C2%A710). "Accepted false positive" reads differently as "URLs in your own language grade above LOW". Two distinct axes, which an earlier revision of this file conflated. For third-party link corpora — URLs an ingester collects from other people's sites — language does predict: the corpus of Norwegian legal/government sources carried escapes where two link corpora of mostly English-language domains carried none. For generated paths the predictor is not language but slugger class: a whitelist slugger ([^a-z0-9]+ → "-") cannot emit an escape in any language, because it discards the character before anything encodes it — measured structurally, not statistically, by a consumer whose content is Norwegian and whose slugger output is "Løkkene i produksjonslinja" → "l-kkene-i-produksjonslinja". A path built with encodeURIComponent produces escapes systematically the moment titles are non-ASCII. So non-ASCII language is a confounder for encode-vs-whitelist, and the slugger class is the testable thing at a consumer — a one-line code read, not a corpus census. Both consumer sluggers we have now read are whitelists, so their generated-path exposure is structurally zero rather than measured-zero; the second was read on 2026-07-31 and reduces the same Norwegian input to the same output. What this does not give us is the other half of the axis: we have never seen an encodeURIComponent-class slugger in the field, so "produces escapes systematically" above remains a prediction from the transform, not an observation. One sub-class worth separating: {tenant}/{agent-id} template placeholders in API code samples encode to %7B/%7D (26 of that corpus's 36 hits — an artifact of harvesting, not of prose). On the generated-path side those same placeholders cannot reach that shape at all: braces and / are outside the grammar, so "{tenant}/{agent-id} mal" reduces to tenant-agent-id-mal.

  • A percent-escaped path also defeats tokenization, which feeds the entropy branch. % is not in the separator class, so NSMs%20Grunnprinsipper%20for%20IKT is one 34-character token where the same title with literal spaces would be four short ones. Measured at H=4.04 — under the 4.4 floor, and moot in practice because the % rule already disqualifies the URL. It is recorded because it means the length floor does less work in non-ASCII paths than the calibration assumed. A field re-measurement (2026-07-31) sharpens this by roughly 3x. The highest-entropy legitimate token in a consumer's 2400-URL corpus is not the one above but a 78-character percent-escaped lovdata title, Fra%20%C3%A5ndsverk%20…, at H=4.301 — leaving 0.099 of headroom to the 4.4 floor, not 0.36. Re-run here through the real _URL_TOKEN_RE and shannon_entropy rather than a reconstruction: our tokenizer does emit it as one 78-character token. Still moot per URL (is_ordinary_url is False on the % rule) and still not a reason to move the threshold — but the margin this bullet reports is much thinner than 4.04 implies. Note the corpus count: 2400 is a rebuilt harvester's, where the 2401 in the next bullet is the earlier run's. Same knowledge base, one URL apart. An earlier version of this note offered a second moving variable — a corpus that "grew from 389 files to 394" — and that explanation is withdrawn: the consumer corrected it the same day, and the counts reproduce here against their tree. 394 is every .md under skills/; 389 is the references/** path the measurement actually scoped to; the 5 SKILL.md files between them contributed no unique URLs. That is one snapshot counted two ways, not two snapshots. So only the script moved, and the one-URL gap stays unexplained — the original is gone from a scratchpad and nobody has chased it. The counts remain non-interchangeable: never summed, never quoted as one figure.

  • Legitimate CDN content-asset ids trip the hex branch, permanently. A ≥32-char hex path segment is opaque by design, and several public CMSes mint asset URLs that way (measured: 9 distinct on regjeringen.no, ks.no, datatilsynet.no). This is the "legitimate build hash or doc id" case the calibration predicted, now confirmed present in the field. The class does not decay — it is how those systems generate URLs — so it is a standing false positive rather than a transient one. The branch is otherwise precise (no other false positives in 2401 distinct URLs) and stays.

  • A non-empty query is graded as data-carrying — the over-block that actually occurs in the field. Three corpora have now measured it, and each found a disjoint benign population: (1) 16 of 16 query-carrying external URLs in a vendor-docs corpus were publisher-authored campaign tracking (utm_* on the publisher's own domains); (2) 28 of 28 in a capture store were content identity (?v=, ?channel_id=, ?all=true) where the parameter is the resource; (3) 149 of 1694 distinct learn.microsoft.com URLs in a reference corpus carried ?view=, Microsoft Learn's own documentation-version selector, plus 23 ?api-version=. No parameter-level remedy covers any two of them, let alone all three: an allowlist keyed on tracking-parameter names resolves (1) entirely and (2) and (3) not at all; stripping the query is lossless for (1), dereferences nothing for (2), and silently changes which document is cited for (3) — the worst failure mode of the three, because the result stays plausible. This is a settled constraint on any future middle tier, not a hypothesis. The cost is bounded — a query-carrying link is MEDIUM, so it disposes QUARANTINE_REVIEW under PRESET_USER_UPLOAD and WARN under PRESET_TRUSTED_SOURCE: held or warned, never hard-failed (pinned in tests/test_wiring.py, because both consumers inferred a hard block rather than running it). An image keeps HIGH, and that is the carrier where this would bite — the two corpora that reported a carrier breakdown contained zero remote images and the third did not report one, so the image row of this limitation remains unmeasured in the field.

  • What active:raw-html still costs benign documentation is mostly ordinary HTML, not over-reach — and that residual is not a narrowing away. Two over-reach classes were closed in 0.6.0: the URL-attribute branch now requires an external target (an element outside the active name set carrying href="/en/agent-sdk/quickstart" reaches no attacker-controlled host), and <base> left the active name set (APIM's attribute-less <base /> means "run the inherited policy"; HTML's <base> has its whole affordance in an href the attribute branch still catches). <base /> appears in 25 of reference-corpus's 389 documents — count it with grep -rlE '<base\s*/?>', because a loose grep '<base' says 30 and is wrong: it also matches the literal <base64_string> placeholder, which is not a tag this detector ever fired on. Measured before and after in one session against one corpus state — docs/rawhtml-census.py, whose PRODUCTION row re-measures the shipped predicate rather than a hypothesis: reference-corpus 133 → 108 of 389, vendor-harvest 100 → 98 of 187, generated-notes 90 → 88 of 550. Against a ceiling of 26 / 38 / 41 (raw-HTML detection switched off entirely) that is 96% of the achievable reduction in reference-corpus and 5% in the two wiki corpora — the over-reach was nearly the whole raw-html cost in APIM policy XML and nearly none of it in vendor documentation. The classes had to be measured together. Alone, the URL-attribute branch frees 3 documents in reference-corpus and <base> frees 13; together they free 25. A document carrying one usually carries the other, so closing either alone leaves it blocked by its twin — one-at-a-time measurement understates both, and reading "frees 3" as "this over-reach is cheap" is how this document got the attribution wrong for three releases. What remains is real HTML in vendor documentation, flagged by the name branch, correctly: per wiki corpus <a> 298 occurrences, <frame> 94, <img> 63 — identical in both because 184 of their documents share ancestors. <frame> is the one arguable member: names are lower-cased and frame is in the active set for legacy HTML framesets, which appear in essentially no modern documentation, while Frame is a common MDX component name. Case is not an available discriminator: HTML is case-insensitive, so PascalCase cannot be treated as "component, not tag". Recovering the rest is not a further narrowing — it needs a carrier split and a new label (active:raw-html-link), because raw HTML has no ordinary form and the class collapses to one finding per document. The external-target test now gates one more surface, and it is a literal one. _has_external_target matches ^(?:[A-Za-z][A-Za-z0-9+.\-]*:|//), so a backslash-separated authority — href="\\evil.example/x" — reads as relative and now grades down, although WHATWG URL parsing normalizes backslashes to slashes for special schemes and a browser would resolve it externally. This is inherited, not introduced: the markdown paths have applied the same predicate since 0.3.1. It is recorded here rather than fixed because the fix belongs to the predicate, not to the raw-HTML branch that newly depends on it. The scanner and the mutator no longer share a predicate. Until 0.6.0 neutralize imported is_active_tag by name, so any narrowing moved the opt-in mutator too, and no test discriminated the two halves. They are now is_active_tag and is_defangable_tag; the mutator kept the broader behaviour deliberately, pinned by tests/test_neutralize.py::test_mutator_still_defangs_what_the_scanner_now_lets_pass.

  • The carrier split TIGHTENS the trusted tier when both carrier classes are present. 0.7.0 is sold as a loosening of the upload door, and on that door it is one. But splitting one class into two means a document carrying both an <img src> and an <a href> now emits two findings at MEDIUM+ where it emitted one, which trips the compound overlay (>=2 findings at MEDIUM+ -> escalated one tier). Such a document was WARN through 0.6.1 and is quarantine_review from 0.7.0 under PRESET_TRUSTED_SOURCE. On the trusted preset nothing was hard-failed to begin with, so this is the only direction the split can move it. Pinned by tests/test_wiring.py::test_split_tightens_the_trusted_tier_when_both_carriers_are_present, and docs/rawhtml-census.py now reports a TIGHTENS column against the previously-shipped row on both trust tiers — "frees N" without "tightens M" is a one-sided number. Measured on all three populations, as shipped: 0 documents tightened, on both trust tiers — reference-corpus (389), vendor-harvest (187), generated-notes (552). That zero is empirical, not structural, and the census shows exactly how thin it is. The split measured alone tightens 13 documents on the trusted tier in vendor-harvest and 14 in generated-notes. Adding the no-URL narrowing takes each of them back to 0: in these populations the document's second, HIGH-class carrier was itself a tag naming no target, which the narrowing makes inert, so the compound overlay never sees two findings. That is the census reporting a cancellation, not this repo proving one — a population whose second carrier is a real <img src> would still escalate, which is precisely the case test_split_tightens_the_trusted_tier_when_both_carriers_are_present constructs and pins. Read the zero as "not observed in any of the three populations, each counted against its own denominator", never as "cannot happen".

  • The split also LOOSENS the upload door for a lone anchor — the direction it was built for, and the one with a residual worth naming. Measured as shipped: <a href="https://ext.example/p">t</a> on its own emits one active:raw-html-link at MEDIUM and disposes quarantine_review under PRESET_USER_UPLOAD (warn under PRESET_TRUSTED_SOURCE); through 0.6.1 the same document graded HIGH and fail_secured. An <img src> to the same host is untouched — active:raw-html, HIGH, fail_secure. So an anchor pointing at an attacker-controlled host, arriving on an untrusted upload, is now a human decision rather than a halt. The trade is deliberate and it removes an asymmetry that came from syntax rather than affordance: following an anchor needs a click, exactly like the markdown inline link that has graded MEDIUM since 0.3.1, so the same URL no longer grades two different ways depending on which syntax carries it. It is recorded here so 0.7.0's "frees N documents" is not read as free — what was freed is the click-required class, and MEDIUM is a real grade drop on the door where every finding is trust-escalated.

  • Raw-HTML findings no longer count end tags, and that moved a published field. Through 0.6.1 </a> was active by name on its own, so count ran roughly 1.6× the opening-tag total (measured on one corpus) and a start/end pair counted 2. 0.7.0's no-URL narrowing makes an end tag inert — it names no target — so count is now the opening-tag total. A consumer reading count will see it drop for every document carrying </a>, on a field whose meaning did not change. The finding count is unaffected: the class still collapses to one finding per class per document, and count was never a document count.

  • Which tags the no-URL narrowing may render inert is a judgement about affordance, and no test in this repo can derive it. _URL_AFFORDANCE_TAGS holds the nine names whose entire active affordance is the URL they name — a, area, img, video, audio, source, track, frame, frameset — so carrying no URL attribute they name no target and grade inert. Every other name in the active set stays active with no attributes at all, because it does something a URL cannot describe: <script> executes its body, <style> restyles, <form> submits. That boundary is asserted, not measured. A name placed in the set whose affordance does not reduce to its URL would go silently invisible, and no corpus can catch it, because what it produces is an absence — the census counts findings, and a tag that stopped firing contributes nothing to count. The fail-secure choice one branch further in holds the other way and is worth reading beside it: a URL attribute whose value this module cannot resolve keeps the tag active (active_content.py:320-324), a branch the three corpora exercise 0 times. That zero is empirical, so the predicate is written not to depend on it.

  • "Clean" means graded, no finding raised — never cleaned bytes — and at one measured consumer's door, warn is the floor a document must clear to be persisted rather than rejected. The word is the library's own: a WARN disposition with nothing to report carries the reason string "clean: no findings" (disposition.py:264), and this project has repeated that word in the tables it sends consumers. screen_output is a judgement API — its DispositionResult carries assessment / disposition / max_severity / reasons, with no sanitized-text field to read off it. Defanging lives in a separate, deliberate call — neutralize — that a caller must invoke itself; nothing upstream of that call transforms a byte. Measured against llm-ingestion-okf's 0.7.0 pin (2026-08-13): inbox.py:139 sets its persist floor to warn, and inbox.py:323 persists anything carrying that disposition into the bundle; its adapter (guard_adapter.py:70) forwards the original extracted text, because nothing upstream ever handed it a transformed one. Four raw-HTML carrier forms the 0.7.0 no-URL narrowing grades inert — an <a aria-label> with no href, a bare </a>, <Frame>, <video /> — verified here (screen_output(..., PRESET_USER_UPLOAD)) to dispose warn, clean: no findings; at that consumer's door the same four land written into the bundle, carrier present verbatim. Neither library is wrong: screen_output never promised transformed bytes, and the consumer never called neutralize for them. The gap is in reading "clean" as "sanitized" rather than "no finding raised" — a reading this project's own reports invite, and one that will mislead any caller that persists on warn without calling neutralize itself.

  • Measured, document by document: a large minority of benign documents do not persist unattended at the upload door. The bullets above bound single rules on single URLs. This one bounds the thing a consumer actually feels — how often an ordinary document costs a human — and the honest answer is often, on corpora of technical documentation. Three benign populations, each reported against its own denominator (docs/fp-sweep.py, run on the post-0.4.0 tree carrying the axis separation, which renders every 0.4.0 disposition identically):

    population provenance n disposed non-WARN
    vendor-harvest vendor-published doc pages, harvested verbatim 185 98 (53.0%) — 64 fail-secure, 34 held
    generated-notes model-written notes at their own persist gate 547 88 (16.1%) — 61 fail-secure, 27 held
    reference-corpus first-party authored reference material 389 133 (34.2%) — 80 fail-secure, 53 held

    These three numbers predate the 0.6.0 raw-HTML narrowing and are left as published, because two of the three corpora are living — re-harvested by their owning repo, now 187 and 550 documents against the 185 and 547 measured here — so rewriting the cells would mix a code change with corpus drift. The narrowing's effect was measured separately, before and after against one corpus state: reference-corpus (static, and reproduced at exactly 389/133) drops to 108, the two wiki rows to 98 and 88 on their current state. The bullet above carries the method; docs/rawhtml-census.py LABEL=<path> reproduces it.

    The unit is a document and the gate is the strict one: screen_output(doc, PRESET_USER_UPLOAD), counting disposition is not WARN. WARN is the benign outcome (persisted, with a note), so a finding is not a false positive — only a document the pipeline cannot persist unattended is. Under the default action map that count is equivalent to assessed ELEVATED or worse, and the equivalence is pinned by tests/test_corpus.py::test_the_published_fp_metric_is_a_risk_statement so an action_map override cannot silently redefine the published number. These are not comparable to the URL-level measurements above (16 of 16, 28 of 28, 149 of 1694): different unit, different corpora, and they must never be combined or read as an update to each other. Nor are the three rows summable — different provenance, different denominators. What moved them is mostly residuals this document already concedes, counted by the labels at each document's worst severity (a histogram of every label present would credit the over-block to whatever else happened to be in the document). In vendor-harvest, active:raw-html is a top driver in 52 of the 98, markdown:link-anchor-injection in 23, and only about ten documents are moved by genuinely injection-shaped text, which is what security-adjacent documentation contains honestly. Generated-notes tracks it almost exactly, as 184 shared ancestors imply — active:raw-html in 53 of its 88, markdown:link-anchor-injection in 23 — so read those two rows as one observation, not two. In reference-corpus, which shares no upstream with either, the same shape holds with a different mix: active:markdown-link 38 (largely the ?view= documentation-version class from the query bullet above), active:data-uri 36, active:raw-html 27, markdown:link-anchor-injection 27, and eleven injection-shaped. Ground truth for "benign" is provenance, not inspection: nobody hand-read these corpora — each is benign by where it came from. A planted injection sitting in a harvested corpus is scored here as a false positive, which is a real caveat and not a formality. The populations are disjoint as documents but not independent as content: 184 of generated-notes' 547 are same-named derivatives of vendor-harvest's 185, which is most of why their driver labels agree. The third population shares no upstream with either, and is the one whose provenance is first-party. The trusted door cannot produce this number and is printed only as a footnote (162 of 185, 527 of 547, 366 of 389 WARN there): every non-CRITICAL finding WARNs under trust, which is the structural blindness that let the 0.3.0 active-content regression pass a green suite. Read the contrast as the intended one — the same corpus is cheap to persist from a source you trust and expensive from one you do not. Every population was swept twice and reproduced its counts exactly within the sweep that produced them, but two of the three are living corpora and no longer reproduce. vendor-harvest and generated-notes are directories inside claude-code-llm-wiki, which re-harvests per Claude Code release; re-run on 2026-08-11 against that repo at commit aba87e2 they give 100 of 187 and 90 of 550. The added documents are ordinary vendor documentation, not a detection change — the guard's behaviour did not move between the two runs. reference-corpus is static and reproduced 133 of 389 exactly. The table above is not restated to today's counts: it is a measurement with a date, and measurement provenance is never silently bumped. What was missing was the provenance itself — a reader who re-ran the first two rows got different numbers and had nothing in this file to explain why. Pin the corpus commit when you reproduce, or expect a drifting denominator. The largest document in any of them is 362 kB — no document approached the 1 000 000-character input cap, so truncation confounds nothing here. The sharpest datum needs no corpus at all: this repository's own eight published documents are 8 of 8 fail-secure at the upload door (python docs/fp-sweep.py self-docs=docs --ext=.md on a clone, which also picks up any untracked local notes). It is the first bullet of this file at full strength — a document that describes attacks carries the constructs it describes — and anyone can reproduce it. It is deliberately not a fourth population in the claim above: one eight-document corpus, chosen because it is the worst case, is an illustration and not a rate. The rate is stable across the versions consumers actually pin: the same population measured against the v0.3.1 tag gives 99 of 185 (53.5%) — one document more than today's 98. That one document is a removed false positive, not a regression: markdown:link-anchor-injection used to match 300 characters of ordinary prose by opening at one construct's [ and closing at another's ](…), and 0.3.3's ReDoS fix telescoped it shut. See the CHANGELOG correction — 0.3.3 reported two behaviour-change classes and there were three.

  • URL fragments are not graded. A fragment is never sent to the server, so it cannot carry data to the host a renderer auto-fetches, and …/overview#section is the most common shape in real documentation. The residual: a clicked link to an attacker-controlled page can have its location.hash read by that page's script, so a fragment payload on a link (not an image) is uncovered.

  • Secret egress: base64-wrapped is caught, hex-wrapped is not. The output gate decodes base64 blobs and re-scans the plaintext, so a base64-wrapped secret surfaces as decoded:egress:*. entropy exposes decoded plaintext for base64 only, so hex (and other encodings, or nested wraps) is a deliberate boundary — decode the transport layer first if you need it scanned.

  • Prose that merely mentions <script> fires hybrid-xss:script-tag. The pattern matches the opening tag and no longer requires </script>, so a document about XSS is flagged alongside a document that carries it. This is a deliberate trade made twice over: requiring the closing tag was a fail-open (an unclosed <script>alert(1) was silently missed by this label) and it removed a quadratic-backtracking site on the output path. It was not the last one — 0.3.2 said so and that claim was wrong. The 0.3.3 sweep of all 83 lexicon patterns found two more, and because scan_lexicon runs on the output path too, they were reachable through scan_output: "[" * 100_000 took 334.7s through the gate. The claim was too broad because the sweep behind it drove [ only through scan_active_content, never through the lexicon. 0.3.4 then found three more outside the lexicon — two of them on the input path, where no cap applies at all — so "the output path" was never the whole surface either. Measured, both claims are narrower than they read. The new label costs no consumer a disposition: any text containing a literal <script> already produced active:raw-html at HIGH on 0.3.1 — so the same prose disposed fail_secure under PRESET_USER_UPLOAD before this change and after it. The fail-open was equally confined to scan_lexicon called on its own; through either composed gate, active:raw-html already caught the unclosed tag. What changed is the label, not the outcome. And the outcome is not "a review": HIGH under a low-trust preset is fail_secure. Report-only means the text is never mutated — it does not mean the finding cannot block.

  • The two labels a <script> tag raises come from patterns that do not match the same strings. hybrid-xss:script-tag is <script\b[^><]*>; active:raw-html reads the same tag through HTML_TAG_RE, which consumes quoted attribute runs atomically and so tolerates both < and > inside a quoted value. The lexicon's < exclusion is not a modelling choice — it is the 0.3.3 ReDoS fix, and widening it back to [^>] restores a quadratic arm (the row below carries the numbers). Measured through both scanners: <script src="a<b"> raises active:raw-html and no XSS label, while <script data-t="a>b"> raises both, the lexicon's match simply ending at the quoted >. The disposition never moves — the raw-HTML branch grades <script> HIGH with no attributes at all, so every shape here still fail_secures under PRESET_USER_UPLOAD — so what the divergence costs is the label: a consumer filtering findings on the XSS id sees a subset of the script tags the gate actually caught, and must not read that id as the gate's script-tag census. The residual is practically dead in prose (a < inside a script tag's quoted attribute region is not an ordinary shape) and is recorded because the asymmetry is invisible from either scanner alone.

  • A connection-string password longer than 256 chars is not matched. The password run in the *-connstr egress patterns is bounded by MAX_CONNSTR_VALUE; unbounded, it sits in front of a mandatory @ and makes crafted input quadratic. Excluding the anchor character instead — the fix the active-content table uses — is unavailable here because that character is /, and a password containing / is the common case. What the residual actually costs, measured at the 257-char boundary: a generic long password still trips entropy:base64-blob at CRITICAL, so the disposition is unchanged. A JWT used as a DB password is the case that moves — the remaining detections (entropy:base64-blob HIGH, egress:jwt-token MEDIUM) top out below CRITICAL, so the any-tier CRITICAL block is lost: under PRESET_TRUSTED_SOURCE such a document drops from fail_secure to quarantine_review. Under PRESET_USER_UPLOAD it still fail_secures. The credential is never silently missed; on one preset it is held for review instead of halted.

  • The ReDoS sweep has a measured sensitivity floor, not a clean bill of health. All 152 compiled patterns across all eleven regex-bearing modules are swept arm by arm — payloads synthesised per run from each pattern's own skeleton, so [, [system] and [system]( are each probed separately rather than relying on generic units, and each pattern is timed in the call mode the production code uses (.sub()/.finditer() visit every start position where .match() cannot). Five patterns were quadratic across 0.3.3 and 0.3.4; all are fixed. But the sweep flags on timing, and it ignores measurements below a 1.5 ms noise floor at N=8000 — process CPU time, re-derived on that instrument (see the clock bullet below) rather than inherited from the wall clock the script used through 1.1.0. A quadratic arm sitting just under that floor would still cost up to ~23 s at the 1 000 000-char cap — that figure is arithmetic and not a measurement: a quadratic arm costs the square of the length ratio, and 1.5 ms × 125 × 125 is 23.4 s. So the claim this sweep supports is "no arm worse than ~23 s at the cap", not "no quadratic arm remains". The method's blind spot is real and has now been demonstrated twice: a generic-payload pass found only one of 0.3.3's two patterns, and 0.3.2's hand-written rows missed all three of 0.3.4's — including one on sanitize, the first thing every ingested document touches. Two arm shapes the unit- repetition payloads cannot express are pinned by hand as a result: a tag that closes around a long body, and a run of plain characters carrying no anchor at all.

  • A green ReDoS row is evidence only if it has been seen red, and three rows in this suite had never been. The class is not a bad bound but a payload that cannot reach the defect, and it leaves the row passing under the vulnerable form too. The sub-agent row is the clearest case: the seed's unbounded lazy run costs per prefix match, not per character — each start position where spawn an agent that matches drives its own O(N) scan to end-of-string looking for a capability keyword the payload never supplies, so K prefix matches cost K×O(N), and the bounded (?:\S+\s+){0,12}? port caps each scan at 12 tokens for K×O(1). A payload that matches the prefix once and then pads pays a single lazy run and is linear however long the pad is — two earlier shapes did exactly that, and the row sat measured-dead at 1.2× until the payload was rebuilt as "spawn an agent that " * 3000. (The nesting an older comment blamed is a red herring: the inner .*? sits in an optional group, never a repeated one.) Measured through scan_lexicon with the seed form patched back in — exponent 1.92 against the shipped 1.01, and 4.091s vs 0.190s at 12 000 words, the seed breaking the 2.0s bound outright. Two siblings were dead for different reasons. lexicon-script-tag had to be given its own N=200 000: at the shared N=100 000 the vulnerable [^>] form measured only ~1.21.4s — under the 2.0s assert, so the row was green under both forms and proved nothing. And test_gate_is_bounded_on_the_long_attribute_arm was killed by this repo's own narrowing: 0.7.0 put <a> in _URL_AFFORDANCE_TAGS, so its <a + 100k + > payload became inert and returned before the body ever reached the arm the row exists to guard — separation 1.0×, 0.028s and no findings, against 12.475s for the same payload carried by <script . The carrier was moved to <script , which is active by name with no attributes, so no future URL-shaped narrowing can hollow it out the same way. The general rule the three share: a payload must deny the literal the vulnerable run sits in front of — a unit that supplies it matches immediately and never exercises the run. Nothing but hand measurement finds this class. The row is green either way, so the suite cannot report its own blind spot, and every bound in it should be read as "verified red under the vulnerable form" only where a comment says it was.

  • Every ReDoS bound in the suite is measured on process CPU time, and so is the sweep that sets the published sensitivity floor. tests/redos_clock.py is the one clock all six test files import — time.process_time() — because a blowup is spent cycles while a loaded machine steals wall clock without adding any. On time.monotonic() two 0.7.0 rows failed at 2.24s / 3.66s against a 2.0s bound while two census processes held the CPU, and passed 3/3 on an idle machine: they had been descheduled, not slowed. It lives in one module rather than five copies because test_output.py::test_the_redos_clock_ignores_time_this_process_did_not_spend pins one implementation, and four unpinned copies would be free to drift back to a wall clock with nothing going red. What the CPU clock gives up, stated: a scan that BLOCKS forever burns no CPU, so it would hang the suite instead of failing it. That is acceptable only because every scanner it measures is pure regex over an in-memory string, with no I/O and no locks — the last wall-clock holdout was retired by auditing its path for anything that could block, not by assumption, and a wall clock guarding a mode that cannot occur still charges the false-red premium (measured there at 21.6s against a 10.0s bound, on a scan that spent 7.6s). The sweep now runs on the same clock, and closing that divergence bought no sensitivity. docs/redos-sweep.py imports scan_seconds instead of timing on time.monotonic(), so the 1.5 ms floor at N=8000 and the "~23 s at the cap" figure above are finally in the same currency as the bounds they justify. What the move did not do is quiet the sweep, and the floor came back unchanged. Measured over twelve full runs of all 2585 arms: the median ratio sits at 1.952.03 in every size bucket above 50 µs — the whole surface measures linear — while two-point excursions past the 2.6 flag threshold survive at every magnitude, p99 ratio 2.93.3 even above 1 ms. Flagged arms per run by floor: 6.9 at 0.5 ms, 1.1 at 1.0 ms, 0.33 at 1.5 ms (02 per run), so 1.5 ms is still the knee. Descheduling was never what made this sweep noisy — a ratio computed from two points is. Four distinct arms flagged at the shipped floor across those twelve runs, each in exactly one of them, and nine of the twelve runs were clean; eight consecutive runs of the shipped script immediately after a full test run flagged 03 arms each, so machine load still moves the count even on a CPU clock. Six flagged arms re-measured over six doublings give exponent 0.971.09 and at most 1.2 s at the 1 000 000-char cap. A single clean run of this sweep is therefore not evidence either — and neither is a single flagged one.

  • Every surface now bounds its input, but not all of them the same way. sanitize, fence and neutralize raise OversizeInputError above MAX_INPUT_CHARS (1 000 000) rather than returning a partially transformed document, which bounds the whole input path — sanitize is step 1 of prepare_input, and it only ever removes, so everything after it is already under the cap. They reject rather than truncate because they return content: a shortened document is silent data loss, and a transformed prefix followed by an untransformed tail is a bypass an attacker positions the payload into. The detection surfaces truncate instead, which costs only detection in the tail — scan_lexicon / scan_output always have, and as of 0.4.0 scan_active_content called directly does too (reached through scan_output it inherits that surface's cap and is not flagged twice), as does okf.link_graph, whose cost was a bundle-wide findall over every document body. What truncation costs is worth naming: past the cap, "no finding" means "not looked at". Each says so rather than staying silent — the scanners emit an oversize-input finding (active:oversize-input, OWASP LLM10), and link_graph records (from_id, body_length) in LinkGraphResult.truncated, which is what lets a caller tell "no links past here" apart from "no links read past here". The split itself is conceded for 1.x, not deferred. It is a surface property, not calibration: making the two halves agree later means either raising where a caller gets a value today, or returning a truncated value where one raises — a change to an exported symbol's contract in either direction, therefore 2.0.0. 1.x keeps the rule as stated: a surface that returns content rejects at the cap, a surface that returns findings truncates and says so.

The six documented gaps (tracked by the coverage matrix)

These are asserted to still hold by tests/test_coverage_matrix.py — a closed gap fails the test, forcing this doc to be updated:

  1. Hex-wrapped secret egressentropy decodes base64 only (above).
  2. Semantic / factual poisoning — invisible to token analysis (above).
  3. A lone HIGH in trusted prose → WARN — the §4.7 trust-scaling design (above).
  4. Lexicon dedup (count=1) — first offset only, by design (above).
  5. Pure beaconing — a bare-path remote image on a hostile host is LOW (above).
  6. Short opaque URL segment (<24 chars) — below what entropy can resolve (above).
  • Severity still carries disposition intent on the detection side. The 0.5.0 axis separation split the assessment (Risk) from the action (Disposition), but only on the caller side of the boundary. Inside the detectors, a finding's Severity is still calibrated partly for the disposition it will produce rather than purely for what was observed, and two places in the tree say so outright: calibration.py's ACTIVE_CONTENT_ORDINARY_SEVERITY = LOW exists because grading an ordinary external image HIGH fail-secured ordinary uploads (measured on v0.3.0), and disposition.py's quarantine floor was raised from any finding to MEDIUM+ to repair the same regression from the other end. Both fixes are correct for the dispositions they produce; the cost is that an ordinary external image is recorded as low-severity rather than as a real outward-fetch capability that is not evidence of an attack — so no policy, however strict, can act on that capability, because the detector already decided it did not matter. Closing this means giving detectors a channel that says what was seen separately from how bad it is, which changes the grading and therefore fires the consumer-notification promise in docs/PLAN-v1.md. Conceded for the whole of 1.x, not deferred. That channel changes Finding and Severity, which is a 2.0.0 change under the version contract, so 1.x ships with the coupling intact by decision rather than by omission. A caller that needs the capability separately from the grade must read the finding idactive:markdown-image names the outward fetch whatever severity it carries — and must not infer "nothing was seen" from a low Severity.

  • Under the default action map the assessment axis carries exactly one judgement the disposition does not. DEFAULT_ACTION_MAP sends NONE and LOW to WARN, ELEVATED to QUARANTINE_REVIEW and SEVERE to FAIL_SECURE — the last two 1:1. So for any document that carries a finding at all, assessment is a relabelling of disposition and nothing more; the only thing it adds is clean versus findings present, none dispositive in this context, which 0.4.0 rendered identically. That collapse is the point (the map is what keeps the separation additive, so a caller ignoring the new axis sees no change), and it is also the limitation: reading assessment buys a consumer nothing until it supplies its own action_map or needs the clean/low distinction. The second consequence is on this document. The published false-positive rates are counts of documents disposed non-WARN, and they are a statement about assessed risk only while NONE + LOW are exactly the WARN pre-image. tests/test_corpus.py::test_the_published_fp_metric_is_a_risk_statement pins that equivalence — but it pins it for DEFAULT_ACTION_MAP. A caller running its own map makes "disposed non-WARN" a different claim from the one measured here, with nothing in either repo failing to say so.

  • A ZWJ hidden between two emoji is exempt, and ZWNJ's own false-positive class is untouched. U+200D composes emoji (👩‍💻 is WOMAN + ZWJ + PERSONAL COMPUTER), so testing it on codepoint membership alone flagged and stripped the joiner in any document containing a ZWJ-composed emoji — an any-tier FAIL_SECURE carrier, plus silent decomposition of the emoji, on ordinary first-party content. The joiner is now judged by context instead: exempt only when both neighbours are emoji-context codepoints, measured to cover 122/122 of the codepoints adjacent to a ZWJ across Unicode 17.0's 1614 RGI sequences. Two residuals follow. First, a joiner placed between two emoji is now invisible to the carrier check and could carry a covert channel — it costs an emoji per bit and cannot split a word, which is the shape the word-splitting attack needs, so the narrowing is deliberate rather than complete. Second, U+200C (ZWNJ) still has no context test, and it is orthographically required in Persian, Arabic and Devanagari — those documents remain hard-blocked. That is a separate criterion (script-based, not pictographic) and no corpus is available here to verify it against, so it is parked as a known false-positive class rather than guessed at.

  • That context test is one predicate on two surfaces, and the symbol carrying it is private. sanitize owns _is_joiner_in_emoji_sequence; output imports it (output.py:74) instead of restating it, because the same defect had to be fixed on both surfaces and a split would let the input side stop flagging while the output side kept hard-blocking — or the reverse, which is how a carrier reaches a persisted artifact after passing the input gate. The agreement is pinned by tests/test_output.py::test_output_zwj_narrowing_matches_the_sanitize_side, which asserts stripped == flagged across six shapes — half-context on either side, a leading and a trailing joiner, one genuine in-sequence joiner, and a word split. Two things that pin does not give. The six shapes are hand-written rather than drawn from a corpus, so everywhere outside them the surfaces agree by shared implementation, not by test — which is the stronger guarantee only for as long as the import survives. And the leading underscore means the predicate is not part of the surface frozen under semver: a consumer that imports it is pinning a private name 1.x makes no promise about.

Out-of-scope (documented boundary)

Embedding/vector-layer defenses (OWASP LLM08, downstream of persist); multimodal steganography; query-time / runtime guardrails; semantic factuality verification.