Three consumers reconstructed is_ordinary_url from prose we sent in coordination messages and each produced a different wrong number on a real corpus: one omitted the base64 20-char floor and fired on path words like /blog/; one omitted the opaque-token condition entirely and undercounted; one computed Shannon entropy over whole filenames instead of tokens and concluded the 4.4 floor over-blocks ordinary documents. Same cause each time -- our prose described the rules without their tokenizer. docs/URL-SHAPE.md states the algorithm in order, spells out the separator class and all three length floors, and lists the three reconstruction errors as worked counter-examples. Its example table is parsed and asserted against the real predicate by tests/test_url_shape_doc.py, so the reference cannot drift from the code -- all 18 rows verified load-bearing. LIMITATIONS.md brought current with the field measurements: - Percent-escape is no longer zero. Two English corpora measured 0; a 389-file Norwegian/Microsoft corpus found 10, all Norwegian (%C3%B8, %C3%A5 are just o-slash and a-ring). It is a non-ASCII-language tax, and both zero-measuring corpora being English was a sampling bias invisible from inside. - The query over-block now has THREE disjoint benign populations: utm_* tracking, content identity (?v=, ?channel_id=), and Microsoft Learn's ?view= version selector. No parameter-level remedy covers any two, which moves this from a conclusion to a settled constraint on 0.4.0. - Legitimate CDN asset ids trip the hex branch permanently; the branch is otherwise precise (no other FP in 2401 distinct URLs) and stays. - Raw HTML with a relative URL attribute is HIGH though it reaches no external host, and end tags are counted. - OKF frontmatter: a one-key block-sequence item is silently misparsed to a string where two keys hard-reject, so a pointer can ride past the resource allowlist. Consequence: a conformant OKF v0.2 concept cannot traverse door C at all, since both backward-breaking migration targets are nested. Fail-secure, but a compatibility wall that needs a deliberate parse-safety decision. - A persist gate cannot cover execution risk, and that boundary is unowned. New behaviour claims are pinned by tests so a closed concession fails and forces this doc to be updated. 593 -> 631 passed.
16 KiB
16 KiB
Honest limitations
Conceding these plainly is itself a control — it prevents the false assurance that a green scan means safe content. The README carries a summary of the highest-impact items; this is the full list, each with the mechanism.
- Structural unsolvability at the text layer. Pattern/lexicon detection is bypassable in isolation; character-injection and novel phrasings evade it. The contract (tool-less transform, capability isolation, fail-secure) carries the security — the lexicon is defense-in-depth, not a wall.
- A lone HIGH finding in trusted prose disposes to WARN, not quarantine.
Under
PRESET_TRUSTED_SOURCE, trust-scaling downgrades a single HIGH to WARN, and one HIGH is not "compound" (escalation needs ≥2 findings at MEDIUM+). So a HIGH injection reproduced verbatim under a trusted policy persists with a WARN. This is by design: if your "trusted" sources can carry attacker-influenced text, run them as untrusted (or add a quarantine floor). - The quarantine floor fires at MEDIUM+, and is a no-op under the shipped upload
preset. Through 0.3.0
quarantine_defaultfloored any finding to QUARANTINE_REVIEW. That premise ("a finding is the exception") broke when the active-content detector made every ordinary markdown link a finding, so 0.3.1 raised the floor to MEDIUM+. UnderPRESET_USER_UPLOAD(untrusted) a MEDIUM already escalates on trust alone, so the floor still changes no outcome there; it is live only for a caller-defined trusted policy that opts intoquarantine_default. Documented so the preset is not over-read: a LOW finding on an upload now disposes WARN. - Semantic / factual poisoning is invisible to lexicon + entropy: a false claim
in clean prose carries no suspicious token. Highest impact for a wiki. The
groundingmodule ships only aSourceGroundingCheckseam — the deterministic core does not judge semantics; a[judge]implementation must be plugged in. - Adversarial-ML evasion can survive normalization; tokenizer mismatch between scanner and model leaves gaps. Latent / dormant memory poisoning is not judgeable at write time.
- Dormant / broken-link injection in a linked corpus (e.g. an OKF bundle): a
link to a not-yet-existing target passes a per-concept write-time scan clean — the
payload is planted later, when that target is written.
link_graphsurfaces the dangling edge as the signal, but catching the payload needs cross-write re-scan over time (the caller's disposition call). - OKF reserved files (
index.md/log.md). In a received bundle these are legitimate structure, so mode-bimport_bundlescans their body and frontmatter (an injection in a directory listing is caught) rather than path-rejecting the conformant bundle. A front-end materialising individual uploads keeps the opposite rule (allow_reserved=False): a reserved basename is a listing-shadow and refused. - OKF frontmatter is a restricted grammar, and a one-key block-sequence item is
silently misparsed. Gate T2 accepts a line-oriented subset deliberately — full
YAML is a larger parse-attack surface than a write-time gate needs. Nested mappings
and flow collections (
[a, b],{k: v}) are rejected outright, which fails secure. The defect is between those two outcomes: a block sequence whose items carry exactly one key parses "successfully" into the wrong type —sources:\n - uri: https://e.com/ayields the string'uri: https://e.com/a', not a mapping, while the same list with two keys per item hard-rejects. A pointer can therefore ride through in a key theresourceallowlist never inspects (attester:\n - resource: attesters/sql_equality.py→ WARN), whereas a top-levelresource:with a relative path correctly fails secure. The shape is not conformant OKF, so a well-formed bundle will not produce it; a malformed or hostile one can, and mode-bimport_bundlewrites the merged concept verbatim. - Consequence: an OKF v0.2 concept cannot traverse the external-import path. Both
of v0.2's backward-breaking migration targets are nested —
timestamp→generated.at, and body# Citations→ asourcesblock list of mappings — so a conformant v0.2 concept fails secure at the frontmatter gate. This is the correct direction but it is a compatibility wall, not a policy: v0.2 support requires a deliberate parse-safety decision about widening the grammar, and the dangling-or-substitutedexecutor/attesterpointer question only becomes live once that decision is made. - A persist gate cannot cover execution risk. OKF v0.2 introduces concepts whose
purpose is to name code to be run (
runtime,executor.resource,attester.resource). This library answers "is this safe to store"; executable code carries its risk at run. A file that is harmless to persist can be harmful to point at. Upstream defers the attester ABI and sandboxing to a future revision, so there is no runtime contract to gate against — the execution boundary is unowned across the stack rather than covered by anyone's roadmap, and no tightening of a write-time scanner would change that. - A document that describes attacks is a false positive. Content documenting prompt-injection payloads (security notes, this project's own corpus) trips carrier-strip / fail-secure. At the text layer "about an attack" and "carrying an attack" are indistinguishable; such content needs a deliberate, explicitly escaped path, never a silent allow.
- Bilingual text trips the Cyrillic/Latin homoglyph rule.
homoglyph:cyrillic-latin-mix(MEDIUM) flags a Latin letter adjacent to a Cyrillic look-alike, so genuine bilingual prose → MEDIUM → under untrusted → QUARANTINE_REVIEW — a real false positive for an inbox that expects multilingual content. A calibration fix is pending. - Insider in-place edits by a trusted author are out of the untrusted-content threat model.
- Text-only. The core is
text -> findings: it parses no files (nopypdf/python-docx/archive deps). Extract text first, then scan it with the high-untrust upload provenance. OCR-embedded instructions and multimodal stego are out of scope beyond the sanitizer's character-layer stripping. - Uploaded files: only the extracted text is scanned. The dev-scoped OKF
inbox showcase (
tests/test_okf_inbox_uploads.py+tests/inbox_frontend.py; parsers in the[dev]extra, never coredependencies) reads.txt/.md/.csv/.docx/.pptx/.xlsx, folders and.zip, materializes an OKF bundle, then guards it. What survives extraction is out of scope: macros, OLE/embedded objects, OCR-needing images, font/render stego, encrypted files — the binary layer needs a separate scanner. The front-end owns the container threats it can see (zip-slip → path gate, zip-bomb → size cap, symlink refusal, CSV/XLSX formula-lead cells)..pdfis a deliberate concession: a top-level.pdfis refused as unsupported rather than half-scanned (a PDF parser is disproportionate for a dev showcase, and the OCR/stego it would smuggle is already out of scope). One known gap: the numeric-/+CSV false positive (a typed XLSX numeric cell does not trip it). - Lexicon findings are deduplicated by pattern id —
count=1and the first offset are reported, so a class matched across several channels collapses to one finding at its first location: a deliberate readability tradeoff. - Active-content severity grades on URL shape, so a pure beacon is only LOW.
Since 0.3.1 a URL that merely names a remote document (bare path, no query, no
userinfo, no percent-escapes, no opaque segment) is LOW, and only a URL that can
move bytes outward keeps HIGH/MEDIUM. The deliberate hole:
on an attacker-controlled host still fetches when a renderer touches it, leaking reader IP, user-agent and timing. Grading the fetch itself would re-block every ordinary document, which is precisely the 0.3.0 regression this replaced — so beaconing is conceded, not covered. - Short opaque URL segments slip through the same grading. Opacity is decided by
entropy's primitives: base64 that decodes to text (≥20 chars), a hex id (≥32 chars), or Shannon entropy ≥4.4 at ≥24 chars. A shorter payload segment —https://evil.test/aGVsbG8gd29ybGQ— cannot be told from a name, because entropy is bounded bylog2(length)at short lengths. Mitigation in depth, not in this detector: a literal credential in a URL is still caught by the LLM02 egress patterns in the samescan_outputpass, whatever severity the carrier gets. - Percent-escapes count as data-carrying — a
%20in a path is a false positive. An ordinary link with an encoded space grades as carrying and reaches QUARANTINE_REVIEW / FAIL_SECURE on an untrusted upload. Obfuscated encoding is a core exfil primitive and the ambiguous case is put on the review side deliberately; it is listed here because it is the same class of over-block that 0.3.1 fixed, in a rarer shape. It is a non-ASCII-language tax, and that is the finding. Three consumer corpora measured it (2026-07-25/26). The two English ones found zero — 0 of 347 external URLs in a 527-document vendor-docs corpus, 0 of 81 in a capture store. The third, a 389-file Norwegian/Microsoft reference corpus, found 10 distinct real escapes and every one of them Norwegian:%C3%B8and%C3%A5are simply ø and å in UTF-8, and legal/government sources turn titles into paths (lovdata.no/…/kap2/%C2%A710). "Accepted false positive" reads differently as "URLs in your own language grade above LOW". Our two zero-measuring corpora were both English; that was a sampling bias invisible from inside. Two sub-classes worth separating:{tenant}/{agent-id}template placeholders in API code samples encode to%7B/%7D(26 of that corpus's 36 hits — an artifact of harvesting, not of prose), and a slugger deriving filenames from titles would produce escapes systematically rather than incidentally — still unmeasured. - A percent-escaped path also defeats tokenization, which feeds the entropy
branch.
%is not in the separator class, soNSMs%20Grunnprinsipper%20for%20IKTis one 34-character token where the same title with literal spaces would be four short ones. Measured at H=4.04 — under the 4.4 floor, and moot in practice because the%rule already disqualifies the URL. It is recorded because it means the length floor does less work in non-ASCII paths than the calibration assumed. - Legitimate CDN content-asset ids trip the hex branch, permanently. A ≥32-char
hex path segment is opaque by design, and several public CMSes mint asset URLs that
way (measured: 9 distinct on
regjeringen.no,ks.no,datatilsynet.no). This is the "legitimate build hash or doc id" case the calibration predicted, now confirmed present in the field. The class does not decay — it is how those systems generate URLs — so it is a standing false positive rather than a transient one. The branch is otherwise precise (no other false positives in 2401 distinct URLs) and stays. - A non-empty query is graded as data-carrying — the over-block that actually
occurs in the field. Three corpora have now measured it, and each found a
disjoint benign population:
(1) 16 of 16 query-carrying external URLs in a vendor-docs corpus were
publisher-authored campaign tracking (
utm_*on the publisher's own domains); (2) 28 of 28 in a capture store were content identity (?v=,?channel_id=,?all=true) where the parameter is the resource; (3) 149 of 1694 distinctlearn.microsoft.comURLs in a reference corpus carried?view=, Microsoft Learn's own documentation-version selector, plus 23?api-version=. No parameter-level remedy covers any two of them, let alone all three: an allowlist keyed on tracking-parameter names resolves (1) entirely and (2) and (3) not at all; stripping the query is lossless for (1), dereferences nothing for (2), and silently changes which document is cited for (3) — the worst failure mode of the three, because the result stays plausible. This is a settled constraint on any future middle tier, not a hypothesis. The cost is bounded — a query-carrying link is MEDIUM, so it disposes QUARANTINE_REVIEW underPRESET_USER_UPLOADand WARN underPRESET_TRUSTED_SOURCE: held or warned, never hard-failed (pinned intests/test_wiring.py, because both consumers inferred a hard block rather than running it). An image keeps HIGH, and that is the carrier where this would bite — the two corpora that reported a carrier breakdown contained zero remote images and the third did not report one, so the image row of this limitation remains unmeasured in the field. - Raw HTML with a relative URL attribute is HIGH, though it can reach no
attacker-controlled host. The markdown paths test for an external target before
flagging; the raw-HTML path deliberately does not, because an active element needs
no URL at all (an
on*=handler executes on its own). That reasoning covers event handlers but over-reaches on the URL-attribute branch: an element outside the active name set carryinghref="/en/agent-sdk/quickstart"— an internal doc route — grades HIGH. Measured on a vendor-docs corpus, where it lands on MDX components:<Card href="/…">fires this way, and<Frame>fires on the name branch alone because names are lower-cased andframeis in the active set — legacy HTML framesets, which appear in essentially no modern documentation, whileFrameis a common component name. Case is not an available discriminator: HTML is case-insensitive, so PascalCase cannot be treated as "component, not tag". - Raw-HTML findings count end tags.
</a>is active by name on its own, so a corpus census that counts only opening tags understates what this detector reports by roughly the ratio of closing to opening active tags (measured at 1.6× on one corpus). Severity and finding count are unaffected — the class collapses to one finding — but thecountfield is not a document count. - URL fragments are not graded. A fragment is never sent to the server, so it
cannot carry data to the host a renderer auto-fetches, and
…/overview#sectionis the most common shape in real documentation. The residual: a clicked link to an attacker-controlled page can have itslocation.hashread by that page's script, so a fragment payload on a link (not an image) is uncovered. - Secret egress: base64-wrapped is caught, hex-wrapped is not. The output gate
decodes base64 blobs and re-scans the plaintext, so a base64-wrapped secret
surfaces as
decoded:egress:*.entropyexposes decoded plaintext for base64 only, so hex (and other encodings, or nested wraps) is a deliberate boundary — decode the transport layer first if you need it scanned.
The six documented gaps (tracked by the coverage matrix)
These are asserted to still hold by tests/test_coverage_matrix.py — a closed gap
fails the test, forcing this doc to be updated:
- Hex-wrapped secret egress —
entropydecodes base64 only (above). - Semantic / factual poisoning — invisible to token analysis (above).
- A lone HIGH in trusted prose → WARN — the §4.7 trust-scaling design (above).
- Lexicon dedup (
count=1) — first offset only, by design (above). - Pure beaconing — a bare-path remote image on a hostile host is LOW (above).
- Short opaque URL segment (<24 chars) — below what entropy can resolve (above).
Out-of-scope (documented boundary)
Embedding/vector-layer defenses (OWASP LLM08, downstream of persist); multimodal steganography; query-time / runtime guardrails; semantic factuality verification.