test(docs-scripts): the census re-derived the URL reader it measures, and nothing would have caught the drift
`docs/fp-sweep.py` and `docs/rawhtml-census.py` produce the numbers published in
`docs/LIMITATIONS.md`, and both reach past the public API into private module
state -- `active_content._ACTIVE_TAGS`, `_EVENT_ATTR_RE`, `_URL_ATTR_RE`,
`_has_external_target`, `calibration.RISK_RANK`. A rename inside `src/` broke
them while the suite stayed green, and the breakage would have surfaced months
later, at the moment someone tried to re-measure a published claim. This was the
last uncovered contract in the repo.
The coverage found a live one. `rawhtml-census.has_external_url_attr` parsed
attributes with its own pattern instead of calling the shipped reader, and the
copy had drifted on two shapes:
- a URL attribute reached through a prefix. `_URL_ATTR_RE` matches `src=`
inside `data-src=` on a word boundary, so the shipped predicate reads the
value and blocks; the census's own name table saw `data-src` and skipped it.
- a multi-candidate `srcset`. The shipped reader splits on `[,\s]+` so a
relative first candidate cannot mask an external one behind it; the census
tested the whole attribute value as a single URL.
Both made the `A` candidate rows free documents the shipped predicate keeps --
under-counting against the `PRODUCTION` row printed directly beside them, which
that row exists to expose. The census now delegates to
`active_content._url_attr_is_external`, so there is one reader, not two. This
repo already carries the general form of that lesson in `docs/URL-SHAPE.md`:
three consumers reconstructed a predicate from prose and each got a different
wrong answer.
NO PUBLISHED NUMBER MOVED. Re-measured against all three live populations after
the fix, in one session each: reference-corpus 389 docs (A frees 3, base-url 13,
both 25 -- the numbers in `docs/LIMITATIONS.md`, unchanged), vendor-harvest 187,
generated-notes 550. `PRODUCTION` equals `A + base-url` in all three (108/108,
98/98, 88/88), and vendor-harvest exercises the corrected branch for real (8
external `Card` attributes). The defect was latent, not published.
The tests are scoped to what a test here can honestly hold. The corpora live
outside this repo in private consumer repos, so neither script can be run end to
end from the suite and a stand-in corpus would only pin a fiction. What is
pinned: every imported name still exists with the shape used; `fp-sweep`'s
metric guard fails when the action map is re-mapped (a guard that cannot fail
protects nothing); `measure` still reads `.disposition` and `.assessment` and
excludes empty files from the denominator; the census's in-process patch point
still moves the gate, so a census patching a dead symbol cannot print six
identical rows and read as a finding; `PRODUCTION` equals `A + base-url` across
every branch of the predicate; and both scripts still refuse an argument-less
run rather than measuring nothing.
759 passing (was 736). Coverage matrix unchanged at 128/128 with 6/6 documented
gaps holding. The `736` in `docs/ADOPTION-BRIEF.md` is scoped "as of v0.6.1" and
is correct for that tag; it moves to 759 at the next version bump.
This commit is contained in:
parent
0903785187
commit
0df7e87c2f
2 changed files with 242 additions and 22 deletions
|
|
@ -40,7 +40,6 @@ consumer repos and their paths must not reach a public mirror:
|
|||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import sys
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
|
@ -54,28 +53,19 @@ from llm_ingestion_guard import active_content as ac # noqa: E402
|
|||
|
||||
BENIGN = Disposition.WARN
|
||||
|
||||
# Attribute parser — only needed to read a URL attribute's VALUE, which
|
||||
# `_URL_ATTR_RE` (a presence test) deliberately does not capture.
|
||||
_ATTR_KV_RE = re.compile(
|
||||
r"""\b(?P<k>[A-Za-z_:][\w:.\-]*)\s*=\s*(?P<v>"[^"]*"|'[^']*'|[^\s>]+)"""
|
||||
)
|
||||
_URL_ATTR_NAMES = frozenset({
|
||||
"src", "href", "xlink:href", "srcset", "data", "poster", "formaction",
|
||||
"action", "background", "cite", "codebase", "longdesc",
|
||||
})
|
||||
|
||||
|
||||
def has_external_url_attr(attrs: str) -> bool:
|
||||
"""True if any URL-bearing attribute points at an attacker-reachable target."""
|
||||
for m in _ATTR_KV_RE.finditer(attrs):
|
||||
if m.group("k").lower() not in _URL_ATTR_NAMES:
|
||||
continue
|
||||
value = m.group("v")
|
||||
if value[:1] in "\"'":
|
||||
value = value[1:-1]
|
||||
if ac._has_external_target(value.strip()):
|
||||
return True
|
||||
return False
|
||||
"""True if any URL-bearing attribute points at an attacker-reachable target.
|
||||
|
||||
Delegates to the shipped reader instead of re-deriving it. This function used
|
||||
to parse attributes itself, and the copy drifted: it read attribute names with
|
||||
its own pattern (so `data-src="//evil"` was invisible to it while
|
||||
`_URL_ATTR_RE` matched it) and treated a value as one URL (so an external
|
||||
candidate later in a multi-candidate `srcset` was missed). Both shapes made
|
||||
the `A` rows under-count against the PRODUCTION row printed beside them —
|
||||
exactly the drift the PRODUCTION row exists to expose. Pinned by
|
||||
`tests/test_docs_measurement_scripts.py`.
|
||||
"""
|
||||
return ac._url_attr_is_external(attrs)
|
||||
|
||||
|
||||
def _variant(*, drop: frozenset[str] = frozenset(), external_only: bool = False):
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue