1
0
Fork 0
llm-ingestion-pipeline-secu.../docs/fp-sweep.py
Kjell Tore Guttormsen d1bff6047d measure(fp): the upload door costs a human on 16-53% of benign documents
Every field measurement this repo had published was per URL. None answered
what a consumer actually feels: how often does an ordinary document fail to
persist unattended? Three benign populations, each against its own
denominator, run through screen_output under PRESET_USER_UPLOAD and counted
at document granularity:

  vendor-harvest    98 of 185 (53.0%) non-WARN -- 64 fail-secure, 34 held
  generated-notes   88 of 547 (16.1%) non-WARN -- 61 fail-secure, 27 held
  reference-corpus 133 of 389 (34.2%) non-WARN -- 80 fail-secure, 53 held

The number is bad and ships as measured; PLAN-v1 committed to that in advance
("et roedt FP-resultat er like verdifullt"). The response is a documented
limitation, not a recalibration: moving the grading fires the locked
linkedin-studio notification promise, and the drivers are residuals
LIMITATIONS already concedes. Counted at each document's worst severity,
active:raw-html -- the MDX-component over-reach -- is a top driver in 52 of
vendor-harvest's 98 and 53 of generated-notes' 88; about ten per population
are genuinely injection-shaped text, which security-adjacent documentation
honestly contains.

Method traps closed rather than stepped in:

- The unit is in the number. Document-level rates are NOT comparable to the
  URL-level 16/16, 28/28, 149/1694 above them, and the three rows are not
  summable -- the 2400 != 2401 defect class one level up.
- The gate is the strict one. The trusted door WARNs every non-CRITICAL
  finding, so it would have handed back a beautiful, meaningless near-zero;
  it is printed as a footnote and labelled structurally blind.
- "not WARN" is only a risk statement while the default action map sends
  exactly NONE and LOW to WARN. action_map became a supported override last
  commit, so the equivalence is pinned in the suite and the sweep aborts if
  it breaks.
- Ground truth for benign is provenance, not inspection, and says so.
- The populations are disjoint as documents but not independent as content:
  184 of generated-notes' 547 are same-named derivatives of vendor-harvest.
  Measured, not assumed, and the two rows read as one observation.

Also fixed: tests/test_wiring.py credited a consumer's capture store with
35 of 35 query-carrying URLs. That consumer retracted the number the next day
and re-measured 28 of 28 on the same 81-URL corpus. LIMITATIONS was corrected
then; the comment was not, so a retracted figure has been sitting beside a
live one since 07-27.

717 tests (was 716, none changed), coverage matrix 128/128 with 6/6 gaps
holding, every population swept twice with identical counts.
2026-08-10 21:32:56 +02:00

233 lines
9.6 KiB
Python

"""False-positive sweep — run *benign* document populations through the persist
gate and count the ones it does not wave through.
WHAT THE NUMBER MEANS, EXACTLY. The unit is a **document**, not a URL, and the
metric is `screen_output(doc, PRESET_USER_UPLOAD).disposition is not WARN`. Every
word of that is load-bearing:
- **Document, not URL.** `docs/LIMITATIONS.md` also carries URL-level field
measurements (16 of 16, 28 of 28, 149 of 1694). Those are a different unit over
partly-overlapping corpora. A document rate is NOT comparable to them and must
never be combined with them, or quoted as an update to them.
- **The upload preset, not the trusted one.** `PRESET_TRUSTED_SOURCE` will hand
you a beautiful near-zero and mean nothing: every non-CRITICAL finding WARNs
under trust, which is exactly the structural blindness that let the 0.3.0
active-content regression ship through a green suite (see the docstring on
`tests/test_corpus.py::test_false_positive_is_not_blocked_on_the_upload_gate`).
The trusted door is printed as a footnote, never as the headline.
- **`screen_output`, not `_scan_input`.** The output gate is where
`active_content` lives; the input path never reaches it.
- **not WARN**, not "has findings". A finding is not a false positive — the
library reports and the pipeline decides (BRIEF design principle 4). WARN means
*persisted, with a note*, which is the benign outcome.
GROUND TRUTH for "benign" is **provenance, not inspection**: nobody hand-read
these documents. Each population is benign by where it came from — vendor-
published documentation, this machine's own generated notes, first-party authored
reference material. That is the only ground truth available at this scale, and it
is a real caveat, not a formality: an injected document sitting in a harvested
corpus would be scored as a false positive here.
POPULATIONS ARE NEVER SUMMED. Every population has its own denominator and its
own provenance; a pooled rate would be arithmetic over incommensurable things and
would inherit the `2400 != 2401` defect one level up. This script refuses to print
a total.
USAGE — corpus roots are arguments, never hardcoded; the corpora live in private
consumer repos and their paths must not reach a public mirror:
python docs/fp-sweep.py LABEL=/path/to/corpus [LABEL=/path ...]
[--ext=.md,.txt] [--include=/subtree/]
Each LABEL should name the population's *class* (`vendor-harvest`,
`generated-notes`, `reference-corpus`), not the repo it came from.
"""
from __future__ import annotations
import sys
from collections import Counter
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "src"))
from llm_ingestion_guard import ( # noqa: E402
DEFAULT_ACTION_MAP,
PRESET_TRUSTED_SOURCE,
PRESET_USER_UPLOAD,
Disposition,
Risk,
Source,
__version__,
scan_output,
screen_output,
severity_rank,
)
from llm_ingestion_guard.calibration import RISK_RANK # noqa: E402
BENIGN = Disposition.WARN
"""The benign outcome: persisted, with a note. Anything else costs a human."""
def check_metric_is_a_risk_statement() -> None:
"""Fail loudly if "not WARN" has stopped meaning "assessed ELEVATED or worse".
The published number is a count of non-WARN documents, but what it *claims*
is a statement about assessed risk. The two are the same statement only while
the default action map sends exactly `NONE` and `LOW` to WARN. Re-map that
and the published number silently changes meaning with no test failing — the
method trap this script exists to stay out of. Mirrored in the suite by
`tests/test_corpus.py::test_the_published_fp_metric_is_a_risk_statement`.
"""
elevated = RISK_RANK[Risk.ELEVATED.value]
for risk in Risk:
benign = DEFAULT_ACTION_MAP[risk] is BENIGN
below = RISK_RANK[risk.value] < elevated
if benign != below:
raise SystemExit(
f"metric invalid: {risk.value} maps to "
f"{DEFAULT_ACTION_MAP[risk].value}; 'not WARN' no longer means "
"'assessed ELEVATED or worse' and the published rate would be "
"a different claim than the doc makes"
)
def documents(root: Path, exts: tuple[str, ...], include: str = "") -> list[Path]:
"""Every non-hidden file under ``root`` with a wanted extension.
``include`` is a substring the *relative* path must contain, so a population
can be scoped to a subtree (`--include=/references/`) without pretending a
differently-scoped count is the same population. Two scopings of one tree are
two counts, and the difference between them is exactly the kind of thing
`docs/LIMITATIONS.md` has had to correct in public before.
"""
files = []
for p in sorted(root.rglob("*")):
if not p.is_file() or p.suffix not in exts:
continue
rel = p.relative_to(root)
if any(part.startswith(".") for part in rel.parts):
continue
if include and include not in f"/{rel}":
continue
files.append(p)
return files
def measure(label: str, root: Path, exts: tuple[str, ...], include: str = "") -> dict:
paths = documents(root, exts, include)
dispositions: Counter[str] = Counter()
assessments: Counter[str] = Counter()
trusted_dispositions: Counter[str] = Counter()
labels: Counter[str] = Counter()
drivers: Counter[str] = Counter()
offenders: list[tuple[str, str, tuple[str, ...]]] = []
empty = 0
for path in paths:
text = path.read_text(encoding="utf-8", errors="replace")
if not text.strip():
empty += 1
continue
result = screen_output(text, PRESET_USER_UPLOAD)
dispositions[result.disposition.value] += 1
assessments[result.assessment.value] += 1
trusted_dispositions[
screen_output(text, PRESET_TRUSTED_SOURCE).disposition.value
] += 1
if result.disposition is not BENIGN:
findings = scan_output(text, source=Source.OUTPUT).findings
found = tuple(sorted({f.label for f in findings}))
labels.update(found)
# What actually moved this document: the labels at its *worst*
# severity. A histogram of every label present would credit the
# over-block to whatever else happened to be in the document, which
# is how a residual gets blamed on the lexicon.
worst = max((severity_rank(f.severity) for f in findings), default=-1)
drivers[" + ".join(sorted({
f.label for f in findings if severity_rank(f.severity) == worst
})) or "(no findings)"] += 1
offenders.append((path.name, result.disposition.value, found))
n = sum(dispositions.values())
return {
"label": label,
"n": n,
"empty": empty,
"dispositions": dispositions,
"assessments": assessments,
"trusted": trusted_dispositions,
"labels": labels,
"drivers": drivers,
"offenders": offenders,
"non_warn": n - dispositions[BENIGN.value],
}
def report(m: dict) -> None:
n, non_warn = m["n"], m["non_warn"]
rate = f"{non_warn / n:.1%}" if n else "n/a"
print(f"\n## {m['label']}{non_warn} of {n} documents disposed non-WARN ({rate})")
if m["empty"]:
print(f" ({m['empty']} empty file(s) skipped — no document, no verdict)")
print(" upload gate :", dict(m["dispositions"]))
print(" assessment :", dict(m["assessments"]))
print(" trusted gate :", dict(m["trusted"]), " <- footnote only, structurally blind")
if m["drivers"]:
print(" what MOVED them (labels at each document's worst severity):")
for label, count in m["drivers"].most_common():
print(f" {count:5d} {label}")
if m["labels"]:
print(" what was merely present on them:")
for label, count in m["labels"].most_common():
print(f" {count:5d} {label}")
for name, disposition, found in m["offenders"][:10]:
print(f" - {name}: {disposition} via {', '.join(found) or '(no labels)'}")
if len(m["offenders"]) > 10:
print(f" ... and {len(m['offenders']) - 10} more")
def main() -> None:
argv = sys.argv[1:]
exts = (".md", ".txt")
include = ""
rest = []
for arg in argv:
if arg.startswith("--ext="):
exts = tuple(e if e.startswith(".") else f".{e}"
for e in arg.split("=", 1)[1].split(","))
elif arg.startswith("--include="):
include = arg.split("=", 1)[1]
else:
rest.append(arg)
if not rest:
print(__doc__)
raise SystemExit(2)
check_metric_is_a_risk_statement()
print(f"llm-ingestion-guard {__version__} — persist gate, PRESET_USER_UPLOAD")
print("metric: disposition is not WARN (== assessed ELEVATED or worse)")
print(f"extensions: {', '.join(exts)}")
measurements = []
for spec in rest:
if "=" not in spec:
raise SystemExit(f"expected LABEL=PATH, got {spec!r}")
label, _, path = spec.partition("=")
root = Path(path).expanduser()
if not root.is_dir():
raise SystemExit(f"{label}: {root} is not a directory")
measurements.append(measure(label, root, exts, include))
for m in measurements:
report(m)
print("\n---")
print("Populations are reported separately by construction. They have "
"different\nprovenance and different denominators; a pooled rate would "
"be arithmetic over\nincommensurable things. This script prints no total.")
if __name__ == "__main__":
main()