llm-ingestion-okf/tests/test_cid_measure.py
Kjell Tore Guttormsen 36c201cc8a chore(ruff): the acceptance was whatever the default happened to be [skip-docs]
`uv sync --frozen` resolved ruff 0.15.22 and the tree read clean. A loose
install resolves 0.16.6, under which the SAME untouched code reports 148
findings -- 4 more than round 9 counted, because this round added four files.
All of them are new rules rather than new defects: 0.16 widened the default
rule set to whole families (YTT, ASYNC, PL, ISC, C4, UP, B, SIM, FURB, ...).

(`[skip-docs]` is for CLAUDE.md, which a lint-configuration change does not
reach. README's developer section IS updated in this commit.)

THE DEFECT IS NOT THE 148, IT IS THAT NOBODY CHOSE THEM. `[tool.ruff]` set only
`line-length` and `target-version`, so the acceptance was ruff's default, and
the tree stayed green only as long as the lockfile froze an old ruff. `select`
is now written down: `E4`, `E7`, `E9`, `F` (the historical default), `I`
because this tree already keeps imports sorted, and `RUF100` so a `noqa` that
has stopped meaning anything is caught rather than left as decoration. Pin
`ruff>=0.9` -> `ruff>=0.16.6,<0.17`.

Per rule, before -> after: RUF100 50 -> 0, I001 20 -> 0, ISC004 19, PLW1510 8,
C408 8, EXE001 6, RUF007 5, PLE2515 4, UP031 3, B017 3, and fourteen more with
2 or fewer -- the families out of the declared set are 0 by selection, and 148
is the number to start from if they are adopted, which is a separate decision
and not one to take inside a version-pin commit. 57 were auto-fixed; one E402
was reintroduced by the import-sorting fix merging a block away from its
`noqa`, and got the directive back rather than a bare one.

`S` IS MEASURED OUT, NOT ASSUMED OUT: it reports 2657 `S101` on a suite whose
every assertion is an `assert`, and `S603` flags 19 subprocess calls of which
one was ever marked -- selecting it buys 18 suppressions and no defect. Two
`noqa` directives naming non-selected rules were dropped with that reason
recorded in the configuration instead.

THE TWO FILES 0.16 WOULD REFORMAT ARE MARKDOWN, NOT PYTHON: `README.md` and
`docs/2026-09-08-blindsone-below-k-k2.md`. 0.16 formats fenced Python inside
markdown, and both blocks are RECORDS -- the second is a quotation of
`COST_VOCABULARY` as it stood when that measurement was taken. Reformatting a
quotation makes it stop being one, so markdown is excluded from the formatter
and `ruff format --check .` stays in the acceptance over `.py`.

`tools/okf_consume_measure.py` is fenced by the order as run-not-edited, so its
three findings are exempted by path with the reason and the debt named, and its
bytes are untouched.

THE LOCKFILE TRAP IS CLOSED, NOT AVOIDED. `uv.lock` predated the `[ocr]` extra,
so any unlocked resolve wrote that extra's transitive tree back into it -- 681
insertions over 4 deletions, twice now, and round 9 recorded the cause as
`uv run` OUTSIDE the project when it is `uv run` without `--frozen` INSIDE it.
The relock is complete for every declared extra (703 insertions, 26 deletions),
and measured after it, an unfrozen `uv run` leaves the file alone.

`ruff check src tests tools`, `ruff format --check .` (0.16.6), `mypy src` over
21 files and 1535 tests, all green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 23:15:17 +02:00

50 lines
1.8 KiB
Python

"""The CID-share instrument: order 20260904T172353Z-6290714297-from-.claude.
Measures how much of a document's extracted text is undecoded `(cid:N)` glyph
codes -- the failure mode `docs/2026-09-04-k3-arm-c.md` found on Bilag 9.1
(95.1 %, 98 alphabetic words of four or more letters survive). This instrument
answers whether that document is alone or whether the corpus has more of them.
The negative control matters here as much as the positive one: an instrument
that reports a nonzero CID share on text that has none would inflate every
number it ever produces, the same reason `test_fidelity.py` pins its own
negative control.
"""
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
import okf_cid_measure
def test_known_cid_fraction() -> None:
text = "abcd efgh " + "(cid:12)" * 2
result = okf_cid_measure.measure(text)
assert result.total_chars == len(text) == 26
assert result.cid_chars == len("(cid:12)") * 2 == 16
assert result.word_count == 2 # "abcd", "efgh" -- "cid" itself is 3 letters
assert abs(result.pct - (16 / 26 * 100)) < 1e-9
def test_no_cid_is_zero_with_nonzero_total() -> None:
result = okf_cid_measure.measure("plain readable text with no cid codes at all")
assert result.total_chars > 0
assert result.cid_chars == 0
assert result.pct == 0.0
def test_all_cid_is_full_share_and_zero_words() -> None:
result = okf_cid_measure.measure("(cid:1)(cid:2)(cid:3)")
assert result.pct == 100.0
assert result.word_count == 0
def test_empty_text_reports_zero_percent_not_a_crash() -> None:
result = okf_cid_measure.measure("")
assert result.total_chars == 0
assert result.cid_chars == 0
assert result.pct == 0.0