`uv sync --frozen` resolved ruff 0.15.22 and the tree read clean. A loose install resolves 0.16.6, under which the SAME untouched code reports 148 findings -- 4 more than round 9 counted, because this round added four files. All of them are new rules rather than new defects: 0.16 widened the default rule set to whole families (YTT, ASYNC, PL, ISC, C4, UP, B, SIM, FURB, ...). (`[skip-docs]` is for CLAUDE.md, which a lint-configuration change does not reach. README's developer section IS updated in this commit.) THE DEFECT IS NOT THE 148, IT IS THAT NOBODY CHOSE THEM. `[tool.ruff]` set only `line-length` and `target-version`, so the acceptance was ruff's default, and the tree stayed green only as long as the lockfile froze an old ruff. `select` is now written down: `E4`, `E7`, `E9`, `F` (the historical default), `I` because this tree already keeps imports sorted, and `RUF100` so a `noqa` that has stopped meaning anything is caught rather than left as decoration. Pin `ruff>=0.9` -> `ruff>=0.16.6,<0.17`. Per rule, before -> after: RUF100 50 -> 0, I001 20 -> 0, ISC004 19, PLW1510 8, C408 8, EXE001 6, RUF007 5, PLE2515 4, UP031 3, B017 3, and fourteen more with 2 or fewer -- the families out of the declared set are 0 by selection, and 148 is the number to start from if they are adopted, which is a separate decision and not one to take inside a version-pin commit. 57 were auto-fixed; one E402 was reintroduced by the import-sorting fix merging a block away from its `noqa`, and got the directive back rather than a bare one. `S` IS MEASURED OUT, NOT ASSUMED OUT: it reports 2657 `S101` on a suite whose every assertion is an `assert`, and `S603` flags 19 subprocess calls of which one was ever marked -- selecting it buys 18 suppressions and no defect. Two `noqa` directives naming non-selected rules were dropped with that reason recorded in the configuration instead. THE TWO FILES 0.16 WOULD REFORMAT ARE MARKDOWN, NOT PYTHON: `README.md` and `docs/2026-09-08-blindsone-below-k-k2.md`. 0.16 formats fenced Python inside markdown, and both blocks are RECORDS -- the second is a quotation of `COST_VOCABULARY` as it stood when that measurement was taken. Reformatting a quotation makes it stop being one, so markdown is excluded from the formatter and `ruff format --check .` stays in the acceptance over `.py`. `tools/okf_consume_measure.py` is fenced by the order as run-not-edited, so its three findings are exempted by path with the reason and the debt named, and its bytes are untouched. THE LOCKFILE TRAP IS CLOSED, NOT AVOIDED. `uv.lock` predated the `[ocr]` extra, so any unlocked resolve wrote that extra's transitive tree back into it -- 681 insertions over 4 deletions, twice now, and round 9 recorded the cause as `uv run` OUTSIDE the project when it is `uv run` without `--frozen` INSIDE it. The relock is complete for every declared extra (703 insertions, 26 deletions), and measured after it, an unfrozen `uv run` leaves the file alone. `ruff check src tests tools`, `ruff format --check .` (0.16.6), `mypy src` over 21 files and 1535 tests, all green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
50 lines
1.8 KiB
Python
50 lines
1.8 KiB
Python
"""The CID-share instrument: order 20260904T172353Z-6290714297-from-.claude.
|
|
|
|
Measures how much of a document's extracted text is undecoded `(cid:N)` glyph
|
|
codes -- the failure mode `docs/2026-09-04-k3-arm-c.md` found on Bilag 9.1
|
|
(95.1 %, 98 alphabetic words of four or more letters survive). This instrument
|
|
answers whether that document is alone or whether the corpus has more of them.
|
|
|
|
The negative control matters here as much as the positive one: an instrument
|
|
that reports a nonzero CID share on text that has none would inflate every
|
|
number it ever produces, the same reason `test_fidelity.py` pins its own
|
|
negative control.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
|
|
|
|
import okf_cid_measure
|
|
|
|
|
|
def test_known_cid_fraction() -> None:
|
|
text = "abcd efgh " + "(cid:12)" * 2
|
|
result = okf_cid_measure.measure(text)
|
|
assert result.total_chars == len(text) == 26
|
|
assert result.cid_chars == len("(cid:12)") * 2 == 16
|
|
assert result.word_count == 2 # "abcd", "efgh" -- "cid" itself is 3 letters
|
|
assert abs(result.pct - (16 / 26 * 100)) < 1e-9
|
|
|
|
|
|
def test_no_cid_is_zero_with_nonzero_total() -> None:
|
|
result = okf_cid_measure.measure("plain readable text with no cid codes at all")
|
|
assert result.total_chars > 0
|
|
assert result.cid_chars == 0
|
|
assert result.pct == 0.0
|
|
|
|
|
|
def test_all_cid_is_full_share_and_zero_words() -> None:
|
|
result = okf_cid_measure.measure("(cid:1)(cid:2)(cid:3)")
|
|
assert result.pct == 100.0
|
|
assert result.word_count == 0
|
|
|
|
|
|
def test_empty_text_reports_zero_percent_not_a_crash() -> None:
|
|
result = okf_cid_measure.measure("")
|
|
assert result.total_chars == 0
|
|
assert result.cid_chars == 0
|
|
assert result.pct == 0.0
|