docs(front-page): the gate numbers the gate actually prints, and a breaking point that was measured

Four claims on the front page were false on this commit, and one of them was a
number no division ever produced.

**The retrieval gate.** README reported it RED on rows 3, 4, 5, 7, 8 and 9,
with row 3 at 2 of 5 and row 4 at 3 of 6. Run on this commit it is RED on rows
5, 7, 8 and 9, with row 3 at 5 of 5 and row 4 at 6 of 6: `f81683e` made a
withheld concept carry the rule that actually decided it, and `05cb190` gave
the payload a `coverage` block, and neither updated the table. Row 8 is `0 of 3
| NOT RUN` on the default run and was published as `44 of 64 questions`, which
is what it scores the day all three private sets are handed to it -- now
labelled with the day and the machine rather than printed as a row. The same
four figures were stale in `CLAUDE.md`.

**The breaking point in a generated skill.** `int(LIMIT / per_withheld) if
per_withheld else 0` printed `At roughly 0 concepts the bookkeeping alone
reaches the 120000-byte limit` whenever the generation run withheld nothing --
the absence of a measurement, rendered as one, and read as a bundle that breaks
before it holds anything. A run with no withheld entry has no slope to
extrapolate from, so the sentence is withheld with its reason. The shipped
`skills/okf-consume/SKILL.md` is generated with the question its
`references/README.md` names, withholds nothing, and carried exactly that `0`;
it is regenerated. Two arms in the test, because one would pass on an empty
set: the bundles that withhold something must still state a positive figure.

The sentence for that arm also stopped saying `**4 bytes** for 3 concepts`
where the 4 bytes were the cost of 0 withheld entries. It is now `for N of M
concepts`, which moves two generated skills' line counts and therefore the
published comparison: 280 of 312 and 310 -> 281 of 313 and 311, re-measured,
with the 62 differing lines unchanged.

**Four tools.** A single-bundle server exposes three: `okf_list` is absent
where there is nothing to list. README's table already said so in a cell; the
heading and the CHANGELOG did not.

**What `--accounting` accounts for.** The account is over the element classes
each format's vocabulary names, verified against `accounting._READERS` rather
than against the report: a file whose suffix has no reader is accounted at file
level only, `.docx` reads `document.xml` and `footnotes.xml` (so headers,
footers, endnotes and comments are outside), `.pptx` reads the slides (so
speaker notes are outside), `.xlsx` reads the worksheets (so cell comments are
outside and a cell contributes its cached value, never its formula), and `.rtf`
skips its header and footer groups. A hidden slide or sheet IS counted -- it
lives in the same part as a visible one. Nothing is built for this; the list is
what `0 unaccounted` does not claim.

Gates re-run on the commit: retrieval `GATE RED: rows 5, 7, 8, 9` (exit 1),
MCP `GATE RED: rows 2` (exit 1), both matching what is now written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-20 15:40:20 +02:00
commit d300338e4d
Signed by: ktg
SSH key fingerprint: SHA256:JakMjO6FTBBzN0Bhfj9saOoEjaFxlSdYuZQQpM/lF9Q
9 changed files with 203 additions and 754 deletions

View file

@ -1,416 +0,0 @@
"""The adjudication command: it records a judgement, and never invents one.
A proposal a human has not looked at must never be replayable as an
adjudication, because replay is exactly what the run path does with a plan --
deterministically and forever. So this command writes a SIBLING record and
leaves the proposal untouched: the two files together say who judged what,
when, and how long it took, and either can be re-read against the other.
Three properties are pinned here rather than described:
- **The model leg is OFF by default.** Pre-annotation has been measured
LOWERING a good annotator's accuracy, from 98.1 % to 95.8 %, so a leg that
cannot be switched off is a leg whose value can never be measured. With it
off, no process is spawned at all -- asserted by breaking `subprocess.run`.
- **The CLI is named, and the other one is excluded BY NAME.** The model leg
shells out to the `claude` CLI. `gemini` is not merely unmentioned; its
absence from the module is a test, because "we did not use it" and "nothing
stops us using it" look identical in a review.
- **Dwell time travels with the verdict** (PM decision B2). A ratified flag
with no per-item time is unfalsifiable, and it is the same number that makes
adjudication throughput measurable at all.
It lives outside `src/`, so it never enters a wheel and no consumer's install
surface changes because it exists.
"""
from __future__ import annotations
import json
import subprocess
import sys
from pathlib import Path
from typing import Any
import pytest
from llm_ingestion_okf.errors import SegmentationError
from llm_ingestion_okf.segmentation import parse_segmentation_plan
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
import okf_adjudicate
from llm_ingestion_okf import propose as okf_propose_segments
DOCUMENT = """# N500 Vegbygging
Innledende tekst om vegbygging og dens omfang.
## 3.1 Brannkonsept
Krav til seksjonering av bygget.
## 3.2 Roemning
To uavhengige roemningsveier.
"""
ADJUDICATOR = "ktg"
AT = "2026-09-02T10:00:00Z"
def proposal(tmp_path: Path) -> Path:
source = tmp_path / "n500.md"
source.write_text(DOCUMENT, encoding="utf-8", newline="")
out = tmp_path / "plan.json"
assert okf_propose_segments.main([str(source), "--out", str(out), "--proposed-at", AT]) == 0
return out
def adjudicate(tmp_path: Path, *extra: str) -> tuple[int, Path]:
verdict = tmp_path / "adjudicated.json"
code = okf_adjudicate.main(
[
"--plan",
str(proposal(tmp_path)),
"--out",
str(verdict),
"--adjudicator",
ADJUDICATOR,
"--adjudicated-at",
AT,
*extra,
]
)
return code, verdict
def payload(path: Path) -> dict[str, Any]:
return json.loads(path.read_text(encoding="utf-8"))
def test_the_proposal_survives_untouched(tmp_path: Path) -> None:
plan_path = proposal(tmp_path)
before = plan_path.read_bytes()
okf_adjudicate.main(
[
"--plan",
str(plan_path),
"--out",
str(tmp_path / "adjudicated.json"),
"--adjudicator",
ADJUDICATOR,
"--adjudicated-at",
AT,
]
)
assert plan_path.read_bytes() == before
def test_the_verdict_records_adjudicator_timestamp_and_dwell(tmp_path: Path) -> None:
code, verdict = adjudicate(tmp_path)
assert code == 0
written = payload(verdict)
assert written["adjudicated"] is True
for entry in written["entries"]:
record = entry["adjudication"]
assert record["adjudicated_by"] == ADJUDICATOR
assert record["adjudicated_at"] == AT
assert isinstance(record["adjudication_dwell_s"], int)
assert not isinstance(record["adjudication_dwell_s"], bool)
def test_the_verdict_parses_as_a_segmentation_plan(tmp_path: Path) -> None:
_, verdict = adjudicate(tmp_path)
parsed = parse_segmentation_plan(payload(verdict))
assert parsed.adjudicated is True
assert all(entry.adjudication is not None for entry in parsed.entries)
def test_replaying_the_same_verdict_produces_identical_bytes(tmp_path: Path) -> None:
"""K4a's mechanism: an adjudication is data, so a re-run is a copy."""
_, first = adjudicate(tmp_path)
kept = first.read_bytes()
second = tmp_path / "again.json"
okf_adjudicate.main(
[
"--plan",
str(tmp_path / "plan.json"),
"--out",
str(second),
"--adjudicator",
ADJUDICATOR,
"--adjudicated-at",
AT,
]
)
assert second.read_bytes() == kept
def test_with_the_model_leg_off_no_process_is_spawned(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""Asserted by BREAKING the spawn, not by reading the code.
A test that merely inspects the default would pass just as happily if the
default were ignored.
"""
def refuse(*args: object, **kwargs: object) -> None:
raise AssertionError("the model leg spawned a process while switched off")
monkeypatch.setattr(subprocess, "run", refuse)
code, _ = adjudicate(tmp_path)
assert code == 0
def test_the_model_leg_is_off_unless_asked_for(tmp_path: Path) -> None:
assert (
okf_adjudicate.parse_args(
["--plan", "p", "--out", "o", "--adjudicator", "a", "--adjudicated-at", AT]
).model
is None
)
def test_the_resolved_argv_starts_with_the_claude_binary() -> None:
argv = okf_adjudicate.model_argv("claude-opus-5", "spoersmaal")
assert argv[0] == okf_adjudicate.CLAUDE_CLI
assert Path(argv[0]).name == "claude"
assert "--model" in argv
assert argv[argv.index("--model") + 1] == "claude-opus-5"
def test_the_other_cli_is_excluded_by_name_not_merely_unused() -> None:
""" "We did not use it" and "nothing stops us using it" look identical in a
review. This is the difference, as a measurement."""
module = Path(okf_adjudicate.__file__).read_text(encoding="utf-8")
assert "gemini" not in module.lower()
def test_the_gemini_check_can_actually_fire() -> None:
"""The negative control for the check above: prove it can find the word."""
assert "gemini" in "a line naming gemini".lower()
def test_a_missing_plan_exits_two_and_says_so(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
code = okf_adjudicate.main(
[
"--plan",
str(tmp_path / "nothing.json"),
"--out",
str(tmp_path / "out.json"),
"--adjudicator",
ADJUDICATOR,
"--adjudicated-at",
AT,
]
)
assert code == 2
assert "nothing.json" in capsys.readouterr().err
# --- the empty plan: a judgement with nothing to keep -----------------------
#
# Measured on the K3 corpus: 4 of 12 judgements produced no artifact at all,
# because the adjudicator's verdict was "none of these segments should be
# persisted" and the parser refuses a plan with zero entries. That refusal is
# CORRECT for the run path -- an empty plan would silently persist nothing for a
# document that was dropped -- so the grammar is left alone and the recording
# tool is taught to record a rejection. The two are different acts: refusing to
# materialize is about a bundle, recording a judgement is about a person.
def empty_proposal(tmp_path: Path) -> Path:
"""A real proposal with its entries removed -- the plan-level fields stay
exactly as the proposer wrote them, so this is a rejection and not a stub."""
plan_path = proposal(tmp_path)
written = payload(plan_path)
written["entries"] = []
rejected = tmp_path / "rejected.json"
rejected.write_text(json.dumps(written, indent=2) + "\n", encoding="utf-8", newline="")
return rejected
def adjudicate_empty(
tmp_path: Path, plan_path: Path, out_name: str = "verdict.json"
) -> tuple[int, Path]:
verdict = tmp_path / out_name
code = okf_adjudicate.main(
[
"--plan",
str(plan_path),
"--out",
str(verdict),
"--adjudicator",
ADJUDICATOR,
"--adjudicated-at",
AT,
]
)
return code, verdict
def test_a_judgement_over_an_empty_plan_gets_an_artifact(tmp_path: Path) -> None:
"""The defect this closes: the judgement happened and left no trace."""
code, verdict = adjudicate_empty(tmp_path, empty_proposal(tmp_path))
assert code == 0
assert verdict.is_file()
written = payload(verdict)
assert written["entries"] == []
assert written["adjudicated"] is True
assert written["adjudicated_by"] == ADJUDICATOR
assert written["adjudicated_at"] == AT
def test_the_empty_verdict_carries_the_dwell_time_at_the_top(tmp_path: Path) -> None:
"""There is no entry to hang it on, and a ratified rejection with no time on
it is as unfalsifiable as a ratified acceptance with none."""
_, verdict = adjudicate_empty(tmp_path, empty_proposal(tmp_path))
written = payload(verdict)
assert isinstance(written["adjudication_dwell_s"], int)
assert not isinstance(written["adjudication_dwell_s"], bool)
assert written["adjudication_dwell_s"] > 0
def test_the_empty_verdict_is_not_replayable_by_the_run_path(tmp_path: Path) -> None:
"""The grammar is UNCHANGED. Recording a rejection and materializing from it
are different acts, and only the first one is now possible."""
_, verdict = adjudicate_empty(tmp_path, empty_proposal(tmp_path))
with pytest.raises(SegmentationError) as excinfo:
parse_segmentation_plan(payload(verdict))
assert excinfo.value.code == "segmentation_plan_invalid"
def test_the_rejected_proposal_survives_untouched(tmp_path: Path) -> None:
plan_path = empty_proposal(tmp_path)
before = plan_path.read_bytes()
adjudicate_empty(tmp_path, plan_path)
assert plan_path.read_bytes() == before
def test_replaying_an_empty_verdict_produces_identical_bytes(tmp_path: Path) -> None:
"""K4a over the arm that had no artifact to compare before."""
plan_path = empty_proposal(tmp_path)
_, first = adjudicate_empty(tmp_path, plan_path, "first.json")
kept = first.read_bytes()
_, second = adjudicate_empty(tmp_path, plan_path, "second.json")
assert second.read_bytes() == kept
def test_an_empty_plan_missing_a_required_field_is_still_refused(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""The empty branch is not a hole in the validation: a plan is still a plan,
and only its entry list is allowed to be empty."""
plan_path = empty_proposal(tmp_path)
written = payload(plan_path)
del written["source_sha256"]
plan_path.write_text(json.dumps(written), encoding="utf-8", newline="")
code, verdict = adjudicate_empty(tmp_path, plan_path)
assert code == 2
assert "source_sha256" in capsys.readouterr().err
assert not verdict.exists()
def test_an_entries_value_that_is_not_a_list_is_still_refused(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""Empty is a judgement; the wrong TYPE is a malformed plan, and the two
must not collapse. The malformed one still meets the unchanged grammar.
It reaches the caller as EXIT 2, the same code every other malformed plan
already got. A grammar refusal used to escape as a traceback and exit 1,
which said "this command crashed" where the truth was "this file is not a
plan" -- and exit codes are the interface callers script against."""
plan_path = empty_proposal(tmp_path)
written = payload(plan_path)
written["entries"] = "none"
plan_path.write_text(json.dumps(written), encoding="utf-8", newline="")
code, _ = adjudicate_empty(tmp_path, plan_path)
assert code == 2
assert "segmentation_plan_invalid" in capsys.readouterr().err
assert not (tmp_path / "verdict.json").exists()
def test_a_verdict_this_tool_cannot_read_back_is_not_blamed_on_the_plan(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""The plan is parsed BEFORE the verdict is built, and the verdict is parsed
after. Only the first failure is the plan's.
A verdict that will not parse back is THIS command failing on what it was
told to stamp -- here an empty `--adjudicator`. Reporting that as a
malformed plan sends the operator to fix the one artifact that was fine,
which is worse than the traceback it replaced: a clean, confident, wrong
diagnosis."""
plan_path = proposal(tmp_path)
capsys.readouterr() # the proposer's own report is not what is under test
code = okf_adjudicate.main(
[
"--plan",
str(plan_path),
"--out",
str(tmp_path / "verdict.json"),
"--adjudicator",
"",
"--adjudicated-at",
AT,
]
)
assert code == 2
stderr = capsys.readouterr().err
assert "malformed plan" not in stderr
assert "verdict" in stderr
assert not (tmp_path / "verdict.json").exists()
def test_a_malformed_non_empty_plan_exits_two_with_one_stderr_line(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""The same refusal on the other branch: a plan with entries is parsed
BEFORE anything is written, and that parse failing is a malformed plan too.
One line, because a caller reading stderr to tell malformed from missing
should not have to parse a traceback to do it."""
plan_path = proposal(tmp_path)
written = payload(plan_path)
written["entries"][0]["span"] = [10, 3]
plan_path.write_text(json.dumps(written), encoding="utf-8", newline="")
capsys.readouterr() # the proposer's own report is not what is under test
code = okf_adjudicate.main(
[
"--plan",
str(plan_path),
"--out",
str(tmp_path / "verdict.json"),
"--adjudicator",
ADJUDICATOR,
"--adjudicated-at",
AT,
]
)
assert code == 2
stderr = capsys.readouterr().err
assert len(stderr.strip().splitlines()) == 1
assert "segmentation_span_invalid" in stderr
assert not (tmp_path / "verdict.json").exists()

View file

@ -363,3 +363,55 @@ def test_a_project_skill_still_passes_the_contract_checker(tmp_path: Path) -> No
payload = okf_consume.build_payload(GOLDEN, question="hva er kravet til pris?")
report = okf_contract_check.check(written.read_text(encoding="utf-8"), payload)
assert report.findings == ()
@pytest.mark.parametrize("bundle", BUNDLES, ids=lambda path: path.name)
def test_the_breaking_point_is_a_measurement_or_it_is_withheld(
bundle: Path, tmp_path: Path
) -> None:
"""`0 concepts` was a division that never happened, printed as a number.
The figure is EXTRAPOLATED from what one `withheld` entry costs, so a
generation run that withheld nothing has no slope to extrapolate from:
`per_withheld` was `0.0`, the guard returned the literal `0`, and the
document told its reader the bundle's bookkeeping fills a 120000-byte
budget at zero concepts -- before the bundle holds anything at all.
Driven from both sides so a generator that simply stopped stating the
figure would fail: the bundle that withholds nothing must say it could not
measure it, and a bundle that withholds something must still print a
positive count.
"""
written = _generate(bundle, tmp_path / "out")
text = written.read_text(encoding="utf-8")
payload = json.loads((tmp_path / "out" / "references" / "example-payload.json").read_text())
assert payload["withheld"], "the known-positive arm withheld nothing to extrapolate from"
assert "**0 concepts**" not in text
stated = re.search(r"At roughly\s+\*\*(\d+) concepts\*\*", text)
assert stated is not None, "a bundle that withheld something states no figure"
assert int(stated.group(1)) > 0
def test_a_generation_that_withheld_nothing_says_so_instead_of_printing_zero(
tmp_path: Path,
) -> None:
"""The arm the SHIPPED skill is on, and the one that was wrong.
`okf skill --example-question "Hva sier veiledningen om krav?"` delivers
all three concepts of the golden bundle, so `withheld` is empty and there
is no per-entry cost. The question is part of what the shipped file is
(`skills/okf-consume/references/README.md`), which is why the defect was
in the repository rather than only reachable in theory.
"""
written = okf_skill.generate(
GOLDEN,
out=tmp_path / "out",
question="Hva sier veiledningen om krav?",
force=True,
)
text = written.read_text(encoding="utf-8")
payload = json.loads((tmp_path / "out" / "references" / "example-payload.json").read_text())
assert payload["withheld"] == [], "the premise of this arm no longer holds"
assert "**0 concepts**" not in text
assert "breaking point could not be measured" in text
assert "At roughly" not in text