docs(consume): the Claude Code recipe, measured end to end on two bundles
Four questions, two bundles, one run each, in a scratch project outside this repository with a generated skill per bundle. All four passed, and zero numbers or identifiers appeared in any answer that were not in the delivered set or in the payload's own identities (62, 45 and 35 unique numeric tokens checked). The skill triggered WITHOUT being named in the prompt and selected the right one of two installed skills from the question alone, so no special invocation syntax is needed: the generated `description`, which carries the bundle id, the concept count and the ref, is enough to route on. One defect the runs found, and it was in the prose rather than the payload. The citation guidance listed the four locator keys this library writes, so on the 270-concept third-party bundle the model reported "no page locator, the address is at document level" while the excerpt in front of it carried `source_element_id` - that bundle's own locator, correctly delivered by the prefix rule. The guidance now tells the reader to cite whichever `source_*` keys are present. On the re-run the same question returned the element id. Two runs of one question, the second measuring a changed artefact and not retrying the first. One finding that is not a defect in this chain: the first attempt at a known-negative was not one. The bundle covers water and frost protection on 17 of its 270 concepts and the ranker put none of them in the cut. The consumer behaved exactly as the contract asks - refused, named its denominator, reported its own zero as unmeasured because `withheld` entries carry no titles, and did not go around the cut. Recorded as a retrieval miss rather than replaced, and it is the same shape as the open fusion finding. A correction to this session's own measurement is in the record too: a first sweep used `grep -rhoE "^source_[a-z_]+:"`, whose character class excludes digits, and so missed `source_sha256` on 270 of 270 concepts. A pattern that cannot match what it is looking for returns a zero that reads like a fact. README gains "Consume in Claude Code": folder to answer in three commands, every one of them run in this session. A test holds that the recipe invokes only scripts this repository ships, at the paths it names. Suite 1373 (1339 at the session baseline), ruff clean, mypy src clean. No version bump, no tag, no push. Co-Authored-By: Claude <claude-opus-5>
This commit is contained in:
parent
c95d18905a
commit
171798ed32
5 changed files with 477 additions and 15 deletions
|
|
@ -1118,6 +1118,7 @@ def test_no_corpus_document_name_reaches_any_file_this_work_tracks() -> None:
|
|||
PROJECT_ROOT / "docs" / "2026-09-08-prisform-og-loggen-k2.md",
|
||||
PROJECT_ROOT / "docs" / "2026-09-08-kravnummer-tokenisering.md",
|
||||
PROJECT_ROOT / "docs" / "2026-09-08-sjeldenhetsvekt.md",
|
||||
PROJECT_ROOT / "docs" / "2026-09-08-claude-code-skill-vilkaarlig-bundle.md",
|
||||
PROJECT_ROOT / "README.md",
|
||||
PROJECT_ROOT / "CLAUDE.md",
|
||||
]
|
||||
|
|
@ -1128,12 +1129,30 @@ def test_no_corpus_document_name_reaches_any_file_this_work_tracks() -> None:
|
|||
def test_the_readme_consume_section_states_the_rule_count_the_code_emits() -> None:
|
||||
# A published number must have a test that goes red when it goes false.
|
||||
readme = (PROJECT_ROOT / "README.md").read_text(encoding="utf-8")
|
||||
assert readme.count("## Consume") == 1
|
||||
assert readme.count("## Consume\n") == 1
|
||||
assert readme.count("## Consume in Claude Code\n") == 1
|
||||
assert len(okf_consume.WITHHOLDING_RULES) == 6
|
||||
assert "closed set of six" in readme
|
||||
assert "tools/okf_consume.py" in readme
|
||||
|
||||
|
||||
def test_the_readme_recipe_names_only_commands_this_repository_ships() -> None:
|
||||
# Every command in the "Consume in Claude Code" section was run in the
|
||||
# session that wrote it. This test cannot re-run them; what it can hold is
|
||||
# that each script the recipe invokes still exists under the path it names.
|
||||
readme = (PROJECT_ROOT / "README.md").read_text(encoding="utf-8")
|
||||
recipe = readme.split("## Consume in Claude Code", 1)[1].split("\n## ", 1)[0]
|
||||
scripts = set(re.findall(r"python3 (tools/\S+\.py)", recipe))
|
||||
assert scripts == {
|
||||
"tools/okf_skill.py",
|
||||
"tools/okf_consume.py",
|
||||
"tools/okf_contract_check.py",
|
||||
}, scripts
|
||||
for script in scripts:
|
||||
assert (PROJECT_ROOT / script).is_file(), script
|
||||
assert "okf build " in recipe
|
||||
|
||||
|
||||
# --- Step 11: the measurement scorer -----------------------------------------
|
||||
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue