feat(prepass,explore): padding dies at the prompt; a refusal the model can act on is a return value [skip-docs]

Order 20260908T195801Z. Findings 4 and 5 from the S7 acid test, then the two things
finding 99 measured and deliberately did not fix (D3, D2).

No user-facing surface changes: no new flag, no new command, no changed output
contract. Both seams are internal (the pre-pass rendering, and the shape a tool
answers a model with), so [skip-docs] rather than a README edit that would describe
nothing an operator can do differently.

FINDING 4 -- MEASURED, NOTHING BUILT. K2's price schedule IS readable without
guessing (8 column spans, 71 of 91 non-blank rows give >= 2 cells, the split stable
for K = 2..64). But 0 of 92 rows name all three of code/quantity/unit_cost -- also
under a looser substring match -- and 0 of 91 data rows carry code + quantity +
amount. The triple is not formatted away; it is not in the document. It is a price
SUMMARY plus nine rate cards whose unit-price columns are empty (pre-award). The
order's binding decision rule therefore falls against building:
--derive-cost-baseline keeps refusing, and MAJOR-4's own honesty limit holds.

FINDING 5 -- BUILT. Measured on the actual rendering path (concept_text, not the
raw file): the delivered excerpt is 104 lines / 67 245 chars, carrying 208 interior
whitespace runs, 117 of them >= 100 and the longest 887 -- 56 806 of 67 245
characters = 84.5 %, over 72 of 104 lines. collapse_padding, called from
_data_blocks (the one renderer both arms share, and therefore AFTER
verify_against_bundle -- collapsing in concept_text would break every payload's own
digest), gives -72.4 %: line count invariant, non-whitespace byte-identical, leading
indentation untouched, no number changed.

F99-D3 -- read_file / read_dir / read_bundle now RETURN their refusal. MAF turns a
tool raise into "Error: Function failed." (_tools.py:1410-1432, :1427) and counts it
against DEFAULT_MAX_CONSECUTIVE_ERRORS_PER_REQUEST = 3, so everything the refusing
arm knows is destroyed on the way out. The gates are unchanged; the property they
exist for -- the reason travels, the bytes never do -- is now asserted explicitly on
the returned value. The arm is keyed on named classes, never bare Exception, because
ExplorationError is itself a RuntimeError subclass.

F99-D2 -- the invariant row, plus one for finding 5 (a stated deviation from "one
row only": finding 5 is a separately built seam and the ledger's standing rule
requires its own row).

19 existing arms rewritten, never deleted and never weakened: where the class
carried a distinction, the refusal KIND carries it now.

13 mutations, all red against the whole suite (W1-W5, M1-M8), each restored from
scratchpad with shasum -c. Control 1543 passed / 5 skipped (from 1529/5, a strict
superset, 0 removed). Golden demo-transcript.stdout unchanged
(shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Measurement: docs/2026-09-08-funn-4-5-og-read-nekt.md

Co-Authored-By: Claude <Opus 5>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-08 23:36:54 +02:00
commit eb4137415c
13 changed files with 975 additions and 69 deletions

View file

@ -758,8 +758,13 @@ def test_an_unknown_knowledge_base_is_refused_by_name() -> None:
failed."`` and counts three of them as grounds to stop function calling for the whole request —
so the one thing this refusal knows and the model does not (which ids exist) never reached it.
The property T20 was written for is untouched and asserted here in both halves: the answer is
still a refusal, and it still names what IS configured. The ``navigator_tools`` half is what
keeps the change SCOPED those three tools still raise.
still a refusal, and it still names what IS configured.
REWRITTEN A SECOND TIME (F99-D3, order 20260908T195801Z), for the reason the note above gave
for the first rewrite: the three navigator tools no longer raise either, so the same
measurement now applies to them and the arm asserts the same two halves through their return
value. What is NOT relaxed is the class the arm is keyed on a failure nobody named still
propagates, gated by ``test_explore_read_tools_refusal``'s known-negative.
"""
validate = explore.quick_validate_tool(("/tmp/base-a",))
verdict = validate.func(bundle_id="base-b", proposal_json="{}")
@ -767,9 +772,8 @@ def test_an_unknown_knowledge_base_is_refused_by_name() -> None:
assert "base-b" in verdict["reason"] and "base-a" in verdict["reason"]
read_file = next(t for t in explore.navigator_tools(("/tmp/base-a",)) if t.name == "read_file")
with pytest.raises(explore.ExplorationError) as excinfo:
read_file.func(bundle_id="base-b", path="x.md")
assert "base-a" in str(excinfo.value)
answer = read_file.func(bundle_id="base-b", path="x.md")
assert answer.startswith("REFUSED (") and "base-a" in answer
def test_two_bases_with_the_same_name_are_refused() -> None: