feat(prepass,explore): padding dies at the prompt; a refusal the model can act on is a return value [skip-docs]
Order 20260908T195801Z. Findings 4 and 5 from the S7 acid test, then the two things finding 99 measured and deliberately did not fix (D3, D2). No user-facing surface changes: no new flag, no new command, no changed output contract. Both seams are internal (the pre-pass rendering, and the shape a tool answers a model with), so [skip-docs] rather than a README edit that would describe nothing an operator can do differently. FINDING 4 -- MEASURED, NOTHING BUILT. K2's price schedule IS readable without guessing (8 column spans, 71 of 91 non-blank rows give >= 2 cells, the split stable for K = 2..64). But 0 of 92 rows name all three of code/quantity/unit_cost -- also under a looser substring match -- and 0 of 91 data rows carry code + quantity + amount. The triple is not formatted away; it is not in the document. It is a price SUMMARY plus nine rate cards whose unit-price columns are empty (pre-award). The order's binding decision rule therefore falls against building: --derive-cost-baseline keeps refusing, and MAJOR-4's own honesty limit holds. FINDING 5 -- BUILT. Measured on the actual rendering path (concept_text, not the raw file): the delivered excerpt is 104 lines / 67 245 chars, carrying 208 interior whitespace runs, 117 of them >= 100 and the longest 887 -- 56 806 of 67 245 characters = 84.5 %, over 72 of 104 lines. collapse_padding, called from _data_blocks (the one renderer both arms share, and therefore AFTER verify_against_bundle -- collapsing in concept_text would break every payload's own digest), gives -72.4 %: line count invariant, non-whitespace byte-identical, leading indentation untouched, no number changed. F99-D3 -- read_file / read_dir / read_bundle now RETURN their refusal. MAF turns a tool raise into "Error: Function failed." (_tools.py:1410-1432, :1427) and counts it against DEFAULT_MAX_CONSECUTIVE_ERRORS_PER_REQUEST = 3, so everything the refusing arm knows is destroyed on the way out. The gates are unchanged; the property they exist for -- the reason travels, the bytes never do -- is now asserted explicitly on the returned value. The arm is keyed on named classes, never bare Exception, because ExplorationError is itself a RuntimeError subclass. F99-D2 -- the invariant row, plus one for finding 5 (a stated deviation from "one row only": finding 5 is a separately built seam and the ledger's standing rule requires its own row). 19 existing arms rewritten, never deleted and never weakened: where the class carried a distinction, the refusal KIND carries it now. 13 mutations, all red against the whole suite (W1-W5, M1-M8), each restored from scratchpad with shasum -c. Control 1543 passed / 5 skipped (from 1529/5, a strict superset, 0 removed). Golden demo-transcript.stdout unchanged (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f). Measurement: docs/2026-09-08-funn-4-5-og-read-nekt.md Co-Authored-By: Claude <Opus 5>
This commit is contained in:
parent
078a099898
commit
eb4137415c
13 changed files with 975 additions and 69 deletions
|
|
@ -758,8 +758,13 @@ def test_an_unknown_knowledge_base_is_refused_by_name() -> None:
|
|||
failed."`` and counts three of them as grounds to stop function calling for the whole request —
|
||||
so the one thing this refusal knows and the model does not (which ids exist) never reached it.
|
||||
The property T20 was written for is untouched and asserted here in both halves: the answer is
|
||||
still a refusal, and it still names what IS configured. The ``navigator_tools`` half is what
|
||||
keeps the change SCOPED — those three tools still raise.
|
||||
still a refusal, and it still names what IS configured.
|
||||
|
||||
REWRITTEN A SECOND TIME (F99-D3, order 20260908T195801Z), for the reason the note above gave
|
||||
for the first rewrite: the three navigator tools no longer raise either, so the same
|
||||
measurement now applies to them and the arm asserts the same two halves through their return
|
||||
value. What is NOT relaxed is the class the arm is keyed on — a failure nobody named still
|
||||
propagates, gated by ``test_explore_read_tools_refusal``'s known-negative.
|
||||
"""
|
||||
validate = explore.quick_validate_tool(("/tmp/base-a",))
|
||||
verdict = validate.func(bundle_id="base-b", proposal_json="{}")
|
||||
|
|
@ -767,9 +772,8 @@ def test_an_unknown_knowledge_base_is_refused_by_name() -> None:
|
|||
assert "base-b" in verdict["reason"] and "base-a" in verdict["reason"]
|
||||
|
||||
read_file = next(t for t in explore.navigator_tools(("/tmp/base-a",)) if t.name == "read_file")
|
||||
with pytest.raises(explore.ExplorationError) as excinfo:
|
||||
read_file.func(bundle_id="base-b", path="x.md")
|
||||
assert "base-a" in str(excinfo.value)
|
||||
answer = read_file.func(bundle_id="base-b", path="x.md")
|
||||
assert answer.startswith("REFUSED (") and "base-a" in answer
|
||||
|
||||
|
||||
def test_two_bases_with_the_same_name_are_refused() -> None:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue