feat(step5): the falsification that informed the next hypothesis now leaves the loop
generate_via_llm consumed each validator Rejection internally (`last`), fed it into the next attempt's prompt, and dropped it. So Step 5 was real but unobservable: a caller could see THAT a proposal validated, never that it validated on attempt 2 after the deterministic validator falsified attempt 1. It was the one step of the eight with no output to show. The seam is a typed return value -- GenerationResult(outcome, refinements) -- rather than an out-parameter or a callback: a returned value cannot be silently lost by a caller that forgets to pass a collector, and mypy forces every call site to acknowledge it. refinements carries ONLY rejections that were actually fed back. When the attempt budget runs out the final rejection IS outcome; counting it here would be double-counting, and the bounded control test goes red on the collect-everything implementation that gets this wrong. The loop's bound is untouched: max_attempts and meter.tick_round stand, and `last` still drives the prompt alone, so prompt growth is unchanged. run.py accumulates across _evaluate calls, so _evaluate_mandate is untouched; RunResult.refinements defaults (the coverage precedent) and is concatenated across approaches rather than keyed per approach -- stated as an honesty limit. The simulation now shows it: the scripted proposer overclaims 250000, which the validator falsifies against P90 = 90000, and the corrected 30000 validates. Only the overclaim is scripted -- the rejection is computed. scripted_factory takes a per-role reply selector so this needs no second scripted client body. README records the two accuracy changes only (Step 5 is now inspectable; the simulation trace shows the correction). The level-2 publishing claim stays deferred until after the demo (O4). Load-bearing MEASURED against the full suite with a control, four mutations all red: detach the returned history (4 tests) - collect-everything (control only) - detach the run wiring (2 tests) - revert the simulation's proposer to a constant (the demo-protection test). Control: 759 passed / 4 skipped; ruff, format and mypy clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CcWFcREUi6YPjEpN3ACDP
This commit is contained in:
parent
cd011c4ac7
commit
d6f3359fae
11 changed files with 381 additions and 26 deletions
|
|
@ -34,14 +34,14 @@ def _meter() -> TokenMeter:
|
|||
async def test_wellformed_reply_yields_validated_proposal(project) -> None:
|
||||
client = FakeChatClient(scripted=[_VALID], default_reply=_VALID)
|
||||
result = await generate_via_llm(client, project, "", _meter(), max_attempts=3)
|
||||
assert isinstance(result, ValidatedProposal)
|
||||
assert isinstance(result.proposal, SavingsProposal)
|
||||
assert isinstance(result.outcome, ValidatedProposal)
|
||||
assert isinstance(result.outcome.proposal, SavingsProposal)
|
||||
|
||||
|
||||
async def test_malformed_reply_is_retried_not_silently_accepted(project) -> None:
|
||||
client = FakeChatClient(scripted=["not json at all {{{", _VALID], default_reply=_VALID)
|
||||
result = await generate_via_llm(client, project, "", _meter(), max_attempts=3)
|
||||
assert isinstance(result, ValidatedProposal) # the malformed reply was NOT accepted
|
||||
assert isinstance(result.outcome, ValidatedProposal) # the malformed reply was NOT accepted
|
||||
assert client.call_count >= 2 # it retried past the malformed reply
|
||||
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue