feat(step5): the falsification that informed the next hypothesis now leaves the loop
generate_via_llm consumed each validator Rejection internally (`last`), fed it into the next attempt's prompt, and dropped it. So Step 5 was real but unobservable: a caller could see THAT a proposal validated, never that it validated on attempt 2 after the deterministic validator falsified attempt 1. It was the one step of the eight with no output to show. The seam is a typed return value -- GenerationResult(outcome, refinements) -- rather than an out-parameter or a callback: a returned value cannot be silently lost by a caller that forgets to pass a collector, and mypy forces every call site to acknowledge it. refinements carries ONLY rejections that were actually fed back. When the attempt budget runs out the final rejection IS outcome; counting it here would be double-counting, and the bounded control test goes red on the collect-everything implementation that gets this wrong. The loop's bound is untouched: max_attempts and meter.tick_round stand, and `last` still drives the prompt alone, so prompt growth is unchanged. run.py accumulates across _evaluate calls, so _evaluate_mandate is untouched; RunResult.refinements defaults (the coverage precedent) and is concatenated across approaches rather than keyed per approach -- stated as an honesty limit. The simulation now shows it: the scripted proposer overclaims 250000, which the validator falsifies against P90 = 90000, and the corrected 30000 validates. Only the overclaim is scripted -- the rejection is computed. scripted_factory takes a per-role reply selector so this needs no second scripted client body. README records the two accuracy changes only (Step 5 is now inspectable; the simulation trace shows the correction). The level-2 publishing claim stays deferred until after the demo (O4). Load-bearing MEASURED against the full suite with a control, four mutations all red: detach the returned history (4 tests) - collect-everything (control only) - detach the run wiring (2 tests) - revert the simulation's proposer to a constant (the demo-protection test). Control: 759 passed / 4 skipped; ruff, format and mypy clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CcWFcREUi6YPjEpN3ACDP
This commit is contained in:
parent
cd011c4ac7
commit
d6f3359fae
11 changed files with 381 additions and 26 deletions
|
|
@ -111,7 +111,7 @@ async def test_refinement_feeds_prior_falsification_into_next_prompt() -> None:
|
|||
# context="" so the flip token cannot pre-exist in attempt 1's prompt.
|
||||
result = await generate_via_llm(client, project, "", _meter(), max_attempts=3)
|
||||
|
||||
assert isinstance(result, ValidatedProposal), (
|
||||
assert isinstance(result.outcome, ValidatedProposal), (
|
||||
"the proposer corrected on attempt 2 but the loop did not validate -- the falsification "
|
||||
"never reached the next prompt"
|
||||
)
|
||||
|
|
@ -134,5 +134,5 @@ async def test_refinement_loop_stays_bounded_when_never_fixed() -> None:
|
|||
|
||||
result = await generate_via_llm(client, project, "", _meter(), max_attempts=3)
|
||||
|
||||
assert isinstance(result, Rejection)
|
||||
assert isinstance(result.outcome, Rejection)
|
||||
assert client.call_count == 3
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue