feat(step5): the falsification that informed the next hypothesis now leaves the loop
generate_via_llm consumed each validator Rejection internally (`last`), fed it into the next attempt's prompt, and dropped it. So Step 5 was real but unobservable: a caller could see THAT a proposal validated, never that it validated on attempt 2 after the deterministic validator falsified attempt 1. It was the one step of the eight with no output to show. The seam is a typed return value -- GenerationResult(outcome, refinements) -- rather than an out-parameter or a callback: a returned value cannot be silently lost by a caller that forgets to pass a collector, and mypy forces every call site to acknowledge it. refinements carries ONLY rejections that were actually fed back. When the attempt budget runs out the final rejection IS outcome; counting it here would be double-counting, and the bounded control test goes red on the collect-everything implementation that gets this wrong. The loop's bound is untouched: max_attempts and meter.tick_round stand, and `last` still drives the prompt alone, so prompt growth is unchanged. run.py accumulates across _evaluate calls, so _evaluate_mandate is untouched; RunResult.refinements defaults (the coverage precedent) and is concatenated across approaches rather than keyed per approach -- stated as an honesty limit. The simulation now shows it: the scripted proposer overclaims 250000, which the validator falsifies against P90 = 90000, and the corrected 30000 validates. Only the overclaim is scripted -- the rejection is computed. scripted_factory takes a per-role reply selector so this needs no second scripted client body. README records the two accuracy changes only (Step 5 is now inspectable; the simulation trace shows the correction). The level-2 publishing claim stays deferred until after the demo (O4). Load-bearing MEASURED against the full suite with a control, four mutations all red: detach the returned history (4 tests) - collect-everything (control only) - detach the run wiring (2 tests) - revert the simulation's proposer to a constant (the demo-protection test). Control: 759 passed / 4 skipped; ruff, format and mypy clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CcWFcREUi6YPjEpN3ACDP
This commit is contained in:
parent
cd011c4ac7
commit
d6f3359fae
11 changed files with 381 additions and 26 deletions
|
|
@ -85,7 +85,7 @@ async def test_commissioned_approach_gets_no_validator_discount(project) -> None
|
|||
result = await generate_via_llm(
|
||||
client, project, "", _meter(), max_attempts=1, approach=_APPROACH
|
||||
)
|
||||
assert isinstance(result, Rejection)
|
||||
assert isinstance(result.outcome, Rejection)
|
||||
|
||||
|
||||
async def test_commissioned_approach_still_validates_when_the_numbers_hold(project) -> None:
|
||||
|
|
@ -96,5 +96,5 @@ async def test_commissioned_approach_still_validates_when_the_numbers_hold(proje
|
|||
result = await generate_via_llm(
|
||||
client, project, "", _meter(), max_attempts=1, approach=_APPROACH
|
||||
)
|
||||
assert isinstance(result, ValidatedProposal)
|
||||
assert isinstance(result.outcome, ValidatedProposal)
|
||||
assert _APPROACH.label in client.received_texts[0][0] # it went through the commissioned prompt
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue