feat(step5): the falsification that informed the next hypothesis now leaves the loop
generate_via_llm consumed each validator Rejection internally (`last`), fed it into the next attempt's prompt, and dropped it. So Step 5 was real but unobservable: a caller could see THAT a proposal validated, never that it validated on attempt 2 after the deterministic validator falsified attempt 1. It was the one step of the eight with no output to show. The seam is a typed return value -- GenerationResult(outcome, refinements) -- rather than an out-parameter or a callback: a returned value cannot be silently lost by a caller that forgets to pass a collector, and mypy forces every call site to acknowledge it. refinements carries ONLY rejections that were actually fed back. When the attempt budget runs out the final rejection IS outcome; counting it here would be double-counting, and the bounded control test goes red on the collect-everything implementation that gets this wrong. The loop's bound is untouched: max_attempts and meter.tick_round stand, and `last` still drives the prompt alone, so prompt growth is unchanged. run.py accumulates across _evaluate calls, so _evaluate_mandate is untouched; RunResult.refinements defaults (the coverage precedent) and is concatenated across approaches rather than keyed per approach -- stated as an honesty limit. The simulation now shows it: the scripted proposer overclaims 250000, which the validator falsifies against P90 = 90000, and the corrected 30000 validates. Only the overclaim is scripted -- the rejection is computed. scripted_factory takes a per-role reply selector so this needs no second scripted client body. README records the two accuracy changes only (Step 5 is now inspectable; the simulation trace shows the correction). The level-2 publishing claim stays deferred until after the demo (O4). Load-bearing MEASURED against the full suite with a control, four mutations all red: detach the returned history (4 tests) - collect-everything (control only) - detach the run wiring (2 tests) - revert the simulation's proposer to a constant (the demo-protection test). Control: 759 passed / 4 skipped; ruff, format and mypy clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CcWFcREUi6YPjEpN3ACDP
This commit is contained in:
parent
cd011c4ac7
commit
d6f3359fae
11 changed files with 381 additions and 26 deletions
|
|
@ -131,6 +131,15 @@ class RunResult:
|
|||
#: report is honest there, because nothing was ordered. It defaults so every existing
|
||||
#: constructor call and every frozen aggregate over ``RunResult`` is unaffected.
|
||||
coverage: tuple[ApproachOutcome, ...] = ()
|
||||
#: Step 5 (målbilde §5/§7): the validator falsifications that informed a LATER generation
|
||||
#: attempt, in attempt order — what ``generate_via_llm`` corrected in response to, rather than
|
||||
#: only what it ended up with. EMPTY on the common path where the first candidate validates:
|
||||
#: nothing was falsified, so there is nothing to show. Honesty limit: with a mandate this is
|
||||
#: the run's refinements CONCATENATED across every commissioned approach, not keyed per
|
||||
#: approach — ``coverage`` is the per-approach report, and hanging proposals off its rows is
|
||||
#: what ``_evaluate_mandate`` deliberately avoids. It defaults, so every existing constructor
|
||||
#: call is unaffected (mirrors ``coverage``).
|
||||
refinements: tuple[Rejection, ...] = ()
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
|
|
@ -635,10 +644,18 @@ async def run_project(
|
|||
# each under the SAME meter — no new unbounded loop; the caps already in force are the bound.
|
||||
proposer_client = factory("proposer")
|
||||
|
||||
# Step 5 (målbilde §5/§7): generation now returns its falsification history alongside the
|
||||
# outcome. ``_evaluate`` keeps its ``ValidatedProposal | Rejection`` shape so ``_evaluate_mandate``
|
||||
# is untouched, and the history is accumulated here in call order — one entry per approach that
|
||||
# needed correcting, concatenated (see ``RunResult.refinements`` for that honesty limit).
|
||||
refinements: list[Rejection] = []
|
||||
|
||||
async def _evaluate(approach: Approach | None) -> ValidatedProposal | Rejection:
|
||||
return await generate_via_llm(
|
||||
generated = await generate_via_llm(
|
||||
proposer_client, project, gen_context, meter, baseline=baseline, approach=approach
|
||||
)
|
||||
refinements.extend(generated.refinements)
|
||||
return generated.outcome
|
||||
|
||||
coverage: tuple[ApproachOutcome, ...] = ()
|
||||
evaluated: tuple[tuple[str, ValidatedProposal | Rejection], ...] = ()
|
||||
|
|
@ -774,6 +791,7 @@ async def run_project(
|
|||
debate_output=debate_output,
|
||||
checker_verdict=checker_decision,
|
||||
coverage=coverage,
|
||||
refinements=tuple(refinements),
|
||||
)
|
||||
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue