feat(major2): the reviewer sits after the validator in the attempt loop; the outcome is the validator's last ruling [skip-docs]

Ordre 20260904T173146Z-8102814273-from-portfolio-optimiser, steg 3 av 10.

generate_via_llm faar fire keyword-only parametre: reviewer, reviews (KALLER-EID sink),
review_key og checker_verdict. Revieweren kalles synkront i det validate_proposal
AKSEPTERER en kandidat - aldri paa en avvist (den er alt matet tilbake informert; aa be
et menneske kommentere tall maskinen nettopp gjendrev bruker mennesket paa maskinens jobb).

EXIT-KONTRAKTEN ER ENDRET. Foer dette hvilte utgangen paa `last`, som KUN en validator-
avvisning setter - saa "validert -> revise" paa hvert forsoek naadde slutten av loekka med
`last is None` og doede paa `assert last is not None` (og under -O paa None.proposal).
`last_ruling` er naa en eksplisitt baerer, og D6 leses rett av den.

Sinken er kaller-eid av parse_failures' MAALTE grunn, ett hakk skarpere: meter.tick_round
reiser inne i _fetch_parsed paa forsoeket en revise kjoepte, saa paa noeyaktig den kjoeringen
posten betyr mest returnerer funksjonen INGENTING. Et felt paa GenerationResult ville vaert
blindt for det.

attempts_remaining = min(max_attempts - i - 1, meter.budget.max_rounds - meter.rounds) -
LEDGER-BEVISST, fordi rundeboka deles av hver approach i et mandat og ofte er det som binder.

honoured betyr at forsoeket revisen kjoepte FAKTISK HENTET et svar, ikke at det ble kjoept:
posten settes False og forfremmes foerst naar _fetch_parsed har returnert.

RODT foer impl: 11 armer. TO AV PLANENS EGNE TALL BLE FALSIFISERT AV MAALINGEN og staar
korrigert i testen: (1) planens T3 (max_rounds=2, max_attempts=10 => BudgetExceeded
observed=3) er ikke naabar under den ledger-bevisste remaining fra planens egen revisjon 5 -
loekka stopper etter to hentinger UTEN unntak; armen er delt i T3a (ledgeren stopper
revisjonene, M38s diskriminator) og T3b (honoured=False naar den kjoepte hentingen aldri
returnerte, M40s vitne, drevet via parse-retryen som gir noeyaktig rounds/2/3). (2) planens
T-ledger sier attempts_remaining == 0 ved max_rounds=2/max_attempts=3; maalt er det 1 -
armen bruker max_rounds=1, der ledgeren faktisk binder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-05 06:52:16 +02:00
commit bbf4d3b5b6
2 changed files with 495 additions and 12 deletions

View file

@ -23,7 +23,7 @@ from __future__ import annotations
import json
from collections.abc import Callable, Mapping
from dataclasses import dataclass, field
from dataclasses import dataclass, field, replace
from typing import Any
from agent_framework import BaseChatClient, Message
@ -32,6 +32,11 @@ from pydantic import BaseModel, ValidationError
from portfolio_optimiser.budget import TokenMeter
from portfolio_optimiser.ir import CostBaseline, SavingsProposal
from portfolio_optimiser.mandate import Approach
from portfolio_optimiser.proposal_review import (
ProposalReview,
ProposalReviewer,
ProposalReviewRequest,
)
from portfolio_optimiser.reference_domain import Project
from portfolio_optimiser.validator import (
Rejection,
@ -421,6 +426,10 @@ async def generate_via_llm(
baseline: CostBaseline | None = None,
approach: Approach | None = None,
parse_failures: list[ParseFailure] | None = None,
reviewer: ProposalReviewer | None = None,
reviews: list[ProposalReview] | None = None,
review_key: tuple[str | None, str | None] = (None, None),
checker_verdict: str = "absent",
) -> GenerationResult:
"""Async LLM path: non-streaming chat -> parse -> validate, with TWO bounded retry kinds,
the meter checked in this loop:
@ -460,6 +469,34 @@ async def generate_via_llm(
REACHES the caller; here it does not, so the rule is cited and departed from deliberately. That
a caller can forget is answered by a test on the wiring, not by a shape that cannot work.
``reviewer`` (MAJOR-2) is the THIRD falsifier a human, and the only one that gates nothing.
It is called synchronously the moment ``validate_proposal`` ACCEPTS a candidate, never on a
rejected one (a rejection is already fed back informed above; asking a person to comment on
numbers the machine just refuted spends the person on the machine's job). It answers
``approve`` take it as it stands or ``revise(feedback)``, which buys exactly ONE more
attempt out of the budget this loop already has: no new loop, no second cap, and the words go
into the next prompt through ``prior_feedback`` (verbatim, never the proposal JSON).
**D6: the outcome is the validator's LAST ruling, and the reviewer never selects among
attempts.** An honoured revise whose follow-up the validator then rejects, with nothing left
to buy, ends in that ``Rejection`` exactly as an exhausted loop always has. Falling back to
the earlier validated proposal would hand the run the very candidate the expert asked to
change. The candidate they were LOOKING at is not lost: it travels in the review record.
``reviews`` is a CALLER-OWNED sink for the same measured reason ``parse_failures`` is one, one
level sharper: ``meter.tick_round`` raises inside ``_fetch_parsed`` on the attempt a revise
bought, so on exactly the run whose record matters most this function returns NOTHING. A field
on ``GenerationResult`` would be blind to it. ``review_key`` is the caller's ``(approach_id,
approach_label)`` the loop records what the caller keyed, so the artefact can say WHICH
candidate a human answered about; a ``None`` label falls back to the project id so a direct
library caller never sees an empty header. ``checker_verdict`` (D3) is the run-level reasoning
gate's answer, passed down READ-ONLY: it informs the expert, and never enters the record.
**Honesty limit inherited by Step 5:** a revise consumes one of the same ``max_attempts``
iterations a validator rejection would, so a run can spend its attempts on expert revisions
and never reach a second validator falsification ``refinements`` then under-reports by
construction on that run. Stated, not repaired.
Returns a ``GenerationResult``: the ``ValidatedProposal | Rejection`` outcome plus every
rejection that was fed back into a later attempt's prompt. Surfacing that history changes
nothing about the loop's BOUND — ``max_attempts`` and ``meter.tick_round`` are exactly as
@ -497,20 +534,95 @@ async def generate_via_llm(
# prompt growth is unchanged. This list is a record for the CALLER, appended to only once a
# rejection is about to inform a further attempt; it is never read back into a prompt.
fed_back: list[Rejection] = []
for _ in range(max_attempts):
# The most recent ruling of EITHER kind. Before MAJOR-2 the exit rested on ``last``, which only
# a validator REJECTION sets -- so "validated -> revise" on every attempt reached the end of the
# loop with ``last is None`` and died on ``assert last is not None`` (and, under -O, on
# ``None.proposal``). The carrier is explicit, and D6 reads straight off it: whatever the
# validator ruled LAST is what the run carries, never a proposal the reviewer picked.
last_ruling: ValidatedProposal | Rejection | None = None
# The expert's STANDING instruction. Unlike ``last`` it is not cleared per attempt: it holds
# until the reviewer next answers, because a human's request survives one machine round trip.
feedback: str | None = None
# Index into ``reviews`` of a revise whose bought attempt has not fetched yet. ``honoured``
# means the follow-up actually FETCHED, not that it was bought -- so it is promoted only once
# ``_fetch_parsed`` returns, and a revise the ledger cut before then stays false.
pending_revise: int | None = None
approach_id, approach_label = review_key
for i in range(max_attempts):
# Informed refinement: feed the PREVIOUS attempt's validator rejection into this
# attempt's prompt. ``last`` is None on attempt 1 -> the unchanged base prompt; it is
# overwritten each round -> only the most-recent falsification ("forrige"), never an
# accumulated history (bounded prompt growth).
if last is not None:
fed_back.append(last)
messages = _build_messages(project, context, prior_rejection=last, approach=approach)
messages = _build_messages(
project,
context,
prior_rejection=last,
approach=approach,
prior_feedback=feedback,
)
candidate = await _fetch_parsed(messages)
if pending_revise is not None and reviews is not None:
reviews[pending_revise] = replace(reviews[pending_revise], honoured=True)
pending_revise = None
result = validate_proposal(candidate, baseline=baseline)
if isinstance(result, ValidatedProposal):
last_ruling = result
if isinstance(result, Rejection):
last = result
continue
if reviewer is None:
return GenerationResult(outcome=result, refinements=tuple(fed_back))
last = result
assert last is not None # max_attempts >= 1, so at least one validation ran
# Validation never passed within the attempt budget -> typed Rejection. ``last`` is the outcome
# and was never fed back, so it is deliberately absent from ``refinements``.
return GenerationResult(outcome=last, refinements=tuple(fed_back))
# What a ``revise`` can ACTUALLY buy: the smaller of this call's own attempt headroom and
# the SHARED round ledger's. The ledger (``max_rounds*4`` at run level) spans every approach
# of a mandate and is frequently what binds; a number computed from ``max_attempts`` alone
# would be a claim the terminal makes about itself that the meter then refutes.
remaining = min(max_attempts - i - 1, meter.budget.max_rounds - meter.rounds)
decision = reviewer(
ProposalReviewRequest(
project_id=project.id,
approach_id=approach_id,
approach_label=approach_label or project.id,
attempt=i,
attempts_remaining=remaining,
proposal=result,
checker_verdict=checker_verdict,
)
)
if decision.feedback is None:
if reviews is not None:
reviews.append(
ProposalReview(
approach_id=approach_id,
attempt=i,
decision="approve",
feedback="",
honoured=True,
proposal=result,
)
)
return GenerationResult(outcome=result, refinements=tuple(fed_back))
if reviews is not None:
reviews.append(
ProposalReview(
approach_id=approach_id,
attempt=i,
decision="revise",
feedback=decision.feedback,
honoured=False,
proposal=result,
)
)
if remaining <= 0:
# An un-honoured revise buys nothing, so the last ruling -- the validated one -- stands
# (criterion 12). One rule with D6, both cases.
return GenerationResult(outcome=result, refinements=tuple(fed_back))
pending_revise = None if reviews is None else len(reviews) - 1
feedback = decision.feedback
# The machine's reason does not apply to a proposal the machine ACCEPTED.
last = None
if last_ruling is None: # pragma: no cover - every call site passes max_attempts >= 1
raise ValueError(f"max_attempts must be positive, got {max_attempts}")
# The validator's LAST ruling, whichever kind it was. When it is a ``Rejection`` it was never
# fed back, so it is deliberately absent from ``refinements``.
return GenerationResult(outcome=last_ruling, refinements=tuple(fed_back))