feat(major2): the reviewer sits after the validator in the attempt loop; the outcome is the validator's last ruling [skip-docs]
Ordre 20260904T173146Z-8102814273-from-portfolio-optimiser, steg 3 av 10. generate_via_llm faar fire keyword-only parametre: reviewer, reviews (KALLER-EID sink), review_key og checker_verdict. Revieweren kalles synkront i det validate_proposal AKSEPTERER en kandidat - aldri paa en avvist (den er alt matet tilbake informert; aa be et menneske kommentere tall maskinen nettopp gjendrev bruker mennesket paa maskinens jobb). EXIT-KONTRAKTEN ER ENDRET. Foer dette hvilte utgangen paa `last`, som KUN en validator- avvisning setter - saa "validert -> revise" paa hvert forsoek naadde slutten av loekka med `last is None` og doede paa `assert last is not None` (og under -O paa None.proposal). `last_ruling` er naa en eksplisitt baerer, og D6 leses rett av den. Sinken er kaller-eid av parse_failures' MAALTE grunn, ett hakk skarpere: meter.tick_round reiser inne i _fetch_parsed paa forsoeket en revise kjoepte, saa paa noeyaktig den kjoeringen posten betyr mest returnerer funksjonen INGENTING. Et felt paa GenerationResult ville vaert blindt for det. attempts_remaining = min(max_attempts - i - 1, meter.budget.max_rounds - meter.rounds) - LEDGER-BEVISST, fordi rundeboka deles av hver approach i et mandat og ofte er det som binder. honoured betyr at forsoeket revisen kjoepte FAKTISK HENTET et svar, ikke at det ble kjoept: posten settes False og forfremmes foerst naar _fetch_parsed har returnert. RODT foer impl: 11 armer. TO AV PLANENS EGNE TALL BLE FALSIFISERT AV MAALINGEN og staar korrigert i testen: (1) planens T3 (max_rounds=2, max_attempts=10 => BudgetExceeded observed=3) er ikke naabar under den ledger-bevisste remaining fra planens egen revisjon 5 - loekka stopper etter to hentinger UTEN unntak; armen er delt i T3a (ledgeren stopper revisjonene, M38s diskriminator) og T3b (honoured=False naar den kjoepte hentingen aldri returnerte, M40s vitne, drevet via parse-retryen som gir noeyaktig rounds/2/3). (2) planens T-ledger sier attempts_remaining == 0 ved max_rounds=2/max_attempts=3; maalt er det 1 - armen bruker max_rounds=1, der ledgeren faktisk binder. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
d4c8691326
commit
bbf4d3b5b6
2 changed files with 495 additions and 12 deletions
|
|
@ -23,7 +23,7 @@ from __future__ import annotations
|
|||
|
||||
import json
|
||||
from collections.abc import Callable, Mapping
|
||||
from dataclasses import dataclass, field
|
||||
from dataclasses import dataclass, field, replace
|
||||
from typing import Any
|
||||
|
||||
from agent_framework import BaseChatClient, Message
|
||||
|
|
@ -32,6 +32,11 @@ from pydantic import BaseModel, ValidationError
|
|||
from portfolio_optimiser.budget import TokenMeter
|
||||
from portfolio_optimiser.ir import CostBaseline, SavingsProposal
|
||||
from portfolio_optimiser.mandate import Approach
|
||||
from portfolio_optimiser.proposal_review import (
|
||||
ProposalReview,
|
||||
ProposalReviewer,
|
||||
ProposalReviewRequest,
|
||||
)
|
||||
from portfolio_optimiser.reference_domain import Project
|
||||
from portfolio_optimiser.validator import (
|
||||
Rejection,
|
||||
|
|
@ -421,6 +426,10 @@ async def generate_via_llm(
|
|||
baseline: CostBaseline | None = None,
|
||||
approach: Approach | None = None,
|
||||
parse_failures: list[ParseFailure] | None = None,
|
||||
reviewer: ProposalReviewer | None = None,
|
||||
reviews: list[ProposalReview] | None = None,
|
||||
review_key: tuple[str | None, str | None] = (None, None),
|
||||
checker_verdict: str = "absent",
|
||||
) -> GenerationResult:
|
||||
"""Async LLM path: non-streaming chat -> parse -> validate, with TWO bounded retry kinds,
|
||||
the meter checked in this loop:
|
||||
|
|
@ -460,6 +469,34 @@ async def generate_via_llm(
|
|||
REACHES the caller; here it does not, so the rule is cited and departed from deliberately. That
|
||||
a caller can forget is answered by a test on the wiring, not by a shape that cannot work.
|
||||
|
||||
``reviewer`` (MAJOR-2) is the THIRD falsifier — a human, and the only one that gates nothing.
|
||||
It is called synchronously the moment ``validate_proposal`` ACCEPTS a candidate, never on a
|
||||
rejected one (a rejection is already fed back informed above; asking a person to comment on
|
||||
numbers the machine just refuted spends the person on the machine's job). It answers
|
||||
``approve`` — take it as it stands — or ``revise(feedback)``, which buys exactly ONE more
|
||||
attempt out of the budget this loop already has: no new loop, no second cap, and the words go
|
||||
into the next prompt through ``prior_feedback`` (verbatim, never the proposal JSON).
|
||||
|
||||
**D6: the outcome is the validator's LAST ruling, and the reviewer never selects among
|
||||
attempts.** An honoured revise whose follow-up the validator then rejects, with nothing left
|
||||
to buy, ends in that ``Rejection`` — exactly as an exhausted loop always has. Falling back to
|
||||
the earlier validated proposal would hand the run the very candidate the expert asked to
|
||||
change. The candidate they were LOOKING at is not lost: it travels in the review record.
|
||||
|
||||
``reviews`` is a CALLER-OWNED sink for the same measured reason ``parse_failures`` is one, one
|
||||
level sharper: ``meter.tick_round`` raises inside ``_fetch_parsed`` on the attempt a revise
|
||||
bought, so on exactly the run whose record matters most this function returns NOTHING. A field
|
||||
on ``GenerationResult`` would be blind to it. ``review_key`` is the caller's ``(approach_id,
|
||||
approach_label)`` — the loop records what the caller keyed, so the artefact can say WHICH
|
||||
candidate a human answered about; a ``None`` label falls back to the project id so a direct
|
||||
library caller never sees an empty header. ``checker_verdict`` (D3) is the run-level reasoning
|
||||
gate's answer, passed down READ-ONLY: it informs the expert, and never enters the record.
|
||||
|
||||
**Honesty limit inherited by Step 5:** a revise consumes one of the same ``max_attempts``
|
||||
iterations a validator rejection would, so a run can spend its attempts on expert revisions
|
||||
and never reach a second validator falsification — ``refinements`` then under-reports by
|
||||
construction on that run. Stated, not repaired.
|
||||
|
||||
Returns a ``GenerationResult``: the ``ValidatedProposal | Rejection`` outcome plus every
|
||||
rejection that was fed back into a later attempt's prompt. Surfacing that history changes
|
||||
nothing about the loop's BOUND — ``max_attempts`` and ``meter.tick_round`` are exactly as
|
||||
|
|
@ -497,20 +534,95 @@ async def generate_via_llm(
|
|||
# prompt growth is unchanged. This list is a record for the CALLER, appended to only once a
|
||||
# rejection is about to inform a further attempt; it is never read back into a prompt.
|
||||
fed_back: list[Rejection] = []
|
||||
for _ in range(max_attempts):
|
||||
# The most recent ruling of EITHER kind. Before MAJOR-2 the exit rested on ``last``, which only
|
||||
# a validator REJECTION sets -- so "validated -> revise" on every attempt reached the end of the
|
||||
# loop with ``last is None`` and died on ``assert last is not None`` (and, under -O, on
|
||||
# ``None.proposal``). The carrier is explicit, and D6 reads straight off it: whatever the
|
||||
# validator ruled LAST is what the run carries, never a proposal the reviewer picked.
|
||||
last_ruling: ValidatedProposal | Rejection | None = None
|
||||
# The expert's STANDING instruction. Unlike ``last`` it is not cleared per attempt: it holds
|
||||
# until the reviewer next answers, because a human's request survives one machine round trip.
|
||||
feedback: str | None = None
|
||||
# Index into ``reviews`` of a revise whose bought attempt has not fetched yet. ``honoured``
|
||||
# means the follow-up actually FETCHED, not that it was bought -- so it is promoted only once
|
||||
# ``_fetch_parsed`` returns, and a revise the ledger cut before then stays false.
|
||||
pending_revise: int | None = None
|
||||
approach_id, approach_label = review_key
|
||||
for i in range(max_attempts):
|
||||
# Informed refinement: feed the PREVIOUS attempt's validator rejection into this
|
||||
# attempt's prompt. ``last`` is None on attempt 1 -> the unchanged base prompt; it is
|
||||
# overwritten each round -> only the most-recent falsification ("forrige"), never an
|
||||
# accumulated history (bounded prompt growth).
|
||||
if last is not None:
|
||||
fed_back.append(last)
|
||||
messages = _build_messages(project, context, prior_rejection=last, approach=approach)
|
||||
messages = _build_messages(
|
||||
project,
|
||||
context,
|
||||
prior_rejection=last,
|
||||
approach=approach,
|
||||
prior_feedback=feedback,
|
||||
)
|
||||
candidate = await _fetch_parsed(messages)
|
||||
if pending_revise is not None and reviews is not None:
|
||||
reviews[pending_revise] = replace(reviews[pending_revise], honoured=True)
|
||||
pending_revise = None
|
||||
result = validate_proposal(candidate, baseline=baseline)
|
||||
if isinstance(result, ValidatedProposal):
|
||||
last_ruling = result
|
||||
if isinstance(result, Rejection):
|
||||
last = result
|
||||
continue
|
||||
if reviewer is None:
|
||||
return GenerationResult(outcome=result, refinements=tuple(fed_back))
|
||||
last = result
|
||||
assert last is not None # max_attempts >= 1, so at least one validation ran
|
||||
# Validation never passed within the attempt budget -> typed Rejection. ``last`` is the outcome
|
||||
# and was never fed back, so it is deliberately absent from ``refinements``.
|
||||
return GenerationResult(outcome=last, refinements=tuple(fed_back))
|
||||
# What a ``revise`` can ACTUALLY buy: the smaller of this call's own attempt headroom and
|
||||
# the SHARED round ledger's. The ledger (``max_rounds*4`` at run level) spans every approach
|
||||
# of a mandate and is frequently what binds; a number computed from ``max_attempts`` alone
|
||||
# would be a claim the terminal makes about itself that the meter then refutes.
|
||||
remaining = min(max_attempts - i - 1, meter.budget.max_rounds - meter.rounds)
|
||||
decision = reviewer(
|
||||
ProposalReviewRequest(
|
||||
project_id=project.id,
|
||||
approach_id=approach_id,
|
||||
approach_label=approach_label or project.id,
|
||||
attempt=i,
|
||||
attempts_remaining=remaining,
|
||||
proposal=result,
|
||||
checker_verdict=checker_verdict,
|
||||
)
|
||||
)
|
||||
if decision.feedback is None:
|
||||
if reviews is not None:
|
||||
reviews.append(
|
||||
ProposalReview(
|
||||
approach_id=approach_id,
|
||||
attempt=i,
|
||||
decision="approve",
|
||||
feedback="",
|
||||
honoured=True,
|
||||
proposal=result,
|
||||
)
|
||||
)
|
||||
return GenerationResult(outcome=result, refinements=tuple(fed_back))
|
||||
if reviews is not None:
|
||||
reviews.append(
|
||||
ProposalReview(
|
||||
approach_id=approach_id,
|
||||
attempt=i,
|
||||
decision="revise",
|
||||
feedback=decision.feedback,
|
||||
honoured=False,
|
||||
proposal=result,
|
||||
)
|
||||
)
|
||||
if remaining <= 0:
|
||||
# An un-honoured revise buys nothing, so the last ruling -- the validated one -- stands
|
||||
# (criterion 12). One rule with D6, both cases.
|
||||
return GenerationResult(outcome=result, refinements=tuple(fed_back))
|
||||
pending_revise = None if reviews is None else len(reviews) - 1
|
||||
feedback = decision.feedback
|
||||
# The machine's reason does not apply to a proposal the machine ACCEPTED.
|
||||
last = None
|
||||
if last_ruling is None: # pragma: no cover - every call site passes max_attempts >= 1
|
||||
raise ValueError(f"max_attempts must be positive, got {max_attempts}")
|
||||
# The validator's LAST ruling, whichever kind it was. When it is a ``Rejection`` it was never
|
||||
# fed back, so it is deliberately absent from ``refinements``.
|
||||
return GenerationResult(outcome=last_ruling, refinements=tuple(fed_back))
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue