feat(fase3): make the maker-checker checker actually gate the reasoning
Closes gap #3 (maalbilde §5): the GroupChat checker critiqued into the void — output_from=[proposer] surfaced only the proposer, so an explicit checker rejection was ignored and the deterministic validator was the sole gate. Two falsifiers now act on the same candidate: the validator gates the NUMBERS (blocking, unchanged), the checker gates the REASONING (maalbilde §2/§6). - workflow.py: output_from=agents surfaces both participants; the checker instruction ends with a VERDICT: APPROVE / VERDICT: REJECT - <reason> line. - run.py: _authored_texts() reads author_name through out.messages (MAF 1.9.0 puts it there, not on the AgentResponse); _debate_text() now selects the PROPOSER-authored output (fixes a latent texts[-1] regression that would feed the checker's verdict to generation at even round counts); _checker_verdict() parses the gate decision. An explicit REJECT overrides an otherwise-validated outcome to a checker-sourced Rejection. Opt-in-reject (fail-open on a missing marker). RunResult gains checker_verdict; provenance.validator_decision is stamped from the validator outcome BEFORE the override, so it never conflates the two falsifiers (provenance honesty). Load-bearing (maalbilde §7): tests/test_checker_gate_loadbearing.py is a PAIR — an explicit checker REJECT on a VALIDATOR-VALID proposal yields a Rejection whose reason carries the checker's reason while validator_decision stays "validated"; the causality control (checker APPROVE, same proposer) validates normally. Proven RED on BOTH detach points (revert output_from, or drop the override). Suite 134->136 passed, 4 skipped; mypy + ruff check clean. Pre-existing ruff-format drift (backends/budget/verdicts/test_contracts) left untouched for a surgical diff. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MHR8iKxJRxDiDfNw8HZmWE
This commit is contained in:
parent
8814a698c2
commit
4ec778c855
5 changed files with 194 additions and 32 deletions
|
|
@ -26,7 +26,12 @@ from agent_framework.orchestrations import GroupChatBuilder
|
|||
_MAKER_CHECKER_ROLES = ("proposer", "checker")
|
||||
_INSTRUCTIONS = {
|
||||
"proposer": "You propose one concrete cost-saving measure for the project.",
|
||||
"checker": "You critique the proposal and flag any constraint violation.",
|
||||
"checker": (
|
||||
"You critique the proposal's reasoning and flag any constraint violation. "
|
||||
"End your reply with exactly one verdict line: 'VERDICT: APPROVE' if the reasoning "
|
||||
"holds, or 'VERDICT: REJECT - <short reason>' if it does not. The validator gates the "
|
||||
"numbers; your verdict gates the reasoning (Step 3/4 — målbilde §2)."
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
|
|
@ -79,10 +84,11 @@ def fresh_workflow(
|
|||
is called once per role, so each run owns its own clients — no state survives between runs.
|
||||
|
||||
``tools`` + ``middleware`` are attached to each agent (F2/F7; constructed by the
|
||||
orchestrator). ``output_from=[proposer]`` makes ``WorkflowRunResult.get_outputs()`` surface
|
||||
the proposer's converged output — without it, ``get_outputs()`` yields only the
|
||||
orchestrator's "reached max rounds" notice, so the F1 debate->generation dataflow could not
|
||||
read the debate result (verified against installed 1.9.0).
|
||||
orchestrator). ``output_from=agents`` makes ``WorkflowRunResult.get_outputs()`` surface BOTH
|
||||
participants' converged outputs — the proposer's (fed to generation, F1) AND the checker's
|
||||
gate verdict (Step 3/4); ``run.py`` separates them by per-message ``author_name``. Without it,
|
||||
``get_outputs()`` yields only the orchestrator's "reached max rounds" notice (verified against
|
||||
installed 1.9.0).
|
||||
"""
|
||||
agents = maker_checker_agents(client_factory, tools=tools, middleware=middleware)
|
||||
# Agents are built from _MAKER_CHECKER_ROLES in order with name=role, so the role tuple
|
||||
|
|
@ -100,9 +106,10 @@ def fresh_workflow(
|
|||
selection_func=select,
|
||||
# Safety net well above the hard cap; with_max_rounds is the binding bound (B4).
|
||||
termination_condition=make_termination(max_rounds * len(names) + 1),
|
||||
# Surface the PROPOSER's converged output so get_outputs() carries the debate
|
||||
# result (F1); the default surfaces only the orchestrator's termination notice.
|
||||
output_from=[agents[0]],
|
||||
# Surface BOTH participants so get_outputs() carries the proposer's converged output
|
||||
# (fed to generation, F1) AND the checker's gate verdict (Step 3/4); run.py separates
|
||||
# them by author_name. The default surfaces only the orchestrator's termination notice.
|
||||
output_from=agents,
|
||||
).with_max_rounds(max_rounds)
|
||||
if enable_layer1_hitl:
|
||||
# Layer-1: in-run synchronous review on the checker (no checkpoint — research 01).
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue