portfolio-optimiser/tests/test_explore_callsites_loadbearing.py
Kjell Tore Guttormsen 84e8de8679 feat(explore): --plan-review gjoer "be om svar, bruke svarene" naabar fra CLI (F4, ORDRE 20260825T133139Z)
F4 fra misjonsreviewen: begge operatorflatene nektet enable_plan_review, og eneste doer var
explore(..., plan_reviewer=...) i bibliotek-APIet. Maalbildets HITL-loop var dermed unaabar for
enhver som ikke importerte pakka. Reviewens tre fil:linje-paastander ble verifisert mot kilden foer
bygging og stemte.

Flate: CLI. `--plan-review` bygger en terminal_plan_reviewer() og gir den til den UENDREDE sloeyfa.
Operatoren vises planen og svarer "approve" eller "revise <hva>"; en revisjon gaar tilbake til
manageren, som replanlegger og spoer IGJEN om den NYE planen.

Gaten er den ANDRE halvdelen av setningen. En doer som printer planen, leser linja og kaster den
bestaar "operatoren ble spurt" og feiler maalbildet -- repoets vakuoes-gate-klasse. T1 er derfor
test_explore_loadbearing sin T15 loeftet til CLI-niva og er ROED mot en alltid-godkjenn-reviewer.
Vitnet er {run_id}-exploration.json (skrevet fra en finally), ikke skrapet stdout.

Fail-closed paa operatorens egen input: alt utenfor vokabularet spoerres paa nytt, og EOF raiser
PlanReviewInputError -- stillhet er aldri en signatur.

Fire nekter ved navn, hvorav to lukket et stille dropp ingen test dekket: report_forbidden (report-
modus returnerer FOER hver utforsknings-nekt) og portefoelje-partisjonen. De to konfig-avhengige
nektene deler tokenet enable_plan_review og har derfor ulik saertekst; den eksisterende testen
asserterte paa det delte tokenet og er rettet (oekt-57-mutasjonen, niende gang).

Hosting nekter fortsatt -- reviewen er synkron og ville blokkert bade HTTP-requesten og event-loekka
som svarer /readiness -- men meldingen navngir na CLI-doeren i stedet for aa paasta at biblioteket er
den eneste.

Mid-loep-spoersmaal er IKKE bygget, og fravaeret er MAALT: _magentic.py har noeyaktig ETT
ctx.request_info (:1044, plan review) i hele modulen. Reviewens "kun plan-review FOER loepet" er
derimot upresist -- samme forespoersel fyrer ogsaa ved re-plan etter en stall.

Load-bearing MAALT (tests/test_plan_review_cli_door_loadbearing.py, 12 tester), elleve mutasjoner
alle roede mot HELE suiten + groenn kontroll 1040/5 og golden demo-transcript.stdout BYTE-UENDRET
(ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LtMDsh2zfp4Bmw8KGJ4aLD
2026-08-26 00:27:36 +02:00

766 lines
32 KiB
Python

"""U4 + U13, part 2 — the CALL SITES. Load-bearing proofs for the seams econ 56 left open.
Three things are proved here, and each of them is a seam a mutation can detach.
**1. The trace is a CALLER-OWNED accumulator, for the reason the parse-failure sink is one
(Fase 1b, funn 1).** ``explore()`` raises ``BudgetExceeded`` on its round cap, and a token cap
fires from inside the middleware mid-run — on both paths ``ExplorationResult`` never returns, so a
``ledger_log`` that existed only as a return value would be destroyed by exactly the endings § C.2
requires the artefact to be readable after ("så en stoppet utforskning er lesbar uansett hvilken
vakt som fyrte"). The accumulator the caller holds survives however the loop ended.
``ExplorationResult.ledger_log`` is BUILT FROM that accumulator rather than alongside it: two lists
holding one fact is the kø-(p) drift class, one layer up.
**2. The CLI door refuses everything it cannot honour, by name.** ``--explore`` and ``--mandate``
are two sources of ONE mandate and are REFUSED together rather than merged: ``explore()`` sets the
objective from the prompt and hardcodes ``allow_own_proposals=True``, so composing them would
silently overwrite three fields an operator wrote by hand. The refusal names the library API
(``explore(seed_approaches=…)``) because § C.6 door 1 is a real need this surface does not serve.
**3. ``{run_id}-exploration.json`` is written from a ``finally``**, so the run that most needs the
evidence — the one a cap cut short — is the one that has it.
The client is the repo's own ``ScriptedChatClient`` throughout (a bare ``BaseChatClient`` no-ops
``BudgetMiddleware``), and every tool assertion calls the tool's ``func`` DIRECTLY: measured in
econ 56, a scripted run returns TEXT and never emits a tool call, so no scripted exploration
reaches a tool body and a gate that only drove ``explore()`` would be vacuous.
"""
from __future__ import annotations
import json
from collections.abc import Callable
from pathlib import Path
from typing import Any
import pytest
from agent_framework import BaseChatClient
import portfolio_optimiser
from portfolio_optimiser import explore, hosting, okf, run, simulation
from portfolio_optimiser.budget import BudgetExceeded
from portfolio_optimiser.explore import ExplorationContract, ExplorationTrace
from portfolio_optimiser.mandate import Approach, Mandate
from portfolio_optimiser.simulation import ScriptedChatClient
_BUNDLE_DIR = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
_PID = "BYGG-KONTOR-NORD"
#: The hypothesiser's marked line. The label is the marker the end-to-end arm looks for on stdout:
#: it appears nowhere in the bundle, in the reference projects or in any other test, so its presence
#: in the settlement can only have come through the mandate the exploration shaped.
_LABEL = "SENTINEL-EXPLORE-7c1d33"
_ENERGY_REPLY = (
'{"measure":"LED-retrofit av kontorbelysning","affected_items":'
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],"claimed_saving_nok":30000}'
)
_CONTRACT_JSON: dict[str, Any] = {
"max_rounds": 4,
"max_tokens": 100_000,
"max_stall_count": 2,
"max_reset_count": 1,
"max_plan_revisions": 0,
"enable_plan_review": False,
}
@pytest.fixture(autouse=True)
def _isolate_model_env(monkeypatch: pytest.MonkeyPatch) -> None:
"""Hermetic env: no arm here may read the operator's Foundry configuration."""
monkeypatch.delenv("PORTFOLIO_MODEL_MAP", raising=False)
monkeypatch.delenv("PORTFOLIO_FOUNDRY_PROJECT_ENDPOINT", raising=False)
def _ledger_json(*, satisfied: bool, speaker: str = "hypothesiser") -> str:
return json.dumps(
{
"is_request_satisfied": {"reason": "r", "answer": satisfied},
"is_in_loop": {"reason": "r", "answer": False},
"is_progress_being_made": {"reason": "r", "answer": True},
"next_speaker": {"reason": "r", "answer": speaker},
"instruction_or_question": {"reason": "r", "answer": "Shape one hypothesis."},
}
)
def _manager_script(ledgers: list[str]) -> Callable[[str, str], str]:
"""Route a manager prompt blob to its scripted reply (the econ-56 helper, verbatim in shape).
The stage ORDER is load-bearing (§ F, A6): the selector sees the CONCATENATION of the call's
messages, so a later-stage prompt still carries the earlier stage's text.
"""
def _select(blob: str, _role: str) -> str:
if "provide the final answer" in blob:
return "FINAL: exploration done."
if "pure JSON format" in blob:
return ledgers.pop(0) if ledgers else _ledger_json(satisfied=True)
if "went wrong on this last run" in blob:
return "PLAN-UPDATE: revised plan."
if "rewrite the following fact sheet" in blob:
return "FACTS-UPDATE: revised facts."
if "bullet-point plan" in blob:
return "PLAN: - ask the hypothesiser"
if "pre-survey" in blob:
return "FACTS: the bundle is anchored."
return "{}"
return _select
def _hypothesis_line(label: str, rationale: str) -> str:
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps({"label": label, "rationale": rationale})
def _factory(
*,
ledgers: list[str],
hypothesiser: list[str],
fallback: str = "ok",
sink: list[str] | None = None,
) -> Callable[[str], BaseChatClient]:
"""One fresh ``ScriptedChatClient`` per role — the exploration's three plus everyone else.
``fallback`` serves the roles the PIPELINE builds (proposer/checker), so one factory can drive
an exploration and the run it hands its mandate to. That is what makes the end-to-end arm an
end-to-end arm rather than two half-proofs.
"""
def factory(role: str) -> BaseChatClient:
if role == explore.MANAGER_ROLE:
return ScriptedChatClient(sink=sink, reply_selector=_manager_script(ledgers), role=role)
if role == explore.HYPOTHESISER_ROLE:
replies = list(hypothesiser)
def _hyp(_blob: str, _role: str) -> str:
return replies.pop(0) if replies else "nothing further."
return ScriptedChatClient(sink=sink, reply_selector=_hyp, role=role)
if role == explore.NAVIGATOR_ROLE:
return ScriptedChatClient("NAVIGATOR: index read.", sink, role=role)
return ScriptedChatClient(fallback, sink, role=role)
return factory
def _contract(**overrides: Any) -> ExplorationContract:
return ExplorationContract(**{**_CONTRACT_JSON, **overrides})
# ---------------------------------------------------------------------------------------------
# 1. The caller-owned accumulator (the funn-1 sink shape, applied to the exploration)
# ---------------------------------------------------------------------------------------------
def _micro_bundle_dir() -> str:
return str(
Path(portfolio_optimiser.__file__).parent
/ "data"
/ "bundles"
/ "bygg-energi-baseline-mikro"
)
def test_quick_validate_verdicts_reach_the_callers_trace() -> None:
"""T1: every advisory verdict the hypothesiser asked for is recorded where the caller can read
it — the ONE thing ``ExplorationResult`` deliberately does not carry.
Called DIRECTLY, because a scripted run never reaches a tool body (measured, econ 56): an arm
that drove ``explore()`` and then asserted on an empty list would be green against every
implementation, including one with no sink at all.
Detach point: drop the ``sink`` append in ``quick_validate_tool`` → RED.
"""
base = _micro_bundle_dir()
projection = dict(okf.load_ir_projection(base))
projection.pop("_note", None)
trace = ExplorationTrace()
validate = explore.quick_validate_tool((base,), sink=trace.quick_validations)
honest = validate.func(
bundle_id="bygg-energi-baseline-mikro", proposal_json=json.dumps(projection)
)
invented = dict(projection)
invented["affected_items"] = [
{**dict(projection["affected_items"][0]), "code": "CODE-THAT-DOES-NOT-EXIST"}
]
validate.func(bundle_id="bygg-energi-baseline-mikro", proposal_json=json.dumps(invented))
assert len(trace.quick_validations) == 2, "both calls must be recorded, in call order"
first, second = trace.quick_validations
assert first.bundle_id == "bygg-energi-baseline-mikro"
assert first.verdict == honest, "the recorded verdict must be the one the tool ANSWERED"
assert second.verdict["decision"] == "rejected"
assert "CODE-THAT-DOES-NOT-EXIST" in second.proposal_json
@pytest.mark.asyncio
async def test_the_returned_ledger_log_is_the_traces_own_entries() -> None:
"""T2: ``ExplorationResult.ledger_log`` is BUILT FROM the accumulator, never alongside it.
Two lists holding one fact drift (kø-(p)), and a drifted pair would let the returned result and
the written artefact describe different runs.
Detach point: accumulate rounds in a second local list → RED.
"""
trace = ExplorationTrace()
result = await explore.explore(
"Find a saving.",
contract=_contract(),
bundle_dirs=(str(_BUNDLE_DIR),),
client_factory=_factory(
ledgers=[_ledger_json(satisfied=False), _ledger_json(satisfied=True)],
hypothesiser=[_hypothesis_line(_LABEL, "because the bundle says so")],
),
trace=trace,
)
assert result.stop is None
assert len(trace.ledger) == 2
assert tuple(trace.ledger) == result.ledger_log
@pytest.mark.asyncio
async def test_the_trace_survives_the_budget_exception_that_destroys_the_result() -> None:
"""T3: the round cap raises, and the caller STILL holds every round the loop recorded.
This is the whole reason the accumulator is caller-owned. The round cap leaves as a typed
``BudgetExceeded`` (econ 56), so nothing is returned — and § C.2 requires the artefact to be
readable no matter which guard fired.
Detach point: return the log only, keeping no caller-visible accumulator → RED.
"""
trace = ExplorationTrace()
with pytest.raises(BudgetExceeded) as excinfo:
await explore.explore(
"Find a saving.",
contract=_contract(max_rounds=2),
bundle_dirs=(str(_BUNDLE_DIR),),
client_factory=_factory(
ledgers=[_ledger_json(satisfied=False), _ledger_json(satisfied=False)],
hypothesiser=["still thinking."],
),
trace=trace,
)
assert excinfo.value.kind == "exploration_rounds"
assert len(trace.ledger) == 2, (
"the rounds the exploration DID record were destroyed with the result — the artefact a "
"capped run needs most would be empty"
)
# ---------------------------------------------------------------------------------------------
# 2. The CLI door — every refusal by name, never a silent merge or a silent drop
# ---------------------------------------------------------------------------------------------
def _config_file(tmp_path: Path, **overrides: Any) -> str:
path = tmp_path / "exploration.json"
path.write_text(json.dumps({**_CONTRACT_JSON, **overrides}), encoding="utf-8")
return str(path)
def _mandate_file(tmp_path: Path) -> str:
path = tmp_path / "mandate.json"
path.write_text(
json.dumps(
{
"objective": "cut energy cost",
"approaches": [{"id": "a1", "label": "LED", "description": "swap the fittings"}],
"allow_own_proposals": False,
}
),
encoding="utf-8",
)
return str(path)
def _base_argv(tmp_path: Path) -> list[str]:
return [
_PID,
"--docs-dir",
str(_BUNDLE_DIR),
"--bundle-dir",
str(_BUNDLE_DIR),
"--explore",
"Find the cheapest saving.",
"--explore-config",
_config_file(tmp_path),
]
def test_an_exploration_config_without_an_exploration_is_refused_by_name(tmp_path, capsys) -> None:
"""T4: ``--explore-config`` alone would be loaded and then dropped on the floor — the exact
silent-ignore ``--embedder-config requires --semantic-retrieval`` exists to prevent.
Detach point: drop the refusal → RED.
"""
rc = run.main(
[_PID, "--docs-dir", str(_BUNDLE_DIR), "--explore-config", _config_file(tmp_path)]
)
assert rc == 1
assert "--explore-config" in capsys.readouterr().err
def test_an_exploration_without_its_bounds_is_refused_rather_than_defaulted(
tmp_path, capsys
) -> None:
"""T5: ``--explore`` alone is refused — the CLI may not invent bounds.
Every ``ExplorationContract`` field is required WITHOUT a default precisely because
``MagenticBuilder`` falls back to unbounded, and a CLI that supplied its own numbers would undo
that decision one layer up.
"""
rc = run.main(
[_PID, "--docs-dir", str(_BUNDLE_DIR), "--bundle-dir", str(_BUNDLE_DIR), "--explore", "go"]
)
assert rc == 1
assert "--explore-config" in capsys.readouterr().err
def test_explore_and_mandate_are_two_sources_of_one_mandate_and_are_refused_together(
tmp_path, capsys
) -> None:
"""T6: the decision, made deliberately and stated: REFUSE, never merge.
``explore()`` takes the objective from the prompt and hardcodes ``allow_own_proposals=True``, so
composing the two would silently overwrite fields the operator wrote by hand. The message names
the library door (``seed_approaches``) so the refusal teaches instead of only forbidding.
Detach point: let one source silently win → RED.
"""
rc = run.main(_base_argv(tmp_path) + ["--mandate", _mandate_file(tmp_path)])
assert rc == 1
err = capsys.readouterr().err
assert "--explore" in err and "--mandate" in err
assert "seed_approaches" in err, "the refusal must name the door that DOES serve door 1"
def test_explore_and_live_dry_run_contradict_and_are_refused(tmp_path, capsys) -> None:
"""T7: ``--live-dry-run`` stops before the first model call; an exploration IS model calls."""
rc = run.main(_base_argv(tmp_path) + ["--live-dry-run"])
assert rc == 1
assert "--live-dry-run" in capsys.readouterr().err
def test_an_exploration_with_no_knowledge_base_is_refused(tmp_path, capsys) -> None:
"""T8: without ``--bundle-dir`` the navigator has nothing to open — the loop would run, cost
tokens and read nothing. Refused rather than run empty (the ``--semantic-retrieval`` shape).
Detach point: drop the requirement → RED.
"""
rc = run.main(
[
_PID,
"--docs-dir",
str(_BUNDLE_DIR),
"--explore",
"go",
"--explore-config",
_config_file(tmp_path),
]
)
assert rc == 1
assert "--bundle-dir" in capsys.readouterr().err
def test_a_plan_review_nobody_can_answer_is_refused_at_the_cli(tmp_path, capsys) -> None:
"""T9: ``enable_plan_review`` is the U13 SYNCHRONOUS door, and a run must never stop at a
review nobody offered to answer.
Refused HERE rather than left to ``explore()``: ``ExplorationError`` is a ``RuntimeError``, so
it is outside ``main()``'s ``(ValueError, FileNotFoundError, ValidationError)`` refusal tuple
and would leave as a traceback instead of the rc-1 line every other misconfiguration produces.
**The assertion names wording unique to THIS branch.** Since F4 the CLI has a second refusal
carrying ``enable_plan_review`` (``--plan-review`` against a config that asks for no review),
so asserting on the shared token would pass against a surface missing this branch entirely —
the økt-57 mutation, in the form this repo keeps meeting it.
Detach point: let the flag through to ``explore()`` → RED (traceback, not rc 1).
"""
rc = run.main(
_base_argv(tmp_path)[:-1] + [_config_file(tmp_path, enable_plan_review=True)],
)
assert rc == 1
err = capsys.readouterr().err
assert "no reviewer was offered" in err
assert "--plan-review" in err, "the refusal must name the door that answers it (F4)"
def test_explore_belongs_to_single_project_mode(tmp_path, capsys) -> None:
"""T10: portfolio mode is a documented partition, and ``--explore`` is on the single-project
side of it — one exploration shapes ONE mandate against ONE knowledge base.
The assertion names ``--portfolio``, and that was MEASURED rather than chosen: asserting only
that the message mentions ``--explore`` passed against an implementation with no partition
entry at all, because the run then fell through to ``--explore requires --bundle-dir``, which
names ``--explore`` too. Two refusals sharing a substring is this repo's "assert never on
wording two branches share" rule, caught by its own mutation.
Detach point: drop ``--explore`` from the portfolio ``single_only`` partition → RED.
"""
rc = run.main(["--portfolio", "--explore", "go", "--explore-config", _config_file(tmp_path)])
assert rc == 1
err = capsys.readouterr().err
assert "--explore" in err and "--portfolio" in err
@pytest.fixture()
def _explored_main(monkeypatch: pytest.MonkeyPatch) -> list[str]:
"""Inject the role-dispatching scripted factory into the seam ``main()`` resolves through.
``main()`` passes no ``client_factory``, and ``explore()`` imports ``run._default_factory``
lazily at call time, so this ONE patch covers both the exploration and the pipeline it feeds —
which is what makes the arm below end-to-end rather than a wiring spy.
"""
sink: list[str] = []
factory = _factory(
ledgers=[_ledger_json(satisfied=False), _ledger_json(satisfied=True)],
hypothesiser=[_hypothesis_line(_LABEL, "the index says the fittings are old")],
fallback=_ENERGY_REPLY,
sink=sink,
)
monkeypatch.setattr("portfolio_optimiser.run._default_factory", lambda profile: factory)
return sink
def test_the_shaped_mandate_reaches_the_pipeline(tmp_path, capsys, _explored_main) -> None:
"""T11: the approach the hypothesiser shaped is SETTLED by the run — the whole point of (1).
The settlement is printed only for a run that HAS a mandate, and the label appears nowhere in
the bundle or the reference projects, so it can have reached stdout only by travelling
prompt → ``explore()`` → ``Mandate`` → ``run_project(mandate=…)`` → ``settle``.
Detach point: drop ``mandate=`` from the exploring branch's ``run_project`` call → RED.
"""
rc = run.main(_base_argv(tmp_path))
assert rc == 0
out = capsys.readouterr().out
assert _LABEL in out, "the exploration's mandate never reached the pipeline's settlement"
# ---------------------------------------------------------------------------------------------
# 3. The artefact — written from a ``finally``, because a capped run is what it exists for
# ---------------------------------------------------------------------------------------------
def test_the_exploration_artefact_carries_the_rounds_and_the_advisory_verdicts(
tmp_path, _explored_main
) -> None:
"""T12: ``{run_id}-exploration.json`` holds the per-round ledger AND the ``quick_validate``
verdicts — the level-1 evidence ``ExplorationResult`` deliberately does not carry (§ C.2).
Detach point: drop the artefact write → RED.
"""
outbox = tmp_path / "outbox"
rc = run.main(_base_argv(tmp_path) + ["--outbox-dir", str(outbox), "--run-id", "r1"])
assert rc == 0
payload = json.loads((outbox / "r1-exploration.json").read_text(encoding="utf-8"))
assert payload["run_id"] == "r1"
assert payload["completed"] is True
assert payload["stop"] is None
assert [row["round_index"] for row in payload["rounds"]] == [1, 2]
assert payload["rounds"][-1]["is_request_satisfied"] is True
assert payload["rounds"][0]["next_speaker"] == "hypothesiser"
assert "quick_validations" in payload
def test_the_artefact_is_written_even_when_the_exploration_was_cut_short(
tmp_path, monkeypatch
) -> None:
"""T13: a capped exploration is the run whose evidence matters MOST, and it is the one that
returns nothing — so the write lives in a ``finally`` (the ``write_parse_failures`` precedent).
``completed`` is a field rather than an inference: with no result there is no ``stop``, and a
``stop: null`` that meant BOTH "concluded normally" and "we never found out" would be the kind
of silence this repo writes required fields to close.
Detach point: move the write out of the ``finally`` → RED.
"""
factory = _factory(
ledgers=[_ledger_json(satisfied=False), _ledger_json(satisfied=False)],
hypothesiser=["still thinking."],
fallback=_ENERGY_REPLY,
)
monkeypatch.setattr("portfolio_optimiser.run._default_factory", lambda profile: factory)
outbox = tmp_path / "outbox"
with pytest.raises(BudgetExceeded):
run.main(
[
_PID,
"--docs-dir",
str(_BUNDLE_DIR),
"--bundle-dir",
str(_BUNDLE_DIR),
"--explore",
"go",
"--explore-config",
_config_file(tmp_path, max_rounds=2),
"--outbox-dir",
str(outbox),
"--run-id",
"r2",
]
)
payload = json.loads((outbox / "r2-exploration.json").read_text(encoding="utf-8"))
assert payload["completed"] is False
assert payload["stop"] is None
assert len(payload["rounds"]) == 2
def test_the_artefact_payload_is_byte_deterministic() -> None:
"""T14 (control): the same trace renders the same bytes, so the artefact is diff-stable like
every other outbox file. Drives the renderer directly — the CLI arms above prove it is CALLED,
this proves what it produces."""
trace = ExplorationTrace()
trace.ledger.append(
explore.LedgerEntry(
round_index=1,
is_request_satisfied=True,
is_in_loop=False,
is_progress_being_made=True,
next_speaker="hypothesiser",
instruction_or_question="Shape one hypothesis.",
speaker_known=True,
)
)
trace.quick_validations.append(
explore.QuickValidation(
bundle_id="b", proposal_json="{}", verdict={"decision": "unparseable"}
)
)
first = explore.trace_payload(trace, stop=None, completed=True)
second = explore.trace_payload(trace, stop=None, completed=True)
assert json.dumps(first, sort_keys=True) == json.dumps(second, sort_keys=True)
def test_a_seeded_mandate_still_leads_the_shaped_one() -> None:
"""T15 (control for T6's refusal): the library door the refusal names actually works.
A refusal that pointed at a door which did not open would be worse than no message at all.
"""
seed = Approach(id="expert-1", label="expert's own", description="the domain expert asked")
minted = explore._mint_approaches((seed,), [(_LABEL, "shaped in the loop", "")])
assert [a.id for a in minted] == ["expert-1", "hypothesis-1"]
assert isinstance(Mandate(objective="o", approaches=minted, allow_own_proposals=True), Mandate)
# ---------------------------------------------------------------------------------------------
# 4. The hosted surface — a THREE-way whitelist, and the Fase 4e rule extended to cover it
# ---------------------------------------------------------------------------------------------
def _hosted_payload(**extra: Any) -> dict[str, Any]:
return {
"project_id": _PID,
"docs_dir": str(_BUNDLE_DIR),
"verdict_input": {"decision": "approved", "rationale": "expert reviewed (explore)"},
# LOCAL, never the hosted AZURE default: the AZURE arm resolves a Foundry deployment name
# from the model map before any client is built, so it cannot complete offline.
"profile": "local",
**extra,
}
@pytest.fixture()
def _hosted_backend(monkeypatch: pytest.MonkeyPatch) -> list[str]:
"""The scripted backend behind a hosted invocation, plus the prompt sink that proves it ran.
``client_factory`` is refused by the invocations whitelist on purpose — the caller of a hosted
agent never chooses the server's model client — so ``run._default_factory`` is the only
injection point the surface leaves, and ``explore()`` resolves through the same one.
"""
sink: list[str] = []
factory = _factory(
ledgers=[_ledger_json(satisfied=False), _ledger_json(satisfied=True)],
hypothesiser=[_hypothesis_line(_LABEL, "the index says the fittings are old")],
fallback=_ENERGY_REPLY,
sink=sink,
)
monkeypatch.setattr("portfolio_optimiser.run._default_factory", lambda profile: factory)
return sink
@pytest.mark.asyncio
async def test_a_hosted_exploration_shapes_the_mandate_the_run_evaluates(_hosted_backend) -> None:
"""H1: ``explore_prompt`` over the hosted surface reaches the pipeline as a mandate.
Driven through ``hosting.invoke`` and the REAL ``run_project`` (Fase 4e): every other
invocations test hands ``invoke`` a recorder that swallows ``**kwargs`` and therefore cannot
see whether a new field composes with the signature at all.
The proof is the PROMPT, not the status code: an ``Approach``'s description reaches the
proposer VERBATIM, and this label exists nowhere in the bundle or the reference projects — so
finding it in a generation prompt means it travelled prompt → ``explore()`` → ``Mandate`` →
``run_project(mandate=…)``.
Detach point: stop passing the shaped mandate into ``run_project`` → RED.
"""
body = await hosting.invoke(
_hosted_payload(
bundle_dir=str(_BUNDLE_DIR),
explore_prompt="Find the cheapest saving.",
explore_contract=dict(_CONTRACT_JSON),
)
)
assert body["outcome_type"] in {"validated", "rejected"}
assert any(_LABEL in prompt for prompt in _hosted_backend), (
"the shaped approach never reached a prompt — the hosted door does not wire the mandate"
)
@pytest.mark.asyncio
async def test_a_consumed_field_is_never_forwarded_to_run_project() -> None:
"""H2: the whitelist is a THREE-way partition, and the consumed half is proved NEGATIVELY.
``explore_prompt``/``explore_contract`` are accepted by the surface and consumed BY it — they
are not ``run_project`` parameters, and forwarding one would be a ``TypeError`` answered as a
500. The positive half of Fase 4e (every forwarded field reaches the real signature) cannot
see that; without this arm a field sliding from consumed to forwarded is exactly the drift 4e
exists to catch.
Detach point: build ``kwargs`` from the whole payload again → RED.
"""
import inspect
_, kwargs, consumed = hosting._run_kwargs(
_hosted_payload(
bundle_dir=str(_BUNDLE_DIR),
explore_prompt="p",
explore_contract=dict(_CONTRACT_JSON),
)
)
assert set(hosting._CONSUMED_FIELDS).isdisjoint(kwargs), (
"a consumed field was forwarded to run_project, which does not take it"
)
assert set(consumed) == set(hosting._CONSUMED_FIELDS)
parameters = inspect.signature(run.run_project).parameters
assert set(hosting._CONSUMED_FIELDS).isdisjoint(parameters), (
"a CONSUMED field is a run_project parameter — it belongs in the forwarded half"
)
for name in (*hosting._REQUIRED_FIELDS, *hosting._OPTIONAL_FIELDS):
# project_id is positional; every other forwarded field must be a real keyword.
assert name in parameters, f"whitelisted field {name!r} is not a run_project parameter"
@pytest.mark.asyncio
@pytest.mark.parametrize(
("payload", "expected"),
[
pytest.param(
{"bundle_dir": str(_BUNDLE_DIR), "explore_contract": dict(_CONTRACT_JSON)},
"explore_prompt",
id="bounds-without-an-exploration",
),
pytest.param(
{"bundle_dir": str(_BUNDLE_DIR), "explore_prompt": "p"},
"explore_contract",
id="exploration-without-bounds",
),
pytest.param(
{"explore_prompt": "p", "explore_contract": dict(_CONTRACT_JSON)},
"bundle_dir",
id="exploration-without-a-knowledge-base",
),
pytest.param(
{
"bundle_dir": str(_BUNDLE_DIR),
"explore_prompt": "p",
"explore_contract": {**_CONTRACT_JSON, "enable_plan_review": True},
},
"enable_plan_review",
id="a-review-nobody-can-answer",
),
],
)
async def test_the_hosted_door_refuses_by_name_on_the_callers_channel(payload, expected) -> None:
"""H3: each hosted refusal names the field, and each is a ``ValueError`` — the 400 arm.
``enable_plan_review`` is the one that had to be refused HERE rather than in ``explore()``:
``ExplorationError`` is a ``RuntimeError``, so leaving it to the loop would answer a caller's
configuration mistake on the crash channel (500), which is where a fallen-over endpoint lives.
A synchronous plan review would also block the HTTP request on a reviewer that does not exist.
Detach point: drop any one of the four guards → RED.
"""
with pytest.raises(ValueError) as excinfo:
await hosting.invoke(_hosted_payload(**payload))
assert expected in str(excinfo.value)
# ---------------------------------------------------------------------------------------------
# 5. The demo scenario — a THIRD entry, reachable only by name
# ---------------------------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_the_demo_scenario_lets_a_shaped_direction_reach_the_hypothesis(tmp_path) -> None:
"""S1: the offline walkthrough of U4 — a prompt and a knowledge base become a mandate, and the
direction the loop shaped reaches the proposer VERBATIM.
The same honesty limit the rest of the demo carries applies here and is worth restating: this
proves the plumbing and that the data flow closes, NOT that a live model would shape a good
direction. Every reply is scripted.
Detach point: drop ``mandate=`` from the scenario's ``run_project`` call → RED.
"""
result = await simulation.simulate_exploration(str(_BUNDLE_DIR), str(tmp_path), max_rounds=3)
assert [a.label for a in result.exploration.mandate.approaches] == [result.label]
assert result.label_in_generation_prompt, (
"the shaped direction never reached the hypothesis prompt — the demo would show a mandate "
"the pipeline ignored"
)
assert result.trace.ledger, "the exploration recorded no rounds"
def test_an_outbox_without_a_run_id_is_refused_before_the_exploration_spends_anything(
tmp_path, _explored_main
) -> None:
"""T16: an argv that cannot finish is refused BEFORE the loop costs anything.
``run_project`` refuses ``outbox_dir`` without ``run_id`` at its very first statement, which is
early enough for every path that existed before U4. The exploration runs AHEAD of that call, so
without this guard the run spends its whole exploration budget on model calls and only then
refuses — and the artefact write is skipped too, so not even the evidence of what was spent
survives. Exactly the hoist ``main()`` already performs twice ("an incomplete argv is refused
BEFORE the honesty banner could claim a scripted run happened").
The assertion is that NO model call happened, not merely that rc is 1: a refusal that arrives
after the spend looks identical at the exit code.
Detach point: drop the guard from the exploration block → RED.
"""
rc = run.main(_base_argv(tmp_path) + ["--outbox-dir", str(tmp_path / "outbox")])
assert rc == 1
assert not _explored_main, (
"the exploration made model calls before the run was refused — the budget was spent on an "
"argv that could never finish"
)
@pytest.mark.asyncio
async def test_a_direction_the_base_already_states_is_refused_as_vacuous(tmp_path) -> None:
"""S2: a label the knowledge base ALREADY contains is refused, not demonstrated.
Exactly the guard ``simulate_learning_loop`` raises on when its two markers coincide: the
scenario's whole claim is that the direction came from the LOOP, and a label the bundle states
on its own would reach the prompt as ordinary context — a demonstration that demonstrates
nothing, which is this repo's vacuous-gate class in demo form.
Detach point: drop the guard → RED.
"""
stated = "LED-retrofit" # present in the bundle's own text
with pytest.raises(ValueError) as excinfo:
await simulation.simulate_exploration(str(_BUNDLE_DIR), str(tmp_path), label=stated)
assert stated in str(excinfo.value)