feat(explore): U4+U13 synkron - utforskningssloeyfa som mandat-former (ORDRE 20260823T185602Z) [skip-docs]

Magentic legges OVER den normative sloeyfa, aldri inni Steg 3: prompt +
kunnskapsbaser -> Mandate -> run_project(mandate=...) UENDRET. Manageren velger
VEI; det som forlater friheten er et Mandate, aldri et forslag. explore() skriver
ingenting - niva 3 (skriverettigheter) tilhoerer pipelinen alene.

Levert i denne oekten (kjernen; kallstedene staar til oekt 57):
- ExplorationContract: seks paakrevde felt uten default. max_reset_count=0 nektes
  paa en MAALING - reset_count >= max_reset_count mot en teller som starter paa 0
  terminerer kjoeringen FOER foerste runde med null ledger-events, altsaa en
  utforskning som utforsket ingenting, forkledd som en stall som aldri skjedde.
- explore() + fresh_exploration_workflow(): fersk builder per utforskning,
  BudgetMiddleware paa HVER agent inkl. manageren, synkron plan review via
  request_info, og max_plan_revisions som binder den ubundne revise-loekka.
- Tre kanaler: tokens OG runder raiser BudgetExceeded (rundene oversatt av vaart
  lag som kind="exploration_rounds", fordi orkestreringen maalt ikke raiser ved
  sitt eget rundetak); alt semantisk er en VERDI i stop.
- quick_validate (niva 1, raadgivende) + navigator-verktoey over safe_resolve.
- U14s tre utsatte events landet som span-events paa EN exploration-span.

Load-bearing MAALT mot HELE suiten, groenn kontroll 975/5, golden ea8c534
uendret: tolv mutasjoner alle roede. TO av dem falsifiserte testen foerst -
skrivefrihets-testen naadde aldri en verktoeykropp (ScriptedChatClient emitterer
ingen verktoeykall), og stdout-testens capsys er blind for ConsoleSpanExporter,
hvis out-default bindes ved modulimport. Begge er rettet; stdout-armen er naa en
subprosess, som er P4-presedensen.

[skip-docs] fordi flaten ikke er naabar for en bruker enna: --explore, det
whitelistede hosting-feltet og sim-scenarioet bygges i oekt 57, og en
README-oppfoering naa ville vaert en paastand om en inngang som ikke finnes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-08-23 22:39:00 +02:00
commit f0c54cc8dc
4 changed files with 1892 additions and 4 deletions

View file

@ -772,6 +772,70 @@ Python ≥3.10. MAF (`agent-framework-core` 1.9.0). Pakkehåndtering: `uv`. To b
fire-linjers stderr og `test_portfolio_cli_offline`s stille-pass — altså er omisjonen gatet av
uavhengige vitner) · detach demo-wiringen + detach CLI-wiringen (5 røde) · detach hosting-wiringen
(1 rød, KUN subprosess-testen — P4-presedensen).
- **Utforskningssløyfa er en MANDAT-FORMER, og de tre garantinivåene er strukturelle (U4+U13
synkron, økt 56):** `explore.py` legger en Magentic-manager OVER den normative sløyfa — `prompt +
kunnskapsbaser → Mandate → run_project(mandate=…)` UENDRET, Steg 3s maker-checker urørt (commons-eid
og normativ). Manageren velger VEI; det som forlater friheten er `mandate.Mandate`, aldri et forslag.
**Nivå 1** = `quick_validate`-verktøyet (SAMME `validate_proposal`, SAMME baseline, men rådgivende —
når ALDRI provenance); **nivå 2** = pipelinen som stempler; **nivå 3** = skriverettigheter, som kun
pipelinen har. `explore()` skriver INGENTING. **Mandatet bygges fra hypotesiserens MERKEDE turer**
(`HYPOTHESIS: {"label","rationale"}`), aldri fra sluttsvaret — sluttsvaret er RÅTT per design, og en
parser på det ville gjort det til et forslag. Markøren er dét som gjør fail-closed mulig: en umerket
tur er ikke en påstand (ingen stillhet å lukke), mens en MERKET-men-uleselig linje raiser
(`write_concept_file`-regelen). **Frø-approaches bevares ALLTID og FØRST** — også når sløyfa fant
ingenting og også ved stopp (§ C.6 dør 1 er en bevaringsregel, ikke en belønning for å bli ferdig).
**Tre kanaler, aldri én:** tokens OG runder raiser `BudgetExceeded` (rundene som
`kind="exploration_rounds"`, oversatt av VÅRT lag fordi orkestreringen MÅLT ikke raiser ved sitt eget
rundetak — den returnerer en kanonisk assistent-melding som ved transporten er uskillbar fra suksess),
mens alt semantisk er en VERDI i `stop` (S3.4-splitten: utmattelse og utfall er ikke samme sak).
Diskriminatoren mellom «nådde taket» og «ble kappet av taket» er SISTE ledgers
`is_request_satisfied` — samme felt orkestratoren selv forgrener på (`:1106`) — aldri
termineringsmeldingen, som er en inline f-string (`:1253`) uten konstant å pinne mot og som en modells
eget sluttsvar kan inneholde. **`speaker_known` sjekkes FØRST og slår alt annet:** en `next_speaker`
uten treff gir stille sluttsvar med NULL deltakerarbeid (`:1128-1131`), altså et plausibelt svar
ingen jobbet for (E2-klassen) — sløyfas egne funn holdes da tilbake, frøene ikke.
**Kontrakten nekter tre ting ved konstruksjon:** hvert av seks felt er PÅKREVD uten default (MAF
defaulter `max_round_count`/`max_reset_count` til ubegrenset, så et utelatt felt faller ikke tilbake
til noe forsiktig, men til dét `method-spec` §8 forbyr); `max_reset_count=0` nektes — **MÅLT**, ikke
resonnert: `reset_count >= max_reset_count` mot en teller som starter på 0 gjør at kjøringen
terminerer FØR første runde med kun `facts`+`plan`, null ledger-events og «maximum reset count», altså
en utforskning som utforsket ingenting, forkledd som en stall som aldri skjedde (`max_stall_count=0`
er derimot LOVLIG — strengt `>` gjør 0 til «reset ved første stallede runde»); og
`max_plan_revisions>0` med `enable_plan_review=False` nektes (en cap på en hendelse som ikke kan skje).
`max_plan_revisions` finnes fordi A3 MÅLTE at en `revise` koster 2 manager-kall, **null** ledger-kall
og **null** runder og så spør PÅ NYTT — under rundetaket alene er en alltid-reviderende ekspert
ubundet forbruk under vakter som alle ser tilfredse ut. Ved cap: typet stopp, ALDRI en påtvunget
approve (repair av et menneskes beslutning er den verste sorten). **U14s tre utsatte events er
landet** (`plan_created`/`replanned`/`progress_ledger_updated` som span-events på ÉN
`exploration`-span) — emisjon er UBETINGET og «av» betyr at OTel kaster dem, samme form MAFs egen
instrumentering alt har; en flagget emitter ville vært en andre oppløsning av regelen `tracing.py`
eier. Uttalt ærlighets-grense: eventene registreres når event-strømmen foldes, så REKKEFØLGEN er
tro og tidsstemplene er ikke øyeblikkene manageren handlet.
**Load-bearing MÅLT** (`tests/test_explore_loadbearing.py`, 32 tester), tolv mutasjoner alle røde mot
HELE suiten + grønn kontroll 975/5: detach rundetak-oversettelsen (1 rød) · test rundetaket FØR
tilfredsstillelse (1 rød — den motsatte feilen, som gjør en fullført utforskning til en budsjettfeil) ·
detach ukjent-taler-sjekken (1) · `BudgetMiddleware` av manageren, deltakerne beholder den (1) ·
detach revisjons-capen (1) · dropp frø-bevaringen (3) · la en ukjent-taler-kjøring levere funnene
videre (1) · detach alle tre span-events (1) · tillat `max_reset_count=0` (1) · gjør en merket-men-
uleselig hypotese tolerant (1) · la `read_bundle` skrive i basen den leser (1) · send spans til OTels
default-sink (2). **TO av dem FALSIFISERTE testen først, og begge er repoets vakuøs-gate-klasse:**
(i) skrivefrihets-testen drev kun `explore()`, men en `ScriptedChatClient` returnerer TEKST og
emitterer aldri et verktøykall — så ingen scriptet kjøring når en verktøykropp, og hele lesesømmen
(eneste sted en skriving realistisk kan komme fra) lå utenfor gaten; testen kaller nå hvert verktøy
DIREKTE. (ii) stdout-testen brukte `capsys`, men `ConsoleSpanExporter`s `out`-default bindes når
`opentelemetry.sdk.trace.export` FØRST importeres — under pytest er det stdout ved COLLECTION, som
`capsys` aldri ser; spans lå faktisk på stdout mens asserten var grønn. Dét er ikke en test-quirk å
omgå, det er nøyaktig faktumet U14 finnes for, og arven er P4-presedensen: **subprosessen er
målingen**. Begge armene kjøres nå i et barn (av: null trace-data noe sted; `PORTFOLIO_OTEL=console`:
spanet + `progress_ledger_updated`**stderr** og stdout tomt), med `EXPLORATION-OK` på stderr som
kontroll — uten den ville «stdout var tomt» vært like sant om et barn som krasjet ved import.
**Ærlighets-grenser, uttalt:** multi-base-dispatch (`Approach.bundle_id`, § C.7) venter til
`run_project` tar mer enn én `bundle_dir` — å shippe feltet før konsumenten er en form gjettet i
stedet for målt; `quick_validate`-dommene hypotesiseren så bor ikke i `ExplorationResult`, de er nivå
1 og hører hjemme i `{run_id}-exploration.json` som CLI-wiringen skriver; utforskningsrollene løses
via `resolve_model`s `default`-fallback til en operatør mapper dem eksplisitt; at en LEVENDE modell
kaller verktøyene er ikke bevist offline (samme klasse som structured-output-grensen). `--explore` i
`run.py`, `explore_prompt` i `hosting.py` og sim-scenarioet er IKKE bygget (økt 57).
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -0,0 +1,804 @@
"""U4 + U13-synchronous — the Magentic exploration loop, sitting OVER the normative pipeline.
**One sentence of design, and it is not to be reopened without a measurement.** Magentic is laid
over the eight-step loop, never inside Step 3: ``prompt + knowledge bases -> Mandate ->
run_project(mandate=...)``, with that last call byte-for-byte the one it already was. The manager
is free to choose which base to open and which hypothesis to shape next; what leaves that freedom
is a ``mandate.Mandate`` approaches worth *testing*. The exploration's own final answer is RAW,
never a proposal, and ``explore()`` writes to neither the outbox nor the wiki. Every number that
survives is gated by ``validate_proposal`` inside ``run_project``, in the same blocking gate as
today, and ``shared/method-spec.md`` §3's maker-checker debate is untouched (it is commons-owned
and normative; replacing it would need an amendment, not a module).
**Three levels of guarantee, stated so nobody reads more into the loop than is there.** The
``quick_validate`` tool the hypothesiser calls in-loop is level 1: the SAME ``validate_proposal``,
against the SAME baseline, but its verdict is advisory it never becomes provenance. Level 2 is
the pipeline, where each approach is validated and stamped. Level 3 is write access, which only
the pipeline has.
This module imports ``agent_framework.orchestrations`` and therefore may never be imported from
``okf.py`` / ``mandate.py`` / ``hitl.py``, which are held framework-neutral by
``tests/test_okf.py::test_okf_is_maf_free``.
"""
from __future__ import annotations
import json
from collections.abc import Callable, Mapping, Sequence
from dataclasses import dataclass
from pathlib import Path
from typing import Any, Final, Literal
from agent_framework import Agent, BaseChatClient, FunctionTool, tool
from agent_framework.orchestrations import (
MagenticBuilder,
MagenticOrchestratorEventType,
MagenticPlanReviewResponse,
)
from pydantic import BaseModel, Field, ValidationError, model_validator
from portfolio_optimiser import okf
from portfolio_optimiser.backends import Profile
from portfolio_optimiser.budget import Budget, BudgetExceeded, BudgetMiddleware, TokenMeter
from portfolio_optimiser.ir import SavingsProposal
from portfolio_optimiser.mandate import OWN_PROPOSAL_ID, Approach, Mandate
from portfolio_optimiser.retrieval import safe_resolve
from portfolio_optimiser.tracing import exploration_tracer
from portfolio_optimiser.validator import Rejection, validate_proposal
class ExplorationContract(BaseModel):
"""The stated bounds of ONE exploration. Every field is required; none has a default.
Mirrors ``contracts.TerminationContract`` in spirit and goes further in one respect: there is
nothing to inherit. ``MagenticBuilder`` defaults ``max_round_count`` to ``None`` (unbounded)
and ``max_reset_count`` to ``None`` (unlimited), so a field left out here would not fall back
to something conservative it would fall back to the one shape ``shared/method-spec.md`` §8
forbids outright, an unbounded loop. A default would also assert an intent nobody stated,
which is the ground ``ProvenanceStamp.cost_baseline_anchored`` is required on.
``max_plan_revisions`` earns its place by measurement, not symmetry (§ F, A3): a plan-review
``revise`` costs two manager calls, emits **no** progress ledger and consumes **no** round,
then asks again. Under the round cap alone an always-revising expert is an unbounded spend
that the round counter never sees. ``0`` is meaningful and allowed: the plan must be approved
as first written or the exploration stops.
"""
#: Hard cap on orchestration rounds (``MagenticBuilder(max_round_count=...)``).
max_rounds: int = Field(gt=0)
#: Hard cap on tokens, enforced by ``BudgetMiddleware`` on EVERY agent, manager included.
max_tokens: int = Field(gt=0)
#: Consecutive no-progress rounds tolerated before the manager resets and replans. ``0`` is
#: allowed and means "reset on the first round that reports no progress": the orchestrator's
#: check is STRICT (``stall_count > max_stall_count``, ``_magentic.py:1118``) and the counter is
#: incremented ahead of it, so zero is strictness rather than self-defeat.
max_stall_count: int = Field(ge=0)
#: Resets tolerated before the exploration is declared stalled. Must be POSITIVE, and that is a
#: MEASUREMENT: the limit check is ``reset_count >= max_reset_count`` (``:1243``) against a
#: counter starting at zero, so a cap of zero is already met before the first round. Measured
#: against the installed stack, ``max_reset_count=0`` produces only the ``facts`` and ``plan``
#: manager calls, **zero** progress-ledger events, and the canonical "maximum reset count"
#: termination — an exploration that explored nothing, dressed as a stall that never happened.
max_reset_count: int = Field(gt=0)
#: Plan revisions the reviewer may ask for before the exploration stops. See the class note.
max_plan_revisions: int = Field(ge=0)
#: Whether a human/persona signs the plan off before the loop runs (the U13 synchronous door).
enable_plan_review: bool
@model_validator(mode="after")
def _a_revision_cap_needs_a_review_to_cap(self) -> "ExplorationContract":
"""A revision cap with the review switched off bounds an event that cannot occur.
Refused rather than dropped: a caller who wrote ``max_plan_revisions=3`` believes they
bounded something, and silently ignoring it is the failure mode this repo's flag contract
exists to prevent (``--embedder-config requires --semantic-retrieval``, same rule). The
coherent way to say "no reviews" is ``enable_plan_review=False, max_plan_revisions=0``.
"""
if not self.enable_plan_review and self.max_plan_revisions:
raise ValueError(
f"max_plan_revisions={self.max_plan_revisions} bounds plan revisions, but "
"enable_plan_review is false so no plan review is ever requested and no revision "
"can occur; set max_plan_revisions=0 or enable the review"
)
return self
#: The manager's role name. Not a participant: the manager plans, picks the next speaker and
#: keeps the progress ledger, and ``BudgetMiddleware`` rides on it like on everyone else (A1,
#: measured green — without that, the most talkative agent in the loop would be the one outside
#: the token cap).
MANAGER_ROLE: Final = "manager"
#: Reads the knowledge bases PROGRESSIVELY (§3 Steg 1) — index, then whatever the index links.
NAVIGATOR_ROLE: Final = "navigator"
#: Shapes ONE candidate direction at a time and may have its numbers advisory-checked in-loop.
HYPOTHESISER_ROLE: Final = "hypothesiser"
#: The participants, in build order. The manager's ``next_speaker`` is validated against exactly
#: these names: a ledger naming anyone else makes the orchestrator emit a final answer having
#: asked NOBODY (measured, ``_magentic.py:1128-1131``), which is a plausible answer produced by
#: zero work — the hazard class this repo retired ``E2`` for.
PARTICIPANT_ROLES: Final = (NAVIGATOR_ROLE, HYPOTHESISER_ROLE)
#: The span every exploration records its decisions under. One span per exploration, with the
#: manager's decisions as span EVENTS on it, because those are moments INSIDE one activity rather
#: than activities of their own — and because a reader wants them in order, on one timeline.
EXPLORATION_SPAN: Final = "exploration"
#: The line prefix that turns a hypothesiser turn into a candidate direction. A marker rather than
#: "parse every turn as JSON" because most turns legitimately are not hypotheses — the agent also
#: reasons out loud. That distinction is what lets an UNPARSEABLE marked line be a hard error
#: instead of a silence: with no marker there is nothing to be silent about.
HYPOTHESIS_MARKER: Final = "HYPOTHESIS:"
_INSTRUCTIONS: Final = {
NAVIGATOR_ROLE: (
"You read the project's knowledge bases. Use list_bundles to see what exists, then "
"read_bundle to open ONE at a time and read_file to follow a specific document. Quote "
"what you found; never guess at content you have not read."
),
HYPOTHESISER_ROLE: (
"You shape ONE candidate cost-saving direction at a time from what the navigator found. "
"You may call quick_validate to sanity-check a candidate's numbers; its verdict is "
"ADVISORY and is not the project's decision. When you commit to a direction, end your "
f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", '
'"rationale": "<why this project, in your own words>"}'
),
}
class ExplorationError(RuntimeError):
"""The exploration cannot be honoured as configured, or produced something unreadable."""
@dataclass(frozen=True)
class LedgerEntry:
"""One progress-ledger round, reduced to the five fields the manager steers on (C.2).
``speaker_known`` is not part of the ledger MAF produces it is this layer's reading of it,
and the reason the whole entry is recorded rather than counted. A ``next_speaker`` matching no
participant does not raise: the orchestrator logs a warning and jumps straight to a final
answer, so the run returns a plausible result that nobody worked for. Recording the fact is
what lets ``explore()`` refuse to hand that result onward as if it had been explored.
"""
round_index: int
is_request_satisfied: bool
is_in_loop: bool
is_progress_being_made: bool
next_speaker: str
instruction_or_question: str
speaker_known: bool
@dataclass(frozen=True)
class PlanReview:
"""One synchronous plan-review round trip (U13, door 2 of § C.6).
``is_stalled`` is carried because the same request type serves two very different moments:
the initial sign-off, and a re-review after the manager reset and replanned. A reviewer that
cannot tell them apart cannot answer the second one usefully.
"""
index: int
plan: str
is_stalled: bool
decision: Literal["approve", "revise"]
feedback: str = ""
#: Why an exploration ended without a satisfied request. ``None`` means it concluded normally.
#: Resource exhaustion is NOT here: tokens and rounds raise ``BudgetExceeded`` (the 429 channel),
#: because "we ran out" and "we finished, unsatisfied" are the two things S3.4 split apart and a
#: single field would fuse back together.
ExplorationStop = Literal["stalled", "plan_revisions_exhausted", "unknown_speaker"]
@dataclass(frozen=True)
class ExplorationResult:
"""What one exploration produced. The mandate is the product; the rest is the evidence.
``mandate`` is ALWAYS present, including on a stop: an exploration that was cut short still
carries the expert's seed approaches forward, because door 1 of § C.6 is a preservation rule
and not a reward for finishing. ``stop`` is what says the mandate is smaller than it might
have been, and why.
"""
mandate: Mandate
ledger_log: tuple[LedgerEntry, ...]
stop: ExplorationStop | None
plan_reviews: tuple[PlanReview, ...]
@dataclass(frozen=True)
class PlanReviewRequest:
"""What the reviewer is shown before the loop is allowed to run."""
index: int
plan: str
current_progress: str
is_stalled: bool
@dataclass(frozen=True)
class PlanReviewDecision:
"""The reviewer's answer. ``feedback is None`` means approve.
Two named constructors mirroring ``MagenticPlanReviewResponse.approve()/.revise()``, so a
persona or operator answering a review never has to import ``agent_framework`` the adapter
is one-directional and lives in exactly one place.
"""
feedback: str | None
@staticmethod
def approve() -> PlanReviewDecision:
return PlanReviewDecision(feedback=None)
@staticmethod
def revise(feedback: str) -> PlanReviewDecision:
if not feedback.strip():
raise ValueError("a revision must say what to revise; use approve() to sign off")
return PlanReviewDecision(feedback=feedback)
#: The synchronous HITL seam (U13). Given the request, answer it. Called in-process, so the
#: exploration blocks on it exactly as a human at a terminal would block the loop.
PlanReviewer = Callable[[PlanReviewRequest], PlanReviewDecision]
# ---------------------------------------------------------------------------------------------
# The tools. Level 1 of the three-guarantee table: real computation, ADVISORY verdicts.
# ---------------------------------------------------------------------------------------------
def _bundle_index(bundle_dirs: Sequence[str]) -> dict[str, str]:
"""Map each knowledge base's id to its directory. The id is the directory's BASENAME.
A duplicate basename is REFUSED rather than resolved by order: the id is what the manager
names a base by, and two bases answering to one name would let it read A while believing it
read B the S3.2 key-collision class, one layer up.
"""
index: dict[str, str] = {}
for raw in bundle_dirs:
bundle_id = Path(raw).name
if bundle_id in index:
raise ExplorationError(
f"two knowledge bases share the id {bundle_id!r} ({index[bundle_id]!r} and "
f"{raw!r}); the manager names a base by that id, so it must be unique"
)
index[bundle_id] = raw
return index
def _resolve_bundle(index: Mapping[str, str], bundle_id: str) -> str:
if bundle_id not in index:
known = ", ".join(sorted(index)) or "(none configured)"
raise ExplorationError(f"unknown knowledge base {bundle_id!r}; configured: {known}")
return index[bundle_id]
def navigator_tools(bundle_dirs: Sequence[str]) -> list[FunctionTool]:
"""The navigator's three tools: survey the catalogue, open one base, read one document.
Progressive disclosure, not stuffing (målbilde §2/§4): ``list_bundles`` never returns content,
only what each base IS and crucially whether it ships a ``cost-baseline.json``, because a base
without one cannot have its hypotheses reconciled against the project's real cost lines, and a
manager that does not know which bases are anchored cannot plan around it (§ C.7).
``read_bundle`` returns ``okf.bundle_context``, which EXCLUDES the ``type: verdict`` layer by
construction prior verdicts reach a hypothesis only through the gated ExpeL fold inside
``run_project``, never by being read as context here.
"""
index = _bundle_index(bundle_dirs)
@tool(
name="list_bundles",
description=(
"List the knowledge bases available to this exploration: id, what the index says the "
"base is about, how many prior expert verdicts it holds, and whether it ships a cost "
"baseline (without one, numbers cannot be reconciled against the project's own)."
),
)
def list_bundles() -> list[dict[str, Any]]:
catalogue: list[dict[str, Any]] = []
for bundle_id, bundle_dir in index.items():
bundle = okf.navigate_bundle(bundle_dir)
catalogue.append(
{
"id": bundle_id,
"index_summary": bundle.index_summary,
"verdict_count": len(bundle.verdicts),
# Tolerant on CONTENT, fail-fast on the PATH: an operator's bad directory is
# refused by navigate_bundle above, while a navigable base that simply has no
# baseline is legitimate (load_optional_cost_baseline's own contract).
"cost_baseline": okf.load_optional_cost_baseline(bundle_dir) is not None,
"skipped_links": [
{"from_file": s.from_file, "target": s.target, "reason": s.reason}
for s in bundle.skipped
],
}
)
return catalogue
@tool(
name="read_bundle",
description="Open ONE knowledge base by id and read its navigated context.",
)
def read_bundle(bundle_id: str) -> str:
return okf.bundle_context(okf.navigate_bundle(_resolve_bundle(index, bundle_id)))
@tool(
name="read_file",
description="Read ONE document inside a knowledge base, by base id and relative path.",
)
def read_file(bundle_id: str, path: str) -> str:
bundle_dir = _resolve_bundle(index, bundle_id)
# safe_resolve is the ONE in-/out-of-bundle test in this repo, and it is fail-closed. A
# model-chosen path is untrusted input by definition, so it goes through the same gate the
# navigation walk uses rather than a second, laxer check.
return Path(safe_resolve(bundle_dir, path)).read_text(encoding="utf-8")
return [list_bundles, read_bundle, read_file]
def quick_validate_tool(bundle_dirs: Sequence[str]) -> FunctionTool:
"""The hypothesiser's in-loop deterministic check — level 1, and advisory by construction.
It is the SAME ``validate_proposal`` against the SAME baseline the pipeline will use, so the
numbers it reports are real; what it never becomes is provenance. ``ProvenanceStamp`` is
written by ``run_project`` alone, and nothing here reaches it.
Measured cheap (S5: 13.6 ms median through stage 0 + the CBC solve + a 512-sample Monte Carlo,
on an anchored base with assumption bands), so it may be called freely in the loop and the
contract carries no latency budget.
``anchored`` is reported alongside the verdict for the same reason
``ProvenanceStamp.cost_baseline_anchored`` is a required field: a verdict reached without the
project's own cost lines is a weaker claim, and one that does not say so is the silence S4.0's
visibility work closed.
"""
index = _bundle_index(bundle_dirs)
@tool(
name="quick_validate",
description=(
"Run the deterministic validator over a candidate, as an ADVISORY check. Pass the "
"knowledge base id and the candidate as IR JSON. The verdict is not the project's "
"decision — the pipeline re-validates and stamps."
),
)
def quick_validate(bundle_id: str, proposal_json: str) -> dict[str, Any]:
bundle_dir = _resolve_bundle(index, bundle_id)
baseline = okf.load_optional_cost_baseline(bundle_dir)
try:
proposal = SavingsProposal.model_validate_json(proposal_json)
except ValidationError as exc:
return {"decision": "unparseable", "reason": str(exc), "anchored": baseline is not None}
outcome = validate_proposal(proposal, baseline=baseline)
if isinstance(outcome, Rejection):
return {
"decision": "rejected",
"reason": outcome.reason,
"anchored": baseline is not None,
}
return {
"decision": "validated",
"reason": "",
"anchored": baseline is not None,
"p10": outcome.p10,
"p50": outcome.p50,
"p90": outcome.p90,
}
return quick_validate
# ---------------------------------------------------------------------------------------------
# The workflow. One builder, one build, ONE run per exploration (B7 / § C.4).
# ---------------------------------------------------------------------------------------------
def fresh_exploration_workflow(
client_factory: Callable[[str], BaseChatClient],
*,
contract: ExplorationContract,
bundle_dirs: Sequence[str] = (),
middleware: Sequence[Any] | None = None,
) -> Any:
"""Build a FRESH Magentic workflow with fresh agents and fresh clients (mirrors
``workflow.fresh_workflow``).
**Fresh builder per exploration, never a reused one.** A built Magentic workflow is
single-use: a second ``run()`` on it raises ``RuntimeError`` having made zero model calls
(measured, E1) which is a safe failure, unlike GroupChat 1.9.0's silently empty second run.
Reuse is prevented here rather than relied upon to fail.
``manager_agent_factory=`` rather than ``manager_agent=``: the eager form constructs the
manager once and hands the same instance to every ``build()`` (``:1683``, ``:1729-1730``),
which on orchestrations 1.0.0 leaked one exploration's task into the next. Orchestrations
1.0.1 removed the manager's persistent session, so that leak is GONE and **no test here can
tell the two forms apart today** (measured, § F A8: 4/4 0/5). The factory form is used
anyway because it costs nothing and does not depend on an upstream regression fix staying
fixed but it is stated plainly rather than gated, because a gate that cannot go red proves
nothing.
Every agent carries the SAME ``middleware`` list, manager included. That is the whole of the
token guarantee: agent-level ``ChatMiddleware`` does fire on the manager's own calls (measured
A1), and the manager is the most talkative participant in the loop.
"""
hypothesiser_tools: list[Any] = [quick_validate_tool(bundle_dirs)]
tools_by_role: dict[str, list[Any]] = {
NAVIGATOR_ROLE: list(navigator_tools(bundle_dirs)),
HYPOTHESISER_ROLE: hypothesiser_tools,
}
participants = [
Agent(
client_factory(role),
_INSTRUCTIONS[role],
name=role,
description=_INSTRUCTIONS[role],
tools=tools_by_role[role],
middleware=middleware,
)
for role in PARTICIPANT_ROLES
]
def _manager_agent() -> Agent:
return Agent(
client_factory(MANAGER_ROLE),
"You plan and coordinate an exploration of a project's knowledge bases to find "
"cost-saving directions worth testing. You never state a saving figure yourself.",
name=MANAGER_ROLE,
description="plans the exploration",
middleware=middleware,
)
return MagenticBuilder(
participants=participants,
manager_agent_factory=_manager_agent,
max_round_count=contract.max_rounds,
max_stall_count=contract.max_stall_count,
max_reset_count=contract.max_reset_count,
enable_plan_review=contract.enable_plan_review,
).build()
def _truthy(answer: Any) -> bool:
"""Read a ledger answer exactly as the orchestrator does.
``MagenticProgressLedgerItem.answer`` is ``str | bool`` and is NOT normalised per field
(``:292``); the orchestrator then tests it with plain Python truthiness (``:1106``, ``:1112``).
A smarter coercion here treating the string ``"false"`` as false, say would describe a run
the orchestrator did not have. This layer reports what the loop DID, so it copies the loop's
own rule rather than improving on it.
"""
return bool(answer)
def _plan_text(content: Any) -> str:
"""The task-ledger plan as text. ``PLAN_CREATED``/``REPLANNED`` carry a ``Message``."""
return str(getattr(content, "text", "") or "")
def _absorb(
result: Any,
*,
ledger_log: list[LedgerEntry],
hypotheses: list[str],
seen: set[int],
span: Any,
) -> int:
"""Fold one ``run()``'s events into the running log and onto the trace. Returns this call's
replans.
``seen`` de-duplicates by object identity because MAF may surface the same payload under more
than one event type; the alternative pinning one event-type name would silently under- or
over-count if that changed. Participant output is matched on ``executor_id``, which is the
agent's name, exactly as ``run._authored_texts`` matches the debate's ``author_name``.
**The three span events land here** (U14's deferred half, now that it has a call site). They
carry the manager's DECISIONS — which plan, who was asked, whether the request was satisfied —
as separate typed attributes rather than a rendered sentence, so a collector can query them.
Stated honesty limit: they are recorded as the event stream is folded, i.e. after ``run()``
returns, so their ORDER is faithful and their timestamps are not the moments the manager acted.
Emitting live would mean driving the loop through ``run_stream``, which is a different
measurement from the one the plan-review round trip was proved against.
"""
replans = 0
for event in result:
data = getattr(event, "data", None)
if data is None or id(data) in seen:
continue
event_type = getattr(data, "event_type", None)
if event_type is not None:
seen.add(id(data))
if event_type == MagenticOrchestratorEventType.PLAN_CREATED:
span.add_event("plan_created", {"plan": _plan_text(data.content)})
elif event_type == MagenticOrchestratorEventType.REPLANNED:
replans += 1
span.add_event("replanned", {"plan": _plan_text(data.content)})
elif event_type == MagenticOrchestratorEventType.PROGRESS_LEDGER_UPDATED:
ledger = data.content
speaker = str(ledger.next_speaker.answer)
entry = LedgerEntry(
round_index=len(ledger_log) + 1,
is_request_satisfied=_truthy(ledger.is_request_satisfied.answer),
is_in_loop=_truthy(ledger.is_in_loop.answer),
is_progress_being_made=_truthy(ledger.is_progress_being_made.answer),
next_speaker=speaker,
instruction_or_question=str(ledger.instruction_or_question.answer),
speaker_known=speaker in PARTICIPANT_ROLES,
)
ledger_log.append(entry)
span.add_event(
"progress_ledger_updated",
{
"round_index": entry.round_index,
"is_request_satisfied": entry.is_request_satisfied,
"is_in_loop": entry.is_in_loop,
"is_progress_being_made": entry.is_progress_being_made,
"next_speaker": entry.next_speaker,
"speaker_known": entry.speaker_known,
"instruction_or_question": entry.instruction_or_question,
},
)
continue
if getattr(data, "executor_id", None) == HYPOTHESISER_ROLE:
response = getattr(data, "agent_response", None)
text = getattr(response, "text", None)
if text:
seen.add(id(data))
hypotheses.append(text)
return replans
class HypothesisParseError(ExplorationError):
"""A hypothesiser turn carried the marker and then something unreadable.
Fail-closed, and the marker is what makes that affordable: most turns legitimately are not
hypotheses, so an unmarked turn is not a failure and there is nothing to be silent about. A
MARKED line that will not parse is a claim the loop tried to make and could not raising it
is the ``write_concept_file`` rule (validation, never repair) rather than the tolerant RAW
inbox rule, because this is the product of the run, not a folder anyone may drop things in.
"""
def _parse_hypotheses(texts: Sequence[str]) -> list[tuple[str, str]]:
"""Every marked ``(label, rationale)`` pair the hypothesiser committed to, in turn order."""
found: list[tuple[str, str]] = []
for text in texts:
for line in text.splitlines():
stripped = line.strip()
if not stripped.startswith(HYPOTHESIS_MARKER):
continue
payload = stripped[len(HYPOTHESIS_MARKER) :].strip()
try:
data = json.loads(payload)
except json.JSONDecodeError as exc:
raise HypothesisParseError(
f"hypothesiser marked a line as a hypothesis and it is not JSON: {exc}; "
f"line was: {stripped}"
) from exc
if not isinstance(data, dict) or not data.get("label") or not data.get("rationale"):
raise HypothesisParseError(
"a marked hypothesis needs a non-empty 'label' and 'rationale'; got: "
f"{stripped}"
)
found.append((str(data["label"]), str(data["rationale"])))
return found
def _mint_approaches(
seeds: Sequence[Approach], discovered: Sequence[tuple[str, str]]
) -> tuple[Approach, ...]:
"""Seeds FIRST, untouched, then one approach per discovered direction.
Seeds lead because the pipeline evaluates approaches in order under one shared budget: a
direction the domain expert asked for must be reached before the loop's own findings spend
what is left. The minted ids skip anything a seed already claims, so a seed literally named
``hypothesis-1`` cannot collide ``Mandate`` refuses duplicate ids at construction, and
turning an expert's naming choice into a hard failure would be a refusal that teaches nothing.
"""
taken = {approach.id for approach in seeds}
minted: list[Approach] = list(seeds)
counter = 0
for label, rationale in discovered:
counter += 1
candidate = f"hypothesis-{counter}"
while candidate in taken or candidate == OWN_PROPOSAL_ID:
counter += 1
candidate = f"hypothesis-{counter}"
taken.add(candidate)
# description is the hypothesiser's own words, VERBATIM: ``generate._build_messages`` feeds
# an Approach.description to the proposer unchanged, and the reason a direction is worth
# trying is exactly the half a model cannot re-derive from the cost table.
minted.append(Approach(id=candidate, label=label, description=rationale))
return tuple(minted)
def _pending_plan_reviews(result: Any) -> list[Any]:
return [event for event in result if event.type == "request_info"]
async def explore(
prompt: str,
*,
contract: ExplorationContract,
bundle_dirs: Sequence[str] = (),
profile: Profile | str = Profile.LOCAL,
client_factory: Callable[[str], BaseChatClient] | None = None,
seed_approaches: Sequence[Approach] = (),
plan_reviewer: PlanReviewer | None = None,
meter: TokenMeter | None = None,
success_criteria: str = "",
) -> ExplorationResult:
"""Explore the knowledge bases and return the ``Mandate`` the pipeline should evaluate.
Writes NOTHING. No outbox artefact, no wiki promotion, no verdict level 3 of the guarantee
table belongs to ``run_project`` alone, and an exploration that could write would be a route
around the gate that makes an answer checkable.
**Three stops, three channels, deliberately not one.** Tokens raise ``BudgetExceeded`` exactly
as the debate does, so the hosted surface's 429 arm needs no new case. Rounds raise it too,
with ``kind="exploration_rounds"`` the orchestration itself raises nothing at its round cap
(measured, § F): it returns a canonical assistant message that is indistinguishable from
success at the transport, so this layer produces the typed stop. Everything semantic
stalling out, an exhausted revision cap, a ledger naming nobody is a VALUE in ``stop``,
because those are outcomes of the exploration rather than exhaustion of a resource, and S3.4
split those two apart for a reason.
``seed_approaches`` are the domain expert's own hypotheses (door 1 of § C.6). They are in the
returned mandate whatever the loop found, including when the loop found nothing and including
when it stopped early.
Honesty limits, stated rather than implied. (1) A run is bound to ONE mandate; multi-base
dispatch (``Approach.bundle_id``, § C.7) waits for ``run_project`` to accept more than one
``bundle_dir``, and shipping the field before its consumer would be a shape guessed instead of
measured. (2) The ``quick_validate`` verdicts the hypothesiser saw are not in
``ExplorationResult``: they are level-1 advisory, and their home is the
``{run_id}-exploration.json`` artefact the CLI wiring writes. (3) The exploration roles resolve
through ``resolve_model``'s ``default`` fallback unless an operator maps them explicitly.
"""
if contract.enable_plan_review and plan_reviewer is None:
raise ExplorationError(
"enable_plan_review is set but no plan_reviewer was given: the exploration would stop "
"at a review nobody can answer, which is a hang rather than a result"
)
if plan_reviewer is not None and not contract.enable_plan_review:
raise ExplorationError(
"a plan_reviewer was given but enable_plan_review is false, so no review is ever "
"requested and the reviewer would never be called (refused, never silently ignored)"
)
if meter is None:
meter = TokenMeter(Budget(max_tokens=contract.max_tokens, max_rounds=contract.max_rounds))
if client_factory is None:
# Imported HERE, not at module scope: ``run`` imports this module for its ``--explore``
# door, so a top-level import would be circular. One copy of the backend factory, never a
# second (the (p) rule) — the same move ``run.main`` makes for ``scripted_factory``.
from portfolio_optimiser.run import _default_factory
client_factory = _default_factory(profile)
workflow = fresh_exploration_workflow(
client_factory,
contract=contract,
bundle_dirs=bundle_dirs,
middleware=[BudgetMiddleware(meter)],
)
ledger_log: list[LedgerEntry] = []
hypothesis_texts: list[str] = []
plan_reviews: list[PlanReview] = []
seen: set[int] = set()
replans = 0
stop: ExplorationStop | None = None
# ONE span for the whole exploration, opened before the first model call and closed however
# the loop ends — including on a BudgetExceeded, which the span records rather than swallows.
# With no provider installed this is a no-op tracer and every event is discarded, which is
# exactly what "tracing is off" has meant since U14.
with exploration_tracer().start_as_current_span(EXPLORATION_SPAN) as span:
result = await workflow.run(prompt)
while True:
replans += _absorb(
result,
ledger_log=ledger_log,
hypotheses=hypothesis_texts,
seen=seen,
span=span,
)
pending = _pending_plan_reviews(result)
if not pending:
break
assert plan_reviewer is not None # guarded above; the review implies a reviewer
request = pending[0]
review = request.data
decision = plan_reviewer(
PlanReviewRequest(
index=len(plan_reviews),
plan=str(review.plan),
current_progress=str(review.current_progress),
is_stalled=_truthy(review.is_stalled),
)
)
if decision.feedback is None:
plan_reviews.append(
PlanReview(
index=len(plan_reviews),
plan=str(review.plan),
is_stalled=_truthy(review.is_stalled),
decision="approve",
)
)
response = MagenticPlanReviewResponse.approve()
else:
# The revision is recorded whether or not it is APPLIED: ``plan_reviews`` is the
# record of what the reviewer DECIDED, and ``stop`` is what says the last one was
# refused.
applied = sum(1 for entry in plan_reviews if entry.decision == "revise")
plan_reviews.append(
PlanReview(
index=len(plan_reviews),
plan=str(review.plan),
is_stalled=_truthy(review.is_stalled),
decision="revise",
feedback=decision.feedback,
)
)
if applied >= contract.max_plan_revisions:
# Measured (§ F, A3): a revise costs two manager calls, emits no progress
# ledger and consumes no round, then asks AGAIN. Under the round cap alone this
# loop never terminates. Stopping is the honest move — forcing an approve the
# reviewer did not give would be repair, and repair of a human's decision most
# of all.
stop = "plan_revisions_exhausted"
break
response = MagenticPlanReviewResponse.revise(decision.feedback)
result = await workflow.run(responses={request.request_id: response})
if stop is None:
stop = _classify_stop(ledger_log, replans=replans, contract=contract)
discovered = _parse_hypotheses(hypothesis_texts) if stop != "unknown_speaker" else []
return ExplorationResult(
mandate=Mandate(
objective=prompt,
approaches=_mint_approaches(seed_approaches, discovered),
allow_own_proposals=True,
success_criteria=success_criteria,
),
ledger_log=tuple(ledger_log),
stop=stop,
plan_reviews=tuple(plan_reviews),
)
def _classify_stop(
ledger_log: Sequence[LedgerEntry], *, replans: int, contract: ExplorationContract
) -> ExplorationStop | None:
"""Decide what a finished ``run()`` actually was. Structural, never prose-matched.
The orchestration ends four ways and reports three of them with an ordinary-looking result:
a satisfied request, the round cap, the reset cap, and a ``next_speaker`` matching nobody.
Only the first is success. The discriminators used here are counts and ledger flags the
termination MESSAGE is deliberately not matched, because it is an f-string built inline
(``:1253``) with no constant to pin against, and a model's own final answer could contain the
same words.
A ledger naming nobody is checked FIRST and outranks everything else. It is the only ending
in which the orchestrator produced an answer having asked no participant at all (``:1128``),
so the run has a result that no work stands behind the same class as the retired E2 finding,
and not something a later, softer verdict should be allowed to paper over.
"""
if any(not entry.speaker_known for entry in ledger_log):
return "unknown_speaker"
if ledger_log and ledger_log[-1].is_request_satisfied:
return None
if len(ledger_log) >= contract.max_rounds:
raise BudgetExceeded("exploration_rounds", contract.max_rounds, len(ledger_log))
# Whatever is left ended without satisfying the request and without exhausting the round cap:
# the manager reset until it ran out of resets, or terminated with nothing to show. Both are
# the same fact for a caller — the exploration gave up — so they share one token rather than
# inventing a distinction the ledger cannot support.
return "stalled"

View file

@ -37,10 +37,14 @@ cause.
(``opentelemetry-exporter-otlp-proto-grpc`` / ``-http``) are not declared dependencies. They are
egress, they drag grpc and protobuf into a published wheel for a mode that is off by default, and
MAF already raises an ``ImportError`` that names the package to install. Stated honesty limit:
``PORTFOLIO_OTEL=otlp`` works only after the operator installs one of them. Equally absent are the
``PLAN_CREATED`` / ``REPLANNED`` / ``PROGRESS_LEDGER_UPDATED`` events the plan names they belong
to the exploration loop (U4), which does not exist yet, and an emitter written before its call site
is a shape guessed rather than measured.
``PORTFOLIO_OTEL=otlp`` works only after the operator installs one of them.
**The ``PLAN_CREATED`` / ``REPLANNED`` / ``PROGRESS_LEDGER_UPDATED`` events now exist** (U4, økt
56). They were held back here in økt 55 on the ground that an emitter written before its call site
is a shape guessed rather than measured; the call site is ``explore._absorb``, and the events are
recorded on the exploration span this module's ``exploration_tracer`` hands out. Nothing about the
contract above changed: with tracing off there is no provider, so those events are discarded like
every other span this process makes.
MAF-touching by construction, so this module never enters the framework-neutral context layer
(``okf.py``); the ``test_okf_is_maf_free`` guard keeps that boundary.
@ -233,3 +237,29 @@ def tracing_notice(setup: TracingSetup) -> str | None:
f" Tracing: {TRACING_ENV}={MODE_OTLP} — OpenTelemetry spans are EXPORTED OVER THE NETWORK "
f"to the endpoints declared below\n{rows}"
)
#: The instrumentation scope every exploration span is created under. One name, so a collector
#: can select this framework's own spans apart from MAF's (``invoke_agent``, ``workflow.run``)
#: without matching on span names that MAF owns and may rename.
EXPLORATION_TRACER_NAME: Final = "portfolio_optimiser.explore"
def exploration_tracer() -> Any:
"""The tracer the exploration loop records its decisions on.
``get_tracer`` is safe to call whether or not a provider was installed: with none, OpenTelemetry
hands back a no-op tracer and every span and event is discarded. That is the SAME shape MAF's
own instrumentation already has (``ENABLE_INSTRUMENTATION`` defaults to True and its spans are
thrown away for want of a provider), and it is what lets the exploration emit unconditionally.
Gating emission on ``PORTFOLIO_OTEL`` would be a second resolution of a rule this module owns,
free to disagree with the providers actually installed.
A FUNCTION rather than a module-level tracer, and the reason is ordering: ``configure_tracing``
runs at process startup, and a tracer bound at import time would have been taken from the
global provider that existed BEFORE it a no-op one, permanently. It is also the seam a test
substitutes a local provider through, without installing anything globally.
"""
from opentelemetry import trace
return trace.get_tracer(EXPLORATION_TRACER_NAME)

View file

@ -0,0 +1,990 @@
"""U4 + U13-synchronous (økt 56) — the Magentic exploration loop as a MANDATE-FORMER.
**What this loop is, and what it deliberately is not.** ``explore()`` puts a Magentic manager
*over* the normative pipeline, never inside it: the manager is free to choose which knowledge base
to open and which hypothesis to shape next, and what leaves that freedom is a
``mandate.Mandate`` a list of approaches worth *testing*. It is never a proposal. Every number
that survives is still gated by ``validate_proposal`` inside ``run_project``, in the same blocking
gate as today, and the exploration itself can write to neither the outbox nor the wiki. Step 3's
maker-checker debate is untouched (``shared/method-spec.md`` §3 is commons-owned and normative).
**Everything asserted here was measured before it was built** (plan
``docs/plan/2026-08-23-magentic-utforskningssloeyfe.md`` § F, spikes S0S6 in økt 54, plus three
probes run at the head of økt 56):
* a plan-review ``revise`` costs two manager calls, **zero** rounds, and asks *again* so an
always-revising expert is unbounded spend under a round cap that never ticks. That is the whole
reason ``max_plan_revisions`` is a required contract field rather than a nicety.
* the round cap and the reset cap **raise nothing**. Both end the run with a canonical assistant
message and a normal-looking result (measured: ``max_round_count=2`` two ledger events and
``'Workflow terminated due to reaching maximum round count.'``; a stalling ledger with
``max_reset_count=1`` one ``REPLANNED`` event and ``'…maximum reset count.'``). At the
transport both are indistinguishable from success, so this layer produces the typed stop itself.
* a ``next_speaker`` naming nobody produces a **silent final answer with zero participant work**
(``_magentic.py:1128-1131``) a plausible answer produced by no work at all, which is the
hazard class E2 was retired for. The names are therefore validated, never assumed.
**The client is the repo's own ``ScriptedChatClient``.** A bare ``BaseChatClient`` silently no-ops
``BudgetMiddleware`` (measured, ``simulation.py:373-375``), so a budget claim proved against one
would prove nothing.
"""
from __future__ import annotations
import asyncio
import json
import os
import shutil
import subprocess
import sys
from collections.abc import Callable
from pathlib import Path
from typing import Any
import pytest
from agent_framework import BaseChatClient
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from pydantic import ValidationError
import portfolio_optimiser
from portfolio_optimiser import explore, okf
from portfolio_optimiser.budget import Budget, BudgetExceeded, TokenMeter
from portfolio_optimiser.explore import ExplorationContract
from portfolio_optimiser.mandate import Approach
from portfolio_optimiser.simulation import ScriptedChatClient
# ---------------------------------------------------------------------------------------------
# C.3 — the contract: an exploration without stated bounds refuses to start
# ---------------------------------------------------------------------------------------------
_FULL_CONTRACT = {
"max_rounds": 4,
"max_tokens": 5_000,
"max_stall_count": 2,
"max_reset_count": 1,
"max_plan_revisions": 1,
"enable_plan_review": True,
}
@pytest.mark.parametrize("omitted", sorted(_FULL_CONTRACT))
def test_every_bound_is_required_with_no_default(omitted: str) -> None:
"""T1: each of the six fields is REQUIRED — dropping any one refuses construction.
Not a style point. ``MagenticBuilder`` defaults ``max_round_count`` to ``None`` (unbounded)
and ``max_reset_count`` to ``None`` (unlimited), and inheriting either would give this repo
the one thing ``shared/method-spec.md`` §8 forbids outright: a loop with no stated end. A
default here would also be a claim about the operator's intent that nobody made — the same
ground on which ``ProvenanceStamp.cost_baseline_anchored`` is required without one.
"""
payload = {k: v for k, v in _FULL_CONTRACT.items() if k != omitted}
with pytest.raises(ValidationError):
ExplorationContract(**payload)
def test_full_contract_constructs() -> None:
"""T2: the control for T1 — the complete payload IS valid.
Without it, T1 would pass on a model that refuses everything, which is the vacuous-gate class
this repo has paid for six times.
"""
contract = ExplorationContract(**_FULL_CONTRACT)
assert contract.max_rounds == 4
assert contract.enable_plan_review is True
def test_a_revision_cap_without_plan_review_is_refused_not_ignored() -> None:
"""T3: ``max_plan_revisions > 0`` with ``enable_plan_review=False`` refuses.
A plan revision can only arise from a plan review with the review off, the cap bounds an
event that cannot occur, and a caller who set it believes they bounded something. This repo
refuses a setting that cannot take effect rather than dropping it silently (the same partition
``--embedder-config requires --semantic-retrieval`` enforces on the CLI).
"""
with pytest.raises(ValidationError):
ExplorationContract(**{**_FULL_CONTRACT, "enable_plan_review": False})
def test_review_off_with_zero_revisions_is_the_coherent_form() -> None:
"""T4: the control for T3 — review off and the cap at ``0`` is a consistent statement, and
must construct. Without this arm T3 would pass on a model that simply forbade
``enable_plan_review=False`` outright, which is a different (and wrong) rule.
"""
contract = ExplorationContract(
**{**_FULL_CONTRACT, "enable_plan_review": False, "max_plan_revisions": 0}
)
assert contract.enable_plan_review is False
assert contract.max_plan_revisions == 0
# ---------------------------------------------------------------------------------------------
# The scripted stand-ins. ScriptedChatClient, never a bare BaseChatClient: the latter no-ops
# BudgetMiddleware (measured, simulation.py:373-375), so a budget assertion made against one
# would assert nothing.
# ---------------------------------------------------------------------------------------------
PROMPT = "Find the cheapest saving available in the energy bundle."
def _ledger_json(
*, satisfied: bool, speaker: str, instruction: str = "Shape one hypothesis."
) -> str:
"""A progress ledger naming ``speaker``.
The name is a PARAMETER, never a literal, because a ``next_speaker`` matching no participant
is the measured footgun this module defends against: the orchestrator does not error, it
quietly emits a final answer having asked nobody (``_magentic.py:1128-1131``).
"""
return json.dumps(
{
"is_request_satisfied": {"reason": "r", "answer": satisfied},
"is_in_loop": {"reason": "r", "answer": False},
"is_progress_being_made": {"reason": "r", "answer": True},
"next_speaker": {"reason": "r", "answer": speaker},
"instruction_or_question": {"reason": "r", "answer": instruction},
}
)
def _stalling_ledger_json(speaker: str) -> str:
"""A ledger reporting NO progress and a loop — the two flags that drive ``stall_count`` up."""
return json.dumps(
{
"is_request_satisfied": {"reason": "r", "answer": False},
"is_in_loop": {"reason": "circles", "answer": True},
"is_progress_being_made": {"reason": "none", "answer": False},
"next_speaker": {"reason": "r", "answer": speaker},
"instruction_or_question": {"reason": "r", "answer": "Try again."},
}
)
def _manager_script(
ledgers: list[str], calls: list[str] | None = None
) -> Callable[[str, str], str]:
"""Route a manager prompt blob to its scripted reply, consuming ``ledgers`` in order.
The ORDER of these tests is load-bearing and was measured (§ F, A6): the selector receives the
CONCATENATION of every message in the call, so a later-stage prompt still carries the earlier
stage's text — one manager call in five carries two markers. Testing the later stage FIRST is
what resolves it; reversing two of these silently reattributes a reply to the wrong stage.
"""
def _select(blob: str, _role: str) -> str:
if calls is not None:
calls.append(blob[:40])
if "provide the final answer" in blob:
return "FINAL: exploration done."
if "pure JSON format" in blob:
return ledgers.pop(0) if ledgers else _ledger_json(satisfied=True, speaker="navigator")
if "went wrong on this last run" in blob:
return "PLAN-UPDATE: revised plan."
if "rewrite the following fact sheet" in blob:
return "FACTS-UPDATE: revised facts."
if "bullet-point plan" in blob:
return "PLAN: - ask the hypothesiser"
if "pre-survey" in blob:
return "FACTS: the bundle is anchored."
return "{}"
return _select
def _factory(
*, ledgers: list[str], hypothesiser: list[str], navigator: str = "NAVIGATOR: index read."
) -> Callable[[str], BaseChatClient]:
"""One fresh ``ScriptedChatClient`` per role, exactly as the real factory hands out one per
role. ``hypothesiser`` is a list consumed in order, so a run can shape several candidates."""
def factory(role: str) -> BaseChatClient:
if role == explore.MANAGER_ROLE:
return ScriptedChatClient(reply_selector=_manager_script(ledgers), role=role)
if role == explore.HYPOTHESISER_ROLE:
replies = list(hypothesiser)
def _hyp(_blob: str, _role: str) -> str:
return replies.pop(0) if replies else "nothing further."
return ScriptedChatClient(reply_selector=_hyp, role=role)
return ScriptedChatClient(navigator, role=role)
return factory
def _hypothesis_line(label: str, rationale: str) -> str:
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps({"label": label, "rationale": rationale})
#: The no-review base every stop test derives from. ``enable_plan_review`` and
#: ``max_plan_revisions`` move together — ``ExplorationContract`` refuses them apart — so a test
#: about round or stall behaviour has to say so explicitly rather than inherit ``_FULL_CONTRACT``.
_NO_REVIEW = {**_FULL_CONTRACT, "enable_plan_review": False, "max_plan_revisions": 0}
_CONTRACT = ExplorationContract(
max_rounds=6,
max_tokens=100_000,
max_stall_count=2,
max_reset_count=1,
max_plan_revisions=0,
enable_plan_review=False,
)
# ---------------------------------------------------------------------------------------------
# C.0 / C.6 — the exploration is a MANDATE-FORMER, and a seed never disappears
# ---------------------------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_hypotheses_become_the_mandate_in_the_order_they_were_shaped() -> None:
"""T5: what the hypothesiser MARKED becomes ``Mandate.approaches``, rationale VERBATIM.
The rationale is the half a model cannot re-derive from cost data ``mandate.Approach``
already feeds ``description`` to the proposer verbatim (``generate._build_messages``), so
paraphrasing it here would drop precisely the part the exploration exists to carry forward.
"""
ledgers = [
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
]
result = await explore.explore(
PROMPT,
contract=_CONTRACT,
bundle_dirs=(),
client_factory=_factory(
ledgers=ledgers,
hypothesiser=[
"Looking at the bundle.\n"
+ _hypothesis_line("LED retrofit", "the fixtures are 1990s fluorescent")
],
),
)
assert [a.label for a in result.mandate.approaches] == ["LED retrofit"]
assert result.mandate.approaches[0].description == "the fixtures are 1990s fluorescent"
assert result.mandate.objective == PROMPT
assert result.stop is None
@pytest.mark.asyncio
async def test_a_seed_approach_survives_whatever_the_manager_found() -> None:
"""T6: an expert's own hypothesis is in the output mandate, FIRST, untouched.
Door 1 of § C.6, and the ``not_evaluated`` rule applied one stage earlier: a direction the
domain expert asked for may never vanish because an autonomous loop preferred its own. Seeds
lead so the pipeline reaches them before spending its budget on discovered ones.
"""
seed = Approach(id="fagperson-1", label="Night setback", description="the expert's own words")
ledgers = [
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
]
result = await explore.explore(
PROMPT,
contract=_CONTRACT,
bundle_dirs=(),
seed_approaches=(seed,),
client_factory=_factory(
ledgers=ledgers,
hypothesiser=[_hypothesis_line("LED retrofit", "fluorescent fixtures")],
),
)
assert [a.id for a in result.mandate.approaches] == ["fagperson-1", "hypothesis-1"]
assert result.mandate.approaches[0] == seed
@pytest.mark.asyncio
async def test_an_exploration_that_shaped_nothing_still_returns_the_seeds() -> None:
"""T7: the control for T6 — with the hypothesiser silent, the seed is still the mandate.
This is what makes T6 a statement about PRESERVATION rather than about ordering: a test that
only ever saw seeds alongside discoveries could not tell "seeds are kept" from "seeds sort
first".
"""
seed = Approach(id="fagperson-1", label="Night setback", description="the expert's own words")
result = await explore.explore(
PROMPT,
contract=_CONTRACT,
bundle_dirs=(),
seed_approaches=(seed,),
client_factory=_factory(
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
hypothesiser=[],
),
)
assert result.mandate.approaches == (seed,)
assert result.mandate.allow_own_proposals is True
def test_zero_resets_is_refused_because_it_silently_explores_nothing() -> None:
"""T8: ``max_reset_count=0`` refuses — MEASURED, not reasoned.
The orchestrator's limit check is ``reset_count >= max_reset_count`` (``_magentic.py:1243``)
and ``reset_count`` starts at zero, so a cap of zero is already met before the first round.
Measured against the installed stack: the run makes only the ``facts`` and ``plan`` manager
calls, emits **zero** progress-ledger events, and returns
``'Workflow terminated due to reaching maximum reset count.'`` an exploration that explored
nothing, reported as a stall that never happened. An operator writing "allow no resets" would
get "do no work", quietly. So it is refused at construction, where the reason can be said.
"""
with pytest.raises(ValidationError):
ExplorationContract(**{**_FULL_CONTRACT, "max_reset_count": 0})
def test_zero_stalls_is_allowed_because_it_means_something() -> None:
"""T9: the control for T8 — ``max_stall_count=0`` is a real setting and must construct.
The stall check is STRICT (``stall_count > max_stall_count``, ``:1118``) and the counter is
incremented before it, so zero means "reset on the first round that reports no progress".
That is strictness, not self-defeat, and refusing both zeroes on symmetry would have banned a
usable configuration on the strength of a measurement about a different field.
"""
contract = ExplorationContract(**{**_FULL_CONTRACT, "max_stall_count": 0})
assert contract.max_stall_count == 0
# ---------------------------------------------------------------------------------------------
# C.2 / C.3 — three endings the orchestration reports as if they were success
# ---------------------------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_the_round_cap_leaves_as_a_typed_budget_stop() -> None:
"""T10: the round cap becomes ``BudgetExceeded(kind="exploration_rounds")``.
Measured (§ F, E5, re-measured at the head of this økt): ``max_round_count`` raises NOTHING.
The run ends with the assistant message ``'Workflow terminated due to reaching maximum round
count.'`` and a result that ``get_outputs()`` answers like any other — at the transport it is
indistinguishable from a finished exploration. Left alone, a caller would read a run that
explored two rounds of a six-round question as a completed answer. The triple is the one
-(y) defends: WHICH cap bound, what it was, and how far the run actually got.
"""
contract = ExplorationContract(**{**_NO_REVIEW, "max_rounds": 2})
with pytest.raises(BudgetExceeded) as excinfo:
await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
client_factory=_factory(
ledgers=[_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE)] * 4,
hypothesiser=[_hypothesis_line("LED", "worth a look")] * 4,
),
)
assert excinfo.value.kind == "exploration_rounds"
assert excinfo.value.limit == 2
assert excinfo.value.observed == 2
@pytest.mark.asyncio
async def test_a_request_satisfied_on_the_last_allowed_round_is_success() -> None:
"""T11: the discriminator for T10 — reaching the cap is not the same as being cut off by it.
Both runs end with exactly ``max_rounds`` progress-ledger events, so a check written on the
count alone would raise on this one too and turn a completed exploration into a budget error.
What separates them is the LAST ledger's ``is_request_satisfied``, which is also what the
orchestrator itself branches on (``:1106``). Without this arm, T10 would pass on an
implementation that refuses every exploration that uses its whole allowance.
"""
contract = ExplorationContract(**{**_NO_REVIEW, "max_rounds": 2})
result = await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
client_factory=_factory(
ledgers=[
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
],
hypothesiser=[_hypothesis_line("LED", "worth a look")],
),
)
assert len(result.ledger_log) == contract.max_rounds
assert result.stop is None
assert [a.label for a in result.mandate.approaches] == ["LED"]
@pytest.mark.asyncio
async def test_stalling_out_is_a_typed_value_never_an_exception() -> None:
"""T12: stall → reset → out of resets is ``stop="stalled"``, and the run still returns.
Kept as a VALUE while the round cap RAISES, and the split is S3.4's, not a preference: a
stalled exploration is an outcome (the manager tried and got nowhere), whereas an exhausted
round or token cap is resource exhaustion. Fusing them would leave a caller unable to tell
"there was nothing here" from "we could not afford to look".
"""
contract = ExplorationContract(
**{**_NO_REVIEW, "max_rounds": 6, "max_stall_count": 1, "max_reset_count": 1}
)
result = await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
client_factory=_factory(
ledgers=[_stalling_ledger_json(explore.HYPOTHESISER_ROLE)] * 6,
hypothesiser=["still nothing."] * 6,
),
)
assert result.stop == "stalled"
assert len(result.ledger_log) < contract.max_rounds
assert all(entry.is_in_loop for entry in result.ledger_log)
@pytest.mark.asyncio
async def test_a_ledger_naming_nobody_withholds_what_the_run_produced() -> None:
"""T13: a ``next_speaker`` matching no participant stops the exploration and drops its finds.
The measured footgun (``_magentic.py:1128-1131``): the orchestrator neither raises nor retries
on an unknown speaker it logs a warning and jumps to ``_prepare_final_answer``. The run
therefore returns a plausible answer that no participant was asked for. This is the shape E2
was retired over ("a plausible verdict produced by zero work"), so the mandate is NOT built
from what such a run said it found.
The scripted run reaches the bad ledger on round TWO, after a good round in which the
hypothesiser really did commit to a direction. That ordering is what makes the assertion
sharp: with the bad ledger first, nobody would ever have spoken and "nothing was carried
forward" would be true of any implementation at all.
"""
seed = Approach(id="fagperson-1", label="Night setback", description="expert's own")
result = await explore.explore(
PROMPT,
contract=_CONTRACT,
bundle_dirs=(),
seed_approaches=(seed,),
client_factory=_factory(
ledgers=[
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
_ledger_json(satisfied=False, speaker="a-name-nobody-answers-to"),
],
hypothesiser=[_hypothesis_line("LED retrofit", "fluorescent fixtures")],
),
)
assert result.stop == "unknown_speaker"
assert result.ledger_log[0].speaker_known is True
assert result.ledger_log[-1].speaker_known is False
# The seed survives — preservation is unconditional (§ C.6 door 1) — while the loop's own
# find does not, because nothing stands behind the turn that ended the run.
assert result.mandate.approaches == (seed,)
# ---------------------------------------------------------------------------------------------
# C.2 — the token cap covers the MANAGER, which is the loop's most talkative agent
# ---------------------------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_the_token_cap_binds_the_manager_before_any_participant_speaks() -> None:
"""T14: a one-token budget stops the exploration on the MANAGER's own first call.
Agent-level ``ChatMiddleware`` does fire on the manager's calls (§ F, A1, measured green), and
the manager talks more than anyone else in a Magentic loop it extracts facts, writes the
plan, and writes a progress ledger every single round. A cap fastened only to the participants
would be a cap in name.
The assertion is deliberately not "something raised". ``kind == "tokens"`` separates it from
the round-cap stop, ``meter.tokens == 8`` shows the charge came from a call that was actually
made and metered, and the EMPTY ledger log shows it landed before the loop had run a single
round which is exactly what a manager-attached middleware does and a participant-only one
cannot.
"""
meter = TokenMeter(Budget(max_tokens=1, max_rounds=6))
contract = ExplorationContract(**{**_NO_REVIEW, "max_tokens": 1})
with pytest.raises(BudgetExceeded) as excinfo:
await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
meter=meter,
client_factory=_factory(
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
hypothesiser=[],
),
)
assert excinfo.value.kind == "tokens"
assert meter.tokens == 8, "the manager's own call must have been charged to the meter"
# ---------------------------------------------------------------------------------------------
# C.5 / U13 — the synchronous plan review, and the cap the measurement forced
# ---------------------------------------------------------------------------------------------
def _reviewer(script: list[explore.PlanReviewDecision], seen: list[explore.PlanReviewRequest]):
def review(request: explore.PlanReviewRequest) -> explore.PlanReviewDecision:
seen.append(request)
return script.pop(0) if script else explore.PlanReviewDecision.approve()
return review
@pytest.mark.asyncio
async def test_a_revision_reaches_the_manager_and_the_review_is_asked_again() -> None:
"""T15: revise → replan → asked AGAIN → approve → the loop runs.
This is målbilde's "ask the question, use the answer, carry on" on the installed stack: the
expert's words go into the manager's history, the manager replans, and the human is asked to
sign off on the NEW plan rather than the old one. Both round trips are recorded, in order,
with the feedback verbatim an audit of what a human actually told an autonomous loop is
worth nothing paraphrased.
"""
seen: list[explore.PlanReviewRequest] = []
contract = ExplorationContract(
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 2}
)
result = await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
plan_reviewer=_reviewer(
[explore.PlanReviewDecision.revise("Also test night setback.")], seen
),
client_factory=_factory(
ledgers=[
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
],
hypothesiser=[_hypothesis_line("Night setback", "the expert asked for it")],
),
)
assert [r.decision for r in result.plan_reviews] == ["revise", "approve"]
assert result.plan_reviews[0].feedback == "Also test night setback."
assert len(seen) == 2, "a revision must produce a SECOND review, not resume silently"
assert seen[1].plan != "", "the second review must show the revised plan"
assert result.stop is None
assert [a.label for a in result.mandate.approaches] == ["Night setback"]
@pytest.mark.asyncio
async def test_an_always_revising_reviewer_is_stopped_by_the_cap() -> None:
"""T16: the cap terminates a reviewer that never signs off — the reason it exists.
Measured (§ F, A3): a revise costs two manager calls, emits NO progress ledger and consumes
NO round, then asks again. The round cap therefore never ticks, and without
``max_plan_revisions`` this is an unbounded spend under caps that all look satisfied
precisely what ``shared/method-spec.md`` §8 forbids. The stop is typed and the exploration
still returns; the reviewer's last (refused) revision is recorded, because the record is of
what the human decided and ``stop`` is what says it was not applied.
"""
seen: list[explore.PlanReviewRequest] = []
always_revise = [explore.PlanReviewDecision.revise(f"Again #{n}.") for n in range(10)]
contract = ExplorationContract(
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 1}
)
result = await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
plan_reviewer=_reviewer(always_revise, seen),
client_factory=_factory(
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
hypothesiser=[],
),
)
assert result.stop == "plan_revisions_exhausted"
assert [r.decision for r in result.plan_reviews] == ["revise", "revise"]
assert result.ledger_log == (), "the loop must never have run: the plan was never approved"
@pytest.mark.asyncio
async def test_a_reviewer_that_signs_off_at_once_is_not_capped() -> None:
"""T17: the control for T16 — the same cap, a reviewer that approves, and no stop.
Without it, T16 would pass on an implementation that refuses every plan review it is given,
which would stop the runaway loop and every legitimate one with it.
"""
seen: list[explore.PlanReviewRequest] = []
contract = ExplorationContract(
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 1}
)
result = await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
plan_reviewer=_reviewer([], seen),
client_factory=_factory(
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
hypothesiser=[],
),
)
assert result.stop is None
assert [r.decision for r in result.plan_reviews] == ["approve"]
assert len(result.ledger_log) >= 1, "an approved plan must let the loop actually run"
# ---------------------------------------------------------------------------------------------
# C.0 level 3 — the exploration has no write access, and level 1 is advisory
# ---------------------------------------------------------------------------------------------
def _tree(root: Path) -> dict[str, bytes]:
return {
str(p.relative_to(root)): p.read_bytes() for p in sorted(root.rglob("*")) if p.is_file()
}
@pytest.mark.asyncio
async def test_an_exploration_leaves_the_knowledge_base_byte_identical(tmp_path: Path) -> None:
"""T18: ``explore()`` writes NOTHING — not to the base, not anywhere under it.
Level 3 of the guarantee table: only the pipeline may write an outbox artefact, and only the
gated ``promote_verdict`` may write to the wiki. An exploration that could write would be a
route around the gate that makes an answer checkable and, promoting into the base it reads,
the self-contamination loop the Step-8 gate exists to prevent.
Compared BYTE for byte over the whole subtree rather than by listing names, so a rewritten
``index.md`` of the same length would still fail.
**The tools are exercised DIRECTLY, and that is a correction, not thoroughness.** A first
version of this test drove only ``explore()`` and a mutation that made ``read_bundle`` write
a file into the base it reads left the WHOLE suite green (measured: 974 passed). A
``ScriptedChatClient`` returns text and never emits a tool call, so no scripted run reaches a
tool body: the read surface, which is the only place a write could plausibly come from, was
outside the gate entirely.
"""
base = tmp_path / "bygg-energi-baseline-mikro"
shutil.copytree(
Path(portfolio_optimiser.__file__).parent / "data" / "bundles" / base.name, base
)
before = _tree(tmp_path)
result = await explore.explore(
PROMPT,
contract=_CONTRACT,
bundle_dirs=(str(base),),
client_factory=_factory(
# Three ledgers, and the third is what makes the second one matter: the orchestrator
# tests ``is_request_satisfied`` BEFORE it reads ``next_speaker`` (``:1106``), so a
# satisfied ledger naming the hypothesiser never actually asks it anything.
ledgers=[
_ledger_json(satisfied=False, speaker=explore.NAVIGATOR_ROLE),
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
],
hypothesiser=[_hypothesis_line("LED retrofit", "fluorescent fixtures")],
),
)
assert result.mandate.approaches[0].label == "LED retrofit"
# Every read tool, called on the same base, with model-shaped arguments.
tools = {t.name: t for t in explore.navigator_tools((str(base),))}
assert tools["list_bundles"].func()[0]["id"] == base.name
assert tools["read_bundle"].func(bundle_id=base.name) != ""
assert tools["read_file"].func(bundle_id=base.name, path="index.md") != ""
explore.quick_validate_tool((str(base),)).func(
bundle_id=base.name, proposal_json=json.dumps(_micro_projection())
)
assert _tree(tmp_path) == before
def _micro_bundle_dir() -> str:
return str(
Path(portfolio_optimiser.__file__).parent
/ "data"
/ "bundles"
/ "bygg-energi-baseline-mikro"
)
def _micro_projection() -> dict[str, Any]:
projection = dict(okf.load_ir_projection(_micro_bundle_dir()))
projection.pop("_note", None)
return projection
def test_quick_validate_reports_the_real_verdict_and_says_whether_it_was_anchored() -> None:
"""T19: the in-loop check is the SAME validator, and it declares its own anchoring.
Level 1 is advisory but never fake: it runs ``validate_proposal`` against the base's own
``cost-baseline.json``, so stage 0 reconciliation is live and a fabricated cost line is caught
in the loop rather than three steps later. ``anchored`` rides along for the reason
``ProvenanceStamp.cost_baseline_anchored`` is a required field a verdict reached without the
project's real cost lines is a weaker claim, and one that does not say so is a silence.
"""
base = _micro_bundle_dir()
projection = _micro_projection()
validate = explore.quick_validate_tool((base,))
honest = validate.func(
bundle_id="bygg-energi-baseline-mikro", proposal_json=json.dumps(projection)
)
assert honest["decision"] == "validated"
assert honest["anchored"] is True
assert honest["p90"] >= honest["p50"] >= honest["p10"]
# A cost code the project does not have is refused by stage 0 — the one stage that can tell a
# fabricated line from a real one, and the reason `anchored` is worth reporting at all.
invented = dict(projection)
invented["affected_items"] = [
{**dict(projection["affected_items"][0]), "code": "CODE-THAT-DOES-NOT-EXIST"}
]
fabricated = validate.func(
bundle_id="bygg-energi-baseline-mikro", proposal_json=json.dumps(invented)
)
assert fabricated["decision"] == "rejected"
assert "CODE-THAT-DOES-NOT-EXIST" in fabricated["reason"]
def test_an_unknown_knowledge_base_is_refused_by_name() -> None:
"""T20: a tool call naming a base nobody configured refuses, and says what IS configured.
Model-chosen arguments are untrusted input. Answering an unknown id with an empty result would
let the manager conclude the base is empty rather than absent the fourth face of the
verification law, arrived at through a tool rather than a query.
"""
validate = explore.quick_validate_tool(("/tmp/base-a",))
with pytest.raises(explore.ExplorationError) as excinfo:
validate.func(bundle_id="base-b", proposal_json="{}")
assert "base-a" in str(excinfo.value)
def test_two_bases_with_the_same_name_are_refused() -> None:
"""T21: duplicate ids refuse at construction — the S3.2 key-collision class, one layer up.
The id is how the manager names a base. Two bases answering to one name would let it read A
while believing it read B, and every quotation it produced afterwards would be attributed to
the wrong project.
"""
with pytest.raises(explore.ExplorationError):
explore.navigator_tools(("/tmp/one/shared-name", "/tmp/two/shared-name"))
# ---------------------------------------------------------------------------------------------
# U14 — the three events the tracing seam was landed for, now that they have a call site
# ---------------------------------------------------------------------------------------------
def _recording_tracer() -> tuple[Any, InMemorySpanExporter]:
"""A REAL OpenTelemetry tracer over an in-memory exporter — not a spy.
A recorder standing in for ``add_event`` would prove that this module calls something shaped
like OTel; this proves the events survive the actual SDK, with the attribute types it will
accept. The provider is LOCAL and is never installed globally, so the pytest process keeps
whatever tracing configuration it had (the same restraint U14's own tests exercise).
"""
provider = TracerProvider()
exporter = InMemorySpanExporter()
provider.add_span_processor(SimpleSpanProcessor(exporter))
return provider.get_tracer("test"), exporter
@pytest.mark.asyncio
async def test_the_three_orchestrator_events_reach_the_trace(monkeypatch: Any) -> None:
"""T22: ``plan_created``, ``replanned`` and ``progress_ledger_updated`` are recorded.
These are the events U14 deliberately did NOT build in økt 55 "an emitter with no call site
is a shape guessed instead of measured". This is the call site. A Magentic manager decides
which base to open and who speaks next; without these, the only trace of that reasoning is
MAF's own ``invoke_agent`` spans, which say a call happened and nothing about what it decided.
The ledger event carries the decision fields rather than a rendered sentence, for the reason
``SkippedLink`` is structured and ``BudgetExceeded`` carries three fields: "who was asked" and
"was the request satisfied" are separate operative questions, and a reader who has to re-parse
prose to tell them apart has a trace they cannot query.
"""
tracer, exporter = _recording_tracer()
# Patched where the name is BOUND (the ``hosting.run_project`` precedent): ``explore``
# imports it by name, so patching ``tracing`` would leave that binding untouched and this
# test would quietly measure nothing.
monkeypatch.setattr(explore, "exploration_tracer", lambda: tracer)
seen: list[explore.PlanReviewRequest] = []
contract = ExplorationContract(
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 2}
)
await explore.explore(
PROMPT,
contract=contract,
bundle_dirs=(),
plan_reviewer=_reviewer([explore.PlanReviewDecision.revise("Test night setback.")], seen),
client_factory=_factory(
ledgers=[
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
],
hypothesiser=[_hypothesis_line("Night setback", "the expert asked")],
),
)
spans = exporter.get_finished_spans()
assert [s.name for s in spans] == [explore.EXPLORATION_SPAN]
events = [(e.name, dict(e.attributes or {})) for e in spans[0].events]
names = [name for name, _ in events]
assert names.count("plan_created") == 1
assert names.count("replanned") == 1, "the human's revision must be visible in the trace"
assert names.count("progress_ledger_updated") == 2
ledger_events = [attrs for name, attrs in events if name == "progress_ledger_updated"]
assert [a["round_index"] for a in ledger_events] == [1, 2]
assert [a["next_speaker"] for a in ledger_events] == [explore.HYPOTHESISER_ROLE] * 2
assert [a["is_request_satisfied"] for a in ledger_events] == [False, True]
assert all(a["speaker_known"] for a in ledger_events)
#: A complete exploration in a CHILD interpreter. The stdout/stderr question cannot be answered
#: in-process: ``ConsoleSpanExporter``'s ``out`` default is bound when
#: ``opentelemetry.sdk.trace.export`` is first imported, so under pytest it is whatever stdout was
#: at COLLECTION time — and ``capsys``, which replaces ``sys.stdout`` later, never sees it. That is
#: not a testing quirk to work around; it is precisely the fact U14 exists for, and the reason
#: ``configure_tracing`` passes ``out=`` explicitly instead of trusting the default. Measured: a
#: mutation routing exploration spans to that default left an in-process ``capsys`` assertion
#: GREEN while the spans really were on stdout.
_CHILD_EXPLORATION = """
import asyncio, json, sys
from portfolio_optimiser import explore
from portfolio_optimiser.simulation import ScriptedChatClient
from portfolio_optimiser.tracing import configure_tracing
configure_tracing()
LEDGER = json.dumps({
"is_request_satisfied": {"reason": "r", "answer": True},
"is_in_loop": {"reason": "r", "answer": False},
"is_progress_being_made": {"reason": "r", "answer": True},
"next_speaker": {"reason": "r", "answer": "navigator"},
"instruction_or_question": {"reason": "r", "answer": "none"},
})
def _select(blob, _role):
if "provide the final answer" in blob:
return "FINAL: done."
if "pure JSON format" in blob:
return LEDGER
if "bullet-point plan" in blob:
return "PLAN: - ask the navigator"
if "pre-survey" in blob:
return "FACTS: none."
return "{}"
def factory(role):
if role == explore.MANAGER_ROLE:
return ScriptedChatClient(reply_selector=_select, role=role)
return ScriptedChatClient("ok", role=role)
contract = explore.ExplorationContract(
max_rounds=4, max_tokens=100000, max_stall_count=2,
max_reset_count=1, max_plan_revisions=0, enable_plan_review=False,
)
result = asyncio.run(
explore.explore("probe", contract=contract, bundle_dirs=(), client_factory=factory)
)
assert result.stop is None, result.stop
print("EXPLORATION-OK", file=sys.stderr)
"""
def _run_child(**env: str) -> subprocess.CompletedProcess[str]:
return subprocess.run(
[sys.executable, "-c", _CHILD_EXPLORATION],
capture_output=True,
text=True,
cwd=str(Path(__file__).resolve().parent.parent),
env={**os.environ, **env},
)
def test_an_untraced_exploration_writes_nothing_to_stdout_or_stderr() -> None:
"""T23: in a real process, with tracing off, an exploration prints NOTHING.
A subprocess and not ``capsys``, for the reason recorded above ``_CHILD_EXPLORATION`` and the
stakes are the pinned artefacts: ``tests/golden/demo-transcript.stdout`` is byte-fixed and the
demo's stderr is fixed at four lines, so one stray span dump would break both.
``EXPLORATION-OK`` on stderr is the control. Without it, "stdout was empty" would be equally
true of a child that crashed on import, which is the fourth face of the verification law: an
absence is only evidence once you have shown the measurement could have found something.
"""
proc = _run_child(PORTFOLIO_OTEL="")
assert proc.returncode == 0, proc.stderr
assert "EXPLORATION-OK" in proc.stderr, "the child must really have run an exploration"
assert proc.stdout == ""
# Not an exact-equality assertion on stderr: MAF emits two ``ExperimentalWarning`` lines while
# importing, under every run form, and they are the same pair the demo's pinned stderr already
# carries. What must be absent is TRACE data, so that is what is asserted.
assert '"name": "exploration"' not in proc.stderr
assert "progress_ledger_updated" not in proc.stderr
def test_a_traced_exploration_puts_its_span_on_stderr_and_leaves_stdout_clean() -> None:
"""T24: the positive arm — ``PORTFOLIO_OTEL=console`` and the exploration span is on STDERR.
This is what the whole U14 seam was landed for, now carrying the events U4 gave it a call site
for. Both halves are asserted: the span and its ``progress_ledger_updated`` event ARE exported
(so tracing is real), and stdout is STILL empty (so the byte-pinned transcript survives a
traced run). Asserting only the first would pass on an exporter writing to stdout which is
OpenTelemetry's own default, and therefore the mistake actually available to make.
"""
proc = _run_child(PORTFOLIO_OTEL="console")
assert proc.returncode == 0, proc.stderr
assert "EXPLORATION-OK" in proc.stderr
assert proc.stdout == "", "a traced run must not put one byte on stdout"
assert '"name": "exploration"' in proc.stderr
assert "progress_ledger_updated" in proc.stderr
def test_a_marked_line_that_will_not_parse_is_a_hard_error() -> None:
"""T24: the marker is what makes fail-closed affordable here.
Most hypothesiser turns legitimately are not hypotheses the agent reasons out loud so
"parse every turn or fail" would refuse a normal exploration. The marker separates a turn that
is not a claim from a claim that cannot be read. The second is the run's own product coming
back unreadable, so it raises (``write_concept_file``'s rule: validation, never repair) rather
than following the tolerant RAW-inbox rule, which belongs to folders anyone may drop files in.
"""
with pytest.raises(explore.HypothesisParseError):
explore._parse_hypotheses([f"{explore.HYPOTHESIS_MARKER} not json at all"])
with pytest.raises(explore.HypothesisParseError):
explore._parse_hypotheses([f'{explore.HYPOTHESIS_MARKER} {{"label": "no rationale"}}'])
def test_unmarked_prose_is_not_a_failure() -> None:
"""T25: the control for T24 — ordinary reasoning yields no hypothesis and no error.
Without it, T24 would pass on an implementation that refused every hypothesiser turn that was
not a hypothesis, which would make the loop unusable and the strictness meaningless.
"""
assert explore._parse_hypotheses(["I looked at the index and nothing stands out yet."]) == []
def test_a_review_nobody_can_answer_is_refused_before_the_first_model_call() -> None:
"""T26: plan review without a reviewer refuses; a reviewer without plan review refuses too.
The first would hang: the workflow stops at a ``request_info`` and nothing ever answers it, and
a hang is the one failure mode that reports nothing at all. The second is the silent-ignore the
repo's flag partition forbids — a caller who supplied a reviewer believes a human is in the
loop. Both are refused BEFORE anything is built, so neither costs a model call.
"""
with pytest.raises(explore.ExplorationError):
asyncio.run(
explore.explore(
PROMPT,
contract=ExplorationContract(**_FULL_CONTRACT),
bundle_dirs=(),
client_factory=_factory(ledgers=[], hypothesiser=[]),
)
)
with pytest.raises(explore.ExplorationError):
asyncio.run(
explore.explore(
PROMPT,
contract=ExplorationContract(**_NO_REVIEW),
bundle_dirs=(),
plan_reviewer=lambda _r: explore.PlanReviewDecision.approve(),
client_factory=_factory(ledgers=[], hypothesiser=[]),
)
)