feat(explore): U4+U13 synkron - utforskningssloeyfa som mandat-former (ORDRE 20260823T185602Z) [skip-docs]
Magentic legges OVER den normative sloeyfa, aldri inni Steg 3: prompt + kunnskapsbaser -> Mandate -> run_project(mandate=...) UENDRET. Manageren velger VEI; det som forlater friheten er et Mandate, aldri et forslag. explore() skriver ingenting - niva 3 (skriverettigheter) tilhoerer pipelinen alene. Levert i denne oekten (kjernen; kallstedene staar til oekt 57): - ExplorationContract: seks paakrevde felt uten default. max_reset_count=0 nektes paa en MAALING - reset_count >= max_reset_count mot en teller som starter paa 0 terminerer kjoeringen FOER foerste runde med null ledger-events, altsaa en utforskning som utforsket ingenting, forkledd som en stall som aldri skjedde. - explore() + fresh_exploration_workflow(): fersk builder per utforskning, BudgetMiddleware paa HVER agent inkl. manageren, synkron plan review via request_info, og max_plan_revisions som binder den ubundne revise-loekka. - Tre kanaler: tokens OG runder raiser BudgetExceeded (rundene oversatt av vaart lag som kind="exploration_rounds", fordi orkestreringen maalt ikke raiser ved sitt eget rundetak); alt semantisk er en VERDI i stop. - quick_validate (niva 1, raadgivende) + navigator-verktoey over safe_resolve. - U14s tre utsatte events landet som span-events paa EN exploration-span. Load-bearing MAALT mot HELE suiten, groenn kontroll 975/5, golden ea8c534 uendret: tolv mutasjoner alle roede. TO av dem falsifiserte testen foerst - skrivefrihets-testen naadde aldri en verktoeykropp (ScriptedChatClient emitterer ingen verktoeykall), og stdout-testens capsys er blind for ConsoleSpanExporter, hvis out-default bindes ved modulimport. Begge er rettet; stdout-armen er naa en subprosess, som er P4-presedensen. [skip-docs] fordi flaten ikke er naabar for en bruker enna: --explore, det whitelistede hosting-feltet og sim-scenarioet bygges i oekt 57, og en README-oppfoering naa ville vaert en paastand om en inngang som ikke finnes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
4a19d39e63
commit
f0c54cc8dc
4 changed files with 1892 additions and 4 deletions
64
CLAUDE.md
64
CLAUDE.md
|
|
@ -772,6 +772,70 @@ Python ≥3.10. MAF (`agent-framework-core` 1.9.0). Pakkehåndtering: `uv`. To b
|
|||
fire-linjers stderr og `test_portfolio_cli_offline`s stille-pass — altså er omisjonen gatet av
|
||||
uavhengige vitner) · detach demo-wiringen + detach CLI-wiringen (5 røde) · detach hosting-wiringen
|
||||
(1 rød, KUN subprosess-testen — P4-presedensen).
|
||||
- **Utforskningssløyfa er en MANDAT-FORMER, og de tre garantinivåene er strukturelle (U4+U13
|
||||
synkron, økt 56):** `explore.py` legger en Magentic-manager OVER den normative sløyfa — `prompt +
|
||||
kunnskapsbaser → Mandate → run_project(mandate=…)` UENDRET, Steg 3s maker-checker urørt (commons-eid
|
||||
og normativ). Manageren velger VEI; det som forlater friheten er `mandate.Mandate`, aldri et forslag.
|
||||
**Nivå 1** = `quick_validate`-verktøyet (SAMME `validate_proposal`, SAMME baseline, men rådgivende —
|
||||
når ALDRI provenance); **nivå 2** = pipelinen som stempler; **nivå 3** = skriverettigheter, som kun
|
||||
pipelinen har. `explore()` skriver INGENTING. **Mandatet bygges fra hypotesiserens MERKEDE turer**
|
||||
(`HYPOTHESIS: {"label","rationale"}`), aldri fra sluttsvaret — sluttsvaret er RÅTT per design, og en
|
||||
parser på det ville gjort det til et forslag. Markøren er dét som gjør fail-closed mulig: en umerket
|
||||
tur er ikke en påstand (ingen stillhet å lukke), mens en MERKET-men-uleselig linje raiser
|
||||
(`write_concept_file`-regelen). **Frø-approaches bevares ALLTID og FØRST** — også når sløyfa fant
|
||||
ingenting og også ved stopp (§ C.6 dør 1 er en bevaringsregel, ikke en belønning for å bli ferdig).
|
||||
**Tre kanaler, aldri én:** tokens OG runder raiser `BudgetExceeded` (rundene som
|
||||
`kind="exploration_rounds"`, oversatt av VÅRT lag fordi orkestreringen MÅLT ikke raiser ved sitt eget
|
||||
rundetak — den returnerer en kanonisk assistent-melding som ved transporten er uskillbar fra suksess),
|
||||
mens alt semantisk er en VERDI i `stop` (S3.4-splitten: utmattelse og utfall er ikke samme sak).
|
||||
Diskriminatoren mellom «nådde taket» og «ble kappet av taket» er SISTE ledgers
|
||||
`is_request_satisfied` — samme felt orkestratoren selv forgrener på (`:1106`) — aldri
|
||||
termineringsmeldingen, som er en inline f-string (`:1253`) uten konstant å pinne mot og som en modells
|
||||
eget sluttsvar kan inneholde. **`speaker_known` sjekkes FØRST og slår alt annet:** en `next_speaker`
|
||||
uten treff gir stille sluttsvar med NULL deltakerarbeid (`:1128-1131`), altså et plausibelt svar
|
||||
ingen jobbet for (E2-klassen) — sløyfas egne funn holdes da tilbake, frøene ikke.
|
||||
**Kontrakten nekter tre ting ved konstruksjon:** hvert av seks felt er PÅKREVD uten default (MAF
|
||||
defaulter `max_round_count`/`max_reset_count` til ubegrenset, så et utelatt felt faller ikke tilbake
|
||||
til noe forsiktig, men til dét `method-spec` §8 forbyr); `max_reset_count=0` nektes — **MÅLT**, ikke
|
||||
resonnert: `reset_count >= max_reset_count` mot en teller som starter på 0 gjør at kjøringen
|
||||
terminerer FØR første runde med kun `facts`+`plan`, null ledger-events og «maximum reset count», altså
|
||||
en utforskning som utforsket ingenting, forkledd som en stall som aldri skjedde (`max_stall_count=0`
|
||||
er derimot LOVLIG — strengt `>` gjør 0 til «reset ved første stallede runde»); og
|
||||
`max_plan_revisions>0` med `enable_plan_review=False` nektes (en cap på en hendelse som ikke kan skje).
|
||||
`max_plan_revisions` finnes fordi A3 MÅLTE at en `revise` koster 2 manager-kall, **null** ledger-kall
|
||||
og **null** runder og så spør PÅ NYTT — under rundetaket alene er en alltid-reviderende ekspert
|
||||
ubundet forbruk under vakter som alle ser tilfredse ut. Ved cap: typet stopp, ALDRI en påtvunget
|
||||
approve (repair av et menneskes beslutning er den verste sorten). **U14s tre utsatte events er
|
||||
landet** (`plan_created`/`replanned`/`progress_ledger_updated` som span-events på ÉN
|
||||
`exploration`-span) — emisjon er UBETINGET og «av» betyr at OTel kaster dem, samme form MAFs egen
|
||||
instrumentering alt har; en flagget emitter ville vært en andre oppløsning av regelen `tracing.py`
|
||||
eier. Uttalt ærlighets-grense: eventene registreres når event-strømmen foldes, så REKKEFØLGEN er
|
||||
tro og tidsstemplene er ikke øyeblikkene manageren handlet.
|
||||
**Load-bearing MÅLT** (`tests/test_explore_loadbearing.py`, 32 tester), tolv mutasjoner alle røde mot
|
||||
HELE suiten + grønn kontroll 975/5: detach rundetak-oversettelsen (1 rød) · test rundetaket FØR
|
||||
tilfredsstillelse (1 rød — den motsatte feilen, som gjør en fullført utforskning til en budsjettfeil) ·
|
||||
detach ukjent-taler-sjekken (1) · `BudgetMiddleware` av manageren, deltakerne beholder den (1) ·
|
||||
detach revisjons-capen (1) · dropp frø-bevaringen (3) · la en ukjent-taler-kjøring levere funnene
|
||||
videre (1) · detach alle tre span-events (1) · tillat `max_reset_count=0` (1) · gjør en merket-men-
|
||||
uleselig hypotese tolerant (1) · la `read_bundle` skrive i basen den leser (1) · send spans til OTels
|
||||
default-sink (2). **TO av dem FALSIFISERTE testen først, og begge er repoets vakuøs-gate-klasse:**
|
||||
(i) skrivefrihets-testen drev kun `explore()`, men en `ScriptedChatClient` returnerer TEKST og
|
||||
emitterer aldri et verktøykall — så ingen scriptet kjøring når en verktøykropp, og hele lesesømmen
|
||||
(eneste sted en skriving realistisk kan komme fra) lå utenfor gaten; testen kaller nå hvert verktøy
|
||||
DIREKTE. (ii) stdout-testen brukte `capsys`, men `ConsoleSpanExporter`s `out`-default bindes når
|
||||
`opentelemetry.sdk.trace.export` FØRST importeres — under pytest er det stdout ved COLLECTION, som
|
||||
`capsys` aldri ser; spans lå faktisk på stdout mens asserten var grønn. Dét er ikke en test-quirk å
|
||||
omgå, det er nøyaktig faktumet U14 finnes for, og arven er P4-presedensen: **subprosessen er
|
||||
målingen**. Begge armene kjøres nå i et barn (av: null trace-data noe sted; `PORTFOLIO_OTEL=console`:
|
||||
spanet + `progress_ledger_updated` på **stderr** og stdout tomt), med `EXPLORATION-OK` på stderr som
|
||||
kontroll — uten den ville «stdout var tomt» vært like sant om et barn som krasjet ved import.
|
||||
**Ærlighets-grenser, uttalt:** multi-base-dispatch (`Approach.bundle_id`, § C.7) venter til
|
||||
`run_project` tar mer enn én `bundle_dir` — å shippe feltet før konsumenten er en form gjettet i
|
||||
stedet for målt; `quick_validate`-dommene hypotesiseren så bor ikke i `ExplorationResult`, de er nivå
|
||||
1 og hører hjemme i `{run_id}-exploration.json` som CLI-wiringen skriver; utforskningsrollene løses
|
||||
via `resolve_model`s `default`-fallback til en operatør mapper dem eksplisitt; at en LEVENDE modell
|
||||
kaller verktøyene er ikke bevist offline (samme klasse som structured-output-grensen). `--explore` i
|
||||
`run.py`, `explore_prompt` i `hosting.py` og sim-scenarioet er IKKE bygget (økt 57).
|
||||
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
||||
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
||||
|
||||
|
|
|
|||
804
src/portfolio_optimiser/explore.py
Normal file
804
src/portfolio_optimiser/explore.py
Normal file
|
|
@ -0,0 +1,804 @@
|
|||
"""U4 + U13-synchronous — the Magentic exploration loop, sitting OVER the normative pipeline.
|
||||
|
||||
**One sentence of design, and it is not to be reopened without a measurement.** Magentic is laid
|
||||
over the eight-step loop, never inside Step 3: ``prompt + knowledge bases -> Mandate ->
|
||||
run_project(mandate=...)``, with that last call byte-for-byte the one it already was. The manager
|
||||
is free to choose which base to open and which hypothesis to shape next; what leaves that freedom
|
||||
is a ``mandate.Mandate`` — approaches worth *testing*. The exploration's own final answer is RAW,
|
||||
never a proposal, and ``explore()`` writes to neither the outbox nor the wiki. Every number that
|
||||
survives is gated by ``validate_proposal`` inside ``run_project``, in the same blocking gate as
|
||||
today, and ``shared/method-spec.md`` §3's maker-checker debate is untouched (it is commons-owned
|
||||
and normative; replacing it would need an amendment, not a module).
|
||||
|
||||
**Three levels of guarantee, stated so nobody reads more into the loop than is there.** The
|
||||
``quick_validate`` tool the hypothesiser calls in-loop is level 1: the SAME ``validate_proposal``,
|
||||
against the SAME baseline, but its verdict is advisory — it never becomes provenance. Level 2 is
|
||||
the pipeline, where each approach is validated and stamped. Level 3 is write access, which only
|
||||
the pipeline has.
|
||||
|
||||
This module imports ``agent_framework.orchestrations`` and therefore may never be imported from
|
||||
``okf.py`` / ``mandate.py`` / ``hitl.py``, which are held framework-neutral by
|
||||
``tests/test_okf.py::test_okf_is_maf_free``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from collections.abc import Callable, Mapping, Sequence
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Final, Literal
|
||||
|
||||
from agent_framework import Agent, BaseChatClient, FunctionTool, tool
|
||||
from agent_framework.orchestrations import (
|
||||
MagenticBuilder,
|
||||
MagenticOrchestratorEventType,
|
||||
MagenticPlanReviewResponse,
|
||||
)
|
||||
from pydantic import BaseModel, Field, ValidationError, model_validator
|
||||
|
||||
from portfolio_optimiser import okf
|
||||
from portfolio_optimiser.backends import Profile
|
||||
from portfolio_optimiser.budget import Budget, BudgetExceeded, BudgetMiddleware, TokenMeter
|
||||
from portfolio_optimiser.ir import SavingsProposal
|
||||
from portfolio_optimiser.mandate import OWN_PROPOSAL_ID, Approach, Mandate
|
||||
from portfolio_optimiser.retrieval import safe_resolve
|
||||
from portfolio_optimiser.tracing import exploration_tracer
|
||||
from portfolio_optimiser.validator import Rejection, validate_proposal
|
||||
|
||||
|
||||
class ExplorationContract(BaseModel):
|
||||
"""The stated bounds of ONE exploration. Every field is required; none has a default.
|
||||
|
||||
Mirrors ``contracts.TerminationContract`` in spirit and goes further in one respect: there is
|
||||
nothing to inherit. ``MagenticBuilder`` defaults ``max_round_count`` to ``None`` (unbounded)
|
||||
and ``max_reset_count`` to ``None`` (unlimited), so a field left out here would not fall back
|
||||
to something conservative — it would fall back to the one shape ``shared/method-spec.md`` §8
|
||||
forbids outright, an unbounded loop. A default would also assert an intent nobody stated,
|
||||
which is the ground ``ProvenanceStamp.cost_baseline_anchored`` is required on.
|
||||
|
||||
``max_plan_revisions`` earns its place by measurement, not symmetry (§ F, A3): a plan-review
|
||||
``revise`` costs two manager calls, emits **no** progress ledger and consumes **no** round,
|
||||
then asks again. Under the round cap alone an always-revising expert is an unbounded spend
|
||||
that the round counter never sees. ``0`` is meaningful and allowed: the plan must be approved
|
||||
as first written or the exploration stops.
|
||||
"""
|
||||
|
||||
#: Hard cap on orchestration rounds (``MagenticBuilder(max_round_count=...)``).
|
||||
max_rounds: int = Field(gt=0)
|
||||
#: Hard cap on tokens, enforced by ``BudgetMiddleware`` on EVERY agent, manager included.
|
||||
max_tokens: int = Field(gt=0)
|
||||
#: Consecutive no-progress rounds tolerated before the manager resets and replans. ``0`` is
|
||||
#: allowed and means "reset on the first round that reports no progress": the orchestrator's
|
||||
#: check is STRICT (``stall_count > max_stall_count``, ``_magentic.py:1118``) and the counter is
|
||||
#: incremented ahead of it, so zero is strictness rather than self-defeat.
|
||||
max_stall_count: int = Field(ge=0)
|
||||
#: Resets tolerated before the exploration is declared stalled. Must be POSITIVE, and that is a
|
||||
#: MEASUREMENT: the limit check is ``reset_count >= max_reset_count`` (``:1243``) against a
|
||||
#: counter starting at zero, so a cap of zero is already met before the first round. Measured
|
||||
#: against the installed stack, ``max_reset_count=0`` produces only the ``facts`` and ``plan``
|
||||
#: manager calls, **zero** progress-ledger events, and the canonical "maximum reset count"
|
||||
#: termination — an exploration that explored nothing, dressed as a stall that never happened.
|
||||
max_reset_count: int = Field(gt=0)
|
||||
#: Plan revisions the reviewer may ask for before the exploration stops. See the class note.
|
||||
max_plan_revisions: int = Field(ge=0)
|
||||
#: Whether a human/persona signs the plan off before the loop runs (the U13 synchronous door).
|
||||
enable_plan_review: bool
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _a_revision_cap_needs_a_review_to_cap(self) -> "ExplorationContract":
|
||||
"""A revision cap with the review switched off bounds an event that cannot occur.
|
||||
|
||||
Refused rather than dropped: a caller who wrote ``max_plan_revisions=3`` believes they
|
||||
bounded something, and silently ignoring it is the failure mode this repo's flag contract
|
||||
exists to prevent (``--embedder-config requires --semantic-retrieval``, same rule). The
|
||||
coherent way to say "no reviews" is ``enable_plan_review=False, max_plan_revisions=0``.
|
||||
"""
|
||||
if not self.enable_plan_review and self.max_plan_revisions:
|
||||
raise ValueError(
|
||||
f"max_plan_revisions={self.max_plan_revisions} bounds plan revisions, but "
|
||||
"enable_plan_review is false so no plan review is ever requested and no revision "
|
||||
"can occur; set max_plan_revisions=0 or enable the review"
|
||||
)
|
||||
return self
|
||||
|
||||
|
||||
#: The manager's role name. Not a participant: the manager plans, picks the next speaker and
|
||||
#: keeps the progress ledger, and ``BudgetMiddleware`` rides on it like on everyone else (A1,
|
||||
#: measured green — without that, the most talkative agent in the loop would be the one outside
|
||||
#: the token cap).
|
||||
MANAGER_ROLE: Final = "manager"
|
||||
#: Reads the knowledge bases PROGRESSIVELY (§3 Steg 1) — index, then whatever the index links.
|
||||
NAVIGATOR_ROLE: Final = "navigator"
|
||||
#: Shapes ONE candidate direction at a time and may have its numbers advisory-checked in-loop.
|
||||
HYPOTHESISER_ROLE: Final = "hypothesiser"
|
||||
|
||||
#: The participants, in build order. The manager's ``next_speaker`` is validated against exactly
|
||||
#: these names: a ledger naming anyone else makes the orchestrator emit a final answer having
|
||||
#: asked NOBODY (measured, ``_magentic.py:1128-1131``), which is a plausible answer produced by
|
||||
#: zero work — the hazard class this repo retired ``E2`` for.
|
||||
PARTICIPANT_ROLES: Final = (NAVIGATOR_ROLE, HYPOTHESISER_ROLE)
|
||||
|
||||
#: The span every exploration records its decisions under. One span per exploration, with the
|
||||
#: manager's decisions as span EVENTS on it, because those are moments INSIDE one activity rather
|
||||
#: than activities of their own — and because a reader wants them in order, on one timeline.
|
||||
EXPLORATION_SPAN: Final = "exploration"
|
||||
|
||||
#: The line prefix that turns a hypothesiser turn into a candidate direction. A marker rather than
|
||||
#: "parse every turn as JSON" because most turns legitimately are not hypotheses — the agent also
|
||||
#: reasons out loud. That distinction is what lets an UNPARSEABLE marked line be a hard error
|
||||
#: instead of a silence: with no marker there is nothing to be silent about.
|
||||
HYPOTHESIS_MARKER: Final = "HYPOTHESIS:"
|
||||
|
||||
_INSTRUCTIONS: Final = {
|
||||
NAVIGATOR_ROLE: (
|
||||
"You read the project's knowledge bases. Use list_bundles to see what exists, then "
|
||||
"read_bundle to open ONE at a time and read_file to follow a specific document. Quote "
|
||||
"what you found; never guess at content you have not read."
|
||||
),
|
||||
HYPOTHESISER_ROLE: (
|
||||
"You shape ONE candidate cost-saving direction at a time from what the navigator found. "
|
||||
"You may call quick_validate to sanity-check a candidate's numbers; its verdict is "
|
||||
"ADVISORY and is not the project's decision. When you commit to a direction, end your "
|
||||
f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", '
|
||||
'"rationale": "<why this project, in your own words>"}'
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
class ExplorationError(RuntimeError):
|
||||
"""The exploration cannot be honoured as configured, or produced something unreadable."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LedgerEntry:
|
||||
"""One progress-ledger round, reduced to the five fields the manager steers on (C.2).
|
||||
|
||||
``speaker_known`` is not part of the ledger MAF produces — it is this layer's reading of it,
|
||||
and the reason the whole entry is recorded rather than counted. A ``next_speaker`` matching no
|
||||
participant does not raise: the orchestrator logs a warning and jumps straight to a final
|
||||
answer, so the run returns a plausible result that nobody worked for. Recording the fact is
|
||||
what lets ``explore()`` refuse to hand that result onward as if it had been explored.
|
||||
"""
|
||||
|
||||
round_index: int
|
||||
is_request_satisfied: bool
|
||||
is_in_loop: bool
|
||||
is_progress_being_made: bool
|
||||
next_speaker: str
|
||||
instruction_or_question: str
|
||||
speaker_known: bool
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PlanReview:
|
||||
"""One synchronous plan-review round trip (U13, door 2 of § C.6).
|
||||
|
||||
``is_stalled`` is carried because the same request type serves two very different moments:
|
||||
the initial sign-off, and a re-review after the manager reset and replanned. A reviewer that
|
||||
cannot tell them apart cannot answer the second one usefully.
|
||||
"""
|
||||
|
||||
index: int
|
||||
plan: str
|
||||
is_stalled: bool
|
||||
decision: Literal["approve", "revise"]
|
||||
feedback: str = ""
|
||||
|
||||
|
||||
#: Why an exploration ended without a satisfied request. ``None`` means it concluded normally.
|
||||
#: Resource exhaustion is NOT here: tokens and rounds raise ``BudgetExceeded`` (the 429 channel),
|
||||
#: because "we ran out" and "we finished, unsatisfied" are the two things S3.4 split apart and a
|
||||
#: single field would fuse back together.
|
||||
ExplorationStop = Literal["stalled", "plan_revisions_exhausted", "unknown_speaker"]
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ExplorationResult:
|
||||
"""What one exploration produced. The mandate is the product; the rest is the evidence.
|
||||
|
||||
``mandate`` is ALWAYS present, including on a stop: an exploration that was cut short still
|
||||
carries the expert's seed approaches forward, because door 1 of § C.6 is a preservation rule
|
||||
and not a reward for finishing. ``stop`` is what says the mandate is smaller than it might
|
||||
have been, and why.
|
||||
"""
|
||||
|
||||
mandate: Mandate
|
||||
ledger_log: tuple[LedgerEntry, ...]
|
||||
stop: ExplorationStop | None
|
||||
plan_reviews: tuple[PlanReview, ...]
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PlanReviewRequest:
|
||||
"""What the reviewer is shown before the loop is allowed to run."""
|
||||
|
||||
index: int
|
||||
plan: str
|
||||
current_progress: str
|
||||
is_stalled: bool
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PlanReviewDecision:
|
||||
"""The reviewer's answer. ``feedback is None`` means approve.
|
||||
|
||||
Two named constructors mirroring ``MagenticPlanReviewResponse.approve()/.revise()``, so a
|
||||
persona or operator answering a review never has to import ``agent_framework`` — the adapter
|
||||
is one-directional and lives in exactly one place.
|
||||
"""
|
||||
|
||||
feedback: str | None
|
||||
|
||||
@staticmethod
|
||||
def approve() -> PlanReviewDecision:
|
||||
return PlanReviewDecision(feedback=None)
|
||||
|
||||
@staticmethod
|
||||
def revise(feedback: str) -> PlanReviewDecision:
|
||||
if not feedback.strip():
|
||||
raise ValueError("a revision must say what to revise; use approve() to sign off")
|
||||
return PlanReviewDecision(feedback=feedback)
|
||||
|
||||
|
||||
#: The synchronous HITL seam (U13). Given the request, answer it. Called in-process, so the
|
||||
#: exploration blocks on it exactly as a human at a terminal would block the loop.
|
||||
PlanReviewer = Callable[[PlanReviewRequest], PlanReviewDecision]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# The tools. Level 1 of the three-guarantee table: real computation, ADVISORY verdicts.
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _bundle_index(bundle_dirs: Sequence[str]) -> dict[str, str]:
|
||||
"""Map each knowledge base's id to its directory. The id is the directory's BASENAME.
|
||||
|
||||
A duplicate basename is REFUSED rather than resolved by order: the id is what the manager
|
||||
names a base by, and two bases answering to one name would let it read A while believing it
|
||||
read B — the S3.2 key-collision class, one layer up.
|
||||
"""
|
||||
index: dict[str, str] = {}
|
||||
for raw in bundle_dirs:
|
||||
bundle_id = Path(raw).name
|
||||
if bundle_id in index:
|
||||
raise ExplorationError(
|
||||
f"two knowledge bases share the id {bundle_id!r} ({index[bundle_id]!r} and "
|
||||
f"{raw!r}); the manager names a base by that id, so it must be unique"
|
||||
)
|
||||
index[bundle_id] = raw
|
||||
return index
|
||||
|
||||
|
||||
def _resolve_bundle(index: Mapping[str, str], bundle_id: str) -> str:
|
||||
if bundle_id not in index:
|
||||
known = ", ".join(sorted(index)) or "(none configured)"
|
||||
raise ExplorationError(f"unknown knowledge base {bundle_id!r}; configured: {known}")
|
||||
return index[bundle_id]
|
||||
|
||||
|
||||
def navigator_tools(bundle_dirs: Sequence[str]) -> list[FunctionTool]:
|
||||
"""The navigator's three tools: survey the catalogue, open one base, read one document.
|
||||
|
||||
Progressive disclosure, not stuffing (målbilde §2/§4): ``list_bundles`` never returns content,
|
||||
only what each base IS — and crucially whether it ships a ``cost-baseline.json``, because a base
|
||||
without one cannot have its hypotheses reconciled against the project's real cost lines, and a
|
||||
manager that does not know which bases are anchored cannot plan around it (§ C.7).
|
||||
|
||||
``read_bundle`` returns ``okf.bundle_context``, which EXCLUDES the ``type: verdict`` layer by
|
||||
construction — prior verdicts reach a hypothesis only through the gated ExpeL fold inside
|
||||
``run_project``, never by being read as context here.
|
||||
"""
|
||||
index = _bundle_index(bundle_dirs)
|
||||
|
||||
@tool(
|
||||
name="list_bundles",
|
||||
description=(
|
||||
"List the knowledge bases available to this exploration: id, what the index says the "
|
||||
"base is about, how many prior expert verdicts it holds, and whether it ships a cost "
|
||||
"baseline (without one, numbers cannot be reconciled against the project's own)."
|
||||
),
|
||||
)
|
||||
def list_bundles() -> list[dict[str, Any]]:
|
||||
catalogue: list[dict[str, Any]] = []
|
||||
for bundle_id, bundle_dir in index.items():
|
||||
bundle = okf.navigate_bundle(bundle_dir)
|
||||
catalogue.append(
|
||||
{
|
||||
"id": bundle_id,
|
||||
"index_summary": bundle.index_summary,
|
||||
"verdict_count": len(bundle.verdicts),
|
||||
# Tolerant on CONTENT, fail-fast on the PATH: an operator's bad directory is
|
||||
# refused by navigate_bundle above, while a navigable base that simply has no
|
||||
# baseline is legitimate (load_optional_cost_baseline's own contract).
|
||||
"cost_baseline": okf.load_optional_cost_baseline(bundle_dir) is not None,
|
||||
"skipped_links": [
|
||||
{"from_file": s.from_file, "target": s.target, "reason": s.reason}
|
||||
for s in bundle.skipped
|
||||
],
|
||||
}
|
||||
)
|
||||
return catalogue
|
||||
|
||||
@tool(
|
||||
name="read_bundle",
|
||||
description="Open ONE knowledge base by id and read its navigated context.",
|
||||
)
|
||||
def read_bundle(bundle_id: str) -> str:
|
||||
return okf.bundle_context(okf.navigate_bundle(_resolve_bundle(index, bundle_id)))
|
||||
|
||||
@tool(
|
||||
name="read_file",
|
||||
description="Read ONE document inside a knowledge base, by base id and relative path.",
|
||||
)
|
||||
def read_file(bundle_id: str, path: str) -> str:
|
||||
bundle_dir = _resolve_bundle(index, bundle_id)
|
||||
# safe_resolve is the ONE in-/out-of-bundle test in this repo, and it is fail-closed. A
|
||||
# model-chosen path is untrusted input by definition, so it goes through the same gate the
|
||||
# navigation walk uses rather than a second, laxer check.
|
||||
return Path(safe_resolve(bundle_dir, path)).read_text(encoding="utf-8")
|
||||
|
||||
return [list_bundles, read_bundle, read_file]
|
||||
|
||||
|
||||
def quick_validate_tool(bundle_dirs: Sequence[str]) -> FunctionTool:
|
||||
"""The hypothesiser's in-loop deterministic check — level 1, and advisory by construction.
|
||||
|
||||
It is the SAME ``validate_proposal`` against the SAME baseline the pipeline will use, so the
|
||||
numbers it reports are real; what it never becomes is provenance. ``ProvenanceStamp`` is
|
||||
written by ``run_project`` alone, and nothing here reaches it.
|
||||
|
||||
Measured cheap (S5: 13.6 ms median through stage 0 + the CBC solve + a 512-sample Monte Carlo,
|
||||
on an anchored base with assumption bands), so it may be called freely in the loop and the
|
||||
contract carries no latency budget.
|
||||
|
||||
``anchored`` is reported alongside the verdict for the same reason
|
||||
``ProvenanceStamp.cost_baseline_anchored`` is a required field: a verdict reached without the
|
||||
project's own cost lines is a weaker claim, and one that does not say so is the silence S4.0's
|
||||
visibility work closed.
|
||||
"""
|
||||
index = _bundle_index(bundle_dirs)
|
||||
|
||||
@tool(
|
||||
name="quick_validate",
|
||||
description=(
|
||||
"Run the deterministic validator over a candidate, as an ADVISORY check. Pass the "
|
||||
"knowledge base id and the candidate as IR JSON. The verdict is not the project's "
|
||||
"decision — the pipeline re-validates and stamps."
|
||||
),
|
||||
)
|
||||
def quick_validate(bundle_id: str, proposal_json: str) -> dict[str, Any]:
|
||||
bundle_dir = _resolve_bundle(index, bundle_id)
|
||||
baseline = okf.load_optional_cost_baseline(bundle_dir)
|
||||
try:
|
||||
proposal = SavingsProposal.model_validate_json(proposal_json)
|
||||
except ValidationError as exc:
|
||||
return {"decision": "unparseable", "reason": str(exc), "anchored": baseline is not None}
|
||||
outcome = validate_proposal(proposal, baseline=baseline)
|
||||
if isinstance(outcome, Rejection):
|
||||
return {
|
||||
"decision": "rejected",
|
||||
"reason": outcome.reason,
|
||||
"anchored": baseline is not None,
|
||||
}
|
||||
return {
|
||||
"decision": "validated",
|
||||
"reason": "",
|
||||
"anchored": baseline is not None,
|
||||
"p10": outcome.p10,
|
||||
"p50": outcome.p50,
|
||||
"p90": outcome.p90,
|
||||
}
|
||||
|
||||
return quick_validate
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# The workflow. One builder, one build, ONE run per exploration (B7 / § C.4).
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def fresh_exploration_workflow(
|
||||
client_factory: Callable[[str], BaseChatClient],
|
||||
*,
|
||||
contract: ExplorationContract,
|
||||
bundle_dirs: Sequence[str] = (),
|
||||
middleware: Sequence[Any] | None = None,
|
||||
) -> Any:
|
||||
"""Build a FRESH Magentic workflow with fresh agents and fresh clients (mirrors
|
||||
``workflow.fresh_workflow``).
|
||||
|
||||
**Fresh builder per exploration, never a reused one.** A built Magentic workflow is
|
||||
single-use: a second ``run()`` on it raises ``RuntimeError`` having made zero model calls
|
||||
(measured, E1) — which is a safe failure, unlike GroupChat 1.9.0's silently empty second run.
|
||||
Reuse is prevented here rather than relied upon to fail.
|
||||
|
||||
``manager_agent_factory=`` rather than ``manager_agent=``: the eager form constructs the
|
||||
manager once and hands the same instance to every ``build()`` (``:1683``, ``:1729-1730``),
|
||||
which on orchestrations 1.0.0 leaked one exploration's task into the next. Orchestrations
|
||||
1.0.1 removed the manager's persistent session, so that leak is GONE and **no test here can
|
||||
tell the two forms apart today** (measured, § F A8: 4/4 → 0/5). The factory form is used
|
||||
anyway because it costs nothing and does not depend on an upstream regression fix staying
|
||||
fixed — but it is stated plainly rather than gated, because a gate that cannot go red proves
|
||||
nothing.
|
||||
|
||||
Every agent carries the SAME ``middleware`` list, manager included. That is the whole of the
|
||||
token guarantee: agent-level ``ChatMiddleware`` does fire on the manager's own calls (measured
|
||||
A1), and the manager is the most talkative participant in the loop.
|
||||
"""
|
||||
hypothesiser_tools: list[Any] = [quick_validate_tool(bundle_dirs)]
|
||||
tools_by_role: dict[str, list[Any]] = {
|
||||
NAVIGATOR_ROLE: list(navigator_tools(bundle_dirs)),
|
||||
HYPOTHESISER_ROLE: hypothesiser_tools,
|
||||
}
|
||||
participants = [
|
||||
Agent(
|
||||
client_factory(role),
|
||||
_INSTRUCTIONS[role],
|
||||
name=role,
|
||||
description=_INSTRUCTIONS[role],
|
||||
tools=tools_by_role[role],
|
||||
middleware=middleware,
|
||||
)
|
||||
for role in PARTICIPANT_ROLES
|
||||
]
|
||||
|
||||
def _manager_agent() -> Agent:
|
||||
return Agent(
|
||||
client_factory(MANAGER_ROLE),
|
||||
"You plan and coordinate an exploration of a project's knowledge bases to find "
|
||||
"cost-saving directions worth testing. You never state a saving figure yourself.",
|
||||
name=MANAGER_ROLE,
|
||||
description="plans the exploration",
|
||||
middleware=middleware,
|
||||
)
|
||||
|
||||
return MagenticBuilder(
|
||||
participants=participants,
|
||||
manager_agent_factory=_manager_agent,
|
||||
max_round_count=contract.max_rounds,
|
||||
max_stall_count=contract.max_stall_count,
|
||||
max_reset_count=contract.max_reset_count,
|
||||
enable_plan_review=contract.enable_plan_review,
|
||||
).build()
|
||||
|
||||
|
||||
def _truthy(answer: Any) -> bool:
|
||||
"""Read a ledger answer exactly as the orchestrator does.
|
||||
|
||||
``MagenticProgressLedgerItem.answer`` is ``str | bool`` and is NOT normalised per field
|
||||
(``:292``); the orchestrator then tests it with plain Python truthiness (``:1106``, ``:1112``).
|
||||
A smarter coercion here — treating the string ``"false"`` as false, say — would describe a run
|
||||
the orchestrator did not have. This layer reports what the loop DID, so it copies the loop's
|
||||
own rule rather than improving on it.
|
||||
"""
|
||||
return bool(answer)
|
||||
|
||||
|
||||
def _plan_text(content: Any) -> str:
|
||||
"""The task-ledger plan as text. ``PLAN_CREATED``/``REPLANNED`` carry a ``Message``."""
|
||||
return str(getattr(content, "text", "") or "")
|
||||
|
||||
|
||||
def _absorb(
|
||||
result: Any,
|
||||
*,
|
||||
ledger_log: list[LedgerEntry],
|
||||
hypotheses: list[str],
|
||||
seen: set[int],
|
||||
span: Any,
|
||||
) -> int:
|
||||
"""Fold one ``run()``'s events into the running log and onto the trace. Returns this call's
|
||||
replans.
|
||||
|
||||
``seen`` de-duplicates by object identity because MAF may surface the same payload under more
|
||||
than one event type; the alternative — pinning one event-type name — would silently under- or
|
||||
over-count if that changed. Participant output is matched on ``executor_id``, which is the
|
||||
agent's name, exactly as ``run._authored_texts`` matches the debate's ``author_name``.
|
||||
|
||||
**The three span events land here** (U14's deferred half, now that it has a call site). They
|
||||
carry the manager's DECISIONS — which plan, who was asked, whether the request was satisfied —
|
||||
as separate typed attributes rather than a rendered sentence, so a collector can query them.
|
||||
Stated honesty limit: they are recorded as the event stream is folded, i.e. after ``run()``
|
||||
returns, so their ORDER is faithful and their timestamps are not the moments the manager acted.
|
||||
Emitting live would mean driving the loop through ``run_stream``, which is a different
|
||||
measurement from the one the plan-review round trip was proved against.
|
||||
"""
|
||||
replans = 0
|
||||
for event in result:
|
||||
data = getattr(event, "data", None)
|
||||
if data is None or id(data) in seen:
|
||||
continue
|
||||
event_type = getattr(data, "event_type", None)
|
||||
if event_type is not None:
|
||||
seen.add(id(data))
|
||||
if event_type == MagenticOrchestratorEventType.PLAN_CREATED:
|
||||
span.add_event("plan_created", {"plan": _plan_text(data.content)})
|
||||
elif event_type == MagenticOrchestratorEventType.REPLANNED:
|
||||
replans += 1
|
||||
span.add_event("replanned", {"plan": _plan_text(data.content)})
|
||||
elif event_type == MagenticOrchestratorEventType.PROGRESS_LEDGER_UPDATED:
|
||||
ledger = data.content
|
||||
speaker = str(ledger.next_speaker.answer)
|
||||
entry = LedgerEntry(
|
||||
round_index=len(ledger_log) + 1,
|
||||
is_request_satisfied=_truthy(ledger.is_request_satisfied.answer),
|
||||
is_in_loop=_truthy(ledger.is_in_loop.answer),
|
||||
is_progress_being_made=_truthy(ledger.is_progress_being_made.answer),
|
||||
next_speaker=speaker,
|
||||
instruction_or_question=str(ledger.instruction_or_question.answer),
|
||||
speaker_known=speaker in PARTICIPANT_ROLES,
|
||||
)
|
||||
ledger_log.append(entry)
|
||||
span.add_event(
|
||||
"progress_ledger_updated",
|
||||
{
|
||||
"round_index": entry.round_index,
|
||||
"is_request_satisfied": entry.is_request_satisfied,
|
||||
"is_in_loop": entry.is_in_loop,
|
||||
"is_progress_being_made": entry.is_progress_being_made,
|
||||
"next_speaker": entry.next_speaker,
|
||||
"speaker_known": entry.speaker_known,
|
||||
"instruction_or_question": entry.instruction_or_question,
|
||||
},
|
||||
)
|
||||
continue
|
||||
if getattr(data, "executor_id", None) == HYPOTHESISER_ROLE:
|
||||
response = getattr(data, "agent_response", None)
|
||||
text = getattr(response, "text", None)
|
||||
if text:
|
||||
seen.add(id(data))
|
||||
hypotheses.append(text)
|
||||
return replans
|
||||
|
||||
|
||||
class HypothesisParseError(ExplorationError):
|
||||
"""A hypothesiser turn carried the marker and then something unreadable.
|
||||
|
||||
Fail-closed, and the marker is what makes that affordable: most turns legitimately are not
|
||||
hypotheses, so an unmarked turn is not a failure and there is nothing to be silent about. A
|
||||
MARKED line that will not parse is a claim the loop tried to make and could not — raising it
|
||||
is the ``write_concept_file`` rule (validation, never repair) rather than the tolerant RAW
|
||||
inbox rule, because this is the product of the run, not a folder anyone may drop things in.
|
||||
"""
|
||||
|
||||
|
||||
def _parse_hypotheses(texts: Sequence[str]) -> list[tuple[str, str]]:
|
||||
"""Every marked ``(label, rationale)`` pair the hypothesiser committed to, in turn order."""
|
||||
found: list[tuple[str, str]] = []
|
||||
for text in texts:
|
||||
for line in text.splitlines():
|
||||
stripped = line.strip()
|
||||
if not stripped.startswith(HYPOTHESIS_MARKER):
|
||||
continue
|
||||
payload = stripped[len(HYPOTHESIS_MARKER) :].strip()
|
||||
try:
|
||||
data = json.loads(payload)
|
||||
except json.JSONDecodeError as exc:
|
||||
raise HypothesisParseError(
|
||||
f"hypothesiser marked a line as a hypothesis and it is not JSON: {exc}; "
|
||||
f"line was: {stripped}"
|
||||
) from exc
|
||||
if not isinstance(data, dict) or not data.get("label") or not data.get("rationale"):
|
||||
raise HypothesisParseError(
|
||||
"a marked hypothesis needs a non-empty 'label' and 'rationale'; got: "
|
||||
f"{stripped}"
|
||||
)
|
||||
found.append((str(data["label"]), str(data["rationale"])))
|
||||
return found
|
||||
|
||||
|
||||
def _mint_approaches(
|
||||
seeds: Sequence[Approach], discovered: Sequence[tuple[str, str]]
|
||||
) -> tuple[Approach, ...]:
|
||||
"""Seeds FIRST, untouched, then one approach per discovered direction.
|
||||
|
||||
Seeds lead because the pipeline evaluates approaches in order under one shared budget: a
|
||||
direction the domain expert asked for must be reached before the loop's own findings spend
|
||||
what is left. The minted ids skip anything a seed already claims, so a seed literally named
|
||||
``hypothesis-1`` cannot collide — ``Mandate`` refuses duplicate ids at construction, and
|
||||
turning an expert's naming choice into a hard failure would be a refusal that teaches nothing.
|
||||
"""
|
||||
taken = {approach.id for approach in seeds}
|
||||
minted: list[Approach] = list(seeds)
|
||||
counter = 0
|
||||
for label, rationale in discovered:
|
||||
counter += 1
|
||||
candidate = f"hypothesis-{counter}"
|
||||
while candidate in taken or candidate == OWN_PROPOSAL_ID:
|
||||
counter += 1
|
||||
candidate = f"hypothesis-{counter}"
|
||||
taken.add(candidate)
|
||||
# description is the hypothesiser's own words, VERBATIM: ``generate._build_messages`` feeds
|
||||
# an Approach.description to the proposer unchanged, and the reason a direction is worth
|
||||
# trying is exactly the half a model cannot re-derive from the cost table.
|
||||
minted.append(Approach(id=candidate, label=label, description=rationale))
|
||||
return tuple(minted)
|
||||
|
||||
|
||||
def _pending_plan_reviews(result: Any) -> list[Any]:
|
||||
return [event for event in result if event.type == "request_info"]
|
||||
|
||||
|
||||
async def explore(
|
||||
prompt: str,
|
||||
*,
|
||||
contract: ExplorationContract,
|
||||
bundle_dirs: Sequence[str] = (),
|
||||
profile: Profile | str = Profile.LOCAL,
|
||||
client_factory: Callable[[str], BaseChatClient] | None = None,
|
||||
seed_approaches: Sequence[Approach] = (),
|
||||
plan_reviewer: PlanReviewer | None = None,
|
||||
meter: TokenMeter | None = None,
|
||||
success_criteria: str = "",
|
||||
) -> ExplorationResult:
|
||||
"""Explore the knowledge bases and return the ``Mandate`` the pipeline should evaluate.
|
||||
|
||||
Writes NOTHING. No outbox artefact, no wiki promotion, no verdict — level 3 of the guarantee
|
||||
table belongs to ``run_project`` alone, and an exploration that could write would be a route
|
||||
around the gate that makes an answer checkable.
|
||||
|
||||
**Three stops, three channels, deliberately not one.** Tokens raise ``BudgetExceeded`` exactly
|
||||
as the debate does, so the hosted surface's 429 arm needs no new case. Rounds raise it too,
|
||||
with ``kind="exploration_rounds"`` — the orchestration itself raises nothing at its round cap
|
||||
(measured, § F): it returns a canonical assistant message that is indistinguishable from
|
||||
success at the transport, so this layer produces the typed stop. Everything semantic —
|
||||
stalling out, an exhausted revision cap, a ledger naming nobody — is a VALUE in ``stop``,
|
||||
because those are outcomes of the exploration rather than exhaustion of a resource, and S3.4
|
||||
split those two apart for a reason.
|
||||
|
||||
``seed_approaches`` are the domain expert's own hypotheses (door 1 of § C.6). They are in the
|
||||
returned mandate whatever the loop found, including when the loop found nothing and including
|
||||
when it stopped early.
|
||||
|
||||
Honesty limits, stated rather than implied. (1) A run is bound to ONE mandate; multi-base
|
||||
dispatch (``Approach.bundle_id``, § C.7) waits for ``run_project`` to accept more than one
|
||||
``bundle_dir``, and shipping the field before its consumer would be a shape guessed instead of
|
||||
measured. (2) The ``quick_validate`` verdicts the hypothesiser saw are not in
|
||||
``ExplorationResult``: they are level-1 advisory, and their home is the
|
||||
``{run_id}-exploration.json`` artefact the CLI wiring writes. (3) The exploration roles resolve
|
||||
through ``resolve_model``'s ``default`` fallback unless an operator maps them explicitly.
|
||||
"""
|
||||
if contract.enable_plan_review and plan_reviewer is None:
|
||||
raise ExplorationError(
|
||||
"enable_plan_review is set but no plan_reviewer was given: the exploration would stop "
|
||||
"at a review nobody can answer, which is a hang rather than a result"
|
||||
)
|
||||
if plan_reviewer is not None and not contract.enable_plan_review:
|
||||
raise ExplorationError(
|
||||
"a plan_reviewer was given but enable_plan_review is false, so no review is ever "
|
||||
"requested and the reviewer would never be called (refused, never silently ignored)"
|
||||
)
|
||||
|
||||
if meter is None:
|
||||
meter = TokenMeter(Budget(max_tokens=contract.max_tokens, max_rounds=contract.max_rounds))
|
||||
if client_factory is None:
|
||||
# Imported HERE, not at module scope: ``run`` imports this module for its ``--explore``
|
||||
# door, so a top-level import would be circular. One copy of the backend factory, never a
|
||||
# second (the (p) rule) — the same move ``run.main`` makes for ``scripted_factory``.
|
||||
from portfolio_optimiser.run import _default_factory
|
||||
|
||||
client_factory = _default_factory(profile)
|
||||
|
||||
workflow = fresh_exploration_workflow(
|
||||
client_factory,
|
||||
contract=contract,
|
||||
bundle_dirs=bundle_dirs,
|
||||
middleware=[BudgetMiddleware(meter)],
|
||||
)
|
||||
|
||||
ledger_log: list[LedgerEntry] = []
|
||||
hypothesis_texts: list[str] = []
|
||||
plan_reviews: list[PlanReview] = []
|
||||
seen: set[int] = set()
|
||||
replans = 0
|
||||
stop: ExplorationStop | None = None
|
||||
|
||||
# ONE span for the whole exploration, opened before the first model call and closed however
|
||||
# the loop ends — including on a BudgetExceeded, which the span records rather than swallows.
|
||||
# With no provider installed this is a no-op tracer and every event is discarded, which is
|
||||
# exactly what "tracing is off" has meant since U14.
|
||||
with exploration_tracer().start_as_current_span(EXPLORATION_SPAN) as span:
|
||||
result = await workflow.run(prompt)
|
||||
while True:
|
||||
replans += _absorb(
|
||||
result,
|
||||
ledger_log=ledger_log,
|
||||
hypotheses=hypothesis_texts,
|
||||
seen=seen,
|
||||
span=span,
|
||||
)
|
||||
pending = _pending_plan_reviews(result)
|
||||
if not pending:
|
||||
break
|
||||
assert plan_reviewer is not None # guarded above; the review implies a reviewer
|
||||
request = pending[0]
|
||||
review = request.data
|
||||
decision = plan_reviewer(
|
||||
PlanReviewRequest(
|
||||
index=len(plan_reviews),
|
||||
plan=str(review.plan),
|
||||
current_progress=str(review.current_progress),
|
||||
is_stalled=_truthy(review.is_stalled),
|
||||
)
|
||||
)
|
||||
if decision.feedback is None:
|
||||
plan_reviews.append(
|
||||
PlanReview(
|
||||
index=len(plan_reviews),
|
||||
plan=str(review.plan),
|
||||
is_stalled=_truthy(review.is_stalled),
|
||||
decision="approve",
|
||||
)
|
||||
)
|
||||
response = MagenticPlanReviewResponse.approve()
|
||||
else:
|
||||
# The revision is recorded whether or not it is APPLIED: ``plan_reviews`` is the
|
||||
# record of what the reviewer DECIDED, and ``stop`` is what says the last one was
|
||||
# refused.
|
||||
applied = sum(1 for entry in plan_reviews if entry.decision == "revise")
|
||||
plan_reviews.append(
|
||||
PlanReview(
|
||||
index=len(plan_reviews),
|
||||
plan=str(review.plan),
|
||||
is_stalled=_truthy(review.is_stalled),
|
||||
decision="revise",
|
||||
feedback=decision.feedback,
|
||||
)
|
||||
)
|
||||
if applied >= contract.max_plan_revisions:
|
||||
# Measured (§ F, A3): a revise costs two manager calls, emits no progress
|
||||
# ledger and consumes no round, then asks AGAIN. Under the round cap alone this
|
||||
# loop never terminates. Stopping is the honest move — forcing an approve the
|
||||
# reviewer did not give would be repair, and repair of a human's decision most
|
||||
# of all.
|
||||
stop = "plan_revisions_exhausted"
|
||||
break
|
||||
response = MagenticPlanReviewResponse.revise(decision.feedback)
|
||||
result = await workflow.run(responses={request.request_id: response})
|
||||
|
||||
if stop is None:
|
||||
stop = _classify_stop(ledger_log, replans=replans, contract=contract)
|
||||
|
||||
discovered = _parse_hypotheses(hypothesis_texts) if stop != "unknown_speaker" else []
|
||||
return ExplorationResult(
|
||||
mandate=Mandate(
|
||||
objective=prompt,
|
||||
approaches=_mint_approaches(seed_approaches, discovered),
|
||||
allow_own_proposals=True,
|
||||
success_criteria=success_criteria,
|
||||
),
|
||||
ledger_log=tuple(ledger_log),
|
||||
stop=stop,
|
||||
plan_reviews=tuple(plan_reviews),
|
||||
)
|
||||
|
||||
|
||||
def _classify_stop(
|
||||
ledger_log: Sequence[LedgerEntry], *, replans: int, contract: ExplorationContract
|
||||
) -> ExplorationStop | None:
|
||||
"""Decide what a finished ``run()`` actually was. Structural, never prose-matched.
|
||||
|
||||
The orchestration ends four ways and reports three of them with an ordinary-looking result:
|
||||
a satisfied request, the round cap, the reset cap, and a ``next_speaker`` matching nobody.
|
||||
Only the first is success. The discriminators used here are counts and ledger flags — the
|
||||
termination MESSAGE is deliberately not matched, because it is an f-string built inline
|
||||
(``:1253``) with no constant to pin against, and a model's own final answer could contain the
|
||||
same words.
|
||||
|
||||
A ledger naming nobody is checked FIRST and outranks everything else. It is the only ending
|
||||
in which the orchestrator produced an answer having asked no participant at all (``:1128``),
|
||||
so the run has a result that no work stands behind — the same class as the retired E2 finding,
|
||||
and not something a later, softer verdict should be allowed to paper over.
|
||||
"""
|
||||
if any(not entry.speaker_known for entry in ledger_log):
|
||||
return "unknown_speaker"
|
||||
if ledger_log and ledger_log[-1].is_request_satisfied:
|
||||
return None
|
||||
if len(ledger_log) >= contract.max_rounds:
|
||||
raise BudgetExceeded("exploration_rounds", contract.max_rounds, len(ledger_log))
|
||||
# Whatever is left ended without satisfying the request and without exhausting the round cap:
|
||||
# the manager reset until it ran out of resets, or terminated with nothing to show. Both are
|
||||
# the same fact for a caller — the exploration gave up — so they share one token rather than
|
||||
# inventing a distinction the ledger cannot support.
|
||||
return "stalled"
|
||||
|
|
@ -37,10 +37,14 @@ cause.
|
|||
(``opentelemetry-exporter-otlp-proto-grpc`` / ``-http``) are not declared dependencies. They are
|
||||
egress, they drag grpc and protobuf into a published wheel for a mode that is off by default, and
|
||||
MAF already raises an ``ImportError`` that names the package to install. Stated honesty limit:
|
||||
``PORTFOLIO_OTEL=otlp`` works only after the operator installs one of them. Equally absent are the
|
||||
``PLAN_CREATED`` / ``REPLANNED`` / ``PROGRESS_LEDGER_UPDATED`` events the plan names — they belong
|
||||
to the exploration loop (U4), which does not exist yet, and an emitter written before its call site
|
||||
is a shape guessed rather than measured.
|
||||
``PORTFOLIO_OTEL=otlp`` works only after the operator installs one of them.
|
||||
|
||||
**The ``PLAN_CREATED`` / ``REPLANNED`` / ``PROGRESS_LEDGER_UPDATED`` events now exist** (U4, økt
|
||||
56). They were held back here in økt 55 on the ground that an emitter written before its call site
|
||||
is a shape guessed rather than measured; the call site is ``explore._absorb``, and the events are
|
||||
recorded on the exploration span this module's ``exploration_tracer`` hands out. Nothing about the
|
||||
contract above changed: with tracing off there is no provider, so those events are discarded like
|
||||
every other span this process makes.
|
||||
|
||||
MAF-touching by construction, so this module never enters the framework-neutral context layer
|
||||
(``okf.py``); the ``test_okf_is_maf_free`` guard keeps that boundary.
|
||||
|
|
@ -233,3 +237,29 @@ def tracing_notice(setup: TracingSetup) -> str | None:
|
|||
f" Tracing: {TRACING_ENV}={MODE_OTLP} — OpenTelemetry spans are EXPORTED OVER THE NETWORK "
|
||||
f"to the endpoints declared below\n{rows}"
|
||||
)
|
||||
|
||||
|
||||
#: The instrumentation scope every exploration span is created under. One name, so a collector
|
||||
#: can select this framework's own spans apart from MAF's (``invoke_agent``, ``workflow.run``)
|
||||
#: without matching on span names that MAF owns and may rename.
|
||||
EXPLORATION_TRACER_NAME: Final = "portfolio_optimiser.explore"
|
||||
|
||||
|
||||
def exploration_tracer() -> Any:
|
||||
"""The tracer the exploration loop records its decisions on.
|
||||
|
||||
``get_tracer`` is safe to call whether or not a provider was installed: with none, OpenTelemetry
|
||||
hands back a no-op tracer and every span and event is discarded. That is the SAME shape MAF's
|
||||
own instrumentation already has (``ENABLE_INSTRUMENTATION`` defaults to True and its spans are
|
||||
thrown away for want of a provider), and it is what lets the exploration emit unconditionally.
|
||||
Gating emission on ``PORTFOLIO_OTEL`` would be a second resolution of a rule this module owns,
|
||||
free to disagree with the providers actually installed.
|
||||
|
||||
A FUNCTION rather than a module-level tracer, and the reason is ordering: ``configure_tracing``
|
||||
runs at process startup, and a tracer bound at import time would have been taken from the
|
||||
global provider that existed BEFORE it — a no-op one, permanently. It is also the seam a test
|
||||
substitutes a local provider through, without installing anything globally.
|
||||
"""
|
||||
from opentelemetry import trace
|
||||
|
||||
return trace.get_tracer(EXPLORATION_TRACER_NAME)
|
||||
|
|
|
|||
990
tests/test_explore_loadbearing.py
Normal file
990
tests/test_explore_loadbearing.py
Normal file
|
|
@ -0,0 +1,990 @@
|
|||
"""U4 + U13-synchronous (økt 56) — the Magentic exploration loop as a MANDATE-FORMER.
|
||||
|
||||
**What this loop is, and what it deliberately is not.** ``explore()`` puts a Magentic manager
|
||||
*over* the normative pipeline, never inside it: the manager is free to choose which knowledge base
|
||||
to open and which hypothesis to shape next, and what leaves that freedom is a
|
||||
``mandate.Mandate`` — a list of approaches worth *testing*. It is never a proposal. Every number
|
||||
that survives is still gated by ``validate_proposal`` inside ``run_project``, in the same blocking
|
||||
gate as today, and the exploration itself can write to neither the outbox nor the wiki. Step 3's
|
||||
maker-checker debate is untouched (``shared/method-spec.md`` §3 is commons-owned and normative).
|
||||
|
||||
**Everything asserted here was measured before it was built** (plan
|
||||
``docs/plan/2026-08-23-magentic-utforskningssloeyfe.md`` § F, spikes S0–S6 in økt 54, plus three
|
||||
probes run at the head of økt 56):
|
||||
|
||||
* a plan-review ``revise`` costs two manager calls, **zero** rounds, and asks *again* — so an
|
||||
always-revising expert is unbounded spend under a round cap that never ticks. That is the whole
|
||||
reason ``max_plan_revisions`` is a required contract field rather than a nicety.
|
||||
* the round cap and the reset cap **raise nothing**. Both end the run with a canonical assistant
|
||||
message and a normal-looking result (measured: ``max_round_count=2`` → two ledger events and
|
||||
``'Workflow terminated due to reaching maximum round count.'``; a stalling ledger with
|
||||
``max_reset_count=1`` → one ``REPLANNED`` event and ``'…maximum reset count.'``). At the
|
||||
transport both are indistinguishable from success, so this layer produces the typed stop itself.
|
||||
* a ``next_speaker`` naming nobody produces a **silent final answer with zero participant work**
|
||||
(``_magentic.py:1128-1131``) — a plausible answer produced by no work at all, which is the
|
||||
hazard class E2 was retired for. The names are therefore validated, never assumed.
|
||||
|
||||
**The client is the repo's own ``ScriptedChatClient``.** A bare ``BaseChatClient`` silently no-ops
|
||||
``BudgetMiddleware`` (measured, ``simulation.py:373-375``), so a budget claim proved against one
|
||||
would prove nothing.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
from collections.abc import Callable
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
from agent_framework import BaseChatClient
|
||||
from opentelemetry.sdk.trace import TracerProvider
|
||||
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
|
||||
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
|
||||
from pydantic import ValidationError
|
||||
|
||||
import portfolio_optimiser
|
||||
from portfolio_optimiser import explore, okf
|
||||
from portfolio_optimiser.budget import Budget, BudgetExceeded, TokenMeter
|
||||
from portfolio_optimiser.explore import ExplorationContract
|
||||
from portfolio_optimiser.mandate import Approach
|
||||
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# C.3 — the contract: an exploration without stated bounds refuses to start
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
_FULL_CONTRACT = {
|
||||
"max_rounds": 4,
|
||||
"max_tokens": 5_000,
|
||||
"max_stall_count": 2,
|
||||
"max_reset_count": 1,
|
||||
"max_plan_revisions": 1,
|
||||
"enable_plan_review": True,
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("omitted", sorted(_FULL_CONTRACT))
|
||||
def test_every_bound_is_required_with_no_default(omitted: str) -> None:
|
||||
"""T1: each of the six fields is REQUIRED — dropping any one refuses construction.
|
||||
|
||||
Not a style point. ``MagenticBuilder`` defaults ``max_round_count`` to ``None`` (unbounded)
|
||||
and ``max_reset_count`` to ``None`` (unlimited), and inheriting either would give this repo
|
||||
the one thing ``shared/method-spec.md`` §8 forbids outright: a loop with no stated end. A
|
||||
default here would also be a claim about the operator's intent that nobody made — the same
|
||||
ground on which ``ProvenanceStamp.cost_baseline_anchored`` is required without one.
|
||||
"""
|
||||
payload = {k: v for k, v in _FULL_CONTRACT.items() if k != omitted}
|
||||
with pytest.raises(ValidationError):
|
||||
ExplorationContract(**payload)
|
||||
|
||||
|
||||
def test_full_contract_constructs() -> None:
|
||||
"""T2: the control for T1 — the complete payload IS valid.
|
||||
|
||||
Without it, T1 would pass on a model that refuses everything, which is the vacuous-gate class
|
||||
this repo has paid for six times.
|
||||
"""
|
||||
contract = ExplorationContract(**_FULL_CONTRACT)
|
||||
assert contract.max_rounds == 4
|
||||
assert contract.enable_plan_review is True
|
||||
|
||||
|
||||
def test_a_revision_cap_without_plan_review_is_refused_not_ignored() -> None:
|
||||
"""T3: ``max_plan_revisions > 0`` with ``enable_plan_review=False`` refuses.
|
||||
|
||||
A plan revision can only arise from a plan review — with the review off, the cap bounds an
|
||||
event that cannot occur, and a caller who set it believes they bounded something. This repo
|
||||
refuses a setting that cannot take effect rather than dropping it silently (the same partition
|
||||
``--embedder-config requires --semantic-retrieval`` enforces on the CLI).
|
||||
"""
|
||||
with pytest.raises(ValidationError):
|
||||
ExplorationContract(**{**_FULL_CONTRACT, "enable_plan_review": False})
|
||||
|
||||
|
||||
def test_review_off_with_zero_revisions_is_the_coherent_form() -> None:
|
||||
"""T4: the control for T3 — review off and the cap at ``0`` is a consistent statement, and
|
||||
must construct. Without this arm T3 would pass on a model that simply forbade
|
||||
``enable_plan_review=False`` outright, which is a different (and wrong) rule.
|
||||
"""
|
||||
contract = ExplorationContract(
|
||||
**{**_FULL_CONTRACT, "enable_plan_review": False, "max_plan_revisions": 0}
|
||||
)
|
||||
assert contract.enable_plan_review is False
|
||||
assert contract.max_plan_revisions == 0
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# The scripted stand-ins. ScriptedChatClient, never a bare BaseChatClient: the latter no-ops
|
||||
# BudgetMiddleware (measured, simulation.py:373-375), so a budget assertion made against one
|
||||
# would assert nothing.
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
PROMPT = "Find the cheapest saving available in the energy bundle."
|
||||
|
||||
|
||||
def _ledger_json(
|
||||
*, satisfied: bool, speaker: str, instruction: str = "Shape one hypothesis."
|
||||
) -> str:
|
||||
"""A progress ledger naming ``speaker``.
|
||||
|
||||
The name is a PARAMETER, never a literal, because a ``next_speaker`` matching no participant
|
||||
is the measured footgun this module defends against: the orchestrator does not error, it
|
||||
quietly emits a final answer having asked nobody (``_magentic.py:1128-1131``).
|
||||
"""
|
||||
return json.dumps(
|
||||
{
|
||||
"is_request_satisfied": {"reason": "r", "answer": satisfied},
|
||||
"is_in_loop": {"reason": "r", "answer": False},
|
||||
"is_progress_being_made": {"reason": "r", "answer": True},
|
||||
"next_speaker": {"reason": "r", "answer": speaker},
|
||||
"instruction_or_question": {"reason": "r", "answer": instruction},
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def _stalling_ledger_json(speaker: str) -> str:
|
||||
"""A ledger reporting NO progress and a loop — the two flags that drive ``stall_count`` up."""
|
||||
return json.dumps(
|
||||
{
|
||||
"is_request_satisfied": {"reason": "r", "answer": False},
|
||||
"is_in_loop": {"reason": "circles", "answer": True},
|
||||
"is_progress_being_made": {"reason": "none", "answer": False},
|
||||
"next_speaker": {"reason": "r", "answer": speaker},
|
||||
"instruction_or_question": {"reason": "r", "answer": "Try again."},
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def _manager_script(
|
||||
ledgers: list[str], calls: list[str] | None = None
|
||||
) -> Callable[[str, str], str]:
|
||||
"""Route a manager prompt blob to its scripted reply, consuming ``ledgers`` in order.
|
||||
|
||||
The ORDER of these tests is load-bearing and was measured (§ F, A6): the selector receives the
|
||||
CONCATENATION of every message in the call, so a later-stage prompt still carries the earlier
|
||||
stage's text — one manager call in five carries two markers. Testing the later stage FIRST is
|
||||
what resolves it; reversing two of these silently reattributes a reply to the wrong stage.
|
||||
"""
|
||||
|
||||
def _select(blob: str, _role: str) -> str:
|
||||
if calls is not None:
|
||||
calls.append(blob[:40])
|
||||
if "provide the final answer" in blob:
|
||||
return "FINAL: exploration done."
|
||||
if "pure JSON format" in blob:
|
||||
return ledgers.pop(0) if ledgers else _ledger_json(satisfied=True, speaker="navigator")
|
||||
if "went wrong on this last run" in blob:
|
||||
return "PLAN-UPDATE: revised plan."
|
||||
if "rewrite the following fact sheet" in blob:
|
||||
return "FACTS-UPDATE: revised facts."
|
||||
if "bullet-point plan" in blob:
|
||||
return "PLAN: - ask the hypothesiser"
|
||||
if "pre-survey" in blob:
|
||||
return "FACTS: the bundle is anchored."
|
||||
return "{}"
|
||||
|
||||
return _select
|
||||
|
||||
|
||||
def _factory(
|
||||
*, ledgers: list[str], hypothesiser: list[str], navigator: str = "NAVIGATOR: index read."
|
||||
) -> Callable[[str], BaseChatClient]:
|
||||
"""One fresh ``ScriptedChatClient`` per role, exactly as the real factory hands out one per
|
||||
role. ``hypothesiser`` is a list consumed in order, so a run can shape several candidates."""
|
||||
|
||||
def factory(role: str) -> BaseChatClient:
|
||||
if role == explore.MANAGER_ROLE:
|
||||
return ScriptedChatClient(reply_selector=_manager_script(ledgers), role=role)
|
||||
if role == explore.HYPOTHESISER_ROLE:
|
||||
replies = list(hypothesiser)
|
||||
|
||||
def _hyp(_blob: str, _role: str) -> str:
|
||||
return replies.pop(0) if replies else "nothing further."
|
||||
|
||||
return ScriptedChatClient(reply_selector=_hyp, role=role)
|
||||
return ScriptedChatClient(navigator, role=role)
|
||||
|
||||
return factory
|
||||
|
||||
|
||||
def _hypothesis_line(label: str, rationale: str) -> str:
|
||||
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps({"label": label, "rationale": rationale})
|
||||
|
||||
|
||||
#: The no-review base every stop test derives from. ``enable_plan_review`` and
|
||||
#: ``max_plan_revisions`` move together — ``ExplorationContract`` refuses them apart — so a test
|
||||
#: about round or stall behaviour has to say so explicitly rather than inherit ``_FULL_CONTRACT``.
|
||||
_NO_REVIEW = {**_FULL_CONTRACT, "enable_plan_review": False, "max_plan_revisions": 0}
|
||||
|
||||
_CONTRACT = ExplorationContract(
|
||||
max_rounds=6,
|
||||
max_tokens=100_000,
|
||||
max_stall_count=2,
|
||||
max_reset_count=1,
|
||||
max_plan_revisions=0,
|
||||
enable_plan_review=False,
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# C.0 / C.6 — the exploration is a MANDATE-FORMER, and a seed never disappears
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_hypotheses_become_the_mandate_in_the_order_they_were_shaped() -> None:
|
||||
"""T5: what the hypothesiser MARKED becomes ``Mandate.approaches``, rationale VERBATIM.
|
||||
|
||||
The rationale is the half a model cannot re-derive from cost data — ``mandate.Approach``
|
||||
already feeds ``description`` to the proposer verbatim (``generate._build_messages``), so
|
||||
paraphrasing it here would drop precisely the part the exploration exists to carry forward.
|
||||
"""
|
||||
ledgers = [
|
||||
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
|
||||
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
|
||||
]
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(),
|
||||
client_factory=_factory(
|
||||
ledgers=ledgers,
|
||||
hypothesiser=[
|
||||
"Looking at the bundle.\n"
|
||||
+ _hypothesis_line("LED retrofit", "the fixtures are 1990s fluorescent")
|
||||
],
|
||||
),
|
||||
)
|
||||
|
||||
assert [a.label for a in result.mandate.approaches] == ["LED retrofit"]
|
||||
assert result.mandate.approaches[0].description == "the fixtures are 1990s fluorescent"
|
||||
assert result.mandate.objective == PROMPT
|
||||
assert result.stop is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_seed_approach_survives_whatever_the_manager_found() -> None:
|
||||
"""T6: an expert's own hypothesis is in the output mandate, FIRST, untouched.
|
||||
|
||||
Door 1 of § C.6, and the ``not_evaluated`` rule applied one stage earlier: a direction the
|
||||
domain expert asked for may never vanish because an autonomous loop preferred its own. Seeds
|
||||
lead so the pipeline reaches them before spending its budget on discovered ones.
|
||||
"""
|
||||
seed = Approach(id="fagperson-1", label="Night setback", description="the expert's own words")
|
||||
ledgers = [
|
||||
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
|
||||
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
|
||||
]
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(),
|
||||
seed_approaches=(seed,),
|
||||
client_factory=_factory(
|
||||
ledgers=ledgers,
|
||||
hypothesiser=[_hypothesis_line("LED retrofit", "fluorescent fixtures")],
|
||||
),
|
||||
)
|
||||
|
||||
assert [a.id for a in result.mandate.approaches] == ["fagperson-1", "hypothesis-1"]
|
||||
assert result.mandate.approaches[0] == seed
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_exploration_that_shaped_nothing_still_returns_the_seeds() -> None:
|
||||
"""T7: the control for T6 — with the hypothesiser silent, the seed is still the mandate.
|
||||
|
||||
This is what makes T6 a statement about PRESERVATION rather than about ordering: a test that
|
||||
only ever saw seeds alongside discoveries could not tell "seeds are kept" from "seeds sort
|
||||
first".
|
||||
"""
|
||||
seed = Approach(id="fagperson-1", label="Night setback", description="the expert's own words")
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(),
|
||||
seed_approaches=(seed,),
|
||||
client_factory=_factory(
|
||||
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
|
||||
hypothesiser=[],
|
||||
),
|
||||
)
|
||||
|
||||
assert result.mandate.approaches == (seed,)
|
||||
assert result.mandate.allow_own_proposals is True
|
||||
|
||||
|
||||
def test_zero_resets_is_refused_because_it_silently_explores_nothing() -> None:
|
||||
"""T8: ``max_reset_count=0`` refuses — MEASURED, not reasoned.
|
||||
|
||||
The orchestrator's limit check is ``reset_count >= max_reset_count`` (``_magentic.py:1243``)
|
||||
and ``reset_count`` starts at zero, so a cap of zero is already met before the first round.
|
||||
Measured against the installed stack: the run makes only the ``facts`` and ``plan`` manager
|
||||
calls, emits **zero** progress-ledger events, and returns
|
||||
``'Workflow terminated due to reaching maximum reset count.'`` — an exploration that explored
|
||||
nothing, reported as a stall that never happened. An operator writing "allow no resets" would
|
||||
get "do no work", quietly. So it is refused at construction, where the reason can be said.
|
||||
"""
|
||||
with pytest.raises(ValidationError):
|
||||
ExplorationContract(**{**_FULL_CONTRACT, "max_reset_count": 0})
|
||||
|
||||
|
||||
def test_zero_stalls_is_allowed_because_it_means_something() -> None:
|
||||
"""T9: the control for T8 — ``max_stall_count=0`` is a real setting and must construct.
|
||||
|
||||
The stall check is STRICT (``stall_count > max_stall_count``, ``:1118``) and the counter is
|
||||
incremented before it, so zero means "reset on the first round that reports no progress".
|
||||
That is strictness, not self-defeat, and refusing both zeroes on symmetry would have banned a
|
||||
usable configuration on the strength of a measurement about a different field.
|
||||
"""
|
||||
contract = ExplorationContract(**{**_FULL_CONTRACT, "max_stall_count": 0})
|
||||
assert contract.max_stall_count == 0
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# C.2 / C.3 — three endings the orchestration reports as if they were success
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_round_cap_leaves_as_a_typed_budget_stop() -> None:
|
||||
"""T10: the round cap becomes ``BudgetExceeded(kind="exploration_rounds")``.
|
||||
|
||||
Measured (§ F, E5, re-measured at the head of this økt): ``max_round_count`` raises NOTHING.
|
||||
The run ends with the assistant message ``'Workflow terminated due to reaching maximum round
|
||||
count.'`` and a result that ``get_outputs()`` answers like any other — at the transport it is
|
||||
indistinguishable from a finished exploration. Left alone, a caller would read a run that
|
||||
explored two rounds of a six-round question as a completed answer. The triple is the one
|
||||
kø-(y) defends: WHICH cap bound, what it was, and how far the run actually got.
|
||||
"""
|
||||
contract = ExplorationContract(**{**_NO_REVIEW, "max_rounds": 2})
|
||||
with pytest.raises(BudgetExceeded) as excinfo:
|
||||
await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
client_factory=_factory(
|
||||
ledgers=[_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE)] * 4,
|
||||
hypothesiser=[_hypothesis_line("LED", "worth a look")] * 4,
|
||||
),
|
||||
)
|
||||
|
||||
assert excinfo.value.kind == "exploration_rounds"
|
||||
assert excinfo.value.limit == 2
|
||||
assert excinfo.value.observed == 2
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_request_satisfied_on_the_last_allowed_round_is_success() -> None:
|
||||
"""T11: the discriminator for T10 — reaching the cap is not the same as being cut off by it.
|
||||
|
||||
Both runs end with exactly ``max_rounds`` progress-ledger events, so a check written on the
|
||||
count alone would raise on this one too and turn a completed exploration into a budget error.
|
||||
What separates them is the LAST ledger's ``is_request_satisfied``, which is also what the
|
||||
orchestrator itself branches on (``:1106``). Without this arm, T10 would pass on an
|
||||
implementation that refuses every exploration that uses its whole allowance.
|
||||
"""
|
||||
contract = ExplorationContract(**{**_NO_REVIEW, "max_rounds": 2})
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
client_factory=_factory(
|
||||
ledgers=[
|
||||
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
|
||||
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
|
||||
],
|
||||
hypothesiser=[_hypothesis_line("LED", "worth a look")],
|
||||
),
|
||||
)
|
||||
|
||||
assert len(result.ledger_log) == contract.max_rounds
|
||||
assert result.stop is None
|
||||
assert [a.label for a in result.mandate.approaches] == ["LED"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_stalling_out_is_a_typed_value_never_an_exception() -> None:
|
||||
"""T12: stall → reset → out of resets is ``stop="stalled"``, and the run still returns.
|
||||
|
||||
Kept as a VALUE while the round cap RAISES, and the split is S3.4's, not a preference: a
|
||||
stalled exploration is an outcome (the manager tried and got nowhere), whereas an exhausted
|
||||
round or token cap is resource exhaustion. Fusing them would leave a caller unable to tell
|
||||
"there was nothing here" from "we could not afford to look".
|
||||
"""
|
||||
contract = ExplorationContract(
|
||||
**{**_NO_REVIEW, "max_rounds": 6, "max_stall_count": 1, "max_reset_count": 1}
|
||||
)
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
client_factory=_factory(
|
||||
ledgers=[_stalling_ledger_json(explore.HYPOTHESISER_ROLE)] * 6,
|
||||
hypothesiser=["still nothing."] * 6,
|
||||
),
|
||||
)
|
||||
|
||||
assert result.stop == "stalled"
|
||||
assert len(result.ledger_log) < contract.max_rounds
|
||||
assert all(entry.is_in_loop for entry in result.ledger_log)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_ledger_naming_nobody_withholds_what_the_run_produced() -> None:
|
||||
"""T13: a ``next_speaker`` matching no participant stops the exploration and drops its finds.
|
||||
|
||||
The measured footgun (``_magentic.py:1128-1131``): the orchestrator neither raises nor retries
|
||||
on an unknown speaker — it logs a warning and jumps to ``_prepare_final_answer``. The run
|
||||
therefore returns a plausible answer that no participant was asked for. This is the shape E2
|
||||
was retired over ("a plausible verdict produced by zero work"), so the mandate is NOT built
|
||||
from what such a run said it found.
|
||||
|
||||
The scripted run reaches the bad ledger on round TWO, after a good round in which the
|
||||
hypothesiser really did commit to a direction. That ordering is what makes the assertion
|
||||
sharp: with the bad ledger first, nobody would ever have spoken and "nothing was carried
|
||||
forward" would be true of any implementation at all.
|
||||
"""
|
||||
seed = Approach(id="fagperson-1", label="Night setback", description="expert's own")
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(),
|
||||
seed_approaches=(seed,),
|
||||
client_factory=_factory(
|
||||
ledgers=[
|
||||
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
|
||||
_ledger_json(satisfied=False, speaker="a-name-nobody-answers-to"),
|
||||
],
|
||||
hypothesiser=[_hypothesis_line("LED retrofit", "fluorescent fixtures")],
|
||||
),
|
||||
)
|
||||
|
||||
assert result.stop == "unknown_speaker"
|
||||
assert result.ledger_log[0].speaker_known is True
|
||||
assert result.ledger_log[-1].speaker_known is False
|
||||
# The seed survives — preservation is unconditional (§ C.6 door 1) — while the loop's own
|
||||
# find does not, because nothing stands behind the turn that ended the run.
|
||||
assert result.mandate.approaches == (seed,)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# C.2 — the token cap covers the MANAGER, which is the loop's most talkative agent
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_token_cap_binds_the_manager_before_any_participant_speaks() -> None:
|
||||
"""T14: a one-token budget stops the exploration on the MANAGER's own first call.
|
||||
|
||||
Agent-level ``ChatMiddleware`` does fire on the manager's calls (§ F, A1, measured green), and
|
||||
the manager talks more than anyone else in a Magentic loop — it extracts facts, writes the
|
||||
plan, and writes a progress ledger every single round. A cap fastened only to the participants
|
||||
would be a cap in name.
|
||||
|
||||
The assertion is deliberately not "something raised". ``kind == "tokens"`` separates it from
|
||||
the round-cap stop, ``meter.tokens == 8`` shows the charge came from a call that was actually
|
||||
made and metered, and the EMPTY ledger log shows it landed before the loop had run a single
|
||||
round — which is exactly what a manager-attached middleware does and a participant-only one
|
||||
cannot.
|
||||
"""
|
||||
meter = TokenMeter(Budget(max_tokens=1, max_rounds=6))
|
||||
contract = ExplorationContract(**{**_NO_REVIEW, "max_tokens": 1})
|
||||
|
||||
with pytest.raises(BudgetExceeded) as excinfo:
|
||||
await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
meter=meter,
|
||||
client_factory=_factory(
|
||||
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
|
||||
hypothesiser=[],
|
||||
),
|
||||
)
|
||||
|
||||
assert excinfo.value.kind == "tokens"
|
||||
assert meter.tokens == 8, "the manager's own call must have been charged to the meter"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# C.5 / U13 — the synchronous plan review, and the cap the measurement forced
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _reviewer(script: list[explore.PlanReviewDecision], seen: list[explore.PlanReviewRequest]):
|
||||
def review(request: explore.PlanReviewRequest) -> explore.PlanReviewDecision:
|
||||
seen.append(request)
|
||||
return script.pop(0) if script else explore.PlanReviewDecision.approve()
|
||||
|
||||
return review
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_revision_reaches_the_manager_and_the_review_is_asked_again() -> None:
|
||||
"""T15: revise → replan → asked AGAIN → approve → the loop runs.
|
||||
|
||||
This is målbilde's "ask the question, use the answer, carry on" on the installed stack: the
|
||||
expert's words go into the manager's history, the manager replans, and the human is asked to
|
||||
sign off on the NEW plan rather than the old one. Both round trips are recorded, in order,
|
||||
with the feedback verbatim — an audit of what a human actually told an autonomous loop is
|
||||
worth nothing paraphrased.
|
||||
"""
|
||||
seen: list[explore.PlanReviewRequest] = []
|
||||
contract = ExplorationContract(
|
||||
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 2}
|
||||
)
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
plan_reviewer=_reviewer(
|
||||
[explore.PlanReviewDecision.revise("Also test night setback.")], seen
|
||||
),
|
||||
client_factory=_factory(
|
||||
ledgers=[
|
||||
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
|
||||
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
|
||||
],
|
||||
hypothesiser=[_hypothesis_line("Night setback", "the expert asked for it")],
|
||||
),
|
||||
)
|
||||
|
||||
assert [r.decision for r in result.plan_reviews] == ["revise", "approve"]
|
||||
assert result.plan_reviews[0].feedback == "Also test night setback."
|
||||
assert len(seen) == 2, "a revision must produce a SECOND review, not resume silently"
|
||||
assert seen[1].plan != "", "the second review must show the revised plan"
|
||||
assert result.stop is None
|
||||
assert [a.label for a in result.mandate.approaches] == ["Night setback"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_always_revising_reviewer_is_stopped_by_the_cap() -> None:
|
||||
"""T16: the cap terminates a reviewer that never signs off — the reason it exists.
|
||||
|
||||
Measured (§ F, A3): a revise costs two manager calls, emits NO progress ledger and consumes
|
||||
NO round, then asks again. The round cap therefore never ticks, and without
|
||||
``max_plan_revisions`` this is an unbounded spend under caps that all look satisfied —
|
||||
precisely what ``shared/method-spec.md`` §8 forbids. The stop is typed and the exploration
|
||||
still returns; the reviewer's last (refused) revision is recorded, because the record is of
|
||||
what the human decided and ``stop`` is what says it was not applied.
|
||||
"""
|
||||
seen: list[explore.PlanReviewRequest] = []
|
||||
always_revise = [explore.PlanReviewDecision.revise(f"Again #{n}.") for n in range(10)]
|
||||
contract = ExplorationContract(
|
||||
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 1}
|
||||
)
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
plan_reviewer=_reviewer(always_revise, seen),
|
||||
client_factory=_factory(
|
||||
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
|
||||
hypothesiser=[],
|
||||
),
|
||||
)
|
||||
|
||||
assert result.stop == "plan_revisions_exhausted"
|
||||
assert [r.decision for r in result.plan_reviews] == ["revise", "revise"]
|
||||
assert result.ledger_log == (), "the loop must never have run: the plan was never approved"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_reviewer_that_signs_off_at_once_is_not_capped() -> None:
|
||||
"""T17: the control for T16 — the same cap, a reviewer that approves, and no stop.
|
||||
|
||||
Without it, T16 would pass on an implementation that refuses every plan review it is given,
|
||||
which would stop the runaway loop and every legitimate one with it.
|
||||
"""
|
||||
seen: list[explore.PlanReviewRequest] = []
|
||||
contract = ExplorationContract(
|
||||
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 1}
|
||||
)
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
plan_reviewer=_reviewer([], seen),
|
||||
client_factory=_factory(
|
||||
ledgers=[_ledger_json(satisfied=True, speaker=explore.NAVIGATOR_ROLE)],
|
||||
hypothesiser=[],
|
||||
),
|
||||
)
|
||||
|
||||
assert result.stop is None
|
||||
assert [r.decision for r in result.plan_reviews] == ["approve"]
|
||||
assert len(result.ledger_log) >= 1, "an approved plan must let the loop actually run"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# C.0 level 3 — the exploration has no write access, and level 1 is advisory
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _tree(root: Path) -> dict[str, bytes]:
|
||||
return {
|
||||
str(p.relative_to(root)): p.read_bytes() for p in sorted(root.rglob("*")) if p.is_file()
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_exploration_leaves_the_knowledge_base_byte_identical(tmp_path: Path) -> None:
|
||||
"""T18: ``explore()`` writes NOTHING — not to the base, not anywhere under it.
|
||||
|
||||
Level 3 of the guarantee table: only the pipeline may write an outbox artefact, and only the
|
||||
gated ``promote_verdict`` may write to the wiki. An exploration that could write would be a
|
||||
route around the gate that makes an answer checkable — and, promoting into the base it reads,
|
||||
the self-contamination loop the Step-8 gate exists to prevent.
|
||||
|
||||
Compared BYTE for byte over the whole subtree rather than by listing names, so a rewritten
|
||||
``index.md`` of the same length would still fail.
|
||||
|
||||
**The tools are exercised DIRECTLY, and that is a correction, not thoroughness.** A first
|
||||
version of this test drove only ``explore()`` — and a mutation that made ``read_bundle`` write
|
||||
a file into the base it reads left the WHOLE suite green (measured: 974 passed). A
|
||||
``ScriptedChatClient`` returns text and never emits a tool call, so no scripted run reaches a
|
||||
tool body: the read surface, which is the only place a write could plausibly come from, was
|
||||
outside the gate entirely.
|
||||
"""
|
||||
base = tmp_path / "bygg-energi-baseline-mikro"
|
||||
shutil.copytree(
|
||||
Path(portfolio_optimiser.__file__).parent / "data" / "bundles" / base.name, base
|
||||
)
|
||||
before = _tree(tmp_path)
|
||||
|
||||
result = await explore.explore(
|
||||
PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(base),),
|
||||
client_factory=_factory(
|
||||
# Three ledgers, and the third is what makes the second one matter: the orchestrator
|
||||
# tests ``is_request_satisfied`` BEFORE it reads ``next_speaker`` (``:1106``), so a
|
||||
# satisfied ledger naming the hypothesiser never actually asks it anything.
|
||||
ledgers=[
|
||||
_ledger_json(satisfied=False, speaker=explore.NAVIGATOR_ROLE),
|
||||
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
|
||||
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
|
||||
],
|
||||
hypothesiser=[_hypothesis_line("LED retrofit", "fluorescent fixtures")],
|
||||
),
|
||||
)
|
||||
|
||||
assert result.mandate.approaches[0].label == "LED retrofit"
|
||||
|
||||
# Every read tool, called on the same base, with model-shaped arguments.
|
||||
tools = {t.name: t for t in explore.navigator_tools((str(base),))}
|
||||
assert tools["list_bundles"].func()[0]["id"] == base.name
|
||||
assert tools["read_bundle"].func(bundle_id=base.name) != ""
|
||||
assert tools["read_file"].func(bundle_id=base.name, path="index.md") != ""
|
||||
explore.quick_validate_tool((str(base),)).func(
|
||||
bundle_id=base.name, proposal_json=json.dumps(_micro_projection())
|
||||
)
|
||||
|
||||
assert _tree(tmp_path) == before
|
||||
|
||||
|
||||
def _micro_bundle_dir() -> str:
|
||||
return str(
|
||||
Path(portfolio_optimiser.__file__).parent
|
||||
/ "data"
|
||||
/ "bundles"
|
||||
/ "bygg-energi-baseline-mikro"
|
||||
)
|
||||
|
||||
|
||||
def _micro_projection() -> dict[str, Any]:
|
||||
projection = dict(okf.load_ir_projection(_micro_bundle_dir()))
|
||||
projection.pop("_note", None)
|
||||
return projection
|
||||
|
||||
|
||||
def test_quick_validate_reports_the_real_verdict_and_says_whether_it_was_anchored() -> None:
|
||||
"""T19: the in-loop check is the SAME validator, and it declares its own anchoring.
|
||||
|
||||
Level 1 is advisory but never fake: it runs ``validate_proposal`` against the base's own
|
||||
``cost-baseline.json``, so stage 0 reconciliation is live and a fabricated cost line is caught
|
||||
in the loop rather than three steps later. ``anchored`` rides along for the reason
|
||||
``ProvenanceStamp.cost_baseline_anchored`` is a required field — a verdict reached without the
|
||||
project's real cost lines is a weaker claim, and one that does not say so is a silence.
|
||||
"""
|
||||
base = _micro_bundle_dir()
|
||||
projection = _micro_projection()
|
||||
validate = explore.quick_validate_tool((base,))
|
||||
|
||||
honest = validate.func(
|
||||
bundle_id="bygg-energi-baseline-mikro", proposal_json=json.dumps(projection)
|
||||
)
|
||||
assert honest["decision"] == "validated"
|
||||
assert honest["anchored"] is True
|
||||
assert honest["p90"] >= honest["p50"] >= honest["p10"]
|
||||
|
||||
# A cost code the project does not have is refused by stage 0 — the one stage that can tell a
|
||||
# fabricated line from a real one, and the reason `anchored` is worth reporting at all.
|
||||
invented = dict(projection)
|
||||
invented["affected_items"] = [
|
||||
{**dict(projection["affected_items"][0]), "code": "CODE-THAT-DOES-NOT-EXIST"}
|
||||
]
|
||||
fabricated = validate.func(
|
||||
bundle_id="bygg-energi-baseline-mikro", proposal_json=json.dumps(invented)
|
||||
)
|
||||
assert fabricated["decision"] == "rejected"
|
||||
assert "CODE-THAT-DOES-NOT-EXIST" in fabricated["reason"]
|
||||
|
||||
|
||||
def test_an_unknown_knowledge_base_is_refused_by_name() -> None:
|
||||
"""T20: a tool call naming a base nobody configured refuses, and says what IS configured.
|
||||
|
||||
Model-chosen arguments are untrusted input. Answering an unknown id with an empty result would
|
||||
let the manager conclude the base is empty rather than absent — the fourth face of the
|
||||
verification law, arrived at through a tool rather than a query.
|
||||
"""
|
||||
validate = explore.quick_validate_tool(("/tmp/base-a",))
|
||||
with pytest.raises(explore.ExplorationError) as excinfo:
|
||||
validate.func(bundle_id="base-b", proposal_json="{}")
|
||||
assert "base-a" in str(excinfo.value)
|
||||
|
||||
|
||||
def test_two_bases_with_the_same_name_are_refused() -> None:
|
||||
"""T21: duplicate ids refuse at construction — the S3.2 key-collision class, one layer up.
|
||||
|
||||
The id is how the manager names a base. Two bases answering to one name would let it read A
|
||||
while believing it read B, and every quotation it produced afterwards would be attributed to
|
||||
the wrong project.
|
||||
"""
|
||||
with pytest.raises(explore.ExplorationError):
|
||||
explore.navigator_tools(("/tmp/one/shared-name", "/tmp/two/shared-name"))
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
# U14 — the three events the tracing seam was landed for, now that they have a call site
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _recording_tracer() -> tuple[Any, InMemorySpanExporter]:
|
||||
"""A REAL OpenTelemetry tracer over an in-memory exporter — not a spy.
|
||||
|
||||
A recorder standing in for ``add_event`` would prove that this module calls something shaped
|
||||
like OTel; this proves the events survive the actual SDK, with the attribute types it will
|
||||
accept. The provider is LOCAL and is never installed globally, so the pytest process keeps
|
||||
whatever tracing configuration it had (the same restraint U14's own tests exercise).
|
||||
"""
|
||||
provider = TracerProvider()
|
||||
exporter = InMemorySpanExporter()
|
||||
provider.add_span_processor(SimpleSpanProcessor(exporter))
|
||||
return provider.get_tracer("test"), exporter
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_three_orchestrator_events_reach_the_trace(monkeypatch: Any) -> None:
|
||||
"""T22: ``plan_created``, ``replanned`` and ``progress_ledger_updated`` are recorded.
|
||||
|
||||
These are the events U14 deliberately did NOT build in økt 55 — "an emitter with no call site
|
||||
is a shape guessed instead of measured". This is the call site. A Magentic manager decides
|
||||
which base to open and who speaks next; without these, the only trace of that reasoning is
|
||||
MAF's own ``invoke_agent`` spans, which say a call happened and nothing about what it decided.
|
||||
|
||||
The ledger event carries the decision fields rather than a rendered sentence, for the reason
|
||||
``SkippedLink`` is structured and ``BudgetExceeded`` carries three fields: "who was asked" and
|
||||
"was the request satisfied" are separate operative questions, and a reader who has to re-parse
|
||||
prose to tell them apart has a trace they cannot query.
|
||||
"""
|
||||
tracer, exporter = _recording_tracer()
|
||||
# Patched where the name is BOUND (the ``hosting.run_project`` precedent): ``explore``
|
||||
# imports it by name, so patching ``tracing`` would leave that binding untouched and this
|
||||
# test would quietly measure nothing.
|
||||
monkeypatch.setattr(explore, "exploration_tracer", lambda: tracer)
|
||||
|
||||
seen: list[explore.PlanReviewRequest] = []
|
||||
contract = ExplorationContract(
|
||||
**{**_FULL_CONTRACT, "enable_plan_review": True, "max_plan_revisions": 2}
|
||||
)
|
||||
await explore.explore(
|
||||
PROMPT,
|
||||
contract=contract,
|
||||
bundle_dirs=(),
|
||||
plan_reviewer=_reviewer([explore.PlanReviewDecision.revise("Test night setback.")], seen),
|
||||
client_factory=_factory(
|
||||
ledgers=[
|
||||
_ledger_json(satisfied=False, speaker=explore.HYPOTHESISER_ROLE),
|
||||
_ledger_json(satisfied=True, speaker=explore.HYPOTHESISER_ROLE),
|
||||
],
|
||||
hypothesiser=[_hypothesis_line("Night setback", "the expert asked")],
|
||||
),
|
||||
)
|
||||
|
||||
spans = exporter.get_finished_spans()
|
||||
assert [s.name for s in spans] == [explore.EXPLORATION_SPAN]
|
||||
events = [(e.name, dict(e.attributes or {})) for e in spans[0].events]
|
||||
names = [name for name, _ in events]
|
||||
assert names.count("plan_created") == 1
|
||||
assert names.count("replanned") == 1, "the human's revision must be visible in the trace"
|
||||
assert names.count("progress_ledger_updated") == 2
|
||||
|
||||
ledger_events = [attrs for name, attrs in events if name == "progress_ledger_updated"]
|
||||
assert [a["round_index"] for a in ledger_events] == [1, 2]
|
||||
assert [a["next_speaker"] for a in ledger_events] == [explore.HYPOTHESISER_ROLE] * 2
|
||||
assert [a["is_request_satisfied"] for a in ledger_events] == [False, True]
|
||||
assert all(a["speaker_known"] for a in ledger_events)
|
||||
|
||||
|
||||
#: A complete exploration in a CHILD interpreter. The stdout/stderr question cannot be answered
|
||||
#: in-process: ``ConsoleSpanExporter``'s ``out`` default is bound when
|
||||
#: ``opentelemetry.sdk.trace.export`` is first imported, so under pytest it is whatever stdout was
|
||||
#: at COLLECTION time — and ``capsys``, which replaces ``sys.stdout`` later, never sees it. That is
|
||||
#: not a testing quirk to work around; it is precisely the fact U14 exists for, and the reason
|
||||
#: ``configure_tracing`` passes ``out=`` explicitly instead of trusting the default. Measured: a
|
||||
#: mutation routing exploration spans to that default left an in-process ``capsys`` assertion
|
||||
#: GREEN while the spans really were on stdout.
|
||||
_CHILD_EXPLORATION = """
|
||||
import asyncio, json, sys
|
||||
from portfolio_optimiser import explore
|
||||
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||
from portfolio_optimiser.tracing import configure_tracing
|
||||
|
||||
configure_tracing()
|
||||
|
||||
LEDGER = json.dumps({
|
||||
"is_request_satisfied": {"reason": "r", "answer": True},
|
||||
"is_in_loop": {"reason": "r", "answer": False},
|
||||
"is_progress_being_made": {"reason": "r", "answer": True},
|
||||
"next_speaker": {"reason": "r", "answer": "navigator"},
|
||||
"instruction_or_question": {"reason": "r", "answer": "none"},
|
||||
})
|
||||
|
||||
def _select(blob, _role):
|
||||
if "provide the final answer" in blob:
|
||||
return "FINAL: done."
|
||||
if "pure JSON format" in blob:
|
||||
return LEDGER
|
||||
if "bullet-point plan" in blob:
|
||||
return "PLAN: - ask the navigator"
|
||||
if "pre-survey" in blob:
|
||||
return "FACTS: none."
|
||||
return "{}"
|
||||
|
||||
def factory(role):
|
||||
if role == explore.MANAGER_ROLE:
|
||||
return ScriptedChatClient(reply_selector=_select, role=role)
|
||||
return ScriptedChatClient("ok", role=role)
|
||||
|
||||
contract = explore.ExplorationContract(
|
||||
max_rounds=4, max_tokens=100000, max_stall_count=2,
|
||||
max_reset_count=1, max_plan_revisions=0, enable_plan_review=False,
|
||||
)
|
||||
result = asyncio.run(
|
||||
explore.explore("probe", contract=contract, bundle_dirs=(), client_factory=factory)
|
||||
)
|
||||
assert result.stop is None, result.stop
|
||||
print("EXPLORATION-OK", file=sys.stderr)
|
||||
"""
|
||||
|
||||
|
||||
def _run_child(**env: str) -> subprocess.CompletedProcess[str]:
|
||||
return subprocess.run(
|
||||
[sys.executable, "-c", _CHILD_EXPLORATION],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
cwd=str(Path(__file__).resolve().parent.parent),
|
||||
env={**os.environ, **env},
|
||||
)
|
||||
|
||||
|
||||
def test_an_untraced_exploration_writes_nothing_to_stdout_or_stderr() -> None:
|
||||
"""T23: in a real process, with tracing off, an exploration prints NOTHING.
|
||||
|
||||
A subprocess and not ``capsys``, for the reason recorded above ``_CHILD_EXPLORATION`` — and the
|
||||
stakes are the pinned artefacts: ``tests/golden/demo-transcript.stdout`` is byte-fixed and the
|
||||
demo's stderr is fixed at four lines, so one stray span dump would break both.
|
||||
|
||||
``EXPLORATION-OK`` on stderr is the control. Without it, "stdout was empty" would be equally
|
||||
true of a child that crashed on import, which is the fourth face of the verification law: an
|
||||
absence is only evidence once you have shown the measurement could have found something.
|
||||
"""
|
||||
proc = _run_child(PORTFOLIO_OTEL="")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
assert "EXPLORATION-OK" in proc.stderr, "the child must really have run an exploration"
|
||||
assert proc.stdout == ""
|
||||
# Not an exact-equality assertion on stderr: MAF emits two ``ExperimentalWarning`` lines while
|
||||
# importing, under every run form, and they are the same pair the demo's pinned stderr already
|
||||
# carries. What must be absent is TRACE data, so that is what is asserted.
|
||||
assert '"name": "exploration"' not in proc.stderr
|
||||
assert "progress_ledger_updated" not in proc.stderr
|
||||
|
||||
|
||||
def test_a_traced_exploration_puts_its_span_on_stderr_and_leaves_stdout_clean() -> None:
|
||||
"""T24: the positive arm — ``PORTFOLIO_OTEL=console`` and the exploration span is on STDERR.
|
||||
|
||||
This is what the whole U14 seam was landed for, now carrying the events U4 gave it a call site
|
||||
for. Both halves are asserted: the span and its ``progress_ledger_updated`` event ARE exported
|
||||
(so tracing is real), and stdout is STILL empty (so the byte-pinned transcript survives a
|
||||
traced run). Asserting only the first would pass on an exporter writing to stdout — which is
|
||||
OpenTelemetry's own default, and therefore the mistake actually available to make.
|
||||
"""
|
||||
proc = _run_child(PORTFOLIO_OTEL="console")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
assert "EXPLORATION-OK" in proc.stderr
|
||||
assert proc.stdout == "", "a traced run must not put one byte on stdout"
|
||||
assert '"name": "exploration"' in proc.stderr
|
||||
assert "progress_ledger_updated" in proc.stderr
|
||||
|
||||
|
||||
def test_a_marked_line_that_will_not_parse_is_a_hard_error() -> None:
|
||||
"""T24: the marker is what makes fail-closed affordable here.
|
||||
|
||||
Most hypothesiser turns legitimately are not hypotheses — the agent reasons out loud — so
|
||||
"parse every turn or fail" would refuse a normal exploration. The marker separates a turn that
|
||||
is not a claim from a claim that cannot be read. The second is the run's own product coming
|
||||
back unreadable, so it raises (``write_concept_file``'s rule: validation, never repair) rather
|
||||
than following the tolerant RAW-inbox rule, which belongs to folders anyone may drop files in.
|
||||
"""
|
||||
with pytest.raises(explore.HypothesisParseError):
|
||||
explore._parse_hypotheses([f"{explore.HYPOTHESIS_MARKER} not json at all"])
|
||||
with pytest.raises(explore.HypothesisParseError):
|
||||
explore._parse_hypotheses([f'{explore.HYPOTHESIS_MARKER} {{"label": "no rationale"}}'])
|
||||
|
||||
|
||||
def test_unmarked_prose_is_not_a_failure() -> None:
|
||||
"""T25: the control for T24 — ordinary reasoning yields no hypothesis and no error.
|
||||
|
||||
Without it, T24 would pass on an implementation that refused every hypothesiser turn that was
|
||||
not a hypothesis, which would make the loop unusable and the strictness meaningless.
|
||||
"""
|
||||
assert explore._parse_hypotheses(["I looked at the index and nothing stands out yet."]) == []
|
||||
|
||||
|
||||
def test_a_review_nobody_can_answer_is_refused_before_the_first_model_call() -> None:
|
||||
"""T26: plan review without a reviewer refuses; a reviewer without plan review refuses too.
|
||||
|
||||
The first would hang: the workflow stops at a ``request_info`` and nothing ever answers it, and
|
||||
a hang is the one failure mode that reports nothing at all. The second is the silent-ignore the
|
||||
repo's flag partition forbids — a caller who supplied a reviewer believes a human is in the
|
||||
loop. Both are refused BEFORE anything is built, so neither costs a model call.
|
||||
"""
|
||||
with pytest.raises(explore.ExplorationError):
|
||||
asyncio.run(
|
||||
explore.explore(
|
||||
PROMPT,
|
||||
contract=ExplorationContract(**_FULL_CONTRACT),
|
||||
bundle_dirs=(),
|
||||
client_factory=_factory(ledgers=[], hypothesiser=[]),
|
||||
)
|
||||
)
|
||||
with pytest.raises(explore.ExplorationError):
|
||||
asyncio.run(
|
||||
explore.explore(
|
||||
PROMPT,
|
||||
contract=ExplorationContract(**_NO_REVIEW),
|
||||
bundle_dirs=(),
|
||||
plan_reviewer=lambda _r: explore.PlanReviewDecision.approve(),
|
||||
client_factory=_factory(ledgers=[], hypothesiser=[]),
|
||||
)
|
||||
)
|
||||
Loading…
Add table
Add a link
Reference in a new issue