feat(p17b): ONE commission, SEVERAL bases -- reachable from the command line

``run_mandate_across_bundles`` has existed since session 58, reachable from FIVE
test files and from NO command line (measured: ``grep -n across-bundle run.py``
= 0 hits). ``--across-bundle <dir>``, repeated once per base, is that door.

The engine takes a CALLBACK rather than an outbox directory. Its own docstring
has always said N runs need N ``run_id``s and that minting them there would
default a key this repo requires a caller to supply -- so ``outbox_for`` is that
contract KEPT, not relaxed, and the operator-chosen ``<run-id>-<bundle_id>``
rule lives in ``main()`` where the decision was made. The order's alternative (a
caller running ``run_project`` itself over ``route_by_bundle``'s sub-mandates)
would be a second copy of the loop's id reconciliation, shared store, per-base
project resolution, collision accounting and both budget teeth.

``resolve_bundle_routing`` is ONE resolution shared by the engine and the
dry-run arm: a free trip answering with a different project id, or tolerating a
duplicate id the paid dispatch refuses, would rehearse a different run.

``{run-id}-multibase.json`` is written from a ``finally`` and every row is built
from the resolution plus disk, so the pass a cap cut short still leaves the
record. ``completed`` is a required field for ``ExplorationTrace.completed``'s
reason. ``stop_reason`` is read BACK from each base's own coverage artefact.

Load-bearing MEASURED (17 arms), four mutations all red against the WHOLE suite,
green control 1761/5 (from 1744/5, superset, 0 removed), golden byte-unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-15 04:24:48 +02:00
commit 5e4c497a84
5 changed files with 994 additions and 37 deletions

View file

@ -2728,6 +2728,47 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
runde betaler første approach, den andre er den taket kutter). **Ærlighets-grense, uttalt:**
`token_usage` er kjøringens ENE teller — den skiller ikke debatt fra generering, og en
per-fase-fordeling ville krevd en andre måler.
- **ÉN kommisjon, FLERE baser er nåbar fra CLI-en — og utboksens nøkkel er KALLERENS, aldri
motorens (P17b DEL 1, 15.09):** `run_mandate_across_bundles` har eksistert siden økt 58, nåbar
fra FEM testfiler og fra INGEN kommandolinje (MÅLT: `grep -n across-bundle run.py` = 0 treff).
`--across-bundle <dir>` (repeterbart; `--bundle-dir` forblir ÉN katalog og er NEKTET her) krever
`--mandate`, `--run-id` og `--outbox-dir`. **Motoren fikk en CALLBACK, ikke en `outbox_dir`:**
dens egen docstring har alltid sagt at N kjøringer trenger N `run_id`-er og at å mynte dem der
ville defaultet en nøkkel repoet krever at en kaller oppgir — så `outbox_for(bundle_id) ->
(dir, run_id)` er dét kravet OPPFYLT, ikke slakket, og myntingsregelen `<run-id>-<bundle_id>`
(OPERATØRVALGT 14.09) bor i `main()` der beslutningen ble tatt. Ordrens andre alternativ — en
kaller som kjører `run_project` selv over `route_by_bundle`s sub-mandater — ville vært en ANDRE
kopi av løkkas id-avstemming, delte store, per-base-prosjektoppslag, kollisjonsregnskap og
BEGGE budsjett-tenner: kø-(p) over fem regler som hver har ett hjem.
**`resolve_bundle_routing` er ÉN oppløsning, delt av motoren og dry-run-armen:** en gratis tur
som svarte med en annen `project_id`, eller tolererte en duplisert id den betalte kjøringen
nekter, ville vært en generalprøve på en annen kjøring. **Samlefila skrives fra en `finally`**
(`write_parse_failures`-presedensen) og hver rad bygges av RESOLUSJONEN + DISK — de konfigurerte
basene, kallerens egen myntingsregel, og hver bases egen `{run_id}-coverage.json` — så den
kjøringen som mest trenger regnskapet, den et tak kappet, etterlater det. `completed` er et
EGET påkrevd felt (`ExplorationTrace.completed`s grunn ordrett): «ingenting ble uoppnådd» og
«vi fikk aldri vite» må ikke være samme verdi. `stop_reason` LESES TILBAKE fra coverage-fila,
aldri utledet på nytt — P19 D2 la faktumet der, og en andre utledning her ville stått fritt til
å være uenig med den dommeren leser. **`BudgetExceeded` er i nekt-tuppelen** av
enkeltprosjekt-stiens MÅLTE grunn: den er en `RuntimeError`, og den FØRSTE tilnærmingen som
treffer taket re-raiser ved design — over flere baser er dét ikke et kanttilfelle (runde 3 målte
`stop_reason: rounds` i 5 av 5), så uten armen er en multi-base-kjørings vanligste utfall en
traceback. Tre partisjons-rader: `report_forbidden` og `--portfolio` NEKTER ved navn, mens
live-dry-run-raden er en WIRING — den driller HVER konfigurert base og printer hver bases egne
varsler, fordi en kjøring som sa dem én gang bare kunne snakket om én av N. Load-bearing MÅLT
(`tests/test_across_bundles_cli_loadbearing.py`, 17 armer), **fire mutasjoner alle røde mot HELE
suiten** + grønn kontroll **1761/5** (fra 1744/5, supersett, 0 fjernet) og golden
`demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): (i) samlefila droppes (2 røde) · (ii) myntingen
kollapser til bart `<run-id>` (3) · (iii) `report_forbidden` slipper flagget stille (1) ·
(iv) `opened`/`requirements`-sinkene deles mellom basene (5, hvorav FIRE i tester eldre enn
dette arbeidet — uavhengige vitner på at sinkene er per kjøring). **Ærlighets-grenser, uttalt:**
en base som RAISER propagerer fortsatt (motorens egen dokumenterte grense — `collect-and-continue`
tilhører `run_portfolio`), så de etterfølgende basene kjøres ikke og samlefila sier `completed:
false`; `announce` sier fortsatt «the portfolio» når ingen `project_id` er gitt; den hostede
flaten er BEVISST urørt (feltet er i ingen av hostings tre sett); og `--proposal-review` trås
gjennom til ÉN terminal delt av alle basene (motorens egen begrunnelse — dispatchen er
sekvensiell).
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -412,7 +412,7 @@ when the seam is detached, so the loop cannot silently degrade into theater.
are still made and discarded, and nothing can leave the process. The OTLP exporter packages are
not declared dependencies (they are egress, and heavy in a published wheel); install one yourself
if you use that mode.
- **Run:** the `run.py` CLI has **three modes** — a documented partition, since one invocation
- **Run:** the `run.py` CLI has **four modes** — a documented partition, since one invocation
cannot exercise every flag:
- **Single-project**`PROJECT_ID --docs-dir <dir>`, plus optional `--bundle-dir`,
`--verdict-dir`, `--outbox-dir` (which requires `--run-id`), `--dimension-config`,
@ -425,6 +425,19 @@ when the seam is detached, so the loop cannot silently degrade into theater.
shape the mandate this run evaluates — see below), and `--derive-cost-baseline` (opt-in:
anchor the validator on a priced schedule already inside `--bundle-dir` instead of a
hand-written `cost-baseline.json` — see below).
- **Multi-base (P17b)**`--across-bundle <dir>` repeated once per knowledge base, plus
`--mandate <file>`, `--run-id <id>` and `--outbox-dir <dir>` (all three required). One
commission, several bases: the mandate is partitioned by each approach's `bundle_id` and the
existing pipeline runs once per base, sequentially, threading ONE verdict store so a verdict
minted against base *k* reaches base *k+1*'s hypothesis. Each base writes its own full
artefact set under `<run-id>-<bundle_id>`, and one `<run-id>-multibase.json` beside them
records the spend order, the per-base `run_id`, `unreached`, `collisions`, `budget_stop` and
each base's own `stop_reason` — written even when a base is cut short, with a `completed`
field so "nothing was left unreached" cannot be read as "we never found out". `--live-dry-run`
drills every configured base and stops before the first model call. `--bundle-dir` stays ONE
directory and is refused here, as are `--portfolio`, `--explore`, `--prepass-payload` and
`--proposals-from-mandate` — each of those resolves one base, and picking which of N was meant
is not this layer's to decide.
- **Portfolio**`--portfolio`, plus optional `--goals`, `--ledger`, `--dimension-config`,
`--semantic-retrieval`; it stops early and prints a `goal reached: …` line when the
accumulated ledger meets a goal.

View file

@ -230,6 +230,55 @@ def write_debate_tools(
return path
def write_multibase(
outbox_dir: str,
run_id: str,
*,
runs: Sequence[Mapping[str, Any]],
completed: bool,
unreached: Sequence[Mapping[str, Any]],
collisions: Sequence[Mapping[str, Any]],
stopped_early: bool,
budget_stop: Mapping[str, Any] | None,
) -> Path:
"""Write ``{run_id}-multibase.json`` — what ONE commission did across SEVERAL bases (P17b).
The question no per-base artefact can answer. Each base writes its own full set under its own
minted ``run_id``, but nothing in that set says in which ORDER the bases were spent, which id
each one was given, which approaches were never reached, or which candidates two bases both
described and a reader who has to reconstruct the ``<run-id>-<bundle_id>`` convention to
pair the files back to the pass has been handed a naming rule instead of a record.
Written IFF the pass was given an outbox, exactly like its neighbours, and the per-base
``stop_reason`` rows are READ BACK from each base's own ``{run_id}-coverage.json`` by the
caller rather than recomputed here: P19 D2 put that fact in that file, and a second derivation
of it would be free to disagree with the one the judge reads.
Plain data only, so the RAW output layer stays MAF-free (``write_debate_tools``' own rule)."""
directory = Path(outbox_dir)
directory.mkdir(parents=True, exist_ok=True)
path = directory / f"{run_id}-multibase.json"
path.write_text(
_dump(
{
"run_id": run_id,
# REQUIRED, never inferred from an empty ``unreached``: a pass a cap or a provider
# cut short never got to say what it did not reach, and "nothing was left
# unreached" must not be the value that means "we never found out"
# (``ExplorationTrace.completed``'s own reason).
"completed": completed,
"runs": [dict(row) for row in runs],
"unreached": [dict(row) for row in unreached],
"collisions": [dict(row) for row in collisions],
"stopped_early": stopped_early,
"budget_stop": dict(budget_stop) if budget_stop is not None else None,
}
),
encoding="utf-8",
)
return path
def write_prepass(
outbox_dir: str,
run_id: str,

View file

@ -29,7 +29,7 @@ import asyncio
import json
from collections.abc import Awaitable, Callable, Iterable, Sequence
from contextlib import AsyncExitStack
from dataclasses import dataclass, replace
from dataclasses import asdict, dataclass, replace
from pathlib import Path
from typing import Any, Literal, cast
@ -2181,6 +2181,12 @@ class BundleRun:
bundle_dir: str
project_id: str
result: RunResult
#: The ``run_id`` this base's artefacts were written under, or ``""`` when the pass wrote no
#: outbox at all. Carried here rather than re-derived by the caller for the reason the two
#: fields above are: the CALLER minted it (P17b — N runs need N ids, and this engine refuses
#: to default a key the repo requires a caller to supply), so a reader pairing an artefact
#: back to a base must not have to reconstruct the naming convention to do it.
run_id: str = ""
@dataclass(frozen=True)
@ -2213,6 +2219,116 @@ class MultiBaseResult:
collisions: tuple[VerdictCollision, ...] = ()
def _coverage_stop_reason(outbox_dir: str, run_id: str) -> str:
"""``BudgetExceeded.kind`` this base recorded, read back off its OWN coverage artefact.
Never recomputed from the ``RunResult``: P19 D2 put "why did this run stop" in
``{run_id}-coverage.json`` precisely because ``_evaluate_mandate`` swallows the exception once
something has been produced, so the dispatcher never sees it. A second derivation here would
be free to disagree with the file the judge reads.
``"absent"`` is a THIRD value, and not the same as ``""``: a base that wrote no coverage file
at all is a different finding from one that finished with nothing stopping it the
``stress`` judge's own vocabulary, reused rather than re-invented."""
path = Path(outbox_dir) / f"{run_id}-coverage.json"
if not path.is_file():
return "absent"
try:
return str(json.loads(path.read_text(encoding="utf-8")).get("stop_reason", ""))
except (OSError, json.JSONDecodeError):
return "absent"
def _write_multibase_summary(
outbox_dir: str,
run_id: str,
*,
resolved: Sequence[tuple[str, str, str]],
mint: Callable[[str], tuple[str, str]],
multi: MultiBaseResult | None,
) -> None:
"""Write ``{run_id}-multibase.json`` for a multi-base pass, COMPLETED or not (P17b).
Called from a ``finally``, which is ``write_parse_failures``' rule applied one layer up: the
pass that most needs a record of what it spent is the one a cap or a provider cut short, and
the engine's documented limit is that a base which RAISES propagates. Every per-base row is
therefore built from the RESOLUTION and from DISK the configured bases, the caller's own
minting rule, and each base's own ``{run_id}-coverage.json`` — none of which needs the
dispatch to have returned.
``completed`` is a REQUIRED field of the artefact and not an inference from an empty
``unreached``: ``ExplorationTrace.completed``'s reason verbatim, because "nothing was left
unreached" and "we never found out" must not be the same value. When the pass did not
complete, ``unreached``/``collisions``/``budget_stop`` are what the dispatch never got to say,
and they are written as empty/``None`` UNDER that flag rather than as findings.
"""
rows = []
for bundle_id, bundle_dir, project_id in resolved:
_, base_run_id = mint(bundle_id)
rows.append(
{
"bundle_id": bundle_id,
"bundle_dir": bundle_dir,
"project_id": project_id,
"run_id": base_run_id,
"stop_reason": _coverage_stop_reason(outbox_dir, base_run_id),
}
)
outbox.write_multibase(
outbox_dir,
run_id,
runs=rows,
completed=multi is not None,
unreached=[asdict(row) for row in multi.unreached] if multi is not None else [],
collisions=[asdict(row) for row in multi.collisions] if multi is not None else [],
stopped_early=multi.stopped_early if multi is not None else False,
budget_stop=(
asdict(multi.budget_stop)
if multi is not None and multi.budget_stop is not None
else None
),
)
def resolve_bundle_routing(
bundle_dirs: Sequence[str],
) -> tuple[tuple[str, str, str], ...]:
"""``(bundle_id, bundle_dir, project_id)`` per configured base, in CONFIGURED order.
The ONE resolution shared by ``run_mandate_across_bundles`` and the CLI's multi-base dry run
(P17b). Extracted rather than copied for the reason the id derivation itself was unified in
Step 10: a drill that answered with a different project id, or tolerated a duplicate id the
paid dispatch refuses, would be a free trip that fails to measure the very run it precedes.
``bundle_id`` is ``okf.reconcile_bundle_id``'s — the declared id wins over the mount (S7a-3).
Two bases answering to ONE id refuse here, the same refusal ``explore._bundle_index`` makes
and for the same reason: the id is how the mandate NAMES a base, so a collision would let an
approach be evaluated against A while the report says B (the S3.2 key-collision class).
``project_id`` is S7b søm 1's precedence, and it is load-bearing in BOTH directions: the
hand-written IR projection FIRST (that file is what every existing base has always been routed
by), the base's own DECLARED id as the fallback (an ingested corpus carries no projection, so
file-only could not route it at all). No caller-supplied constant is admitted: it could only
ever be right for one base out of N.
:raises MandateRoutingError: two configured bases share one id.
"""
out: list[tuple[str, str, str]] = []
seen: dict[str, str] = {}
for raw in bundle_dirs:
bundle_id = okf.reconcile_bundle_id(raw).id
if bundle_id in seen:
raise MandateRoutingError(
f"two knowledge bases share the id {bundle_id!r} ({seen[bundle_id]!r} and "
f"{raw!r}); an approach names a base by that id, so it must be unique"
)
seen[bundle_id] = raw
declared_ir = okf.load_optional_ir_projection(raw)
project_id = str(declared_ir["project_id"]) if declared_ir is not None else bundle_id
out.append((bundle_id, raw, project_id))
return tuple(out)
async def run_mandate_across_bundles(
mandate: Mandate,
bundle_dirs: Sequence[str],
@ -2233,6 +2349,17 @@ async def run_mandate_across_bundles(
#: AND ``project_id`` — each base's own, read by ``_project_from_bundle`` — so bases can be
#: told apart even when a multi-base commission reuses approach ids.
proposal_reviewer: ProposalReviewer | None = None,
#: P17b. Where THIS base's artefacts go, and under which ``run_id`` — supplied by the caller
#: per base, never minted here. That is the engine's own long-standing contract kept rather
#: than relaxed: N runs need N ``run_id``s, and minting one here would default a key this repo
#: requires a caller to supply, for byte-determinism. A CALLBACK rather than an
#: ``outbox_dir``/``run_id`` pair because the naming rule is an OPERATOR decision
#: (``<run-id>-<bundle_id>``, chosen 14.09) and belongs at the call site that made it; the
#: alternative the order offered — a caller running ``run_project`` itself over
#: ``route_by_bundle``'s sub-mandates — would be a SECOND copy of this loop's id
#: reconciliation, shared store, per-base project resolution, collision accounting and both
#: budget teeth (kø-(p), over five rules that each have exactly one home).
outbox_for: Callable[[str], tuple[str, str]] | None = None,
) -> MultiBaseResult:
"""Evaluate ONE commission across SEVERAL knowledge bases — the multi-base dispatch (§ C.7).
@ -2267,30 +2394,18 @@ async def run_mandate_across_bundles(
independent projects, whereas here the caller asked for ONE commission to be evaluated.
(2) Without a ``portfolio_meter`` the pass's ceiling is the number of routed bases times
``max_tokens``, each run bounded on its own the global ledger is opt-in, and this does not
re-implement it (``_run_meter`` is the one copy of the binding rule). (3) The outbox is NOT
wired: N runs need N ``run_id``s, and minting them here would default a key this repo requires
a caller to supply, for byte-determinism. A caller who needs artefacts per base calls
``run_project`` itself with the sub-mandates ``route_by_bundle`` hands back.
re-implement it (``_run_meter`` is the one copy of the binding rule). (3) The outbox is
CALLER-KEYED: N runs need N ``run_id``s, and minting them here would default a key this repo
requires a caller to supply, for byte-determinism so ``outbox_for`` hands each base its
directory and its id, and a caller that offers no callback still writes nothing (every
pre-P17b call site, unchanged).
:raises MandateRoutingError: the commission cannot be routed against ``bundle_dirs``.
:raises BudgetRefused: a global remainder that cannot fund a single run.
"""
by_id: dict[str, str] = {}
for raw in bundle_dirs:
# The ONE derivation rule (Step 10) — this used to be a second private copy of
# ``Path(raw).name``, free to drift from ``explore``'s. The REFUSAL below stays local:
# ``MandateRoutingError`` is this door's class, ``ExplorationError`` is explore's, and
# unifying the derivation is not the same as unifying the two doors' error vocabularies.
bundle_id = okf.reconcile_bundle_id(raw).id
if bundle_id in by_id:
# The same refusal ``explore._bundle_index`` makes, for the same reason: the id is how
# the mandate names a base, so two bases answering to one name would let an approach be
# evaluated against A while the report says B (the S3.2 key-collision class).
raise MandateRoutingError(
f"two knowledge bases share the id {bundle_id!r} ({by_id[bundle_id]!r} and "
f"{raw!r}); an approach names a base by that id, so it must be unique"
)
by_id[bundle_id] = raw
resolved = resolve_bundle_routing(bundle_dirs)
by_id = {bundle_id: bundle_dir for bundle_id, bundle_dir, _ in resolved}
project_by_id = {bundle_id: project_id for bundle_id, _, project_id in resolved}
routed = route_by_bundle(mandate, tuple(by_id))
@ -2327,19 +2442,7 @@ async def run_mandate_across_bundles(
break
bundle_dir = by_id[bundle_id]
# ONE reading of the base's own project id, used both to ADDRESS the run and to LABEL it.
# A second lookup for the label would be the kø-(p) duplicate free to drift from the value
# the run was actually dispatched with.
#
# S7b søm 1: the hand-written projection FIRST, the base's DECLARED id as the fallback. The
# precedence is load-bearing in both directions. Declaration-first would re-address every
# existing base whose ``project_id`` differs from its ``bundle_id`` — the file is what those
# bases have always been routed by. File-only was the refusal this seam removes: an ingested
# corpus carries no projection, so it could not be routed at all. ``bundle_id`` is the
# identity every other door already resolves through (S7a-3), so the fallback introduces no
# third notion of what a base is called; ``by_id`` above is that same resolution, reused.
declared_ir = okf.load_optional_ir_projection(bundle_dir)
project_id = str(declared_ir["project_id"]) if declared_ir is not None else bundle_id
project_id = project_by_id[bundle_id]
# D2: the id of the verdict THIS base minted, taken from ``run_project``'s existing
# ``notify`` seam rather than off the returned ``RunResult``. ``notify`` fires inside the
# capture block, so it is called exactly when a verdict exists (F2: never when nobody
@ -2347,6 +2450,7 @@ async def run_mandate_across_bundles(
# each iteration — a shared accumulator would let a later base read the previous base's
# verdict and manufacture a collision that never happened.
minted_here: list[str] = []
base_outbox, base_run_id = outbox_for(bundle_id) if outbox_for is not None else (None, "")
result = cast(
RunResult,
await run_project(
@ -2354,6 +2458,8 @@ async def run_mandate_across_bundles(
profile,
docs_dir=bundle_dir,
bundle_dir=bundle_dir,
outbox_dir=base_outbox,
run_id=base_run_id or None,
notify=lambda verdict: minted_here.append(verdict.id),
verdict_input=verdict_input,
verdict_dir=verdict_dir,
@ -2397,6 +2503,7 @@ async def run_mandate_across_bundles(
bundle_dir=bundle_dir,
project_id=project_id,
result=result,
run_id=base_run_id,
)
)
@ -2535,6 +2642,18 @@ def main(argv: list[str] | None = None) -> int:
parser.add_argument(
"--bundle-dir", default=None, help="OKF bundle dir (enables the Step-1 fold)"
)
parser.add_argument(
"--across-bundle",
action="append",
default=None,
metavar="DIR",
help="P17b: run ONE commission across SEVERAL knowledge bases — repeat the flag once per "
"base. The mandate is partitioned by each approach's bundle_id and the existing pipeline "
"runs once per base, sequentially, threading ONE verdict store so a verdict minted against "
"base k reaches base k+1. Each base writes its own artefact set under <run-id>-<bundle_id>, "
"plus one <run-id>-multibase.json summary. Requires --mandate, --run-id and --outbox-dir; "
"--bundle-dir stays ONE directory and is refused here",
)
parser.add_argument(
"--verdict-dir",
default=None,
@ -2888,6 +3007,10 @@ def main(argv: list[str] | None = None) -> int:
"--goals": args.goals is not None,
"--docs-dir": args.docs_dir is not None,
"--bundle-dir": args.bundle_dir is not None,
# P17b, and for its neighbours' reason: report mode returns ABOVE every dispatch,
# including the multi-base one, so an omission here is a SILENT DROP of a whole pass
# rather than a refusal (the F4 class).
"--across-bundle": bool(args.across_bundle),
"--verdict-dir": args.verdict_dir is not None,
"--outbox-dir": args.outbox_dir is not None,
"--run-id": args.run_id is not None,
@ -2969,6 +3092,13 @@ def main(argv: list[str] | None = None) -> int:
single_only = {
"--docs-dir": args.docs_dir,
"--bundle-dir": args.bundle_dir,
# P17b. A portfolio pass keys on PROJECTS and reads each project's base off its own
# row, so a run-level list of bases has nowhere to go; the two are different axes
# (``MultiBaseResult`` is a distinct type from ``PortfolioResult`` for exactly that
# reason). BY NAME, like its neighbours: falling through to "--across-bundle requires
# --mandate" would tell an operator who wrote --portfolio --across-bundle to add a
# flag that is legal in both modes, which answers the wrong question.
"--across-bundle": bool(args.across_bundle),
"--verdict-dir": args.verdict_dir,
"--outbox-dir": args.outbox_dir,
"--run-id": args.run_id,
@ -3037,13 +3167,63 @@ def main(argv: list[str] | None = None) -> int:
)
return 1
# P17b — the multi-base door's own refusals, at FUNCTION level and never nested under another
# flag's branch, for the F4 reason its neighbours are: under one, a bare combination falls
# straight through to a dispatch that drops the flag in silence. Placed ABOVE the required-args
# guard because this mode takes NO ``PROJECT_ID`` at all — each base's project is read from
# THAT base's own IR projection, which is the whole reason the dispatch has no such parameter.
if args.across_bundle:
# The three things a multi-base pass cannot invent. ``--mandate`` because the commission IS
# the partition key (without ``Approach.bundle_id`` there is nothing to route on), and the
# outbox pair because N runs need N ``run_id``s: the engine refuses to default that key,
# so the caller must supply the stem it mints them from.
required = {
"--mandate": args.mandate,
"--run-id": args.run_id,
"--outbox-dir": args.outbox_dir,
}
missing = [name for name, value in required.items() if not value]
if missing:
print(
f"run refused: --across-bundle requires {', '.join(missing)} (the commission is "
"what partitions the pass by base, and each base writes its own artefact set "
"under <run-id>-<bundle_id> — a key this repo requires a caller to supply)",
file=sys.stderr,
)
return 1
# Four single-base modes, each refused BY NAME rather than by falling through. Every one
# of them resolves something from THE base — one directory, one cut, one exploration, one
# derived schedule — and silently picking which of N that means is the guessed-shape class
# this repo refuses outright.
conflicting = {
"--bundle-dir": args.bundle_dir,
"--explore": args.explore,
"--prepass-payload": args.prepass_payload,
"--proposals-from-mandate": args.proposals_from_mandate,
}
clash = [name for name, value in conflicting.items() if value]
if clash:
print(
f"run refused: --across-bundle cannot be combined with {', '.join(clash)} (each "
"of those resolves ONE knowledge base — a directory, a declared cut, an "
"exploration or a derived schedule — and this mode configures several; which one "
"was meant is not something this layer may decide)",
file=sys.stderr,
)
return 1
# Single-project mode requires PROJECT_ID + --docs-dir (compensating for the relaxed argparse
# required/positional so the legacy contract keeps failing loudly via the refusal surface).
# HOISTED above the scripted door (below) so an incomplete argv is refused BEFORE the honesty
# banner could claim a scripted run happened; the refusal ORDER within single-project mode
# (required args -> semantic-retrieval -> scripted) is unchanged.
if not args.portfolio and (
args.project_id is None or (args.docs_dir is None and args.bundle_dir is None)
#
# ``not args.across_bundle`` is a MODE test, not a relaxation: the multi-base pass takes no
# PROJECT_ID and no single ``--bundle-dir``, and both of those are refused above by name.
if (
not args.portfolio
and not args.across_bundle
and (args.project_id is None or (args.docs_dir is None and args.bundle_dir is None))
):
print(
"run refused: single-project mode requires PROJECT_ID and either --docs-dir or "
@ -3892,6 +4072,158 @@ def main(argv: list[str] | None = None) -> int:
print(settle(coverage))
return 0
# P17b — the multi-base dispatch. Placed BELOW the announcement (the commission is declared
# before the work it commissions) and ABOVE the portfolio dispatch, because it is a MODE
# rather than a modifier: it runs the pass and returns. The dry-run arm lives INSIDE this
# block rather than in the generic ``--live-dry-run`` branch further down, which addresses
# ``args.project_id``/``args.bundle_dir`` — neither of which this argv has — and would
# therefore drop the whole pass in silence (the F4 class).
if args.across_bundle:
assert mandate is not None # narrowed by the required-flags refusal above
assert args.run_id is not None and args.outbox_dir is not None # same refusal
bases = list(args.across_bundle)
try:
# Resolved and ROUTED on the free trip too: a commission that cannot be executed as
# written must be refused while it is still free (the økt-57 hoist), and a drill that
# tolerated a routing error the paid pass refuses would be a rehearsal of a different
# run.
resolved = resolve_bundle_routing(bases)
route_by_bundle(mandate, tuple(bundle_id for bundle_id, _, _ in resolved))
except (MandateRoutingError, okf.BundleIdMismatch, FileNotFoundError, ValueError) as exc:
print(f"run refused: {exc}", file=sys.stderr)
return 1
if args.live_dry_run:
for bundle_id, bundle_dir, project_id in resolved:
try:
report = asyncio.run(
run_project(
project_id,
args.profile,
docs_dir=bundle_dir,
bundle_dir=bundle_dir,
verdict_dir=args.verdict_dir,
dimension=(
load_dimension(args.dimension_config)
if args.dimension_config
else None
),
max_rounds=args.max_rounds,
max_tokens=args.max_tokens,
require_cost_baseline=args.require_cost_baseline,
mcp_servers=mcp_servers,
live_dry_run=True,
)
)
except (ValueError, FileNotFoundError, ValidationError) as exc:
print(f"live-dry-run refused: {bundle_id}: {exc}", file=sys.stderr)
return 1
assert isinstance(report, DryRunReport)
print(
f"{bundle_id} ({project_id}): LIVE-DRY-RUN OK (profile={report.profile}, "
f"models={report.resolved_models}, max_rounds={report.max_rounds}, "
f"max_tokens={report.max_tokens}, top_k={report.top_k}) — "
"ingen modellkall gjort (stoppet før første debate.run)"
)
# Every notice the single-base drill prints, per base: a pass whose SECOND base
# cannot be anchored is exactly as unanchored as one whose first cannot, and a
# drill that said it once would leave the operator guessing which base it meant.
for notice in (
cost_baseline_notice(report.cost_baseline_anchored),
grounding_offer_notice(report.grounding_offer),
skipped_links_notice(report.skipped_links),
bundle_id_notice(report.bundle_id_source),
):
if notice is not None:
print(notice)
return 0
outbox_dir = args.outbox_dir
stem = args.run_id
def outbox_for(bundle_id: str) -> tuple[str, str]:
"""The OPERATOR-CHOSEN minting rule (14.09): ``<run-id>-<bundle_id>``, one stem per
base. Written here, at the call site that made the decision, rather than inside the
dispatch which is precisely why the engine takes a callback and not a directory."""
return outbox_dir, f"{stem}-{bundle_id}"
#: Bound by the dispatch when it RETURNS; still ``None`` when a base raised. The summary
#: is written from a ``finally`` either way (``write_parse_failures``' precedent), because
#: the pass that most needs a record of what it spent is the one a cap or a provider cut
#: short — and the engine's documented limit is that a base which RAISES propagates.
multi: MultiBaseResult | None = None
try:
multi = asyncio.run(
run_mandate_across_bundles(
mandate,
tuple(bases),
args.profile,
verdict_dir=args.verdict_dir,
dimension=(
load_dimension(args.dimension_config) if args.dimension_config else None
),
client_factory=scripted_client_factory,
max_rounds=args.max_rounds,
max_tokens=args.max_tokens,
verdict_input=_verdict_input_from_args(args),
proposal_reviewer=(
terminal_proposal_reviewer() if args.proposal_review else None
),
outbox_for=outbox_for,
)
)
except ChatClientException as exc:
print(f"run stopped: {exc}", file=sys.stderr)
return 1
except ProposalReviewInputError as exc:
print(f"run stopped: {exc}", file=sys.stderr)
return 1
except (
MandateRoutingError,
okf.BundleIdMismatch,
FileNotFoundError,
ValidationError,
ValueError,
BudgetExceeded,
) as exc:
# ``BudgetExceeded`` is in the tuple for the single-project path's measured reason: it
# is a ``RuntimeError``, and the FIRST approach to hit the cap re-raises by design
# (``_evaluate_mandate`` only swallows it mid-list). Over several bases that is not an
# edge case — round 3 measured ``stop_reason: rounds`` in 5 of 5 runs — so without this
# arm the ordinary outcome of a multi-base pass is a traceback.
print(f"run refused: {exc}", file=sys.stderr)
return 1
finally:
_write_multibase_summary(
outbox_dir,
args.run_id,
resolved=resolved,
mint=outbox_for,
multi=multi,
)
for bundle_run in multi.runs:
print(f"--- {bundle_run.bundle_id} ({bundle_run.project_id}) ---")
print(
f"{bundle_run.project_id}: {type(bundle_run.result.outcome).__name__} "
f"({verdict_notice(bundle_run.result)})"
)
# Every per-run notice the single-base path prints, read off THIS base's own stamp —
# a pass that said them once could only be talking about one of N bases, and the
# reader could not tell which.
for notice in (
cost_baseline_notice(bundle_run.result.provenance.cost_baseline_anchored),
bundle_id_notice(bundle_run.result.provenance.bundle_id_source),
skipped_links_notice(bundle_run.result.skipped_links),
unkeyed_verdicts_notice(bundle_run.result.unkeyed_verdicts),
):
if notice is not None:
print(notice)
notice = collision_notice(multi.collisions)
if notice is not None:
print(notice)
return 0
if args.portfolio:
# Portfolio mode (Step 3): dispatch to the EXISTING run_portfolio via the fail-fast loaders
# (run_portfolio itself is unchanged). Loader/ValueError failures surface through the same

View file

@ -0,0 +1,522 @@
"""P17b DEL 1 — ONE run, SEVERAL knowledge bases, from the CLI (order ``20260915T014020Z``).
``run_mandate_across_bundles`` has existed since økt 58 and was reachable from FIVE test files and
from NO command line (measured 15.09: ``grep -n across-bundle run.py`` = 0 hits). The operator
directive of 14.09 is that ``po`` must demonstrably work against THE BASES plural a given run
says it will use, and a library function no operator can invoke is not that demonstration.
**Three things this file pins, and each is a different claim.**
*The minting rule.* N runs need N ``run_id``s the engine's own docstring says so, and refuses to
default a key this repo requires a caller to supply. The rule is ``<run-id>-<bundle_id>`` and it is
OPERATOR-CHOSEN (14.09), not a preference argued here. The engine therefore takes an
``outbox_for(bundle_id) -> (outbox_dir, run_id)`` CALLBACK rather than an ``outbox_dir``: the
caller supplies the ids, which is exactly the contract the engine documents, and the naming
convention stays in ``main()`` where the operator's decision lives.
*Why a callback rather than a caller-side loop.* The order offered both. A caller that ran
``run_project`` itself over ``route_by_bundle``'s sub-mandates would have to re-implement the id
reconciliation, the shared ``VerdictStore``, the per-base ``project_id`` resolution, the collision
accounting and both budget teeth five rules that already have exactly one home. Two copies of a
dispatch loop is -(p) with a much larger surface than the one-parameter alternative.
*The summary file.* ``<outbox>/<run-id>-multibase.json`` answers the question no per-base artefact
can: in which ORDER the bases were spent, which ``run_id`` each one got, what was never reached,
which candidates two bases both described, and why each base stopped. Its ``stop_reason`` is READ
BACK from each base's own ``{run_id}-coverage.json`` rather than recomputed, because that file is
where P19 D2 put the fact and a second derivation of it would be free to disagree.
"""
from __future__ import annotations
import json
import shutil
from pathlib import Path
from typing import Any
import pytest
from portfolio_optimiser import explore, okf, run as run_module
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
_BYGG = _EXAMPLES / "bygg-energi-mikro" # BYGG-KONTOR-NORD
_TUNNEL = _EXAMPLES / "tunnel-hauglia" # TUNNEL-HAUGLIA
_PROPOSAL = json.dumps(
{
"measure": "Redusert omfang",
"affected_items": [{"code": "ENERGI-TOTAL-EL", "quantity": 300000.0, "unit_cost": 1.0}],
"claimed_saving_nok": 30000,
}
)
#: The CHECKER carries the tool-calling step list, never the proposer: a proposer script is
#: consumed a SECOND time by the fresh client ``generate_via_llm`` builds, so a leading
#: ``function_call`` there would answer the generation call too and burn an attempt on a parse
#: failure (``_load_scripted_replies``' own stated honesty limit). The checker runs in the debate
#: and nowhere else, so one entry there buys exactly one tool call per base.
_REPLIES: dict[str, Any] = {
"proposer": _PROPOSAL,
"checker": [{"call": "list_bundles", "args": {}}, "VERDICT: APPROVE"],
}
def _mount(tmp_path: Path, *sources: Path) -> list[str]:
out = []
for src in sources:
dst = tmp_path / src.name
shutil.copytree(src, dst)
out.append(str(dst))
return out
def _mandate_file(
tmp_path: Path, *, rows: list[tuple[str, str]], name: str = "mandate.json"
) -> str:
path = tmp_path / name
path.write_text(
json.dumps(
{
"objective": "kutt kostnad i begge basene",
"success_criteria": "minst én tilnærming validerer",
"allow_own_proposals": False,
"approaches": [
{
"id": approach_id,
"label": f"Tilnærming {approach_id}",
"affected_codes": ["ENERGI-TOTAL-EL"],
"claimed_saving_nok": 30000.0,
"bundle_id": bundle_id,
}
for approach_id, bundle_id in rows
],
}
),
encoding="utf-8",
)
return str(path)
def _replies_file(tmp_path: Path, payload: dict[str, Any] | None = None) -> str:
path = tmp_path / "replies.json"
path.write_text(json.dumps(payload if payload is not None else _REPLIES), encoding="utf-8")
return str(path)
def _argv(
bases: list[str], *, mandate: str, run_id: str, outbox: Path, replies: str, extra: list[str]
) -> list[str]:
argv = []
for base in bases:
argv += ["--across-bundle", base]
return argv + [
"--mandate",
mandate,
"--run-id",
run_id,
"--outbox-dir",
str(outbox),
"--scripted-replies",
replies,
"--max-rounds",
"2",
*extra,
]
# ---------------------------------------------------------------------------------------------
# (i) Two outboxes and one summary, each keyed by the operator-chosen minting rule.
# ---------------------------------------------------------------------------------------------
def test_each_base_writes_its_own_artefact_set_under_its_own_minted_run_id(tmp_path: Path) -> None:
"""(i): ``<run-id>-<bundle_id>`` per base, the full artefact set, and ONE summary beside them.
The discriminator against a single shared ``run_id`` is not that files exist but that they are
DISTINCT: with one key the second base would overwrite the first and the outbox would hold one
base's answers under a name claiming to cover both.
"""
bases = _mount(tmp_path, _BYGG, _TUNNEL)
outbox = tmp_path / "out"
rc = run_module.main(
_argv(
bases,
mandate=_mandate_file(
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
),
run_id="X",
outbox=outbox,
replies=_replies_file(tmp_path),
extra=[],
)
)
assert rc == 0
for bundle_id, approach in (("bygg-energi-mikro", "a"), ("tunnel-hauglia", "b")):
stem = f"X-{bundle_id}"
assert (outbox / f"{stem}-{approach}-proposal.json").is_file()
assert (outbox / f"{stem}-{approach}-outcome.json").is_file()
assert (outbox / f"{stem}-debate.json").is_file()
assert (outbox / f"{stem}-runconfig.json").is_file()
assert (outbox / f"{stem}-coverage.json").is_file()
summary = json.loads((outbox / "X-multibase.json").read_text(encoding="utf-8"))
assert summary["run_id"] == "X"
assert [row["bundle_id"] for row in summary["runs"]] == [
"bygg-energi-mikro",
"tunnel-hauglia",
]
assert [row["run_id"] for row in summary["runs"]] == [
"X-bygg-energi-mikro",
"X-tunnel-hauglia",
]
assert summary["unreached"] == []
assert summary["collisions"] == []
assert summary["stopped_early"] is False
assert summary["budget_stop"] is None
assert summary["completed"] is True
# Read BACK from each base's own coverage file, never recomputed here.
assert [row["stop_reason"] for row in summary["runs"]] == ["", ""]
def test_the_summary_reports_the_stop_reason_each_base_recorded(tmp_path: Path) -> None:
"""(i), the half a clean pass cannot show: a base the round cap cut short says so.
An unparseable proposer burns the round cap, the FIRST approach to hit it re-raises by design
(``_evaluate_mandate`` only swallows mid-list), and the pass therefore ends rc 1 which is
exactly what round 3 measured on ``kontrakt-sorasen-2027-04``. The summary must still be
there, must say ``rounds`` rather than the empty string that means "the run finished", and
must say ``completed: false`` rather than leaving an empty ``unreached`` to be read as
"nothing was left unreached".
"""
bases = _mount(tmp_path, _BYGG)
outbox = tmp_path / "out"
rc = run_module.main(
[
"--across-bundle",
bases[0],
"--mandate",
_mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro")]),
"--run-id",
"X",
"--outbox-dir",
str(outbox),
"--scripted-replies",
_replies_file(tmp_path, {"proposer": "not json at all", "checker": "VERDICT: APPROVE"}),
"--max-rounds",
"1",
]
)
assert rc == 1
summary = json.loads((outbox / "X-multibase.json").read_text(encoding="utf-8"))
assert summary["runs"][0]["stop_reason"] == "rounds"
assert summary["completed"] is False
assert summary["runs"][0]["run_id"] == "X-bygg-energi-mikro"
# ---------------------------------------------------------------------------------------------
# (ii) + (iii) The two refusals the dispatch owns, both as rc 1 with a readable message.
# ---------------------------------------------------------------------------------------------
def test_an_approach_naming_a_base_this_run_does_not_configure_is_refused(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""(ii): ``MandateRoutingError`` reaches the operator as rc 1 and names the base."""
bases = _mount(tmp_path, _BYGG)
rc = run_module.main(
_argv(
bases,
mandate=_mandate_file(tmp_path, rows=[("a", "en-base-som-ikke-er-med")]),
run_id="X",
outbox=tmp_path / "out",
replies=_replies_file(tmp_path),
extra=[],
)
)
assert rc == 1
err = capsys.readouterr().err
assert "en-base-som-ikke-er-med" in err
assert "bygg-energi-mikro" in err
def test_two_bases_declaring_one_id_are_refused(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""(iii): the id is how a mandate NAMES a base, so two bases answering to it is a refusal.
Both copies DECLARE the same id on their root ``index.md`` mount-derived ids could never
collide under two directory names, so a test built on directory names alone would prove
nothing about the branch that fires here.
"""
bases = _mount(tmp_path, _BYGG)
second = tmp_path / "andre-base"
shutil.copytree(_BYGG, second)
bases.append(str(second))
for base in bases:
index = Path(base) / "index.md"
text = index.read_text(encoding="utf-8")
index.write_text(text.replace("---\n", "---\nbundle_id: delt-id\n", 1), encoding="utf-8")
rc = run_module.main(
_argv(
bases,
mandate=_mandate_file(tmp_path, rows=[("a", "delt-id")]),
run_id="X",
outbox=tmp_path / "out",
replies=_replies_file(tmp_path),
extra=[],
)
)
assert rc == 1
assert "delt-id" in capsys.readouterr().err
# ---------------------------------------------------------------------------------------------
# (iv) ONE store across the pass — the event, not a substring both branches share.
# ---------------------------------------------------------------------------------------------
def test_the_second_base_sees_the_verdict_the_first_base_minted(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""(iv): identity of the store AND the verdict ids present when each base started.
``VerdictStore`` is a pydantic model with VALUE equality, so three distinct EMPTY stores are
all ``==`` (økt 58's measured vacuity). Identity is the claim, and it is paired with the
EVENT: the set of verdict ids already in the store when base 2 began must contain the id base
1 minted, which a fresh-store-per-base implementation cannot produce.
"""
bases = _mount(tmp_path, _BYGG, _TUNNEL)
seen: list[tuple[int, frozenset[str]]] = []
real = run_module.run_project
async def recording(*args: Any, **kwargs: Any) -> Any:
store = kwargs["store"]
seen.append((id(store), frozenset(v.id for v in store.verdicts)))
return await real(*args, **kwargs)
monkeypatch.setattr(run_module, "run_project", recording)
rc = run_module.main(
_argv(
bases,
mandate=_mandate_file(
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
),
run_id="X",
outbox=tmp_path / "out",
replies=_replies_file(tmp_path),
extra=["--decision", "approved", "--rationale", "expert reviewed (scripted)"],
)
)
assert rc == 0
assert len(seen) == 2
assert seen[0][0] == seen[1][0], "each base was handed a FRESH store, so nothing carries over"
assert seen[0][1] == frozenset(), "base 1 started with a store that already held something"
assert seen[1][1], "base 2 started with an EMPTY store — base 1's verdict never reached it"
# ---------------------------------------------------------------------------------------------
# (v) Two bases, two ``opened`` sinks — the requirement gate reads the base it is asked about.
# ---------------------------------------------------------------------------------------------
def test_the_requirement_gate_reads_the_second_bases_own_opened_list(tmp_path: Path) -> None:
"""(v), driven through the tool bodies DIRECTLY.
A ``ScriptedChatClient`` returns TEXT and never emits a ``function_call`` for a constant
reply, so no scripted run reaches a tool body økt 56's measured vacuity, and the reason the
repo's own answer is to call each tool by name. Base 1 opens a document; declaring THAT path
against base 2 must refuse, because base 2's own sink never saw it.
"""
first, second = _mount(tmp_path, _BYGG, _TUNNEL)
opened_1: list[explore.ToolCall] = []
opened_2: list[explore.ToolCall] = []
reqs_1: list[explore.DeclaredRequirement] = []
reqs_2: list[explore.DeclaredRequirement] = []
tools_1 = {
t.name: t for t in explore.navigator_tools([first], opened=opened_1, requirements=reqs_1)
}
tools_2 = {
t.name: t for t in explore.navigator_tools([second], opened=opened_2, requirements=reqs_2)
}
bundle_1 = okf.reconcile_bundle_id(first).id
bundle_2 = okf.reconcile_bundle_id(second).id
doc_1 = okf.navigate_bundle(first).context_files[0].name
doc_2 = okf.navigate_bundle(second).context_files[0].name
# ``opened`` is filled by the RECORDER middleware, never by the tool bodies (S2c/MAJOR-1:
# what a run read is a property of the invocation, not of the tool's return value), so the
# arm records the call the way a run does and then asks the gate about it.
tools_1["read_file"].func(bundle_1, doc_1)
opened_1.append(explore.ToolCall(name="read_file", bundle_id=bundle_1, path=doc_1))
assert not opened_2, "the two bases shared one opened sink"
refusal = tools_2["declare_requirement"].func(bundle_2, doc_1, "Krav 1")
assert refusal["refusal"] == "RequirementNotRead", refusal
assert not reqs_2, "base 2 recorded a requirement it never read"
tools_2["read_file"].func(bundle_2, doc_2)
opened_2.append(explore.ToolCall(name="read_file", bundle_id=bundle_2, path=doc_2))
accepted = tools_2["declare_requirement"].func(bundle_2, doc_2, "Krav 1")
assert accepted.get("declared") is True, accepted
assert [r.path for r in reqs_2] == [doc_2]
assert not reqs_1, "the two bases shared one requirements sink"
def test_each_bases_debate_artefact_lists_only_its_own_tool_calls(tmp_path: Path) -> None:
"""(v), the half the direct call cannot see: ``run_project`` builds the sinks PER BASE.
The arm above proves the gate reads the sink it is given; this proves the dispatch gives each
base its own. With one shared sink base 2's artefact would carry base 1's call as well, so the
discriminator is the COUNT and not the presence.
"""
bases = _mount(tmp_path, _BYGG, _TUNNEL)
outbox = tmp_path / "out"
rc = run_module.main(
_argv(
bases,
mandate=_mandate_file(
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
),
run_id="X",
outbox=outbox,
replies=_replies_file(tmp_path),
extra=[],
)
)
assert rc == 0
counts = []
for bundle_id in ("bygg-energi-mikro", "tunnel-hauglia"):
payload = json.loads((outbox / f"X-{bundle_id}-debate.json").read_text(encoding="utf-8"))
counts.append([call["name"] for call in payload["tool_calls"]])
assert counts == [["list_bundles"], ["list_bundles"]], counts
# ---------------------------------------------------------------------------------------------
# The three partition rows. Each refusal is paired with an rc-0 control on an argv that is
# otherwise ACCEPTED — without it, rc 1 could come from anywhere in the argv.
# ---------------------------------------------------------------------------------------------
def test_the_flag_is_refused_in_portfolio_mode_by_name(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
assert (
run_module.main(
["--portfolio", "--across-bundle", str(tmp_path), "--goals", str(tmp_path / "g.json")]
)
== 1
)
assert "--across-bundle" in capsys.readouterr().err
def test_the_flag_is_refused_in_report_mode_by_name(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
ledger = tmp_path / "ledger.json"
ledger.write_text("[]", encoding="utf-8")
assert run_module.main(["--report", "--ledger", str(ledger)]) == 0, "control: report mode works"
capsys.readouterr()
assert (
run_module.main(["--report", "--ledger", str(ledger), "--across-bundle", str(tmp_path)])
== 1
)
assert "mode-exclusive" in capsys.readouterr().err
def test_a_dry_run_drills_every_base_and_stops_before_the_first_call(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""The third row is a WIRING, not a refusal: the drill must cover every configured base.
Silently dropping the flag here is the F4 class the dry run would fall through to the
single-project branch, which has no ``PROJECT_ID`` in this argv at all.
"""
bases = _mount(tmp_path, _BYGG, _TUNNEL)
rc = run_module.main(
[
*sum([["--across-bundle", b] for b in bases], []),
"--mandate",
_mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]),
"--run-id",
"X",
"--outbox-dir",
str(tmp_path / "out"),
"--live-dry-run",
]
)
out = capsys.readouterr().out
assert rc == 0
assert out.count("LIVE-DRY-RUN OK") == 2
assert "bygg-energi-mikro" in out and "tunnel-hauglia" in out
# The per-base notices are the discriminator against ONE drill that merely names two bases:
# ``bygg-energi-mikro`` ships no ``cost-baseline.json`` and ``tunnel-hauglia`` does, so exactly
# one unanchored notice and exactly one grounding offer must appear — and a drill of only the
# first, or only the second, produces a different count either way.
assert out.count("Grounding offer") == 1
assert out.count("Cost baseline: NONE") == 1
@pytest.mark.parametrize(
"missing, token",
[("--mandate", "--mandate"), ("--run-id", "--run-id"), ("--outbox-dir", "--outbox-dir")],
)
def test_the_flag_requires_the_three_things_a_multi_base_pass_cannot_invent(
tmp_path: Path, capsys: pytest.CaptureFixture[str], missing: str, token: str
) -> None:
base = _mount(tmp_path, _BYGG)[0]
full = {
"--mandate": _mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro")]),
"--run-id": "X",
"--outbox-dir": str(tmp_path / "out"),
}
argv = ["--across-bundle", base]
for name, value in full.items():
if name != missing:
argv += [name, value]
assert run_module.main(argv) == 1
assert token in capsys.readouterr().err
@pytest.mark.parametrize(
"extra",
[
["--bundle-dir", "somewhere"],
["--explore", "en prompt"],
["--prepass-payload", "payload.json"],
["--proposals-from-mandate"],
],
)
def test_the_flag_refuses_every_single_base_mode_by_name(
tmp_path: Path, capsys: pytest.CaptureFixture[str], extra: list[str]
) -> None:
base = _mount(tmp_path, _BYGG)[0]
argv = [
"--across-bundle",
base,
"--mandate",
_mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro")]),
"--run-id",
"X",
"--outbox-dir",
str(tmp_path / "out"),
*extra,
]
assert run_module.main(argv) == 1
err = capsys.readouterr().err
assert "--across-bundle" in err and extra[0] in err