feat(p17b): ONE commission, SEVERAL bases -- reachable from the command line
``run_mandate_across_bundles`` has existed since session 58, reachable from FIVE
test files and from NO command line (measured: ``grep -n across-bundle run.py``
= 0 hits). ``--across-bundle <dir>``, repeated once per base, is that door.
The engine takes a CALLBACK rather than an outbox directory. Its own docstring
has always said N runs need N ``run_id``s and that minting them there would
default a key this repo requires a caller to supply -- so ``outbox_for`` is that
contract KEPT, not relaxed, and the operator-chosen ``<run-id>-<bundle_id>``
rule lives in ``main()`` where the decision was made. The order's alternative (a
caller running ``run_project`` itself over ``route_by_bundle``'s sub-mandates)
would be a second copy of the loop's id reconciliation, shared store, per-base
project resolution, collision accounting and both budget teeth.
``resolve_bundle_routing`` is ONE resolution shared by the engine and the
dry-run arm: a free trip answering with a different project id, or tolerating a
duplicate id the paid dispatch refuses, would rehearse a different run.
``{run-id}-multibase.json`` is written from a ``finally`` and every row is built
from the resolution plus disk, so the pass a cap cut short still leaves the
record. ``completed`` is a required field for ``ExplorationTrace.completed``'s
reason. ``stop_reason`` is read BACK from each base's own coverage artefact.
Load-bearing MEASURED (17 arms), four mutations all red against the WHOLE suite,
green control 1761/5 (from 1744/5, superset, 0 removed), golden byte-unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
fc26d4f0a0
commit
5e4c497a84
5 changed files with 994 additions and 37 deletions
41
CLAUDE.md
41
CLAUDE.md
|
|
@ -2728,6 +2728,47 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
|
||||||
runde betaler første approach, den andre er den taket kutter). **Ærlighets-grense, uttalt:**
|
runde betaler første approach, den andre er den taket kutter). **Ærlighets-grense, uttalt:**
|
||||||
`token_usage` er kjøringens ENE teller — den skiller ikke debatt fra generering, og en
|
`token_usage` er kjøringens ENE teller — den skiller ikke debatt fra generering, og en
|
||||||
per-fase-fordeling ville krevd en andre måler.
|
per-fase-fordeling ville krevd en andre måler.
|
||||||
|
- **ÉN kommisjon, FLERE baser er nåbar fra CLI-en — og utboksens nøkkel er KALLERENS, aldri
|
||||||
|
motorens (P17b DEL 1, 15.09):** `run_mandate_across_bundles` har eksistert siden økt 58, nåbar
|
||||||
|
fra FEM testfiler og fra INGEN kommandolinje (MÅLT: `grep -n across-bundle run.py` = 0 treff).
|
||||||
|
`--across-bundle <dir>` (repeterbart; `--bundle-dir` forblir ÉN katalog og er NEKTET her) krever
|
||||||
|
`--mandate`, `--run-id` og `--outbox-dir`. **Motoren fikk en CALLBACK, ikke en `outbox_dir`:**
|
||||||
|
dens egen docstring har alltid sagt at N kjøringer trenger N `run_id`-er og at å mynte dem der
|
||||||
|
ville defaultet en nøkkel repoet krever at en kaller oppgir — så `outbox_for(bundle_id) ->
|
||||||
|
(dir, run_id)` er dét kravet OPPFYLT, ikke slakket, og myntingsregelen `<run-id>-<bundle_id>`
|
||||||
|
(OPERATØRVALGT 14.09) bor i `main()` der beslutningen ble tatt. Ordrens andre alternativ — en
|
||||||
|
kaller som kjører `run_project` selv over `route_by_bundle`s sub-mandater — ville vært en ANDRE
|
||||||
|
kopi av løkkas id-avstemming, delte store, per-base-prosjektoppslag, kollisjonsregnskap og
|
||||||
|
BEGGE budsjett-tenner: kø-(p) over fem regler som hver har ett hjem.
|
||||||
|
**`resolve_bundle_routing` er ÉN oppløsning, delt av motoren og dry-run-armen:** en gratis tur
|
||||||
|
som svarte med en annen `project_id`, eller tolererte en duplisert id den betalte kjøringen
|
||||||
|
nekter, ville vært en generalprøve på en annen kjøring. **Samlefila skrives fra en `finally`**
|
||||||
|
(`write_parse_failures`-presedensen) og hver rad bygges av RESOLUSJONEN + DISK — de konfigurerte
|
||||||
|
basene, kallerens egen myntingsregel, og hver bases egen `{run_id}-coverage.json` — så den
|
||||||
|
kjøringen som mest trenger regnskapet, den et tak kappet, etterlater det. `completed` er et
|
||||||
|
EGET påkrevd felt (`ExplorationTrace.completed`s grunn ordrett): «ingenting ble uoppnådd» og
|
||||||
|
«vi fikk aldri vite» må ikke være samme verdi. `stop_reason` LESES TILBAKE fra coverage-fila,
|
||||||
|
aldri utledet på nytt — P19 D2 la faktumet der, og en andre utledning her ville stått fritt til
|
||||||
|
å være uenig med den dommeren leser. **`BudgetExceeded` er i nekt-tuppelen** av
|
||||||
|
enkeltprosjekt-stiens MÅLTE grunn: den er en `RuntimeError`, og den FØRSTE tilnærmingen som
|
||||||
|
treffer taket re-raiser ved design — over flere baser er dét ikke et kanttilfelle (runde 3 målte
|
||||||
|
`stop_reason: rounds` i 5 av 5), så uten armen er en multi-base-kjørings vanligste utfall en
|
||||||
|
traceback. Tre partisjons-rader: `report_forbidden` og `--portfolio` NEKTER ved navn, mens
|
||||||
|
live-dry-run-raden er en WIRING — den driller HVER konfigurert base og printer hver bases egne
|
||||||
|
varsler, fordi en kjøring som sa dem én gang bare kunne snakket om én av N. Load-bearing MÅLT
|
||||||
|
(`tests/test_across_bundles_cli_loadbearing.py`, 17 armer), **fire mutasjoner alle røde mot HELE
|
||||||
|
suiten** + grønn kontroll **1761/5** (fra 1744/5, supersett, 0 fjernet) og golden
|
||||||
|
`demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
|
||||||
|
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): (i) samlefila droppes (2 røde) · (ii) myntingen
|
||||||
|
kollapser til bart `<run-id>` (3) · (iii) `report_forbidden` slipper flagget stille (1) ·
|
||||||
|
(iv) `opened`/`requirements`-sinkene deles mellom basene (5, hvorav FIRE i tester eldre enn
|
||||||
|
dette arbeidet — uavhengige vitner på at sinkene er per kjøring). **Ærlighets-grenser, uttalt:**
|
||||||
|
en base som RAISER propagerer fortsatt (motorens egen dokumenterte grense — `collect-and-continue`
|
||||||
|
tilhører `run_portfolio`), så de etterfølgende basene kjøres ikke og samlefila sier `completed:
|
||||||
|
false`; `announce` sier fortsatt «the portfolio» når ingen `project_id` er gitt; den hostede
|
||||||
|
flaten er BEVISST urørt (feltet er i ingen av hostings tre sett); og `--proposal-review` trås
|
||||||
|
gjennom til ÉN terminal delt av alle basene (motorens egen begrunnelse — dispatchen er
|
||||||
|
sekvensiell).
|
||||||
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
||||||
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
||||||
|
|
||||||
|
|
|
||||||
15
README.md
15
README.md
|
|
@ -412,7 +412,7 @@ when the seam is detached, so the loop cannot silently degrade into theater.
|
||||||
are still made and discarded, and nothing can leave the process. The OTLP exporter packages are
|
are still made and discarded, and nothing can leave the process. The OTLP exporter packages are
|
||||||
not declared dependencies (they are egress, and heavy in a published wheel); install one yourself
|
not declared dependencies (they are egress, and heavy in a published wheel); install one yourself
|
||||||
if you use that mode.
|
if you use that mode.
|
||||||
- **Run:** the `run.py` CLI has **three modes** — a documented partition, since one invocation
|
- **Run:** the `run.py` CLI has **four modes** — a documented partition, since one invocation
|
||||||
cannot exercise every flag:
|
cannot exercise every flag:
|
||||||
- **Single-project** — `PROJECT_ID --docs-dir <dir>`, plus optional `--bundle-dir`,
|
- **Single-project** — `PROJECT_ID --docs-dir <dir>`, plus optional `--bundle-dir`,
|
||||||
`--verdict-dir`, `--outbox-dir` (which requires `--run-id`), `--dimension-config`,
|
`--verdict-dir`, `--outbox-dir` (which requires `--run-id`), `--dimension-config`,
|
||||||
|
|
@ -425,6 +425,19 @@ when the seam is detached, so the loop cannot silently degrade into theater.
|
||||||
shape the mandate this run evaluates — see below), and `--derive-cost-baseline` (opt-in:
|
shape the mandate this run evaluates — see below), and `--derive-cost-baseline` (opt-in:
|
||||||
anchor the validator on a priced schedule already inside `--bundle-dir` instead of a
|
anchor the validator on a priced schedule already inside `--bundle-dir` instead of a
|
||||||
hand-written `cost-baseline.json` — see below).
|
hand-written `cost-baseline.json` — see below).
|
||||||
|
- **Multi-base (P17b)** — `--across-bundle <dir>` repeated once per knowledge base, plus
|
||||||
|
`--mandate <file>`, `--run-id <id>` and `--outbox-dir <dir>` (all three required). One
|
||||||
|
commission, several bases: the mandate is partitioned by each approach's `bundle_id` and the
|
||||||
|
existing pipeline runs once per base, sequentially, threading ONE verdict store so a verdict
|
||||||
|
minted against base *k* reaches base *k+1*'s hypothesis. Each base writes its own full
|
||||||
|
artefact set under `<run-id>-<bundle_id>`, and one `<run-id>-multibase.json` beside them
|
||||||
|
records the spend order, the per-base `run_id`, `unreached`, `collisions`, `budget_stop` and
|
||||||
|
each base's own `stop_reason` — written even when a base is cut short, with a `completed`
|
||||||
|
field so "nothing was left unreached" cannot be read as "we never found out". `--live-dry-run`
|
||||||
|
drills every configured base and stops before the first model call. `--bundle-dir` stays ONE
|
||||||
|
directory and is refused here, as are `--portfolio`, `--explore`, `--prepass-payload` and
|
||||||
|
`--proposals-from-mandate` — each of those resolves one base, and picking which of N was meant
|
||||||
|
is not this layer's to decide.
|
||||||
- **Portfolio** — `--portfolio`, plus optional `--goals`, `--ledger`, `--dimension-config`,
|
- **Portfolio** — `--portfolio`, plus optional `--goals`, `--ledger`, `--dimension-config`,
|
||||||
`--semantic-retrieval`; it stops early and prints a `goal reached: …` line when the
|
`--semantic-retrieval`; it stops early and prints a `goal reached: …` line when the
|
||||||
accumulated ledger meets a goal.
|
accumulated ledger meets a goal.
|
||||||
|
|
|
||||||
|
|
@ -230,6 +230,55 @@ def write_debate_tools(
|
||||||
return path
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def write_multibase(
|
||||||
|
outbox_dir: str,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
runs: Sequence[Mapping[str, Any]],
|
||||||
|
completed: bool,
|
||||||
|
unreached: Sequence[Mapping[str, Any]],
|
||||||
|
collisions: Sequence[Mapping[str, Any]],
|
||||||
|
stopped_early: bool,
|
||||||
|
budget_stop: Mapping[str, Any] | None,
|
||||||
|
) -> Path:
|
||||||
|
"""Write ``{run_id}-multibase.json`` — what ONE commission did across SEVERAL bases (P17b).
|
||||||
|
|
||||||
|
The question no per-base artefact can answer. Each base writes its own full set under its own
|
||||||
|
minted ``run_id``, but nothing in that set says in which ORDER the bases were spent, which id
|
||||||
|
each one was given, which approaches were never reached, or which candidates two bases both
|
||||||
|
described — and a reader who has to reconstruct the ``<run-id>-<bundle_id>`` convention to
|
||||||
|
pair the files back to the pass has been handed a naming rule instead of a record.
|
||||||
|
|
||||||
|
Written IFF the pass was given an outbox, exactly like its neighbours, and the per-base
|
||||||
|
``stop_reason`` rows are READ BACK from each base's own ``{run_id}-coverage.json`` by the
|
||||||
|
caller rather than recomputed here: P19 D2 put that fact in that file, and a second derivation
|
||||||
|
of it would be free to disagree with the one the judge reads.
|
||||||
|
|
||||||
|
Plain data only, so the RAW output layer stays MAF-free (``write_debate_tools``' own rule)."""
|
||||||
|
directory = Path(outbox_dir)
|
||||||
|
directory.mkdir(parents=True, exist_ok=True)
|
||||||
|
path = directory / f"{run_id}-multibase.json"
|
||||||
|
path.write_text(
|
||||||
|
_dump(
|
||||||
|
{
|
||||||
|
"run_id": run_id,
|
||||||
|
# REQUIRED, never inferred from an empty ``unreached``: a pass a cap or a provider
|
||||||
|
# cut short never got to say what it did not reach, and "nothing was left
|
||||||
|
# unreached" must not be the value that means "we never found out"
|
||||||
|
# (``ExplorationTrace.completed``'s own reason).
|
||||||
|
"completed": completed,
|
||||||
|
"runs": [dict(row) for row in runs],
|
||||||
|
"unreached": [dict(row) for row in unreached],
|
||||||
|
"collisions": [dict(row) for row in collisions],
|
||||||
|
"stopped_early": stopped_early,
|
||||||
|
"budget_stop": dict(budget_stop) if budget_stop is not None else None,
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
def write_prepass(
|
def write_prepass(
|
||||||
outbox_dir: str,
|
outbox_dir: str,
|
||||||
run_id: str,
|
run_id: str,
|
||||||
|
|
|
||||||
|
|
@ -29,7 +29,7 @@ import asyncio
|
||||||
import json
|
import json
|
||||||
from collections.abc import Awaitable, Callable, Iterable, Sequence
|
from collections.abc import Awaitable, Callable, Iterable, Sequence
|
||||||
from contextlib import AsyncExitStack
|
from contextlib import AsyncExitStack
|
||||||
from dataclasses import dataclass, replace
|
from dataclasses import asdict, dataclass, replace
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Literal, cast
|
from typing import Any, Literal, cast
|
||||||
|
|
||||||
|
|
@ -2181,6 +2181,12 @@ class BundleRun:
|
||||||
bundle_dir: str
|
bundle_dir: str
|
||||||
project_id: str
|
project_id: str
|
||||||
result: RunResult
|
result: RunResult
|
||||||
|
#: The ``run_id`` this base's artefacts were written under, or ``""`` when the pass wrote no
|
||||||
|
#: outbox at all. Carried here rather than re-derived by the caller for the reason the two
|
||||||
|
#: fields above are: the CALLER minted it (P17b — N runs need N ids, and this engine refuses
|
||||||
|
#: to default a key the repo requires a caller to supply), so a reader pairing an artefact
|
||||||
|
#: back to a base must not have to reconstruct the naming convention to do it.
|
||||||
|
run_id: str = ""
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
|
|
@ -2213,6 +2219,116 @@ class MultiBaseResult:
|
||||||
collisions: tuple[VerdictCollision, ...] = ()
|
collisions: tuple[VerdictCollision, ...] = ()
|
||||||
|
|
||||||
|
|
||||||
|
def _coverage_stop_reason(outbox_dir: str, run_id: str) -> str:
|
||||||
|
"""``BudgetExceeded.kind`` this base recorded, read back off its OWN coverage artefact.
|
||||||
|
|
||||||
|
Never recomputed from the ``RunResult``: P19 D2 put "why did this run stop" in
|
||||||
|
``{run_id}-coverage.json`` precisely because ``_evaluate_mandate`` swallows the exception once
|
||||||
|
something has been produced, so the dispatcher never sees it. A second derivation here would
|
||||||
|
be free to disagree with the file the judge reads.
|
||||||
|
|
||||||
|
``"absent"`` is a THIRD value, and not the same as ``""``: a base that wrote no coverage file
|
||||||
|
at all is a different finding from one that finished with nothing stopping it — the
|
||||||
|
``stress`` judge's own vocabulary, reused rather than re-invented."""
|
||||||
|
path = Path(outbox_dir) / f"{run_id}-coverage.json"
|
||||||
|
if not path.is_file():
|
||||||
|
return "absent"
|
||||||
|
try:
|
||||||
|
return str(json.loads(path.read_text(encoding="utf-8")).get("stop_reason", ""))
|
||||||
|
except (OSError, json.JSONDecodeError):
|
||||||
|
return "absent"
|
||||||
|
|
||||||
|
|
||||||
|
def _write_multibase_summary(
|
||||||
|
outbox_dir: str,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
resolved: Sequence[tuple[str, str, str]],
|
||||||
|
mint: Callable[[str], tuple[str, str]],
|
||||||
|
multi: MultiBaseResult | None,
|
||||||
|
) -> None:
|
||||||
|
"""Write ``{run_id}-multibase.json`` for a multi-base pass, COMPLETED or not (P17b).
|
||||||
|
|
||||||
|
Called from a ``finally``, which is ``write_parse_failures``' rule applied one layer up: the
|
||||||
|
pass that most needs a record of what it spent is the one a cap or a provider cut short, and
|
||||||
|
the engine's documented limit is that a base which RAISES propagates. Every per-base row is
|
||||||
|
therefore built from the RESOLUTION and from DISK — the configured bases, the caller's own
|
||||||
|
minting rule, and each base's own ``{run_id}-coverage.json`` — none of which needs the
|
||||||
|
dispatch to have returned.
|
||||||
|
|
||||||
|
``completed`` is a REQUIRED field of the artefact and not an inference from an empty
|
||||||
|
``unreached``: ``ExplorationTrace.completed``'s reason verbatim, because "nothing was left
|
||||||
|
unreached" and "we never found out" must not be the same value. When the pass did not
|
||||||
|
complete, ``unreached``/``collisions``/``budget_stop`` are what the dispatch never got to say,
|
||||||
|
and they are written as empty/``None`` UNDER that flag rather than as findings.
|
||||||
|
"""
|
||||||
|
rows = []
|
||||||
|
for bundle_id, bundle_dir, project_id in resolved:
|
||||||
|
_, base_run_id = mint(bundle_id)
|
||||||
|
rows.append(
|
||||||
|
{
|
||||||
|
"bundle_id": bundle_id,
|
||||||
|
"bundle_dir": bundle_dir,
|
||||||
|
"project_id": project_id,
|
||||||
|
"run_id": base_run_id,
|
||||||
|
"stop_reason": _coverage_stop_reason(outbox_dir, base_run_id),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
outbox.write_multibase(
|
||||||
|
outbox_dir,
|
||||||
|
run_id,
|
||||||
|
runs=rows,
|
||||||
|
completed=multi is not None,
|
||||||
|
unreached=[asdict(row) for row in multi.unreached] if multi is not None else [],
|
||||||
|
collisions=[asdict(row) for row in multi.collisions] if multi is not None else [],
|
||||||
|
stopped_early=multi.stopped_early if multi is not None else False,
|
||||||
|
budget_stop=(
|
||||||
|
asdict(multi.budget_stop)
|
||||||
|
if multi is not None and multi.budget_stop is not None
|
||||||
|
else None
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_bundle_routing(
|
||||||
|
bundle_dirs: Sequence[str],
|
||||||
|
) -> tuple[tuple[str, str, str], ...]:
|
||||||
|
"""``(bundle_id, bundle_dir, project_id)`` per configured base, in CONFIGURED order.
|
||||||
|
|
||||||
|
The ONE resolution shared by ``run_mandate_across_bundles`` and the CLI's multi-base dry run
|
||||||
|
(P17b). Extracted rather than copied for the reason the id derivation itself was unified in
|
||||||
|
Step 10: a drill that answered with a different project id, or tolerated a duplicate id the
|
||||||
|
paid dispatch refuses, would be a free trip that fails to measure the very run it precedes.
|
||||||
|
|
||||||
|
``bundle_id`` is ``okf.reconcile_bundle_id``'s — the declared id wins over the mount (S7a-3).
|
||||||
|
Two bases answering to ONE id refuse here, the same refusal ``explore._bundle_index`` makes
|
||||||
|
and for the same reason: the id is how the mandate NAMES a base, so a collision would let an
|
||||||
|
approach be evaluated against A while the report says B (the S3.2 key-collision class).
|
||||||
|
|
||||||
|
``project_id`` is S7b søm 1's precedence, and it is load-bearing in BOTH directions: the
|
||||||
|
hand-written IR projection FIRST (that file is what every existing base has always been routed
|
||||||
|
by), the base's own DECLARED id as the fallback (an ingested corpus carries no projection, so
|
||||||
|
file-only could not route it at all). No caller-supplied constant is admitted: it could only
|
||||||
|
ever be right for one base out of N.
|
||||||
|
|
||||||
|
:raises MandateRoutingError: two configured bases share one id.
|
||||||
|
"""
|
||||||
|
out: list[tuple[str, str, str]] = []
|
||||||
|
seen: dict[str, str] = {}
|
||||||
|
for raw in bundle_dirs:
|
||||||
|
bundle_id = okf.reconcile_bundle_id(raw).id
|
||||||
|
if bundle_id in seen:
|
||||||
|
raise MandateRoutingError(
|
||||||
|
f"two knowledge bases share the id {bundle_id!r} ({seen[bundle_id]!r} and "
|
||||||
|
f"{raw!r}); an approach names a base by that id, so it must be unique"
|
||||||
|
)
|
||||||
|
seen[bundle_id] = raw
|
||||||
|
declared_ir = okf.load_optional_ir_projection(raw)
|
||||||
|
project_id = str(declared_ir["project_id"]) if declared_ir is not None else bundle_id
|
||||||
|
out.append((bundle_id, raw, project_id))
|
||||||
|
return tuple(out)
|
||||||
|
|
||||||
|
|
||||||
async def run_mandate_across_bundles(
|
async def run_mandate_across_bundles(
|
||||||
mandate: Mandate,
|
mandate: Mandate,
|
||||||
bundle_dirs: Sequence[str],
|
bundle_dirs: Sequence[str],
|
||||||
|
|
@ -2233,6 +2349,17 @@ async def run_mandate_across_bundles(
|
||||||
#: AND ``project_id`` — each base's own, read by ``_project_from_bundle`` — so bases can be
|
#: AND ``project_id`` — each base's own, read by ``_project_from_bundle`` — so bases can be
|
||||||
#: told apart even when a multi-base commission reuses approach ids.
|
#: told apart even when a multi-base commission reuses approach ids.
|
||||||
proposal_reviewer: ProposalReviewer | None = None,
|
proposal_reviewer: ProposalReviewer | None = None,
|
||||||
|
#: P17b. Where THIS base's artefacts go, and under which ``run_id`` — supplied by the caller
|
||||||
|
#: per base, never minted here. That is the engine's own long-standing contract kept rather
|
||||||
|
#: than relaxed: N runs need N ``run_id``s, and minting one here would default a key this repo
|
||||||
|
#: requires a caller to supply, for byte-determinism. A CALLBACK rather than an
|
||||||
|
#: ``outbox_dir``/``run_id`` pair because the naming rule is an OPERATOR decision
|
||||||
|
#: (``<run-id>-<bundle_id>``, chosen 14.09) and belongs at the call site that made it; the
|
||||||
|
#: alternative the order offered — a caller running ``run_project`` itself over
|
||||||
|
#: ``route_by_bundle``'s sub-mandates — would be a SECOND copy of this loop's id
|
||||||
|
#: reconciliation, shared store, per-base project resolution, collision accounting and both
|
||||||
|
#: budget teeth (kø-(p), over five rules that each have exactly one home).
|
||||||
|
outbox_for: Callable[[str], tuple[str, str]] | None = None,
|
||||||
) -> MultiBaseResult:
|
) -> MultiBaseResult:
|
||||||
"""Evaluate ONE commission across SEVERAL knowledge bases — the multi-base dispatch (§ C.7).
|
"""Evaluate ONE commission across SEVERAL knowledge bases — the multi-base dispatch (§ C.7).
|
||||||
|
|
||||||
|
|
@ -2267,30 +2394,18 @@ async def run_mandate_across_bundles(
|
||||||
independent projects, whereas here the caller asked for ONE commission to be evaluated.
|
independent projects, whereas here the caller asked for ONE commission to be evaluated.
|
||||||
(2) Without a ``portfolio_meter`` the pass's ceiling is the number of routed bases times
|
(2) Without a ``portfolio_meter`` the pass's ceiling is the number of routed bases times
|
||||||
``max_tokens``, each run bounded on its own — the global ledger is opt-in, and this does not
|
``max_tokens``, each run bounded on its own — the global ledger is opt-in, and this does not
|
||||||
re-implement it (``_run_meter`` is the one copy of the binding rule). (3) The outbox is NOT
|
re-implement it (``_run_meter`` is the one copy of the binding rule). (3) The outbox is
|
||||||
wired: N runs need N ``run_id``s, and minting them here would default a key this repo requires
|
CALLER-KEYED: N runs need N ``run_id``s, and minting them here would default a key this repo
|
||||||
a caller to supply, for byte-determinism. A caller who needs artefacts per base calls
|
requires a caller to supply, for byte-determinism — so ``outbox_for`` hands each base its
|
||||||
``run_project`` itself with the sub-mandates ``route_by_bundle`` hands back.
|
directory and its id, and a caller that offers no callback still writes nothing (every
|
||||||
|
pre-P17b call site, unchanged).
|
||||||
|
|
||||||
:raises MandateRoutingError: the commission cannot be routed against ``bundle_dirs``.
|
:raises MandateRoutingError: the commission cannot be routed against ``bundle_dirs``.
|
||||||
:raises BudgetRefused: a global remainder that cannot fund a single run.
|
:raises BudgetRefused: a global remainder that cannot fund a single run.
|
||||||
"""
|
"""
|
||||||
by_id: dict[str, str] = {}
|
resolved = resolve_bundle_routing(bundle_dirs)
|
||||||
for raw in bundle_dirs:
|
by_id = {bundle_id: bundle_dir for bundle_id, bundle_dir, _ in resolved}
|
||||||
# The ONE derivation rule (Step 10) — this used to be a second private copy of
|
project_by_id = {bundle_id: project_id for bundle_id, _, project_id in resolved}
|
||||||
# ``Path(raw).name``, free to drift from ``explore``'s. The REFUSAL below stays local:
|
|
||||||
# ``MandateRoutingError`` is this door's class, ``ExplorationError`` is explore's, and
|
|
||||||
# unifying the derivation is not the same as unifying the two doors' error vocabularies.
|
|
||||||
bundle_id = okf.reconcile_bundle_id(raw).id
|
|
||||||
if bundle_id in by_id:
|
|
||||||
# The same refusal ``explore._bundle_index`` makes, for the same reason: the id is how
|
|
||||||
# the mandate names a base, so two bases answering to one name would let an approach be
|
|
||||||
# evaluated against A while the report says B (the S3.2 key-collision class).
|
|
||||||
raise MandateRoutingError(
|
|
||||||
f"two knowledge bases share the id {bundle_id!r} ({by_id[bundle_id]!r} and "
|
|
||||||
f"{raw!r}); an approach names a base by that id, so it must be unique"
|
|
||||||
)
|
|
||||||
by_id[bundle_id] = raw
|
|
||||||
|
|
||||||
routed = route_by_bundle(mandate, tuple(by_id))
|
routed = route_by_bundle(mandate, tuple(by_id))
|
||||||
|
|
||||||
|
|
@ -2327,19 +2442,7 @@ async def run_mandate_across_bundles(
|
||||||
break
|
break
|
||||||
|
|
||||||
bundle_dir = by_id[bundle_id]
|
bundle_dir = by_id[bundle_id]
|
||||||
# ONE reading of the base's own project id, used both to ADDRESS the run and to LABEL it.
|
project_id = project_by_id[bundle_id]
|
||||||
# A second lookup for the label would be the kø-(p) duplicate free to drift from the value
|
|
||||||
# the run was actually dispatched with.
|
|
||||||
#
|
|
||||||
# S7b søm 1: the hand-written projection FIRST, the base's DECLARED id as the fallback. The
|
|
||||||
# precedence is load-bearing in both directions. Declaration-first would re-address every
|
|
||||||
# existing base whose ``project_id`` differs from its ``bundle_id`` — the file is what those
|
|
||||||
# bases have always been routed by. File-only was the refusal this seam removes: an ingested
|
|
||||||
# corpus carries no projection, so it could not be routed at all. ``bundle_id`` is the
|
|
||||||
# identity every other door already resolves through (S7a-3), so the fallback introduces no
|
|
||||||
# third notion of what a base is called; ``by_id`` above is that same resolution, reused.
|
|
||||||
declared_ir = okf.load_optional_ir_projection(bundle_dir)
|
|
||||||
project_id = str(declared_ir["project_id"]) if declared_ir is not None else bundle_id
|
|
||||||
# D2: the id of the verdict THIS base minted, taken from ``run_project``'s existing
|
# D2: the id of the verdict THIS base minted, taken from ``run_project``'s existing
|
||||||
# ``notify`` seam rather than off the returned ``RunResult``. ``notify`` fires inside the
|
# ``notify`` seam rather than off the returned ``RunResult``. ``notify`` fires inside the
|
||||||
# capture block, so it is called exactly when a verdict exists (F2: never when nobody
|
# capture block, so it is called exactly when a verdict exists (F2: never when nobody
|
||||||
|
|
@ -2347,6 +2450,7 @@ async def run_mandate_across_bundles(
|
||||||
# each iteration — a shared accumulator would let a later base read the previous base's
|
# each iteration — a shared accumulator would let a later base read the previous base's
|
||||||
# verdict and manufacture a collision that never happened.
|
# verdict and manufacture a collision that never happened.
|
||||||
minted_here: list[str] = []
|
minted_here: list[str] = []
|
||||||
|
base_outbox, base_run_id = outbox_for(bundle_id) if outbox_for is not None else (None, "")
|
||||||
result = cast(
|
result = cast(
|
||||||
RunResult,
|
RunResult,
|
||||||
await run_project(
|
await run_project(
|
||||||
|
|
@ -2354,6 +2458,8 @@ async def run_mandate_across_bundles(
|
||||||
profile,
|
profile,
|
||||||
docs_dir=bundle_dir,
|
docs_dir=bundle_dir,
|
||||||
bundle_dir=bundle_dir,
|
bundle_dir=bundle_dir,
|
||||||
|
outbox_dir=base_outbox,
|
||||||
|
run_id=base_run_id or None,
|
||||||
notify=lambda verdict: minted_here.append(verdict.id),
|
notify=lambda verdict: minted_here.append(verdict.id),
|
||||||
verdict_input=verdict_input,
|
verdict_input=verdict_input,
|
||||||
verdict_dir=verdict_dir,
|
verdict_dir=verdict_dir,
|
||||||
|
|
@ -2397,6 +2503,7 @@ async def run_mandate_across_bundles(
|
||||||
bundle_dir=bundle_dir,
|
bundle_dir=bundle_dir,
|
||||||
project_id=project_id,
|
project_id=project_id,
|
||||||
result=result,
|
result=result,
|
||||||
|
run_id=base_run_id,
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
@ -2535,6 +2642,18 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--bundle-dir", default=None, help="OKF bundle dir (enables the Step-1 fold)"
|
"--bundle-dir", default=None, help="OKF bundle dir (enables the Step-1 fold)"
|
||||||
)
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--across-bundle",
|
||||||
|
action="append",
|
||||||
|
default=None,
|
||||||
|
metavar="DIR",
|
||||||
|
help="P17b: run ONE commission across SEVERAL knowledge bases — repeat the flag once per "
|
||||||
|
"base. The mandate is partitioned by each approach's bundle_id and the existing pipeline "
|
||||||
|
"runs once per base, sequentially, threading ONE verdict store so a verdict minted against "
|
||||||
|
"base k reaches base k+1. Each base writes its own artefact set under <run-id>-<bundle_id>, "
|
||||||
|
"plus one <run-id>-multibase.json summary. Requires --mandate, --run-id and --outbox-dir; "
|
||||||
|
"--bundle-dir stays ONE directory and is refused here",
|
||||||
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--verdict-dir",
|
"--verdict-dir",
|
||||||
default=None,
|
default=None,
|
||||||
|
|
@ -2888,6 +3007,10 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
"--goals": args.goals is not None,
|
"--goals": args.goals is not None,
|
||||||
"--docs-dir": args.docs_dir is not None,
|
"--docs-dir": args.docs_dir is not None,
|
||||||
"--bundle-dir": args.bundle_dir is not None,
|
"--bundle-dir": args.bundle_dir is not None,
|
||||||
|
# P17b, and for its neighbours' reason: report mode returns ABOVE every dispatch,
|
||||||
|
# including the multi-base one, so an omission here is a SILENT DROP of a whole pass
|
||||||
|
# rather than a refusal (the F4 class).
|
||||||
|
"--across-bundle": bool(args.across_bundle),
|
||||||
"--verdict-dir": args.verdict_dir is not None,
|
"--verdict-dir": args.verdict_dir is not None,
|
||||||
"--outbox-dir": args.outbox_dir is not None,
|
"--outbox-dir": args.outbox_dir is not None,
|
||||||
"--run-id": args.run_id is not None,
|
"--run-id": args.run_id is not None,
|
||||||
|
|
@ -2969,6 +3092,13 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
single_only = {
|
single_only = {
|
||||||
"--docs-dir": args.docs_dir,
|
"--docs-dir": args.docs_dir,
|
||||||
"--bundle-dir": args.bundle_dir,
|
"--bundle-dir": args.bundle_dir,
|
||||||
|
# P17b. A portfolio pass keys on PROJECTS and reads each project's base off its own
|
||||||
|
# row, so a run-level list of bases has nowhere to go; the two are different axes
|
||||||
|
# (``MultiBaseResult`` is a distinct type from ``PortfolioResult`` for exactly that
|
||||||
|
# reason). BY NAME, like its neighbours: falling through to "--across-bundle requires
|
||||||
|
# --mandate" would tell an operator who wrote --portfolio --across-bundle to add a
|
||||||
|
# flag that is legal in both modes, which answers the wrong question.
|
||||||
|
"--across-bundle": bool(args.across_bundle),
|
||||||
"--verdict-dir": args.verdict_dir,
|
"--verdict-dir": args.verdict_dir,
|
||||||
"--outbox-dir": args.outbox_dir,
|
"--outbox-dir": args.outbox_dir,
|
||||||
"--run-id": args.run_id,
|
"--run-id": args.run_id,
|
||||||
|
|
@ -3037,13 +3167,63 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
)
|
)
|
||||||
return 1
|
return 1
|
||||||
|
|
||||||
|
# P17b — the multi-base door's own refusals, at FUNCTION level and never nested under another
|
||||||
|
# flag's branch, for the F4 reason its neighbours are: under one, a bare combination falls
|
||||||
|
# straight through to a dispatch that drops the flag in silence. Placed ABOVE the required-args
|
||||||
|
# guard because this mode takes NO ``PROJECT_ID`` at all — each base's project is read from
|
||||||
|
# THAT base's own IR projection, which is the whole reason the dispatch has no such parameter.
|
||||||
|
if args.across_bundle:
|
||||||
|
# The three things a multi-base pass cannot invent. ``--mandate`` because the commission IS
|
||||||
|
# the partition key (without ``Approach.bundle_id`` there is nothing to route on), and the
|
||||||
|
# outbox pair because N runs need N ``run_id``s: the engine refuses to default that key,
|
||||||
|
# so the caller must supply the stem it mints them from.
|
||||||
|
required = {
|
||||||
|
"--mandate": args.mandate,
|
||||||
|
"--run-id": args.run_id,
|
||||||
|
"--outbox-dir": args.outbox_dir,
|
||||||
|
}
|
||||||
|
missing = [name for name, value in required.items() if not value]
|
||||||
|
if missing:
|
||||||
|
print(
|
||||||
|
f"run refused: --across-bundle requires {', '.join(missing)} (the commission is "
|
||||||
|
"what partitions the pass by base, and each base writes its own artefact set "
|
||||||
|
"under <run-id>-<bundle_id> — a key this repo requires a caller to supply)",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1
|
||||||
|
# Four single-base modes, each refused BY NAME rather than by falling through. Every one
|
||||||
|
# of them resolves something from THE base — one directory, one cut, one exploration, one
|
||||||
|
# derived schedule — and silently picking which of N that means is the guessed-shape class
|
||||||
|
# this repo refuses outright.
|
||||||
|
conflicting = {
|
||||||
|
"--bundle-dir": args.bundle_dir,
|
||||||
|
"--explore": args.explore,
|
||||||
|
"--prepass-payload": args.prepass_payload,
|
||||||
|
"--proposals-from-mandate": args.proposals_from_mandate,
|
||||||
|
}
|
||||||
|
clash = [name for name, value in conflicting.items() if value]
|
||||||
|
if clash:
|
||||||
|
print(
|
||||||
|
f"run refused: --across-bundle cannot be combined with {', '.join(clash)} (each "
|
||||||
|
"of those resolves ONE knowledge base — a directory, a declared cut, an "
|
||||||
|
"exploration or a derived schedule — and this mode configures several; which one "
|
||||||
|
"was meant is not something this layer may decide)",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1
|
||||||
|
|
||||||
# Single-project mode requires PROJECT_ID + --docs-dir (compensating for the relaxed argparse
|
# Single-project mode requires PROJECT_ID + --docs-dir (compensating for the relaxed argparse
|
||||||
# required/positional so the legacy contract keeps failing loudly via the refusal surface).
|
# required/positional so the legacy contract keeps failing loudly via the refusal surface).
|
||||||
# HOISTED above the scripted door (below) so an incomplete argv is refused BEFORE the honesty
|
# HOISTED above the scripted door (below) so an incomplete argv is refused BEFORE the honesty
|
||||||
# banner could claim a scripted run happened; the refusal ORDER within single-project mode
|
# banner could claim a scripted run happened; the refusal ORDER within single-project mode
|
||||||
# (required args -> semantic-retrieval -> scripted) is unchanged.
|
# (required args -> semantic-retrieval -> scripted) is unchanged.
|
||||||
if not args.portfolio and (
|
#
|
||||||
args.project_id is None or (args.docs_dir is None and args.bundle_dir is None)
|
# ``not args.across_bundle`` is a MODE test, not a relaxation: the multi-base pass takes no
|
||||||
|
# PROJECT_ID and no single ``--bundle-dir``, and both of those are refused above by name.
|
||||||
|
if (
|
||||||
|
not args.portfolio
|
||||||
|
and not args.across_bundle
|
||||||
|
and (args.project_id is None or (args.docs_dir is None and args.bundle_dir is None))
|
||||||
):
|
):
|
||||||
print(
|
print(
|
||||||
"run refused: single-project mode requires PROJECT_ID and either --docs-dir or "
|
"run refused: single-project mode requires PROJECT_ID and either --docs-dir or "
|
||||||
|
|
@ -3892,6 +4072,158 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
print(settle(coverage))
|
print(settle(coverage))
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
|
# P17b — the multi-base dispatch. Placed BELOW the announcement (the commission is declared
|
||||||
|
# before the work it commissions) and ABOVE the portfolio dispatch, because it is a MODE
|
||||||
|
# rather than a modifier: it runs the pass and returns. The dry-run arm lives INSIDE this
|
||||||
|
# block rather than in the generic ``--live-dry-run`` branch further down, which addresses
|
||||||
|
# ``args.project_id``/``args.bundle_dir`` — neither of which this argv has — and would
|
||||||
|
# therefore drop the whole pass in silence (the F4 class).
|
||||||
|
if args.across_bundle:
|
||||||
|
assert mandate is not None # narrowed by the required-flags refusal above
|
||||||
|
assert args.run_id is not None and args.outbox_dir is not None # same refusal
|
||||||
|
bases = list(args.across_bundle)
|
||||||
|
try:
|
||||||
|
# Resolved and ROUTED on the free trip too: a commission that cannot be executed as
|
||||||
|
# written must be refused while it is still free (the økt-57 hoist), and a drill that
|
||||||
|
# tolerated a routing error the paid pass refuses would be a rehearsal of a different
|
||||||
|
# run.
|
||||||
|
resolved = resolve_bundle_routing(bases)
|
||||||
|
route_by_bundle(mandate, tuple(bundle_id for bundle_id, _, _ in resolved))
|
||||||
|
except (MandateRoutingError, okf.BundleIdMismatch, FileNotFoundError, ValueError) as exc:
|
||||||
|
print(f"run refused: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
if args.live_dry_run:
|
||||||
|
for bundle_id, bundle_dir, project_id in resolved:
|
||||||
|
try:
|
||||||
|
report = asyncio.run(
|
||||||
|
run_project(
|
||||||
|
project_id,
|
||||||
|
args.profile,
|
||||||
|
docs_dir=bundle_dir,
|
||||||
|
bundle_dir=bundle_dir,
|
||||||
|
verdict_dir=args.verdict_dir,
|
||||||
|
dimension=(
|
||||||
|
load_dimension(args.dimension_config)
|
||||||
|
if args.dimension_config
|
||||||
|
else None
|
||||||
|
),
|
||||||
|
max_rounds=args.max_rounds,
|
||||||
|
max_tokens=args.max_tokens,
|
||||||
|
require_cost_baseline=args.require_cost_baseline,
|
||||||
|
mcp_servers=mcp_servers,
|
||||||
|
live_dry_run=True,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
except (ValueError, FileNotFoundError, ValidationError) as exc:
|
||||||
|
print(f"live-dry-run refused: {bundle_id}: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
assert isinstance(report, DryRunReport)
|
||||||
|
print(
|
||||||
|
f"{bundle_id} ({project_id}): LIVE-DRY-RUN OK (profile={report.profile}, "
|
||||||
|
f"models={report.resolved_models}, max_rounds={report.max_rounds}, "
|
||||||
|
f"max_tokens={report.max_tokens}, top_k={report.top_k}) — "
|
||||||
|
"ingen modellkall gjort (stoppet før første debate.run)"
|
||||||
|
)
|
||||||
|
# Every notice the single-base drill prints, per base: a pass whose SECOND base
|
||||||
|
# cannot be anchored is exactly as unanchored as one whose first cannot, and a
|
||||||
|
# drill that said it once would leave the operator guessing which base it meant.
|
||||||
|
for notice in (
|
||||||
|
cost_baseline_notice(report.cost_baseline_anchored),
|
||||||
|
grounding_offer_notice(report.grounding_offer),
|
||||||
|
skipped_links_notice(report.skipped_links),
|
||||||
|
bundle_id_notice(report.bundle_id_source),
|
||||||
|
):
|
||||||
|
if notice is not None:
|
||||||
|
print(notice)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
outbox_dir = args.outbox_dir
|
||||||
|
stem = args.run_id
|
||||||
|
|
||||||
|
def outbox_for(bundle_id: str) -> tuple[str, str]:
|
||||||
|
"""The OPERATOR-CHOSEN minting rule (14.09): ``<run-id>-<bundle_id>``, one stem per
|
||||||
|
base. Written here, at the call site that made the decision, rather than inside the
|
||||||
|
dispatch — which is precisely why the engine takes a callback and not a directory."""
|
||||||
|
return outbox_dir, f"{stem}-{bundle_id}"
|
||||||
|
|
||||||
|
#: Bound by the dispatch when it RETURNS; still ``None`` when a base raised. The summary
|
||||||
|
#: is written from a ``finally`` either way (``write_parse_failures``' precedent), because
|
||||||
|
#: the pass that most needs a record of what it spent is the one a cap or a provider cut
|
||||||
|
#: short — and the engine's documented limit is that a base which RAISES propagates.
|
||||||
|
multi: MultiBaseResult | None = None
|
||||||
|
try:
|
||||||
|
multi = asyncio.run(
|
||||||
|
run_mandate_across_bundles(
|
||||||
|
mandate,
|
||||||
|
tuple(bases),
|
||||||
|
args.profile,
|
||||||
|
verdict_dir=args.verdict_dir,
|
||||||
|
dimension=(
|
||||||
|
load_dimension(args.dimension_config) if args.dimension_config else None
|
||||||
|
),
|
||||||
|
client_factory=scripted_client_factory,
|
||||||
|
max_rounds=args.max_rounds,
|
||||||
|
max_tokens=args.max_tokens,
|
||||||
|
verdict_input=_verdict_input_from_args(args),
|
||||||
|
proposal_reviewer=(
|
||||||
|
terminal_proposal_reviewer() if args.proposal_review else None
|
||||||
|
),
|
||||||
|
outbox_for=outbox_for,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
except ChatClientException as exc:
|
||||||
|
print(f"run stopped: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
except ProposalReviewInputError as exc:
|
||||||
|
print(f"run stopped: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
except (
|
||||||
|
MandateRoutingError,
|
||||||
|
okf.BundleIdMismatch,
|
||||||
|
FileNotFoundError,
|
||||||
|
ValidationError,
|
||||||
|
ValueError,
|
||||||
|
BudgetExceeded,
|
||||||
|
) as exc:
|
||||||
|
# ``BudgetExceeded`` is in the tuple for the single-project path's measured reason: it
|
||||||
|
# is a ``RuntimeError``, and the FIRST approach to hit the cap re-raises by design
|
||||||
|
# (``_evaluate_mandate`` only swallows it mid-list). Over several bases that is not an
|
||||||
|
# edge case — round 3 measured ``stop_reason: rounds`` in 5 of 5 runs — so without this
|
||||||
|
# arm the ordinary outcome of a multi-base pass is a traceback.
|
||||||
|
print(f"run refused: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
finally:
|
||||||
|
_write_multibase_summary(
|
||||||
|
outbox_dir,
|
||||||
|
args.run_id,
|
||||||
|
resolved=resolved,
|
||||||
|
mint=outbox_for,
|
||||||
|
multi=multi,
|
||||||
|
)
|
||||||
|
|
||||||
|
for bundle_run in multi.runs:
|
||||||
|
print(f"--- {bundle_run.bundle_id} ({bundle_run.project_id}) ---")
|
||||||
|
print(
|
||||||
|
f"{bundle_run.project_id}: {type(bundle_run.result.outcome).__name__} "
|
||||||
|
f"({verdict_notice(bundle_run.result)})"
|
||||||
|
)
|
||||||
|
# Every per-run notice the single-base path prints, read off THIS base's own stamp —
|
||||||
|
# a pass that said them once could only be talking about one of N bases, and the
|
||||||
|
# reader could not tell which.
|
||||||
|
for notice in (
|
||||||
|
cost_baseline_notice(bundle_run.result.provenance.cost_baseline_anchored),
|
||||||
|
bundle_id_notice(bundle_run.result.provenance.bundle_id_source),
|
||||||
|
skipped_links_notice(bundle_run.result.skipped_links),
|
||||||
|
unkeyed_verdicts_notice(bundle_run.result.unkeyed_verdicts),
|
||||||
|
):
|
||||||
|
if notice is not None:
|
||||||
|
print(notice)
|
||||||
|
notice = collision_notice(multi.collisions)
|
||||||
|
if notice is not None:
|
||||||
|
print(notice)
|
||||||
|
return 0
|
||||||
|
|
||||||
if args.portfolio:
|
if args.portfolio:
|
||||||
# Portfolio mode (Step 3): dispatch to the EXISTING run_portfolio via the fail-fast loaders
|
# Portfolio mode (Step 3): dispatch to the EXISTING run_portfolio via the fail-fast loaders
|
||||||
# (run_portfolio itself is unchanged). Loader/ValueError failures surface through the same
|
# (run_portfolio itself is unchanged). Loader/ValueError failures surface through the same
|
||||||
|
|
|
||||||
522
tests/test_across_bundles_cli_loadbearing.py
Normal file
522
tests/test_across_bundles_cli_loadbearing.py
Normal file
|
|
@ -0,0 +1,522 @@
|
||||||
|
"""P17b DEL 1 — ONE run, SEVERAL knowledge bases, from the CLI (order ``20260915T014020Z``).
|
||||||
|
|
||||||
|
``run_mandate_across_bundles`` has existed since økt 58 and was reachable from FIVE test files and
|
||||||
|
from NO command line (measured 15.09: ``grep -n across-bundle run.py`` = 0 hits). The operator
|
||||||
|
directive of 14.09 is that ``po`` must demonstrably work against THE BASES — plural — a given run
|
||||||
|
says it will use, and a library function no operator can invoke is not that demonstration.
|
||||||
|
|
||||||
|
**Three things this file pins, and each is a different claim.**
|
||||||
|
|
||||||
|
*The minting rule.* N runs need N ``run_id``s — the engine's own docstring says so, and refuses to
|
||||||
|
default a key this repo requires a caller to supply. The rule is ``<run-id>-<bundle_id>`` and it is
|
||||||
|
OPERATOR-CHOSEN (14.09), not a preference argued here. The engine therefore takes an
|
||||||
|
``outbox_for(bundle_id) -> (outbox_dir, run_id)`` CALLBACK rather than an ``outbox_dir``: the
|
||||||
|
caller supplies the ids, which is exactly the contract the engine documents, and the naming
|
||||||
|
convention stays in ``main()`` where the operator's decision lives.
|
||||||
|
|
||||||
|
*Why a callback rather than a caller-side loop.* The order offered both. A caller that ran
|
||||||
|
``run_project`` itself over ``route_by_bundle``'s sub-mandates would have to re-implement the id
|
||||||
|
reconciliation, the shared ``VerdictStore``, the per-base ``project_id`` resolution, the collision
|
||||||
|
accounting and both budget teeth — five rules that already have exactly one home. Two copies of a
|
||||||
|
dispatch loop is kø-(p) with a much larger surface than the one-parameter alternative.
|
||||||
|
|
||||||
|
*The summary file.* ``<outbox>/<run-id>-multibase.json`` answers the question no per-base artefact
|
||||||
|
can: in which ORDER the bases were spent, which ``run_id`` each one got, what was never reached,
|
||||||
|
which candidates two bases both described, and why each base stopped. Its ``stop_reason`` is READ
|
||||||
|
BACK from each base's own ``{run_id}-coverage.json`` rather than recomputed, because that file is
|
||||||
|
where P19 D2 put the fact and a second derivation of it would be free to disagree.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import shutil
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from portfolio_optimiser import explore, okf, run as run_module
|
||||||
|
|
||||||
|
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||||
|
_BYGG = _EXAMPLES / "bygg-energi-mikro" # BYGG-KONTOR-NORD
|
||||||
|
_TUNNEL = _EXAMPLES / "tunnel-hauglia" # TUNNEL-HAUGLIA
|
||||||
|
|
||||||
|
_PROPOSAL = json.dumps(
|
||||||
|
{
|
||||||
|
"measure": "Redusert omfang",
|
||||||
|
"affected_items": [{"code": "ENERGI-TOTAL-EL", "quantity": 300000.0, "unit_cost": 1.0}],
|
||||||
|
"claimed_saving_nok": 30000,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
#: The CHECKER carries the tool-calling step list, never the proposer: a proposer script is
|
||||||
|
#: consumed a SECOND time by the fresh client ``generate_via_llm`` builds, so a leading
|
||||||
|
#: ``function_call`` there would answer the generation call too and burn an attempt on a parse
|
||||||
|
#: failure (``_load_scripted_replies``' own stated honesty limit). The checker runs in the debate
|
||||||
|
#: and nowhere else, so one entry there buys exactly one tool call per base.
|
||||||
|
_REPLIES: dict[str, Any] = {
|
||||||
|
"proposer": _PROPOSAL,
|
||||||
|
"checker": [{"call": "list_bundles", "args": {}}, "VERDICT: APPROVE"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _mount(tmp_path: Path, *sources: Path) -> list[str]:
|
||||||
|
out = []
|
||||||
|
for src in sources:
|
||||||
|
dst = tmp_path / src.name
|
||||||
|
shutil.copytree(src, dst)
|
||||||
|
out.append(str(dst))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def _mandate_file(
|
||||||
|
tmp_path: Path, *, rows: list[tuple[str, str]], name: str = "mandate.json"
|
||||||
|
) -> str:
|
||||||
|
path = tmp_path / name
|
||||||
|
path.write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"objective": "kutt kostnad i begge basene",
|
||||||
|
"success_criteria": "minst én tilnærming validerer",
|
||||||
|
"allow_own_proposals": False,
|
||||||
|
"approaches": [
|
||||||
|
{
|
||||||
|
"id": approach_id,
|
||||||
|
"label": f"Tilnærming {approach_id}",
|
||||||
|
"affected_codes": ["ENERGI-TOTAL-EL"],
|
||||||
|
"claimed_saving_nok": 30000.0,
|
||||||
|
"bundle_id": bundle_id,
|
||||||
|
}
|
||||||
|
for approach_id, bundle_id in rows
|
||||||
|
],
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
return str(path)
|
||||||
|
|
||||||
|
|
||||||
|
def _replies_file(tmp_path: Path, payload: dict[str, Any] | None = None) -> str:
|
||||||
|
path = tmp_path / "replies.json"
|
||||||
|
path.write_text(json.dumps(payload if payload is not None else _REPLIES), encoding="utf-8")
|
||||||
|
return str(path)
|
||||||
|
|
||||||
|
|
||||||
|
def _argv(
|
||||||
|
bases: list[str], *, mandate: str, run_id: str, outbox: Path, replies: str, extra: list[str]
|
||||||
|
) -> list[str]:
|
||||||
|
argv = []
|
||||||
|
for base in bases:
|
||||||
|
argv += ["--across-bundle", base]
|
||||||
|
return argv + [
|
||||||
|
"--mandate",
|
||||||
|
mandate,
|
||||||
|
"--run-id",
|
||||||
|
run_id,
|
||||||
|
"--outbox-dir",
|
||||||
|
str(outbox),
|
||||||
|
"--scripted-replies",
|
||||||
|
replies,
|
||||||
|
"--max-rounds",
|
||||||
|
"2",
|
||||||
|
*extra,
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# (i) Two outboxes and one summary, each keyed by the operator-chosen minting rule.
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_each_base_writes_its_own_artefact_set_under_its_own_minted_run_id(tmp_path: Path) -> None:
|
||||||
|
"""(i): ``<run-id>-<bundle_id>`` per base, the full artefact set, and ONE summary beside them.
|
||||||
|
|
||||||
|
The discriminator against a single shared ``run_id`` is not that files exist but that they are
|
||||||
|
DISTINCT: with one key the second base would overwrite the first and the outbox would hold one
|
||||||
|
base's answers under a name claiming to cover both.
|
||||||
|
"""
|
||||||
|
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
|
||||||
|
rc = run_module.main(
|
||||||
|
_argv(
|
||||||
|
bases,
|
||||||
|
mandate=_mandate_file(
|
||||||
|
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
|
||||||
|
),
|
||||||
|
run_id="X",
|
||||||
|
outbox=outbox,
|
||||||
|
replies=_replies_file(tmp_path),
|
||||||
|
extra=[],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 0
|
||||||
|
for bundle_id, approach in (("bygg-energi-mikro", "a"), ("tunnel-hauglia", "b")):
|
||||||
|
stem = f"X-{bundle_id}"
|
||||||
|
assert (outbox / f"{stem}-{approach}-proposal.json").is_file()
|
||||||
|
assert (outbox / f"{stem}-{approach}-outcome.json").is_file()
|
||||||
|
assert (outbox / f"{stem}-debate.json").is_file()
|
||||||
|
assert (outbox / f"{stem}-runconfig.json").is_file()
|
||||||
|
assert (outbox / f"{stem}-coverage.json").is_file()
|
||||||
|
|
||||||
|
summary = json.loads((outbox / "X-multibase.json").read_text(encoding="utf-8"))
|
||||||
|
assert summary["run_id"] == "X"
|
||||||
|
assert [row["bundle_id"] for row in summary["runs"]] == [
|
||||||
|
"bygg-energi-mikro",
|
||||||
|
"tunnel-hauglia",
|
||||||
|
]
|
||||||
|
assert [row["run_id"] for row in summary["runs"]] == [
|
||||||
|
"X-bygg-energi-mikro",
|
||||||
|
"X-tunnel-hauglia",
|
||||||
|
]
|
||||||
|
assert summary["unreached"] == []
|
||||||
|
assert summary["collisions"] == []
|
||||||
|
assert summary["stopped_early"] is False
|
||||||
|
assert summary["budget_stop"] is None
|
||||||
|
assert summary["completed"] is True
|
||||||
|
# Read BACK from each base's own coverage file, never recomputed here.
|
||||||
|
assert [row["stop_reason"] for row in summary["runs"]] == ["", ""]
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_summary_reports_the_stop_reason_each_base_recorded(tmp_path: Path) -> None:
|
||||||
|
"""(i), the half a clean pass cannot show: a base the round cap cut short says so.
|
||||||
|
|
||||||
|
An unparseable proposer burns the round cap, the FIRST approach to hit it re-raises by design
|
||||||
|
(``_evaluate_mandate`` only swallows mid-list), and the pass therefore ends rc 1 — which is
|
||||||
|
exactly what round 3 measured on ``kontrakt-sorasen-2027-04``. The summary must still be
|
||||||
|
there, must say ``rounds`` rather than the empty string that means "the run finished", and
|
||||||
|
must say ``completed: false`` rather than leaving an empty ``unreached`` to be read as
|
||||||
|
"nothing was left unreached".
|
||||||
|
"""
|
||||||
|
bases = _mount(tmp_path, _BYGG)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
|
||||||
|
rc = run_module.main(
|
||||||
|
[
|
||||||
|
"--across-bundle",
|
||||||
|
bases[0],
|
||||||
|
"--mandate",
|
||||||
|
_mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro")]),
|
||||||
|
"--run-id",
|
||||||
|
"X",
|
||||||
|
"--outbox-dir",
|
||||||
|
str(outbox),
|
||||||
|
"--scripted-replies",
|
||||||
|
_replies_file(tmp_path, {"proposer": "not json at all", "checker": "VERDICT: APPROVE"}),
|
||||||
|
"--max-rounds",
|
||||||
|
"1",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 1
|
||||||
|
summary = json.loads((outbox / "X-multibase.json").read_text(encoding="utf-8"))
|
||||||
|
assert summary["runs"][0]["stop_reason"] == "rounds"
|
||||||
|
assert summary["completed"] is False
|
||||||
|
assert summary["runs"][0]["run_id"] == "X-bygg-energi-mikro"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# (ii) + (iii) The two refusals the dispatch owns, both as rc 1 with a readable message.
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_approach_naming_a_base_this_run_does_not_configure_is_refused(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
"""(ii): ``MandateRoutingError`` reaches the operator as rc 1 and names the base."""
|
||||||
|
bases = _mount(tmp_path, _BYGG)
|
||||||
|
|
||||||
|
rc = run_module.main(
|
||||||
|
_argv(
|
||||||
|
bases,
|
||||||
|
mandate=_mandate_file(tmp_path, rows=[("a", "en-base-som-ikke-er-med")]),
|
||||||
|
run_id="X",
|
||||||
|
outbox=tmp_path / "out",
|
||||||
|
replies=_replies_file(tmp_path),
|
||||||
|
extra=[],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 1
|
||||||
|
err = capsys.readouterr().err
|
||||||
|
assert "en-base-som-ikke-er-med" in err
|
||||||
|
assert "bygg-energi-mikro" in err
|
||||||
|
|
||||||
|
|
||||||
|
def test_two_bases_declaring_one_id_are_refused(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
"""(iii): the id is how a mandate NAMES a base, so two bases answering to it is a refusal.
|
||||||
|
|
||||||
|
Both copies DECLARE the same id on their root ``index.md`` — mount-derived ids could never
|
||||||
|
collide under two directory names, so a test built on directory names alone would prove
|
||||||
|
nothing about the branch that fires here.
|
||||||
|
"""
|
||||||
|
bases = _mount(tmp_path, _BYGG)
|
||||||
|
second = tmp_path / "andre-base"
|
||||||
|
shutil.copytree(_BYGG, second)
|
||||||
|
bases.append(str(second))
|
||||||
|
for base in bases:
|
||||||
|
index = Path(base) / "index.md"
|
||||||
|
text = index.read_text(encoding="utf-8")
|
||||||
|
index.write_text(text.replace("---\n", "---\nbundle_id: delt-id\n", 1), encoding="utf-8")
|
||||||
|
|
||||||
|
rc = run_module.main(
|
||||||
|
_argv(
|
||||||
|
bases,
|
||||||
|
mandate=_mandate_file(tmp_path, rows=[("a", "delt-id")]),
|
||||||
|
run_id="X",
|
||||||
|
outbox=tmp_path / "out",
|
||||||
|
replies=_replies_file(tmp_path),
|
||||||
|
extra=[],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 1
|
||||||
|
assert "delt-id" in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# (iv) ONE store across the pass — the event, not a substring both branches share.
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_second_base_sees_the_verdict_the_first_base_minted(
|
||||||
|
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||||
|
) -> None:
|
||||||
|
"""(iv): identity of the store AND the verdict ids present when each base started.
|
||||||
|
|
||||||
|
``VerdictStore`` is a pydantic model with VALUE equality, so three distinct EMPTY stores are
|
||||||
|
all ``==`` (økt 58's measured vacuity). Identity is the claim, and it is paired with the
|
||||||
|
EVENT: the set of verdict ids already in the store when base 2 began must contain the id base
|
||||||
|
1 minted, which a fresh-store-per-base implementation cannot produce.
|
||||||
|
"""
|
||||||
|
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||||
|
seen: list[tuple[int, frozenset[str]]] = []
|
||||||
|
real = run_module.run_project
|
||||||
|
|
||||||
|
async def recording(*args: Any, **kwargs: Any) -> Any:
|
||||||
|
store = kwargs["store"]
|
||||||
|
seen.append((id(store), frozenset(v.id for v in store.verdicts)))
|
||||||
|
return await real(*args, **kwargs)
|
||||||
|
|
||||||
|
monkeypatch.setattr(run_module, "run_project", recording)
|
||||||
|
|
||||||
|
rc = run_module.main(
|
||||||
|
_argv(
|
||||||
|
bases,
|
||||||
|
mandate=_mandate_file(
|
||||||
|
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
|
||||||
|
),
|
||||||
|
run_id="X",
|
||||||
|
outbox=tmp_path / "out",
|
||||||
|
replies=_replies_file(tmp_path),
|
||||||
|
extra=["--decision", "approved", "--rationale", "expert reviewed (scripted)"],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 0
|
||||||
|
assert len(seen) == 2
|
||||||
|
assert seen[0][0] == seen[1][0], "each base was handed a FRESH store, so nothing carries over"
|
||||||
|
assert seen[0][1] == frozenset(), "base 1 started with a store that already held something"
|
||||||
|
assert seen[1][1], "base 2 started with an EMPTY store — base 1's verdict never reached it"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# (v) Two bases, two ``opened`` sinks — the requirement gate reads the base it is asked about.
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_requirement_gate_reads_the_second_bases_own_opened_list(tmp_path: Path) -> None:
|
||||||
|
"""(v), driven through the tool bodies DIRECTLY.
|
||||||
|
|
||||||
|
A ``ScriptedChatClient`` returns TEXT and never emits a ``function_call`` for a constant
|
||||||
|
reply, so no scripted run reaches a tool body — økt 56's measured vacuity, and the reason the
|
||||||
|
repo's own answer is to call each tool by name. Base 1 opens a document; declaring THAT path
|
||||||
|
against base 2 must refuse, because base 2's own sink never saw it.
|
||||||
|
"""
|
||||||
|
first, second = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||||
|
|
||||||
|
opened_1: list[explore.ToolCall] = []
|
||||||
|
opened_2: list[explore.ToolCall] = []
|
||||||
|
reqs_1: list[explore.DeclaredRequirement] = []
|
||||||
|
reqs_2: list[explore.DeclaredRequirement] = []
|
||||||
|
tools_1 = {
|
||||||
|
t.name: t for t in explore.navigator_tools([first], opened=opened_1, requirements=reqs_1)
|
||||||
|
}
|
||||||
|
tools_2 = {
|
||||||
|
t.name: t for t in explore.navigator_tools([second], opened=opened_2, requirements=reqs_2)
|
||||||
|
}
|
||||||
|
|
||||||
|
bundle_1 = okf.reconcile_bundle_id(first).id
|
||||||
|
bundle_2 = okf.reconcile_bundle_id(second).id
|
||||||
|
doc_1 = okf.navigate_bundle(first).context_files[0].name
|
||||||
|
doc_2 = okf.navigate_bundle(second).context_files[0].name
|
||||||
|
|
||||||
|
# ``opened`` is filled by the RECORDER middleware, never by the tool bodies (S2c/MAJOR-1:
|
||||||
|
# what a run read is a property of the invocation, not of the tool's return value), so the
|
||||||
|
# arm records the call the way a run does and then asks the gate about it.
|
||||||
|
tools_1["read_file"].func(bundle_1, doc_1)
|
||||||
|
opened_1.append(explore.ToolCall(name="read_file", bundle_id=bundle_1, path=doc_1))
|
||||||
|
assert not opened_2, "the two bases shared one opened sink"
|
||||||
|
|
||||||
|
refusal = tools_2["declare_requirement"].func(bundle_2, doc_1, "Krav 1")
|
||||||
|
assert refusal["refusal"] == "RequirementNotRead", refusal
|
||||||
|
assert not reqs_2, "base 2 recorded a requirement it never read"
|
||||||
|
|
||||||
|
tools_2["read_file"].func(bundle_2, doc_2)
|
||||||
|
opened_2.append(explore.ToolCall(name="read_file", bundle_id=bundle_2, path=doc_2))
|
||||||
|
accepted = tools_2["declare_requirement"].func(bundle_2, doc_2, "Krav 1")
|
||||||
|
assert accepted.get("declared") is True, accepted
|
||||||
|
assert [r.path for r in reqs_2] == [doc_2]
|
||||||
|
assert not reqs_1, "the two bases shared one requirements sink"
|
||||||
|
|
||||||
|
|
||||||
|
def test_each_bases_debate_artefact_lists_only_its_own_tool_calls(tmp_path: Path) -> None:
|
||||||
|
"""(v), the half the direct call cannot see: ``run_project`` builds the sinks PER BASE.
|
||||||
|
|
||||||
|
The arm above proves the gate reads the sink it is given; this proves the dispatch gives each
|
||||||
|
base its own. With one shared sink base 2's artefact would carry base 1's call as well, so the
|
||||||
|
discriminator is the COUNT and not the presence.
|
||||||
|
"""
|
||||||
|
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
|
||||||
|
rc = run_module.main(
|
||||||
|
_argv(
|
||||||
|
bases,
|
||||||
|
mandate=_mandate_file(
|
||||||
|
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
|
||||||
|
),
|
||||||
|
run_id="X",
|
||||||
|
outbox=outbox,
|
||||||
|
replies=_replies_file(tmp_path),
|
||||||
|
extra=[],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert rc == 0
|
||||||
|
counts = []
|
||||||
|
for bundle_id in ("bygg-energi-mikro", "tunnel-hauglia"):
|
||||||
|
payload = json.loads((outbox / f"X-{bundle_id}-debate.json").read_text(encoding="utf-8"))
|
||||||
|
counts.append([call["name"] for call in payload["tool_calls"]])
|
||||||
|
assert counts == [["list_bundles"], ["list_bundles"]], counts
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
# The three partition rows. Each refusal is paired with an rc-0 control on an argv that is
|
||||||
|
# otherwise ACCEPTED — without it, rc 1 could come from anywhere in the argv.
|
||||||
|
# ---------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_flag_is_refused_in_portfolio_mode_by_name(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
assert (
|
||||||
|
run_module.main(
|
||||||
|
["--portfolio", "--across-bundle", str(tmp_path), "--goals", str(tmp_path / "g.json")]
|
||||||
|
)
|
||||||
|
== 1
|
||||||
|
)
|
||||||
|
assert "--across-bundle" in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_flag_is_refused_in_report_mode_by_name(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
ledger = tmp_path / "ledger.json"
|
||||||
|
ledger.write_text("[]", encoding="utf-8")
|
||||||
|
assert run_module.main(["--report", "--ledger", str(ledger)]) == 0, "control: report mode works"
|
||||||
|
capsys.readouterr()
|
||||||
|
assert (
|
||||||
|
run_module.main(["--report", "--ledger", str(ledger), "--across-bundle", str(tmp_path)])
|
||||||
|
== 1
|
||||||
|
)
|
||||||
|
assert "mode-exclusive" in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_dry_run_drills_every_base_and_stops_before_the_first_call(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
"""The third row is a WIRING, not a refusal: the drill must cover every configured base.
|
||||||
|
|
||||||
|
Silently dropping the flag here is the F4 class — the dry run would fall through to the
|
||||||
|
single-project branch, which has no ``PROJECT_ID`` in this argv at all.
|
||||||
|
"""
|
||||||
|
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||||
|
|
||||||
|
rc = run_module.main(
|
||||||
|
[
|
||||||
|
*sum([["--across-bundle", b] for b in bases], []),
|
||||||
|
"--mandate",
|
||||||
|
_mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]),
|
||||||
|
"--run-id",
|
||||||
|
"X",
|
||||||
|
"--outbox-dir",
|
||||||
|
str(tmp_path / "out"),
|
||||||
|
"--live-dry-run",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert rc == 0
|
||||||
|
assert out.count("LIVE-DRY-RUN OK") == 2
|
||||||
|
assert "bygg-energi-mikro" in out and "tunnel-hauglia" in out
|
||||||
|
# The per-base notices are the discriminator against ONE drill that merely names two bases:
|
||||||
|
# ``bygg-energi-mikro`` ships no ``cost-baseline.json`` and ``tunnel-hauglia`` does, so exactly
|
||||||
|
# one unanchored notice and exactly one grounding offer must appear — and a drill of only the
|
||||||
|
# first, or only the second, produces a different count either way.
|
||||||
|
assert out.count("Grounding offer") == 1
|
||||||
|
assert out.count("Cost baseline: NONE") == 1
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"missing, token",
|
||||||
|
[("--mandate", "--mandate"), ("--run-id", "--run-id"), ("--outbox-dir", "--outbox-dir")],
|
||||||
|
)
|
||||||
|
def test_the_flag_requires_the_three_things_a_multi_base_pass_cannot_invent(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str], missing: str, token: str
|
||||||
|
) -> None:
|
||||||
|
base = _mount(tmp_path, _BYGG)[0]
|
||||||
|
full = {
|
||||||
|
"--mandate": _mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro")]),
|
||||||
|
"--run-id": "X",
|
||||||
|
"--outbox-dir": str(tmp_path / "out"),
|
||||||
|
}
|
||||||
|
argv = ["--across-bundle", base]
|
||||||
|
for name, value in full.items():
|
||||||
|
if name != missing:
|
||||||
|
argv += [name, value]
|
||||||
|
assert run_module.main(argv) == 1
|
||||||
|
assert token in capsys.readouterr().err
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"extra",
|
||||||
|
[
|
||||||
|
["--bundle-dir", "somewhere"],
|
||||||
|
["--explore", "en prompt"],
|
||||||
|
["--prepass-payload", "payload.json"],
|
||||||
|
["--proposals-from-mandate"],
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_the_flag_refuses_every_single_base_mode_by_name(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str], extra: list[str]
|
||||||
|
) -> None:
|
||||||
|
base = _mount(tmp_path, _BYGG)[0]
|
||||||
|
argv = [
|
||||||
|
"--across-bundle",
|
||||||
|
base,
|
||||||
|
"--mandate",
|
||||||
|
_mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro")]),
|
||||||
|
"--run-id",
|
||||||
|
"X",
|
||||||
|
"--outbox-dir",
|
||||||
|
str(tmp_path / "out"),
|
||||||
|
*extra,
|
||||||
|
]
|
||||||
|
assert run_module.main(argv) == 1
|
||||||
|
err = capsys.readouterr().err
|
||||||
|
assert "--across-bundle" in err and extra[0] in err
|
||||||
Loading…
Add table
Add a link
Reference in a new issue