feat(s33): wave executor with per-project snapshot and deterministic merge barrier

Replaces run_portfolio's sequential loop with a wave loop over _waves(ids, k):
each wave takes a per-project snapshot of the shared store, runs the wave under
one asyncio.gather in a single event loop, then crosses a merge barrier that
folds each project's NEW verdicts back in wave-submission order. Step 2's
contract goes GREEN; runs stays in project_ids order because gather resolves in
argument order, not completion order.

TWO PLAN CORRECTIONS, both found by the RED-first test rather than by reading:

1. The plan specified sorting the merged verdicts on `project_id`. Measured, that
   produces a deterministic order which is the WRONG one: lexicographic gives
   BRU/FV42/RV13 while the sequential pass gives FV42/RV13/BRU. It satisfies
   "deterministic" while breaking "identical to concurrency=1" — and the second is
   the actual contract. The merge preserves submission order instead.

2. The plan named the barrier's sort as the load-bearing seam. It is not — with
   per-project snapshots the wave list is never reordered by completion, so a
   sorted() there would re-sort an already-ordered list and read as a guard while
   guarding nothing. The SNAPSHOT is the half that carries the load. Rather than
   ship a decorative sort, both halves were measured (scratchpad-restore, never
   git checkout):

     detach _wave_snapshot  -> RED (store lands in completion order)
     detach merge ordering  -> RED (reversed wave order diverges)

   Both restored byte-identical (sha 42b01d46).

The snapshot carries `retriever` across deliberately: dropping it would silently
downgrade a caller-owned store's S3.1 semantic-retrieval opt-in mid-pass.

Also strengthens the scripted-client consolidation guard, which the probe broke by
being a legitimate third _inner_get_response def-site. It pinned a literal count
of 2 — the wrong shape: it failed on any new legitimate subclass while still
passing if someone pasted a duplicated body into an already-listed file. It now
pins the property (registered sites, scripted-lineage overrides must delegate via
super(), foreign-lineage doubles must genuinely be foreign). Verified load-bearing:
removing both delegation sites turns it RED.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQapztREtC2mkr5oU811pr
This commit is contained in:
Kjell Tore Guttormsen 2026-07-31 15:56:13 +02:00
commit 756b1d5b5c
3 changed files with 207 additions and 51 deletions

View file

@ -25,6 +25,7 @@ durable learned verdict captured out-of-band in the VerdictStore (D7-portable).
from __future__ import annotations
import asyncio
from collections.abc import Callable, Sequence
from dataclasses import dataclass, replace
from decimal import ROUND_HALF_UP, Decimal
@ -556,6 +557,54 @@ def _goal_limit_if_reached(goal: GoalContract, observed_ore: int, baseline_ore:
return None
def _wave_snapshot(store: VerdictStore) -> VerdictStore:
"""A per-project copy of the shared store's CURRENT verdicts (D-D wave model).
Two coroutines appending to one list would interleave by completion order, which no barrier
could then undo. Giving every project in a wave its own copy removes the race at its source
rather than serializing it away with a lock a lock would order the appends by whoever won,
which is exactly the nondeterminism being eliminated.
``retriever`` is carried across deliberately: it is the S3.1 opt-in seam, and a snapshot that
dropped it would silently downgrade a caller-owned store's semantic retrieval to the
structural default mid-pass."""
return VerdictStore(verdicts=list(store.verdicts), retriever=store.retriever)
def _merge_wave(store: VerdictStore, wave: Sequence[tuple[str, VerdictStore]]) -> None:
"""The deterministic merge barrier — the seam S3.3 rests on.
Each project ran against its own snapshot, so its NEW verdicts are the ones absent from the
wave-start store. They are merged back in **wave-submission order** the order the caller
listed the projects in so the shared store's contents depend on the portfolio's membership
and never on which project's model round-trips happened to finish first.
**Submission order, NOT lexicographic project_id.** The plan specified a sort on ``project_id``;
the Step-2 contract test measured that it produces a deterministic order which is nevertheless
the WRONG one. On the shipped fixture, lexicographic order is BRU/FV42/RV13 while the sequential
pass yields FV42/RV13/BRU, so a ``project_id`` sort satisfies "deterministic" while breaking
"identical to ``concurrency=1``" and the second is the actual contract. ``wave`` arrives in
submission order, so preserving it is the fix.
**Detach point: this function's ordering discipline is only half the seam — see
``_wave_snapshot``, which is the half that carries the load.** Because each project writes to
its own copy, ``wave`` is never reordered by completion, so iterating it in order is already
deterministic. Removing the SNAPSHOT is what turns the shared list back into a race and takes
``test_concurrent_pass_is_byte_identical_to_sequential`` RED. This is recorded plainly rather
than dressing the loop in a ``sorted(...)`` that would re-sort an already-ordered list and read
as a guard while guarding nothing.
Safe to apply as-is because ``VerdictStore.add`` is first-write-wins per content-hash id
(``verdicts.py:303-304``) the wave-start verdicts every snapshot carries are re-offered and
dropped, so only the new ones land."""
seen = {v.id for v in store.verdicts}
for _pid, snapshot in wave:
for verdict in snapshot.verdicts:
if verdict.id not in seen:
store.add(verdict)
seen.add(verdict.id)
def _waves(ids: list[str], k: int) -> list[list[str]]:
"""Partition ``ids`` into consecutive waves of at most ``k``, preserving caller order (D-D).
@ -626,55 +675,73 @@ async def run_portfolio(
runs: list[RunResult] = []
stopped_early = False
stop_reason: GoalReached | None = None
for pid in ids:
if pid not in projects:
raise ValueError(f"unknown project_id: {pid!r}")
project = projects[pid]
for wave_ids in _waves(ids, concurrency):
members: list[str] = []
for pid in wave_ids:
if pid not in projects:
raise ValueError(f"unknown project_id: {pid!r}")
project = projects[pid]
if goals.portfolio is not None:
observed = ledger.portfolio_total()
limit = _goal_limit_if_reached(goals.portfolio, observed, portfolio_baseline_ore)
if limit is not None:
if goals.portfolio.mode == "hard":
stopped_early = True
stop_reason = GoalReached("portfolio", None, limit, observed)
break
if stop_reason is None:
stop_reason = GoalReached("portfolio", None, limit, observed) # soft flag
if goals.portfolio is not None:
observed = ledger.portfolio_total()
limit = _goal_limit_if_reached(goals.portfolio, observed, portfolio_baseline_ore)
if limit is not None:
if goals.portfolio.mode == "hard":
stopped_early = True
stop_reason = GoalReached("portfolio", None, limit, observed)
break
if stop_reason is None:
stop_reason = GoalReached("portfolio", None, limit, observed) # soft flag
per_project_goal = goals.per_project.get(pid)
if per_project_goal is not None:
observed = ledger.per_project_total(pid)
limit = _goal_limit_if_reached(per_project_goal, observed, _to_ore(project.total_cost))
if limit is not None:
if stop_reason is None:
stop_reason = GoalReached("project", pid, limit, observed)
if per_project_goal.mode == "hard":
continue # skip THIS pid; the rest of the pass proceeds
per_project_goal = goals.per_project.get(pid)
if per_project_goal is not None:
observed = ledger.per_project_total(pid)
limit = _goal_limit_if_reached(
per_project_goal, observed, _to_ore(project.total_cost)
)
if limit is not None:
if stop_reason is None:
stop_reason = GoalReached("project", pid, limit, observed)
if per_project_goal.mode == "hard":
continue # skip THIS pid; the rest of the pass proceeds
members.append(pid)
# Every project in the wave reads the SAME wave-start state and writes only its own copy,
# so no two coroutines touch one list. The snapshot is what makes the barrier sufficient.
snapshots = [(pid, _wave_snapshot(store)) for pid in members]
# run_portfolio only drives full runs (never dry-run), so the return narrows to RunResult;
# the cast keeps the widened run_project signature honest without an @overload duplication.
result = cast(
RunResult,
await run_project(
pid,
profile,
docs_dir=project.docs_dir,
verdict_input=project.verdict_input,
bundle_dir=project.bundle_dir,
verdict_dir=project.verdict_dir,
dimension=dimension,
store=store,
client_factory=client_factory,
max_rounds=max_rounds,
max_tokens=max_tokens,
top_k=top_k,
semantic_retrieval=semantic_retrieval,
embedder=embedder,
meter=meter_factory() if meter_factory is not None else None,
),
wave_results = await asyncio.gather(
*(
run_project(
pid,
profile,
docs_dir=projects[pid].docs_dir,
verdict_input=projects[pid].verdict_input,
bundle_dir=projects[pid].bundle_dir,
verdict_dir=projects[pid].verdict_dir,
dimension=dimension,
store=snapshot,
client_factory=client_factory,
max_rounds=max_rounds,
max_tokens=max_tokens,
top_k=top_k,
semantic_retrieval=semantic_retrieval,
embedder=embedder,
meter=meter_factory() if meter_factory is not None else None,
)
for pid, snapshot in snapshots
)
)
runs.append(result)
# ``gather`` resolves in ARGUMENT order, not completion order, and waves follow
# ``project_ids`` — so ``runs`` stays in caller order however the schedule interleaved.
runs.extend(cast(RunResult, r) for r in wave_results)
_merge_wave(store, snapshots)
if stopped_early:
break
base = _aggregate(tuple(runs), store)
if stopped_early or stop_reason is not None: