DEL C. P18 gave read_dir a window (filter/offset/limit) and then measured its
own paid round without being able to see it used: five of 31 documents read
lay outside the default window, so the window HAD been widened and the trace
could not say with which knob. ToolCall now carries the three arguments,
always present and empty/zero when not passed -- an absent key and "not
narrowed" must not read the same -- and the judge counts filter_calls and
paged_calls. _number_argument is a SIBLING of _string_argument, not a widening
of it: a model may send limit as 10 or as "10", and a reader that knew one
shape would report a paged call as unpaged.
DEL D. P18's finding 4 was WRONG AS WRITTEN. provenance.token_usage has been
stamped on every proposal artefact since S3.4 and stands in every one of round
2's; what was missing is a READER. The judge reads it now (round 2 measured:
289 054 tokens against round 1's 2 679 305, -89 %), and the P18 report gets a
dated correction UNDER its original paragraph rather than instead of it.
What was genuinely absent is {run_id}-coverage.json. settle prints the
coverage report and ApproachOutcome has carried not_evaluated since Trekk A3,
but neither ever reached a file, so a judge could see an approach had no
artefact and could not tell a budget stop from an approach nobody ordered.
Written from the finally IFF a mandate was given. stop_reason comes from a
CALLER-OWNED sink rather than from in_flight, and that is a measurement:
_evaluate_mandate SWALLOWS BudgetExceeded once something has been produced, so
run_project's own in_flight never sees it.
Load-bearing measured (10 arms), four mutations all red against the whole
suite, green control 1744/5 and the golden byte-unchanged. D-i stood GREEN
first -- the vacuous-gate class, 25th time: the arm called write_coverage
itself and therefore chose the reason it then asserted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
410 lines
16 KiB
Python
410 lines
16 KiB
Python
"""P19 DEL C + DEL D — the trace says HOW, and a run says what it spent and why it stopped.
|
|
|
|
**DEL C, the measured silence.** P18 gave ``read_dir`` a window — ``filter`` / ``offset`` /
|
|
``limit`` — and then measured its own paid round without being able to see it used: five of 31
|
|
documents read lay OUTSIDE the default window, so the window had been widened, and the trace could
|
|
not say with which knob (``docs/2026-09-14-p18-stressrunde-2.md`` § 1, finding 1). The recorder
|
|
kept ``name`` / ``bundle_id`` / ``path`` and nothing else. "Did the model narrow the level, or page
|
|
through it?" is the operative question about a corpus of 2 756 documents, and it was unanswerable
|
|
from the artefact the run leaves behind.
|
|
|
|
**DEL D, and one of the two findings it fixes was WRONG AS WRITTEN.** P18's finding 4 said a
|
|
successful run does not say what it used. Measured 14.09: ``provenance.token_usage`` has been
|
|
stamped on every proposal artefact since S3.4 and stands in every one of round 2's. What was
|
|
missing is a JUDGE that reads it — so the report's sentence is corrected in its own § 7 addendum
|
|
rather than the field being invented a second time. What was genuinely absent is the OTHER half:
|
|
``settle`` prints the coverage report, ``ApproachOutcome`` has carried ``not_evaluated`` since
|
|
Trekk A3, and neither ever reached a file. A judge reading an outbox could see that an approach had
|
|
no artefact and could not tell a budget stop from an approach nobody ordered — the silence
|
|
``ApproachOutcome`` exists to remove, one layer out.
|
|
|
|
What each arm pins:
|
|
|
|
(a) C — the three window arguments are RECORDED, in the shape the wire actually carries;
|
|
(b) C — a call that passes none of them yields empty/zero fields, never a refusal and never an
|
|
absent key: "not narrowed" and "we did not look" must not read the same;
|
|
(c) C — a numeric argument sent as a STRING is coerced. A model may send either, and a recorder
|
|
that read only one shape would report a paged call as unpaged;
|
|
(d) C — the judge counts them, so "did the model use the filter" is readable directly rather than
|
|
inferred from which documents happened to lie outside a default window;
|
|
(e) D2 — ``{run_id}-coverage.json`` is written IFF a mandate was given, INCLUDING when a cap cut
|
|
the run short before a single approach was evaluated. A mandate-less run leaves the outbox
|
|
byte-identical, which two older tests pin as an exact listing;
|
|
(f) D2 — the stop reason is ``BudgetExceeded.kind``, carried as a field rather than inferred, and
|
|
the judge turns it into each unevaluated row's ``not_evaluated_reason``;
|
|
(g) D1 — the judge reads ``provenance.token_usage`` and reports it, with ``absent`` distinguishing
|
|
a run that wrote no coverage file (every run before today) from one that finished cleanly.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
from pathlib import Path
|
|
from typing import Any
|
|
|
|
import pytest
|
|
from agent_framework import BaseChatClient
|
|
from conftest import SyntheticUsageChatClient
|
|
|
|
from portfolio_optimiser import explore as ex
|
|
from portfolio_optimiser import outbox, stress
|
|
from portfolio_optimiser.budget import Budget, TokenMeter
|
|
from portfolio_optimiser.mandate import Approach, Mandate
|
|
from portfolio_optimiser.run import run_project
|
|
|
|
_BUNDLE_DIR = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
|
|
_PID = "BYGG-KONTOR-NORD"
|
|
_PROPOSER = (
|
|
'{"measure":"LED-retrofit av kontorbelysning","affected_items":'
|
|
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],"claimed_saving_nok":30000}'
|
|
)
|
|
_CHECKER = "Supported by the cited documents. VERDICT: APPROVE"
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _isolate_model_env(monkeypatch: pytest.MonkeyPatch) -> None:
|
|
monkeypatch.delenv("PORTFOLIO_MODEL_MAP", raising=False)
|
|
monkeypatch.delenv("PORTFOLIO_FOUNDRY_PROJECT_ENDPOINT", raising=False)
|
|
|
|
|
|
def _factory(role: str) -> BaseChatClient:
|
|
return SyntheticUsageChatClient(default_reply=_CHECKER if role == "checker" else _PROPOSER)
|
|
|
|
|
|
class _Fn:
|
|
def __init__(self, name: str) -> None:
|
|
self.name = name
|
|
|
|
|
|
class _Ctx:
|
|
def __init__(self, name: str, arguments: Any) -> None:
|
|
self.function = _Fn(name)
|
|
self.arguments = arguments
|
|
|
|
|
|
def _record(*calls: tuple[str, Any]) -> list[ex.ToolCall]:
|
|
import asyncio
|
|
|
|
sink: list[ex.ToolCall] = []
|
|
recorder = ex.ExplorationToolRecorder(sink)
|
|
|
|
async def _noop() -> None:
|
|
return None
|
|
|
|
for name, args in calls:
|
|
asyncio.run(recorder.process(_Ctx(name, args), _noop)) # type: ignore[arg-type]
|
|
return sink
|
|
|
|
|
|
# ---------------------------------------------------------------------------------------------
|
|
# (a)-(d) DEL C
|
|
# ---------------------------------------------------------------------------------------------
|
|
|
|
|
|
def test_the_window_arguments_are_recorded() -> None:
|
|
"""(a) HOW the level was asked for, not only which one."""
|
|
(call,) = _record(
|
|
("read_dir", {"bundle_id": "k2", "path": "krav/N100", "filter": "rundkjoring", "limit": 25})
|
|
)
|
|
assert (call.filter, call.offset, call.limit) == ("rundkjoring", 0, 25)
|
|
|
|
|
|
def test_a_call_that_passes_none_of_them_yields_empty_fields() -> None:
|
|
"""(b) Never a refusal, never an absent key — the ``path``/``bundle_id`` rule one field over."""
|
|
(call,) = _record(("list_bundles", {}))
|
|
assert (call.filter, call.offset, call.limit) == ("", 0, 0)
|
|
payload = ex.tool_call_payload([call])
|
|
assert payload[0]["filter"] == "" and payload[0]["offset"] == 0 and payload[0]["limit"] == 0
|
|
|
|
|
|
def test_a_numeric_argument_sent_as_a_string_is_coerced() -> None:
|
|
"""(c) A model may send either shape; a recorder that read one would misreport the other."""
|
|
(call,) = _record(("read_dir", {"bundle_id": "k2", "path": "p", "offset": "30", "limit": "10"}))
|
|
assert (call.offset, call.limit) == (30, 10)
|
|
# …and a value that is neither is 0, never an invented number.
|
|
(other,) = _record(("read_dir", {"bundle_id": "k2", "path": "p", "limit": True}))
|
|
assert other.limit == 0
|
|
|
|
|
|
# ---------------------------------------------------------------------------------------------
|
|
# a synthetic base + outbox, self-contained (nothing here touches a real bundle or a model)
|
|
# ---------------------------------------------------------------------------------------------
|
|
|
|
_GOOD = "krav/N1/id-good.md"
|
|
|
|
|
|
def _minibase(root: Path) -> Path:
|
|
base = root / "minibase"
|
|
(base / "krav" / "N1").mkdir(parents=True)
|
|
(base / "index.md").write_text(
|
|
"---\nbundle_id: minibase\n---\n\n- [good](krav/N1/id-good.md)\n", encoding="utf-8"
|
|
)
|
|
(base / _GOOD).write_text(
|
|
'---\ntype: concept\ntitle: "T"\nreq_number: "Krav 1.2.3-4"\n---\n\nBody.\n',
|
|
encoding="utf-8",
|
|
)
|
|
return base
|
|
|
|
|
|
def _context_dir(root: Path) -> Path:
|
|
ctx = root / "ctx"
|
|
(ctx / "docs").mkdir(parents=True)
|
|
(ctx / "bundle.txt").write_text("name: minibase\nbundle_id: minibase\n", encoding="utf-8")
|
|
(ctx / "mandate.json").write_text(
|
|
json.dumps(
|
|
{
|
|
"objective": "o",
|
|
"success_criteria": "s",
|
|
"approaches": [
|
|
{"id": "a1", "label": "L", "affected_codes": ["CODE-1"]},
|
|
{"id": "a2", "label": "M", "affected_codes": ["CODE-2"]},
|
|
],
|
|
}
|
|
),
|
|
encoding="utf-8",
|
|
)
|
|
(ctx / "fasit.json").write_text(
|
|
json.dumps(
|
|
{
|
|
"project_id": "proj",
|
|
"bundle": "minibase",
|
|
"bundle_id": "minibase",
|
|
"must_cite": [
|
|
{
|
|
"approach_id": "a1",
|
|
"rationale": "why",
|
|
"concepts": [{"path": _GOOD, "title": "T", "ref": "Krav 1.2.3-4"}],
|
|
}
|
|
],
|
|
"must_refuse": [],
|
|
"honesty": "synthetic",
|
|
}
|
|
),
|
|
encoding="utf-8",
|
|
)
|
|
return ctx
|
|
|
|
|
|
def _outbox(root: Path, *, tokens: int, tool_calls: list[dict[str, Any]]) -> Path:
|
|
out = root / "out"
|
|
out.mkdir(parents=True, exist_ok=True)
|
|
(out / "r1-a1-proposal.json").write_text(
|
|
json.dumps(
|
|
{
|
|
"proposal": {
|
|
"project_id": "proj",
|
|
"measure": "m",
|
|
"affected_items": [{"code": "CODE-1", "quantity": 1.0, "unit_cost": 2.0}],
|
|
"claimed_saving_nok": 1.0,
|
|
"assumptions": {},
|
|
},
|
|
"provenance": {"citations": [], "token_usage": tokens},
|
|
}
|
|
),
|
|
encoding="utf-8",
|
|
)
|
|
(out / "r1-a1-outcome.json").write_text(
|
|
json.dumps({"outcome_type": "rejected", "reason": "no"}), encoding="utf-8"
|
|
)
|
|
(out / "r1-debate.json").write_text(
|
|
json.dumps({"run_id": "r1", "tool_calls": tool_calls}), encoding="utf-8"
|
|
)
|
|
return out
|
|
|
|
|
|
def test_the_judge_counts_filtered_and_paged_calls(tmp_path: Path) -> None:
|
|
"""(d) Readable directly, instead of inferred from which documents fell outside a default."""
|
|
calls = [
|
|
{
|
|
"name": "read_dir",
|
|
"bundle_id": "b",
|
|
"path": "p",
|
|
"filter": "rund",
|
|
"offset": 0,
|
|
"limit": 0,
|
|
},
|
|
{"name": "read_dir", "bundle_id": "b", "path": "p", "filter": "", "offset": 30, "limit": 0},
|
|
{
|
|
"name": "read_file",
|
|
"bundle_id": "b",
|
|
"path": _GOOD,
|
|
"filter": "",
|
|
"offset": 0,
|
|
"limit": 0,
|
|
},
|
|
]
|
|
verdict = stress.score_context_set(
|
|
_context_dir(tmp_path),
|
|
_outbox(tmp_path, tokens=1234, tool_calls=calls),
|
|
"r1",
|
|
_minibase(tmp_path),
|
|
)
|
|
assert (verdict.filter_calls, verdict.paged_calls) == (1, 1)
|
|
assert verdict.tool_calls_seen == 3
|
|
|
|
|
|
# ---------------------------------------------------------------------------------------------
|
|
# (e)-(g) DEL D
|
|
# ---------------------------------------------------------------------------------------------
|
|
|
|
|
|
def test_the_coverage_artefact_carries_the_rows_and_the_stop_reason(tmp_path: Path) -> None:
|
|
"""(e)/(f) A required stop reason, never an inferred one: "finished" and "we never found out"
|
|
must not be the same value."""
|
|
outbox.write_coverage(
|
|
str(tmp_path),
|
|
"r1",
|
|
rows=[
|
|
{"id": "a1", "label": "L", "status": "validated", "detail": "", "saving_nok": 10.0},
|
|
{
|
|
"id": "a2",
|
|
"label": "M",
|
|
"status": "not_evaluated",
|
|
"detail": "budget",
|
|
"saving_nok": None,
|
|
},
|
|
],
|
|
stop_reason="tokens",
|
|
)
|
|
payload = json.loads((tmp_path / "r1-coverage.json").read_text(encoding="utf-8"))
|
|
assert payload["stop_reason"] == "tokens"
|
|
assert [r["status"] for r in payload["rows"]] == ["validated", "not_evaluated"]
|
|
# Byte-deterministic: the same inputs twice produce the same file (write_run_config's rule).
|
|
first = (tmp_path / "r1-coverage.json").read_bytes()
|
|
outbox.write_coverage(
|
|
str(tmp_path),
|
|
"r1",
|
|
rows=[
|
|
{"id": "a1", "label": "L", "status": "validated", "detail": "", "saving_nok": 10.0},
|
|
{
|
|
"id": "a2",
|
|
"label": "M",
|
|
"status": "not_evaluated",
|
|
"detail": "budget",
|
|
"saving_nok": None,
|
|
},
|
|
],
|
|
stop_reason="tokens",
|
|
)
|
|
assert (tmp_path / "r1-coverage.json").read_bytes() == first
|
|
|
|
|
|
def test_the_judge_reads_the_stop_reason_onto_every_unevaluated_row(tmp_path: Path) -> None:
|
|
"""(f) The second half: a row with no artefact says WHY, instead of leaving a reader to guess."""
|
|
ctx = _context_dir(tmp_path)
|
|
out = _outbox(tmp_path, tokens=45_642, tool_calls=[])
|
|
base = _minibase(tmp_path)
|
|
|
|
# Before any coverage file exists — every run written before today.
|
|
blind = stress.score_context_set(ctx, out, "r1", base)
|
|
assert blind.stop_reason == "absent"
|
|
assert [r.not_evaluated_reason for r in blind.approaches] == ["", "absent"]
|
|
|
|
outbox.write_coverage(
|
|
str(out),
|
|
"r1",
|
|
rows=[
|
|
{"id": "a2", "label": "M", "status": "not_evaluated", "detail": "b", "saving_nok": None}
|
|
],
|
|
stop_reason="rounds",
|
|
)
|
|
informed = stress.score_context_set(ctx, out, "r1", base)
|
|
assert informed.stop_reason == "rounds"
|
|
assert [r.not_evaluated_reason for r in informed.approaches] == ["", "rounds"]
|
|
|
|
|
|
def test_the_judge_reports_what_the_run_spent(tmp_path: Path) -> None:
|
|
"""(g) D1. P18's finding 4 was wrong as written: the field was always there, the READER was not.
|
|
|
|
The figure is the one measured on round 2's own artefacts (tunnel-hauglia-2027-02).
|
|
"""
|
|
verdict = stress.score_context_set(
|
|
_context_dir(tmp_path),
|
|
_outbox(tmp_path, tokens=45_642, tool_calls=[]),
|
|
"r1",
|
|
_minibase(tmp_path),
|
|
)
|
|
assert verdict.token_usage == 45_642
|
|
|
|
|
|
# ---------------------------------------------------------------------------------------------
|
|
# (e) the run itself writes it -- without this arm the wiring has no witness
|
|
# ---------------------------------------------------------------------------------------------
|
|
|
|
|
|
async def test_a_commissioned_run_writes_the_coverage_artefact(tmp_path: Path) -> None:
|
|
"""(e) END-TO-END. The arm above drives ``write_coverage`` directly and would stay green with
|
|
the ``run_project`` wiring detached -- the vacuous-gate shape this repo has now met two dozen
|
|
times. This one runs the real pipeline."""
|
|
out = tmp_path / "with-mandate"
|
|
await run_project(
|
|
_PID,
|
|
"local",
|
|
docs_dir=str(_BUNDLE_DIR),
|
|
bundle_dir=str(_BUNDLE_DIR),
|
|
client_factory=_factory,
|
|
max_rounds=2,
|
|
outbox_dir=str(out),
|
|
run_id="r1",
|
|
mandate=Mandate(
|
|
objective="o",
|
|
approaches=(Approach(id="a1", label="LED", description="swap the fittings"),),
|
|
allow_own_proposals=False,
|
|
),
|
|
)
|
|
payload = json.loads((out / "r1-coverage.json").read_text(encoding="utf-8"))
|
|
assert payload["stop_reason"] == "", "nothing cut this run short"
|
|
assert [r["id"] for r in payload["rows"]] == ["a1"]
|
|
assert payload["rows"][0]["status"] in {"validated", "rejected"}
|
|
|
|
|
|
async def test_a_mandateless_run_leaves_the_outbox_byte_identical(tmp_path: Path) -> None:
|
|
"""(e), the other half. Coverage is the MANDATE's report; a run without one would be described
|
|
by a file saying nothing, and two older tests pin an exact listing on such a run."""
|
|
out = tmp_path / "no-mandate"
|
|
await run_project(
|
|
_PID,
|
|
"local",
|
|
docs_dir=str(_BUNDLE_DIR),
|
|
bundle_dir=str(_BUNDLE_DIR),
|
|
client_factory=_factory,
|
|
max_rounds=2,
|
|
outbox_dir=str(out),
|
|
run_id="r1",
|
|
)
|
|
assert not (out / "r1-coverage.json").exists()
|
|
|
|
|
|
async def test_a_cap_that_cuts_the_run_short_is_named_in_the_artefact(tmp_path: Path) -> None:
|
|
"""(f) END-TO-END, and this arm exists because the direct one was VACUOUS.
|
|
|
|
MEASURED: mutation D-i — ``stop_reason=""`` whatever happened — left the WHOLE suite green
|
|
(1743/5), because the arm above calls ``write_coverage`` itself and therefore chooses the
|
|
reason it then asserts. Only a run that a cap actually cut short can tell the two apart, and
|
|
it is exactly the run for which the artefact matters: a commission that stopped halfway.
|
|
"""
|
|
out = tmp_path / "capped"
|
|
await run_project(
|
|
_PID,
|
|
"local",
|
|
docs_dir=str(_BUNDLE_DIR),
|
|
bundle_dir=str(_BUNDLE_DIR),
|
|
client_factory=_factory,
|
|
max_rounds=2,
|
|
outbox_dir=str(out),
|
|
run_id="r1",
|
|
# tick_round fires once per generation ATTEMPT: one round pays for the first
|
|
# approach, the second is the one the cap cuts.
|
|
meter=TokenMeter(Budget(max_tokens=10_000_000, max_rounds=1)),
|
|
mandate=Mandate(
|
|
objective="o",
|
|
approaches=(
|
|
Approach(id="a1", label="LED", description="one"),
|
|
Approach(id="a2", label="Sensors", description="two"),
|
|
),
|
|
allow_own_proposals=False,
|
|
),
|
|
)
|
|
payload = json.loads((out / "r1-coverage.json").read_text(encoding="utf-8"))
|
|
assert payload["stop_reason"] == "rounds", payload
|
|
assert [r["status"] for r in payload["rows"]][-1] == "not_evaluated", payload
|