portfolio-optimiser/tests/test_debate_navigation_cost_loadbearing.py
Kjell Tore Guttormsen eb4137415c feat(prepass,explore): padding dies at the prompt; a refusal the model can act on is a return value [skip-docs]
Order 20260908T195801Z. Findings 4 and 5 from the S7 acid test, then the two things
finding 99 measured and deliberately did not fix (D3, D2).

No user-facing surface changes: no new flag, no new command, no changed output
contract. Both seams are internal (the pre-pass rendering, and the shape a tool
answers a model with), so [skip-docs] rather than a README edit that would describe
nothing an operator can do differently.

FINDING 4 -- MEASURED, NOTHING BUILT. K2's price schedule IS readable without
guessing (8 column spans, 71 of 91 non-blank rows give >= 2 cells, the split stable
for K = 2..64). But 0 of 92 rows name all three of code/quantity/unit_cost -- also
under a looser substring match -- and 0 of 91 data rows carry code + quantity +
amount. The triple is not formatted away; it is not in the document. It is a price
SUMMARY plus nine rate cards whose unit-price columns are empty (pre-award). The
order's binding decision rule therefore falls against building:
--derive-cost-baseline keeps refusing, and MAJOR-4's own honesty limit holds.

FINDING 5 -- BUILT. Measured on the actual rendering path (concept_text, not the
raw file): the delivered excerpt is 104 lines / 67 245 chars, carrying 208 interior
whitespace runs, 117 of them >= 100 and the longest 887 -- 56 806 of 67 245
characters = 84.5 %, over 72 of 104 lines. collapse_padding, called from
_data_blocks (the one renderer both arms share, and therefore AFTER
verify_against_bundle -- collapsing in concept_text would break every payload's own
digest), gives -72.4 %: line count invariant, non-whitespace byte-identical, leading
indentation untouched, no number changed.

F99-D3 -- read_file / read_dir / read_bundle now RETURN their refusal. MAF turns a
tool raise into "Error: Function failed." (_tools.py:1410-1432, :1427) and counts it
against DEFAULT_MAX_CONSECUTIVE_ERRORS_PER_REQUEST = 3, so everything the refusing
arm knows is destroyed on the way out. The gates are unchanged; the property they
exist for -- the reason travels, the bytes never do -- is now asserted explicitly on
the returned value. The arm is keyed on named classes, never bare Exception, because
ExplorationError is itself a RuntimeError subclass.

F99-D2 -- the invariant row, plus one for finding 5 (a stated deviation from "one
row only": finding 5 is a separately built seam and the ledger's standing rule
requires its own row).

19 existing arms rewritten, never deleted and never weakened: where the class
carried a distinction, the refusal KIND carries it now.

13 mutations, all red against the whole suite (W1-W5, M1-M8), each restored from
scratchpad with shasum -c. Control 1543 passed / 5 skipped (from 1529/5, a strict
superset, 0 removed). Golden demo-transcript.stdout unchanged
(shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Measurement: docs/2026-09-08-funn-4-5-og-read-nekt.md

Co-Authored-By: Claude <Opus 5>
2026-09-08 23:36:54 +02:00

387 lines
18 KiB
Python

"""S2c LOAD-BEARING: the DEBATE navigates the knowledge base; it is never handed the whole of it.
MAJOR-3 and S7a-3 made the EXPLORATION cheap — 18 355 o200k tokens over the whole of K2 — and
left the pipeline's own two phases stuffing. Measured on K2 (630 concepts,
``docs/2026-09-04-syretest-s7b-k2.md`` § 3.4, re-measured with the same instrument before this
change): ``okf.bundle_context`` is **648 962 o200k tokens** and rides in **three** prompts (two
proposer turns + the checker's), so debate + generation is **1 947 342 tokens = 99,1 %** of a
run's prompt cost. Ninety-nine percent of what a run pays for is context nobody asked for twice.
The move is S7a-3's own, one seam over: ``run_project``'s bundle path stops rendering the base
into the task message and instead hands the debate the SAME four navigator tools the exploration
uses (``navigator_tools``) plus a POINTER naming the base. Generation is covered by the same
change: ``gen_context = debate_output or context``, so the last-resort fallback is bounded by
construction rather than being the whole corpus.
Arms, each with a named detach point:
* **(a) no prompt carries the base.** A sentinel that lives only in a concept body is ABSENT from
every prompt the debate and the generation call see. CONTROL: the same sentinel IS in
``bundle_context``, so its absence is caused by the seam and not by an empty fixture.
* **(b) bounded, and not vacuously so.** Every prompt is under a ceiling that lives HERE, never in
``run.py`` (the ``read_bundle``/catalogue rule: raising the constant is the regression the gate
exists to catch), while ``bundle_context`` for the same base is over FIVE times it — without the
flat control a green bound could just mean the fixture is small. And the pointer must NAME the
base, because a bounded prompt that omits the id is a debate that cannot call the tools at all.
* **(c) the tools are actually there.** The four navigator tools reach the built workflow.
* **(d) the tool trace is a first-class artefact.** ``RunResult.debate_tool_calls`` and
``{run_id}-debate.json`` carry name + bundle_id + path in CALL ORDER — S7a-3 pkt. 3's rule, on
the second surface that now opens a base: over 629 concepts a trace reading ``read_file`` twice
answers nothing about which two.
* **(e) the dimension scope survives the move.** §4.1a promised the agents read ONLY
dimension-matched bundle knowledge. That promise used to be kept by ``bundle_context``'s filter;
with navigation it has to be kept by the TOOLS, on BOTH rungs — a listing that hides a foreign
document while ``read_file`` still serves it is a filter in name only.
"""
from __future__ import annotations
import json
import shutil
from collections.abc import Callable
from pathlib import Path
from typing import Any
from agent_framework import BaseChatClient
from portfolio_optimiser import okf
from portfolio_optimiser.dimension import Dimension
from portfolio_optimiser.explore import DimensionScopeRefused, navigator_tools
from portfolio_optimiser.run import run_project
from portfolio_optimiser.simulation import ScriptedChatClient, scripted_factory
BUNDLE_DIR = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
_PID = "BYGG-KONTOR-NORD"
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (sim)"}
#: The prompt budget one debate/generation call may spend on bundle context, in CHARACTERS. It
#: lives in the TEST for ``_CATALOGUE_EXCERPT_CHARS``' reason: a ceiling imported from the
#: implementation moves with it, and raising it is precisely the regression this gate catches.
#: Characters rather than o200k tokens because ``tiktoken`` is not a project dependency, and a gate
#: that skips when an optional package is missing is a gate that can be silently absent
#: (MAJOR-3's own stated deviation, same reason).
_CEILING_CHARS = 1_500
_VALID_REPLY = (
'{"measure":"LED-retrofit","affected_items":'
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],'
'"claimed_saving_nok":30000}'
)
_CHECKER_REPLY = "Reasoning holds.\nVERDICT: APPROVE"
def _prompt_blob(messages: Any) -> str:
"""Everything a prompt actually carries: text PLUS ``function_call``/``function_result``
contents. ``Message.text`` alone measures a context-bearing prompt at a few characters — the
corrected S7a-2 instrument, and the reason S2c could be measured at all."""
parts: list[str] = []
for message in messages:
text = getattr(message, "text", "") or ""
if text:
parts.append(text)
for content in getattr(message, "contents", ()) or ():
for attr in ("result", "arguments"):
value = getattr(content, attr, None)
if value:
parts.append(str(value))
return "\n".join(parts)
def _recording_factory(
sink: list[str], *, script: dict[str, Any] | None = None
) -> Callable[[str], BaseChatClient]:
"""A scripted client factory whose every prompt lands in ``sink`` as the FULL blob."""
def factory(role: str) -> BaseChatClient:
if script is not None and role in script:
client: BaseChatClient = scripted_factory(script, [])(role)
else:
client = ScriptedChatClient(
_CHECKER_REPLY if role == "checker" else _VALID_REPLY, role=role
)
original = client._inner_get_response # type: ignore[attr-defined]
def recording(*, messages, options, stream=False, **kwargs): # type: ignore[no-untyped-def]
sink.append(_prompt_blob(messages))
return original(messages=messages, options=options, stream=stream, **kwargs)
client._inner_get_response = recording # type: ignore[attr-defined,method-assign]
return client
return factory
async def _run(**kwargs: Any) -> tuple[Any, list[str]]:
sink: list[str] = []
script = kwargs.pop("script", None)
result = await run_project(
_PID,
"local",
docs_dir=str(BUNDLE_DIR),
bundle_dir=str(BUNDLE_DIR),
verdict_input=_VERDICT_INPUT,
client_factory=_recording_factory(sink, script=script),
**kwargs,
)
return result, sink
def _sentinel_from_the_base() -> str:
"""A string that exists ONLY inside a concept body of the fixture base — the leak probe."""
bundle = okf.navigate_bundle(str(BUNDLE_DIR))
reference = next(f for f in bundle.context_files if f.name == "kilder-realiseringsgap.md")
line = next(row for row in reference.body.splitlines() if len(row.strip()) > 60)
return line.strip()
# ---------------------------------------------------------------- (a) nothing carries the base
async def test_no_debate_or_generation_prompt_carries_the_whole_base() -> None:
"""LOAD-BEARING (a): a sentinel living only in a concept body reaches NO prompt.
Detach point: restore ``context = okf.bundle_context(bundle, ...)`` in ``run_project``'s
bundle arm → RED (the sentinel is back in all three prompts)."""
sentinel = _sentinel_from_the_base()
_, sink = await _run()
assert sink, "the debate never ran — no prompt was captured"
leaking = [i for i, prompt in enumerate(sink) if sentinel in prompt]
assert not leaking, (
f"the whole knowledge base is still being stuffed into prompt(s) {leaking}: "
"the debate is handed the corpus instead of navigating it"
)
def test_the_sentinel_is_really_in_the_base() -> None:
"""CAUSALITY CONTROL for (a): the sentinel IS what ``bundle_context`` renders, so its absence
above is caused by the seam rather than by a fixture that never held it."""
sentinel = _sentinel_from_the_base()
assert sentinel in okf.bundle_context(okf.navigate_bundle(str(BUNDLE_DIR)))
# ------------------------------------------------------------------- (b) bounded, not vacuous
async def test_every_prompt_stays_under_the_ceiling() -> None:
"""LOAD-BEARING (b): the debate + generation prompts are bounded by the POINTER, so their cost
follows the number of bases configured (which the operator chose) and not the size of the
corpus (which they did not)."""
_, sink = await _run()
worst = max(len(prompt) for prompt in sink)
assert worst <= _CEILING_CHARS, (
f"the widest debate/generation prompt is {worst} characters, over the {_CEILING_CHARS} "
"ceiling — the corpus is riding along again"
)
def test_the_flat_form_is_far_over_the_ceiling() -> None:
"""FLAT CONTROL for (b): ``bundle_context`` on the SAME base is over five times the ceiling, so
a green bound above cannot just mean the fixture is small."""
whole = okf.bundle_context(okf.navigate_bundle(str(BUNDLE_DIR)))
assert len(whole) > 5 * _CEILING_CHARS, (
"the fixture base is too small for the bound to prove anything"
)
async def test_the_pointer_names_the_base_the_tools_take() -> None:
"""LOAD-BEARING (b, second half): a bounded prompt that does not NAME the base is a debate
that cannot call a single tool — bounded and useless is the vacuous form of this gate.
Detach point: drop the bundle id from the pointer → RED."""
_, sink = await _run()
bundle_id = okf.reconcile_bundle_id(str(BUNDLE_DIR)).id
assert any(bundle_id in prompt for prompt in sink), (
f"no prompt names the knowledge base {bundle_id!r}; the navigator tools take that id, so "
"the debate has been given a bounded prompt it cannot act on"
)
# ------------------------------------------------------------------------- (c) the tools exist
async def test_the_debate_is_given_the_navigator_tools(monkeypatch) -> None:
"""LOAD-BEARING (c): the bundle path hands the debate the SAME four tools the exploration uses.
Detach point: drop ``navigator_tools(...)`` from ``debate_tools`` → RED."""
import portfolio_optimiser.run as run_module
captured: list[list[Any]] = []
original = run_module.fresh_workflow
def spy(*args: Any, **kwargs: Any) -> Any:
captured.append(list(kwargs.get("tools") or []))
return original(*args, **kwargs)
monkeypatch.setattr(run_module, "fresh_workflow", spy)
await _run()
assert captured, "the debate was never built"
names = {getattr(t, "name", "") for t in captured[0]}
assert {"list_bundles", "read_bundle", "read_dir", "read_file"} <= names, (
f"the debate's tool list is {sorted(names)} — it cannot navigate the base it was pointed at"
)
# ------------------------------------------------------------------------- (d) the tool trace
async def test_the_debate_tool_trace_reaches_the_result_and_the_outbox(tmp_path) -> None:
"""LOAD-BEARING (d): what the debate OPENED is recorded, in call order, with the path.
The proposer is driven by a step MANUSCRIPT (MAJOR-1 b) — the only offline form that can make a
scripted role emit a ``function_call``, and therefore the only way a free rehearsal can prove
the debate opens anything at all.
Detach points, each RED on its own: drop the ``ExplorationToolRecorder`` from the debate's
middleware; stop writing ``{run_id}-debate.json``; drop ``path`` from the payload."""
bundle_id = okf.reconcile_bundle_id(str(BUNDLE_DIR)).id
outbox_dir = tmp_path / "outbox"
result, _ = await _run(
outbox_dir=str(outbox_dir),
run_id="s2c",
script={
"proposer": [
{"call": "read_bundle", "args": {"bundle_id": bundle_id}},
{
"call": "read_file",
"args": {"bundle_id": bundle_id, "path": "metode-ipmvp-a.md"},
},
_VALID_REPLY,
],
"checker": _CHECKER_REPLY,
},
)
observed = [(c.name, c.bundle_id, c.path) for c in result.debate_tool_calls]
assert observed[:2] == [
("read_bundle", bundle_id, ""),
("read_file", bundle_id, "metode-ipmvp-a.md"),
], f"the debate's tool trace is {observed} — the call sequence is not recorded as it happened"
payload = json.loads((outbox_dir / "s2c-debate.json").read_text(encoding="utf-8"))
assert payload["tool_calls"][:2] == [
{"name": "read_bundle", "bundle_id": bundle_id, "path": ""},
{"name": "read_file", "bundle_id": bundle_id, "path": "metode-ipmvp-a.md"},
], f"the artefact does not carry the call sequence: {payload['tool_calls']}"
async def test_a_debate_that_opened_nothing_says_so(tmp_path) -> None:
"""CONTROL for (d): a run whose agents called no tool leaves an EMPTY trace rather than no
artefact — "the debate never opened the base" is the S2c regression signal itself, so it must
be readable off the leaving, not inferred from a missing file."""
outbox_dir = tmp_path / "outbox"
result, _ = await _run(outbox_dir=str(outbox_dir), run_id="s2c-quiet")
assert result.debate_tool_calls == ()
payload = json.loads((outbox_dir / "s2c-quiet-debate.json").read_text(encoding="utf-8"))
assert payload["tool_calls"] == []
# ------------------------------------------------------------------ (e) the dimension survives
_ENERGY_DIM = Dimension(
id="energi", label="Energi", allowed_measure_types=frozenset({"energy_efficiency"})
)
_ASFALT_SENTINEL = "ASFALT-LEAK-SENTINEL-x7y8z9"
_ASFALT_FILE = "asfalt-dekke.md"
def _bundle_with_a_foreign_dimension(tmp_path: Path) -> str:
"""A copy of the fixture base plus ONE concept file marked ``dimension: asfalt``, linked from
the index so navigation reaches it."""
copy = tmp_path / "bundle"
shutil.copytree(BUNDLE_DIR, copy)
(copy / _ASFALT_FILE).write_text(
f"---\ntype: reference\ntitle: Asfaltdekke\ndimension: asfalt\n---\n\n{_ASFALT_SENTINEL}\n",
encoding="utf-8",
)
index = copy / "index.md"
index.write_text(
index.read_text(encoding="utf-8") + f"\n- [Asfaltdekke]({_ASFALT_FILE})\n", encoding="utf-8"
)
return str(copy)
def _tools(bundle_dir: str, dimension: str | None) -> dict[str, Any]:
return {t.name: t for t in navigator_tools([bundle_dir], dimension=dimension)}
async def _invoke(tool: Any, **arguments: Any) -> str:
"""A tool's answer as text. ``FunctionTool.invoke`` returns ``[Content]``, so a test that
stringified the list would compare object reprs and pass against anything."""
return "".join(getattr(c, "text", "") or "" for c in await tool.invoke(arguments=arguments))
async def test_a_foreign_dimension_document_is_neither_listed_nor_readable(tmp_path) -> None:
"""LOAD-BEARING (e): under a dimension the navigator can neither SEE nor READ a document from
another one — both rungs, because a listing filter alone is a filter in name only.
The tools are called DIRECTLY: a ``ScriptedChatClient`` returns text and never emits a tool
call, so a test that only drove ``run_project`` would leave the whole tool surface outside the
gate (``test_explore_loadbearing``'s own measured correction).
Detach points: drop ``dimension`` from ``directory_listing``; drop the gate in ``read_file``."""
bundle_dir = _bundle_with_a_foreign_dimension(tmp_path)
bundle_id = okf.reconcile_bundle_id(bundle_dir).id
tools = _tools(bundle_dir, "energi")
listing = await _invoke(tools["read_bundle"], bundle_id=bundle_id)
assert _ASFALT_FILE not in listing, (
"a document from another dimension is still listed to the agents"
)
answer = await _invoke(tools["read_file"], bundle_id=bundle_id, path=_ASFALT_FILE)
assert answer.startswith(f"REFUSED ({DimensionScopeRefused.__name__})")
# F99-D3 returns the refusal, so §4.1a's property is asserted on the value: the reason
# travels, the out-of-scope document's bytes do not.
assert _ASFALT_SENTINEL not in answer
async def test_without_a_dimension_the_same_document_is_listed_and_readable(tmp_path) -> None:
"""CAUSALITY CONTROL for (e): with ``dimension=None`` the SAME file is both listed and read, so
its refusal above is caused by the scope and not by the file being unreachable."""
bundle_dir = _bundle_with_a_foreign_dimension(tmp_path)
bundle_id = okf.reconcile_bundle_id(bundle_dir).id
tools = _tools(bundle_dir, None)
listing = await _invoke(tools["read_bundle"], bundle_id=bundle_id)
assert _ASFALT_FILE in listing
assert _ASFALT_SENTINEL in await _invoke(
tools["read_file"], bundle_id=bundle_id, path=_ASFALT_FILE
)
async def test_a_dimension_scoped_run_gives_the_debate_scoped_tools(tmp_path, monkeypatch) -> None:
"""LOAD-BEARING (e, wiring): ``run_project``'s ``dimension`` reaches the TOOLS, not just the
(now absent) rendered context — otherwise §4.1a's promise is kept by nothing at all.
Detach point: build the debate's navigator tools without ``dimension=`` → RED."""
import portfolio_optimiser.run as run_module
bundle_dir = _bundle_with_a_foreign_dimension(tmp_path)
bundle_id = okf.reconcile_bundle_id(bundle_dir).id
captured: list[list[Any]] = []
original = run_module.fresh_workflow
def spy(*args: Any, **kwargs: Any) -> Any:
captured.append(list(kwargs.get("tools") or []))
return original(*args, **kwargs)
monkeypatch.setattr(run_module, "fresh_workflow", spy)
sink: list[str] = []
await run_project(
_PID,
"local",
docs_dir=bundle_dir,
bundle_dir=bundle_dir,
dimension=_ENERGY_DIM,
verdict_input=_VERDICT_INPUT,
client_factory=_recording_factory(sink),
)
read_file = next(t for t in captured[0] if getattr(t, "name", "") == "read_file")
answer = await _invoke(read_file, bundle_id=bundle_id, path=_ASFALT_FILE)
assert answer.startswith(f"REFUSED ({DimensionScopeRefused.__name__})")
assert _ASFALT_SENTINEL not in answer