feat(p19): a direction must NAME the requirement that binds it, and have READ it

Two paid rounds scored 0 of 26 fasit concepts opened -- the same number twice.
P18 closed the navigation side (a listing is a window, an invented path is
refused by name) and it did not move, which makes it a ROLE question: nothing
in the loop ever asked the model to say what requirement binds the direction it
committed to, so opening one was never on the critical path to an answer.

A PREMISE OF THE ORDER WAS FELLED BEFORE ANYTHING WAS BUILT ON IT. A1 places
the demand in _INSTRUCTIONS[HYPOTHESISER_ROLE] alone. Measured: the stress
command sends --mandate and NOT --explore, the two are refused together by
name, and none of the nine round-1/2 outboxes holds a {run_id}-exploration.json
-- the hypothesiser never runs in a stress round, so A3 would have been
unreachable in exactly the paid runs this order commissions.

A2's own sentence resolves it: the refusal goes to the model "som en tur den
kan rette (samme mekanisme som quick_validate's nekt), ikke som en raise" --
and quick_validate IS a tool. declare_requirement therefore lives in
navigator_tools, held by BOTH roles that navigate (the exploration, and since
S2c the debate). It EXISTS only when the caller offers both sinks, which keeps
every pre-P19 call site byte-identical; one sink without the other is refused
at construction. 'opened' is the SAME list ExplorationToolRecorder fills, so
the refusal reads the run's own read trace.

The marked hypothesis carries 'requirement' as a REQUIRED key: omitted is a
hard error, explicit null is legal and needs 'why_none', a half-named one is
refused. A minted approach carries it; a seed never acquires one. The proposer
prompt names it only when the field exists, and the judge counts a hit against
THIS approach's fasit concepts, never against the base.

Load-bearing measured (12 arms), four mutations all red against the whole
suite, green control 1711/5 and demo-transcript.stdout byte-unchanged.
A-iii's predicted signature was FALSIFIED: the golden stays green because the
demo runs without a mandate, so _build_messages' approach branch is never
taken there. A-iv was GREEN first -- the repo's vacuous-gate class, 24th time:
the arm drove _attributable while the hit is computed at the call site.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-15 01:21:47 +02:00
commit c84e8bf6f1
22 changed files with 903 additions and 35 deletions

View file

@ -2598,6 +2598,49 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
(ignorer scopet) → 1 rød, den nye armen ALENE. **Isolert målt på ekte data:** re-dømming av runde 1 (ignorer scopet) → 1 rød, den nye armen ALENE. **Isolert målt på ekte data:** re-dømming av runde 1
med den nye dommeren tar kontrakt-sorasen fra `named=3` til `named=1`, og den ene som står igjen med den nye dommeren tar kontrakt-sorasen fra `named=3` til `named=1`, og den ene som står igjen
er `named_in_measure` — modellens egne ord. er `named_in_measure` — modellens egne ord.
- **En retning må NAVNGI kravet som binder den, og ha LEST det — og ordrens egen plassering var
strukturelt inert (P19 DEL A, 15.09):** to betalte runder på rad ga **0 av 26** fasit-konsepter
åpnet (P16 § 3, P18 § 1). P18 lukket navigasjonssiden (vindu + navngitt nekt for oppfunnet sti) og
tallet rørte seg ikke — dét er hva som gjør det til en ROLLE-sak: ingenting i sløyfa ba noensinne
modellen si hvilket krav i korpuset som binder retningen den forpliktet seg til, så å åpne ett var
aldri på kritisk sti til et svar. **PREMISSET FELT FØR NOE BLE BYGGET PÅ DET:** ordrens A1 legger
kravet i `_INSTRUCTIONS[HYPOTHESISER_ROLE]` ALENE, men stresskommandoen (P18 D1, gjentatt ordrett i
P19 E1) sender `--mandate` og IKKE `--explore`, de to er nektet sammen ved navn, og **ingen av de
ni runde-1/2-utboksene har en `{run_id}-exploration.json`** — hypotesisereren kjører aldri i en
stressrunde, så A3 («kravet når forslaget») ville vært unåbar i nøyaktig de betalte kjøringene
ordren bestiller. **A2s egen setning er dét som løser det:** nekten skal gå til modellen «som en
tur den kan rette (samme mekanisme som `quick_validate`s nekt), ikke som en `raise`» — og
`quick_validate` ER et verktøy. `declare_requirement` bor derfor i `navigator_tools`, altså hos
BEGGE roller som navigerer (utforskningen, og siden S2c debatten). **Verktøyet FINNES kun når
kalleren gir begge sinkene** (`opened` + `requirements`), og dét er hva som holder hvert
pre-P19-kallsted byte-identisk; én sink uten den andre NEKTES ved konstruksjon, fordi en logg som
ikke ser hva som ble åpnet ville akseptert enhver erklæring (vakuøs-gate-klassen, uttalt på
forhånd). `opened` er den SAMME lista `ExplorationToolRecorder` fyller (alias, aldri kopi —
`_drive`-regelen), så nekten leser kjøringens EGEN lesetrace. `RequirementNotRead` er en
`ValueError` (`BundlePathNotFound`-presedensen). **HYPOTESE-LINJA: `requirement` er en PÅKREVD
nøkkel** — utelatt er en hard feil av samme klasse som en uleselig merket linje, eksplisitt `null`
er lovlig og krever `why_none` (basen uten krav er et FUNN; en ubegrunnet null er en stillhet), og
et halvnavngitt krav nektes (validering, ALDRI reparasjon). Et MYNTET forslag bærer feltet, et FRØ
får det ALDRI (§ C.6 dør 1 er en bevaringsregel). **A3:** linja `Binding requirement: <ref>
(<path>)` finnes i `_build_messages` kun når feltet finnes, så hver prompt skrevet før i dag er
byte-identisk. **A4:** dommerens `requirement_hit` måles mot DENNE approachens fasit-konsepter,
aldri mot basen, og `requirement_source` (`approach` | `run` | `absent`) sier hvilken av de to
stedene den kom fra — debatten erklærer ÉN gang per kjøring og kan derfor ikke tilskrives én
approach alene. Load-bearing MÅLT (`tests/test_binding_requirement_loadbearing.py`, 12 armer),
**fire mutasjoner alle røde mot HELE suiten** + grønn kontroll **1710/5** (fra 1699/5, supersett,
0 fjernet) og golden `demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): A-i feltet valgfritt igjen (1) · A-ii nekten sjekker
ikke mot åpnede stier (1) · A-iii A3-linja fyrer ubetinget (1) · A-iv dommeren teller mot hele
basen (1). **ÉN AV ORDRENS SPÅDDE SIGNATURER BLE FALSIFISERT AV MÅLINGEN:** A-iii skulle gjøre
GOLDEN rød — den er byte-uendret, fordi demoen kjører UTEN mandat og `_build_messages`' approach-
gren derfor aldri tas på golden-stien; vitnet er byte-identitets-halvdelen av arm (g).
**OG ÉN MUTASJON FALSIFISERTE TESTEN FØRST (repoets vakuøs-gate-klasse, TJUEFJERDE gang):** A-iv
sto GRØNN mot HELE suiten, fordi arm (h) drev `_attributable` mens treffet regnes ut på KALLSTEDET
i `score_context_set`; armen driver nå hele dommeren mot en erklæring som ER i basen men IKKE er
fasitens, med en positiv kontroll. **Ærlighets-grenser, uttalt:** ingen LEVENDE modell har kalt
verktøyet (structured-output-grensens klasse); debattens erklæring er RUN-nivå, så en approach
uten eget krav tilskrives kjøringens erklæringer og raden SIER at den gjør det; og feltet når
ikke `ProvenanceStamp` (stempelet beskriver gaten som dømte ÉN kandidat).
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet. - **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase. - Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -460,6 +460,18 @@ when the seam is detached, so the loop cannot silently degrade into theater.
None of them has a default, because an omitted cap falls back to an *unbounded* loop rather than None of them has a default, because an omitted cap falls back to an *unbounded* loop rather than
a conservative one. a conservative one.
**Naming the requirement that binds a direction.** Whoever navigates a knowledge base — the
exploration's hypothesiser and, since the debate started navigating, the proposer — is asked to
name the ONE requirement that binds the direction it commits to, and to declare it with
`declare_requirement(bundle_id, path, ref)` after reading it. A declaration naming a document
the run never opened is **refused by name** (`RequirementNotRead`) and comes back as a turn the
model can correct by going and reading it; nothing is recorded until it has. A marked hypothesis
therefore carries `"requirement": {"path": ..., "ref": ...}`, or an explicit `null` together with
`"why_none"` — a base that holds no requirement for a direction is a finding worth stating, and
the field is never simply omitted. Where an approach carries one, the proposer's prompt names it
and asks for it back verbatim in `measure`, and the declaration is written to
`{run_id}-debate.json` / `{run_id}-exploration.json` under `requirements`.
**Answering the plan review (`--plan-review`).** With `enable_plan_review` set, the exploration **Answering the plan review (`--plan-review`).** With `enable_plan_review` set, the exploration
stops before the loop is allowed to run and asks you to sign the plan off. `--plan-review` stops before the loop is allowed to run and asks you to sign the plan off. `--plan-review`
answers it *at your terminal*: you are shown the plan, and you type `approve` or answers it *at your terminal*: you are shown the plan, and you type `approve` or

View file

@ -50,7 +50,12 @@ from portfolio_optimiser import okf
from portfolio_optimiser.backends import Profile from portfolio_optimiser.backends import Profile
from portfolio_optimiser.budget import Budget, BudgetExceeded, BudgetMiddleware, TokenMeter from portfolio_optimiser.budget import Budget, BudgetExceeded, BudgetMiddleware, TokenMeter
from portfolio_optimiser.ir import SavingsProposal from portfolio_optimiser.ir import SavingsProposal
from portfolio_optimiser.mandate import OWN_PROPOSAL_ID, Approach, Mandate from portfolio_optimiser.mandate import (
OWN_PROPOSAL_ID,
Approach,
BindingRequirement,
Mandate,
)
from portfolio_optimiser.retrieval import safe_resolve from portfolio_optimiser.retrieval import safe_resolve
from portfolio_optimiser.tracing import exploration_tracer from portfolio_optimiser.tracing import exploration_tracer
from portfolio_optimiser.validator import Rejection, validate_proposal from portfolio_optimiser.validator import Rejection, validate_proposal
@ -205,10 +210,17 @@ _INSTRUCTIONS: Final = {
), ),
HYPOTHESISER_ROLE: ( HYPOTHESISER_ROLE: (
"You shape ONE candidate cost-saving direction at a time from what the navigator found. " "You shape ONE candidate cost-saving direction at a time from what the navigator found. "
"BEFORE you commit to a direction, name the ONE requirement in the knowledge base that "
"BINDS it: have the navigator find it with read_dir(filter=...) and read it with "
"read_file, then call declare_requirement with the base id, that path and the "
"requirement's own number. A direction with no requirement behind it is a guess. "
"You may call quick_validate to sanity-check a candidate's numbers; its verdict is " "You may call quick_validate to sanity-check a candidate's numbers; its verdict is "
"ADVISORY and is not the project's decision. When you commit to a direction, end your " "ADVISORY and is not the project's decision. When you commit to a direction, end your "
f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", ' f'turn with a line of the form: {HYPOTHESIS_MARKER} {{"label": "<short name>", '
'"rationale": "<why this project, in your own words>"}' '"rationale": "<why this project, in your own words>", "requirement": '
'{"path": "<the path you read>", "ref": "<the requirement number>"}}. If the base '
'genuinely holds no requirement for this direction, send "requirement": null and '
'"why_none": "<why the base has none>" instead — the field is never omitted.'
), ),
} }
@ -333,6 +345,35 @@ class ToolCall:
path: str path: str
@dataclass(frozen=True)
class DeclaredRequirement:
"""One requirement a navigating role DECLARED as binding, after reading it (P19 DEL A).
Recorded on a CALLER-OWNED sink for the reason ``ToolCall`` is: the declaration is made mid-run
by a tool, and a budget stop after it constructs no result at all so a returned value would
lose exactly the evidence a paid run was bought for.
"""
bundle_id: str
path: str
ref: str
class RequirementNotRead(ValueError):
"""A role declared a binding requirement it never opened (P19 A2).
**A refusal the model can act on, never a raise that ends the run.** The declaration is a claim
about the corpus, and the cheapest falsifier of it is the run's own read trace: a path that is
not among this run's ``read_file`` calls was not read, whatever the declaration says. Answering
that as a refused TURN ``quick_validate``'s mechanism, one field over — leaves the model able
to go and read it; raising would end a run over a mistake that costs one tool call to fix.
A ``ValueError``, the ``BundlePathNotFound``/``DimensionScopeRefused`` precedent: the caller is
a model choosing a path, so if it ever escapes a tool body it belongs on the CLI's refusal
tuple and hosting's 400 arm rather than the crash channel.
"""
def _string_argument(arguments: Any, key: str) -> str: def _string_argument(arguments: Any, key: str) -> str:
"""One named argument of a call, from either shape ``FunctionInvocationContext`` allows. """One named argument of a call, from either shape ``FunctionInvocationContext`` allows.
@ -415,6 +456,11 @@ class ExplorationTrace:
#: like a successful one on every other field, which is what made the dress rehearsal vacuous #: like a successful one on every other field, which is what made the dress rehearsal vacuous
#: by construction and unreadable after the fact. #: by construction and unreadable after the fact.
tool_calls: list[ToolCall] = field(default_factory=list) tool_calls: list[ToolCall] = field(default_factory=list)
#: Every requirement a role DECLARED as binding, in declaration order (P19 DEL A). Beside
#: ``tool_calls`` rather than derived from it: the trace says which documents were opened, this
#: says which one the model committed to as the thing that binds — two different facts, and the
#: second cannot be inferred from the first.
requirements: list[DeclaredRequirement] = field(default_factory=list)
#: Tokens spent so far, refreshed as the loop turns rather than written once at the end. The #: Tokens spent so far, refreshed as the loop turns rather than written once at the end. The
#: meter is internal to ``explore``, so this is the only way the artefact can report a spend — #: meter is internal to ``explore``, so this is the only way the artefact can report a spend —
#: and updating it per iteration is what makes it readable for a run a cap cut short, which is #: and updating it per iteration is what makes it readable for a run a cap cut short, which is
@ -517,6 +563,10 @@ def trace_payload(
for call in trace.quick_validations for call in trace.quick_validations
], ],
"tool_calls": tool_call_payload(trace.tool_calls), "tool_calls": tool_call_payload(trace.tool_calls),
# P19 DEL A: what this run declared as binding, beside what it opened. An empty list is an
# honest positive statement — the run declared nothing — which is ``Bundle.skipped``'s
# empty tuple rather than a field that has to be inferred from an absence.
"requirements": requirement_payload(trace.requirements),
} }
@ -938,6 +988,7 @@ def _index_excerpt(body: str) -> tuple[str, bool]:
#: reachable ``ExplorationError`` is ``_resolve_bundle``'s unknown-base refusal. #: reachable ``ExplorationError`` is ``_resolve_bundle``'s unknown-base refusal.
_RETURNABLE_REFUSALS: Final = ( _RETURNABLE_REFUSALS: Final = (
ExplorationError, ExplorationError,
RequirementNotRead,
okf.BundleIdMismatch, okf.BundleIdMismatch,
okf.BundlePathNotFound, okf.BundlePathNotFound,
okf.DocumentPathRefused, okf.DocumentPathRefused,
@ -1005,9 +1056,14 @@ def _refused_text(exc: Exception) -> str:
def navigator_tools( def navigator_tools(
bundle_dirs: Sequence[str], *, dimension: str | None = None bundle_dirs: Sequence[str],
*,
dimension: str | None = None,
opened: list[ToolCall] | None = None,
requirements: list[DeclaredRequirement] | None = None,
) -> list[FunctionTool]: ) -> list[FunctionTool]:
"""The navigator's three tools: survey the catalogue, open one base, read one document. """The navigator's tools: survey the catalogue, open one base, read one document — and, when
the caller offers the two sinks, DECLARE the requirement that binds a direction.
Progressive disclosure, not stuffing (målbilde §2/§4): ``list_bundles`` never returns content, Progressive disclosure, not stuffing (målbilde §2/§4): ``list_bundles`` never returns content,
only what each base IS and crucially whether it ships a ``cost-baseline.json``, because a base only what each base IS and crucially whether it ships a ``cost-baseline.json``, because a base
@ -1052,7 +1108,21 @@ def navigator_tools(
hides a foreign-dimension document while ``read_file`` still serves it by path is a filter in hides a foreign-dimension document while ``read_file`` still serves it by path is a filter in
name only a model-chosen path is untrusted input, so the gate belongs where the bytes leave. name only a model-chosen path is untrusted input, so the gate belongs where the bytes leave.
``None`` (the exploration's own call) admits everything, byte-identical to before. ``None`` (the exploration's own call) admits everything, byte-identical to before.
**``declare_requirement`` exists only when the caller passes BOTH sinks** (P19 DEL A), and that
is what keeps every other call site byte-identical including the tool-set assertions three
older gates make. ``opened`` is the SAME list ``ExplorationToolRecorder`` appends to (an alias,
never a copy the ``_drive`` rule), because the refusal this tool exists for is answered by
the run's own read trace and a second record of it would be free to disagree with the first.
Passing one without the other is refused at construction: a log that cannot see what was opened
would accept every declaration, which is the vacuous-gate class.
""" """
if (opened is None) != (requirements is None):
raise ExplorationError(
"navigator_tools takes 'opened' and 'requirements' together or not at all: a "
"requirement sink with no read trace could not refuse an undeclared read, and a read "
"trace with no sink would record nothing"
)
index = _bundle_index(bundle_dirs) index = _bundle_index(bundle_dirs)
@tool( @tool(
@ -1239,7 +1309,52 @@ def navigator_tools(
) )
return resolved.read_text(encoding="utf-8") return resolved.read_text(encoding="utf-8")
return [list_bundles, read_bundle, read_dir, read_file] @tool(
name="declare_requirement",
description=(
"Declare the ONE requirement in a knowledge base that BINDS the direction you are "
"about to commit to, by base id, the bundle-relative path read_file gave you, and the "
"requirement's own number as its frontmatter states it. You must have READ the "
"document with read_file first: a declaration naming a path this run never opened is "
"refused, and reading it is the correction. Use read_dir with a 'filter' word to find "
"it, read_file to read it, then declare it."
),
)
def declare_requirement(bundle_id: str, path: str, ref: str) -> dict[str, Any]:
try:
return _declare_requirement(bundle_id, path, ref)
except _RETURNABLE_REFUSALS as exc:
return _refused_mapping(exc)
def _declare_requirement(bundle_id: str, path: str, ref: str) -> dict[str, Any]:
assert opened is not None and requirements is not None # the constructor guard above
# The base is resolved by the SAME index every read rung uses, so an unknown base is
# refused here exactly as it is there rather than being accepted into the record.
_resolve_bundle(index, bundle_id)
read_paths = [call.path for call in opened if call.name == "read_file" and call.path]
if path not in read_paths:
raise RequirementNotRead(
f"{path!r}; this run has opened {len(read_paths)} document(s) with read_file, and "
"this is not one of them. Read it first — a requirement nobody read cannot bind a "
"direction"
)
requirements.append(DeclaredRequirement(bundle_id=bundle_id, path=path, ref=ref))
return {"declared": True, "bundle_id": bundle_id, "path": path, "ref": ref}
tools = [list_bundles, read_bundle, read_dir, read_file]
if requirements is not None:
tools.append(declare_requirement)
return tools
def requirement_payload(declared: Sequence[DeclaredRequirement]) -> list[dict[str, Any]]:
"""The ONE rendering of declared requirements into plain data, in DECLARATION order.
Two artefacts carry it ``{run_id}-exploration.json`` and ``{run_id}-debate.json`` and two
copies of "what a declaration looks like" would drift into two answers about one run, which is
the -(p) defect landing in exactly the files an operator reads after a paid run.
"""
return [{"bundle_id": d.bundle_id, "path": d.path, "ref": d.ref} for d in declared]
def _refused( def _refused(
@ -1369,6 +1484,8 @@ def fresh_exploration_workflow(
bundle_dirs: Sequence[str] = (), bundle_dirs: Sequence[str] = (),
middleware: Sequence[Any] | None = None, middleware: Sequence[Any] | None = None,
quick_validate_sink: list[QuickValidation] | None = None, quick_validate_sink: list[QuickValidation] | None = None,
tool_call_sink: list[ToolCall] | None = None,
requirement_sink: list[DeclaredRequirement] | None = None,
checkpoint_dir: str | None = None, checkpoint_dir: str | None = None,
) -> Any: ) -> Any:
"""Build a FRESH Magentic workflow with fresh agents and fresh clients (mirrors """Build a FRESH Magentic workflow with fresh agents and fresh clients (mirrors
@ -1392,9 +1509,18 @@ def fresh_exploration_workflow(
token guarantee: agent-level ``ChatMiddleware`` does fire on the manager's own calls (measured token guarantee: agent-level ``ChatMiddleware`` does fire on the manager's own calls (measured
A1), and the manager is the most talkative participant in the loop. A1), and the manager is the most talkative participant in the loop.
""" """
# P19 DEL A: the declaration tool reaches BOTH roles, and that is a measurement rather than
# generosity. The instruction that asks for a binding requirement is the HYPOTHESISER's — it is
# the role that commits to a direction — while the ``read_file`` calls the refusal checks are
# the NAVIGATOR's. Giving it to the navigator alone would leave the committing role unable to
# state its own commitment; to the hypothesiser alone, unable to declare what its partner read.
navigator = list(
navigator_tools(bundle_dirs, opened=tool_call_sink, requirements=requirement_sink)
)
hypothesiser_tools: list[Any] = [quick_validate_tool(bundle_dirs, sink=quick_validate_sink)] hypothesiser_tools: list[Any] = [quick_validate_tool(bundle_dirs, sink=quick_validate_sink)]
hypothesiser_tools += [t for t in navigator if getattr(t, "name", "") == "declare_requirement"]
tools_by_role: dict[str, list[Any]] = { tools_by_role: dict[str, list[Any]] = {
NAVIGATOR_ROLE: list(navigator_tools(bundle_dirs)), NAVIGATOR_ROLE: navigator,
HYPOTHESISER_ROLE: hypothesiser_tools, HYPOTHESISER_ROLE: hypothesiser_tools,
} }
participants = [ participants = [
@ -1583,13 +1709,45 @@ def _resolve_hypothesis_bundle(raw: str, bundle_ids: Sequence[str]) -> str:
) )
def _requirement_of(data: Mapping[str, Any], line: str) -> BindingRequirement | None:
"""The marked hypothesis's binding requirement, or ``None`` when the base genuinely has none.
**The field is never OMITTED** (P19 A1): a missing key is a hard error of the same class as an
unreadable marked line, because the marker is what makes fail-closed affordable the model
committed to a direction, and "which requirement binds it" is part of that commitment rather
than an optional extra. An EXPLICIT ``null`` is legal and needs ``why_none``: a base that holds
no requirement for a direction is a finding worth stating, and one stated without a reason is
indistinguishable from the model having skipped the question.
"""
if "requirement" not in data:
raise HypothesisParseError(
"a marked hypothesis must carry a 'requirement' — either "
'{"path": ..., "ref": ...} for the document that binds it, or null together with '
f"'why_none'. A direction with no requirement behind it is a guess; got: {line}"
)
raw = data["requirement"]
if raw is None:
if not data.get("why_none"):
raise HypothesisParseError(
"a marked hypothesis with 'requirement': null must say 'why_none' — the base "
f"holding no requirement is a finding, and an unexplained null is a silence: {line}"
)
return None
if not isinstance(raw, dict) or not raw.get("path") or not raw.get("ref"):
raise HypothesisParseError(
"a marked hypothesis's 'requirement' needs a non-empty 'path' and 'ref'; a half-named "
f"requirement reads as a citation and points at nothing: {line}"
)
return BindingRequirement(path=str(raw["path"]), ref=str(raw["ref"]))
def _parse_hypotheses( def _parse_hypotheses(
texts: Sequence[str], bundle_ids: Sequence[str] texts: Sequence[str], bundle_ids: Sequence[str]
) -> list[tuple[str, str, str]]: ) -> list[tuple[str, str, str, BindingRequirement | None]]:
"""Every marked ``(label, rationale, bundle_id)`` triple the hypothesiser committed to, in turn """Every marked ``(label, rationale, bundle_id, requirement)`` the hypothesiser committed to,
order. The base is resolved HERE rather than at minting time so an unroutable claim is refused in turn order. The base is resolved HERE rather than at minting time so an unroutable claim is
while the line that made it is still in hand for the error message.""" refused while the line that made it is still in hand for the error message."""
found: list[tuple[str, str, str]] = [] found: list[tuple[str, str, str, BindingRequirement | None]] = []
for text in texts: for text in texts:
for line in text.splitlines(): for line in text.splitlines():
stripped = line.strip() stripped = line.strip()
@ -1613,6 +1771,7 @@ def _parse_hypotheses(
str(data["label"]), str(data["label"]),
str(data["rationale"]), str(data["rationale"]),
_resolve_hypothesis_bundle(str(data.get("bundle_id") or ""), bundle_ids), _resolve_hypothesis_bundle(str(data.get("bundle_id") or ""), bundle_ids),
_requirement_of(data, stripped),
) )
) )
return found return found
@ -1649,7 +1808,8 @@ def refuse_unroutable_seeds(seeds: Sequence[Approach], bundle_ids: Sequence[str]
def _mint_approaches( def _mint_approaches(
seeds: Sequence[Approach], discovered: Sequence[tuple[str, str, str]] seeds: Sequence[Approach],
discovered: Sequence[tuple[str, str, str, BindingRequirement | None]],
) -> tuple[Approach, ...]: ) -> tuple[Approach, ...]:
"""Seeds FIRST, untouched, then one approach per discovered direction. """Seeds FIRST, untouched, then one approach per discovered direction.
@ -1662,7 +1822,7 @@ def _mint_approaches(
taken = {approach.id for approach in seeds} taken = {approach.id for approach in seeds}
minted: list[Approach] = list(seeds) minted: list[Approach] = list(seeds)
counter = 0 counter = 0
for label, rationale, bundle_id in discovered: for label, rationale, bundle_id, requirement in discovered:
counter += 1 counter += 1
candidate = f"hypothesis-{counter}" candidate = f"hypothesis-{counter}"
while candidate in taken or candidate == OWN_PROPOSAL_ID: while candidate in taken or candidate == OWN_PROPOSAL_ID:
@ -1675,8 +1835,17 @@ def _mint_approaches(
# #
# ``bundle_id`` is stamped on a MINTED approach because there is nothing here to preserve — # ``bundle_id`` is stamped on a MINTED approach because there is nothing here to preserve —
# the opposite call from the seeds above, which pass through untouched (§ C.6 door 1). # the opposite call from the seeds above, which pass through untouched (§ C.6 door 1).
# ``requirement`` is stamped on a MINTED approach for the same reason ``bundle_id`` is:
# there is nothing here to preserve. A SEED passes through untouched (§ C.6 door 1) — an
# expert who named no requirement is not to be given one on their behalf.
minted.append( minted.append(
Approach(id=candidate, label=label, description=rationale, bundle_id=bundle_id) Approach(
id=candidate,
label=label,
description=rationale,
bundle_id=bundle_id,
requirement=requirement,
)
) )
return tuple(minted) return tuple(minted)
@ -1801,6 +1970,11 @@ async def explore(
bundle_dirs=bundle_dirs, bundle_dirs=bundle_dirs,
middleware=[BudgetMiddleware(meter), ExplorationToolRecorder(trace.tool_calls)], middleware=[BudgetMiddleware(meter), ExplorationToolRecorder(trace.tool_calls)],
quick_validate_sink=trace.quick_validations, quick_validate_sink=trace.quick_validations,
# ALIASES of the trace's own lists, never copies (the ``_drive`` rule): the refusal reads
# the same read trace the recorder writes, so the two can never disagree about what this
# run opened.
tool_call_sink=trace.tool_calls,
requirement_sink=trace.requirements,
checkpoint_dir=checkpoint_dir, checkpoint_dir=checkpoint_dir,
) )
@ -2093,6 +2267,11 @@ async def resume_exploration(
bundle_dirs=parked.bundle_dirs, bundle_dirs=parked.bundle_dirs,
middleware=[BudgetMiddleware(meter), ExplorationToolRecorder(trace.tool_calls)], middleware=[BudgetMiddleware(meter), ExplorationToolRecorder(trace.tool_calls)],
quick_validate_sink=trace.quick_validations, quick_validate_sink=trace.quick_validations,
# ALIASES of the trace's own lists, never copies (the ``_drive`` rule): the refusal reads
# the same read trace the recorder writes, so the two can never disagree about what this
# run opened.
tool_call_sink=trace.tool_calls,
requirement_sink=trace.requirements,
checkpoint_dir=checkpoint_dir, checkpoint_dir=checkpoint_dir,
) )

View file

@ -327,6 +327,17 @@ def _build_messages(
) )
if approach.description: if approach.description:
head += f"Why the expert wants it evaluated: {approach.description}\n" head += f"Why the expert wants it evaluated: {approach.description}\n"
# P19 A3: the ONE requirement of the knowledge base that binds this direction, when one was
# declared. The line exists only when the field does, so every prompt written before today
# — the demo's included, which is what keeps the golden transcript byte-identical — is
# unchanged by construction. ``ref`` leads because that is what the model is asked to
# restate; the path follows so the claim can be checked against what the run opened.
if approach.requirement is not None:
head += (
f"Binding requirement: {approach.requirement.ref} "
f"({approach.requirement.path})\n"
"Name that requirement verbatim in 'measure'.\n"
)
prompt = ( prompt = (
f"{head}" f"{head}"
f"Project: {project.id} - {project.name}\n" f"Project: {project.id} - {project.name}\n"

View file

@ -43,6 +43,31 @@ from portfolio_optimiser.ir import AffectedItem, CostBaseline, SavingsProposal
OWN_PROPOSAL_ID = "own-proposal" OWN_PROPOSAL_ID = "own-proposal"
class BindingRequirement(BaseModel):
"""The ONE document in the knowledge base that BINDS a direction (P19 DEL A).
Measured over two paid stress rounds (P16 § 3, P18 § 1): **0 of 26** fasit concepts were ever
opened, in both rounds, while every run still produced proposals so a direction could be
committed to, quantified and validated without a single requirement of the corpus having been
read. P18 closed the navigation side of that (a window, a named refusal for an invented path)
and the number did not move, which is what made it a ROLE question rather than a ladder one:
nothing in the loop ever asked the model to name what binds it.
``path`` is bundle-relative exactly as a listing gave it the same string ``read_file`` takes,
so it can be checked against what the run actually opened without a second normalisation.
``ref`` is the requirement's OWN number as the document's frontmatter states it, never a
paraphrase: it is what a reader searches the corpus with, and what the judge matches the fasit
on.
Both are required with ``min_length=1``. A half-declared requirement would be worse than none:
it reads as a citation and points at nothing, which is the fabricated-provenance class
``write_concept_file`` refuses and ``RunFailure`` names in its own docstring.
"""
path: str = Field(min_length=1)
ref: str = Field(min_length=1)
class Approach(BaseModel): class Approach(BaseModel):
"""One approach a domain expert wants evaluated for a project. """One approach a domain expert wants evaluated for a project.
@ -84,6 +109,15 @@ class Approach(BaseModel):
#: The field is a ROUTING key, never a claim about content. It says which pipeline the approach #: The field is a ROUTING key, never a claim about content. It says which pipeline the approach
#: must be evaluated in, and ``route_by_bundle`` is the one place that reads it. #: must be evaluated in, and ``route_by_bundle`` is the one place that reads it.
bundle_id: str = "" bundle_id: str = ""
#: The ONE requirement of the knowledge base that binds this direction (P19 DEL A).
#:
#: DEFAULTS to ``None``, which keeps every mandate written before today valid and dispatchable
#: unchanged — the ``bundle_id`` precedent, and the honest reading of a commission whose author
#: named no requirement. ``None`` is therefore "none stated", never "none exists": the loop
#: that MINTS an approach must say which of the two it means (``why_none`` on the hypothesis
#: line), because a direction the base has no requirement for is a finding, while one nobody
#: looked for is a silence.
requirement: BindingRequirement | None = None
class Mandate(BaseModel): class Mandate(BaseModel):

View file

@ -194,8 +194,10 @@ def write_debate_tools(
run_id: str, run_id: str,
*, *,
tool_calls: Sequence[Mapping[str, Any]], tool_calls: Sequence[Mapping[str, Any]],
requirements: Sequence[Mapping[str, Any]] = (),
) -> Path: ) -> Path:
"""Write ``{run_id}-debate.json`` — WHICH documents the debate opened, in call order (S2c). """Write ``{run_id}-debate.json`` — WHICH documents the debate opened, in call order (S2c), and
WHICH requirement it declared as binding (P19 DEL A).
The sibling of ``write_exploration`` one phase over. Since S2c the debate navigates the The sibling of ``write_exploration`` one phase over. Since S2c the debate navigates the
knowledge base instead of being handed it whole, so "what did this run actually read" is a knowledge base instead of being handed it whole, so "what did this run actually read" is a
@ -213,7 +215,16 @@ def write_debate_tools(
directory.mkdir(parents=True, exist_ok=True) directory.mkdir(parents=True, exist_ok=True)
path = directory / f"{run_id}-debate.json" path = directory / f"{run_id}-debate.json"
path.write_text( path.write_text(
_dump({"run_id": run_id, "tool_calls": [dict(call) for call in tool_calls]}), _dump(
{
"run_id": run_id,
"tool_calls": [dict(call) for call in tool_calls],
# DEFAULTS to empty for the same reason ``Bundle.skipped`` does: "this debate
# declared nothing" is an honest positive statement, and it is the one every run
# written before today makes.
"requirements": [dict(r) for r in requirements],
}
),
encoding="utf-8", encoding="utf-8",
) )
return path return path

View file

@ -64,6 +64,7 @@ from portfolio_optimiser.explore import (
ExplorationResult, ExplorationResult,
ExplorationToolRecorder, ExplorationToolRecorder,
ExplorationTrace, ExplorationTrace,
DeclaredRequirement,
ParkedStateError, ParkedStateError,
PlanReviewDecision, PlanReviewDecision,
PlanReviewParked, PlanReviewParked,
@ -75,6 +76,7 @@ from portfolio_optimiser.explore import (
parked_notice, parked_notice,
parked_payload, parked_payload,
navigator_tools, navigator_tools,
requirement_payload,
resume_exploration, resume_exploration,
terminal_plan_reviewer, terminal_plan_reviewer,
tool_call_payload, tool_call_payload,
@ -607,7 +609,11 @@ def _bundle_pointer(bundle: okf.Bundle, bundle_id: str, *, dimension: str | None
"A listing is a WINDOW: it reports 'total' for the level and gives you 'limit' entries " "A listing is a WINDOW: it reports 'total' for the level and gives you 'limit' entries "
"from 'offset'. When 'total' is large, do not page through it — narrow it: " "from 'offset'. When 'total' is large, do not page through it — narrow it: "
f"read_dir({bundle_id!r}, path, filter='<word>') answers with the entries whose title, " f"read_dir({bundle_id!r}, path, filter='<word>') answers with the entries whose title, "
"requirement number or path contains that word, and reports 'total_matches'." "requirement number or path contains that word, and reports 'total_matches'.\n"
"Before you settle on a measure, name the ONE requirement of this base that BINDS it: "
"find it with a filter, read it with read_file, then call "
f"declare_requirement({bundle_id!r}, path, ref) with the requirement's own number. A "
"declaration naming a document this run never opened is refused; reading it is the fix."
) )
@ -1116,6 +1122,21 @@ async def run_project(
# gate's share rule needs, and composing them here — where the base is already walked — is what # gate's share rule needs, and composing them here — where the base is already walked — is what
# keeps them from being a second, drifting reconstruction (kø-(p)). # keeps them from being a second, drifting reconstruction (kø-(p)).
bundle_grounding: tuple[str, ...] = () bundle_grounding: tuple[str, ...] = ()
# S2c: a CALLER-OWNED sink for what the debate opens (the ``parse_failures``/``ExplorationTrace``
# shape). A returned value would be lost on exactly the run that most needs the evidence — a
# budget stop mid-debate raises out of ``debate.run`` and constructs no ``RunResult`` at all.
# ``ExplorationToolRecorder`` is REUSED rather than re-implemented: it is already the recorder
# for in-process navigator calls, ordered and un-deduplicated, which is exactly the question
# here too ("did this run open anything, and in what sequence"). Its sibling
# ``mcp_tools.ToolCallRecorder`` stays what it is — a sorted, de-duplicated EGRESS claim.
#
# BOUND HERE, above the fork, because the bundle arm hands both lists to ``navigator_tools``:
# the declaration rung refuses against the very trace the recorder writes, and a second list
# would be free to disagree with it about what this run opened (kø-(p)).
debate_tool_calls: list[ToolCall] = []
#: P19 DEL A: which requirement the debate declared as binding, in declaration order.
debate_requirements: list[DeclaredRequirement] = []
if bundle_dir is not None: if bundle_dir is not None:
bundle = okf.navigate_bundle(bundle_dir) bundle = okf.navigate_bundle(bundle_dir)
bundle_grounding = tuple( bundle_grounding = tuple(
@ -1192,7 +1213,17 @@ async def run_project(
# both rungs (``navigator_tools``' own gate), because that is now where the bytes # both rungs (``navigator_tools``' own gate), because that is now where the bytes
# leave. Under a payload the SAME two gates are re-raised by # leave. Under a payload the SAME two gates are re-raised by
# ``prepass.verify_against_bundle`` on the mounted documents instead. # ``prepass.verify_against_bundle`` on the mounted documents instead.
debate_tools = list(navigator_tools([bundle_dir], dimension=dimension_id)) # P19 DEL A: the debate gets the declaration rung too, and the sinks are what create
# it. ``debate_tool_calls`` is bound below — this list is the SAME one
# ``ExplorationToolRecorder`` fills, so the refusal reads the run's own read trace.
debate_tools = list(
navigator_tools(
[bundle_dir],
dimension=dimension_id,
opened=debate_tool_calls,
requirements=debate_requirements,
)
)
# What the navigation could NOT reach, taken from the run's ONE walk. The road path below # What the navigation could NOT reach, taken from the run's ONE walk. The road path below
# navigates no bundle at all, so its empty tuple is literally true rather than a stand-in. # navigates no bundle at all, so its empty tuple is literally true rather than a stand-in.
skipped_links: tuple[okf.SkippedLink, ...] = bundle.skipped skipped_links: tuple[okf.SkippedLink, ...] = bundle.skipped
@ -1266,7 +1297,6 @@ async def run_project(
# for in-process navigator calls, ordered and un-deduplicated, which is exactly the question # for in-process navigator calls, ordered and un-deduplicated, which is exactly the question
# here too ("did this run open anything, and in what sequence"). Its sibling # here too ("did this run open anything, and in what sequence"). Its sibling
# ``mcp_tools.ToolCallRecorder`` stays what it is — a sorted, de-duplicated EGRESS claim. # ``mcp_tools.ToolCallRecorder`` stays what it is — a sorted, de-duplicated EGRESS claim.
debate_tool_calls: list[ToolCall] = []
debate_middleware: list[Any] = [budget_mw, ExplorationToolRecorder(debate_tool_calls)] debate_middleware: list[Any] = [budget_mw, ExplorationToolRecorder(debate_tool_calls)]
if call_recorder is not None: if call_recorder is not None:
debate_middleware.append(call_recorder) debate_middleware.append(call_recorder)
@ -1343,7 +1373,10 @@ async def run_project(
declaration=prepass.declaration_payload(prepass_declaration), declaration=prepass.declaration_payload(prepass_declaration),
) )
outbox.write_debate_tools( outbox.write_debate_tools(
outbox_dir, run_id, tool_calls=tool_call_payload(debate_tool_calls) outbox_dir,
run_id,
tool_calls=tool_call_payload(debate_tool_calls),
requirements=requirement_payload(debate_requirements),
) )
# F1: the candidate must derive from the DEBATE. Feed the proposer's converged output into # F1: the candidate must derive from the DEBATE. Feed the proposer's converged output into
# generation (retrieval context is the last-resort fallback only). The checker's verdict # generation (retrieval context is the last-resort fallback only). The checker's verdict

View file

@ -745,7 +745,17 @@ def scripted_exploration_factory(
existing ``_proposer_reply`` / ``_CHECKER_APPROVE`` scaffolding, so nothing about the debate existing ``_proposer_reply`` / ``_CHECKER_APPROVE`` scaffolding, so nothing about the debate
changes.""" changes."""
pipeline = scripted_factory({"proposer": _proposer_reply, "checker": _CHECKER_APPROVE}, sink) pipeline = scripted_factory({"proposer": _proposer_reply, "checker": _CHECKER_APPROVE}, sink)
hypothesis_line = f"{HYPOTHESIS_MARKER} " + json.dumps({"label": label, "rationale": rationale}) hypothesis_line = f"{HYPOTHESIS_MARKER} " + json.dumps(
# P19 A1: ``requirement`` is a required key. The demo's scripted manager routes to the
# hypothesiser without any document having been read, so the explicit-null branch is the
# only honest one here — and it is what keeps the golden transcript byte-identical.
{
"label": label,
"rationale": rationale,
"requirement": None,
"why_none": "scripted demo: no document was read",
}
)
def factory(role: str) -> BaseChatClient: def factory(role: str) -> BaseChatClient:
if role == MANAGER_ROLE: if role == MANAGER_ROLE:

View file

@ -69,6 +69,7 @@ import argparse
import json import json
import os import os
import sys import sys
from collections.abc import Sequence
from dataclasses import asdict, dataclass from dataclasses import asdict, dataclass
from pathlib import Path from pathlib import Path
from typing import Any from typing import Any
@ -110,6 +111,14 @@ class ApproachVerdict:
named_in_snippet: bool named_in_snippet: bool
#: (c) - attributable hallucinations, ``citation:<file>`` / ``code:<code>``. #: (c) - attributable hallucinations, ``citation:<file>`` / ``code:<code>``.
hallucinations: tuple[str, ...] hallucinations: tuple[str, ...]
#: P19 A4 - the binding requirement this row can be attributed, and whether it is one of the
#: fasit's own concepts for this approach. ``requirement_source`` says WHICH of the two places
#: it came from, because they are different claims: ``approach`` is per-approach by
#: construction (the mandate carries it), while ``run`` is the DEBATE's declaration, which is
#: written once per run and therefore cannot be attributed to one approach on its own.
requirement_declared: tuple[str, ...]
requirement_source: str # "approach" | "run" | "absent"
requirement_hit: bool
ferdig: bool ferdig: bool
@ -133,6 +142,11 @@ class ContextSetVerdict:
must_refuse: tuple[RefusalVerdict, ...] must_refuse: tuple[RefusalVerdict, ...]
#: Run-level: read paths the base does not carry (guessed by the navigator). #: Run-level: read paths the base does not carry (guessed by the navigator).
hallucinated_reads: tuple[str, ...] hallucinated_reads: tuple[str, ...]
#: P19 A4: every requirement the RUN declared as binding, in declaration order, with the
#: denominator every other field here carries. Reported even when empty - "the run declared
#: none" is the measurement, and a missing field would be indistinguishable from a judge that
#: did not look.
requirements_declared: tuple[str, ...]
tool_calls_seen: int tool_calls_seen: int
citations_seen: int citations_seen: int
approach_rows_seen: int approach_rows_seen: int
@ -171,6 +185,24 @@ def _inside(base: Path, raw: str) -> Path | None:
return resolved if resolved == root or root in resolved.parents else None return resolved if resolved == root or root in resolved.parents else None
def _attributable(approach: Any, declared: Sequence[str]) -> tuple[tuple[str, ...], str]:
"""Which declared requirement paths this approach may be judged on, and where they came from.
The approach's OWN requirement wins when it has one: the mandate carries it per approach, so
it is unambiguous by construction. Otherwise the RUN's declarations are attributable - the
debate declares once for the whole run, so the row says ``run`` rather than pretending the
declaration was made about it. ``absent`` is the third value and is not the same as "declared
nothing that matched": a run that declared nothing is a different finding from one that
declared the wrong document."""
own = getattr(approach, "requirement", None)
path = "" if own is None else str(getattr(own, "path", "") or "")
if path:
return (path,), "approach"
if declared:
return tuple(declared), "run"
return (), "absent"
def score_context_set( def score_context_set(
context_dir: str | Path, context_dir: str | Path,
outbox_dir: str | Path, outbox_dir: str | Path,
@ -206,6 +238,18 @@ def score_context_set(
) )
opened_paths = {c.get("path", "") for c in tool_calls if c.get("name") == "read_file"} opened_paths = {c.get("path", "") for c in tool_calls if c.get("name") == "read_file"}
opened_paths.discard("") opened_paths.discard("")
# P19 A4: WHERE the declarations live was measured, not assumed. A ``--mandate`` run (which is
# what every stress round has been) writes no ``{run_id}-exploration.json`` at all - the
# hypothesiser never runs - so the debate artefact is the only one that can carry them there.
# Both are read, because an ``--explore`` run carries them in the other.
declared_paths: list[str] = []
for artefact in (debate, outbox / f"{run_id}-exploration.json"):
if artefact.is_file():
declared_paths += [
str(r.get("path", ""))
for r in _read_json(artefact).get("requirements", [])
if r.get("path")
]
hallucinated_reads: list[str] = [] hallucinated_reads: list[str] = []
for call in tool_calls: for call in tool_calls:
raw = str(call.get("path", "")) raw = str(call.get("path", ""))
@ -245,6 +289,9 @@ def score_context_set(
named_in_measure=False, named_in_measure=False,
named_in_snippet=False, named_in_snippet=False,
hallucinations=(), hallucinations=(),
requirement_declared=_attributable(approach, declared_paths)[0],
requirement_source=_attributable(approach, declared_paths)[1],
requirement_hit=bool(set(_attributable(approach, declared_paths)[0]) & wanted),
ferdig=False, ferdig=False,
) )
) )
@ -288,6 +335,9 @@ def score_context_set(
# snippets and gave this row ``named`` without the model having said anything. # snippets and gave this row ``named`` without the model having said anything.
named_in_snippet = scope == "narrowed" and any(m in snippets for m in marks) named_in_snippet = scope == "narrowed" and any(m in snippets for m in marks)
attributable, requirement_source = _attributable(approach, declared_paths)
requirement_hit = bool(set(attributable) & wanted)
halluc = [f"citation:{f}" for f in sorted(cited_files - concept_names)] halluc = [f"citation:{f}" for f in sorted(cited_files - concept_names)]
allowed = set(approach.affected_codes) | baseline_codes allowed = set(approach.affected_codes) | baseline_codes
halluc += [ halluc += [
@ -310,6 +360,9 @@ def score_context_set(
named_in_measure=named_in_measure, named_in_measure=named_in_measure,
named_in_snippet=named_in_snippet, named_in_snippet=named_in_snippet,
hallucinations=tuple(halluc), hallucinations=tuple(halluc),
requirement_declared=attributable,
requirement_source=requirement_source,
requirement_hit=requirement_hit,
ferdig=( ferdig=(
grounded grounded
and (named_in_measure or named_in_snippet) and (named_in_measure or named_in_snippet)
@ -355,6 +408,7 @@ def score_context_set(
approaches=tuple(rows), approaches=tuple(rows),
must_refuse=tuple(refusals), must_refuse=tuple(refusals),
hallucinated_reads=tuple(hallucinated_reads), hallucinated_reads=tuple(hallucinated_reads),
requirements_declared=tuple(declared_paths),
tool_calls_seen=len(tool_calls), tool_calls_seen=len(tool_calls),
citations_seen=citations_seen, citations_seen=citations_seen,
approach_rows_seen=rows_seen, approach_rows_seen=rows_seen,

View file

@ -76,7 +76,10 @@ _REPLIES = {
"checker": "VERDICT: APPROVE", "checker": "VERDICT: APPROVE",
"manager": _MANAGER_REPLY, "manager": _MANAGER_REPLY,
"navigator": "NAVIGATOR: read the index.", "navigator": "NAVIGATOR: read the index.",
"hypothesiser": "HYPOTHESIS: " + json.dumps({"label": "Night setback", "rationale": "y"}), "hypothesiser": "HYPOTHESIS: "
+ json.dumps(
{"label": "Night setback", "rationale": "y", "requirement": None, "why_none": "scripted"}
),
} }
_FEEDBACK = "Also test night setback on the ventilation." _FEEDBACK = "Also test night setback on the ventilation."
@ -583,7 +586,17 @@ def test_what_the_first_process_found_survives_into_the_resumed_mandate(tmp_path
carried = dataclasses.replace( carried = dataclasses.replace(
caught.value.parked, caught.value.parked,
hypotheses=("HYPOTHESIS: " + json.dumps({"label": "Carried", "rationale": "found first"}),), hypotheses=(
"HYPOTHESIS: "
+ json.dumps(
{
"label": "Carried",
"rationale": "found first",
"requirement": None,
"why_none": "scripted",
}
),
),
ledger=( ledger=(
ex.LedgerEntry( ex.LedgerEntry(
round_index=1, round_index=1,

View file

@ -0,0 +1,404 @@
"""P19 DEL A — a direction must NAME the requirement that binds it, and it must have READ it.
**The measured silence, two paid rounds deep.** P16 (`docs/2026-09-12-p14-kontekstsett.md` § 4.1,
`docs/2026-09-14-p18-stressrunde-2.md`) and P18 both scored **0 of 26** fasit concepts opened
the same number twice, over four and then five paid runs, while every run still produced proposals
the deterministic gate then judged. P18 closed the navigation side of it (a listing is a window; an
invented path is refused by name) and the number did not move at all. That is what turns it from a
LADDER question into a ROLE one: nothing anywhere in the loop ever asked the model to say what
requirement of the corpus binds the direction it committed to, so opening one was never on the
critical path of producing an answer.
**A PREMISE OF THE ORDER WAS FELLED BEFORE ANYTHING WAS BUILT ON IT.** A1 places the demand in
``_INSTRUCTIONS[HYPOTHESISER_ROLE]`` alone. MEASURED: the stress command (P18 D1, repeated
verbatim in P19 E1) passes ``--mandate`` and NOT ``--explore``, the two are refused together by
name, and not one of the nine round-1/round-2 outboxes holds a ``{run_id}-exploration.json``. The
hypothesiser therefore never runs in a stress round, and a demand that only it can carry would be
structurally inert in exactly the paid runs this order commissions A3 ("the requirement reaches
the proposal") unreachable along with it.
A2's own sentence is what resolves it: the refusal must go to the model "som en tur den kan rette
(samme mekanisme som ``quick_validate``s nekt), ikke som en ``raise``". ``quick_validate`` is a
TOOL. So the demand is a tool ``declare_requirement`` and it lives in ``navigator_tools``,
which since S2c is held by BOTH roles that navigate: the exploration's navigator/hypothesiser and
the debate's proposer/checker. One instruction, one refusal, one record, two doors.
What each arm pins, and what it refuses:
(a) the refusal itself a declaration naming a path this run never opened comes back REFUSED, in
the funn-99 form (``refusal`` carries the kind ``RequirementNotRead``), and NOTHING is
recorded. Without it the tool accepts any string and the whole demand is decorative;
(b) the correction works the same path, once ``read_file`` has actually returned it, is accepted
and recorded. A gate that can only refuse is the mirror image of one that can only pass, and
both prove nothing;
(c) an unknown base is refused by the SAME index every read rung uses, so a declaration cannot
name a corpus this run was never given;
(d) the tool EXISTS only when the caller offers both sinks which is what keeps every other
``navigator_tools`` call site byte-identical and passing one without the other is refused at
construction, because a sink with no read trace would accept every declaration (the repo's
vacuous-gate class, stated in advance rather than discovered);
(e) the marked hypothesis carries ``requirement`` as a REQUIRED key: omitted is a hard error of the
same class as an unreadable marked line, explicit ``null`` is legal and needs ``why_none``, and
a half-named requirement is refused (validation, never repair);
(f) a MINTED approach carries it and a SEED does not acquire one § C.6 door 1 is a preservation
rule, and filling the field in on an expert's behalf would put their name on a claim about the
corpus they did not make;
(g) A3 the requirement reaches the proposer's prompt VERBATIM, and a prompt whose approach has
none is byte-identical to before. That second half is what keeps ``demo-transcript.stdout``
unchanged, and it is asserted here rather than left to the golden;
(h) A4 the judge counts a hit against THIS APPROACH'S fasit concepts, never against the base. A
judge matching the whole base would mark every declaration a hit on a 2 756-document corpus,
which is the P16 vacuity ``citation_scope`` already exists to refuse;
(i) the declaration reaches ``{run_id}-debate.json``, so a paid run's evidence survives the process
that produced it.
"""
from __future__ import annotations
import json
from pathlib import Path
from typing import Any
import pytest
from portfolio_optimiser import explore, stress
from portfolio_optimiser.explore import (
DeclaredRequirement,
HypothesisParseError,
ToolCall,
navigator_tools,
)
from portfolio_optimiser.generate import _build_messages
from portfolio_optimiser.mandate import Approach, BindingRequirement
from portfolio_optimiser.reference_domain import Project
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
def _wired(
bundle_dir: Path = _TUNNEL,
) -> tuple[dict[str, Any], list[ToolCall], list[DeclaredRequirement]]:
"""The tools as a RUN holds them: the declaration rung reading the recorder's own trace."""
opened: list[ToolCall] = []
declared: list[DeclaredRequirement] = []
tools = navigator_tools((str(bundle_dir),), opened=opened, requirements=declared)
return {t.name: t for t in tools}, opened, declared
def _a_concept(bundle_dir: Path = _TUNNEL) -> str:
"""One real concept path in the base, taken from the navigated listing rather than guessed."""
from portfolio_optimiser import okf
return okf.navigate_bundle(str(bundle_dir)).context_files[0].name
# ---------------------------------------------------------------------------------------------
# (a)-(c) the tool
# ---------------------------------------------------------------------------------------------
def test_a_requirement_the_run_never_opened_is_refused_and_recorded_nowhere() -> None:
"""(a) The declaration's one falsifier is the run's own read trace."""
tools, _opened, declared = _wired()
answer = tools["declare_requirement"].func(
bundle_id="tunnel-hauglia", path=_a_concept(), ref="12.1"
)
assert answer["refusal"] == "RequirementNotRead"
assert "0 document(s)" in answer["refused"]
assert declared == [], "a refused declaration must leave no record"
def test_the_correction_is_to_read_it_and_then_it_is_accepted() -> None:
"""(b) The refusal is a turn the model can correct — the half that makes (a) a gate."""
tools, opened, declared = _wired()
path = _a_concept()
body = tools["read_file"].func(bundle_id="tunnel-hauglia", path=path)
assert not body.startswith("REFUSED"), body[:120]
# The recorder is middleware in a real run; here the trace is appended directly, which is the
# SAME list the tool reads.
opened.append(ToolCall(name="read_file", bundle_id="tunnel-hauglia", path=path))
answer = tools["declare_requirement"].func(
bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1"
)
assert answer == {
"declared": True,
"bundle_id": "tunnel-hauglia",
"path": path,
"ref": "Krav 12.1",
}
assert declared == [DeclaredRequirement(bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1")]
def test_a_declaration_naming_an_unknown_base_is_refused() -> None:
"""(c) The same base index every read rung uses, so no second, laxer resolution exists."""
tools, opened, declared = _wired()
opened.append(ToolCall(name="read_file", bundle_id="nope", path="a.md"))
answer = tools["declare_requirement"].func(bundle_id="nope", path="a.md", ref="1")
assert answer["refusal"] == "ExplorationError"
assert declared == []
def test_the_rung_exists_only_when_both_sinks_are_offered() -> None:
"""(d) Every pre-P19 call site is byte-identical, and a half-wired one is refused."""
plain = {t.name for t in navigator_tools((str(_TUNNEL),))}
assert plain == {"list_bundles", "read_bundle", "read_dir", "read_file"}
wired, _, _ = _wired()
assert set(wired) == plain | {"declare_requirement"}
with pytest.raises(explore.ExplorationError, match="together or not at all"):
navigator_tools((str(_TUNNEL),), requirements=[])
with pytest.raises(explore.ExplorationError, match="together or not at all"):
navigator_tools((str(_TUNNEL),), opened=[])
# ---------------------------------------------------------------------------------------------
# (e)-(f) the marked hypothesis
# ---------------------------------------------------------------------------------------------
def _line(**payload: Any) -> str:
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps({"label": "L", "rationale": "R", **payload})
def test_the_requirement_key_is_never_omitted() -> None:
"""(e) A marked line that skips the question is the unreadable-marked-line class."""
with pytest.raises(HypothesisParseError, match="must carry a 'requirement'"):
explore._parse_hypotheses([_line()], ())
def test_an_explicit_null_is_legal_and_must_say_why() -> None:
"""(e) "the base holds none" is a FINDING; an unexplained null is a silence."""
((label, rationale, _base, requirement),) = explore._parse_hypotheses(
[_line(requirement=None, why_none="the base is a standard, not a price schedule")], ()
)
assert (label, rationale, requirement) == ("L", "R", None)
with pytest.raises(HypothesisParseError, match="why_none"):
explore._parse_hypotheses([_line(requirement=None)], ())
def test_a_half_named_requirement_is_refused() -> None:
"""(e) Validation, never repair: a citation that points at nothing is worse than none."""
with pytest.raises(HypothesisParseError, match="non-empty 'path' and 'ref'"):
explore._parse_hypotheses([_line(requirement={"path": "a.md"})], ())
with pytest.raises(HypothesisParseError, match="non-empty 'path' and 'ref'"):
explore._parse_hypotheses([_line(requirement={"ref": "12.1"})], ())
def test_a_minted_approach_carries_it_and_a_seed_never_acquires_one() -> None:
"""(f) § C.6 door 1 is a preservation rule."""
seed = Approach(id="expert-1", label="Seeded", description="the expert's own")
minted = explore._mint_approaches(
(seed,), [("Shaped", "in the loop", "", BindingRequirement(path="k/12-1.md", ref="12.1"))]
)
assert minted[0] is seed and minted[0].requirement is None
assert minted[1].requirement == BindingRequirement(path="k/12-1.md", ref="12.1")
# ---------------------------------------------------------------------------------------------
# (g) A3 — the requirement reaches the proposal
# ---------------------------------------------------------------------------------------------
def _project() -> Project:
return Project(id="X", name="Y", description="d", currency="NOK", docs_dir=".", cost_items=())
def test_the_binding_requirement_reaches_the_proposer_verbatim() -> None:
"""(g) A3, and the byte-identical half is what keeps the golden transcript unchanged."""
bare = Approach(id="a1", label="L", description="D")
with_req = bare.model_copy(
update={"requirement": BindingRequirement(path="R761/12-1/x.md", ref="12.1")}
)
without = _build_messages(_project(), "ctx", approach=bare)[0].text
withit = _build_messages(_project(), "ctx", approach=with_req)[0].text
assert "Binding requirement" not in without
assert "Binding requirement: 12.1 (R761/12-1/x.md)" in withit
# The ONLY difference is the block — an approach without one is byte-identical to before.
assert (
withit.replace(
"Binding requirement: 12.1 (R761/12-1/x.md)\nName that requirement verbatim in 'measure'.\n",
"",
)
== without
)
# ---------------------------------------------------------------------------------------------
# (h) A4 — the judge
# ---------------------------------------------------------------------------------------------
def test_a_hit_is_counted_against_this_approachs_fasit_never_the_base() -> None:
"""(h) The P16 vacuity, one column over: on a 2 756-document base everything is 'in the base'."""
wanted = {"krav/12-1/a.md", "krav/12-12/b.md"}
approach = Approach(
id="a1",
label="L",
requirement=BindingRequirement(path="krav/12-1/a.md", ref="12.1"),
)
paths, source = stress._attributable(approach, ["krav/99-9/elsewhere.md"])
assert (paths, source) == (("krav/12-1/a.md",), "approach")
assert set(paths) & wanted
# A run-level declaration that is NOT one of this approach's fasit concepts is reported and is
# not a hit — the base holds it, which is exactly what must not count.
bare = Approach(id="a2", label="M")
paths, source = stress._attributable(bare, ["krav/99-9/elsewhere.md"])
assert (paths, source) == (("krav/99-9/elsewhere.md",), "run")
assert not set(paths) & wanted
# And "declared nothing" is a third, distinct finding.
assert stress._attributable(bare, []) == ((), "absent")
_GOOD = "krav/N1/id-good.md"
_OTHER = "krav/N1/id-other.md"
def _minibase(root: Path) -> Path:
base = root / "minibase"
(base / "krav" / "N1").mkdir(parents=True)
(base / "index.md").write_text(
"---\nbundle_id: minibase\n---\n\n- [good](krav/N1/id-good.md)\n"
"- [other](krav/N1/id-other.md)\n",
encoding="utf-8",
)
(base / _GOOD).write_text(
'---\ntype: concept\ntitle: "T"\nreq_number: "Krav 1.2.3-4"\n---\n\nBody.\n',
encoding="utf-8",
)
(base / _OTHER).write_text(
'---\ntype: concept\ntitle: "Other"\nreq_number: "Krav 9.9.9-9"\n---\n\nOther.\n',
encoding="utf-8",
)
return base
def _context_dir(root: Path) -> Path:
ctx = root / "ctx"
(ctx / "docs").mkdir(parents=True)
(ctx / "bundle.txt").write_text("name: minibase\nbundle_id: minibase\n", encoding="utf-8")
(ctx / "mandate.json").write_text(
json.dumps(
{
"objective": "o",
"success_criteria": "s",
"approaches": [
{
"id": "a1",
"label": "L",
"affected_codes": ["CODE-1"],
"bundle_id": "minibase",
}
],
}
),
encoding="utf-8",
)
(ctx / "fasit.json").write_text(
json.dumps(
{
"project_id": "proj",
"bundle": "minibase",
"bundle_id": "minibase",
"must_cite": [
{
"approach_id": "a1",
"rationale": "why",
"concepts": [{"path": _GOOD, "title": "T", "ref": "Krav 1.2.3-4"}],
}
],
"must_refuse": [],
"honesty": "synthetic",
}
),
encoding="utf-8",
)
return ctx
def _outbox(root: Path, *, declared: str) -> Path:
out = root / "out"
out.mkdir(parents=True, exist_ok=True)
(out / "r1-a1-proposal.json").write_text(
json.dumps(
{
"proposal": {
"project_id": "proj",
"measure": "m",
"affected_items": [{"code": "CODE-1", "quantity": 1.0, "unit_cost": 2.0}],
"claimed_saving_nok": 1.0,
"assumptions": {},
},
"provenance": {"citations": [], "token_usage": 10},
}
),
encoding="utf-8",
)
(out / "r1-a1-outcome.json").write_text(
json.dumps({"outcome_type": "rejected", "reason": "no"}), encoding="utf-8"
)
(out / "r1-debate.json").write_text(
json.dumps(
{
"run_id": "r1",
"tool_calls": [],
"requirements": [
{"bundle_id": "minibase", "path": declared, "ref": "Krav 9.9.9-9"}
],
}
),
encoding="utf-8",
)
return out
def test_a_declaration_in_the_base_but_not_in_the_fasit_is_not_a_hit(tmp_path: Path) -> None:
"""(h) END-TO-END, and this arm exists because the helper-only one was VACUOUS.
MEASURED: mutation A5-(iv) the judge matching ``concept_names`` instead of this approach's
own ``wanted`` set left the WHOLE suite green (1710/5), because the arm above drives
``_attributable`` while the hit is computed at the call site in ``score_context_set``. The
declaration here names a document the base really does hold; what it is not is the one the
fasit asks for, which is the only distinction a 2 756-document corpus leaves standing.
"""
base = _minibase(tmp_path)
ctx = _context_dir(tmp_path)
miss = stress.score_context_set(ctx, _outbox(tmp_path, declared=_OTHER), "r1", base)
assert miss.approaches[0].requirement_declared == (_OTHER,)
assert miss.approaches[0].requirement_source == "run"
assert miss.approaches[0].requirement_hit is False
assert miss.requirements_declared == (_OTHER,)
# The positive control: a gate that can only come out false proves nothing either.
hit = stress.score_context_set(ctx, _outbox(tmp_path, declared=_GOOD), "r1", base)
assert hit.approaches[0].requirement_hit is True
# ---------------------------------------------------------------------------------------------
# (i) the artefact
# ---------------------------------------------------------------------------------------------
def test_the_declaration_reaches_the_debate_artefact(tmp_path: Path) -> None:
"""(i) A paid run's evidence must survive the process that produced it."""
from portfolio_optimiser import outbox
outbox.write_debate_tools(
str(tmp_path),
"r1",
tool_calls=explore.tool_call_payload(
[ToolCall(name="read_file", bundle_id="b", path="k/a.md")]
),
requirements=explore.requirement_payload(
[DeclaredRequirement(bundle_id="b", path="k/a.md", ref="12.1")]
),
)
payload = json.loads((tmp_path / "r1-debate.json").read_text(encoding="utf-8"))
assert payload["requirements"] == [{"bundle_id": "b", "path": "k/a.md", "ref": "12.1"}]
# An empty list is the honest positive statement every pre-P19 run makes.
outbox.write_debate_tools(str(tmp_path), "r2", tool_calls=[])
assert (
json.loads((tmp_path / "r2-debate.json").read_text(encoding="utf-8"))["requirements"] == []
)

View file

@ -79,7 +79,14 @@ def _manager_script(ledgers: list[str]) -> Callable[[str, str], str]:
def _hypothesis_line(label: str, rationale: str) -> str: def _hypothesis_line(label: str, rationale: str) -> str:
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps({"label": label, "rationale": rationale}) return f"{explore.HYPOTHESIS_MARKER} " + json.dumps(
{
"label": label,
"rationale": rationale,
"requirement": None,
"why_none": "scripted rehearsal: no document was read",
}
)
def _factory(*, ledgers: list[str], hypothesiser: list[str]) -> Callable[[str], BaseChatClient]: def _factory(*, ledgers: list[str], hypothesiser: list[str]) -> Callable[[str], BaseChatClient]:

View file

@ -111,7 +111,14 @@ def _manager_script(ledgers: list[str]) -> Callable[[str, str], str]:
def _hypothesis_line(label: str, rationale: str) -> str: def _hypothesis_line(label: str, rationale: str) -> str:
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps({"label": label, "rationale": rationale}) return f"{explore.HYPOTHESIS_MARKER} " + json.dumps(
{
"label": label,
"rationale": rationale,
"requirement": None,
"why_none": "scripted rehearsal: no document was read",
}
)
def _factory( def _factory(
@ -563,7 +570,7 @@ def test_a_seeded_mandate_still_leads_the_shaped_one() -> None:
A refusal that pointed at a door which did not open would be worse than no message at all. A refusal that pointed at a door which did not open would be worse than no message at all.
""" """
seed = Approach(id="expert-1", label="expert's own", description="the domain expert asked") seed = Approach(id="expert-1", label="expert's own", description="the domain expert asked")
minted = explore._mint_approaches((seed,), [(_LABEL, "shaped in the loop", "")]) minted = explore._mint_approaches((seed,), [(_LABEL, "shaped in the loop", "", None)])
assert [a.id for a in minted] == ["expert-1", "hypothesis-1"] assert [a.id for a in minted] == ["expert-1", "hypothesis-1"]
assert isinstance(Mandate(objective="o", approaches=minted, allow_own_proposals=True), Mandate) assert isinstance(Mandate(objective="o", approaches=minted, allow_own_proposals=True), Mandate)

View file

@ -216,7 +216,17 @@ def _factory(
def _hypothesis_line(label: str, rationale: str) -> str: def _hypothesis_line(label: str, rationale: str) -> str:
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps({"label": label, "rationale": rationale}) # P19 A1: ``requirement`` is a REQUIRED key of a marked line. These stop/ledger tests are not
# about the requirement, so they take the legal explicit-null branch -- which is itself the
# honest shape for a scripted run whose navigator opened nothing.
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps(
{
"label": label,
"rationale": rationale,
"requirement": None,
"why_none": "scripted rehearsal: no document was read",
}
)
#: The no-review base every stop test derives from. ``enable_plan_review`` and #: The no-review base every stop test derives from. ``enable_plan_review`` and

View file

@ -153,6 +153,10 @@ async def test_run_without_mcp_servers_keeps_the_tool_list_unchanged(captured_to
"read_bundle", "read_bundle",
"read_dir", "read_dir",
"read_file", "read_file",
# P19 DEL A: the declaration rung. It is in-process like the other four — the point of
# this arm is that NOTHING here reaches outside the process — and it is created by the two
# caller-owned sinks ``run_project`` passes, never by configuration.
"declare_requirement",
} }

View file

@ -249,7 +249,12 @@ def _factory(
def _hypothesis_line(label: str, rationale: str, bundle_id: str | None = None) -> str: def _hypothesis_line(label: str, rationale: str, bundle_id: str | None = None) -> str:
payload: dict[str, Any] = {"label": label, "rationale": rationale} payload: dict[str, Any] = {
"label": label,
"rationale": rationale,
"requirement": None,
"why_none": "scripted rehearsal: no document was read",
}
if bundle_id is not None: if bundle_id is not None:
payload["bundle_id"] = bundle_id payload["bundle_id"] = bundle_id
return f"{explore.HYPOTHESIS_MARKER} " + json.dumps(payload) return f"{explore.HYPOTHESIS_MARKER} " + json.dumps(payload)

View file

@ -73,7 +73,10 @@ _REPLIES = {
"checker": "VERDICT: APPROVE", "checker": "VERDICT: APPROVE",
"manager": _MANAGER_REPLY, "manager": _MANAGER_REPLY,
"navigator": "NAVIGATOR: read the index.", "navigator": "NAVIGATOR: read the index.",
"hypothesiser": "HYPOTHESIS: " + json.dumps({"label": "Night setback", "rationale": "y"}), "hypothesiser": "HYPOTHESIS: "
+ json.dumps(
{"label": "Night setback", "rationale": "y", "requirement": None, "why_none": "scripted"}
),
} }
_FEEDBACK = "Also test night setback on the ventilation." _FEEDBACK = "Also test night setback on the ventilation."

View file

@ -171,7 +171,10 @@ _REPLIES = {
} }
), ),
"navigator": "NAVIGATOR: read the index.", "navigator": "NAVIGATOR: read the index.",
"hypothesiser": "HYPOTHESIS: " + json.dumps({"label": "Night setback", "rationale": "y"}), "hypothesiser": "HYPOTHESIS: "
+ json.dumps(
{"label": "Night setback", "rationale": "y", "requirement": None, "why_none": "scripted"}
),
} }

View file

@ -31,7 +31,18 @@ from portfolio_optimiser.verdicts import VerdictStore, seed_store_from_bundle
FIXTURE = Path(__file__).parent / "fixtures" / "prepass" / "bygg-energi-mikro-fixture.payload.json" FIXTURE = Path(__file__).parent / "fixtures" / "prepass" / "bygg-energi-mikro-fixture.payload.json"
SHIPPED_BASE = Path(__file__).parent.parent / "shared" / "examples" / "bygg-energi-mikro" SHIPPED_BASE = Path(__file__).parent.parent / "shared" / "examples" / "bygg-energi-mikro"
PROJECT_ID = "BYGG-KONTOR-NORD" PROJECT_ID = "BYGG-KONTOR-NORD"
NAVIGATOR_TOOLS = {"list_bundles", "read_bundle", "read_dir", "read_file"} #: The navigator rungs the debate holds WITHOUT a payload. ``declare_requirement`` joined them in
#: P19 DEL A: it is created by the two sinks ``run_project`` now passes, and it is withdrawn under
#: a payload by exactly the same line that withdraws the other four — which is what this set is
#: here to pin. Named one by one rather than read off ``navigator_tools``: a set imported from the
#: implementation moves with it, and "the debate holds exactly these" is the claim.
NAVIGATOR_TOOLS = {
"list_bundles",
"read_bundle",
"read_dir",
"read_file",
"declare_requirement",
}
_LADDER = "read it with your tools" _LADDER = "read it with your tools"
_EXCERPT_SENTINEL = "SENTINEL-I-ET-LEVERT-UTDRAG" _EXCERPT_SENTINEL = "SENTINEL-I-ET-LEVERT-UTDRAG"

View file

@ -78,7 +78,9 @@ _MANAGER_STAGES = [
_ledger(satisfied=True), _ledger(satisfied=True),
"FINAL: the exploration is done.", "FINAL: the exploration is done.",
] ]
_HYPOTHESIS = "HYPOTHESIS: " + json.dumps({"label": "Night setback", "rationale": "y"}) _HYPOTHESIS = "HYPOTHESIS: " + json.dumps(
{"label": "Night setback", "rationale": "y", "requirement": None, "why_none": "scripted"}
)
_NAVIGATING_SCRIPT = [ _NAVIGATING_SCRIPT = [
{"call": "list_bundles"}, {"call": "list_bundles"},
{"call": "read_bundle", "args": {"bundle_id": DECLARED_ID}}, {"call": "read_bundle", "args": {"bundle_id": DECLARED_ID}},

View file

@ -1483,7 +1483,14 @@ def _resume_replies_file(tmp_path: Path) -> str:
"manager": _RESUME_MANAGER_REPLY, "manager": _RESUME_MANAGER_REPLY,
"navigator": "NAVIGATOR: read the index.", "navigator": "NAVIGATOR: read the index.",
"hypothesiser": "HYPOTHESIS: " "hypothesiser": "HYPOTHESIS: "
+ json.dumps({"label": "Night setback", "rationale": "y"}), + json.dumps(
{
"label": "Night setback",
"rationale": "y",
"requirement": None,
"why_none": "scripted",
}
),
} }
), ),
encoding="utf-8", encoding="utf-8",

View file

@ -160,7 +160,10 @@ def test_explore_with_the_full_five_role_scripted_replies_does_not_crash(tmp_pat
'"next_speaker": {"reason": "r", "answer": "hypothesiser"}, ' '"next_speaker": {"reason": "r", "answer": "hypothesiser"}, '
'"instruction_or_question": {"reason": "r", "answer": "go"}}', '"instruction_or_question": {"reason": "r", "answer": "go"}}',
"navigator": "NAVIGATOR: read the index.", "navigator": "NAVIGATOR: read the index.",
"hypothesiser": "HYPOTHESIS: " + json.dumps({"label": "x", "rationale": "y"}), "hypothesiser": "HYPOTHESIS: "
+ json.dumps(
{"label": "x", "rationale": "y", "requirement": None, "why_none": "scripted"}
),
}, },
) )
@ -318,7 +321,9 @@ _MANAGER_STAGES = [
_ledger(satisfied=True), _ledger(satisfied=True),
"FINAL: the exploration is done.", "FINAL: the exploration is done.",
] ]
_HYPOTHESIS = "HYPOTHESIS: " + json.dumps({"label": "x", "rationale": "y"}) _HYPOTHESIS = "HYPOTHESIS: " + json.dumps(
{"label": "x", "rationale": "y", "requirement": None, "why_none": "scripted"}
)
def _artefact(tmp_path: Path, run_id: str) -> dict[str, Any]: def _artefact(tmp_path: Path, run_id: str) -> dict[str, Any]: