The egress declaration (Trekk B3) says what a run MAY contact. It cannot say what it DID: after the run, nothing distinguished "the agents queried the price register" from "the agents ignored it", and a proposal resting on an external service should be traceable to it. ToolCallRecorder(FunctionMiddleware) mirrors BudgetMiddleware(ChatMiddleware) one layer down — that one observes the debate's chat calls, this one its tool calls. It observes only: call_next is always awaited, so a trace can never alter the run it traces. The record lands on ProvenanceStamp.external_calls, read AFTER the debate so it is a record rather than an intention. MEASURED, not assumed, before any of it was written: FunctionMiddleware fires for a tool served over a REAL MCP stdio subprocess, and context.function.name carries the BARE tool name with no server prefix. That measurement decided the design — MAF cannot tell us which server a tool came from, so attribution comes from our own config, and a name allowed by two servers is recorded UNATTRIBUTED (server="") rather than credited to the first match. Naming a service that may never have been contacted is the one place a guess must not go. Only CONFIGURED tools are recorded. The middleware fires for every function the agents invoke, including the in-process retrieve_cost_docs on the road path; logging those would turn the record into a false egress claim. An empty list is a positive statement — nothing outside this process was contacted — which is why it is always serialized rather than omitted. Honesty limit, written on ExternalCall itself: this is the call and its source. It is NOT evidence that the service's answer reached the proposal, nor a verified rendering of that answer. One finding, and it is the reason for measuring rather than trusting green: the road-path negative test was VACUOUS. Its scripted tool call named an argument the tool does not declare (code vs query), MAF rejected the call before invocation, and the test asserted an empty record against a run where no tool ran at all — green under the exact mutation it existed to catch. It now spies on the recorder and asserts the invocation genuinely reached it before asserting it was not recorded. This is last session's lesson again: a scenario that cannot distinguish two implementations proves nothing. The tool-call double is registered in test_scripted_client_consolidation.py's _DELEGATING_OVERRIDES — it cannot live in the reply_selector seam, which returns a reply STRING, and a response that is not text is its whole subject. Load-bearing MEASURED (tests/test_b4_mcp_call_trace_loadbearing.py) against the whole 755-test suite, four mutations all red: detach the recorder from the debate middleware · record every function invocation · attribute an ambiguous name to the first server · stop reading the recorder into provenance. Control: a run with no configured servers records nothing, so the empty record is a real answer and not the only one the seam can produce. Ran it, not just tested it: the real recorder against a real MCP server subprocess returns ExternalCall(server='prisregister', tool='lookup_unit_price'), and a scripted CLI run's outbox artefact carries the empty list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VtRd8y1PDPGwkrRXFhubqr
80 lines
3.3 KiB
Python
80 lines
3.3 KiB
Python
"""First-class Pydantic provenance stamp.
|
|
|
|
Provenance is authoritative framework data — **independent** of MAF's ``Annotation`` type,
|
|
which silently drops on the Python streaming path (#4316, research 02 Dim 2). A
|
|
``ProvenanceStamp`` is a Pydantic model that must carry at least one ``Citation`` (a
|
|
``min_length=1`` constraint), the model + role that produced the proposal, the validator's
|
|
decision, and the token usage. ``to_annotations()`` maps to MAF ``Annotation`` dicts for
|
|
DISPLAY only — never the source of truth.
|
|
|
|
The locator type (``TextSpan``) is owned by ``retrieval.py`` (Step 5) and imported here, so
|
|
a citation's span is the same exact object the retriever produced.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from typing import Literal
|
|
|
|
from agent_framework import Annotation, TextSpanRegion
|
|
from pydantic import BaseModel, Field
|
|
|
|
from portfolio_optimiser.retrieval import TextSpan
|
|
|
|
|
|
class Citation(BaseModel):
|
|
"""One exact-span citation into a source document (locator owned by retrieval.py)."""
|
|
|
|
file: str
|
|
locator: TextSpan
|
|
snippet: str
|
|
|
|
|
|
class ExternalCall(BaseModel):
|
|
"""One external service call a run actually made (Trekk B4).
|
|
|
|
**What this is evidence of, and what it is not.** It records that ``tool`` was invoked and which
|
|
configured ``server`` it belongs to. It is NOT evidence that the service's answer reached the
|
|
proposal, and it is not a verified rendering of what the service returned — the framework hands
|
|
the answer to the agent, and what the agent does with it is the agent's. Reading this as "the
|
|
figure came from the price register" would claim more than the record supports.
|
|
|
|
``server`` is ``""`` when the tool name cannot be attributed to exactly one configured server.
|
|
MEASURED against a real MCP stdio subprocess: MAF passes the BARE tool name to function
|
|
middleware, with no server prefix, so two servers exposing one tool name are indistinguishable
|
|
at this seam. Unattributed is the honest answer there; naming the first match would put a
|
|
service in the record that may never have been contacted.
|
|
"""
|
|
|
|
server: str
|
|
tool: str
|
|
|
|
|
|
class ProvenanceStamp(BaseModel):
|
|
"""Authoritative provenance for one proposal — at least one citation is mandatory."""
|
|
|
|
citations: list[Citation] = Field(min_length=1)
|
|
model: str
|
|
role: str
|
|
validator_decision: Literal["validated", "rejected"]
|
|
token_usage: int
|
|
#: External service calls the run made (B4). EMPTY is a positive statement — "nothing outside
|
|
#: this process was contacted" — not an absent field, which is why it is always serialized.
|
|
external_calls: list[ExternalCall] = Field(default_factory=list)
|
|
|
|
def to_annotations(self) -> list[Annotation]:
|
|
"""Map to MAF ``Annotation`` dicts for display only (NOT the source of truth)."""
|
|
return [
|
|
Annotation(
|
|
type="citation",
|
|
snippet=c.snippet,
|
|
annotated_regions=[
|
|
TextSpanRegion(
|
|
type="text_span",
|
|
start_index=c.locator.start_index,
|
|
end_index=c.locator.end_index,
|
|
)
|
|
],
|
|
additional_properties={"file": c.file},
|
|
)
|
|
for c in self.citations
|
|
]
|