feat(1b): skjemaet ER akseptert av det levende endepunktet - malt, ikke resonnert [skip-docs]
Fase 1b siste trinn: forste bundne levende kjoring over HELE run_project-stien mot Foundry (gpt-4-1-mini). Okt 37s uttalte aerlighets-grense er lukket ved maling: ingen -parse-failures.json i outboksen, altsa parset hvert eneste genererings-svar. assumptions-normaliseringen virket ende-til-ende. Utfall: rejected pa P90 (claimed 34500 > feasible 11488), checker approve. Kjoringen KONKLUDERTE - per pre-registreringen det bestatte utfallet. De to falsifisererne skilte lag for forste gang mot en levende modell. FUNN storre enn den gronne testen: modellen fant opp kostkoden EL-LIGHTING-OP-HR (null treff i kunnskapsbasen). Avvisningen var riktig men skjedde pa 30%-cap-en, ikke stage 0 - bundelen shipper ingen cost-baseline.json, sa S4.0-forankringen var inaktiv. Ko-fort, ikke rettet her. Gate-designet er ovis beslutning, tatt for koding: - TREDJE distinkt opt-in PORTFOLIO_LIVE_FULL_RUN (truthiness, 4b-invarianten). Begge eksisterende live-tester gater pa SAMME to Foundry-variabler, sa gjenbruk ville latt den billige proben fyre den dyre kjoringen - stigen i maleprotokollen ville kollapset til ett trinn. MALT: den dyre SKIPPET med begge Foundry-variablene satt. - Asserten i EN kopi (conftest.assert_full_run_contract, ko-(p)), smal med vilje: fravaer av parse-failures-artefaktet + at validatoren avgjorde. En rejected BESTAR - pastanden er schema-aksept, ikke modell-dommekraft. - Iron Law uten a betale to ganger: diskrimineringen bevist OFFLINE av test_live_full_run_contract.py. To mutasjoner, hver sin signatur: detach artefakt-sjekken (T1 rod ALENE) - raise ubetinget (T2 rod ALENE). STATE-premiss korrigert: "test_foundry_profile_live dekker KUN klient-nivaet" var upresist - test_portfolio_live.py dekket allerede fan-outen, men dens len(runs)==1 kan ikke skille validert fra avvist og bar derfor ikke pastanden. 869 passed / 5 skipped (fra 867/4), ruff+format+mypy rene. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GF7va4cpRiuf79kTzAi3vW
This commit is contained in:
parent
986fc19350
commit
2d1264088e
5 changed files with 408 additions and 0 deletions
146
tests/test_live_full_run_contract.py
Normal file
146
tests/test_live_full_run_contract.py
Normal file
|
|
@ -0,0 +1,146 @@
|
|||
"""The Fase 1b full-run contract must DISCRIMINATE — proven offline, for free (Fase 1b, last step).
|
||||
|
||||
``tests/test_full_run_live.py`` spends real money on a real endpoint, exactly once, and cannot be
|
||||
run red-then-green: the paid call is the MEASUREMENT, not the proof that the measuring instrument
|
||||
works. This file is that proof, and it costs nothing.
|
||||
|
||||
The repo's own rule — *"en test som ikke kan skille to implementasjoner beviser ingenting"* — has
|
||||
been violated by this project's own measuring instruments three times (økt 27's slide sweep read
|
||||
``section``'s 100vh and reported the same number for all eleven slides; økt 37's T3 asserted on a
|
||||
schema the scripted client ignores). An assertion that can only ever pass is the same defect class,
|
||||
and a LIVE assertion is the worst place to discover it, because a green result there is precisely
|
||||
what the operator would act on.
|
||||
|
||||
So both arms drive the SAME ``assert_full_run_contract`` helper the live test uses (ONE copy, in
|
||||
``conftest`` — a second copy would drift, kø-(p)), over the SAME bundle and the SAME project as the
|
||||
live run, through the canonical ``ScriptedChatClient``'s ``reply_selector`` seam:
|
||||
|
||||
* T1 — a run in which ONE reply failed to parse must make the contract FAIL. This is the arm that
|
||||
matters: it is the offline stand-in for "the live endpoint ignored ``response_format``", which is
|
||||
the single outcome the paid run exists to rule out. Per the repo rule for negative asserts, it
|
||||
first PROVES the event happened (the artefact exists) rather than inferring it from the failure.
|
||||
* T2 — the CONTROL: a run in which every reply parsed must make the contract PASS. Without it, a
|
||||
helper that raised unconditionally would satisfy T1 and the live test could then only ever be red,
|
||||
which is the mirror-image vacuity.
|
||||
|
||||
Together they show the contract keys on the ACTUAL discriminator (the parse-failure artefact), not
|
||||
on something both runs share.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import Callable
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from conftest import assert_full_run_contract
|
||||
|
||||
from portfolio_optimiser.budget import Budget, TokenMeter
|
||||
from portfolio_optimiser.run import RunResult, run_project
|
||||
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||
from portfolio_optimiser.verdicts import VerdictStore
|
||||
|
||||
BUNDLE_DIR = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
|
||||
_PROJECT_ID = "BYGG-KONTOR-NORD"
|
||||
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (contract probe)"}
|
||||
|
||||
#: The line ``generate._build_messages`` puts in EVERY generation prompt and nowhere else — the one
|
||||
#: identifier separating a generation call from a debate turn (borrowed from
|
||||
#: ``test_parse_failure_capture_loadbearing``, whose seam this file shares).
|
||||
_GENERATION_MARK = "Respond with ONLY a JSON object"
|
||||
|
||||
#: Prose where an object was requested: what a model that ignored ``response_format`` returns.
|
||||
_MALFORMED = (
|
||||
"Sure! Here is what I found for this project.\n\n"
|
||||
"The main opportunity looks like demand-controlled lighting, worth roughly 30 000 NOK.\n"
|
||||
"Let me know if you want that as JSON."
|
||||
)
|
||||
|
||||
#: BYGG-KONTOR-NORD: affected total 300000 x 1.0 -> degenerate Monte Carlo P90 = 90000, so a claim
|
||||
#: of 30000 validates (same arithmetic as ``test_parse_failure_capture_loadbearing``).
|
||||
_VALID_REPLY = (
|
||||
'{"measure":"Behovsstyrt belysning i fellesarealer","affected_items":'
|
||||
'[{"code":"ENERGI-TOTAL-EL","quantity":300000,"unit_cost":1.0}],'
|
||||
'"claimed_saving_nok":30000}'
|
||||
)
|
||||
|
||||
|
||||
def _malformed_then_valid(failures: int) -> Callable[[str, str], str]:
|
||||
"""Fail to parse ``failures`` times, then answer with a proposal that validates. Keyed on the
|
||||
GENERATION prompt only, so debate turns — which are never parsed — do not consume the counter."""
|
||||
seen = {"n": 0}
|
||||
|
||||
def _select(blob: str, _role: str) -> str:
|
||||
if _GENERATION_MARK not in blob:
|
||||
return "ok"
|
||||
seen["n"] += 1
|
||||
return _MALFORMED if seen["n"] <= failures else _VALID_REPLY
|
||||
|
||||
return _select
|
||||
|
||||
|
||||
async def _run(select: Callable[[str, str], str], outbox_dir: Path, run_id: str) -> RunResult:
|
||||
def factory(role: str) -> ScriptedChatClient:
|
||||
return ScriptedChatClient(role=role, reply_selector=select, default_reply="ok")
|
||||
|
||||
result = await run_project(
|
||||
_PROJECT_ID,
|
||||
"local",
|
||||
docs_dir=str(BUNDLE_DIR),
|
||||
bundle_dir=str(BUNDLE_DIR),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
store=VerdictStore(verdicts=[]),
|
||||
client_factory=factory,
|
||||
outbox_dir=str(outbox_dir),
|
||||
run_id=run_id,
|
||||
meter=TokenMeter(Budget(max_tokens=10**9, max_rounds=8)),
|
||||
)
|
||||
assert isinstance(result, RunResult)
|
||||
return result
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------------------------
|
||||
# T1 — the arm that matters: an unparsed reply must FAIL the contract.
|
||||
# --------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
async def test_contract_fails_when_a_reply_did_not_parse(tmp_path: Path) -> None:
|
||||
"""The offline stand-in for "the live endpoint ignored the schema". The run itself SUCCEEDS —
|
||||
the second reply validates — so the contract cannot be keying on the run's outcome; the only
|
||||
thing separating this from T2 is the parse-failure artefact.
|
||||
|
||||
RED (i.e. this test fails) on a contract that ignores the artefact and merely checks that a
|
||||
RunResult came back — which is what ``test_portfolio_live.py``'s ``len(runs) == 1`` does, and
|
||||
the reason that existing gated test could not carry this claim."""
|
||||
outbox_dir = tmp_path / "outbox"
|
||||
run_id = "contract-dirty"
|
||||
|
||||
result = await _run(_malformed_then_valid(1), outbox_dir, run_id)
|
||||
|
||||
# Prove the event happened FIRST — a negative assert that merely observes a failure cannot tell
|
||||
# "the contract caught the artefact" from "the contract failed for some unrelated reason".
|
||||
artefact = outbox_dir / f"{run_id}-parse-failures.json"
|
||||
assert artefact.exists(), "precondition: this arm must actually produce a parse failure"
|
||||
assert result.provenance.validator_decision in {"validated", "rejected"}, (
|
||||
"precondition: the run must otherwise CONCLUDE, so the artefact is the only difference"
|
||||
)
|
||||
|
||||
with pytest.raises(AssertionError, match="did NOT honour the structured schema"):
|
||||
assert_full_run_contract(result, outbox_dir, run_id)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------------------------
|
||||
# T2 — the CONTROL: a clean run must PASS.
|
||||
# --------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
async def test_contract_passes_when_every_reply_parsed(tmp_path: Path) -> None:
|
||||
"""Without this control, a helper that raised unconditionally would satisfy T1, and the live
|
||||
test could then only ever be red — the mirror image of a check that can only ever be green."""
|
||||
outbox_dir = tmp_path / "outbox"
|
||||
run_id = "contract-clean"
|
||||
|
||||
result = await _run(_malformed_then_valid(0), outbox_dir, run_id)
|
||||
|
||||
assert not (outbox_dir / f"{run_id}-parse-failures.json").exists()
|
||||
assert_full_run_contract(result, outbox_dir, run_id)
|
||||
Loading…
Add table
Add a link
Reference in a new issue