test(cost): the money rule exists in four copies, and the gate it serves is switched off
D7 mirror candidate (p) `to_ore`, measured offline with the mutation harness. `to_ore` does not exist here (0 of 80 .py files in src+tests; positive control: `unit_cost` found in 19). Our money axis is the accumulated USD spend, and its conversion is the rounding to six decimals — present in FOUR literal copies with no named source (run.py:159, run_s10.py:110 and :130, costsim.py:125) while the share rounding one file over DOES have one (`_SHARE_DIGITS`). The sibling's drift form is present. Three of the four copies are never executed: replacing the whole expression with `999.0` left all 984 green at each. Their green under a detach was never evidence about the rounding — it was "not measured" (ansikt 4 on the apparatus). The default branch WAS pinned; the value branch was not. The one path that decides with the cost — the C3.5 pre-call USD belt — is a permanent no-op: `max_cost_usd` is set in 0 of 27 src modules (positive control: `max_budget_usd_per_call` is found by the same scan). The gate is built; no path hands it a cap. Seven mutations, all VALUE-PROVED with collateral controls green in both runs. No src change: folding the copies is a refactor, wiring the run-total cap is a feature. The tests pin today's boundary so neither lands silently. 984 -> 997 (strict superset, 0 lost node ids). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
b955bf274e
commit
7b934275d8
2 changed files with 357 additions and 2 deletions
|
|
@ -35,12 +35,12 @@ spørringen som ble kjørt** — ikke fila den ble kjørt mot.
|
||||||
## D7-speilingskøen
|
## D7-speilingskøen
|
||||||
|
|
||||||
Åtte kandidater for speiling mellom D7-søsknene. **Ingen er besluttet** — de står som
|
Åtte kandidater for speiling mellom D7-søsknene. **Ingen er besluttet** — de står som
|
||||||
kandidater, ikke som planlagt arbeid. **Fem er målt, 3 står igjen:**
|
kandidater, ikke som planlagt arbeid. **Seks er målt, 2 står igjen:**
|
||||||
|
|
||||||
- ~~S2.7~~ — **MÅLT 2026-09-07, se under**
|
- ~~S2.7~~ — **MÅLT 2026-09-07, se under**
|
||||||
- ~~S3.2~~ — **MÅLT 2026-09-12, se under**
|
- ~~S3.2~~ — **MÅLT 2026-09-12, se under**
|
||||||
- ~~S4.0~~ (`126807a`) — **MÅLT 2026-09-13, se under**
|
- ~~S4.0~~ (`126807a`) — **MÅLT 2026-09-13, se under**
|
||||||
- (p) `to_ore` — TO kallsteder
|
- ~~(p) `to_ore`~~ — **MÅLT 2026-09-13, se under**
|
||||||
- ~~(a)/(i) `unquote_scalar`~~ — **MÅLT 2026-08-31, se under**
|
- ~~(a)/(i) `unquote_scalar`~~ — **MÅLT 2026-08-31, se under**
|
||||||
- `mandate.py`
|
- `mandate.py`
|
||||||
- A5 — TO halvdeler
|
- A5 — TO halvdeler
|
||||||
|
|
@ -353,6 +353,100 @@ søm-bevis, og er merket som det over.
|
||||||
konvergens.
|
konvergens.
|
||||||
|
|
||||||
|
|
||||||
|
### (p) `to_ore` — regelen finnes i FIRE kopier, tre av dem kjøres aldri, og gaten den skulle tjene er slått av (målt 2026-09-13)
|
||||||
|
|
||||||
|
Kandidaten står i køen fordi søskenet vokste en navngitt `to_ore` etter at en **pengekonvertering
|
||||||
|
hadde driftet i to kopier** — `run.py` re-implementerte `ledger.to_ore` privat, og de to kopiene
|
||||||
|
møttes på hver sin side av én målsammenligning (`_goal_limit_if_reached`). Søskenets beslutning ble
|
||||||
|
«kvantiser per linje, summer heltall», pinnet med `assert run_mod.to_ore is ledger_mod.to_ore`.
|
||||||
|
Speilings-spørsmålet er derfor ikke «har vi funksjonen», men: **er vår pengekonvertering ÉN regel,
|
||||||
|
og er den load-bearing der den beslutter noe?**
|
||||||
|
|
||||||
|
**Populasjonen først.** `to_ore` finnes ikke hos oss: **0 treff av 80 undersøkte `.py`-filer** i
|
||||||
|
`src`+`tests` (positiv kontroll: samme spørring finner `unit_cost` i 19 av dem, så spørringen KAN
|
||||||
|
finne). Vi har ingen NOK→øre-akse i det hele tatt — vår pengeakse er den akkumulerte
|
||||||
|
USD-forbruket `total_cost_usd`, og vår konvertering er avrundingen til seks desimaler. Den finnes
|
||||||
|
i **fire literale kopier uten navngitt kilde**: `run.py:159` (`_client_cost_usd`), `run_s10.py:110`
|
||||||
|
og `:130` (de to persisterings-kallstedene) og `costsim.py:125` (per-celle-estimatet). Én fil
|
||||||
|
unna gjør det samme pakket det motsatte: `valuereport.py` avrunder gjennom `_SHARE_DIGITS`.
|
||||||
|
**Søskenets drift-form finnes altså her** — i den halvdelen som ikke har fått sin ene kilde.
|
||||||
|
|
||||||
|
**Og driften har et møtested.** På en `cost_usd`-stopp skriver ETT kall begge tall: `stop.json`
|
||||||
|
bærer `observed` fra `BudgetExceeded` — den **rå** verdien gaten sammenlignet — mens `usage.json`
|
||||||
|
bærer `cost_usd` **avrundet**. To tall for samme størrelse, fra samme kall, etter to regler.
|
||||||
|
|
||||||
|
**Men det skarpeste funnet er at tre av de fire kopiene aldri kjøres.** Med
|
||||||
|
`scripts/mutation_harness.py`, nevner `tests/` (hele den gamle suiten, 984 tester), hver kjøring
|
||||||
|
sha256-restaurert fra disk — og med positivkontrollen først: å erstatte HELE avrundingsuttrykket
|
||||||
|
med konstanten `999.0` lot **alle 984 stå grønne** på `run.py:159`, `run_s10.py:110` og `:130`.
|
||||||
|
Grenene utøves aldri med en kost. Deres grønne under en detach var derfor aldri bevis om
|
||||||
|
avrundingen — det var «ikke målt» (økt 39, Verifiseringsloven ansikt 4 på måleapparatet). Kun
|
||||||
|
`costsim.py` var dekket. **DEFAULT-en var derimot pinnet:** `getattr(client, "total_cost_usd",
|
||||||
|
None)` → `0.0` er rød i `test_run_entrance_loadbearing.py`. Nøkkelen var pinnet i sitt navn og fri
|
||||||
|
i sin verdi (økt 41) — her: default-grenen pinnet, verdi-grenen ukjørt.
|
||||||
|
|
||||||
|
**Gaten som skulle bruke tallet er slått av på hver eneste sti.** `max_cost_usd` settes i **0 av
|
||||||
|
27 `src`-moduler** (positiv kontroll: samme AST-skann finner `max_budget_usd_per_call`, som ER
|
||||||
|
wiret). Alle fire `BudgetMeter(...)`-konstruksjoner i `src` tar default `None`, og
|
||||||
|
`guard_before_call` returnerer før den sammenligner. C3.5-beltet er altså ikke-no-op **kun i sin
|
||||||
|
egen enhetstest**. Samme klasse som S4.0, ett hakk verre: der var feltet lastet og utenfor
|
||||||
|
dommerens rekkevidde; her er dommeren bygget og ingen sti gir den et tak.
|
||||||
|
|
||||||
|
**SYV mutasjoner, alle VALUE-PROVED** (grønn før / rød etter, hver med kollateral-kontroll grønn i
|
||||||
|
begge kjøringer); anker → erstatning ordrett:
|
||||||
|
|
||||||
|
- **M1 — avrundingen detaches** (`run.py`): `return None if cost is None else round(float(cost), 6)`
|
||||||
|
→ `return None if cost is None else float(cost)`. To RØDE (kvantiseringen og «ikke den rå
|
||||||
|
verdien», splittet); kontroll `test_the_value_branch_is_reached_at_all` GRØNN i begge.
|
||||||
|
- **M2 — sifferantallet drifter** (`run.py`): `return None if cost is None else round(float(cost), 6)`
|
||||||
|
→ `return None if cost is None else round(float(cost), 2)`. Kvantiserings-testen RØD; samme
|
||||||
|
kontroll GRØNN.
|
||||||
|
- **M3 — beltet leser den AVRUNDEDE kosten** (`loop.py`, søskenets drift innført):
|
||||||
|
`meter.guard_before_call(spent_usd)` → `meter.guard_before_call(round(spent_usd, 6))`.
|
||||||
|
Stopp-testen RØD; kontrollen «den avrundede ville ikke ha stoppet» GRØNN i begge.
|
||||||
|
- **M4 — stop-eventet rapporterer den avrundede** (`budget.py`):
|
||||||
|
`raise BudgetExceeded("cost_usd", self._max_cost_usd, spent_usd)` →
|
||||||
|
`raise BudgetExceeded("cost_usd", self._max_cost_usd, round(spent_usd, 6))`. `observed`-testen
|
||||||
|
RØD; at beltet fortsatt STOPPER GRØNN — rødheten attribuerer til rapporteringen, ikke til gaten.
|
||||||
|
- **M5 — én kopi drifter til et annet sifferantall** (`costsim.py`):
|
||||||
|
`cost = round(estimated_tokens * price.usd_per_mtok / _TOKENS_PER_MTOK, 6)` → samme med `, 2)`.
|
||||||
|
Siffer-settet RØDT; kopi-tellingen GRØNN (mutasjonen endrer siffer, ikke antall).
|
||||||
|
- **M6 — én kopi foldes bort** (`run_s10.py`, fiksens retning):
|
||||||
|
` cost_usd=round(client.total_cost_usd, 6),` → ` cost_usd=client.total_cost_usd,`.
|
||||||
|
Kopi-tellingen RØD; siffer-settet og modul-settet GRØNNE.
|
||||||
|
- **M7 — noen wirer run-total-taket** (`run.py`, ratchet-en): ` run_label = run_id or out_dir.name`
|
||||||
|
+ ` meter = BudgetMeter(contracts.termination)` → samme med
|
||||||
|
` meter = BudgetMeter(contracts.termination, max_cost_usd=1.0)`. Fraværs-testen RØD;
|
||||||
|
positivkontrollen og «en meter uten tak stopper aldri» GRØNNE.
|
||||||
|
|
||||||
|
**Fixturen fiksen LESER, bygget** (økt 41). `ScriptedClient` bærer ingen `total_cost_usd` i det
|
||||||
|
hele tatt, og det er nettopp derfor hver eksisterende run-sti-test bare utøver den ærlige
|
||||||
|
null-grenen. `CostingScriptedClient` legger på en — 0.1234567891 USD, mer presisjon enn regelen
|
||||||
|
beholder, så den persisterte verdien forteller hvilket tall som ble skrevet. Belte-testene bruker
|
||||||
|
en forbruks-verdi som ligger STRENGT MELLOM taket og sin egen avrunding (0.10000004 mot tak 0.1,
|
||||||
|
`round(…, 6) == 0.1`): rå verdi stopper, avrundet slipper gjennom. Uten den konstruksjonen er M3 en
|
||||||
|
no-op og ville sett ut som en udekket søm.
|
||||||
|
|
||||||
|
**Pinnet av** `tests/test_cost_usd_quantization_loadbearing.py` (13 tester, 984 → 997, strengt
|
||||||
|
supersett: 0 tapte node-id-er). Tre testpar er SPLITTET fordi en rød test bare beviser sin FØRSTE
|
||||||
|
assert: «kvantisert til seks desimaler» fra «ikke den rå verdien», «beltet stopper» fra «stoppet
|
||||||
|
bærer den rå verdien», og siffer-settet fra kopi-tellingen.
|
||||||
|
|
||||||
|
**Ærlig grense — hva dette IKKE sier.** Ingen `src/`-endring er gjort, og det er en beslutning, ikke
|
||||||
|
en forglemmelse. Å gi de fire kopiene én navngitt kilde er en refaktorering, ikke en fiks på en målt
|
||||||
|
defekt; å wire run-total-taket inn i inngangen er en FEATURE — et CLI-flagg og et kontraktsfelt, med
|
||||||
|
`--help`/README-paritetsvakten bak seg — og per-kall-taket `max_budget_usd_per_call` binder allerede
|
||||||
|
et live forbruk gjennom SDK-en. Testene sier heller ikke at dagens tilstand er ønskelig: de sier at
|
||||||
|
den er MÅLT, og de gjør hver av de tre endringene synlig som en rød test i stedet for en stille
|
||||||
|
landing. Kopi-tellingen er dessuten en tekstlig AST-egenskap, ikke et oppførselsbevis — den fanger
|
||||||
|
en femte kopi og et endret sifferantall, ikke en femte kopi skrevet på en annen form (`f"{x:.6f}"`,
|
||||||
|
`Decimal`). Og `run_s10.py` er byte-frossen og importeres aldri i suiten: dens to kallsteder er
|
||||||
|
pinnet ved AST, ikke ved kjøring, hvilket er den sterkeste sømmen som finnes mot en fil vi ikke
|
||||||
|
kjører.
|
||||||
|
|
||||||
|
**Datering (D7-rammen):** arbeid ETTER 2026-08-09 — skal **ikke** leses som uavhengig konvergens.
|
||||||
|
|
||||||
|
|
||||||
Rammen rundt køen: å lese søskenets kode er tillatt (`3bdf7f0`), men kopiering skal kun skje
|
Rammen rundt køen: å lese søskenets kode er tillatt (`3bdf7f0`), men kopiering skal kun skje
|
||||||
der det tjener løsningen, aldri som snarvei. **Uavhengighets-beviset er DATERT** t.o.m.
|
der det tjener løsningen, aldri som snarvei. **Uavhengighets-beviset er DATERT** t.o.m.
|
||||||
2026-08-09; arbeid etter den datoen kan ikke leses som uavhengig konvergens.
|
2026-08-09; arbeid etter den datoen kan ikke leses som uavhengig konvergens.
|
||||||
|
|
|
||||||
261
tests/test_cost_usd_quantization_loadbearing.py
Normal file
261
tests/test_cost_usd_quantization_loadbearing.py
Normal file
|
|
@ -0,0 +1,261 @@
|
||||||
|
"""Kø-(p) LOAD-BEARING: the USD cost conversion — one rule, four copies, three of them never run.
|
||||||
|
|
||||||
|
The sibling's ``to_ore`` entered the mirror queue because a money conversion had DRIFTED into
|
||||||
|
two copies under two rules, and the two copies met on the two sides of one goal comparison.
|
||||||
|
The mirrored claim is therefore not "do we have the function" but: **is our money conversion
|
||||||
|
one rule, and is that rule load-bearing where it decides something?**
|
||||||
|
|
||||||
|
Measured 2026-09-13. ``to_ore`` does not exist here (0 hits across 80 ``.py`` files in
|
||||||
|
``src``+``tests``; positive control: the same query finds ``unit_cost`` in 19). Our money axis
|
||||||
|
is not NOK->øre but the accumulated USD spend, ``total_cost_usd``, and its conversion is the
|
||||||
|
rounding to six decimals. That rounding exists in FOUR literal copies with no named source —
|
||||||
|
``run.py``'s ``_client_cost_usd``, ``run_s10.py``'s two persistence call sites, and
|
||||||
|
``costsim.py``'s per-cell estimate — while the sibling rule one file over (share rounding in
|
||||||
|
``valuereport.py``) DOES have a named source, ``_SHARE_DIGITS``. The sibling's drift form is
|
||||||
|
present here.
|
||||||
|
|
||||||
|
What the harness then showed is sharper than "pinned only at the edge". Replacing the WHOLE
|
||||||
|
rounding expression with the constant ``999.0`` left all 984 tests green at three of the four
|
||||||
|
sites (``run.py``:159, ``run_s10.py``:110 and :130) — those branches are never executed with a
|
||||||
|
cost at all, so their green under a detach was never evidence about the rounding (økt 39,
|
||||||
|
Verifiseringsloven ansikt 4 applied to the measuring apparatus). Only ``costsim.py`` was
|
||||||
|
covered. The default branch WAS pinned (``getattr(..., None)`` -> ``0.0`` is red in
|
||||||
|
``test_run_entrance``) — the key was pinned in its name and free in its value (økt 41).
|
||||||
|
|
||||||
|
And the one path that DECIDES with the cost, the C3.5 pre-call USD belt, is a permanent no-op:
|
||||||
|
``max_cost_usd`` is set in 0 of 27 ``src`` files (positive control: ``max_budget_usd_per_call``,
|
||||||
|
which IS wired, is found by the same scan). All four ``BudgetMeter(...)`` constructions in
|
||||||
|
``src`` take the default ``None``, so ``guard_before_call`` returns before it ever compares.
|
||||||
|
The gate is built; no path hands it a cap.
|
||||||
|
|
||||||
|
This file pins today's boundary. It makes NO ``src`` change: wiring a run-total USD cap into the
|
||||||
|
entrance is a feature with a CLI flag and a contract field (and the README/``--help`` parity
|
||||||
|
guard behind it), not a fix to a measured defect, and the per-call SDK cap
|
||||||
|
(``max_budget_usd_per_call``) already bounds live spend. What the tests forbid is the change
|
||||||
|
landing SILENTLY — and the rounding rule growing a fifth copy, or drifting to a different digit
|
||||||
|
count, without anyone noticing.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import ast
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Callable, Iterator
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
from _scripted import ScriptedClient, reply
|
||||||
|
|
||||||
|
from portfolio_optimiser_claude.budget import BudgetExceeded, BudgetMeter
|
||||||
|
from portfolio_optimiser_claude.contracts import Contracts, TerminationContract
|
||||||
|
from portfolio_optimiser_claude.ir import load_validator_input
|
||||||
|
from portfolio_optimiser_claude.loop import ModelClient, ModelReply, _guarded_complete
|
||||||
|
from portfolio_optimiser_claude.run import main
|
||||||
|
|
||||||
|
BUNDLE = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
|
||||||
|
SRC = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser_claude"
|
||||||
|
|
||||||
|
# A spend with MORE precision than the rule keeps: the rounded value and the raw value are
|
||||||
|
# different numbers, so a test that reads the persisted figure can tell which one was written.
|
||||||
|
_RAW_SPEND = 0.1234567891
|
||||||
|
_ROUNDED_SPEND = 0.123457 # round(_RAW_SPEND, 6)
|
||||||
|
|
||||||
|
# A spend that sits strictly between the cap and its own rounding: raw > cap (the belt must
|
||||||
|
# stop) but round(raw, 6) == cap (a belt reading the ROUNDED figure would let the call through).
|
||||||
|
_CAP_USD = 0.1
|
||||||
|
_OVERSPEND_RAW = 0.10000004
|
||||||
|
|
||||||
|
|
||||||
|
class CostingScriptedClient(ScriptedClient):
|
||||||
|
"""A scripted stand-in that DOES account USD — the fixture the rule reads.
|
||||||
|
|
||||||
|
``ScriptedClient`` carries no ``total_cost_usd`` at all, which is why every existing run-path
|
||||||
|
test exercises only the honest-null branch of ``_client_cost_usd``. Honesty rule (§1): this
|
||||||
|
is still a scripted stand-in, not a model; the cost is a fixture value, not a provider
|
||||||
|
figure.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, replies: list[ModelReply], *, total_cost_usd: float) -> None:
|
||||||
|
super().__init__(replies=replies)
|
||||||
|
self.total_cost_usd = total_cost_usd
|
||||||
|
|
||||||
|
|
||||||
|
def _costing_factory(
|
||||||
|
replies: list[ModelReply], *, total_cost_usd: float
|
||||||
|
) -> Callable[[Contracts, float], ModelClient]:
|
||||||
|
def factory(contracts: Contracts, max_budget_usd_per_call: float) -> ModelClient:
|
||||||
|
return CostingScriptedClient(list(replies), total_cost_usd=total_cost_usd)
|
||||||
|
|
||||||
|
return factory
|
||||||
|
|
||||||
|
|
||||||
|
def _completed_run_replies() -> list[ModelReply]:
|
||||||
|
return [
|
||||||
|
reply("debate reasoning"),
|
||||||
|
reply("VERDICT: APPROVE"),
|
||||||
|
reply(json.dumps(load_validator_input(BUNDLE).model_dump())),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def _usage_after_run(tmp_path: Path, *, total_cost_usd: float) -> dict[str, object]:
|
||||||
|
out = tmp_path / "out"
|
||||||
|
code = main(
|
||||||
|
["--bundle", str(BUNDLE), "--out", str(out)],
|
||||||
|
client_factory=_costing_factory(_completed_run_replies(), total_cost_usd=total_cost_usd),
|
||||||
|
)
|
||||||
|
assert code == 0, "the fixture must reach the completed-run persistence branch"
|
||||||
|
payload: dict[str, object] = json.loads((out / "usage.json").read_text("utf-8"))
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
class TestTheRoundingRuleIsActuallyApplied:
|
||||||
|
"""The value branch of ``_client_cost_usd`` — never executed by the suite before this file.
|
||||||
|
|
||||||
|
Each assertion is its own test: a red test proves only its FIRST assert, and "the figure is
|
||||||
|
rounded" and "the figure is not the raw value" are two claims about two different mutations.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def test_persisted_cost_is_quantized_to_six_decimals(self, tmp_path: Path) -> None:
|
||||||
|
usage = _usage_after_run(tmp_path, total_cost_usd=_RAW_SPEND)
|
||||||
|
assert usage["cost_usd"] == _ROUNDED_SPEND
|
||||||
|
|
||||||
|
def test_persisted_cost_is_not_the_raw_accumulated_spend(self, tmp_path: Path) -> None:
|
||||||
|
# The discriminating half: a detached rounding persists the raw float, which passes the
|
||||||
|
# test above only if that test is written as an approximate comparison. It is not.
|
||||||
|
usage = _usage_after_run(tmp_path, total_cost_usd=_RAW_SPEND)
|
||||||
|
assert usage["cost_usd"] != _RAW_SPEND
|
||||||
|
|
||||||
|
def test_the_value_branch_is_reached_at_all(self, tmp_path: Path) -> None:
|
||||||
|
# Positive control on the apparatus (ansikt 4): before this file, replacing the whole
|
||||||
|
# expression with a constant left all 984 green, because the branch never ran with a
|
||||||
|
# cost. This pins that a costing client now reaches it — without it, the two tests above
|
||||||
|
# could pass vacuously on a None.
|
||||||
|
usage = _usage_after_run(tmp_path, total_cost_usd=_RAW_SPEND)
|
||||||
|
assert usage["cost_usd"] is not None
|
||||||
|
|
||||||
|
|
||||||
|
class TestTheBeltDecidesOnTheRawSpend:
|
||||||
|
"""What the code COMPUTES WITH is not what it PERSISTS — and the gate must read the raw one.
|
||||||
|
|
||||||
|
Rounding before a comparison moves the threshold: a run that has crossed its cap by less
|
||||||
|
than half a micro-dollar would read as not-crossed. The sibling's defect was exactly this
|
||||||
|
shape — two sides of one money comparison computed under two rules.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def _meter(self) -> BudgetMeter:
|
||||||
|
return BudgetMeter(
|
||||||
|
TerminationContract(max_rounds=1000, max_tokens=10_000), max_cost_usd=_CAP_USD
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_the_belt_stops_on_a_spend_its_own_rounding_would_hide(self) -> None:
|
||||||
|
client = CostingScriptedClient([reply("never reached")], total_cost_usd=_OVERSPEND_RAW)
|
||||||
|
with pytest.raises(BudgetExceeded) as caught:
|
||||||
|
_guarded_complete(client, "prompt", role="proposer", meter=self._meter())
|
||||||
|
assert caught.value.kind == "cost_usd"
|
||||||
|
|
||||||
|
def test_the_stop_event_carries_the_raw_spend_not_the_rounded_one(self) -> None:
|
||||||
|
# The structured stop is the artifact's own record of what the gate saw (§8). If the
|
||||||
|
# belt ever reads a rounded figure, this is the number that changes.
|
||||||
|
client = CostingScriptedClient([reply("never reached")], total_cost_usd=_OVERSPEND_RAW)
|
||||||
|
with pytest.raises(BudgetExceeded) as caught:
|
||||||
|
_guarded_complete(client, "prompt", role="proposer", meter=self._meter())
|
||||||
|
assert caught.value.observed == _OVERSPEND_RAW
|
||||||
|
|
||||||
|
def test_control_the_rounded_spend_would_not_have_stopped(self) -> None:
|
||||||
|
# The positive control that makes the two tests above evidence rather than coincidence:
|
||||||
|
# the SAME meter does not stop on round(_OVERSPEND_RAW, 6). Without this, a belt that
|
||||||
|
# stopped on everything would satisfy them.
|
||||||
|
meter = self._meter()
|
||||||
|
meter.guard_before_call(round(_OVERSPEND_RAW, 6)) # no raise
|
||||||
|
assert round(_OVERSPEND_RAW, 6) == _CAP_USD
|
||||||
|
|
||||||
|
|
||||||
|
def _literal_round_digit_calls() -> Iterator[tuple[str, int, int]]:
|
||||||
|
"""Every ``round(x, <int literal>)`` in ``src``, as (module, lineno, digits).
|
||||||
|
|
||||||
|
A rounding whose digit count is a NAME (``_SHARE_DIGITS``) is a rule with one source and is
|
||||||
|
deliberately not collected here — that is the shape this file says the money rule lacks.
|
||||||
|
"""
|
||||||
|
for path in sorted(SRC.glob("*.py")):
|
||||||
|
tree = ast.parse(path.read_text("utf-8"))
|
||||||
|
for node in ast.walk(tree):
|
||||||
|
if not isinstance(node, ast.Call) or not isinstance(node.func, ast.Name):
|
||||||
|
continue
|
||||||
|
if node.func.id != "round" or len(node.args) != 2:
|
||||||
|
continue
|
||||||
|
digits = node.args[1]
|
||||||
|
if isinstance(digits, ast.Constant) and isinstance(digits.value, int):
|
||||||
|
yield (path.name, node.lineno, digits.value)
|
||||||
|
|
||||||
|
|
||||||
|
class TestTheMoneyRuleHasNotDrifted:
|
||||||
|
"""The copy count IS the finding: four literal copies, no named source (the sibling's form).
|
||||||
|
|
||||||
|
The sibling closed its version with ``assert run_mod.to_ore is ledger_mod.to_ore``. We have
|
||||||
|
no function to compare, so the equivalent guard is over the call sites themselves.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def test_the_scan_finds_the_known_literal_roundings(self) -> None:
|
||||||
|
# Positive control FIRST (before any negative): the scanner CAN find what it looks for,
|
||||||
|
# and the denominator is the whole package, not a file.
|
||||||
|
modules = sorted({name for name, _, _ in _literal_round_digit_calls()})
|
||||||
|
assert len(list(SRC.glob("*.py"))) == 27, "denominator: src modules scanned"
|
||||||
|
assert modules == ["costsim.py", "run.py", "run_s10.py"]
|
||||||
|
|
||||||
|
def test_every_literal_usd_rounding_uses_six_digits(self) -> None:
|
||||||
|
digits = {d for _, _, d in _literal_round_digit_calls()}
|
||||||
|
assert digits == {6}, "one digit count across every copy — a second value IS the drift"
|
||||||
|
|
||||||
|
def test_there_are_exactly_four_copies_of_the_rule(self) -> None:
|
||||||
|
# Red when a fifth copy appears (drift) and red when one is folded into a named source
|
||||||
|
# (the fix) — either way the finding above stops being true and must be re-measured.
|
||||||
|
assert len(list(_literal_round_digit_calls())) == 4
|
||||||
|
|
||||||
|
def test_the_share_rule_by_contrast_has_a_named_source(self) -> None:
|
||||||
|
# The contrast is the point: one file over, the same package already knows how to give a
|
||||||
|
# rounding one source. The money rule is the one that does not.
|
||||||
|
valuereport = (SRC / "valuereport.py").read_text("utf-8")
|
||||||
|
assert "_SHARE_DIGITS" in valuereport
|
||||||
|
assert "round(" in valuereport
|
||||||
|
assert not [n for n, _, _ in _literal_round_digit_calls() if n == "valuereport.py"]
|
||||||
|
|
||||||
|
|
||||||
|
def _call_keywords_in_src() -> Iterator[tuple[str, str]]:
|
||||||
|
"""Every keyword argument NAME used in a call anywhere in ``src``, as (module, keyword)."""
|
||||||
|
for path in sorted(SRC.glob("*.py")):
|
||||||
|
tree = ast.parse(path.read_text("utf-8"))
|
||||||
|
for node in ast.walk(tree):
|
||||||
|
if isinstance(node, ast.Call):
|
||||||
|
for kw in node.keywords:
|
||||||
|
if kw.arg is not None:
|
||||||
|
yield (path.name, kw.arg)
|
||||||
|
|
||||||
|
|
||||||
|
class TestTheRunTotalUsdCapIsWiredNowhere:
|
||||||
|
"""RATCHET on a measured absence — the gate exists and no path gives it a cap.
|
||||||
|
|
||||||
|
``guard_before_call`` returns immediately when ``max_cost_usd is None``, and every
|
||||||
|
``BudgetMeter(...)`` in ``src`` takes that default. The belt is therefore a no-op outside its
|
||||||
|
own unit test. This is not asserted as desirable; it is pinned so that wiring it — a real
|
||||||
|
change to what bounds a live run — cannot land silently while the documented finding still
|
||||||
|
claims otherwise.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def test_the_keyword_scan_finds_a_cap_that_IS_wired(self) -> None:
|
||||||
|
# Positive control before the negative (økt 29): the scan can find a USD cap keyword, so
|
||||||
|
# the null below is a measurement and not a broken query.
|
||||||
|
wired = {
|
||||||
|
module for module, kw in _call_keywords_in_src() if kw == "max_budget_usd_per_call"
|
||||||
|
}
|
||||||
|
assert wired, "the per-call SDK cap must be found by this scan"
|
||||||
|
|
||||||
|
def test_no_src_path_configures_the_run_total_usd_cap(self) -> None:
|
||||||
|
configured = [module for module, kw in _call_keywords_in_src() if kw == "max_cost_usd"]
|
||||||
|
assert configured == [], "denominator: 27 src modules; the belt is a no-op on every path"
|
||||||
|
|
||||||
|
def test_a_meter_without_a_cap_never_stops_however_large_the_spend(self) -> None:
|
||||||
|
# What the absence above MEANS, executed rather than inferred: the default meter — the
|
||||||
|
# one every src path builds — passes a spend far above any plausible run budget.
|
||||||
|
meter = BudgetMeter(TerminationContract(max_rounds=1000, max_tokens=10_000))
|
||||||
|
meter.guard_before_call(9_999_999.0) # no raise: the belt is off
|
||||||
|
assert meter.rounds_used == 0
|
||||||
Loading…
Add table
Add a link
Reference in a new issue