test(loadbearing): close the vacuous-negative class across the whole suite
Oekt 17 found the class on four named files. This sweep ENUMERATES it: 42 negative substring assertions across 21 test files (STATE's "~34 across 23" was a premise -- measured, it is 42/21). Sixteen of them measured an absence without ever having shown presence; all sixteen now carry a positive control asserting the searched-for string PRESENT in the source artifact, in EXACTLY the form the negative looks for. Files touched: test_costsim, test_loop, test_okf (3 sites), test_preflight, test_run_entrance, test_s10_run_layer, test_sdk_version_guard, test_simulation (2 sites), test_step1_expel, test_step5_refine, test_step7_async_loop, test_step8_promotion, test_valuereport. VALUE-PROOF (green-without / red-with, per the oekt-17 rule that a detach proof is not a value proof). Seven source/fixture mutations, each making the negative vacuous: M1 verdict fixture loses the realization signal VALUE-PROVEN M2 decoy fixture loses its text VALUE-PROVEN M3 renderer stops emitting typed section headings VALUE-PROVEN M4 promotion stops writing the marker VALUE-PROVEN (pass 2) M5 fold stops rendering the realization surface VALUE-PROVEN M6 report stops labelling the cost section VALUE-PROVEN M7 preflight stops importing the SDK VALUE-PROVEN M4 needed pass 2: a PRECEDING assertion caught the same mutation, hiding the new control behind it -- the oekt-17 lesson reproduced. The remaining nine controls are vacuity guards (non-emptiness / form-presence) whose mutation would have to break the source artificially; they are stated as guards, not claimed as value-proven. MEASURED FINDING (test_loop): the FIRST-RUN-MARKER negative cannot be given a positive control at all. Within a run only the CHECKER's critique is fed back -- the proposer's own prior reasoning crosses no prompt boundary, not even within a run. So that negative holds trivially. Left in place with the limitation stated in the test rather than dressed up as a controlled seam; the CRITIQUE negative beside it IS controlled and is the real seam. Mutations were in-place on src/ and shared/ with original bytes restored and sha-verified; git status clean before and after. Suite 688 -> 688 (assertions added inside existing tests, no new test cases). ruff + mypy --strict green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Vc5PmZGjwuJypdhzKnJa5
This commit is contained in:
parent
123ecc3113
commit
30ba68a703
13 changed files with 98 additions and 5 deletions
|
|
@ -128,6 +128,11 @@ class TestNoPriceLiteralInSource:
|
|||
def test_no_example_price_appears_as_a_literal(self) -> None:
|
||||
source = (SRC_PKG / "costsim.py").read_text(encoding="utf-8")
|
||||
pricing = load_pricing() # the bundled data/pricing.example.json
|
||||
# Positive controls: an empty price config would make the loop below iterate
|
||||
# zero times, and an unreadable/empty source would make every substring miss —
|
||||
# both green without guarding anything.
|
||||
assert pricing.prices
|
||||
assert "usd_per_mtok" in source
|
||||
for model_id, price in pricing.prices.items():
|
||||
assert str(price.usd_per_mtok) not in source, (
|
||||
f"price for {model_id} is hardcoded in costsim.py — must come from config"
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue