Oekt 17 found the class on four named files. This sweep ENUMERATES it: 42 negative
substring assertions across 21 test files (STATE's "~34 across 23" was a premise --
measured, it is 42/21). Sixteen of them measured an absence without ever having
shown presence; all sixteen now carry a positive control asserting the searched-for
string PRESENT in the source artifact, in EXACTLY the form the negative looks for.
Files touched: test_costsim, test_loop, test_okf (3 sites), test_preflight,
test_run_entrance, test_s10_run_layer, test_sdk_version_guard, test_simulation
(2 sites), test_step1_expel, test_step5_refine, test_step7_async_loop,
test_step8_promotion, test_valuereport.
VALUE-PROOF (green-without / red-with, per the oekt-17 rule that a detach proof is
not a value proof). Seven source/fixture mutations, each making the negative vacuous:
M1 verdict fixture loses the realization signal VALUE-PROVEN
M2 decoy fixture loses its text VALUE-PROVEN
M3 renderer stops emitting typed section headings VALUE-PROVEN
M4 promotion stops writing the marker VALUE-PROVEN (pass 2)
M5 fold stops rendering the realization surface VALUE-PROVEN
M6 report stops labelling the cost section VALUE-PROVEN
M7 preflight stops importing the SDK VALUE-PROVEN
M4 needed pass 2: a PRECEDING assertion caught the same mutation, hiding the new
control behind it -- the oekt-17 lesson reproduced. The remaining nine controls are
vacuity guards (non-emptiness / form-presence) whose mutation would have to break
the source artificially; they are stated as guards, not claimed as value-proven.
MEASURED FINDING (test_loop): the FIRST-RUN-MARKER negative cannot be given a
positive control at all. Within a run only the CHECKER's critique is fed back --
the proposer's own prior reasoning crosses no prompt boundary, not even within a
run. So that negative holds trivially. Left in place with the limitation stated in
the test rather than dressed up as a controlled seam; the CRITIQUE negative beside
it IS controlled and is the real seam.
Mutations were in-place on src/ and shared/ with original bytes restored and
sha-verified; git status clean before and after. Suite 688 -> 688 (assertions added
inside existing tests, no new test cases). ruff + mypy --strict green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Vc5PmZGjwuJypdhzKnJa5
- run.py: compose_run_context (§5: merge inbox -> seed -> fold, read-only on
the inbox) + execute_run (§8 meter, artifacts persisted on BOTH outcomes,
structured exit 3 on budget stop) + thin CLI (python -m ..run). The model
client is injected; only default_client_factory constructs the SDK client
(wired, never executed by the suite). The navigated docs dir comes from the
validated startup contract (resolves review OBS-2 on the shippable path;
run_s10.py stays byte-frozen fasit -> won't-fix there).
- test_run_entrance_loadbearing.py: inbox verdict reaches the composed
context (detach-proven: merge dropped -> red), empty/missing-inbox
controls, read-only inbox byte-proof, R-10 budget-stop binding via the NEW
entrance (detach-proven: stop persistence dropped -> red), happy path
through the CLI with the inbox signal surviving the chain, SDK-wiring test.
- test_ingest_adoption.py (K2.9): the two library guarantees the consumer
relies on, bound through the seam — empty CSV -> typed SourceError with NO
partial bundle on disk; non-SELECT SQL -> SourceError 'returned no columns'
(behavior verified empirically against pin dae0bd1a before binding).
- README: inbox section now points at the shippable entrance; run.py added
to the run layer; stale test count 265 -> 395.
386 -> 395 tests, full gate green (pytest, ruff check+format, mypy strict);
goldens unchanged; runs/s10 and run_s10.py untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>