test(experience): the id was pinned, the keying was not - and the two paths disagree
S3.2 measured with scripts/mutation_harness.py, denominator tests/ (963), every run sha256-restored. The mirroring question answerable offline: is today's boundary - "the id comes from the FILE, the keying comes from the BUNDLE" - load-bearing in both halves? The id half is pinned on both sides (re-minting RED, dropping the mint fallback RED). The keying half is green-but-dead: empty features and the gated C3.2 fix's own shape each left all 963 green. Nothing observed what a seeded verdict is keyed on. One rationale clause too: requiring BOTH learning fields, whose either-or form emits a marker naming the absent field as None (honesty, §1). New beyond C-F5: the sibling's drift form DOES exist here. A verdict is keyed in exactly two places by different rules - the bundle seed (bundle-wide) and the file/inbox path (per-verdict) - so the same verdict id lands in the same first-write-wins slot with a keying decided by LOAD ORDER. And promote writes no candidate features at all, which is why the fix's shape is a no-op against every fixture in the repo and had to be measured against one carrying the fields. Pinned by tests/test_experience_keying_loadbearing.py (11 tests, 963 -> 974, strict superset, 0 node ids lost). No src/ change: C-F5 is GATED on D-A pkt. 4, and these tests exist so the gated work must arrive as a visible red test and a decision. Two behaviour-preserving mutations carry their own controls (a third VerdictRecord site; hoisting the projection read into the loop). Dated under the D7 frame: work after 2026-08-09, never independent convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
544655b4c8
commit
1041687d31
2 changed files with 430 additions and 2 deletions
354
tests/test_experience_keying_loadbearing.py
Normal file
354
tests/test_experience_keying_loadbearing.py
Normal file
|
|
@ -0,0 +1,354 @@
|
|||
"""S3.2: the ExpeL seeding seam — is WHAT EACH VERDICT IS KEYED ON load-bearing?
|
||||
|
||||
The D7 mirroring queue (``docs/2026-08-18-funn-koeer-og-gater.md § D7-speilingskøen``)
|
||||
carries S3.2 because the MAF sibling fixed multi-candidate ExpeL seeding: the seed
|
||||
should read the judged candidate from the verdict file's OWN frontmatter (promotion
|
||||
writing it), with today's keying as fallback. On our side the feature half of that
|
||||
defect is confirmed as C-F5 (``docs/review-2026-07.md``) and GATED on D-A pkt. 4 + a
|
||||
commons pull (parity plan row 14), while the id half is already solved here — a
|
||||
loaded ``verdict_id`` is read VERBATIM (§4.2). So this file does NOT build the fix.
|
||||
It answers the question that is answerable offline today: **is today's boundary —
|
||||
"the id comes from the FILE, the keying comes from the BUNDLE" — load-bearing in
|
||||
BOTH halves?**
|
||||
|
||||
MEASURED 2026-09-12 with ``scripts/mutation_harness.py``, denominator ``tests/``
|
||||
(the whole suite, 963 tests), every run restored sha256-verified:
|
||||
|
||||
- Detaching the seam outright (seeding iterates nothing) is RED. So is dropping the
|
||||
type filter, the decision default, and the ``description`` read.
|
||||
- The ID HALF is pinned on both sides: re-minting instead of reading the frontmatter
|
||||
id is RED (``test_step8``), and dropping the mint fallback so unpromoted verdicts
|
||||
key on an empty id is RED too.
|
||||
- The KEYING HALF is GREEN-BUT-DEAD. Replacing the bundle projection with empty
|
||||
features left all 963 tests green, and so did the gated fix's own shape (read the
|
||||
candidate from the verdict file's frontmatter, fallback today's keying). Nothing in
|
||||
the suite observes what a seeded verdict is keyed ON.
|
||||
- One clause of the rationale is green-but-dead as well: requiring BOTH learning
|
||||
fields before emitting the marker. Weakening it to either-or left the suite green,
|
||||
and the consequence is an honesty defect (§1) — a marker naming a field as None.
|
||||
|
||||
The mutations were shown to change behaviour before their greens were read as holes
|
||||
(the økt 39 trap: the harness reports a no-op and an undetected seam identically).
|
||||
Empty keying drops the seeded record's similarity against its own bundle projection
|
||||
from 1.00 to 0.15 — observable, and ranking-relevant the moment a second entry
|
||||
exists. C-F5's rest-defect was reproduced live: two promoted verdicts about distinct
|
||||
candidates (LED vs ventilation) both seed keyed on the bundle's ONE projection, so a
|
||||
query carrying the ventilation candidate's own features scores BOTH at 0.0 and the
|
||||
tie falls to hex id order, not to structure.
|
||||
|
||||
NEW BEYOND C-F5 — the sibling's drift form DOES exist here, across two paths. A
|
||||
verdict record is keyed in exactly two places with different rules: the bundle seed
|
||||
(``experience.py``, keyed bundle-wide) and the file/inbox path (``inbox.py``, keyed
|
||||
on the verdict's OWN ``proposal_features``). The same verdict — same candidate, same
|
||||
minted id, hence the same first-write-wins store slot — is therefore keyed
|
||||
DIFFERENTLY depending on which path reached the store first. And the fix cannot be
|
||||
one-sided: ``promote`` writes ``verdict_id`` and ``description`` but no candidate
|
||||
features at all, which is why the fix's own shape is invisible to every fixture in
|
||||
this repo today.
|
||||
|
||||
HONEST LIMIT — what this does NOT say. Pinning that all verdicts in a bundle share
|
||||
one keying TODAY is not an endorsement; C-F5 calls it a MAJOR spec-level defect and
|
||||
the fix is gated, not declined. These tests pin the boundary so the gated work must
|
||||
arrive as a visible red test and a decision, never as a silent swap — the same
|
||||
ratchet role ``test_ingest_stamp_conformance_loadbearing.py`` plays upstream. No
|
||||
``src/`` change is made here, no spec text is touched, and the golden fixture is not
|
||||
involved in the keying at all (which is precisely why it cannot help).
|
||||
|
||||
Dated under the D7 frame: this is work AFTER 2026-08-09 and must NOT be read as
|
||||
independent convergence with the sibling.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from portfolio_optimiser_claude.experience import (
|
||||
CandidateFeatures,
|
||||
VerdictRecord,
|
||||
VerdictStore,
|
||||
mint_verdict_id,
|
||||
seed_store_from_bundle,
|
||||
similarity,
|
||||
)
|
||||
from portfolio_optimiser_claude.inbox import VerdictDocument
|
||||
from portfolio_optimiser_claude.ir import load_validator_input
|
||||
from portfolio_optimiser_claude.promotion import promote
|
||||
|
||||
SRC_PKG = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser_claude"
|
||||
SHARED_BUNDLE = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
|
||||
|
||||
# A candidate that is NOT the bundle's: other codes, other measure type, other
|
||||
# magnitude bucket. Keyed on itself it sits at similarity 0.00 from the bundle
|
||||
# projection, which is what makes a miskeying observable at all.
|
||||
OTHER = CandidateFeatures(
|
||||
affected_codes=frozenset({"VENT-AGG-01"}),
|
||||
measure_type="ventilation-schedule",
|
||||
claimed_saving_nok=250_000.0,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def bundle(tmp_path: Path) -> Path:
|
||||
target = tmp_path / "bundle"
|
||||
shutil.copytree(SHARED_BUNDLE, target)
|
||||
return target
|
||||
|
||||
|
||||
def _bundle_keying(bundle_dir: Path) -> CandidateFeatures:
|
||||
return CandidateFeatures.from_proposal(load_validator_input(bundle_dir))
|
||||
|
||||
|
||||
def _link(bundle_dir: Path, filename: str) -> None:
|
||||
index = bundle_dir / "index.md"
|
||||
index.write_text(index.read_text(encoding="utf-8") + f"- [Dom]({filename})\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _write_verdict(bundle_dir: Path, filename: str, extra_frontmatter: str) -> None:
|
||||
(bundle_dir / filename).write_text(
|
||||
"---\n"
|
||||
"type: verdict\n"
|
||||
"title: Dom\n"
|
||||
"decision: approved\n"
|
||||
"description: Dom om en annen kandidat.\n"
|
||||
f"{extra_frontmatter}"
|
||||
"tags: [verdict]\n"
|
||||
"---\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
_link(bundle_dir, filename)
|
||||
|
||||
|
||||
def _self_describing_verdict(bundle_dir: Path, features: CandidateFeatures) -> str:
|
||||
"""A verdict file carrying the candidate it judges IN ITS OWN frontmatter.
|
||||
|
||||
This is the INPUT the gated C3.2 fix reads. Today's seeding ignores it, and
|
||||
``promote`` does not write it — both halves are pinned below.
|
||||
"""
|
||||
verdict_id = mint_verdict_id(features)
|
||||
_write_verdict(
|
||||
bundle_dir,
|
||||
"verdict-self-describing.md",
|
||||
f"verdict_id: {verdict_id}\n"
|
||||
f"measure_type: {features.measure_type}\n"
|
||||
f"claimed_saving_nok: {features.claimed_saving_nok}\n",
|
||||
)
|
||||
return verdict_id
|
||||
|
||||
|
||||
def _record_for(store: VerdictStore, features: CandidateFeatures, verdict_id: str) -> VerdictRecord:
|
||||
return next(r for r in store.retrieve(features, k=20) if r.verdict_id == verdict_id)
|
||||
|
||||
|
||||
def _seeded(bundle_dir: Path) -> VerdictStore:
|
||||
store = VerdictStore()
|
||||
seed_store_from_bundle(store, bundle_dir)
|
||||
return store
|
||||
|
||||
|
||||
def _document(features: CandidateFeatures, rationale: str) -> VerdictDocument:
|
||||
return VerdictDocument.from_candidate(
|
||||
features, decision="approved", rationale=rationale, description="d"
|
||||
)
|
||||
|
||||
|
||||
class TestThereAreExactlyTwoKeyingSites:
|
||||
"""The population control, and it runs first: every clause below is worth only as
|
||||
much as the claim that there are exactly TWO places where a verdict is keyed."""
|
||||
|
||||
def test_a_verdict_record_is_keyed_in_exactly_two_places_with_distinct_roles(self) -> None:
|
||||
# The sibling's S3.2 shape made measurable here. Two keying sites exist BY
|
||||
# DESIGN and answer different specs: the bundle seed (§3 Step 1, keyed on the
|
||||
# bundle's ONE IR projection) and the file/inbox path (§4.2, keyed on the
|
||||
# verdict's own proposal_features). RED the day a third appears — a verdict
|
||||
# keyed somewhere no one is looking is how the miskeying spreads. Read from
|
||||
# the AST, not prose.
|
||||
sites: dict[str, int] = {}
|
||||
for module in sorted(SRC_PKG.glob("*.py")):
|
||||
tree = ast.parse(module.read_text(encoding="utf-8"))
|
||||
found = [
|
||||
node
|
||||
for node in ast.walk(tree)
|
||||
if isinstance(node, ast.Call)
|
||||
and isinstance(node.func, ast.Name)
|
||||
and node.func.id == "VerdictRecord"
|
||||
]
|
||||
if found:
|
||||
sites[module.name] = len(found)
|
||||
assert sites == {"experience.py": 1, "inbox.py": 1}, (
|
||||
f"a verdict record is now keyed in {sites} — the two-path design "
|
||||
"(bundle seed + file/inbox) has drifted"
|
||||
)
|
||||
|
||||
def test_the_bundle_projection_is_read_once_per_bundle_not_once_per_file(self) -> None:
|
||||
# Behaviour-preserving guard: hoisting the projection read INTO the seeding
|
||||
# loop changes no result today, and it is the first half of the gated fix.
|
||||
# Every behaviour test below stays green under that move, so the shape of the
|
||||
# read is pinned structurally or not at all.
|
||||
tree = ast.parse((SRC_PKG / "experience.py").read_text(encoding="utf-8"))
|
||||
seed_fn = next(
|
||||
node
|
||||
for node in ast.walk(tree)
|
||||
if isinstance(node, ast.FunctionDef) and node.name == "seed_store_from_bundle"
|
||||
)
|
||||
reads = [
|
||||
node
|
||||
for node in ast.walk(seed_fn)
|
||||
if isinstance(node, ast.Call)
|
||||
and isinstance(node.func, ast.Name)
|
||||
and node.func.id == "load_validator_input"
|
||||
]
|
||||
assert len(reads) == 1, "the bundle projection must be read once, not per verdict file"
|
||||
|
||||
def test_keying_is_observable_at_all(self, bundle: Path) -> None:
|
||||
# POSITIVE CONTROL FOR EVERY NEGATIVE BELOW. Each test that follows asserts
|
||||
# WHICH features a record carries; that is vacuous unless a different keying
|
||||
# demonstrably scores differently on this bundle.
|
||||
keying = _bundle_keying(bundle)
|
||||
assert similarity(keying, keying) == pytest.approx(1.0)
|
||||
assert similarity(OTHER, keying) == pytest.approx(0.0)
|
||||
|
||||
|
||||
class TestWhatTheSeedKeysOnIsLoadBearing:
|
||||
"""One test per clause of ``features=features`` in ``seed_store_from_bundle`` —
|
||||
each proved RED under exactly the mutation it names, with the fold's realization
|
||||
marker green in both runs (the mutation moved the KEYING, not the seam)."""
|
||||
|
||||
def test_a_seeded_verdict_is_keyed_on_the_bundle_ir_projection(self, bundle: Path) -> None:
|
||||
# RED under replacing the keying with empty features. Measured: that drops
|
||||
# the record's similarity against its own bundle projection from 1.00 to
|
||||
# 0.15, so the mutation is behaviour-changing, not a no-op.
|
||||
keying = _bundle_keying(bundle)
|
||||
(record,) = _seeded(bundle).retrieve(keying, k=1)
|
||||
assert record.features == keying
|
||||
assert similarity(record.features, keying) == pytest.approx(1.0)
|
||||
|
||||
def test_the_verdict_files_own_candidate_frontmatter_is_ignored_today(
|
||||
self, bundle: Path
|
||||
) -> None:
|
||||
# The GATED C3.2 fix, pinned as a boundary: a verdict file may carry the
|
||||
# candidate it judges, and today's seeding still keys it on the bundle. RED
|
||||
# the day the seed reads those fields — which is the point. The fix must
|
||||
# arrive as this red test plus a D-A pkt. 4 decision, never as a silent swap.
|
||||
verdict_id = _self_describing_verdict(bundle, OTHER)
|
||||
keying = _bundle_keying(bundle)
|
||||
record = _record_for(_seeded(bundle), keying, verdict_id)
|
||||
assert record.features == keying
|
||||
assert record.features != OTHER
|
||||
assert record.features.measure_type != OTHER.measure_type
|
||||
assert record.features.claimed_saving_nok != OTHER.claimed_saving_nok
|
||||
|
||||
def test_a_multi_candidate_bundle_keys_every_verdict_on_the_one_projection(
|
||||
self, bundle: Path
|
||||
) -> None:
|
||||
# C-F5's rest-defect, reproduced as a test: two promoted verdicts about
|
||||
# DISTINCT candidates both seed keyed on the bundle's one projection, so
|
||||
# retrieval cannot tell them apart — a query carrying the ventilation
|
||||
# candidate's own features scores every entry identically and the tie falls
|
||||
# to hex id order. RED under any per-verdict keying.
|
||||
promote(
|
||||
_document(OTHER, "dom om ventilasjon"),
|
||||
bundle,
|
||||
approved_by="e",
|
||||
experiment="x",
|
||||
timestamp="2026-09-12",
|
||||
)
|
||||
store = _seeded(bundle)
|
||||
keying = _bundle_keying(bundle)
|
||||
assert len(store) == 2
|
||||
assert {r.features for r in store.retrieve(keying, k=20)} == {keying}
|
||||
scores = [similarity(r.features, OTHER) for r in store.retrieve(OTHER, k=20)]
|
||||
assert scores == [pytest.approx(0.0), pytest.approx(0.0)]
|
||||
|
||||
def test_the_same_verdict_id_is_keyed_differently_by_the_two_paths(self, bundle: Path) -> None:
|
||||
# The sibling's drift form, measured on our side: the SAME verdict — same
|
||||
# candidate, same minted id, hence the same first-write-wins store slot — is
|
||||
# keyed on the bundle via the seed and on its own candidate via the inbox.
|
||||
verdict_id = _self_describing_verdict(bundle, OTHER)
|
||||
document = _document(OTHER, "dom om ventilasjon")
|
||||
assert document.id == verdict_id, "same candidate must mint the same id (§4.2)"
|
||||
keying = _bundle_keying(bundle)
|
||||
via_bundle = _record_for(_seeded(bundle), keying, verdict_id)
|
||||
via_inbox = document.to_record()
|
||||
assert via_bundle.features != via_inbox.features
|
||||
assert via_bundle.features == keying
|
||||
assert via_inbox.features == OTHER
|
||||
|
||||
def test_which_keying_survives_is_decided_by_load_order(self, bundle: Path) -> None:
|
||||
# The consequence of two rules over one store slot: FIRST-write-wins (§4.2)
|
||||
# makes the surviving keying a function of arrival order, not of the verdict.
|
||||
# Split from the test above deliberately — a red test proves only its first
|
||||
# failing assert, so the order claim needs its own.
|
||||
verdict_id = _self_describing_verdict(bundle, OTHER)
|
||||
via_inbox = _document(OTHER, "dom om ventilasjon").to_record()
|
||||
keying = _bundle_keying(bundle)
|
||||
|
||||
bundle_first = _seeded(bundle)
|
||||
bundle_first.add(via_inbox)
|
||||
inbox_first = VerdictStore()
|
||||
inbox_first.add(via_inbox)
|
||||
seed_store_from_bundle(inbox_first, bundle)
|
||||
|
||||
assert _record_for(bundle_first, keying, verdict_id).features == keying
|
||||
assert _record_for(inbox_first, keying, verdict_id).features == OTHER
|
||||
|
||||
|
||||
class TestTheLearningMarkerNeedsBothFields:
|
||||
"""§1 honesty rule at the seeding layer: the marker names two numbers, so it is
|
||||
emitted only when the verdict file supplies both."""
|
||||
|
||||
def test_a_half_specified_learning_pair_emits_no_marker(self, bundle: Path) -> None:
|
||||
# RED under weakening the pair requirement to either-or: the rationale then
|
||||
# carries a marker built from the one field that IS there.
|
||||
record = self._half_pair_record(bundle)
|
||||
assert "realiseringsgrad" not in record.rationale
|
||||
assert record.rationale == "Dom om en annen kandidat."
|
||||
|
||||
def test_a_half_specified_learning_pair_never_fabricates_a_missing_number(
|
||||
self, bundle: Path
|
||||
) -> None:
|
||||
# The §1 half of the same clause, asserted FIRST in its own test: under
|
||||
# either-or the emitted marker names the absent field as None — a fabricated
|
||||
# learning signal, which is the one thing the honesty rule forbids outright.
|
||||
assert "None" not in self._half_pair_record(bundle).rationale
|
||||
|
||||
@staticmethod
|
||||
def _half_pair_record(bundle_dir: Path) -> VerdictRecord:
|
||||
_write_verdict(
|
||||
bundle_dir, "verdict-half.md", "verdict_id: halfpair\nrealization_rate: 0.5\n"
|
||||
)
|
||||
return _record_for(_seeded(bundle_dir), _bundle_keying(bundle_dir), "halfpair")
|
||||
|
||||
|
||||
class TestTheGatedFixCannotBeOneSided:
|
||||
"""The fix reads per-verdict features; nothing writes them. Both halves pinned."""
|
||||
|
||||
def test_promotion_writes_no_candidate_features_today(self, bundle: Path) -> None:
|
||||
# §6 keeps the promoted file minimal: id, decision, rationale-as-description,
|
||||
# provenance. The candidate it judges is NOT among them, so the C3.2 fix must
|
||||
# change promotion too. RED the day promotion starts writing them — the
|
||||
# deliberate signal that the gated work has begun on this side.
|
||||
path = promote(
|
||||
_document(OTHER, "dom om ventilasjon"),
|
||||
bundle,
|
||||
approved_by="e",
|
||||
experiment="x",
|
||||
timestamp="2026-09-12",
|
||||
)
|
||||
keys = {
|
||||
line.split(":", 1)[0]
|
||||
for line in path.read_text(encoding="utf-8").splitlines()
|
||||
if ":" in line
|
||||
}
|
||||
assert keys == {
|
||||
"type",
|
||||
"title",
|
||||
"decision",
|
||||
"description",
|
||||
"verdict_id",
|
||||
"provenance",
|
||||
"tags",
|
||||
}
|
||||
Loading…
Add table
Add a link
Reference in a new issue