feat(verdicts): key each verdict on its own candidate, not the bundle's one IR projection (S3.2)
seed_store_from_bundle keyed EVERY `type: verdict` file on bundle_candidate_features — the single candidate the bundle's validator-input.json describes. A bundle carrying verdicts about several candidates collapsed them onto one key, so a verdict about candidate B scored a perfect structural match against candidate A's query and could be folded into A's hypothesis prompt. The ExpeL substrate was single-candidate by construction. A verdict file may now carry its own structural key in frontmatter (affected_codes / measure_type / claimed_saving_nok); absent, keying falls back to the bundle candidate, so every pre-S3.2 seed keeps working unchanged. promote_verdict writes the three fields, so a promoted verdict — frequently about a different candidate than the target bundle's projection — does not impersonate that candidate. Semantics decided HERE, not pulled: commons' seeding rule (method-spec §3 Steg 1 + bundle example) has not arrived; we said we would build locally first. D7 mirroring stays open. - ALL THREE fields or none. A partial declaration raises VerdictFrontmatterError rather than merging with the bundle candidate, which would mint a key belonging to NEITHER candidate. Validation, never repair (mirrors write_concept_file); the tolerant-skip rule belongs to the RAW inbox layer. - claimed_saving_nok parses via json.loads — the SAME literal rule the IR projection went through — and is written back with str() of the raw value. _mint_id hashes that value, so 30000 and 30000.0 are different keys; a normalising writer would split one candidate's signal across two ids. - The structural key is signal-free, so it does not weaken the Step-8 no-leak property (Test C green). Load-bearing MEASURED, five mutations all red: detach per-verdict keying · detach the fields promote_verdict writes · make a partial/unparseable key tolerant · normalise the magnitude on write · remove the fallback (control — breaks the step1 suite at collection, proving the fallback bears load). 589 -> 597 tests. Full gate green (pytest, ruff, mypy). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QkjvTTxrg9LTrmghebfiij
This commit is contained in:
parent
e8cec2e2c0
commit
012adc0a3c
10 changed files with 496 additions and 17 deletions
137
tests/test_step32_multicandidate_loadbearing.py
Normal file
137
tests/test_step32_multicandidate_loadbearing.py
Normal file
|
|
@ -0,0 +1,137 @@
|
|||
"""S3.2 load-bearing seam: a verdict is keyed on ITS OWN candidate, not on the bundle's single IR
|
||||
projection (sesjonsplan Fase 2-6 §S3.2; målbilde §2 step 1).
|
||||
|
||||
The gap: ``seed_store_from_bundle`` keyed EVERY ``type: verdict`` file in a bundle on
|
||||
``bundle_candidate_features`` — the one candidate the bundle's ``validator-input.json`` describes.
|
||||
A bundle carrying verdicts about several candidates therefore collapsed them all onto one key: a
|
||||
verdict about candidate B scored a perfect structural match against candidate A's query and could
|
||||
be folded into A's hypothesis prompt. The ExpeL substrate was single-candidate by construction.
|
||||
|
||||
S3.2 lets each verdict file carry its own structural fields in frontmatter
|
||||
(``affected_codes`` / ``measure_type`` / ``claimed_saving_nok``); when they are absent, keying falls
|
||||
back to the bundle candidate — so every hand-authored seed written before S3.2 keeps working
|
||||
unchanged (proven by the untouched step1/step7/step8 suites plus Test B here).
|
||||
|
||||
Three load-bearing tests:
|
||||
- Test A (SEPARATION, the RED point): a bundle with verdicts about two disjoint candidates — a
|
||||
verdict about candidate B must NEVER reach candidate A's hypothesis prompt. Goes RED the moment
|
||||
per-verdict keying is detached: both verdicts then carry candidate A's key, score identically,
|
||||
mint an identical id, and the stable sort hands back whichever the bundle links first — which the
|
||||
fixture deliberately makes the B verdict.
|
||||
- Test B (FALLBACK): a pre-S3.2 seed with no structural frontmatter still keys on the bundle
|
||||
candidate and still retrieves — backward compatibility, stated as a test rather than assumed.
|
||||
- Test C (DISTINCT IDENTITY): the two verdicts mint DIFFERENT ids. Ids hash the structural features,
|
||||
so a shared id would silently collapse the two candidates in ``VerdictStore.add`` (first-write-wins)
|
||||
even where retrieval separated them.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import shutil
|
||||
from importlib.resources import files
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from portfolio_optimiser.verdicts import (
|
||||
ExpeLContextProvider,
|
||||
VerdictFrontmatterError,
|
||||
bundle_candidate_features,
|
||||
seed_store_from_bundle,
|
||||
)
|
||||
|
||||
_MULTI_BUNDLE = str(files("portfolio_optimiser").joinpath("data/bundles/multi-kandidat-mikro"))
|
||||
_SINGLE_BUNDLE = str(files("portfolio_optimiser").joinpath("data/bundles/bygg-energi-mikro-a"))
|
||||
|
||||
# Realization markers unique to each verdict file's frontmatter; they reach a prompt only through the
|
||||
# gated ExpeL fold, so their presence/absence in the few-shot IS the retrieval outcome.
|
||||
_MARKER_A = "realiseringsgrad=0.80"
|
||||
_MARKER_B = "realiseringsgrad=0.31"
|
||||
|
||||
|
||||
def _fewshot_for_bundle_candidate(bundle_dir: str, *, k: int = 1) -> str:
|
||||
"""The ExpeL few-shot block Step 1 folds into the hypothesis prompt, for the bundle's OWN
|
||||
candidate — the exact string an agent would read."""
|
||||
store = seed_store_from_bundle(bundle_dir)
|
||||
query = bundle_candidate_features(bundle_dir)
|
||||
return ExpeLContextProvider(store, query, k=k).format_fewshot()
|
||||
|
||||
|
||||
def test_verdict_about_another_candidate_never_reaches_this_candidates_prompt() -> None:
|
||||
"""Test A — the RED point. Candidate A's hypothesis prompt carries A's verdict and NOT B's."""
|
||||
fewshot = _fewshot_for_bundle_candidate(_MULTI_BUNDLE)
|
||||
|
||||
assert _MARKER_A in fewshot, (
|
||||
"candidate A's own verdict must reach A's hypothesis prompt; got:\n" + fewshot
|
||||
)
|
||||
assert _MARKER_B not in fewshot, (
|
||||
"a verdict about candidate B (different cost code, measure and magnitude) reached "
|
||||
"candidate A's hypothesis prompt — the verdicts are keyed on the bundle's single IR "
|
||||
"projection instead of on their own candidate:\n" + fewshot
|
||||
)
|
||||
|
||||
|
||||
def test_verdict_without_structural_frontmatter_still_keys_on_the_bundle_candidate() -> None:
|
||||
"""Test B — fallback. A pre-S3.2 seed (no ``affected_codes``/``measure_type``/
|
||||
``claimed_saving_nok`` in frontmatter) keys on the bundle candidate exactly as before, so it
|
||||
still scores a perfect structural match and still reaches the prompt."""
|
||||
store = seed_store_from_bundle(_SINGLE_BUNDLE)
|
||||
query = bundle_candidate_features(_SINGLE_BUNDLE)
|
||||
|
||||
assert store.verdicts, "the single-candidate fixture must still seed at least one verdict"
|
||||
assert all(v.proposal_features == query for v in store.verdicts), (
|
||||
"verdicts with no structural frontmatter must fall back to the bundle candidate key"
|
||||
)
|
||||
assert "realiseringsgrad=0.80" in ExpeLContextProvider(store, query, k=1).format_fewshot()
|
||||
|
||||
|
||||
def test_the_two_candidates_verdicts_mint_distinct_ids() -> None:
|
||||
"""Test C — distinct identity. Ids hash the structural features, so per-verdict keying must also
|
||||
give the two verdicts different ids; a shared id would collapse them in ``VerdictStore.add``
|
||||
(first-write-wins) even where retrieval kept them apart."""
|
||||
verdicts = seed_store_from_bundle(_MULTI_BUNDLE).verdicts
|
||||
|
||||
assert len(verdicts) == 2, f"expected both candidates' verdicts, got {len(verdicts)}"
|
||||
assert len({v.id for v in verdicts}) == 2, (
|
||||
f"two disjoint candidates minted the same verdict id: {[v.id for v in verdicts]}"
|
||||
)
|
||||
|
||||
|
||||
# --- Test D: a half-declared key is refused, never silently merged -------------------------------
|
||||
|
||||
|
||||
def _bundle_with_patched_verdict(tmp_path: Path, old: str, new: str) -> str:
|
||||
"""A throwaway copy of the multi-candidate bundle with one line of candidate B's verdict
|
||||
frontmatter rewritten — the packaged fixture is never mutated."""
|
||||
dst = tmp_path / "bundle"
|
||||
shutil.copytree(_MULTI_BUNDLE, dst)
|
||||
target = dst / "verdict-b-asfalt.md"
|
||||
text = target.read_text(encoding="utf-8")
|
||||
assert old in text, f"fixture drift: {old!r} not in verdict-b-asfalt.md"
|
||||
target.write_text(text.replace(old, new, 1), encoding="utf-8")
|
||||
return str(dst)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("old", "new", "why"),
|
||||
[
|
||||
("measure_type: ", "measure_type_disabled: ", "two of the three fields declared"),
|
||||
("claimed_saving_nok: 900000", "claimed_saving_nok: nokså mye", "unparseable magnitude"),
|
||||
("affected_codes: [05.2]", "affected_codes: []", "declared but empty code set"),
|
||||
],
|
||||
)
|
||||
def test_a_partial_or_unparseable_structural_key_is_refused(tmp_path, old, new, why) -> None:
|
||||
"""Test D — fail-fast. A verdict file that declares its structural key PARTIALLY or unparseably
|
||||
raises rather than falling back to the bundle candidate.
|
||||
|
||||
The silent alternative is not neutral: it keys the verdict to the WRONG candidate, which is the
|
||||
exact defect S3.2 closes — and a merge of some declared fields with some bundle fields mints a
|
||||
key belonging to NEITHER candidate. The curated context layer is validated, never repaired
|
||||
(mirroring ``write_concept_file``); the tolerant-skip rule belongs to the RAW inbox layer, where
|
||||
anyone may drop anything.
|
||||
|
||||
RED if the refusal is relaxed into a fallback."""
|
||||
bundle_dir = _bundle_with_patched_verdict(tmp_path, old, new)
|
||||
|
||||
with pytest.raises(VerdictFrontmatterError):
|
||||
seed_store_from_bundle(bundle_dir)
|
||||
|
|
@ -33,11 +33,13 @@ import pytest
|
|||
|
||||
from portfolio_optimiser import okf
|
||||
from portfolio_optimiser.verdicts import (
|
||||
ExpeLContextProvider,
|
||||
ProposalFeatures,
|
||||
PromotionRefused,
|
||||
Verdict,
|
||||
bundle_candidate_features,
|
||||
promote_verdict,
|
||||
seed_store_from_bundle,
|
||||
)
|
||||
|
||||
BUNDLE_DIR = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
|
||||
|
|
@ -172,3 +174,87 @@ def test_promoted_signal_stays_out_of_bundle_context(tmp_path) -> None:
|
|||
# though it IS present in the bundle (just excluded from context, like the seed verdict):
|
||||
assert _MARKER in okf.parse_frontmatter(path)["description"]
|
||||
assert path.name in {f.name for f in bundle.verdicts}
|
||||
|
||||
|
||||
# --- Test D (S3.2): the promoted verdict carries its OWN structural key --------------------------
|
||||
|
||||
|
||||
def test_promoted_verdict_about_another_candidate_keeps_its_own_key(tmp_path) -> None:
|
||||
"""S3.2 ROUND-TRIP: a promoted verdict is frequently about a DIFFERENT candidate than the one
|
||||
the target bundle's IR projection describes. ``promote_verdict`` therefore writes the verdict's
|
||||
own structural key, and ``seed_store_from_bundle`` reads it back — so the promoted verdict does
|
||||
not impersonate the bundle candidate in the next run's retrieval.
|
||||
|
||||
RED if ``promote_verdict`` stops writing the three structural fields (the promoted verdict then
|
||||
falls back to the bundle candidate's key and its marker reaches that candidate's prompt)."""
|
||||
bundle_dir = _copy_bundle(tmp_path)
|
||||
other_marker = "realiseringsgrad=0.19"
|
||||
other_candidate = Verdict(
|
||||
id="STEG8-OTHER-CANDIDATE",
|
||||
proposal_features=ProposalFeatures(
|
||||
affected_codes=frozenset({"05.2", "03.1"}),
|
||||
measure_type="Redusert asfalttykkelse",
|
||||
claimed_saving_nok=900000,
|
||||
),
|
||||
decision="approved",
|
||||
rationale=f"asfalttiltak godkjent med kraftig realiseringskorreksjon ({other_marker})",
|
||||
)
|
||||
|
||||
path = promote_verdict(
|
||||
bundle_dir, other_candidate, approver="persona", experiment="exp-D", timestamp="2026-06-30"
|
||||
)
|
||||
|
||||
fm = okf.parse_frontmatter(path)
|
||||
assert fm["affected_codes"] == "[03.1, 05.2]" # sorted -> deterministic bytes
|
||||
assert fm["measure_type"] == "Redusert asfalttykkelse"
|
||||
assert fm["claimed_saving_nok"] == "900000"
|
||||
|
||||
# The round trip: the next run's seed keys it on ITS candidate, so it does not surface for the
|
||||
# bundle's own (LED) candidate.
|
||||
store = seed_store_from_bundle(bundle_dir)
|
||||
query = bundle_candidate_features(bundle_dir)
|
||||
fewshot = ExpeLContextProvider(store, query, k=1).format_fewshot()
|
||||
assert other_marker not in fewshot, (
|
||||
"a promoted verdict about a different candidate reached this candidate's hypothesis "
|
||||
"prompt — the promoted file is not carrying its own structural key:\n" + fewshot
|
||||
)
|
||||
# Keyed on the promoted candidate's STRUCTURAL fields (``description`` is surface text, outside
|
||||
# both the similarity score and the minted id — the seeder fills it from ``measure_type``).
|
||||
keys = {
|
||||
(v.proposal_features.affected_codes, v.proposal_features.measure_type)
|
||||
for v in store.verdicts
|
||||
}
|
||||
assert (frozenset({"05.2", "03.1"}), "Redusert asfalttykkelse") in keys, (
|
||||
f"the promoted verdict was seeded with the wrong structural key; got {keys}"
|
||||
)
|
||||
|
||||
|
||||
def test_promotion_round_trip_preserves_the_key_for_the_bundles_own_candidate(tmp_path) -> None:
|
||||
"""S3.2 IDENTITY: promoting a verdict about the bundle's OWN candidate must re-seed to exactly
|
||||
the key the pre-S3.2 fallback produced — same codes, same measure, same magnitude TYPE.
|
||||
|
||||
Otherwise a candidate's learning signal splits across two keys over time (the promoted verdicts
|
||||
under one, the hand-authored seeds under the fallback), and neither retrieves the other. The
|
||||
magnitude is asserted on its exact value AND type because ``_mint_id`` hashes the raw value:
|
||||
``30000`` and ``30000.0`` are different ids."""
|
||||
bundle_dir = _copy_bundle(tmp_path)
|
||||
candidate = bundle_candidate_features(bundle_dir)
|
||||
promote_verdict(
|
||||
bundle_dir,
|
||||
_approved_verdict(bundle_dir),
|
||||
approver="persona",
|
||||
experiment="exp-E",
|
||||
timestamp="2026-06-30",
|
||||
)
|
||||
|
||||
promoted = [v for v in seed_store_from_bundle(bundle_dir).verdicts if _MARKER in v.rationale]
|
||||
assert len(promoted) == 1, f"expected exactly the promoted verdict, got {len(promoted)}"
|
||||
seeded = promoted[0].proposal_features
|
||||
|
||||
assert seeded.affected_codes == candidate.affected_codes
|
||||
assert seeded.measure_type == candidate.measure_type
|
||||
assert seeded.claimed_saving_nok == candidate.claimed_saving_nok
|
||||
assert type(seeded.claimed_saving_nok) is type(candidate.claimed_saving_nok), (
|
||||
f"magnitude type changed across the round trip: {seeded.claimed_saving_nok!r} vs "
|
||||
f"{candidate.claimed_saving_nok!r} — _mint_id hashes the raw value, so this splits the id"
|
||||
)
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue