test(experience): the id was pinned, the keying was not - and the two paths disagree
S3.2 measured with scripts/mutation_harness.py, denominator tests/ (963), every run sha256-restored. The mirroring question answerable offline: is today's boundary - "the id comes from the FILE, the keying comes from the BUNDLE" - load-bearing in both halves? The id half is pinned on both sides (re-minting RED, dropping the mint fallback RED). The keying half is green-but-dead: empty features and the gated C3.2 fix's own shape each left all 963 green. Nothing observed what a seeded verdict is keyed on. One rationale clause too: requiring BOTH learning fields, whose either-or form emits a marker naming the absent field as None (honesty, §1). New beyond C-F5: the sibling's drift form DOES exist here. A verdict is keyed in exactly two places by different rules - the bundle seed (bundle-wide) and the file/inbox path (per-verdict) - so the same verdict id lands in the same first-write-wins slot with a keying decided by LOAD ORDER. And promote writes no candidate features at all, which is why the fix's shape is a no-op against every fixture in the repo and had to be measured against one carrying the fields. Pinned by tests/test_experience_keying_loadbearing.py (11 tests, 963 -> 974, strict superset, 0 node ids lost). No src/ change: C-F5 is GATED on D-A pkt. 4, and these tests exist so the gated work must arrive as a visible red test and a decision. Two behaviour-preserving mutations carry their own controls (a third VerdictRecord site; hoisting the projection read into the loop). Dated under the D7 frame: work after 2026-08-09, never independent convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
544655b4c8
commit
1041687d31
2 changed files with 430 additions and 2 deletions
|
|
@ -35,10 +35,10 @@ spørringen som ble kjørt** — ikke fila den ble kjørt mot.
|
|||
## D7-speilingskøen
|
||||
|
||||
Åtte kandidater for speiling mellom D7-søsknene. **Ingen er besluttet** — de står som
|
||||
kandidater, ikke som planlagt arbeid. **Tre er målt, 5 står igjen:**
|
||||
kandidater, ikke som planlagt arbeid. **Fire er målt, 4 står igjen:**
|
||||
|
||||
- ~~S2.7~~ — **MÅLT 2026-09-07, se under**
|
||||
- S3.2
|
||||
- ~~S3.2~~ — **MÅLT 2026-09-12, se under**
|
||||
- S4.0 (`126807a`)
|
||||
- (p) `to_ore` — TO kallsteder
|
||||
- ~~(a)/(i) `unquote_scalar`~~ — **MÅLT 2026-08-31, se under**
|
||||
|
|
@ -191,6 +191,80 @@ spec-tekst rørt, og fasiten er ikke berørt.
|
|||
**Datering (D7-rammen):** arbeid ETTER 2026-08-09 — skal **ikke** leses som uavhengig konvergens.
|
||||
|
||||
|
||||
### S3.2 — id-en er pinnet, NØKLINGEN er ikke, og de to nøklings-stiene er uenige (målt 2026-09-12)
|
||||
|
||||
Kandidaten står i køen fordi søskenet fikset flerkandidat-ExpeL-seeding: frøet skal lese den
|
||||
dømte kandidaten fra verdict-filas EGEN frontmatter (promoteringen skriver den), med dagens
|
||||
nøkling som fallback. Feature-halvdelen av den defekten er bekreftet her som C-F5
|
||||
(`docs/review-2026-07.md`) og **GATET** på D-A pkt. 4 + commons-pull (paritetsplanens rad 14),
|
||||
mens id-halvdelen alt er løst: en lastet `verdict_id` leses VERBATIM (§4.2). Speilings-
|
||||
spørsmålet som KAN besvares offline i dag er derfor: **er dagens grense — «id-en kommer fra
|
||||
FILA, nøklingen kommer fra BUNDELEN» — load-bearing i BEGGE halvdeler?**
|
||||
|
||||
**Populasjonen først.** En dom nøkles i nøyaktig TO steder i `src/`, med hver sin spec-rolle:
|
||||
bundle-frøet (`experience.py:146`, §3 Steg 1 — nøklet på bundelens ENE IR-projeksjon) og
|
||||
fil/inbox-stien (`inbox.py:105`, §4.2 — nøklet på dommens egne `proposal_features`). Ingen
|
||||
tredje konstruksjon finnes (AST-målt over hele pakken; positiv kontroll: samme AST-spørring
|
||||
finner `VerdictStore`-konstruksjon i tre moduler, så spørringen KAN finne). **Og her finnes
|
||||
søskenets drift-form — i motsetning til S2.7 og `unquote_scalar`:** de to stiene nøkler ETTER
|
||||
ULIKE REGLER.
|
||||
|
||||
**Dekningen deler seg rent mellom id og nøkling.** Med `scripts/mutation_harness.py`, nevner
|
||||
`tests/` (hele suiten, 963 tester), hver kjøring sha256-restaurert:
|
||||
|
||||
- Å detache sømmen (frøet itererer ingenting) er **RØD**; det samme er type-filteret,
|
||||
decision-defaulten og `description`-lesningen.
|
||||
- **ID-halvdelen er pinnet på BEGGE sider:** å re-minte i stedet for å lese frontmatter-id-en
|
||||
er RØD (`test_step8`), og å droppe mint-fallbacken — slik at ikke-promoterte dommer nøkles på
|
||||
en tom id — er også RØD.
|
||||
- **NØKLINGS-halvdelen er grønn-men-død.** Å bytte bundle-projeksjonen mot tomme features lot
|
||||
alle 963 stå grønne, og det gjorde også den gatede fiksens EGEN form (les kandidaten fra
|
||||
verdict-filas frontmatter, fallback dagens nøkling). **Ingenting i suiten observerte hva en
|
||||
seedet dom er nøklet PÅ.**
|
||||
- Én klausul i rasjonalet var grønn-men-død i tillegg: kravet om BEGGE læringsfeltene før
|
||||
markøren emitteres. Svekket til enten-eller forble suiten grønn — og konsekvensen er en
|
||||
ærlighets-defekt (§1): markøren navngir da det fraværende feltet som `None`.
|
||||
|
||||
**Mutasjonene ble vist å endre oppførsel FØR deres grønne ble lest som hull** (fellen fra økt
|
||||
39). Tom nøkling senker den seedede dommens similarity mot sin egen bundle-projeksjon fra 1.00
|
||||
til 0.15 — observerbart, og rangerings-relevant i det øyeblikket en andre oppføring finnes.
|
||||
C-F5s rest-defekt er reprodusert live: to promoterte dommer om ULIKE kandidater (LED vs.
|
||||
ventilasjon) seedes begge nøklet på bundelens ENE projeksjon, så en spørring som bærer
|
||||
ventilasjons-kandidatens egne features scorer BEGGE til 0.0 og avgjøres av hex-id-rekkefølge,
|
||||
ikke av struktur.
|
||||
|
||||
**NYTT UTOVER C-F5 — nøklingen er sti-avhengig, og dermed rekkefølge-avhengig.** Samme dom —
|
||||
samme kandidat, samme mintede id (`c8c97c9d992229dc`), altså samme first-write-wins-slot —
|
||||
nøkles på `VENT-AGG-01`/250 000 via inbox-stien og på `ENERGI-TOTAL-EL`/30 000 via bundle-frøet.
|
||||
Hvilken nøkling som overlever avgjøres av LASTE-REKKEFØLGEN, ingenting annet. Og fiksen kan
|
||||
ikke være ensidig: `promote` skriver `verdict_id` og `description`, men **ingen** kandidat-felt
|
||||
(målt frontmatter-nøkkelsett: `decision`, `description`, `provenance`, `tags`, `title`, `type`,
|
||||
`verdict_id`) — hvilket er grunnen til at fiksens egen form er en **no-op mot enhver fixture i
|
||||
repoet i dag**, og derfor måtte måles mot en fixture som bærer feltene.
|
||||
|
||||
**Pinnet av** `tests/test_experience_keying_loadbearing.py` (11 tester, 963 → 974, strengt
|
||||
supersett: 0 tapte node-id-er). Value-beviset er kjørt, ikke påstått: hver test er grønn før og
|
||||
rød etter nøyaktig sin egen mutasjon, med fold-markøren grønn i BEGGE kjøringer — som mekanisk
|
||||
viser at mutasjonen flyttet NØKLINGEN, ikke sømmen. To sammensatte tester ble SPLITTET fordi en
|
||||
rød test bare beviser sin FØRSTE assert (§1-halvdelen er i tillegg pinnet med `--red-at None`).
|
||||
|
||||
**To kontroller utover oppførselen.** (1) Populasjonskontrollen (AST) ble bevist mot en
|
||||
oppførselsbevarende mutasjon — en tredje, ubrukt `VerdictRecord`-konstruksjon — med alle
|
||||
oppførselstestene grønne. (2) Å heve projeksjons-lesningen INN i seedings-loopen er likeledes
|
||||
oppførselsbevarende (tre oppførselskontroller grønne), og er samtidig første halvdel av den
|
||||
gatede fiksen; derfor er formen pinnet strukturelt (`load_validator_input` kalles nøyaktig ÉN
|
||||
gang i `seed_store_from_bundle`) eller ikke i det hele tatt.
|
||||
|
||||
**Ærlig grense — hva dette IKKE sier.** Å pinne at alle dommer i en bundle deler ÉN nøkling i
|
||||
dag er ingen godkjenning; C-F5 kaller den en MAJOR spec-nivå-defekt, og fiksen er GATET, ikke
|
||||
avvist. Testene pinner grensen slik at det gatede arbeidet MÅ ankomme som en synlig rød test og
|
||||
en D-A pkt. 4-beslutning — samme ratchet-rolle `test_ingest_stamp_conformance_loadbearing.py`
|
||||
har oppstrøms. Ingen `src/`-endring er gjort, ingen spec-tekst rørt, og **fasiten er ikke
|
||||
involvert i nøklingen i det hele tatt** — hvilket er nettopp derfor den ikke kunne hjelpe.
|
||||
|
||||
**Datering (D7-rammen):** arbeid ETTER 2026-08-09 — skal **ikke** leses som uavhengig konvergens.
|
||||
|
||||
|
||||
Rammen rundt køen: å lese søskenets kode er tillatt (`3bdf7f0`), men kopiering skal kun skje
|
||||
der det tjener løsningen, aldri som snarvei. **Uavhengighets-beviset er DATERT** t.o.m.
|
||||
2026-08-09; arbeid etter den datoen kan ikke leses som uavhengig konvergens.
|
||||
|
|
|
|||
354
tests/test_experience_keying_loadbearing.py
Normal file
354
tests/test_experience_keying_loadbearing.py
Normal file
|
|
@ -0,0 +1,354 @@
|
|||
"""S3.2: the ExpeL seeding seam — is WHAT EACH VERDICT IS KEYED ON load-bearing?
|
||||
|
||||
The D7 mirroring queue (``docs/2026-08-18-funn-koeer-og-gater.md § D7-speilingskøen``)
|
||||
carries S3.2 because the MAF sibling fixed multi-candidate ExpeL seeding: the seed
|
||||
should read the judged candidate from the verdict file's OWN frontmatter (promotion
|
||||
writing it), with today's keying as fallback. On our side the feature half of that
|
||||
defect is confirmed as C-F5 (``docs/review-2026-07.md``) and GATED on D-A pkt. 4 + a
|
||||
commons pull (parity plan row 14), while the id half is already solved here — a
|
||||
loaded ``verdict_id`` is read VERBATIM (§4.2). So this file does NOT build the fix.
|
||||
It answers the question that is answerable offline today: **is today's boundary —
|
||||
"the id comes from the FILE, the keying comes from the BUNDLE" — load-bearing in
|
||||
BOTH halves?**
|
||||
|
||||
MEASURED 2026-09-12 with ``scripts/mutation_harness.py``, denominator ``tests/``
|
||||
(the whole suite, 963 tests), every run restored sha256-verified:
|
||||
|
||||
- Detaching the seam outright (seeding iterates nothing) is RED. So is dropping the
|
||||
type filter, the decision default, and the ``description`` read.
|
||||
- The ID HALF is pinned on both sides: re-minting instead of reading the frontmatter
|
||||
id is RED (``test_step8``), and dropping the mint fallback so unpromoted verdicts
|
||||
key on an empty id is RED too.
|
||||
- The KEYING HALF is GREEN-BUT-DEAD. Replacing the bundle projection with empty
|
||||
features left all 963 tests green, and so did the gated fix's own shape (read the
|
||||
candidate from the verdict file's frontmatter, fallback today's keying). Nothing in
|
||||
the suite observes what a seeded verdict is keyed ON.
|
||||
- One clause of the rationale is green-but-dead as well: requiring BOTH learning
|
||||
fields before emitting the marker. Weakening it to either-or left the suite green,
|
||||
and the consequence is an honesty defect (§1) — a marker naming a field as None.
|
||||
|
||||
The mutations were shown to change behaviour before their greens were read as holes
|
||||
(the økt 39 trap: the harness reports a no-op and an undetected seam identically).
|
||||
Empty keying drops the seeded record's similarity against its own bundle projection
|
||||
from 1.00 to 0.15 — observable, and ranking-relevant the moment a second entry
|
||||
exists. C-F5's rest-defect was reproduced live: two promoted verdicts about distinct
|
||||
candidates (LED vs ventilation) both seed keyed on the bundle's ONE projection, so a
|
||||
query carrying the ventilation candidate's own features scores BOTH at 0.0 and the
|
||||
tie falls to hex id order, not to structure.
|
||||
|
||||
NEW BEYOND C-F5 — the sibling's drift form DOES exist here, across two paths. A
|
||||
verdict record is keyed in exactly two places with different rules: the bundle seed
|
||||
(``experience.py``, keyed bundle-wide) and the file/inbox path (``inbox.py``, keyed
|
||||
on the verdict's OWN ``proposal_features``). The same verdict — same candidate, same
|
||||
minted id, hence the same first-write-wins store slot — is therefore keyed
|
||||
DIFFERENTLY depending on which path reached the store first. And the fix cannot be
|
||||
one-sided: ``promote`` writes ``verdict_id`` and ``description`` but no candidate
|
||||
features at all, which is why the fix's own shape is invisible to every fixture in
|
||||
this repo today.
|
||||
|
||||
HONEST LIMIT — what this does NOT say. Pinning that all verdicts in a bundle share
|
||||
one keying TODAY is not an endorsement; C-F5 calls it a MAJOR spec-level defect and
|
||||
the fix is gated, not declined. These tests pin the boundary so the gated work must
|
||||
arrive as a visible red test and a decision, never as a silent swap — the same
|
||||
ratchet role ``test_ingest_stamp_conformance_loadbearing.py`` plays upstream. No
|
||||
``src/`` change is made here, no spec text is touched, and the golden fixture is not
|
||||
involved in the keying at all (which is precisely why it cannot help).
|
||||
|
||||
Dated under the D7 frame: this is work AFTER 2026-08-09 and must NOT be read as
|
||||
independent convergence with the sibling.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from portfolio_optimiser_claude.experience import (
|
||||
CandidateFeatures,
|
||||
VerdictRecord,
|
||||
VerdictStore,
|
||||
mint_verdict_id,
|
||||
seed_store_from_bundle,
|
||||
similarity,
|
||||
)
|
||||
from portfolio_optimiser_claude.inbox import VerdictDocument
|
||||
from portfolio_optimiser_claude.ir import load_validator_input
|
||||
from portfolio_optimiser_claude.promotion import promote
|
||||
|
||||
SRC_PKG = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser_claude"
|
||||
SHARED_BUNDLE = Path(__file__).resolve().parents[1] / "shared" / "examples" / "bygg-energi-mikro"
|
||||
|
||||
# A candidate that is NOT the bundle's: other codes, other measure type, other
|
||||
# magnitude bucket. Keyed on itself it sits at similarity 0.00 from the bundle
|
||||
# projection, which is what makes a miskeying observable at all.
|
||||
OTHER = CandidateFeatures(
|
||||
affected_codes=frozenset({"VENT-AGG-01"}),
|
||||
measure_type="ventilation-schedule",
|
||||
claimed_saving_nok=250_000.0,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def bundle(tmp_path: Path) -> Path:
|
||||
target = tmp_path / "bundle"
|
||||
shutil.copytree(SHARED_BUNDLE, target)
|
||||
return target
|
||||
|
||||
|
||||
def _bundle_keying(bundle_dir: Path) -> CandidateFeatures:
|
||||
return CandidateFeatures.from_proposal(load_validator_input(bundle_dir))
|
||||
|
||||
|
||||
def _link(bundle_dir: Path, filename: str) -> None:
|
||||
index = bundle_dir / "index.md"
|
||||
index.write_text(index.read_text(encoding="utf-8") + f"- [Dom]({filename})\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _write_verdict(bundle_dir: Path, filename: str, extra_frontmatter: str) -> None:
|
||||
(bundle_dir / filename).write_text(
|
||||
"---\n"
|
||||
"type: verdict\n"
|
||||
"title: Dom\n"
|
||||
"decision: approved\n"
|
||||
"description: Dom om en annen kandidat.\n"
|
||||
f"{extra_frontmatter}"
|
||||
"tags: [verdict]\n"
|
||||
"---\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
_link(bundle_dir, filename)
|
||||
|
||||
|
||||
def _self_describing_verdict(bundle_dir: Path, features: CandidateFeatures) -> str:
|
||||
"""A verdict file carrying the candidate it judges IN ITS OWN frontmatter.
|
||||
|
||||
This is the INPUT the gated C3.2 fix reads. Today's seeding ignores it, and
|
||||
``promote`` does not write it — both halves are pinned below.
|
||||
"""
|
||||
verdict_id = mint_verdict_id(features)
|
||||
_write_verdict(
|
||||
bundle_dir,
|
||||
"verdict-self-describing.md",
|
||||
f"verdict_id: {verdict_id}\n"
|
||||
f"measure_type: {features.measure_type}\n"
|
||||
f"claimed_saving_nok: {features.claimed_saving_nok}\n",
|
||||
)
|
||||
return verdict_id
|
||||
|
||||
|
||||
def _record_for(store: VerdictStore, features: CandidateFeatures, verdict_id: str) -> VerdictRecord:
|
||||
return next(r for r in store.retrieve(features, k=20) if r.verdict_id == verdict_id)
|
||||
|
||||
|
||||
def _seeded(bundle_dir: Path) -> VerdictStore:
|
||||
store = VerdictStore()
|
||||
seed_store_from_bundle(store, bundle_dir)
|
||||
return store
|
||||
|
||||
|
||||
def _document(features: CandidateFeatures, rationale: str) -> VerdictDocument:
|
||||
return VerdictDocument.from_candidate(
|
||||
features, decision="approved", rationale=rationale, description="d"
|
||||
)
|
||||
|
||||
|
||||
class TestThereAreExactlyTwoKeyingSites:
|
||||
"""The population control, and it runs first: every clause below is worth only as
|
||||
much as the claim that there are exactly TWO places where a verdict is keyed."""
|
||||
|
||||
def test_a_verdict_record_is_keyed_in_exactly_two_places_with_distinct_roles(self) -> None:
|
||||
# The sibling's S3.2 shape made measurable here. Two keying sites exist BY
|
||||
# DESIGN and answer different specs: the bundle seed (§3 Step 1, keyed on the
|
||||
# bundle's ONE IR projection) and the file/inbox path (§4.2, keyed on the
|
||||
# verdict's own proposal_features). RED the day a third appears — a verdict
|
||||
# keyed somewhere no one is looking is how the miskeying spreads. Read from
|
||||
# the AST, not prose.
|
||||
sites: dict[str, int] = {}
|
||||
for module in sorted(SRC_PKG.glob("*.py")):
|
||||
tree = ast.parse(module.read_text(encoding="utf-8"))
|
||||
found = [
|
||||
node
|
||||
for node in ast.walk(tree)
|
||||
if isinstance(node, ast.Call)
|
||||
and isinstance(node.func, ast.Name)
|
||||
and node.func.id == "VerdictRecord"
|
||||
]
|
||||
if found:
|
||||
sites[module.name] = len(found)
|
||||
assert sites == {"experience.py": 1, "inbox.py": 1}, (
|
||||
f"a verdict record is now keyed in {sites} — the two-path design "
|
||||
"(bundle seed + file/inbox) has drifted"
|
||||
)
|
||||
|
||||
def test_the_bundle_projection_is_read_once_per_bundle_not_once_per_file(self) -> None:
|
||||
# Behaviour-preserving guard: hoisting the projection read INTO the seeding
|
||||
# loop changes no result today, and it is the first half of the gated fix.
|
||||
# Every behaviour test below stays green under that move, so the shape of the
|
||||
# read is pinned structurally or not at all.
|
||||
tree = ast.parse((SRC_PKG / "experience.py").read_text(encoding="utf-8"))
|
||||
seed_fn = next(
|
||||
node
|
||||
for node in ast.walk(tree)
|
||||
if isinstance(node, ast.FunctionDef) and node.name == "seed_store_from_bundle"
|
||||
)
|
||||
reads = [
|
||||
node
|
||||
for node in ast.walk(seed_fn)
|
||||
if isinstance(node, ast.Call)
|
||||
and isinstance(node.func, ast.Name)
|
||||
and node.func.id == "load_validator_input"
|
||||
]
|
||||
assert len(reads) == 1, "the bundle projection must be read once, not per verdict file"
|
||||
|
||||
def test_keying_is_observable_at_all(self, bundle: Path) -> None:
|
||||
# POSITIVE CONTROL FOR EVERY NEGATIVE BELOW. Each test that follows asserts
|
||||
# WHICH features a record carries; that is vacuous unless a different keying
|
||||
# demonstrably scores differently on this bundle.
|
||||
keying = _bundle_keying(bundle)
|
||||
assert similarity(keying, keying) == pytest.approx(1.0)
|
||||
assert similarity(OTHER, keying) == pytest.approx(0.0)
|
||||
|
||||
|
||||
class TestWhatTheSeedKeysOnIsLoadBearing:
|
||||
"""One test per clause of ``features=features`` in ``seed_store_from_bundle`` —
|
||||
each proved RED under exactly the mutation it names, with the fold's realization
|
||||
marker green in both runs (the mutation moved the KEYING, not the seam)."""
|
||||
|
||||
def test_a_seeded_verdict_is_keyed_on_the_bundle_ir_projection(self, bundle: Path) -> None:
|
||||
# RED under replacing the keying with empty features. Measured: that drops
|
||||
# the record's similarity against its own bundle projection from 1.00 to
|
||||
# 0.15, so the mutation is behaviour-changing, not a no-op.
|
||||
keying = _bundle_keying(bundle)
|
||||
(record,) = _seeded(bundle).retrieve(keying, k=1)
|
||||
assert record.features == keying
|
||||
assert similarity(record.features, keying) == pytest.approx(1.0)
|
||||
|
||||
def test_the_verdict_files_own_candidate_frontmatter_is_ignored_today(
|
||||
self, bundle: Path
|
||||
) -> None:
|
||||
# The GATED C3.2 fix, pinned as a boundary: a verdict file may carry the
|
||||
# candidate it judges, and today's seeding still keys it on the bundle. RED
|
||||
# the day the seed reads those fields — which is the point. The fix must
|
||||
# arrive as this red test plus a D-A pkt. 4 decision, never as a silent swap.
|
||||
verdict_id = _self_describing_verdict(bundle, OTHER)
|
||||
keying = _bundle_keying(bundle)
|
||||
record = _record_for(_seeded(bundle), keying, verdict_id)
|
||||
assert record.features == keying
|
||||
assert record.features != OTHER
|
||||
assert record.features.measure_type != OTHER.measure_type
|
||||
assert record.features.claimed_saving_nok != OTHER.claimed_saving_nok
|
||||
|
||||
def test_a_multi_candidate_bundle_keys_every_verdict_on_the_one_projection(
|
||||
self, bundle: Path
|
||||
) -> None:
|
||||
# C-F5's rest-defect, reproduced as a test: two promoted verdicts about
|
||||
# DISTINCT candidates both seed keyed on the bundle's one projection, so
|
||||
# retrieval cannot tell them apart — a query carrying the ventilation
|
||||
# candidate's own features scores every entry identically and the tie falls
|
||||
# to hex id order. RED under any per-verdict keying.
|
||||
promote(
|
||||
_document(OTHER, "dom om ventilasjon"),
|
||||
bundle,
|
||||
approved_by="e",
|
||||
experiment="x",
|
||||
timestamp="2026-09-12",
|
||||
)
|
||||
store = _seeded(bundle)
|
||||
keying = _bundle_keying(bundle)
|
||||
assert len(store) == 2
|
||||
assert {r.features for r in store.retrieve(keying, k=20)} == {keying}
|
||||
scores = [similarity(r.features, OTHER) for r in store.retrieve(OTHER, k=20)]
|
||||
assert scores == [pytest.approx(0.0), pytest.approx(0.0)]
|
||||
|
||||
def test_the_same_verdict_id_is_keyed_differently_by_the_two_paths(self, bundle: Path) -> None:
|
||||
# The sibling's drift form, measured on our side: the SAME verdict — same
|
||||
# candidate, same minted id, hence the same first-write-wins store slot — is
|
||||
# keyed on the bundle via the seed and on its own candidate via the inbox.
|
||||
verdict_id = _self_describing_verdict(bundle, OTHER)
|
||||
document = _document(OTHER, "dom om ventilasjon")
|
||||
assert document.id == verdict_id, "same candidate must mint the same id (§4.2)"
|
||||
keying = _bundle_keying(bundle)
|
||||
via_bundle = _record_for(_seeded(bundle), keying, verdict_id)
|
||||
via_inbox = document.to_record()
|
||||
assert via_bundle.features != via_inbox.features
|
||||
assert via_bundle.features == keying
|
||||
assert via_inbox.features == OTHER
|
||||
|
||||
def test_which_keying_survives_is_decided_by_load_order(self, bundle: Path) -> None:
|
||||
# The consequence of two rules over one store slot: FIRST-write-wins (§4.2)
|
||||
# makes the surviving keying a function of arrival order, not of the verdict.
|
||||
# Split from the test above deliberately — a red test proves only its first
|
||||
# failing assert, so the order claim needs its own.
|
||||
verdict_id = _self_describing_verdict(bundle, OTHER)
|
||||
via_inbox = _document(OTHER, "dom om ventilasjon").to_record()
|
||||
keying = _bundle_keying(bundle)
|
||||
|
||||
bundle_first = _seeded(bundle)
|
||||
bundle_first.add(via_inbox)
|
||||
inbox_first = VerdictStore()
|
||||
inbox_first.add(via_inbox)
|
||||
seed_store_from_bundle(inbox_first, bundle)
|
||||
|
||||
assert _record_for(bundle_first, keying, verdict_id).features == keying
|
||||
assert _record_for(inbox_first, keying, verdict_id).features == OTHER
|
||||
|
||||
|
||||
class TestTheLearningMarkerNeedsBothFields:
|
||||
"""§1 honesty rule at the seeding layer: the marker names two numbers, so it is
|
||||
emitted only when the verdict file supplies both."""
|
||||
|
||||
def test_a_half_specified_learning_pair_emits_no_marker(self, bundle: Path) -> None:
|
||||
# RED under weakening the pair requirement to either-or: the rationale then
|
||||
# carries a marker built from the one field that IS there.
|
||||
record = self._half_pair_record(bundle)
|
||||
assert "realiseringsgrad" not in record.rationale
|
||||
assert record.rationale == "Dom om en annen kandidat."
|
||||
|
||||
def test_a_half_specified_learning_pair_never_fabricates_a_missing_number(
|
||||
self, bundle: Path
|
||||
) -> None:
|
||||
# The §1 half of the same clause, asserted FIRST in its own test: under
|
||||
# either-or the emitted marker names the absent field as None — a fabricated
|
||||
# learning signal, which is the one thing the honesty rule forbids outright.
|
||||
assert "None" not in self._half_pair_record(bundle).rationale
|
||||
|
||||
@staticmethod
|
||||
def _half_pair_record(bundle_dir: Path) -> VerdictRecord:
|
||||
_write_verdict(
|
||||
bundle_dir, "verdict-half.md", "verdict_id: halfpair\nrealization_rate: 0.5\n"
|
||||
)
|
||||
return _record_for(_seeded(bundle_dir), _bundle_keying(bundle_dir), "halfpair")
|
||||
|
||||
|
||||
class TestTheGatedFixCannotBeOneSided:
|
||||
"""The fix reads per-verdict features; nothing writes them. Both halves pinned."""
|
||||
|
||||
def test_promotion_writes_no_candidate_features_today(self, bundle: Path) -> None:
|
||||
# §6 keeps the promoted file minimal: id, decision, rationale-as-description,
|
||||
# provenance. The candidate it judges is NOT among them, so the C3.2 fix must
|
||||
# change promotion too. RED the day promotion starts writing them — the
|
||||
# deliberate signal that the gated work has begun on this side.
|
||||
path = promote(
|
||||
_document(OTHER, "dom om ventilasjon"),
|
||||
bundle,
|
||||
approved_by="e",
|
||||
experiment="x",
|
||||
timestamp="2026-09-12",
|
||||
)
|
||||
keys = {
|
||||
line.split(":", 1)[0]
|
||||
for line in path.read_text(encoding="utf-8").splitlines()
|
||||
if ":" in line
|
||||
}
|
||||
assert keys == {
|
||||
"type",
|
||||
"title",
|
||||
"decision",
|
||||
"description",
|
||||
"verdict_id",
|
||||
"provenance",
|
||||
"tags",
|
||||
}
|
||||
Loading…
Add table
Add a link
Reference in a new issue