feat(p17b): a context set that spans TWO bases, and a judge told which one [skip-docs]
``contexts/dekke-og-kontrakt-lindaas-2027`` is the first set whose approaches route at more than one knowledge base: a1/a2 at n200-2024 (material requirements) and a3/a4 at r761-2025 (the rig, and the falsification arm). That is the whole reason it exists -- P17b measures that ONE commission can be run across several. ``bundle.txt`` grows a block per base; a set naming one base is one block, so the four pre-P17b files parse byte-identically. The reader now has ONE home (``stress.read_bundle_declarations``): it used to be a private copy in the P14 gate and a second, looser one inside ``stress.main``, and the multi-base form is exactly the change that would have let them drift. Rule U becomes the UNION of every declared base, and that is not a formality. MEASURED 15.09: ``enhetspris`` is absent from n200-2024 and carried by 70 of r761-2025's 2 756 concepts, so anchors admitted per base would have admitted a question the pass as a whole CAN ground. It was dropped from the fifth set's anchors for that reason. ``score_context_set(bundle_id=...)`` restricts the judgement to the approaches routed at THIS base. Without it, judging the n200 outbox reports the r761 approach as ``not_evaluated``/``absent`` -- a false finding, because that approach WAS evaluated, against the other base, under the other run_id. That defect is pinned by its own arm. The judge's CLI refuses to guess when a set declares several bases, with an rc-0 control on ``--bundle``. Arm (d) gained a second half: every DECLARED base must be named by some approach, because a base no approach names is never run. The P19/B2 fasit denominator moved 26 -> 32 and is asserted, not dropped: six new references, two of them bare ``prosessnr`` (12.11, 12.12), so B1's punctuation-and-digits form is now exercised by a fasit and not only by a known-positive. Suite 1774/5, golden byte-unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
5e4c497a84
commit
da0ccd0489
8 changed files with 496 additions and 57 deletions
4
contexts/dekke-og-kontrakt-lindaas-2027/bundle.txt
Normal file
4
contexts/dekke-og-kontrakt-lindaas-2027/bundle.txt
Normal file
|
|
@ -0,0 +1,4 @@
|
||||||
|
name: n200-2024
|
||||||
|
bundle_id: vegnormal-n200-2024
|
||||||
|
name: r761-2025
|
||||||
|
bundle_id: vegnormal-r761-2025
|
||||||
7
contexts/dekke-og-kontrakt-lindaas-2027/docs/README.md
Normal file
7
contexts/dekke-og-kontrakt-lindaas-2027/docs/README.md
Normal file
|
|
@ -0,0 +1,7 @@
|
||||||
|
# Prosjektdokumenter
|
||||||
|
|
||||||
|
TOM med vilje. Scenarioet er MS Office + PDF, men ingen av dem kan mates
|
||||||
|
inn her: `--docs-dir`-omveien er fraradet (P13 - den omgar stigen `list_bundles -> read_bundle
|
||||||
|
-> read_dir -> read_file` og de to gatene, §4.1a-dimensjonen og verdict-laget). Veien et ekte
|
||||||
|
prosjektdokument skal ta er gjennom okf-ingest inn i en kunnskapsbase, altsa en base til - ikke
|
||||||
|
en katalog ved siden av. Se `docs/2026-09-12-p14-kontekstsett.md` § 1.4.
|
||||||
68
contexts/dekke-og-kontrakt-lindaas-2027/fasit.json
Normal file
68
contexts/dekke-og-kontrakt-lindaas-2027/fasit.json
Normal file
|
|
@ -0,0 +1,68 @@
|
||||||
|
{
|
||||||
|
"project_id": "dekke-og-kontrakt-lindaas-2027",
|
||||||
|
"must_cite": [
|
||||||
|
{
|
||||||
|
"approach_id": "a1-tynnere-forsterkningslag",
|
||||||
|
"rationale": "Overbygningen er prosjektert med full forsterkningslagstykkelse over hele strekningen. Et riktig svar må gjengi hva N200 krever av forsterkningslag før tykkelsen kan reduseres.",
|
||||||
|
"concepts": [
|
||||||
|
{
|
||||||
|
"path": "krav/N200/id-13c94f7a-2d24-48a1-b258-db66cb2392d7.md",
|
||||||
|
"title": "Krav 3.3.3—1 Forsterkningslag",
|
||||||
|
"ref": "Krav 3.3.3—1"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"path": "krav/N200/id-222bca41-dff8-4b54-fb28-88da74969c9c.md",
|
||||||
|
"title": "Krav 4.6—1 Forsterkningslag",
|
||||||
|
"ref": "Krav 4.6—1"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"approach_id": "a2-filterlag-sprengstein",
|
||||||
|
"rationale": "Filterlaget er prosjektert med innkjøpt sortert materiale. Et riktig svar må gjengi hva N200 krever av filterlag før stedlig sprengstein kan vurderes.",
|
||||||
|
"concepts": [
|
||||||
|
{
|
||||||
|
"path": "krav/N200/id-6d7520d6-d224-4e49-f7ba-22863b49cf92.md",
|
||||||
|
"title": "Krav 4.3—1 Filterlag",
|
||||||
|
"ref": "Krav 4.3—1"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"path": "krav/N200/id-4f1ff497-68a0-4953-c946-3bdf49e16516.md",
|
||||||
|
"title": "Krav 4.3—5 Filterlag",
|
||||||
|
"ref": "Krav 4.3—5"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"approach_id": "a3-riggomfang",
|
||||||
|
"rationale": "Riggen er priset som en frittstående etablering. Et riktig svar må gjengi hva R761 Prosesskoden legger i tilrigging og i drift av rigg, slik at et delt omfang kan beskrives uten at noe faller mellom to prosesser.",
|
||||||
|
"concepts": [
|
||||||
|
{
|
||||||
|
"path": "R761/12-11/_1_id-377d3c43-eb37-4337-a476-0818a6a30490.md",
|
||||||
|
"title": "Tilrigging",
|
||||||
|
"ref": "12.11"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"path": "R761/12-12/_1_id-0d8e750e-18b1-49b9-b838-f5c29f21ff29.md",
|
||||||
|
"title": "Drift av rigg og midlertidige bygninger",
|
||||||
|
"ref": "12.12"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"honesty": "Prosjektet fv. 218 Lindaas er KONSTRUERT av meg: vegnummer, lengde, AADT og alle fire kostlinjene (LIND-FORST-01, LIND-FILT-01, LIND-RIGG-01, LIND-INDEKS-01) er oppdiktet, og beloepene er satte stoerrelsesordener. Kravene og prosessene i must_cite er lest ordrett ut av basenes egen frontmatter (n200-2024 for a1/a2, r761-2025 for a3). Verken N200 eller R761 baerer priser, saa kodene finnes ikke i noen av basene. Dette er det FOERSTE settet som spenner TO baser: a1/a2 rutes mot n200-2024 og a3/a4 mot r761-2025, og det er hele grunnen til at settet finnes -- P17b maaler at EN kommisjon kan kjoeres over flere kunnskapsbaser. Den fjerde tilnaermingen a4-indeksregulering og dens kostkode LIND-INDEKS-01 er ogsaa KONSTRUERT, og med vilje: den er falsifiseringsarmen, en kostlinje INGEN av de to basene baerer grunnlaget for. MAALT 15.09: 'enhetspris' ble FORKASTET som anker fordi r761-2025 baerer ordet i 70 av 2 756 konsepter -- et anker som holder for ett sett med EN base holder ikke noedvendigvis for et sett med to.",
|
||||||
|
"must_refuse": [
|
||||||
|
{
|
||||||
|
"approach_id": "a4-indeksregulering",
|
||||||
|
"anchors": [
|
||||||
|
"indeksregulering",
|
||||||
|
"konsumprisindeks",
|
||||||
|
"markedspris",
|
||||||
|
"tonnpris",
|
||||||
|
"kalkyle",
|
||||||
|
"prisstigning"
|
||||||
|
],
|
||||||
|
"rationale": "N200 beskriver materialkrav og R761 Prosesskoden beskriver hva en prosess omfatter -- ingen av dem regulerer priser. Verken indeksregulering, konsumprisindeks, markedspris, tonnpris, kalkyle eller prisstigning finnes i noen av de to basene. Spoersmaalene denne approachen staar for: Hvilken indeks skal kontraktssummen reguleres etter? | Hva er markedsprisen paa sprengstein i dette omraadet? | Hvilken prisstigning er lagt til grunn i kalkylen?"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
46
contexts/dekke-og-kontrakt-lindaas-2027/mandate.json
Normal file
46
contexts/dekke-og-kontrakt-lindaas-2027/mandate.json
Normal file
|
|
@ -0,0 +1,46 @@
|
||||||
|
{
|
||||||
|
"objective": "Kutt kostnad i vegprosjektet fv. 218 Lindås (4,1 km, ÅDT 2 100) uten å bryte et eneste materialkrav i N200 eller å beskrive riggen annerledes enn R761 Prosesskoden gjør.",
|
||||||
|
"success_criteria": "Minst én tilnærming validerer, og hver validerte tilnærming peker på kravet eller prosessen i SIN EGEN base som faktisk binder den — materialkravene i N200, riggomfanget i R761.",
|
||||||
|
"approaches": [
|
||||||
|
{
|
||||||
|
"id": "a1-tynnere-forsterkningslag",
|
||||||
|
"label": "Tynnere forsterkningslag på strekningen med fast fjell",
|
||||||
|
"description": "Overbygningen er prosjektert med full forsterkningslagstykkelse over hele strekningen, også der undergrunnen er fast fjell. Vi vil vite hva N200 faktisk krever av forsterkningslag.",
|
||||||
|
"affected_codes": [
|
||||||
|
"LIND-FORST-01"
|
||||||
|
],
|
||||||
|
"claimed_saving_nok": 2100000.0,
|
||||||
|
"bundle_id": "vegnormal-n200-2024"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "a2-filterlag-sprengstein",
|
||||||
|
"label": "Filterlag av stedlig sprengstein i stedet for innkjøpt materiale",
|
||||||
|
"description": "Filterlaget er prosjektert med innkjøpt sortert materiale. Spørsmålet er hvilke krav N200 stiller til filterlag, og om stedlig sprengstein kan tilfredsstille dem.",
|
||||||
|
"affected_codes": [
|
||||||
|
"LIND-FILT-01"
|
||||||
|
],
|
||||||
|
"claimed_saving_nok": 1350000.0,
|
||||||
|
"bundle_id": "vegnormal-n200-2024"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "a3-riggomfang",
|
||||||
|
"label": "Redusert riggomfang: felles rigg med naboentreprisen",
|
||||||
|
"description": "Riggen er priset som en frittstående etablering. Vi vil vite hva R761 Prosesskoden legger i tilrigging og drift av rigg, slik at omfanget kan deles med naboentreprisen uten at noe faller mellom to prosesser.",
|
||||||
|
"affected_codes": [
|
||||||
|
"LIND-RIGG-01"
|
||||||
|
],
|
||||||
|
"claimed_saving_nok": 1750000.0,
|
||||||
|
"bundle_id": "vegnormal-r761-2025"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "a4-indeksregulering",
|
||||||
|
"label": "Lavere indeksregulering av kontraktssummen",
|
||||||
|
"description": "Vi vil kutte ved å legge en lavere indeksregulering og en lavere markedspris til grunn for kontraktssummen.",
|
||||||
|
"affected_codes": [
|
||||||
|
"LIND-INDEKS-01"
|
||||||
|
],
|
||||||
|
"claimed_saving_nok": 900000.0,
|
||||||
|
"bundle_id": "vegnormal-r761-2025"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
|
@ -183,6 +183,41 @@ class ContextSetVerdict:
|
||||||
return asdict(self)
|
return asdict(self)
|
||||||
|
|
||||||
|
|
||||||
|
def read_bundle_declarations(path: str | Path) -> tuple[dict[str, str], ...]:
|
||||||
|
"""Parse a context set's ``bundle.txt`` into ONE declaration per knowledge base.
|
||||||
|
|
||||||
|
``key: value`` lines, nothing else; each ``name:`` OPENS a block and each block must close
|
||||||
|
with its own ``bundle_id:``. A set naming ONE base is one block, so every file written before
|
||||||
|
P17b parses byte-identically — the multi-base form (P17b DEL 2) is an extension, not a new
|
||||||
|
format.
|
||||||
|
|
||||||
|
**The ONE reader.** Before P17b the rule had two private copies — one in ``stress.main``'s
|
||||||
|
argument parsing, one in the P14 gate's own test file — and the multi-base form is exactly the
|
||||||
|
kind of change that would have let them drift into two answers about one set (kø-(p)). Both
|
||||||
|
now call this.
|
||||||
|
|
||||||
|
:raises ValueError: a malformed line, or a block that declares no ``bundle_id``.
|
||||||
|
"""
|
||||||
|
blocks: list[dict[str, str]] = []
|
||||||
|
for line in Path(path).read_text(encoding="utf-8").splitlines():
|
||||||
|
line = line.strip()
|
||||||
|
if not line or line.startswith("#"):
|
||||||
|
continue
|
||||||
|
if ": " not in line:
|
||||||
|
raise ValueError(f"malformed bundle.txt line in {path}: {line!r}")
|
||||||
|
key, value = (part.strip() for part in line.split(": ", 1))
|
||||||
|
if key == "name" or not blocks:
|
||||||
|
blocks.append({})
|
||||||
|
blocks[-1][key] = value
|
||||||
|
if not blocks:
|
||||||
|
raise ValueError(f"{path} declares no knowledge base at all")
|
||||||
|
for block in blocks:
|
||||||
|
for required in ("name", "bundle_id"):
|
||||||
|
if required not in block:
|
||||||
|
raise ValueError(f"{path} declares no {required!r}")
|
||||||
|
return tuple(blocks)
|
||||||
|
|
||||||
|
|
||||||
def _read_json(path: Path) -> dict[str, Any]:
|
def _read_json(path: Path) -> dict[str, Any]:
|
||||||
return json.loads(path.read_text(encoding="utf-8")) # type: ignore[no-any-return]
|
return json.loads(path.read_text(encoding="utf-8")) # type: ignore[no-any-return]
|
||||||
|
|
||||||
|
|
@ -233,9 +268,20 @@ def score_context_set(
|
||||||
outbox_dir: str | Path,
|
outbox_dir: str | Path,
|
||||||
run_id: str,
|
run_id: str,
|
||||||
bundle_dir: str | Path,
|
bundle_dir: str | Path,
|
||||||
|
bundle_id: str | None = None,
|
||||||
) -> ContextSetVerdict:
|
) -> ContextSetVerdict:
|
||||||
"""Judge ONE context set against ONE outbox. ``bundle_dir`` is the MOUNTED base itself (the
|
"""Judge ONE context set against ONE outbox. ``bundle_dir`` is the MOUNTED base itself (the
|
||||||
CLI resolves it from ``--bundle-root`` plus the set's own ``bundle.txt`` name)."""
|
CLI resolves it from ``--bundle-root`` plus the set's own ``bundle.txt`` name).
|
||||||
|
|
||||||
|
``bundle_id`` RESTRICTS the judgement to the approaches a multi-base set routed at THIS base
|
||||||
|
(P17b DEL 2). Without it, judging a two-base set's n200 outbox would report the r761 approach
|
||||||
|
as ``not_evaluated`` with reason ``absent`` — a false finding, because that approach WAS
|
||||||
|
evaluated, against the other base, under the other ``run_id``. ``None`` keeps every single-base
|
||||||
|
set judged exactly as before, which is why this is a restriction rather than a new mode: the
|
||||||
|
order offered a ``--multibase`` summary reader, and MEASURED against the shape the artefacts
|
||||||
|
actually take, the per-base run already has its own full artefact set and its own run_id — so
|
||||||
|
what the judge was missing was not a new file to read but the one thing the mandate already
|
||||||
|
knows, namely which approaches belong here."""
|
||||||
context = Path(context_dir)
|
context = Path(context_dir)
|
||||||
outbox = Path(outbox_dir)
|
outbox = Path(outbox_dir)
|
||||||
base = Path(bundle_dir)
|
base = Path(bundle_dir)
|
||||||
|
|
@ -256,6 +302,19 @@ def score_context_set(
|
||||||
must_cite = {row["approach_id"]: row.get("concepts", []) for row in fasit.get("must_cite", [])}
|
must_cite = {row["approach_id"]: row.get("concepts", []) for row in fasit.get("must_cite", [])}
|
||||||
refuse_ids = {row["approach_id"] for row in fasit.get("must_refuse", [])}
|
refuse_ids = {row["approach_id"] for row in fasit.get("must_refuse", [])}
|
||||||
|
|
||||||
|
# P17b: the approaches THIS base was asked about. Read off the mandate, never off a second
|
||||||
|
# per-row key in the fasit — the routing already has exactly one home (kø-(p)).
|
||||||
|
judged_approaches = mandate.approaches
|
||||||
|
if bundle_id is not None:
|
||||||
|
judged_approaches = tuple(
|
||||||
|
a for a in judged_approaches if (a.bundle_id or bundle_id) == bundle_id
|
||||||
|
)
|
||||||
|
if not judged_approaches:
|
||||||
|
raise EmptyMeasurement(
|
||||||
|
f"no approach in {context} is routed at {bundle_id!r} - a judgement over zero "
|
||||||
|
"rows has no denominator"
|
||||||
|
)
|
||||||
|
|
||||||
# ---- run-level trace ---------------------------------------------------------------------
|
# ---- run-level trace ---------------------------------------------------------------------
|
||||||
debate = outbox / f"{run_id}-debate.json"
|
debate = outbox / f"{run_id}-debate.json"
|
||||||
tool_calls: list[dict[str, Any]] = (
|
tool_calls: list[dict[str, Any]] = (
|
||||||
|
|
@ -303,7 +362,7 @@ def score_context_set(
|
||||||
validated_codes: set[str] = set()
|
validated_codes: set[str] = set()
|
||||||
validated_ids: set[str] = set()
|
validated_ids: set[str] = set()
|
||||||
|
|
||||||
for approach in mandate.approaches:
|
for approach in judged_approaches:
|
||||||
proposal_path, outcome_path = _artefacts(outbox, run_id, approach.id)
|
proposal_path, outcome_path = _artefacts(outbox, run_id, approach.id)
|
||||||
concepts = must_cite.get(approach.id, [])
|
concepts = must_cite.get(approach.id, [])
|
||||||
wanted = {c["path"] for c in concepts}
|
wanted = {c["path"] for c in concepts}
|
||||||
|
|
@ -422,9 +481,14 @@ def score_context_set(
|
||||||
|
|
||||||
# ---- the falsification arm ---------------------------------------------------------------
|
# ---- the falsification arm ---------------------------------------------------------------
|
||||||
refusals: list[RefusalVerdict] = []
|
refusals: list[RefusalVerdict] = []
|
||||||
|
judged_ids = {a.id for a in judged_approaches}
|
||||||
for row in fasit.get("must_refuse", []):
|
for row in fasit.get("must_refuse", []):
|
||||||
rid = row["approach_id"]
|
rid = row["approach_id"]
|
||||||
commissioned = next((a for a in mandate.approaches if a.id == rid), None)
|
# P17b: a falsification arm routed at ANOTHER base was neither asked nor answered here,
|
||||||
|
# and reporting it would put a pass/fail on a run that did not happen in this outbox.
|
||||||
|
if rid not in judged_ids:
|
||||||
|
continue
|
||||||
|
commissioned = next((a for a in judged_approaches if a.id == rid), None)
|
||||||
refuse_codes = set(commissioned.affected_codes) if commissioned is not None else set()
|
refuse_codes = set(commissioned.affected_codes) if commissioned is not None else set()
|
||||||
leaked = sorted(refuse_codes & validated_codes)
|
leaked = sorted(refuse_codes & validated_codes)
|
||||||
if rid in validated_ids:
|
if rid in validated_ids:
|
||||||
|
|
@ -446,7 +510,7 @@ def score_context_set(
|
||||||
return ContextSetVerdict(
|
return ContextSetVerdict(
|
||||||
context_set=context.name,
|
context_set=context.name,
|
||||||
run_id=run_id,
|
run_id=run_id,
|
||||||
bundle_id=str(fasit.get("bundle_id", "")),
|
bundle_id=bundle_id if bundle_id is not None else str(fasit.get("bundle_id", "")),
|
||||||
approaches=tuple(rows),
|
approaches=tuple(rows),
|
||||||
must_refuse=tuple(refusals),
|
must_refuse=tuple(refusals),
|
||||||
hallucinated_reads=tuple(hallucinated_reads),
|
hallucinated_reads=tuple(hallucinated_reads),
|
||||||
|
|
@ -480,18 +544,47 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
default=os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", _DEFAULT_BUNDLE_ROOT),
|
default=os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", _DEFAULT_BUNDLE_ROOT),
|
||||||
help="directory the set's bundle.txt name is mounted under",
|
help="directory the set's bundle.txt name is mounted under",
|
||||||
)
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--bundle",
|
||||||
|
default=None,
|
||||||
|
help="which of the set's declared bases this outbox is for (name or bundle_id). Required "
|
||||||
|
"when the set declares more than one (P17b): each approach is judged against ITS OWN "
|
||||||
|
"base, so a judge that guessed would score one base's run against another's fasit rows",
|
||||||
|
)
|
||||||
args = parser.parse_args(argv)
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
context = Path(args.context_dir)
|
context = Path(args.context_dir)
|
||||||
declared = dict(
|
declared = read_bundle_declarations(context / "bundle.txt")
|
||||||
line.split(":", 1) # type: ignore[misc]
|
if args.bundle is None and len(declared) > 1:
|
||||||
for line in (context / "bundle.txt").read_text(encoding="utf-8").splitlines()
|
print(
|
||||||
if ":" in line
|
f"stress refused: {context} declares {len(declared)} knowledge bases "
|
||||||
)
|
f"({', '.join(d['name'] for d in declared)}); name the one this outbox is for with "
|
||||||
base = Path(args.bundle_root).expanduser() / declared["name"].strip()
|
"--bundle, because a judge that picked would be scoring one base's run against "
|
||||||
|
"another base's fasit rows",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1
|
||||||
|
chosen = declared[0] if args.bundle is None else None
|
||||||
|
for block in declared:
|
||||||
|
if args.bundle in (block["name"], block["bundle_id"]):
|
||||||
|
chosen = block
|
||||||
|
if chosen is None:
|
||||||
|
print(
|
||||||
|
f"stress refused: {args.bundle!r} is not one of the knowledge bases {context} "
|
||||||
|
f"declares ({', '.join(d['name'] for d in declared)})",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1
|
||||||
|
base = Path(args.bundle_root).expanduser() / chosen["name"]
|
||||||
|
|
||||||
try:
|
try:
|
||||||
verdict = score_context_set(context, args.outbox_dir, args.run_id, base)
|
verdict = score_context_set(
|
||||||
|
context,
|
||||||
|
args.outbox_dir,
|
||||||
|
args.run_id,
|
||||||
|
base,
|
||||||
|
bundle_id=chosen["bundle_id"] if len(declared) > 1 else None,
|
||||||
|
)
|
||||||
except EmptyMeasurement as exc:
|
except EmptyMeasurement as exc:
|
||||||
print(f"stress refused: {exc}", file=sys.stderr)
|
print(f"stress refused: {exc}", file=sys.stderr)
|
||||||
return 1
|
return 1
|
||||||
|
|
|
||||||
|
|
@ -47,6 +47,7 @@ from pydantic import ValidationError
|
||||||
|
|
||||||
from portfolio_optimiser import okf
|
from portfolio_optimiser import okf
|
||||||
from portfolio_optimiser.mandate import load_mandate
|
from portfolio_optimiser.mandate import load_mandate
|
||||||
|
from portfolio_optimiser.stress import read_bundle_declarations
|
||||||
|
|
||||||
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||||
_CONTEXT_ROOT = _REPO_ROOT / "contexts"
|
_CONTEXT_ROOT = _REPO_ROOT / "contexts"
|
||||||
|
|
@ -109,21 +110,11 @@ def _bundle_root() -> Path:
|
||||||
return Path(os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", str(_DEFAULT_BUNDLE_ROOT)))
|
return Path(os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", str(_DEFAULT_BUNDLE_ROOT)))
|
||||||
|
|
||||||
|
|
||||||
def read_bundle_txt(path: Path) -> dict[str, str]:
|
#: The ONE reader, imported from production rather than copied here (P17b). It used to be a
|
||||||
"""Parse a set's ``bundle.txt``: ``key: value`` lines, nothing else."""
|
#: private copy in this file and a second, looser one inside ``stress.main`` — and the multi-base
|
||||||
out: dict[str, str] = {}
|
#: form is exactly the change that would have let the two drift into different answers about one
|
||||||
for line in path.read_text(encoding="utf-8").splitlines():
|
#: set. A set declaring ONE base is one block, so the four pre-P17b files parse unchanged.
|
||||||
line = line.strip()
|
read_bundle_txt = read_bundle_declarations
|
||||||
if not line or line.startswith("#"):
|
|
||||||
continue
|
|
||||||
if ": " not in line:
|
|
||||||
raise ValueError(f"malformed bundle.txt line in {path}: {line!r}")
|
|
||||||
key, value = line.split(": ", 1)
|
|
||||||
out[key.strip()] = value.strip()
|
|
||||||
for required in ("name", "bundle_id"):
|
|
||||||
if required not in out:
|
|
||||||
raise ValueError(f"{path} declares no {required!r}")
|
|
||||||
return out
|
|
||||||
|
|
||||||
|
|
||||||
def scan_concepts(base: Path) -> list[tuple[str, dict[str, str], str]]:
|
def scan_concepts(base: Path) -> list[tuple[str, dict[str, str], str]]:
|
||||||
|
|
@ -184,14 +175,29 @@ def _require_base(declared: dict[str, str]) -> Path:
|
||||||
return base
|
return base
|
||||||
|
|
||||||
|
|
||||||
|
def _base_by_approach(set_dir: Path) -> dict[str, Path]:
|
||||||
|
"""Which MOUNTED base each approach was routed at (P17b).
|
||||||
|
|
||||||
|
Read off the mandate's ``bundle_id`` and the set's own declarations — the routing has exactly
|
||||||
|
one home, and a second per-row key in the fasit would be the copy free to drift.
|
||||||
|
"""
|
||||||
|
by_id = {block["bundle_id"]: block for block in read_bundle_txt(set_dir / "bundle.txt")}
|
||||||
|
out: dict[str, Path] = {}
|
||||||
|
for approach in load_mandate(set_dir / "mandate.json").approaches:
|
||||||
|
block = by_id.get(approach.bundle_id)
|
||||||
|
assert block is not None, f"{approach.id} routes at an undeclared base"
|
||||||
|
out[approach.id] = _require_base(block)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
# --------------------------------------------------------------------------------------------
|
# --------------------------------------------------------------------------------------------
|
||||||
# The sets exist at all. Without this, every parametrised arm below would collapse to zero cases
|
# The sets exist at all. Without this, every parametrised arm below would collapse to zero cases
|
||||||
# and the file would pass by having nothing to say.
|
# and the file would pass by having nothing to say.
|
||||||
# --------------------------------------------------------------------------------------------
|
# --------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
def test_the_four_context_sets_are_present() -> None:
|
def test_the_five_context_sets_are_present() -> None:
|
||||||
assert len(_SETS) == 4, f"expected four context sets under {_CONTEXT_ROOT}, found {_SET_IDS}"
|
assert len(_SETS) == 5, f"expected five context sets under {_CONTEXT_ROOT}, found {_SET_IDS}"
|
||||||
|
|
||||||
|
|
||||||
# --------------------------------------------------------------------------------------------
|
# --------------------------------------------------------------------------------------------
|
||||||
|
|
@ -212,13 +218,26 @@ def test_a_mandate_loads_fail_fast(set_dir: Path) -> None:
|
||||||
|
|
||||||
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
def test_d_every_approach_is_routed_at_this_sets_own_base(set_dir: Path) -> None:
|
def test_d_every_approach_is_routed_at_this_sets_own_base(set_dir: Path) -> None:
|
||||||
|
"""Every approach names ONE of the set's declared bases, and every declared base is named.
|
||||||
|
|
||||||
|
Both halves are the claim. The first is the original: an approach routed at a base the set is
|
||||||
|
not for would be evaluated against a corpus nobody commissioned. The second arrived with the
|
||||||
|
multi-base form (P17b) and is what keeps the declaration honest the other way — a base listed
|
||||||
|
in ``bundle.txt`` that no approach names is never run (``route_by_bundle``'s own rule, a run
|
||||||
|
costs money and the commission ordered nothing for it), so a set declaring it would be
|
||||||
|
describing a pass wider than the one it commissions.
|
||||||
|
"""
|
||||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
declared = read_bundle_txt(set_dir / "bundle.txt")
|
||||||
mandate = load_mandate(set_dir / "mandate.json")
|
ids = {block["bundle_id"] for block in declared}
|
||||||
for approach in mandate.approaches:
|
routed = {approach.bundle_id for approach in load_mandate(set_dir / "mandate.json").approaches}
|
||||||
assert approach.bundle_id == declared["bundle_id"], (
|
assert routed <= ids, (
|
||||||
f"{set_dir.name}/{approach.id} routes at {approach.bundle_id!r} but the set declares "
|
f"{set_dir.name}: approaches route at {sorted(routed - ids)}, which the set does not "
|
||||||
f"{declared['bundle_id']!r}"
|
f"declare (declared: {sorted(ids)})"
|
||||||
)
|
)
|
||||||
|
assert ids <= routed, (
|
||||||
|
f"{set_dir.name}: declares {sorted(ids - routed)} that no approach names, so the set "
|
||||||
|
"describes a wider pass than it commissions"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
|
|
@ -244,16 +263,23 @@ def test_the_fasit_names_every_commissioned_approach(set_dir: Path) -> None:
|
||||||
|
|
||||||
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
def test_b_every_fasit_concept_is_in_the_base_as_recorded(set_dir: Path) -> None:
|
def test_b_every_fasit_concept_is_in_the_base_as_recorded(set_dir: Path) -> None:
|
||||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
"""Every cited concept is in the base ITS OWN approach was routed at (P17b).
|
||||||
base = _require_base(declared)
|
|
||||||
|
Resolving per approach rather than per set is the multi-base half: in a set spanning two
|
||||||
|
bases, checking every path against one of them would fail half the fasit while proving
|
||||||
|
nothing about the other, and checking against "either" would let a path meant for N200 be
|
||||||
|
satisfied by a coincidence in R761.
|
||||||
|
"""
|
||||||
|
bases = _base_by_approach(set_dir)
|
||||||
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
||||||
|
|
||||||
seen = 0
|
seen = 0
|
||||||
for row in fasit["must_cite"]:
|
for row in fasit["must_cite"]:
|
||||||
assert row["concepts"], f"{row['approach_id']} cites nothing a right answer must reach"
|
assert row["concepts"], f"{row['approach_id']} cites nothing a right answer must reach"
|
||||||
|
base = bases[row["approach_id"]]
|
||||||
for concept in row["concepts"]:
|
for concept in row["concepts"]:
|
||||||
path = base / concept["path"]
|
path = base / concept["path"]
|
||||||
assert path.is_file(), f"{set_dir.name}: {concept['path']} is not in {declared['name']}"
|
assert path.is_file(), f"{set_dir.name}: {concept['path']} is not in {base.name}"
|
||||||
frontmatter = own_frontmatter(path)
|
frontmatter = own_frontmatter(path)
|
||||||
assert frontmatter.get("title", "") == concept["title"], (
|
assert frontmatter.get("title", "") == concept["title"], (
|
||||||
f"{concept['path']}: the base's own title has drifted from the fasit"
|
f"{concept['path']}: the base's own title has drifted from the fasit"
|
||||||
|
|
@ -269,31 +295,43 @@ def test_b_every_fasit_concept_is_in_the_base_as_recorded(set_dir: Path) -> None
|
||||||
|
|
||||||
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
def test_c_rule_u_every_unanswerable_question_is_unanswerable(set_dir: Path) -> None:
|
def test_c_rule_u_every_unanswerable_question_is_unanswerable(set_dir: Path) -> None:
|
||||||
|
"""Rule U over EVERY base the set declares, as ONE scan.
|
||||||
|
|
||||||
|
For a multi-base set "the base cannot answer this" becomes "NEITHER base can", and the union
|
||||||
|
is the honest reading: an anchor absent from N200 but present in R761 is a question the pass
|
||||||
|
as a whole CAN reach. MEASURED 15.09 and the reason this is not a formality — ``enhetspris``
|
||||||
|
is absent from n200-2024 and carried by 70 of r761-2025's 2 756 concepts, so an anchor set
|
||||||
|
admitted per base would have admitted a question the pass could ground.
|
||||||
|
"""
|
||||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
declared = read_bundle_txt(set_dir / "bundle.txt")
|
||||||
base = _require_base(declared)
|
concepts: list[tuple[str, dict[str, str], str]] = []
|
||||||
concepts = scan_concepts(base)
|
names = []
|
||||||
|
for block in declared:
|
||||||
|
base = _require_base(block)
|
||||||
|
names.append(block["name"])
|
||||||
|
concepts += scan_concepts(base)
|
||||||
assert len(concepts) >= 100, (
|
assert len(concepts) >= 100, (
|
||||||
f"{declared['name']} scanned to {len(concepts)} concepts — too few to be the base itself"
|
f"{', '.join(names)} scanned to {len(concepts)} concepts — too few to be the base(s)"
|
||||||
)
|
)
|
||||||
|
|
||||||
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
||||||
for row in fasit["must_refuse"]:
|
for row in fasit["must_refuse"]:
|
||||||
carried = anchors_are_absent(row["anchors"], concepts)
|
carried = anchors_are_absent(row["anchors"], concepts)
|
||||||
assert not carried, (
|
assert not carried, (
|
||||||
f"{set_dir.name}: {declared['name']} DOES carry {carried} over {len(concepts)} "
|
f"{set_dir.name}: {', '.join(names)} DOES carry {carried} over {len(concepts)} "
|
||||||
f"concepts, so {row['approach_id']!r} is not un-groundable by rule U"
|
f"concepts, so {row['approach_id']!r} is not un-groundable by rule U"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
def test_e_the_declared_bundle_id_is_the_bases_own(set_dir: Path) -> None:
|
def test_e_the_declared_bundle_id_is_the_bases_own(set_dir: Path) -> None:
|
||||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
for block in read_bundle_txt(set_dir / "bundle.txt"):
|
||||||
base = _require_base(declared)
|
base = _require_base(block)
|
||||||
resolved = okf.reconcile_bundle_id(base)
|
resolved = okf.reconcile_bundle_id(base)
|
||||||
assert resolved.id == declared["bundle_id"], (
|
assert resolved.id == block["bundle_id"], (
|
||||||
f"{set_dir.name}: bundle.txt declares {declared['bundle_id']!r} but the base resolves to "
|
f"{set_dir.name}: bundle.txt declares {block['bundle_id']!r} for {block['name']} but "
|
||||||
f"{resolved.id!r} (origin {resolved.origin})"
|
f"the base resolves to {resolved.id!r} (origin {resolved.origin})"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
# --------------------------------------------------------------------------------------------
|
# --------------------------------------------------------------------------------------------
|
||||||
|
|
@ -350,8 +388,12 @@ def test_known_positive_d_a_mandate_routed_at_another_base_is_caught(tmp_path: P
|
||||||
broken["approaches"][1]["bundle_id"] = "vegnormal-n100-2023"
|
broken["approaches"][1]["bundle_id"] = "vegnormal-n100-2023"
|
||||||
set_dir = _broken_set(tmp_path, mandate=broken, bundle=_GOOD_BUNDLE_TXT, fasit={})
|
set_dir = _broken_set(tmp_path, mandate=broken, bundle=_GOOD_BUNDLE_TXT, fasit={})
|
||||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
declared = read_bundle_txt(set_dir / "bundle.txt")
|
||||||
mandate = load_mandate(set_dir / "mandate.json")
|
ids = {block["bundle_id"] for block in declared}
|
||||||
assert any(a.bundle_id != declared["bundle_id"] for a in mandate.approaches)
|
routed = {a.bundle_id for a in load_mandate(set_dir / "mandate.json").approaches}
|
||||||
|
# The SAME two set relations arm (d) asserts, and the broken set must fail the first of them:
|
||||||
|
# an approach routed at a base the set does not declare.
|
||||||
|
assert not routed <= ids
|
||||||
|
assert sorted(routed - ids) == ["vegnormal-n100-2023"]
|
||||||
|
|
||||||
|
|
||||||
def test_known_positive_c_an_anchor_the_base_carries_is_reported() -> None:
|
def test_known_positive_c_an_anchor_the_base_carries_is_reported() -> None:
|
||||||
|
|
@ -394,6 +436,34 @@ def test_known_positive_bundle_txt_must_declare_both_keys(tmp_path: Path) -> Non
|
||||||
read_bundle_txt(path)
|
read_bundle_txt(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_known_positive_a_second_block_needs_its_own_bundle_id(tmp_path: Path) -> None:
|
||||||
|
"""P17b: each ``name:`` OPENS a block, and each block closes with its own id.
|
||||||
|
|
||||||
|
The half a single-base file cannot exercise: a reader that flattened the file into one
|
||||||
|
mapping would let the FIRST block's ``bundle_id`` satisfy the second, and the second base
|
||||||
|
would then be addressed under the first one's name.
|
||||||
|
"""
|
||||||
|
path = tmp_path / "bundle.txt"
|
||||||
|
path.write_text(
|
||||||
|
"name: n200-2024\nbundle_id: vegnormal-n200-2024\nname: r761-2025\n", encoding="utf-8"
|
||||||
|
)
|
||||||
|
with pytest.raises(ValueError, match="bundle_id"):
|
||||||
|
read_bundle_txt(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_multi_base_bundle_txt_parses_into_one_block_per_base(tmp_path: Path) -> None:
|
||||||
|
path = tmp_path / "bundle.txt"
|
||||||
|
path.write_text(
|
||||||
|
"name: n200-2024\nbundle_id: vegnormal-n200-2024\n"
|
||||||
|
"name: r761-2025\nbundle_id: vegnormal-r761-2025\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
assert read_bundle_txt(path) == (
|
||||||
|
{"name": "n200-2024", "bundle_id": "vegnormal-n200-2024"},
|
||||||
|
{"name": "r761-2025", "bundle_id": "vegnormal-r761-2025"},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
@pytest.mark.parametrize("set_dir", _SETS, ids=_SET_IDS)
|
||||||
def test_the_fasit_titles_are_distinct_not_the_collapsed_sources_title(set_dir: Path) -> None:
|
def test_the_fasit_titles_are_distinct_not_the_collapsed_sources_title(set_dir: Path) -> None:
|
||||||
"""The fasit's recorded titles must tell the cited concepts APART.
|
"""The fasit's recorded titles must tell the cited concepts APART.
|
||||||
|
|
@ -409,14 +479,15 @@ def test_the_fasit_titles_are_distinct_not_the_collapsed_sources_title(set_dir:
|
||||||
always wins over a nested one of the same name), so the second half now asserts the opposite —
|
always wins over a nested one of the same name), so the second half now asserts the opposite —
|
||||||
that ``parse_frontmatter`` agrees with the fasit's own distinct titles — as a live regression
|
that ``parse_frontmatter`` agrees with the fasit's own distinct titles — as a live regression
|
||||||
guard against the collapse coming back."""
|
guard against the collapse coming back."""
|
||||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
bases = _base_by_approach(set_dir)
|
||||||
base = _require_base(declared)
|
|
||||||
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
||||||
|
|
||||||
cited = [c for row in fasit["must_cite"] for c in row["concepts"]]
|
cited = [(bases[row["approach_id"]], c) for row in fasit["must_cite"] for c in row["concepts"]]
|
||||||
assert len({c["title"] for c in cited}) == len(cited), "recorded titles do not tell them apart"
|
assert len({c["title"] for _, c in cited}) == len(cited), (
|
||||||
|
"recorded titles do not tell them apart"
|
||||||
|
)
|
||||||
|
|
||||||
titles = {okf.parse_frontmatter(base / c["path"]).get("title", "") for c in cited}
|
titles = {okf.parse_frontmatter(base / c["path"]).get("title", "") for base, c in cited}
|
||||||
assert len(titles) == len(cited), (
|
assert len(titles) == len(cited), (
|
||||||
"okf.parse_frontmatter collapsed these titles onto the sources block again — the P15 fix "
|
"okf.parse_frontmatter collapsed these titles onto the sources block again — the P15 fix "
|
||||||
"in okf._frontmatter_from_text has regressed"
|
"in okf._frontmatter_from_text has regressed"
|
||||||
|
|
|
||||||
|
|
@ -140,10 +140,17 @@ def test_a_code_the_baseline_carries_is_never_refused_for_its_shape() -> None:
|
||||||
|
|
||||||
|
|
||||||
def test_every_fasit_reference_in_every_context_set_is_an_identifier() -> None:
|
def test_every_fasit_reference_in_every_context_set_is_an_identifier() -> None:
|
||||||
"""(d) The 26 the order names, with the denominator, plus this repo's own cost code.
|
"""(d) Every fasit reference across every context set, with the denominator.
|
||||||
|
|
||||||
One of them — ``Krav 3.3.2—1_1`` — is why the second form grew an optional ``_<n>`` suffix.
|
One of them — ``Krav 3.3.2—1_1`` — is why the second form grew an optional ``_<n>`` suffix.
|
||||||
Measured, not anticipated: before that it was the single reference the classifier called prose.
|
Measured, not anticipated: before that it was the single reference the classifier called prose.
|
||||||
|
|
||||||
|
**The denominator MOVED 26 -> 32 with P17b's fifth context set**, and it is asserted rather
|
||||||
|
than dropped for the reason it was written down in the first place: a list comprehension over
|
||||||
|
``contexts/*/fasit.json`` that quietly found fewer rows would make this arm weaker without
|
||||||
|
making it red. The six new ones are four ``Krav x.y.z—n`` from n200-2024 and TWO bare
|
||||||
|
``prosessnr`` from r761-2025 (``12.11``, ``12.12``) — the punctuation-and-digits form B1 added,
|
||||||
|
now exercised by a fasit and not only by a known-positive.
|
||||||
"""
|
"""
|
||||||
refs = [
|
refs = [
|
||||||
concept["ref"]
|
concept["ref"]
|
||||||
|
|
@ -151,7 +158,7 @@ def test_every_fasit_reference_in_every_context_set_is_an_identifier() -> None:
|
||||||
for row in json.loads(open(path, encoding="utf-8").read())["must_cite"]
|
for row in json.loads(open(path, encoding="utf-8").read())["must_cite"]
|
||||||
for concept in row["concepts"]
|
for concept in row["concepts"]
|
||||||
]
|
]
|
||||||
assert len(refs) == 26, f"denominator moved: {len(refs)}"
|
assert len(refs) == 32, f"denominator moved: {len(refs)}"
|
||||||
assert [r for r in refs if not has_identifier_form(r)] == []
|
assert [r for r in refs if not has_identifier_form(r)] == []
|
||||||
assert has_identifier_form("ENERGI-TOTAL-EL"), "this repo's own reference cost code"
|
assert has_identifier_form("ENERGI-TOTAL-EL"), "this repo's own reference cost code"
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -486,3 +486,146 @@ def test_k_the_cli_writes_the_verdict_file_and_prints_it(tmp_path: Path) -> None
|
||||||
payload = json.loads(written.read_text(encoding="utf-8"))
|
payload = json.loads(written.read_text(encoding="utf-8"))
|
||||||
assert payload["ferdig"] is True
|
assert payload["ferdig"] is True
|
||||||
assert json.loads(proc.stdout)["ferdig"] is True
|
assert json.loads(proc.stdout)["ferdig"] is True
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
# P17b DEL 2 — a context set spanning SEVERAL bases is judged ONE base at a time, and the judge
|
||||||
|
# is told which. Without that restriction the other base's approach is reported
|
||||||
|
# ``not_evaluated``/``absent``, which is a FALSE finding: that approach WAS evaluated, against the
|
||||||
|
# other base, under the other ``run_id``.
|
||||||
|
# --------------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _two_base_context(root: Path) -> Path:
|
||||||
|
ctx = root / "ctx2"
|
||||||
|
(ctx / "docs").mkdir(parents=True)
|
||||||
|
(ctx / "bundle.txt").write_text(
|
||||||
|
"name: minibase\nbundle_id: minibase\nname: otherbase\nbundle_id: otherbase\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
(ctx / "mandate.json").write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"objective": "o",
|
||||||
|
"success_criteria": "s",
|
||||||
|
"approaches": [
|
||||||
|
{
|
||||||
|
"id": "a1",
|
||||||
|
"label": "Here",
|
||||||
|
"affected_codes": ["CODE-1"],
|
||||||
|
"claimed_saving_nok": 1000.0,
|
||||||
|
"bundle_id": "minibase",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "a2",
|
||||||
|
"label": "Over there",
|
||||||
|
"affected_codes": ["CODE-2"],
|
||||||
|
"claimed_saving_nok": 2000.0,
|
||||||
|
"bundle_id": "otherbase",
|
||||||
|
},
|
||||||
|
],
|
||||||
|
},
|
||||||
|
indent=2,
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
(ctx / "fasit.json").write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"project_id": "proj",
|
||||||
|
"must_cite": [
|
||||||
|
{
|
||||||
|
"approach_id": "a1",
|
||||||
|
"rationale": "why",
|
||||||
|
"concepts": [{"path": _GOOD, "title": _TITLE, "ref": _REF}],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"approach_id": "a2",
|
||||||
|
"rationale": "why",
|
||||||
|
"concepts": [{"path": _OTHER, "title": "Other", "ref": "Krav 9.9.9-9"}],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
"must_refuse": [],
|
||||||
|
"honesty": "synthetic",
|
||||||
|
},
|
||||||
|
indent=2,
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
return ctx
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_multi_base_set_is_judged_one_base_at_a_time(tmp_path: Path) -> None:
|
||||||
|
"""Told which base this outbox is for, the judge answers for THAT base's approaches only."""
|
||||||
|
base = _minibase(tmp_path)
|
||||||
|
ctx = _two_base_context(tmp_path)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
_write_outbox(outbox, "r1-minibase", approach_id="a1", tool_calls=_opened(_GOOD))
|
||||||
|
|
||||||
|
verdict = stress.score_context_set(ctx, outbox, "r1-minibase", base, bundle_id="minibase")
|
||||||
|
|
||||||
|
assert [row.approach_id for row in verdict.approaches] == ["a1"]
|
||||||
|
assert verdict.bundle_id == "minibase"
|
||||||
|
assert verdict.approaches[0].status == "validated"
|
||||||
|
|
||||||
|
|
||||||
|
def test_without_the_restriction_the_other_bases_approach_is_falsely_reported_absent(
|
||||||
|
tmp_path: Path,
|
||||||
|
) -> None:
|
||||||
|
"""The defect the restriction removes, stated as a measurement rather than a worry.
|
||||||
|
|
||||||
|
This is the UNRESTRICTED call on the same outbox: ``a2`` has no artefact here — it was run
|
||||||
|
against the other base, under the other ``run_id`` — and the judge reports it as an approach
|
||||||
|
nobody evaluated, which is exactly the silence ``not_evaluated`` exists to remove.
|
||||||
|
"""
|
||||||
|
base = _minibase(tmp_path)
|
||||||
|
ctx = _two_base_context(tmp_path)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
_write_outbox(outbox, "r1-minibase", approach_id="a1", tool_calls=_opened(_GOOD))
|
||||||
|
|
||||||
|
verdict = stress.score_context_set(ctx, outbox, "r1-minibase", base)
|
||||||
|
|
||||||
|
rows = {row.approach_id: row for row in verdict.approaches}
|
||||||
|
assert set(rows) == {"a1", "a2"}
|
||||||
|
assert rows["a2"].status == "not_evaluated"
|
||||||
|
assert rows["a2"].not_evaluated_reason == "absent"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_base_no_approach_is_routed_at_has_no_denominator(tmp_path: Path) -> None:
|
||||||
|
base = _minibase(tmp_path)
|
||||||
|
ctx = _two_base_context(tmp_path)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
_write_outbox(outbox, "r1-minibase", approach_id="a1")
|
||||||
|
|
||||||
|
with pytest.raises(stress.EmptyMeasurement, match="routed at"):
|
||||||
|
stress.score_context_set(ctx, outbox, "r1-minibase", base, bundle_id="thirdbase")
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_cli_refuses_to_guess_which_base_a_multi_base_outbox_is_for(
|
||||||
|
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||||
|
) -> None:
|
||||||
|
"""Refused, never guessed: picking would score one base's run against another's fasit rows.
|
||||||
|
|
||||||
|
Paired with the rc-0 control below, so "rc 1" cannot be coming from the rest of the argv.
|
||||||
|
"""
|
||||||
|
_minibase(tmp_path)
|
||||||
|
ctx = _two_base_context(tmp_path)
|
||||||
|
outbox = tmp_path / "out"
|
||||||
|
_write_outbox(outbox, "r1-minibase", approach_id="a1", tool_calls=_opened(_GOOD))
|
||||||
|
|
||||||
|
argv = [
|
||||||
|
str(ctx),
|
||||||
|
"--outbox-dir",
|
||||||
|
str(outbox),
|
||||||
|
"--run-id",
|
||||||
|
"r1-minibase",
|
||||||
|
"--bundle-root",
|
||||||
|
str(tmp_path),
|
||||||
|
]
|
||||||
|
assert stress.main(argv) == 1
|
||||||
|
assert "--bundle" in capsys.readouterr().err
|
||||||
|
|
||||||
|
assert stress.main([*argv, "--bundle", "minibase"]) == 0, "control: naming the base works"
|
||||||
|
|
||||||
|
assert stress.main([*argv, "--bundle", "nowhere"]) == 1
|
||||||
|
assert "nowhere" in capsys.readouterr().err
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue