refactor(examples): replace sector-specific example material with generic, fictitious examples
The context sets, the packaged knowledge bases and the example bundles are replaced by one fictitious example set about IT operations in an invented organisation: three context sets (serverrom-2027, driftsavtale-2027 and the two-base drift-og-avtale-2027), two synthetic knowledge bases under src/portfolio_optimiser/data/kunnskapsbaser and two example bundles under src/portfolio_optimiser/data/bundles. Numbers, codes and structural values in tests and fixtures are kept; names, ids and wording change. Dated measurement documents that only recorded runs on the replaced material are deleted. Gate figures measured on the new set are not comparable with earlier ones. The exclusion gate from the previous commit is green: 0 tracked files hit outside the shared/ subtree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
parent
058dd25570
commit
37547fe292
1147 changed files with 24138 additions and 9503 deletions
|
|
@ -72,14 +72,14 @@ _PORTFOLIO_DEFAULT_REPLY = (
|
|||
|
||||
def _anchored_default_replies() -> dict[str, str]:
|
||||
"""A VALID default proposal PER reference project, quoting that project's OWN first cost line
|
||||
verbatim (S4.0): since the road path anchors the validator to ``project.cost_items``, a generic
|
||||
verbatim (S4.0): since the reference path anchors the validator to ``project.cost_items``, a generic
|
||||
reply carrying an invented magnitude for code ``01.1`` is now — correctly — rejected as a
|
||||
fabricated cost line. Anchoring the fixture is the fix; weakening the gate is not.
|
||||
|
||||
``claimed_saving_nok`` stays ``20_000`` for every project, exactly as the single generic reply
|
||||
claimed before, so every ledger/goal/budget assertion built on that figure is unchanged. Each
|
||||
project's first line is ``01.1 Rigg og drift`` at >= 480 000 NOK, so P90 (>= 144 000) clears the
|
||||
claim on every project."""
|
||||
project's first line is code ``01.1`` (project management and operations, a lump sum) at
|
||||
>= 480 000 NOK, so P90 (>= 144 000) clears the claim on every project."""
|
||||
replies: dict[str, str] = {}
|
||||
for project in load_reference_projects():
|
||||
line = project.cost_items[0]
|
||||
|
|
@ -190,7 +190,7 @@ def docs_dir(tmp_path) -> str:
|
|||
d = tmp_path / "docs"
|
||||
d.mkdir()
|
||||
(d / "cost.txt").write_text(
|
||||
"Asphalt Ab11 unit rate renegotiation reduced the paving cost on the school stretch.",
|
||||
"Licence unit price renegotiation reduced the office-suite cost for the head office.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return str(d)
|
||||
|
|
|
|||
26
tests/fixtures/p13b-block-sources/driftskrav-krav-4-2-5-1-3.md
vendored
Normal file
26
tests/fixtures/p13b-block-sources/driftskrav-krav-4-2-5-1-3.md
vendored
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
---
|
||||
type: Krav
|
||||
title: Krav 4.2.5.1—3 Skjerm og kamera
|
||||
description: Skjerm og kamera skal kunne styres fra ett panel, og oppsettet skal være likt i alle møterom av samme størrelse.
|
||||
kravtype: skal
|
||||
standard: D100
|
||||
utgave: D100:2027
|
||||
req_number: Krav 4.2.5.1—3
|
||||
status: stable
|
||||
trust_tier: unverified
|
||||
seksjon: 4.2.5.1
|
||||
seksjonstittel: Skjerm og kamera
|
||||
ingested_at: 2027-01-15T12:00:00Z
|
||||
source_element_id: id-7aed3659-dbbe-5eb5-9b96-e6a693ff0c3c
|
||||
sources:
|
||||
- resource: https://example.invalid/eksempelvirksomheten/driftskrav/d100
|
||||
title: D100:2027
|
||||
---
|
||||
|
||||
## Krav
|
||||
|
||||
Skjerm og kamera skal kunne styres fra ett panel, og oppsettet skal være likt i alle møterom av samme størrelse.
|
||||
|
||||
## Veiledning (ikke-normativ)
|
||||
|
||||
Kravene gjelder møterom med fast utstyr. Mobile løsninger omfattes ikke.
|
||||
|
|
@ -1,26 +0,0 @@
|
|||
---
|
||||
type: Krav
|
||||
title: Krav 4.2.5.1—3 Gangfelt og tilrettelagte kryssingspunkter
|
||||
description: Ved fartsgrense 40 og 50 km/t skal gangfelt etableres dersom:●Antall fotgjengere > 20 og antall kjøretøy > 200 i dimensjonerende time ●Antall fotgjengere > 10…
|
||||
kravtype: skal
|
||||
normal: N100
|
||||
utgave: N100:2023
|
||||
req_number: Krav 4.2.5.1—3
|
||||
kravdato: 2021-06-22
|
||||
hjemmel: forskrift om anlegg av offentlig veg, jf. vegloven § 13
|
||||
fraviksmyndighet: ikke uttalt i kilden
|
||||
status: stable
|
||||
trust_tier: unverified
|
||||
seksjon: 4.2.5.1
|
||||
seksjonstittel: Gangfelt og tilrettelagte kryssingspunkter
|
||||
ingested_at: 2026-09-08T12:00:00Z
|
||||
source_sha256: c58e8bbc5fa9a5400c111e51b04c05f2cfd9edabd884ef5352a486fdab2cb5ab
|
||||
source_element_id: id-2b69893f-e462-4a75-d1c3-b93a7dbf1667
|
||||
sources:
|
||||
- resource: https://viewers.vegnorm.vegvesen.no/api/nisosts/859984?languageCode=nb
|
||||
title: N100:2023
|
||||
---
|
||||
|
||||
## Krav
|
||||
|
||||
Ved fartsgrense 40 og 50 km/t skal gangfelt etableres dersom:●Antall fotgjengere > 20 og antall kjøretøy > 200 i dimensjonerende time ●Antall fotgjengere > 10 og antall kjøretøy > 800 i dimensjonerende time
|
||||
|
|
@ -1,7 +1,7 @@
|
|||
Propose ONE concrete cost-saving measure for this project.
|
||||
Project: N100 - N100
|
||||
Project: D100 - D100
|
||||
Context (prior verdicts / cited cost docs):
|
||||
A concrete cost-saving measure based on the N100 requirements is to avoid grade-separated (planskilt) crossings between pedestrian/cycle paths and roads when the road has an ÅDT of 4,000 or less. According to Krav 3.3.1—13, planskilt crossings are required only if ÅDT > 4,000, so using at-grade crossings under this threshold reduces construction costs while still complying with N100.
|
||||
A concrete cost-saving measure based on the D100 operating requirements is to avoid a dedicated (separate) power circuit between the server room and the office floors when the room draws a load of 4,000 W or less. According to Krav 3.3.1—13, a dedicated circuit is required only if the load exceeds 4,000 W, so sharing the existing circuit under this threshold reduces installation costs while still complying with D100.
|
||||
|
||||
Respond with ONLY a JSON object for a SavingsProposal with keys: project_id, measure, affected_items (list of {code, quantity, unit_cost}), claimed_saving_nok, and optional assumptions.
|
||||
Each entry in affected_items must restate a cost line as the project's price schedule already carries it: quantity and unit_cost are the unchanged baseline figures, not the reduced quantity or unit cost your measure would produce. The effect of the measure belongs in claimed_saving_nok.
|
||||
|
|
@ -4,29 +4,29 @@ Beviser dataflyten, den deterministiske ryggraden og at læringssløyfa lukkes.
|
|||
Beviser IKKE at en levende modell ville produsert dette — forslag og dom er skriptet.
|
||||
==============================================================================
|
||||
|
||||
KUNNSKAPSBASE: veglys-fv-soer — kostbaseline erklært (ENERGI-VEGLYS-EL 4386150 x 1)
|
||||
KUNNSKAPSBASE: klientpark-energi — kostbaseline erklært (ENERGI-KLIENTPARK-EL 4386150 x 1)
|
||||
validatorens stage 0 avstemmer forslagets kostlinjer mot disse, FØR løseren
|
||||
tallene er levert i kunnskapsbasen — utledet av fagkilder (Håndbok V124, NMFV), ikke av demo-manuset
|
||||
tallene er levert i kunnskapsbasen — utledet av dens egne (fiktive) kilder, ikke av demo-manuset
|
||||
|
||||
KJØRING A (VEGLYS-FV-SOER — fersk kunnskapsbase, ingen tidligere dommer)
|
||||
KJØRING A (KLIENTPARK-ENERGI — fersk kunnskapsbase, ingen tidligere dommer)
|
||||
Steg 1 — FORSTÅ KONTEKSTEN (navigert kunnskapsbase + tidligere dommer)
|
||||
navigerte konseptfiler (5): kilder-veglys-realisering.md, tiltak-adaptiv-styring.md, tiltak-led-utskifting.md, veglys-fv-soer.md, metode-ipmvp-a.md
|
||||
navigerte konseptfiler (5): kilder-klientpark-realisering.md, tiltak-adaptiv-stromstyring.md, tiltak-pc-utskifting.md, klientpark-energi.md, metode-ipmvp-a.md
|
||||
tidligere dommer hentet for kandidaten: 0
|
||||
markør 'realiseringsgrad=0.79' i hypotese-prompten: False
|
||||
Steg 2 — HYPOTESE (kandidat med parametere)
|
||||
tiltak: LED-utskifting av 2 500 eldre HPS-armaturer (114 W -> 70 W)
|
||||
kostlinjer: ENERGI-VEGLYS-EL 4386150 x 1
|
||||
tiltak: Utskifting av 2 500 eldre stasjonære PC-er (114 W -> 70 W)
|
||||
kostlinjer: ENERGI-KLIENTPARK-EL 4386150 x 1
|
||||
påstått besparelse: 2100000 NOK
|
||||
Steg 3 — DEBATT (maker-checker, Group Chat)
|
||||
proposer (konvergert): {"measure":"LED-utskifting av 2 500 eldre HPS-armaturer (114 W -> 70 W)","affected_items":[…
|
||||
proposer (konvergert): {"measure":"Utskifting av 2 500 eldre stasjonære PC-er (114 W -> 70 W)","affected_items":[{…
|
||||
checker (gate på resonnementet): VERDICT=APPROVE
|
||||
Steg 4 — VALIDER / FALSIFISER (deterministisk, blokkerende)
|
||||
hypotese #1: REJECTED (claimed saving 2100000 exceeds P90 feasible 1769915)
|
||||
Steg 5 — FORBEDRE, INFORMERT OG BUNDET
|
||||
#1: grunnen fra 2100000 NOK-hypotesen mates tilbake i neste forsøk (bundet av max_attempts)
|
||||
etter forbedring: VALIDATED (påstått 445500 <= P90 1769915 NOK; tiltak: LED-utskifting av 2 500 eldre HPS-armaturer (114 W -> 70 W))
|
||||
etter forbedring: VALIDATED (påstått 445500 <= P90 1769915 NOK; tiltak: Utskifting av 2 500 eldre stasjonære PC-er (114 W -> 70 W))
|
||||
Steg 6 — FORKAST ELLER FORESLÅ (typet utfall forlater kjøringen)
|
||||
FORESLÅTT — LED-utskifting av 2 500 eldre HPS-armaturer (114 W -> 70 W): 445500 NOK (validator=validated, checker=approve)
|
||||
FORESLÅTT — Utskifting av 2 500 eldre stasjonære PC-er (114 W -> 70 W): 445500 NOK (validator=validated, checker=approve)
|
||||
Steg 7 — SVAR PÅ TILBAKEMELDING (ekspert-persona, kort løkke i kjøringen)
|
||||
dom: approved
|
||||
begrunnelse: Godkjent med realiseringskorreksjon. Den modellerte besparelsen er teknisk korrekt fra parameterne og validatoren bekrefter at den er innenfor feasibelt omraade. Men i drift realiseres erfaringsvis ~79% av en timeplan-stipulert LED-besparelse i tilsvarende anlegg (realiseringsgrad=0.79) pga. overes…
|
||||
|
|
@ -41,7 +41,7 @@ KJØRING A (VEGLYS-FV-SOER — fersk kunnskapsbase, ingen tidligere dommer)
|
|||
neste kjøring merger fila inn i minnet FØR hypotesen formes — dager kan gå
|
||||
|
||||
Steg 8 — PROMOTER GODKJENT KUNNSKAP (gatet wiki-promotering)
|
||||
skrev: promoted-verdict-b26f6501ccf74eef.md (lenket i index.md, nøytral etikett)
|
||||
skrev: promoted-verdict-78b95c71ead8577a.md (lenket i index.md, nøytral etikett)
|
||||
bærer: realiseringsgrad=0.79 (personaens dom fra kjøring A)
|
||||
gaten er fail-closed: kun en godkjent dom promoteres — rå agent-output aldri
|
||||
|
||||
|
|
@ -51,7 +51,7 @@ KJØRING B (re-seedet kunnskapsbase + innboksen lest)
|
|||
av disse fulgte 1 av 3 med kunnskapsbasen; de øvrige 2 er dem demoen lærte i denne økten
|
||||
markør 'realiseringsgrad=0.79' (Steg 8, wiki) i hypotese-prompten: True (forventet True)
|
||||
markør 'realiseringsgrad=0.66' (Steg 7, innboks) i hypotese-prompten: True (forventet True)
|
||||
utfall: FORESLÅTT — LED-utskifting av 2 500 eldre HPS-armaturer (114 W -> 70 W): 445500 NOK (validator=validated, checker=approve)
|
||||
utfall: FORESLÅTT — Utskifting av 2 500 eldre stasjonære PC-er (114 W -> 70 W): 445500 NOK (validator=validated, checker=approve)
|
||||
|
||||
------------------------------------------------------------------------------
|
||||
LÆRINGSSLØYFA ER LUKKET, PÅ BEGGE TIDSSKALAER: kunnskapen eksperten godkjente i
|
||||
|
|
|
|||
|
|
@ -23,7 +23,7 @@ _ASSUMPTIONS = {"05.2": (200.0, 230.0), "03.1": (290.0, 330.0)}
|
|||
|
||||
@pytest.fixture(scope="module")
|
||||
def project():
|
||||
return load_reference_projects()[0] # FV42-GSV-E1
|
||||
return load_reference_projects()[0] # KONTOR-IT-E1
|
||||
|
||||
|
||||
def _valid(project) -> SavingsProposal:
|
||||
|
|
|
|||
|
|
@ -20,7 +20,7 @@ _QUERY = ProposalFeatures(
|
|||
affected_codes=frozenset({"05.2", "03.1"}),
|
||||
measure_type="scope_reduction",
|
||||
claimed_saving_nok=220_000, # bucket [100k, 500k)
|
||||
description="asphalt base course reduction near school",
|
||||
description="licence scope reduction near head office",
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -44,7 +44,7 @@ def _store_with_true_match_and_decoys() -> tuple[VerdictStore, str]:
|
|||
affected_codes=frozenset({"09.1"}),
|
||||
measure_type="rate_renegotiation",
|
||||
claimed_saving_nok=50_000, # bucket [0, 100k)
|
||||
description="asphalt base course reduction near school", # same words as query
|
||||
description="licence scope reduction near head office", # same words as query
|
||||
),
|
||||
decision="rejected",
|
||||
rationale="surface-text decoy",
|
||||
|
|
@ -55,7 +55,7 @@ def _store_with_true_match_and_decoys() -> tuple[VerdictStore, str]:
|
|||
affected_codes=frozenset({"21.2"}),
|
||||
measure_type="material_substitution",
|
||||
claimed_saving_nok=700_000, # bucket [500k, 1M)
|
||||
description="asphalt base course reduction extra words",
|
||||
description="licence scope reduction extra words",
|
||||
),
|
||||
decision="rejected",
|
||||
rationale="surface-text decoy",
|
||||
|
|
|
|||
|
|
@ -40,7 +40,8 @@ from portfolio_optimiser import explore, okf, run as run_module
|
|||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_BYGG = _EXAMPLES / "bygg-energi-mikro" # BYGG-KONTOR-NORD
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia" # TUNNEL-HAUGLIA
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
_KJOLING = _BUNDLES / "driftssenter-kjoling" # DRIFTSSENTER-KJOLING
|
||||
|
||||
_PROPOSAL = json.dumps(
|
||||
{
|
||||
|
|
@ -136,14 +137,14 @@ def test_each_base_writes_its_own_artefact_set_under_its_own_minted_run_id(tmp_p
|
|||
DISTINCT: with one key the second base would overwrite the first and the outbox would hold one
|
||||
base's answers under a name claiming to cover both.
|
||||
"""
|
||||
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||
bases = _mount(tmp_path, _BYGG, _KJOLING)
|
||||
outbox = tmp_path / "out"
|
||||
|
||||
rc = run_module.main(
|
||||
_argv(
|
||||
bases,
|
||||
mandate=_mandate_file(
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "driftssenter-kjoling")]
|
||||
),
|
||||
run_id="X",
|
||||
outbox=outbox,
|
||||
|
|
@ -153,7 +154,7 @@ def test_each_base_writes_its_own_artefact_set_under_its_own_minted_run_id(tmp_p
|
|||
)
|
||||
|
||||
assert rc == 0
|
||||
for bundle_id, approach in (("bygg-energi-mikro", "a"), ("tunnel-hauglia", "b")):
|
||||
for bundle_id, approach in (("bygg-energi-mikro", "a"), ("driftssenter-kjoling", "b")):
|
||||
stem = f"X-{bundle_id}"
|
||||
assert (outbox / f"{stem}-{approach}-proposal.json").is_file()
|
||||
assert (outbox / f"{stem}-{approach}-outcome.json").is_file()
|
||||
|
|
@ -165,11 +166,11 @@ def test_each_base_writes_its_own_artefact_set_under_its_own_minted_run_id(tmp_p
|
|||
assert summary["run_id"] == "X"
|
||||
assert [row["bundle_id"] for row in summary["runs"]] == [
|
||||
"bygg-energi-mikro",
|
||||
"tunnel-hauglia",
|
||||
"driftssenter-kjoling",
|
||||
]
|
||||
assert [row["run_id"] for row in summary["runs"]] == [
|
||||
"X-bygg-energi-mikro",
|
||||
"X-tunnel-hauglia",
|
||||
"X-driftssenter-kjoling",
|
||||
]
|
||||
assert summary["unreached"] == []
|
||||
assert summary["collisions"] == []
|
||||
|
|
@ -185,7 +186,7 @@ def test_the_summary_reports_the_stop_reason_each_base_recorded(tmp_path: Path)
|
|||
|
||||
An unparseable proposer burns the round cap, the FIRST approach to hit it re-raises by design
|
||||
(``_evaluate_mandate`` only swallows mid-list), and the pass therefore ends rc 1 — which is
|
||||
exactly what round 3 measured on ``kontrakt-sorasen-2027-04``. The summary must still be
|
||||
exactly what round 3 measured on one context set's ``…-04`` run. The summary must still be
|
||||
there, must say ``rounds`` rather than the empty string that means "the run finished", and
|
||||
must say ``completed: false`` rather than leaving an empty ``unreached`` to be read as
|
||||
"nothing was left unreached".
|
||||
|
|
@ -293,7 +294,7 @@ def test_the_second_base_sees_the_verdict_the_first_base_minted(
|
|||
EVENT: the set of verdict ids already in the store when base 2 began must contain the id base
|
||||
1 minted, which a fresh-store-per-base implementation cannot produce.
|
||||
"""
|
||||
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||
bases = _mount(tmp_path, _BYGG, _KJOLING)
|
||||
seen: list[tuple[int, frozenset[str]]] = []
|
||||
real = run_module.run_project
|
||||
|
||||
|
|
@ -308,7 +309,7 @@ def test_the_second_base_sees_the_verdict_the_first_base_minted(
|
|||
_argv(
|
||||
bases,
|
||||
mandate=_mandate_file(
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "driftssenter-kjoling")]
|
||||
),
|
||||
run_id="X",
|
||||
outbox=tmp_path / "out",
|
||||
|
|
@ -337,7 +338,7 @@ def test_the_requirement_gate_reads_the_second_bases_own_opened_list(tmp_path: P
|
|||
repo's own answer is to call each tool by name. Base 1 opens a document; declaring THAT path
|
||||
against base 2 must refuse, because base 2's own sink never saw it.
|
||||
"""
|
||||
first, second = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||
first, second = _mount(tmp_path, _BYGG, _KJOLING)
|
||||
|
||||
opened_1: list[explore.ToolCall] = []
|
||||
opened_2: list[explore.ToolCall] = []
|
||||
|
|
@ -384,14 +385,14 @@ def test_each_bases_debate_artefact_lists_only_its_own_tool_calls(tmp_path: Path
|
|||
base its own. With one shared sink base 2's artefact would carry base 1's call as well, so the
|
||||
discriminator is the COUNT and not the presence.
|
||||
"""
|
||||
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||
bases = _mount(tmp_path, _BYGG, _KJOLING)
|
||||
outbox = tmp_path / "out"
|
||||
|
||||
rc = run_module.main(
|
||||
_argv(
|
||||
bases,
|
||||
mandate=_mandate_file(
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "driftssenter-kjoling")]
|
||||
),
|
||||
run_id="X",
|
||||
outbox=outbox,
|
||||
|
|
@ -402,7 +403,7 @@ def test_each_bases_debate_artefact_lists_only_its_own_tool_calls(tmp_path: Path
|
|||
|
||||
assert rc == 0
|
||||
counts = []
|
||||
for bundle_id in ("bygg-energi-mikro", "tunnel-hauglia"):
|
||||
for bundle_id in ("bygg-energi-mikro", "driftssenter-kjoling"):
|
||||
payload = json.loads((outbox / f"X-{bundle_id}-debate.json").read_text(encoding="utf-8"))
|
||||
counts.append([call["name"] for call in payload["tool_calls"]])
|
||||
assert counts == [["list_bundles"], ["list_bundles"]], counts
|
||||
|
|
@ -448,13 +449,15 @@ def test_a_dry_run_drills_every_base_and_stops_before_the_first_call(
|
|||
Silently dropping the flag here is the F4 class — the dry run would fall through to the
|
||||
single-project branch, which has no ``PROJECT_ID`` in this argv at all.
|
||||
"""
|
||||
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||
bases = _mount(tmp_path, _BYGG, _KJOLING)
|
||||
|
||||
rc = run_module.main(
|
||||
[
|
||||
*sum([["--across-bundle", b] for b in bases], []),
|
||||
"--mandate",
|
||||
_mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")]),
|
||||
_mandate_file(
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "driftssenter-kjoling")]
|
||||
),
|
||||
"--run-id",
|
||||
"X",
|
||||
"--outbox-dir",
|
||||
|
|
@ -466,11 +469,11 @@ def test_a_dry_run_drills_every_base_and_stops_before_the_first_call(
|
|||
out = capsys.readouterr().out
|
||||
assert rc == 0
|
||||
assert out.count("LIVE-DRY-RUN OK") == 2
|
||||
assert "bygg-energi-mikro" in out and "tunnel-hauglia" in out
|
||||
assert "bygg-energi-mikro" in out and "driftssenter-kjoling" in out
|
||||
# The per-base notices are the discriminator against ONE drill that merely names two bases:
|
||||
# ``bygg-energi-mikro`` ships no ``cost-baseline.json`` and ``tunnel-hauglia`` does, so exactly
|
||||
# one unanchored notice and exactly one grounding offer must appear — and a drill of only the
|
||||
# first, or only the second, produces a different count either way.
|
||||
# ``bygg-energi-mikro`` ships no ``cost-baseline.json`` and ``driftssenter-kjoling`` does, so
|
||||
# exactly one unanchored notice and exactly one grounding offer must appear — and a drill of
|
||||
# only the first, or only the second, produces a different count either way.
|
||||
assert out.count("Grounding offer") == 1
|
||||
assert out.count("Cost baseline: NONE") == 1
|
||||
|
||||
|
|
@ -565,11 +568,13 @@ def test_the_anchoring_requirement_is_carried_to_every_base_not_dropped(
|
|||
would be the F4 class on the one guarantee an operator asked for by name — so the arm asserts
|
||||
the REFUSAL, and its control asserts that the same argv without the flag runs to rc 0.
|
||||
|
||||
``bygg-energi-mikro`` ships no ``cost-baseline.json`` and ``tunnel-hauglia`` does, so the
|
||||
``bygg-energi-mikro`` ships no ``cost-baseline.json`` and ``driftssenter-kjoling`` does, so the
|
||||
refusal must come from the first base rather than from "neither is anchored".
|
||||
"""
|
||||
bases = _mount(tmp_path, _BYGG, _TUNNEL)
|
||||
mandate = _mandate_file(tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "tunnel-hauglia")])
|
||||
bases = _mount(tmp_path, _BYGG, _KJOLING)
|
||||
mandate = _mandate_file(
|
||||
tmp_path, rows=[("a", "bygg-energi-mikro"), ("b", "driftssenter-kjoling")]
|
||||
)
|
||||
replies = _replies_file(tmp_path)
|
||||
outbox = tmp_path / "out"
|
||||
|
||||
|
|
|
|||
|
|
@ -168,7 +168,7 @@ def test_the_demo_entry_point_runs_the_anchored_delivered_bundle() -> None:
|
|||
check=False,
|
||||
)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
assert "kostbaseline erklært (ENERGI-VEGLYS-EL 4386150 x 1)" in proc.stdout, (
|
||||
assert "kostbaseline erklært (ENERGI-KLIENTPARK-EL 4386150 x 1)" in proc.stdout, (
|
||||
"the demo is not anchored on the DELIVERED cost baseline — either it ran a bundle without "
|
||||
"one (the gate then reasons only about the proposal's own numbers), or it fell back to the "
|
||||
"script-derived reserve baseline"
|
||||
|
|
|
|||
|
|
@ -4,7 +4,7 @@ example (Arm A) + an INDEPENDENT deterministic validator rule (Arm B).
|
|||
Arm A (navigable, dimension-scoped): the persisted, MAF-repo-local method bundle
|
||||
(``data/method_examples/energi/``) carries a ``type: methodology`` file with ``dimension: energi``.
|
||||
Its sentinel reaches ``bundle_context(..., dimension="energi")`` but is ABSENT at
|
||||
``dimension="asfalt"`` — the detach is the Step-3 dimension filter (H4: "text is rendered" is not a
|
||||
``dimension="lisens"`` — the detach is the Step-3 dimension filter (H4: "text is rendered" is not a
|
||||
detachable seam on its own; the dimension SCOPING is).
|
||||
|
||||
Arm B (independent validator rule): a proposal the generic P90 stage PASSES but the stricter energy-
|
||||
|
|
@ -52,9 +52,10 @@ def test_persisted_method_navigable_when_dimension_matches() -> None:
|
|||
|
||||
def test_persisted_method_absent_for_other_dimension() -> None:
|
||||
"""LOAD-BEARING (SC7-A): the energi method is ABSENT when the run is scoped to another dimension.
|
||||
RED if the Step-3 dimension filter is detached (the energi method then leaks into an asfalt run)."""
|
||||
RED if the Step-3 dimension filter is detached (the energi method then leaks into a lisens
|
||||
run)."""
|
||||
bundle = okf.navigate_bundle(str(METHOD_BUNDLE))
|
||||
other = okf.bundle_context(bundle, dimension="asfalt")
|
||||
other = okf.bundle_context(bundle, dimension="lisens")
|
||||
assert _METHOD_SENTINEL not in other
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -19,7 +19,7 @@ Detach points, each RED on its own:
|
|||
|
||||
* drop the recorder from the debate middleware -> a called service leaves no trace, and the run
|
||||
reports the same empty record as a run that called nothing;
|
||||
* record every function invocation -> the road path's IN-PROCESS retrieval tool is reported as an
|
||||
* record every function invocation -> the reference path's IN-PROCESS retrieval tool is reported as an
|
||||
external service call, which turns the record into a false egress claim;
|
||||
* attribute an ambiguous tool name to the first server that allows it -> the record names a service
|
||||
that may never have been contacted.
|
||||
|
|
@ -150,7 +150,7 @@ class _ToolCallingClient(ScriptedChatClient):
|
|||
|
||||
#: The argument name each tool actually declares. MEASURED, and load-bearing: a call whose argument
|
||||
#: names do not match the tool's signature is rejected by MAF BEFORE invocation, so the middleware
|
||||
#: never fires — the first version of the road-path test below passed ``code`` to
|
||||
#: never fires — the first version of the reference-path test below passed ``code`` to
|
||||
#: ``retrieve_cost_docs(query)`` and was therefore VACUOUS. It asserted an empty record against a run
|
||||
#: where no tool ran at all, and stayed green under the mutation it exists to catch.
|
||||
_ARGUMENTS = {
|
||||
|
|
@ -223,7 +223,7 @@ async def test_a_called_mcp_tool_is_recorded_in_provenance(monkeypatch: pytest.M
|
|||
async def test_a_local_tool_call_is_not_recorded_as_an_external_call(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""The road path's ``retrieve_cost_docs`` runs IN PROCESS. Calling it is not egress, and a run
|
||||
"""The reference path's ``retrieve_cost_docs`` runs IN PROCESS. Calling it is not egress, and a run
|
||||
that only called it must not claim it contacted the price register.
|
||||
|
||||
RED when the recorder logs every function invocation rather than the configured ones.
|
||||
|
|
@ -246,12 +246,12 @@ async def test_a_local_tool_call_is_not_recorded_as_an_external_call(
|
|||
|
||||
monkeypatch.setattr(ToolCallRecorder, "note", spy)
|
||||
|
||||
rv13 = {p.id: p for p in load_reference_projects()}["RV13-RAS-TP"]
|
||||
nett = {p.id: p for p in load_reference_projects()}["NETT-SIKR-TP"]
|
||||
result = await run_project(
|
||||
"RV13-RAS-TP",
|
||||
"NETT-SIKR-TP",
|
||||
"local",
|
||||
docs_dir=rv13.docs_dir,
|
||||
verdict_input=rv13.verdict_input,
|
||||
docs_dir=nett.docs_dir,
|
||||
verdict_input=nett.verdict_input,
|
||||
store=VerdictStore(verdicts=[]),
|
||||
client_factory=_factory("retrieve_cost_docs"),
|
||||
mcp_servers=(_SERVER,),
|
||||
|
|
|
|||
|
|
@ -601,7 +601,7 @@ _WRITERS_DEFINED = 10
|
|||
_WRITERS_IN_RUN_PATH = 7
|
||||
#: publiserte filer i git-manifestet (uttrekk og arbeidstre gir SAMME tall — det var hele poenget
|
||||
#: med å slutte å telle filtreet: 435 i arbeidstreet var to gitignorerte .local.md-filer).
|
||||
_PUBLISHED_TODAY = 519
|
||||
_PUBLISHED_TODAY = 1427
|
||||
_UNDECODABLE_TODAY = 1
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -47,7 +47,7 @@ _WHERE_THE_SAVING_GOES = "claimed_saving_nok"
|
|||
_APPROACH = Approach(id="led", label="LED retrofit", description="fixtures past rated life")
|
||||
_REJECTION = Rejection(
|
||||
proposal=SavingsProposal(
|
||||
project_id="FV42-GSV-E1",
|
||||
project_id="KONTOR-IT-E1",
|
||||
measure="Reduce scope",
|
||||
affected_items=[{"code": "05.2", "quantity": 1000, "unit_cost": 215}],
|
||||
claimed_saving_nok=10000,
|
||||
|
|
@ -59,7 +59,7 @@ _FEEDBACK = "Halve the volume rather than the unit price."
|
|||
|
||||
@pytest.fixture(scope="module")
|
||||
def project():
|
||||
return load_reference_projects()[0] # FV42-GSV-E1
|
||||
return load_reference_projects()[0] # KONTOR-IT-E1
|
||||
|
||||
|
||||
def _base(project) -> str:
|
||||
|
|
|
|||
|
|
@ -7,7 +7,7 @@ which is what keeps every commons-owned golden bundle running). The anchoring st
|
|||
purpose — and that is not what this file changes.
|
||||
|
||||
What it changes is that the skip was INVISIBLE. Measured (session 48, ``9d149b3``): four
|
||||
``--live-dry-run``s over copies of the veglys bundle — intact rc 0 · without ``validator-input.json``
|
||||
``--live-dry-run``s over copies of a delivered example bundle — intact rc 0 · without ``validator-input.json``
|
||||
rc 1 · **without ``cost-baseline.json`` rc 0 with no message at all** · corrupt baseline rc 1. And
|
||||
``grep baseline provenance.py outbox.py`` returned 0 hits, so neither the stamp nor the outbox
|
||||
artefacts carried it either. An operator could therefore run the whole gate un-anchored, read a
|
||||
|
|
@ -35,7 +35,7 @@ four bare dry-runs above, none of which had a mandate — would still have print
|
|||
|
||||
Arms:
|
||||
(a) the provenance field is ``False`` on an un-anchored bundle run and ``True`` on an anchored one,
|
||||
end-to-end through ``run_project`` (+ the road path, which is anchored by construction);
|
||||
end-to-end through ``run_project`` (+ the reference path, which is anchored by construction);
|
||||
(b) the notice EXISTS un-anchored and is ABSENT anchored — asserted on a sentinel that the anchored
|
||||
branch cannot contain, because it prints no line at all (never a substring both branches share:
|
||||
the 08-09 class);
|
||||
|
|
@ -138,10 +138,10 @@ async def test_provenance_records_an_anchored_bundle_run(fresh_store) -> None:
|
|||
|
||||
|
||||
async def test_road_path_is_anchored_by_construction(docs_dir, fresh_store) -> None:
|
||||
"""The road path derives its baseline from the reference project's own ``cost_items``, so it is
|
||||
"""The reference path derives its baseline from the reference project's own ``cost_items``, so it is
|
||||
ALWAYS anchored — the stamp says so rather than leaving the reader to know it."""
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
|
|
@ -309,7 +309,7 @@ def test_portfolio_surface_announces_an_unanchored_run(monkeypatch, capsys) -> N
|
|||
|
||||
Driven by a CRAFTED ``PortfolioResult`` (the ``budget_stop`` precedent in
|
||||
``test_portfolio_cli_offline_loadbearing``), and for the same measured reason: no reference
|
||||
project sets ``bundle_dir``, so every portfolio run today takes the road path and is anchored by
|
||||
project sets ``bundle_dir``, so every portfolio run today takes the reference path and is anchored by
|
||||
construction. This arm is therefore DEFENSIVE — it guards the surface for the day a bundle-backed
|
||||
project is wired into a pass, rather than covering a path reachable now."""
|
||||
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
"""P19 DEL A — a direction must NAME the requirement that binds it, and it must have READ it.
|
||||
|
||||
**The measured silence, two paid rounds deep.** P16 (`docs/2026-09-12-p14-kontekstsett.md` § 4.1,
|
||||
`docs/2026-09-14-p18-stressrunde-2.md`) and P18 both scored **0 of 26** fasit concepts opened —
|
||||
**The measured silence, two paid rounds deep.** P16 and P18 (the ledger is
|
||||
`docs/invarianter.md`) both scored **0 of 26** fasit concepts opened —
|
||||
the same number twice, over four and then five paid runs, while every run still produced proposals
|
||||
the deterministic gate then judged. P18 closed the navigation side of it (a listing is a window; an
|
||||
invented path is refused by name) and the number did not move at all. That is what turns it from a
|
||||
|
|
@ -72,12 +72,12 @@ from portfolio_optimiser.generate import _build_messages
|
|||
from portfolio_optimiser.mandate import Approach, BindingRequirement
|
||||
from portfolio_optimiser.reference_domain import Project
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
_KJOLING = _BUNDLES / "driftssenter-kjoling"
|
||||
|
||||
|
||||
def _wired(
|
||||
bundle_dir: Path = _TUNNEL,
|
||||
bundle_dir: Path = _KJOLING,
|
||||
) -> tuple[dict[str, Any], list[ToolCall], list[DeclaredRequirement]]:
|
||||
"""The tools as a RUN holds them: the declaration rung reading the recorder's own trace."""
|
||||
opened: list[ToolCall] = []
|
||||
|
|
@ -86,7 +86,7 @@ def _wired(
|
|||
return {t.name: t for t in tools}, opened, declared
|
||||
|
||||
|
||||
def _a_concept(bundle_dir: Path = _TUNNEL) -> str:
|
||||
def _a_concept(bundle_dir: Path = _KJOLING) -> str:
|
||||
"""One real concept path in the base, taken from the navigated listing rather than guessed."""
|
||||
from portfolio_optimiser import okf
|
||||
|
||||
|
|
@ -102,7 +102,7 @@ def test_a_requirement_the_run_never_opened_is_refused_and_recorded_nowhere() ->
|
|||
"""(a) The declaration's one falsifier is the run's own read trace."""
|
||||
tools, _opened, declared = _wired()
|
||||
answer = tools["declare_requirement"].func(
|
||||
approach_id="a1", bundle_id="tunnel-hauglia", path=_a_concept(), ref="12.1"
|
||||
approach_id="a1", bundle_id="driftssenter-kjoling", path=_a_concept(), ref="12.1"
|
||||
)
|
||||
assert answer["refusal"] == "RequirementNotRead"
|
||||
assert "0 document(s)" in answer["refused"]
|
||||
|
|
@ -120,17 +120,17 @@ def test_the_correction_is_to_read_it_and_then_it_is_accepted() -> None:
|
|||
tools, opened, declared = _wired()
|
||||
from portfolio_optimiser import okf
|
||||
|
||||
files = [f.name for f in okf.navigate_bundle(str(_TUNNEL)).context_files]
|
||||
files = [f.name for f in okf.navigate_bundle(str(_KJOLING)).context_files]
|
||||
path = files[0]
|
||||
body = tools["read_file"].func(bundle_id="tunnel-hauglia", path=path)
|
||||
body = tools["read_file"].func(bundle_id="driftssenter-kjoling", path=path)
|
||||
assert not body.startswith("REFUSED"), body[:120]
|
||||
# The recorder is middleware in a real run; here the trace is appended directly, which is the
|
||||
# SAME list the tool reads.
|
||||
for other in files[1:3]:
|
||||
opened.append(ToolCall(name="read_file", bundle_id="tunnel-hauglia", path=other))
|
||||
opened.append(ToolCall(name="read_file", bundle_id="tunnel-hauglia", path=path))
|
||||
opened.append(ToolCall(name="read_file", bundle_id="driftssenter-kjoling", path=other))
|
||||
opened.append(ToolCall(name="read_file", bundle_id="driftssenter-kjoling", path=path))
|
||||
answer = tools["declare_requirement"].func(
|
||||
approach_id="a1", bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1"
|
||||
approach_id="a1", bundle_id="driftssenter-kjoling", path=path, ref="Krav 12.1"
|
||||
)
|
||||
# P20/A1 widened the reply: the three arguments PLUS the document's own title and number and
|
||||
# the sentence saying what the declaration binds. Asserted key by key rather than by equality,
|
||||
|
|
@ -138,7 +138,7 @@ def test_the_correction_is_to_read_it_and_then_it_is_accepted() -> None:
|
|||
# the one property this arm exists for — that the declaration was ACCEPTED and RECORDED.
|
||||
assert answer["declared"] is True
|
||||
assert (answer["bundle_id"], answer["path"], answer["ref"]) == (
|
||||
"tunnel-hauglia",
|
||||
"driftssenter-kjoling",
|
||||
path,
|
||||
"Krav 12.1",
|
||||
)
|
||||
|
|
@ -154,7 +154,7 @@ def test_the_correction_is_to_read_it_and_then_it_is_accepted() -> None:
|
|||
}
|
||||
assert declared == [
|
||||
DeclaredRequirement(
|
||||
bundle_id="tunnel-hauglia", path=path, ref="Krav 12.1", approach_id="a1"
|
||||
bundle_id="driftssenter-kjoling", path=path, ref="Krav 12.1", approach_id="a1"
|
||||
)
|
||||
]
|
||||
|
||||
|
|
@ -172,14 +172,14 @@ def test_a_declaration_naming_an_unknown_base_is_refused() -> None:
|
|||
|
||||
def test_the_rung_exists_only_when_both_sinks_are_offered() -> None:
|
||||
"""(d) Every pre-P19 call site is byte-identical, and a half-wired one is refused."""
|
||||
plain = {t.name for t in navigator_tools((str(_TUNNEL),))}
|
||||
plain = {t.name for t in navigator_tools((str(_KJOLING),))}
|
||||
assert plain == {"list_bundles", "read_bundle", "read_dir", "read_file"}
|
||||
wired, _, _ = _wired()
|
||||
assert set(wired) == plain | {"declare_requirement"}
|
||||
with pytest.raises(explore.ExplorationError, match="together or not at all"):
|
||||
navigator_tools((str(_TUNNEL),), requirements=[])
|
||||
navigator_tools((str(_KJOLING),), requirements=[])
|
||||
with pytest.raises(explore.ExplorationError, match="together or not at all"):
|
||||
navigator_tools((str(_TUNNEL),), opened=[])
|
||||
navigator_tools((str(_KJOLING),), opened=[])
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
|
@ -238,16 +238,16 @@ def test_the_binding_requirement_reaches_the_proposer_verbatim() -> None:
|
|||
"""(g) A3, and the byte-identical half is what keeps the golden transcript unchanged."""
|
||||
bare = Approach(id="a1", label="L", description="D")
|
||||
with_req = bare.model_copy(
|
||||
update={"requirement": BindingRequirement(path="R761/12-1/x.md", ref="12.1")}
|
||||
update={"requirement": BindingRequirement(path="P900/12-1/x.md", ref="12.1")}
|
||||
)
|
||||
without = _build_messages(_project(), "ctx", approach=bare)[0].text
|
||||
withit = _build_messages(_project(), "ctx", approach=with_req)[0].text
|
||||
assert "Binding requirement" not in without
|
||||
assert "Binding requirement: 12.1 (R761/12-1/x.md)" in withit
|
||||
assert "Binding requirement: 12.1 (P900/12-1/x.md)" in withit
|
||||
# The ONLY difference is the block — an approach without one is byte-identical to before.
|
||||
assert (
|
||||
withit.replace(
|
||||
"Binding requirement: 12.1 (R761/12-1/x.md)\nName that requirement verbatim in 'measure'.\n",
|
||||
"Binding requirement: 12.1 (P900/12-1/x.md)\nName that requirement verbatim in 'measure'.\n",
|
||||
"",
|
||||
)
|
||||
== without
|
||||
|
|
|
|||
|
|
@ -1,17 +1,18 @@
|
|||
"""P13b: a BLOCK sequence of mappings is provenance po can read, not provenance po cannot.
|
||||
|
||||
RED-FIRST, measured 2026-09-12 with the full denominator: every concept file in all four delivered
|
||||
knowledge bases writes ``sources`` as a BLOCK sequence — n100 446/446, n200 1133/1133, n500 270/270,
|
||||
r761 2756/2756, **4605 of 4605, and 0 in flow form**. ``read_provenance`` answered
|
||||
``UnreadableProvenance(reason="block-sequence")`` for every one of them, so ``evidence_for`` reported
|
||||
``state="unreadable"`` on 4605 of 4605 documents: the falsification layer had no address for any
|
||||
document in any base po is about to be stress-tested against.
|
||||
RED-FIRST, measured 2026-09-12 with the full denominator: every concept file in all four
|
||||
knowledge bases delivered at the time wrote ``sources`` as a BLOCK sequence — **4605 of 4605, and 0
|
||||
in flow form**. ``read_provenance`` answered ``UnreadableProvenance(reason="block-sequence")`` for
|
||||
every one of them, so ``evidence_for`` reported ``state="unreadable"`` on 4605 of 4605 documents:
|
||||
the falsification layer had no address for any document in any base po was about to be
|
||||
stress-tested against. The package's two example bases write the same form (``sources`` read as
|
||||
entries in 306 of 306 and 301 of 301 concept files, re-measured below rather than quoted).
|
||||
|
||||
**This widens the READER and nothing else.** The producer still EMITS flow (measured: okf's
|
||||
``materialize._render_sources`` is unchanged in 0.8.5), so no bundle bytes move, and
|
||||
``write_concept_file``/``verified_field`` still refuse what the flow decoder refuses — the writer
|
||||
keeps refusing exactly what the reader could not read, which is the property the round-trip gate
|
||||
exists for. What changes is that a form the producer's own SPEC §5.1 documents, and that four
|
||||
exists for. What changes is that a form the producer's own SPEC §5.1 documents, and that the
|
||||
delivered bases actually use, stops being reported as unreadable.
|
||||
|
||||
**Two refusals are KEPT, and they are this file's discriminators.** ``block-mapping`` (an indented
|
||||
|
|
@ -32,14 +33,16 @@ from pathlib import Path
|
|||
|
||||
import pytest
|
||||
|
||||
from portfolio_optimiser import okf
|
||||
from portfolio_optimiser import frozen_bundles, okf
|
||||
|
||||
_FIXTURES = Path(__file__).resolve().parent / "fixtures" / "p13b-block-sources"
|
||||
|
||||
#: A byte-identical copy of a REAL delivered concept file (n100-2023, `krav/N100/
|
||||
#: id-2b69893f-…md`, sha1 8c2952d1…). Read from the producer's tree, never written there. A
|
||||
#: hand-written approximation would prove the reader handles what I imagined the form to be.
|
||||
_REAL = _FIXTURES / "n100-krav-4-2-5-1-3.md"
|
||||
#: A byte-identical copy of a REAL concept file of the example base (driftskrav-2027,
|
||||
#: ``krav/D100/id-7aed3659-…md``, sha1 0242b9f9…) — pinned byte-for-byte by the arm below, so it
|
||||
#: cannot drift from what the base ships. A hand-written approximation would prove the reader
|
||||
#: handles what I imagined the form to be.
|
||||
_REAL = _FIXTURES / "driftskrav-krav-4-2-5-1-3.md"
|
||||
_REAL_IN_BASE = "krav/D100/id-7aed3659-dbbe-5eb5-9b96-e6a693ff0c3c.md"
|
||||
|
||||
|
||||
def _write(tmp_path: Path, frontmatter: str) -> Path:
|
||||
|
|
@ -54,12 +57,27 @@ def test_the_real_delivered_concept_file_yields_its_address() -> None:
|
|||
assert isinstance(entries, tuple), f"still unreadable: {entries!r}"
|
||||
assert entries == (
|
||||
{
|
||||
"resource": "https://viewers.vegnorm.vegvesen.no/api/nisosts/859984?languageCode=nb",
|
||||
"title": "N100:2023",
|
||||
"resource": "https://example.invalid/eksempelvirksomheten/driftskrav/d100",
|
||||
"title": "D100:2027",
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
def test_the_fixture_is_the_base_file_byte_for_byte_and_every_base_file_reads() -> None:
|
||||
"""The fixture is only evidence while it IS the delivered file: pinned against the base the
|
||||
package ships. And the one file stands for the whole population — every concept file of both
|
||||
example bases yields its ``sources`` as entries, with the denominator named."""
|
||||
base = frozen_bundles.bundle_dir("driftskrav-2027")
|
||||
assert _REAL.read_bytes() == (base / _REAL_IN_BASE).read_bytes()
|
||||
counts = {}
|
||||
for name in ("driftskrav-2027", "prosesskatalog-2027"):
|
||||
root = frozen_bundles.bundle_dir(name)
|
||||
files = okf.navigate_bundle(str(root)).context_files
|
||||
read = sum(isinstance(okf.read_provenance(root / f.name, "sources"), tuple) for f in files)
|
||||
counts[name] = (read, len(files))
|
||||
assert counts == {"driftskrav-2027": (306, 306), "prosesskatalog-2027": (301, 301)}, counts
|
||||
|
||||
|
||||
def test_evidence_for_reports_present_on_the_real_file() -> None:
|
||||
"""The three-state answer a falsification verdict acts on, end to end. ``tier`` stays ``None``
|
||||
because ``sources`` is not the key SPEC §5.3 tiers — the B4 rule, unchanged by this widening."""
|
||||
|
|
@ -99,23 +117,23 @@ def test_the_key_after_a_block_sequence_is_still_read(tmp_path: Path) -> None:
|
|||
|
||||
|
||||
def test_a_value_carrying_a_colon_survives_whole(tmp_path: Path) -> None:
|
||||
"""``title: N100:2023`` is one pair and the value keeps its own colon, as does a URL's ``://``.
|
||||
"""``title: D100:2027`` is one pair and the value keeps its own colon, as does a URL's ``://``.
|
||||
|
||||
**This arm is a REGRESSION GUARD, not a discriminator, and the distinction was measured rather
|
||||
than assumed.** It was written claiming to prove the colon-SPACE rule, and mutation M5 (separator
|
||||
changed to the FIRST colon) left it GREEN: the first colon in ``resource: https://…`` and in
|
||||
``title: N100:2023`` is the same one colon-SPACE finds, so the two rules agree on every value
|
||||
``title: D100:2027`` is the same one colon-SPACE finds, so the two rules agree on every value
|
||||
shaped like a delivered one. The arm that actually tells them apart is
|
||||
``test_an_item_with_no_pair_separator_is_refused`` — under first-colon, ``- https://a.example/d``
|
||||
decodes to ``{'https': '//a.example/d'}``, a key invented out of a URL scheme. Kept because it
|
||||
pins the values four delivered bases actually carry; labelled honestly because a test that
|
||||
pins the values the delivered bases actually carry; labelled honestly because a test that
|
||||
cannot separate two implementations proves nothing about them."""
|
||||
path = _write(
|
||||
tmp_path,
|
||||
"type: concept\nsources:\n - resource: https://x.example/a\n title: N100:2023\n",
|
||||
"type: concept\nsources:\n - resource: https://x.example/a\n title: D100:2027\n",
|
||||
)
|
||||
assert okf.read_provenance(path, "sources") == (
|
||||
{"resource": "https://x.example/a", "title": "N100:2023"},
|
||||
{"resource": "https://x.example/a", "title": "D100:2027"},
|
||||
)
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
"""P16 B2 - the CLI had NO door onto a paid run's cap, and the documented command proved it.
|
||||
|
||||
**The measured silence.** ``STATE.md``, ``docs/2026-09-12-p14-kontekstsett.md § 4.1`` and order
|
||||
**The measured silence.** ``STATE.md``, the context-set documentation of the time and order
|
||||
``20260914T091846Z`` all publish the same stress command, ending ``--max-rounds 8 --max-tokens
|
||||
120000``. Measured 14.09: ``run.py`` accepts neither flag, all four free ``--live-dry-run`` drills
|
||||
refused with ``unrecognized arguments``, and ``main()`` never passed ``max_rounds``/``max_tokens``
|
||||
|
|
|
|||
|
|
@ -225,7 +225,7 @@ async def test_the_road_path_stamps_no_bundle_identity_at_all(
|
|||
"""(e, control) A run with no knowledge base has no bundle identity, and says so by ABSENCE
|
||||
rather than by inventing one. Without this arm a constant ``declared-index`` would pass above."""
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
|
|
@ -362,7 +362,7 @@ async def test_two_mounts_declaring_one_id_collide_in_the_dispatcher(tmp_path: P
|
|||
|
||||
def test_the_stamp_field_has_no_default() -> None:
|
||||
"""(l) ``cost_baseline_anchored``'s rule: a stamp that forgot to say which corpus it judged must
|
||||
not construct at all. ``None`` is a VALUE here (the road path), which is precisely why the
|
||||
not construct at all. ``None`` is a VALUE here (the reference path), which is precisely why the
|
||||
absence of the field cannot be allowed to mean it."""
|
||||
with pytest.raises(Exception):
|
||||
ProvenanceStamp( # type: ignore[call-arg]
|
||||
|
|
|
|||
|
|
@ -1,15 +1,16 @@
|
|||
"""The catalogue call costs O(bases), never O(corpus) — and what it drops, it SAYS it dropped.
|
||||
|
||||
Measured 2026-08-25 (session 60, ``docs/2026-08-25-syretest-vei-ab.md``) and re-measured 26.08 with
|
||||
Measured 2026-08-25 (session 60; the ledger is ``docs/invarianter.md``) and re-measured 26.08 with
|
||||
the same instrument (``tiktoken`` ``o200k_base``, run through ``uv run --with tiktoken``, validated
|
||||
first against the three commons example bundles whose numbers commons itself publishes):
|
||||
|
||||
list_bundles() over the three flat Vegnormal bases -> 201 196 chars / 112 116 tokens
|
||||
list_bundles() over the 171 branch bases -> 234 611 chars / 124 942 tokens
|
||||
list_bundles() over three flat requirement bases -> 201 196 chars / 112 116 tokens
|
||||
list_bundles() over the 171 branch bases -> 234 611 chars / 124 942 tokens
|
||||
|
||||
The branch form (``vegnormal-okf`` ``8145c23``) closed the *bundle* side — ``read_bundle`` fell 82-92
|
||||
percent — and made the *catalogue* side WORSE, exactly as that repo predicted: one call now costs
|
||||
more than a 128k window, before the manager has read a single document.
|
||||
The branch form (a rebuild of those corpora by the project that produces them) closed the
|
||||
*bundle* side — ``read_bundle`` fell 82-92 percent — and made the *catalogue* side WORSE, exactly
|
||||
as that project predicted: one call now costs more than a 128k window, before the manager has read
|
||||
a single document.
|
||||
|
||||
The cause is in this repo. ``list_bundles`` returned ``Bundle.index_summary`` — the WHOLE root index
|
||||
body — for EVERY configured base at once, plus one JSON object per unfollowed cross-link. Both grow
|
||||
|
|
@ -19,8 +20,8 @@ cheapest rung of the ladder, and it was the most expensive.
|
|||
|
||||
**A MEASURED premise, felled before anything was built on it:** "the index body tells a manager what
|
||||
the base is about" is FALSE for machine-imported bases. The branch bases' ``index.md`` carries no
|
||||
frontmatter and no prose — it is a pure link list (measured: ``B-n200-2024-gren-1-1-importert``,
|
||||
959 bytes, first byte is ``-``). So the old field was not merely expensive, it was expensive AND
|
||||
frontmatter and no prose — it is a pure link list (measured on one branch base: 959 bytes, first
|
||||
byte is ``-``). So the old field was not merely expensive, it was expensive AND
|
||||
uninformative there; a truncated prefix loses nothing a manager was using.
|
||||
|
||||
**The ceiling lives in this file, not in ``explore.py``.** A test that imported the implementation's
|
||||
|
|
|
|||
|
|
@ -7,24 +7,23 @@ in the set is constructed rather than real).
|
|||
|
||||
**Five arms, and the split between them is a measurement rather than a taste.** Two are
|
||||
unconditional and can never be silently absent — a mandate that does not load, and a mandate routed
|
||||
at a base the set is not for. Three need the base itself, which lives OUTSIDE this repository
|
||||
(``PORTFOLIO_VEGNORMAL_ROOT``): they SKIP when the root is missing, exactly as MAJOR-3's ceiling
|
||||
gate could not take K2 as a test dependency, and for the same published-package reason — a hard
|
||||
failure would break ``uv run pytest`` for any external recipient of the ``git archive HEAD``
|
||||
handover. The skip NAMES the root it looked for.
|
||||
at a base the set is not for. Three need the base itself. The two example bases ship WITH the
|
||||
package (``data/kunnskapsbaser/``, pinned in ``frozen_bundles.json``), so those arms run in every
|
||||
checkout and every installed wheel; they SKIP only when a user has pointed the store elsewhere
|
||||
(``PORTFOLIO_FROZEN_BUNDLES``) and the pinned copy is not there, and the skip NAMES the store it
|
||||
looked in.
|
||||
|
||||
**Every bundle-reading arm carries its own denominator.** A scan that sees zero concepts is RED
|
||||
rather than vacuously green: "the anchor was not found" is equally true of a base that was never
|
||||
read (Verifiseringsloven, ansikt 4).
|
||||
|
||||
**Rule U** — the measurable form of "the base cannot answer this" (documented in
|
||||
``docs/2026-09-12-p14-kontekstsett.md § 2.3``): each ``must_refuse`` row declares >= 1 ``anchor``,
|
||||
``docs/invarianter.md``): each ``must_refuse`` row declares >= 1 ``anchor``,
|
||||
a lowercase word of >= 4 characters, and is admitted **iff every anchor is absent — case-insensitive
|
||||
substring — from the WHOLE text (frontmatter + body) of EVERY concept document in the base**. Not
|
||||
"shares no keyword with any title": a tunnel question shares "tunnel" with hundreds of titles and
|
||||
"shares no keyword with any title": a cooling question shares "kjøling" with dozens of titles and
|
||||
that proves nothing. What makes a question unanswerable is that the base lacks the SUBJECT, and the
|
||||
anchor is that subject. Titles alone would be a proxy the full text costs nothing more to replace
|
||||
(measured: 0.77 s for r761-2025, the largest base).
|
||||
anchor is that subject. Titles alone would be a proxy the full text costs nothing more to replace.
|
||||
|
||||
**P16 A2 moved rule U from ``unanswerable`` to ``must_refuse``, and that is ONE form rather than
|
||||
two.** ``po`` is not a lookup tool (D-1), so an "unanswerable question" had no runnable form: no
|
||||
|
|
@ -56,9 +55,9 @@ _CONTEXT_ROOT = _REPO_ROOT / "contexts"
|
|||
#: in the set: this repository is published, and an absolute path would pin a set to one machine's
|
||||
#: home directory and ride out in the handover archive.
|
||||
|
||||
#: The concept types the four bases declare. ``index.md`` carries none of them — it is navigation,
|
||||
#: not content — which is why the file count and the concept count differ.
|
||||
_CONCEPT_TYPES = {"Krav", "Prosess", "Kapittel", "Normal", "Håndbok"}
|
||||
#: The concept types the two example bases declare. ``index.md`` carries none of them — it is
|
||||
#: navigation, not content — which is why the file count and the concept count differ.
|
||||
_CONCEPT_TYPES = {"Krav", "Prosess", "Standard", "Katalog"}
|
||||
|
||||
_MIN_ANCHOR_CHARS = 4
|
||||
|
||||
|
|
@ -68,29 +67,30 @@ def own_frontmatter(path: Path) -> dict[str, str]:
|
|||
|
||||
**P15 (2026-09-13) fixed the finding this helper was written against.** Before P15,
|
||||
``okf.parse_frontmatter`` was linewise and last-write-wins over EVERY line regardless of
|
||||
indentation, so a nested block overwrote a top-level key of the same name. Every vegnormal
|
||||
concept ends its frontmatter with
|
||||
indentation, so a nested block overwrote a top-level key of the same name. Every concept of a
|
||||
requirements base ends its frontmatter with
|
||||
|
||||
sources:
|
||||
- resource: https://…
|
||||
title: N500:2024
|
||||
title: D200:2027
|
||||
|
||||
and the indented ``title`` used to replace the concept's own. MEASURED on n500-2024 before the
|
||||
fix: ``okf.navigate_bundle`` yielded 270 concept files carrying **1 distinct title**
|
||||
(``N500:2024``, 270 times). ``okf.parse_frontmatter`` now makes indentation load-bearing —
|
||||
a top-level (unindented) key always wins over a nested one of the same name — and re-measured
|
||||
AFTER the fix, the same base's 269 ``krav/N500`` documents carry **269 distinct titles**.
|
||||
and the indented ``title`` used to replace the concept's own. MEASURED on a delivered
|
||||
requirements corpus during development, before the fix: ``okf.navigate_bundle`` yielded 270
|
||||
concept files carrying **1 distinct title** (the sources title, 270 times).
|
||||
``okf.parse_frontmatter`` now makes indentation load-bearing — a top-level (unindented) key
|
||||
always wins over a nested one of the same name — and re-measured AFTER the fix, the same base's
|
||||
269 requirement documents carried **269 distinct titles**.
|
||||
|
||||
**This helper still isn't a plain call to ``okf.parse_frontmatter``, and that remains
|
||||
measured rather than assumed:** ``own_frontmatter`` also strips one layer of enclosing
|
||||
``'`` quotes (``.strip("'")``) so a value matches the fasit's stored plain-text title
|
||||
verbatim, while ``okf.parse_frontmatter`` deliberately leaves scalars quoted — unquoting is
|
||||
``okf.unquote_scalar``'s ONE job (D1/(a)/(i)), and a second copy of that rule here would be
|
||||
the drifting one. Re-measured across all four bases (29 500 field reads: ``type``, ``title``,
|
||||
``req_number``, ``prosessnr`` on every concept file) the two now agree EXACTLY except for
|
||||
quoted scalars (2 728 of 29 500 checks — every one a quote-stripping difference, none a value
|
||||
difference), so this helper stays for that one reason, not for the nested-override bug P15
|
||||
closed.
|
||||
the drifting one. Measured across the corpora of the time (29 500 field reads: ``type``,
|
||||
``title``, ``req_number``, ``prosessnr`` on every concept file) the two agreed EXACTLY except
|
||||
for quoted scalars (2 728 of 29 500 checks — every one a quote-stripping difference, none a
|
||||
value difference), so this helper stays for that one reason, not for the nested-override bug
|
||||
P15 closed.
|
||||
|
||||
Uses ``okf._split_frontmatter`` deliberately: it is the module's ONE place ``---`` is compared
|
||||
(B4), and a second delimiter rule here would be the copy that drifts.
|
||||
|
|
@ -108,8 +108,8 @@ def own_frontmatter(path: Path) -> dict[str, str]:
|
|||
def _bundle_dir(name: str) -> Path:
|
||||
"""The FROZEN copy this repository pins, resolved at call time.
|
||||
|
||||
Absence SKIPS (MAJOR-3's ceiling: no corpus is mounted in the handover archive), drift is
|
||||
allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
Absence SKIPS (a user's own store, named by ``PORTFOLIO_FROZEN_BUNDLES``, may not hold it),
|
||||
drift is allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
"""
|
||||
try:
|
||||
return frozen_bundles.bundle_dir(name)
|
||||
|
|
@ -120,7 +120,7 @@ def _bundle_dir(name: str) -> Path:
|
|||
#: The ONE reader, imported from production rather than copied here (P17b). It used to be a
|
||||
#: private copy in this file and a second, looser one inside ``stress.main`` — and the multi-base
|
||||
#: form is exactly the change that would have let the two drift into different answers about one
|
||||
#: set. A set declaring ONE base is one block, so the four pre-P17b files parse unchanged.
|
||||
#: set. A set declaring ONE base is one block.
|
||||
read_bundle_txt = read_bundle_declarations
|
||||
|
||||
|
||||
|
|
@ -197,8 +197,8 @@ def _base_by_approach(set_dir: Path) -> dict[str, Path]:
|
|||
# --------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_the_five_context_sets_are_present() -> None:
|
||||
assert len(_SETS) == 5, f"expected five context sets under {_CONTEXT_ROOT}, found {_SET_IDS}"
|
||||
def test_the_three_context_sets_are_present() -> None:
|
||||
assert len(_SETS) == 3, f"expected three context sets under {_CONTEXT_ROOT}, found {_SET_IDS}"
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------------------------
|
||||
|
|
@ -268,8 +268,8 @@ def test_b_every_fasit_concept_is_in_the_base_as_recorded(set_dir: Path) -> None
|
|||
|
||||
Resolving per approach rather than per set is the multi-base half: in a set spanning two
|
||||
bases, checking every path against one of them would fail half the fasit while proving
|
||||
nothing about the other, and checking against "either" would let a path meant for N200 be
|
||||
satisfied by a coincidence in R761.
|
||||
nothing about the other, and checking against "either" would let a path meant for the
|
||||
requirements base be satisfied by a coincidence in the process catalogue.
|
||||
"""
|
||||
bases = _base_by_approach(set_dir)
|
||||
fasit = json.loads((set_dir / "fasit.json").read_text(encoding="utf-8"))
|
||||
|
|
@ -299,10 +299,11 @@ def test_c_rule_u_every_unanswerable_question_is_unanswerable(set_dir: Path) ->
|
|||
"""Rule U over EVERY base the set declares, as ONE scan.
|
||||
|
||||
For a multi-base set "the base cannot answer this" becomes "NEITHER base can", and the union
|
||||
is the honest reading: an anchor absent from N200 but present in R761 is a question the pass
|
||||
as a whole CAN reach. MEASURED 15.09 and the reason this is not a formality — ``enhetspris``
|
||||
is absent from n200-2024 and carried by 70 of r761-2025's 2 756 concepts, so an anchor set
|
||||
admitted per base would have admitted a question the pass could ground.
|
||||
is the honest reading: an anchor absent from the requirements base but present in the process
|
||||
catalogue is a question the pass as a whole CAN reach. MEASURED on the example bases, and the
|
||||
reason this is not a formality — ``enhetspris`` is absent from driftskrav-2027 and carried by 57
|
||||
of prosesskatalog-2027's 301 concepts, so an anchor set admitted per base would have admitted a
|
||||
question the pass could ground.
|
||||
"""
|
||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
||||
concepts: list[tuple[str, dict[str, str], str]] = []
|
||||
|
|
@ -359,18 +360,18 @@ _GOOD_MANDATE = {
|
|||
"label": "One",
|
||||
"affected_codes": ["X-1"],
|
||||
"claimed_saving_nok": 1.0,
|
||||
"bundle_id": "vegnormal-n500-2024",
|
||||
"bundle_id": "eksempel-driftskrav-2027",
|
||||
},
|
||||
{
|
||||
"id": "a2",
|
||||
"label": "Two",
|
||||
"affected_codes": ["X-2"],
|
||||
"claimed_saving_nok": 2.0,
|
||||
"bundle_id": "vegnormal-n500-2024",
|
||||
"bundle_id": "eksempel-driftskrav-2027",
|
||||
},
|
||||
],
|
||||
}
|
||||
_GOOD_BUNDLE_TXT = "name: n500-2024\nbundle_id: vegnormal-n500-2024\n"
|
||||
_GOOD_BUNDLE_TXT = "name: driftskrav-2027\nbundle_id: eksempel-driftskrav-2027\n"
|
||||
|
||||
|
||||
def test_known_positive_a_a_malformed_mandate_is_refused(tmp_path: Path) -> None:
|
||||
|
|
@ -386,7 +387,7 @@ def test_known_positive_a_a_malformed_mandate_is_refused(tmp_path: Path) -> None
|
|||
|
||||
def test_known_positive_d_a_mandate_routed_at_another_base_is_caught(tmp_path: Path) -> None:
|
||||
broken = json.loads(json.dumps(_GOOD_MANDATE))
|
||||
broken["approaches"][1]["bundle_id"] = "vegnormal-n100-2023"
|
||||
broken["approaches"][1]["bundle_id"] = "eksempel-prosesskatalog-2027"
|
||||
set_dir = _broken_set(tmp_path, mandate=broken, bundle=_GOOD_BUNDLE_TXT, fasit={})
|
||||
declared = read_bundle_txt(set_dir / "bundle.txt")
|
||||
ids = {block["bundle_id"] for block in declared}
|
||||
|
|
@ -394,11 +395,11 @@ def test_known_positive_d_a_mandate_routed_at_another_base_is_caught(tmp_path: P
|
|||
# The SAME two set relations arm (d) asserts, and the broken set must fail the first of them:
|
||||
# an approach routed at a base the set does not declare.
|
||||
assert not routed <= ids
|
||||
assert sorted(routed - ids) == ["vegnormal-n100-2023"]
|
||||
assert sorted(routed - ids) == ["eksempel-prosesskatalog-2027"]
|
||||
|
||||
|
||||
def test_known_positive_c_an_anchor_the_base_carries_is_reported() -> None:
|
||||
concepts = [("a.md", {"type": "Krav"}, "en tunnel med ventilasjon og belysning")]
|
||||
concepts = [("a.md", {"type": "Krav"}, "et serverrom med ventilasjon og belysning")]
|
||||
assert anchors_are_absent(["enhetspris"], concepts) == []
|
||||
assert anchors_are_absent(["ventilasjon"], concepts) == ["ventilasjon"]
|
||||
|
||||
|
|
@ -411,7 +412,7 @@ def test_known_positive_c_an_empty_scan_is_refused_never_vacuously_absent() -> N
|
|||
def test_known_positive_c_an_unusable_anchor_is_refused() -> None:
|
||||
concepts = [("a.md", {"type": "Krav"}, "tekst")]
|
||||
with pytest.raises(ValueError, match="at least"):
|
||||
anchors_are_absent(["vei"], concepts)
|
||||
anchors_are_absent(["rom"], concepts)
|
||||
with pytest.raises(ValueError, match="lowercase"):
|
||||
anchors_are_absent(["Enhetspris"], concepts)
|
||||
with pytest.raises(ValueError, match="no anchors"):
|
||||
|
|
@ -432,7 +433,7 @@ def test_known_positive_b_a_fasit_path_the_base_does_not_carry_is_caught(tmp_pat
|
|||
|
||||
def test_known_positive_bundle_txt_must_declare_both_keys(tmp_path: Path) -> None:
|
||||
path = tmp_path / "bundle.txt"
|
||||
path.write_text("name: n500-2024\n", encoding="utf-8")
|
||||
path.write_text("name: driftskrav-2027\n", encoding="utf-8")
|
||||
with pytest.raises(ValueError, match="bundle_id"):
|
||||
read_bundle_txt(path)
|
||||
|
||||
|
|
@ -446,7 +447,8 @@ def test_known_positive_a_second_block_needs_its_own_bundle_id(tmp_path: Path) -
|
|||
"""
|
||||
path = tmp_path / "bundle.txt"
|
||||
path.write_text(
|
||||
"name: n200-2024\nbundle_id: vegnormal-n200-2024\nname: r761-2025\n", encoding="utf-8"
|
||||
"name: driftskrav-2027\nbundle_id: eksempel-driftskrav-2027\nname: prosesskatalog-2027\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
with pytest.raises(ValueError, match="bundle_id"):
|
||||
read_bundle_txt(path)
|
||||
|
|
@ -455,13 +457,13 @@ def test_known_positive_a_second_block_needs_its_own_bundle_id(tmp_path: Path) -
|
|||
def test_a_multi_base_bundle_txt_parses_into_one_block_per_base(tmp_path: Path) -> None:
|
||||
path = tmp_path / "bundle.txt"
|
||||
path.write_text(
|
||||
"name: n200-2024\nbundle_id: vegnormal-n200-2024\n"
|
||||
"name: r761-2025\nbundle_id: vegnormal-r761-2025\n",
|
||||
"name: driftskrav-2027\nbundle_id: eksempel-driftskrav-2027\n"
|
||||
"name: prosesskatalog-2027\nbundle_id: eksempel-prosesskatalog-2027\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
assert read_bundle_txt(path) == (
|
||||
{"name": "n200-2024", "bundle_id": "vegnormal-n200-2024"},
|
||||
{"name": "r761-2025", "bundle_id": "vegnormal-r761-2025"},
|
||||
{"name": "driftskrav-2027", "bundle_id": "eksempel-driftskrav-2027"},
|
||||
{"name": "prosesskatalog-2027", "bundle_id": "eksempel-prosesskatalog-2027"},
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -499,12 +501,12 @@ def test_the_fasit_titles_are_distinct_not_the_collapsed_sources_title(set_dir:
|
|||
# (f) + (g) + (h): the set is ANCHORED (P21 B2).
|
||||
#
|
||||
# The measured reason these exist. Four paid rounds ran entirely un-anchored, because the only
|
||||
# file loader reads ``cost-baseline.json`` out of the BUNDLE and no road normal carries prices — a
|
||||
# vegnormal is knowledge, the price belongs to the PROJECT. With ``--cost-baseline`` the project
|
||||
# supplies its own schedule, so the validator's stage 0 judges again: (f) every answerable approach
|
||||
# has a line to reconcile against, and (g) the falsification arm has NONE, so the code it proposes
|
||||
# is refused as "not in the project's cost baseline" — by stage 0, the one stage that can tell an
|
||||
# invented line from a real one, instead of by the weaker downstream gates.
|
||||
# file loader reads ``cost-baseline.json`` out of the BUNDLE and no requirements base carries
|
||||
# prices — a requirement is knowledge, the price belongs to the PROJECT. With ``--cost-baseline``
|
||||
# the project supplies its own schedule, so the validator's stage 0 judges again: (f) every
|
||||
# answerable approach has a line to reconcile against, and (g) the falsification arm has NONE, so
|
||||
# the code it proposes is refused as "not in the project's cost baseline" — by stage 0, the one
|
||||
# stage that can tell an invented line from a real one, instead of by the weaker downstream gates.
|
||||
#
|
||||
# (f) and (g) are SEPARATE arms rather than one loop over all approaches, because they are opposite
|
||||
# claims about opposite rows: a single arm asserting "exactly the non-refuse codes are present"
|
||||
|
|
@ -580,14 +582,14 @@ def test_h_no_cost_line_smuggles_in_an_uncommissioned_requirement_number(set_dir
|
|||
"""No line of the schedule is a reference number the base declares AND nobody commissions.
|
||||
|
||||
**The order words this arm as "no baseline code is a requirement number the base declares", and
|
||||
that rule was FELLED BY MEASUREMENT before anything was built on it.** Measured 15.09 against
|
||||
``okf.declared_reference_numbers`` over the four mounted bases: the four project-coded sets
|
||||
carry 0 such codes, and ``kontrakt-sorasen-2027`` carries FIVE of five — ``12.1``, ``12.12``,
|
||||
``22.1``, ``52.11``, ``51.1`` are real R761 ``prosessnr``. That is not an accident in the set;
|
||||
it is what R761 Prosesskoden IS. A Norwegian road contract's bill of quantities is priced BY
|
||||
process code, so the project's schedule and the corpus's vocabulary share an identifier
|
||||
namespace by design — and the order's rule would have forced a rewrite of the ONE set P20's
|
||||
decision (e) was chosen to preserve.
|
||||
that rule was FELLED BY MEASUREMENT before anything was built on it.** Measured against
|
||||
``okf.declared_reference_numbers``: the project-coded sets carry 0 such codes, and the set
|
||||
priced by process number (today ``driftsavtale-2027``) carries FIVE of five — ``12.1``,
|
||||
``12.12``, ``22.1``, ``52.11``, ``51.1`` are declared ``prosessnr`` of its catalogue. That is
|
||||
not an accident in the set; it is what a process catalogue IS. An agreement settled process by
|
||||
process is priced BY process number, so the project's schedule and the corpus's vocabulary
|
||||
share an identifier namespace by design — and the order's rule would have forced a rewrite of
|
||||
the ONE kind of set P20's decision (e) was chosen to preserve.
|
||||
|
||||
The COMPLEMENT keeps both: a schedule may price what the commission names, and may not
|
||||
INTRODUCE a corpus identifier as a cost line nobody ordered. The order's own mutation still
|
||||
|
|
@ -622,28 +624,27 @@ def test_h_no_cost_line_smuggles_in_an_uncommissioned_requirement_number(set_dir
|
|||
|
||||
|
||||
def test_known_positive_f_a_missing_cost_line_is_caught(tmp_path: Path) -> None:
|
||||
baseline = _set_baseline(_CONTEXT_ROOT / "gate-nordvik-2027")
|
||||
assert "GATE-KRYSS-01" in baseline.items
|
||||
baseline = _set_baseline(_CONTEXT_ROOT / "serverrom-2027")
|
||||
assert "SRV-KJOL-01" in baseline.items
|
||||
stripped = CostBaseline(
|
||||
project_id=baseline.project_id,
|
||||
items={k: v for k, v in baseline.items.items() if k != "GATE-KRYSS-01"},
|
||||
items={k: v for k, v in baseline.items.items() if k != "SRV-KJOL-01"},
|
||||
)
|
||||
assert "GATE-KRYSS-01" not in stripped.items
|
||||
assert "SRV-KJOL-01" not in stripped.items
|
||||
|
||||
|
||||
def test_known_positive_g_a_line_for_the_falsification_arm_is_caught() -> None:
|
||||
"""The order's mutation (i): give a4 a line, and (g)'s assertion must fail on this set."""
|
||||
baseline = _set_baseline(_CONTEXT_ROOT / "gate-nordvik-2027")
|
||||
baseline = _set_baseline(_CONTEXT_ROOT / "serverrom-2027")
|
||||
priced = dict(baseline.items)
|
||||
priced["GATE-GANG-ENHET"] = CostBaselineLine(quantity=6, unit_cost=50_000.0)
|
||||
fasit = json.loads((_CONTEXT_ROOT / "gate-nordvik-2027" / "fasit.json").read_text("utf-8"))
|
||||
priced["SRV-RACK-ENHET"] = CostBaselineLine(quantity=6, unit_cost=50_000.0)
|
||||
fasit = json.loads((_CONTEXT_ROOT / "serverrom-2027" / "fasit.json").read_text("utf-8"))
|
||||
by_id = {
|
||||
a.id: a
|
||||
for a in load_mandate(_CONTEXT_ROOT / "gate-nordvik-2027" / "mandate.json").approaches
|
||||
a.id: a for a in load_mandate(_CONTEXT_ROOT / "serverrom-2027" / "mandate.json").approaches
|
||||
}
|
||||
for row in fasit["must_refuse"]:
|
||||
carried = [c for c in by_id[row["approach_id"]].affected_codes if c in priced]
|
||||
assert carried == ["GATE-GANG-ENHET"]
|
||||
assert carried == ["SRV-RACK-ENHET"]
|
||||
|
||||
|
||||
def test_known_positive_h_an_uncommissioned_requirement_number_is_caught() -> None:
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@ must refuse rather than invent (MAJOR-4, misjonsreview v2 §7; owner B = the con
|
|||
2026-09-02).
|
||||
|
||||
``ir.CostBaseline``'s docstring names two projections INTO it — a bundle's hand-written
|
||||
``cost-baseline.json`` (``okf.load_cost_baseline``) and the road reference domain's ``cost_items``
|
||||
``cost-baseline.json`` (``okf.load_cost_baseline``) and the reference domain's ``cost_items``
|
||||
(``validator.baseline_from_project``). Neither can be produced from an ingested tender corpus, so
|
||||
K2 could be navigated but never RUN: ``run_project``'s bundle arm resolves an OPTIONAL baseline and
|
||||
a bundle carrying only ingested documents is simply un-anchored. This adds the third: derive the
|
||||
|
|
@ -268,10 +268,12 @@ def test_the_verdict_layer_is_not_scanned(tmp_path: Path) -> None:
|
|||
def test_control_a_bundle_shipping_a_handwritten_baseline_is_untouched() -> None:
|
||||
"""``load_optional_cost_baseline`` still answers, unchanged, and remains what the run path
|
||||
reaches for when the flag is absent."""
|
||||
shipped = okf.load_optional_cost_baseline("shared/examples/tunnel-hauglia")
|
||||
shipped = okf.load_optional_cost_baseline(
|
||||
"src/portfolio_optimiser/data/bundles/driftssenter-kjoling"
|
||||
)
|
||||
assert shipped is not None
|
||||
assert shipped.project_id == "TUNNEL-HAUGLIA"
|
||||
assert "ENERGI-TUNNEL-EL" in shipped.items
|
||||
assert shipped.project_id == "DRIFTSSENTER-KJOLING"
|
||||
assert "ENERGI-DRIFTSSENTER-EL" in shipped.items
|
||||
|
||||
|
||||
def test_control_the_fixture_bundles_ship_no_cost_baseline_json() -> None:
|
||||
|
|
|
|||
|
|
@ -1,13 +1,13 @@
|
|||
"""P21 DEL A - the PROJECT carries the price, so a run against a road normal can be anchored.
|
||||
"""P21 DEL A - the PROJECT carries the price, so a run on a requirements base can be anchored.
|
||||
|
||||
**The measurement this closes.** Four paid stress rounds (P16/P18/P19/P17b/P20) ran ENTIRELY
|
||||
un-anchored. The cause is one line: the only file loader reads ``cost-baseline.json`` out of the
|
||||
BUNDLE directory (``okf.load_optional_cost_baseline``), and no vegnormal ships one — N100, N200,
|
||||
N500 and R761 are knowledge, and knowledge carries requirements, never amounts. The validator's
|
||||
stage 0 — the one stage that can tell an invented cost line from a line this project actually buys
|
||||
— was therefore skipped in every single one, and ``validated`` could not mean what it says: P20
|
||||
G1/G2 measured real R761 process numbers (``12.11`` three times on Søråsen, ``1.1.1`` on Lindås)
|
||||
validating with amounts nobody had anywhere.
|
||||
BUNDLE directory (``okf.load_optional_cost_baseline``), and no requirements base or process
|
||||
catalogue ships one — they are knowledge, and knowledge carries requirements, never amounts. The
|
||||
validator's stage 0 — the one stage that can tell an invented cost line from a line this project
|
||||
actually buys — was therefore skipped in every single one, and ``validated`` could not mean what it
|
||||
says: P20 G1/G2 measured real process-catalogue numbers (``12.11`` three times on one set,
|
||||
``1.1.1`` on another) validating with amounts nobody had anywhere.
|
||||
|
||||
``--cost-baseline FILE`` is PM decision (e), taken over three alternatives P20 wrote down: (a)
|
||||
refusing every requirement-shaped code un-anchored would make the one realistic context set
|
||||
|
|
@ -220,8 +220,9 @@ def test_the_bundle_loader_still_parses_through_the_same_seam(tmp_path: Path) ->
|
|||
|
||||
|
||||
def test_cli_requires_a_knowledge_base(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
|
||||
"""(h) On the road path the baseline IS ``Project.cost_items``, so a file there is a second
|
||||
source for one fact with nothing to break the tie (``--require-cost-baseline``'s reason)."""
|
||||
"""(h) On the reference-project path the baseline IS ``Project.cost_items``, so a file there is
|
||||
a second source for one fact with nothing to break the tie (``--require-cost-baseline``'s
|
||||
reason)."""
|
||||
rc = run.main(["P1", "--docs-dir", "docs", "--cost-baseline", _schedule(tmp_path, RIGG=1000.0)])
|
||||
|
||||
assert rc == 1
|
||||
|
|
@ -313,7 +314,7 @@ def test_cli_wiring_anchors_the_dry_run_and_says_so(
|
|||
assert "Cost baseline: NONE in the bundle" in before
|
||||
assert "Cost baseline: 2 lines from" not in before
|
||||
|
||||
schedule = _schedule(tmp_path, RIGG=1000.0, ASFALT=250.0)
|
||||
schedule = _schedule(tmp_path, RIGG=1000.0, LISENS=250.0)
|
||||
assert run.main([*argv, "--cost-baseline", schedule]) == 0
|
||||
after = capsys.readouterr().out
|
||||
|
||||
|
|
|
|||
|
|
@ -22,15 +22,15 @@ from portfolio_optimiser.provenance import Citation
|
|||
def docs(tmp_path):
|
||||
d = tmp_path / "docs"
|
||||
d.mkdir()
|
||||
(d / "asphalt.txt").write_text(
|
||||
"Asphalt Ab11 unit rate renegotiation reduced the paving cost on the school stretch.",
|
||||
(d / "licence.txt").write_text(
|
||||
"Licence unit price renegotiation reduced the office-suite cost for the head office.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return d
|
||||
|
||||
|
||||
def test_in_process_tool_returns_citation_ready_chunks(docs) -> None:
|
||||
chunks = retrieve_chunks("asphalt paving cost", str(docs), top_k=3)
|
||||
chunks = retrieve_chunks("licence office-suite cost", str(docs), top_k=3)
|
||||
assert chunks # at least one citation-ready chunk
|
||||
for d in chunks:
|
||||
assert set(d) >= {"file", "locator", "snippet", "score"}
|
||||
|
|
@ -47,7 +47,7 @@ def test_make_retrieval_tool_builds_ga_function_tool(docs) -> None:
|
|||
|
||||
|
||||
async def test_mcp_wrapper_returns_same_structuredcontent_shape(docs) -> None:
|
||||
query = "asphalt paving cost"
|
||||
query = "licence office-suite cost"
|
||||
expected = retrieve_chunks(query, str(docs), top_k=3)
|
||||
server = build_mcp_server(str(docs), top_k=3)
|
||||
_content, structured = await server.call_tool("retrieve_cost_docs", {"query": query})
|
||||
|
|
|
|||
|
|
@ -288,22 +288,23 @@ async def test_a_debate_that_opened_nothing_says_so(tmp_path) -> None:
|
|||
_ENERGY_DIM = Dimension(
|
||||
id="energi", label="Energi", allowed_measure_types=frozenset({"energy_efficiency"})
|
||||
)
|
||||
_ASFALT_SENTINEL = "ASFALT-LEAK-SENTINEL-x7y8z9"
|
||||
_ASFALT_FILE = "asfalt-dekke.md"
|
||||
_LISENS_SENTINEL = "LISENS-LEAK-SENTINEL-x7y8z9"
|
||||
_LISENS_FILE = "lisens-kontorpakke.md"
|
||||
|
||||
|
||||
def _bundle_with_a_foreign_dimension(tmp_path: Path) -> str:
|
||||
"""A copy of the fixture base plus ONE concept file marked ``dimension: asfalt``, linked from
|
||||
"""A copy of the fixture base plus ONE concept file marked ``dimension: lisens``, linked from
|
||||
the index so navigation reaches it."""
|
||||
copy = tmp_path / "bundle"
|
||||
shutil.copytree(BUNDLE_DIR, copy)
|
||||
(copy / _ASFALT_FILE).write_text(
|
||||
f"---\ntype: reference\ntitle: Asfaltdekke\ndimension: asfalt\n---\n\n{_ASFALT_SENTINEL}\n",
|
||||
(copy / _LISENS_FILE).write_text(
|
||||
f"---\ntype: reference\ntitle: Lisenser\ndimension: lisens\n---\n\n{_LISENS_SENTINEL}\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
index = copy / "index.md"
|
||||
index.write_text(
|
||||
index.read_text(encoding="utf-8") + f"\n- [Asfaltdekke]({_ASFALT_FILE})\n", encoding="utf-8"
|
||||
index.read_text(encoding="utf-8") + f"\n- [Lisenser]({_LISENS_FILE})\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return str(copy)
|
||||
|
||||
|
|
@ -332,14 +333,14 @@ async def test_a_foreign_dimension_document_is_neither_listed_nor_readable(tmp_p
|
|||
tools = _tools(bundle_dir, "energi")
|
||||
|
||||
listing = await _invoke(tools["read_bundle"], bundle_id=bundle_id)
|
||||
assert _ASFALT_FILE not in listing, (
|
||||
assert _LISENS_FILE not in listing, (
|
||||
"a document from another dimension is still listed to the agents"
|
||||
)
|
||||
answer = await _invoke(tools["read_file"], bundle_id=bundle_id, path=_ASFALT_FILE)
|
||||
answer = await _invoke(tools["read_file"], bundle_id=bundle_id, path=_LISENS_FILE)
|
||||
assert answer.startswith(f"REFUSED ({DimensionScopeRefused.__name__})")
|
||||
# F99-D3 returns the refusal, so §4.1a's property is asserted on the value: the reason
|
||||
# travels, the out-of-scope document's bytes do not.
|
||||
assert _ASFALT_SENTINEL not in answer
|
||||
assert _LISENS_SENTINEL not in answer
|
||||
|
||||
|
||||
async def test_without_a_dimension_the_same_document_is_listed_and_readable(tmp_path) -> None:
|
||||
|
|
@ -350,9 +351,9 @@ async def test_without_a_dimension_the_same_document_is_listed_and_readable(tmp_
|
|||
tools = _tools(bundle_dir, None)
|
||||
|
||||
listing = await _invoke(tools["read_bundle"], bundle_id=bundle_id)
|
||||
assert _ASFALT_FILE in listing
|
||||
assert _ASFALT_SENTINEL in await _invoke(
|
||||
tools["read_file"], bundle_id=bundle_id, path=_ASFALT_FILE
|
||||
assert _LISENS_FILE in listing
|
||||
assert _LISENS_SENTINEL in await _invoke(
|
||||
tools["read_file"], bundle_id=bundle_id, path=_LISENS_FILE
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -385,6 +386,6 @@ async def test_a_dimension_scoped_run_gives_the_debate_scoped_tools(tmp_path, mo
|
|||
)
|
||||
|
||||
read_file = next(t for t in captured[0] if getattr(t, "name", "") == "read_file")
|
||||
answer = await _invoke(read_file, bundle_id=bundle_id, path=_ASFALT_FILE)
|
||||
answer = await _invoke(read_file, bundle_id=bundle_id, path=_LISENS_FILE)
|
||||
assert answer.startswith(f"REFUSED ({DimensionScopeRefused.__name__})")
|
||||
assert _ASFALT_SENTINEL not in answer
|
||||
assert _LISENS_SENTINEL not in answer
|
||||
|
|
|
|||
|
|
@ -9,10 +9,10 @@ ONE document.
|
|||
The order offered two rules and asked which discriminates. Replayed against the real listings:
|
||||
|
||||
* **"the declared document must have been returned by a ``read_dir`` filtered on a word from the
|
||||
approach's label"** refuses **13 of 13** — including Søråsen's ``12.11``, which the order names
|
||||
as the closest any run came. Measured, ZERO of the 13 declarations were reached through a
|
||||
filtered listing at all. A gate that refuses every measured case, right and wrong alike, cannot
|
||||
discriminate; it is the vacuous gate's mirror image.
|
||||
approach's label"** refuses **13 of 13** — including the process-catalogue run's ``12.11``,
|
||||
which the order names as the closest any run came. Measured, ZERO of the 13 declarations were
|
||||
reached through a filtered listing at all. A gate that refuses every measured case, right and
|
||||
wrong alike, cannot discriminate; it is the vacuous gate's mirror image.
|
||||
* **"fewer than k distinct documents opened"** at k=3 refuses **8 of 13** and keeps the five that
|
||||
navigated, ``12.11`` among them. k=3, 4 and 5 refuse the SAME eight — the distribution has a gap
|
||||
between 2 and 5 — so the threshold is not on a cliff, and 3 is the lowest of that plateau.
|
||||
|
|
@ -21,9 +21,10 @@ The second is built. It is capped by the base's own document count, so a two-doc
|
|||
declarable rather than becoming a level nobody can declare on.
|
||||
|
||||
**C2.** Over the same six traces, 18 of 143 path-bearing tool calls named a path the base does not
|
||||
hold, and ELEVEN of them are one run walking ``R761/4-3``, ``4.3``, ``4-2``, ``4-1``, ``4-0``,
|
||||
hold, and ELEVEN of them are one run walking ``P900/4-3``, ``4.3``, ``4-2``, ``4-1``, ``4-0``,
|
||||
``4-5``, ``4-6`` — guessing a chapter-number spelling the corpus does not use, while the real names
|
||||
are ``R761/4``, ``R761/41``, ``R761/42``. The refusal already named the nearest LISTABLE ancestor,
|
||||
are ``P900/4``, ``P900/41``, ``P900/42`` (the process catalogue the traces were taken on, shown here
|
||||
under the example set's ``P900`` name). The refusal already named the nearest LISTABLE ancestor,
|
||||
which is the right rung; what it could not say is which of that rung's names was meant. Measured
|
||||
after: **16 of 18** refusals now name at least one neighbour, and the two that do not are paths
|
||||
whose ancestor holds documents and no subdirectories — an empty list is omitted rather than
|
||||
|
|
@ -40,18 +41,18 @@ import pytest
|
|||
from portfolio_optimiser import okf
|
||||
from portfolio_optimiser.explore import ToolCall, navigator_tools
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
_KJOLING = _BUNDLES / "driftssenter-kjoling"
|
||||
|
||||
|
||||
def _wired(bundle_dir: Path = _TUNNEL) -> tuple[dict[str, Any], list[ToolCall], list[Any]]:
|
||||
def _wired(bundle_dir: Path = _KJOLING) -> tuple[dict[str, Any], list[ToolCall], list[Any]]:
|
||||
opened: list[ToolCall] = []
|
||||
declared: list[Any] = []
|
||||
tools = navigator_tools((str(bundle_dir),), opened=opened, requirements=declared)
|
||||
return {t.name: t for t in tools}, opened, declared
|
||||
|
||||
|
||||
def _concepts(bundle_dir: Path = _TUNNEL) -> list[str]:
|
||||
def _concepts(bundle_dir: Path = _KJOLING) -> list[str]:
|
||||
return [f.name for f in okf.navigate_bundle(str(bundle_dir)).context_files]
|
||||
|
||||
|
||||
|
|
@ -67,10 +68,10 @@ def test_c1_a_declaration_after_one_document_is_refused_with_the_denominator() -
|
|||
"""
|
||||
tools, opened, declared = _wired()
|
||||
path = _concepts()[0]
|
||||
opened.append(ToolCall(name="read_file", bundle_id="tunnel-hauglia", path=path))
|
||||
opened.append(ToolCall(name="read_file", bundle_id="driftssenter-kjoling", path=path))
|
||||
|
||||
answer = tools["declare_requirement"].func(
|
||||
approach_id="a1", bundle_id="tunnel-hauglia", path=path, ref="Krav 1.1—1"
|
||||
approach_id="a1", bundle_id="driftssenter-kjoling", path=path, ref="Krav 1.1—1"
|
||||
)
|
||||
|
||||
assert answer["refusal"] == "RequirementNotRead"
|
||||
|
|
@ -90,10 +91,10 @@ def test_c1_the_same_declaration_after_three_documents_is_accepted() -> None:
|
|||
tools, opened, declared = _wired()
|
||||
files = _concepts()
|
||||
for name in files[:3]:
|
||||
opened.append(ToolCall(name="read_file", bundle_id="tunnel-hauglia", path=name))
|
||||
opened.append(ToolCall(name="read_file", bundle_id="driftssenter-kjoling", path=name))
|
||||
|
||||
answer = tools["declare_requirement"].func(
|
||||
approach_id="a1", bundle_id="tunnel-hauglia", path=files[0], ref="Krav 1.1—1"
|
||||
approach_id="a1", bundle_id="driftssenter-kjoling", path=files[0], ref="Krav 1.1—1"
|
||||
)
|
||||
|
||||
assert answer["declared"] is True
|
||||
|
|
@ -106,10 +107,10 @@ def test_c1_the_same_path_read_three_times_is_still_one_document() -> None:
|
|||
tools, opened, declared = _wired()
|
||||
path = _concepts()[0]
|
||||
for _ in range(3):
|
||||
opened.append(ToolCall(name="read_file", bundle_id="tunnel-hauglia", path=path))
|
||||
opened.append(ToolCall(name="read_file", bundle_id="driftssenter-kjoling", path=path))
|
||||
|
||||
answer = tools["declare_requirement"].func(
|
||||
approach_id="a1", bundle_id="tunnel-hauglia", path=path, ref="Krav 1.1—1"
|
||||
approach_id="a1", bundle_id="driftssenter-kjoling", path=path, ref="Krav 1.1—1"
|
||||
)
|
||||
|
||||
assert answer["refusal"] == "RequirementNotRead"
|
||||
|
|
@ -150,10 +151,10 @@ def test_c1_the_never_opened_refusal_still_fires_first() -> None:
|
|||
tools, opened, declared = _wired()
|
||||
files = _concepts()
|
||||
for name in files[:3]:
|
||||
opened.append(ToolCall(name="read_file", bundle_id="tunnel-hauglia", path=name))
|
||||
opened.append(ToolCall(name="read_file", bundle_id="driftssenter-kjoling", path=name))
|
||||
|
||||
answer = tools["declare_requirement"].func(
|
||||
approach_id="a1", bundle_id="tunnel-hauglia", path=files[4], ref="Krav 1.1—1"
|
||||
approach_id="a1", bundle_id="driftssenter-kjoling", path=files[4], ref="Krav 1.1—1"
|
||||
)
|
||||
|
||||
assert answer["refusal"] == "RequirementNotRead"
|
||||
|
|
@ -164,21 +165,22 @@ def test_c1_the_never_opened_refusal_still_fires_first() -> None:
|
|||
# --- C2 -------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _r761_shaped(root: Path) -> Path:
|
||||
"""A base shaped like the delivered R761: numbered chapter directories under one top level.
|
||||
def _p900_shaped(root: Path) -> Path:
|
||||
"""A base shaped like a process catalogue: numbered chapter directories under one top level.
|
||||
|
||||
Crafted rather than mounted, for MAJOR-3's reason — R761 lives outside this repository and
|
||||
cannot be a test dependency — and shaped from the MEASURED names (``R761/4``, ``R761/41``,
|
||||
``R761/42``) so the arm tests the ranking the real corpus provoked.
|
||||
Crafted rather than mounted, for MAJOR-3's reason — the corpus the traces were taken on lived
|
||||
outside this repository and could not be a test dependency — and shaped from the MEASURED
|
||||
names (``P900/4``, ``P900/41``, ``P900/42``) so the arm tests the ranking the real corpus
|
||||
provoked.
|
||||
"""
|
||||
base = root / "proc"
|
||||
(base / "R761" / "4").mkdir(parents=True)
|
||||
(base / "R761" / "41").mkdir()
|
||||
(base / "R761" / "42").mkdir()
|
||||
links = "".join(f"- [{d}](R761/{d}/k.md) — chapter {d}.\n" for d in ("4", "41", "42"))
|
||||
(base / "P900" / "4").mkdir(parents=True)
|
||||
(base / "P900" / "41").mkdir()
|
||||
(base / "P900" / "42").mkdir()
|
||||
links = "".join(f"- [{d}](P900/{d}/k.md) — chapter {d}.\n" for d in ("4", "41", "42"))
|
||||
(base / "index.md").write_text(f"---\nbundle_id: proc\n---\n\n{links}", encoding="utf-8")
|
||||
for d in ("4", "41", "42"):
|
||||
(base / "R761" / d / "k.md").write_text(
|
||||
(base / "P900" / d / "k.md").write_text(
|
||||
f"---\ntype: Krav\ntitle: Chapter {d}\nprosessnr: '{d}'\n---\n\nbody {d}\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
|
@ -186,54 +188,54 @@ def _r761_shaped(root: Path) -> Path:
|
|||
|
||||
|
||||
def test_c2_a_guessed_chapter_spelling_is_answered_with_the_real_names(tmp_path: Path) -> None:
|
||||
"""The order's known-positive: ``R761/4-3`` must come back with ``R761/4…`` candidates.
|
||||
"""The order's known-positive: ``P900/4-3`` must come back with ``P900/4…`` candidates.
|
||||
|
||||
``R761/4`` is FIRST, because the ranking is longest-common-prefix with the segment that failed
|
||||
``P900/4`` is FIRST, because the ranking is longest-common-prefix with the segment that failed
|
||||
— which is the one thing that distinguishes "the names at this level" from "the name you were
|
||||
reaching for".
|
||||
"""
|
||||
base = _r761_shaped(tmp_path)
|
||||
base = _p900_shaped(tmp_path)
|
||||
bundle = okf.navigate_bundle(str(base))
|
||||
|
||||
neighbours = okf.nearest_subdirectories(bundle, "R761/4-3")
|
||||
neighbours = okf.nearest_subdirectories(bundle, "P900/4-3")
|
||||
|
||||
assert neighbours[0] == "R761/4"
|
||||
assert set(neighbours) == {"R761/4", "R761/41", "R761/42"}
|
||||
assert neighbours[0] == "P900/4"
|
||||
assert set(neighbours) == {"P900/4", "P900/41", "P900/42"}
|
||||
|
||||
|
||||
def test_c2_the_read_file_refusal_carries_them(tmp_path: Path) -> None:
|
||||
"""The seam: the message the MODEL sees. A helper nothing calls would leave every arm above
|
||||
green and the measured defect untouched."""
|
||||
base = _r761_shaped(tmp_path)
|
||||
base = _p900_shaped(tmp_path)
|
||||
tools, _opened, _declared = _wired(base)
|
||||
|
||||
answer = tools["read_file"].func(bundle_id="proc", path="R761/4-3/k.md")
|
||||
answer = tools["read_file"].func(bundle_id="proc", path="P900/4-3/k.md")
|
||||
|
||||
assert answer.startswith("REFUSED (BundlePathNotFound)")
|
||||
assert "'R761'" in answer
|
||||
assert "subdirectories include R761/4, R761/41, R761/42" in answer
|
||||
assert "'P900'" in answer
|
||||
assert "subdirectories include P900/4, P900/41, P900/42" in answer
|
||||
|
||||
|
||||
def test_c2_the_read_dir_refusal_carries_them_too(tmp_path: Path) -> None:
|
||||
"""The OTHER rung, and the reason both were changed: a caller that guesses a directory gets the
|
||||
same help as one that guesses a document. One question about one base, one answer."""
|
||||
base = _r761_shaped(tmp_path)
|
||||
base = _p900_shaped(tmp_path)
|
||||
tools, _opened, _declared = _wired(base)
|
||||
|
||||
answer = tools["read_dir"].func(bundle_id="proc", path="R761/4-3")
|
||||
answer = tools["read_dir"].func(bundle_id="proc", path="P900/4-3")
|
||||
|
||||
assert answer["refusal"] == "BundlePathNotFound"
|
||||
assert "R761/4, R761/41, R761/42" in answer["refused"]
|
||||
assert "P900/4, P900/41, P900/42" in answer["refused"]
|
||||
|
||||
|
||||
def test_c2_an_ancestor_with_no_subdirectories_names_none(tmp_path: Path) -> None:
|
||||
"""Omission, never an empty clause. MEASURED: 2 of the 18 guessed paths land here — their
|
||||
ancestor holds documents and no directories — and a trailing "its subdirectories include "
|
||||
with nothing after it is a sentence that says nothing."""
|
||||
base = _r761_shaped(tmp_path)
|
||||
base = _p900_shaped(tmp_path)
|
||||
tools, _opened, _declared = _wired(base)
|
||||
|
||||
answer = tools["read_file"].func(bundle_id="proc", path="R761/4/missing.md")
|
||||
answer = tools["read_file"].func(bundle_id="proc", path="P900/4/missing.md")
|
||||
|
||||
assert answer.startswith("REFUSED (BundlePathNotFound)")
|
||||
assert "subdirectories include" not in answer
|
||||
|
|
@ -242,11 +244,11 @@ def test_c2_an_ancestor_with_no_subdirectories_names_none(tmp_path: Path) -> Non
|
|||
def test_c2_every_name_it_hands_back_resolves(tmp_path: Path) -> None:
|
||||
"""The property that makes the help worth having (``_index_excerpt``'s rule): a path that never
|
||||
was is worse than no path. Each suggestion is fed straight back to ``read_dir``."""
|
||||
base = _r761_shaped(tmp_path)
|
||||
base = _p900_shaped(tmp_path)
|
||||
tools, _opened, _declared = _wired(base)
|
||||
bundle = okf.navigate_bundle(str(base))
|
||||
|
||||
for name in okf.nearest_subdirectories(bundle, "R761/4-3"):
|
||||
for name in okf.nearest_subdirectories(bundle, "P900/4-3"):
|
||||
listing = tools["read_dir"].func(bundle_id="proc", path=name)
|
||||
assert "refused" not in listing, f"{name} does not resolve: {listing}"
|
||||
|
||||
|
|
@ -277,8 +279,9 @@ def test_c2_the_verdict_layer_is_never_advertised_by_name(tmp_path: Path) -> Non
|
|||
@pytest.mark.parametrize("path", ["", "a"])
|
||||
def test_c2_a_top_level_guess_names_the_top_level_directories(tmp_path: Path, path: str) -> None:
|
||||
"""The fallback: no ancestor at all means the base's own top level, which ``directory_listing``
|
||||
always answers. ``krav`` and ``r761`` were both measured as top-level guesses on R761."""
|
||||
base = _r761_shaped(tmp_path)
|
||||
always answers. ``krav`` and the catalogue's own lower-cased name were both measured as
|
||||
top-level guesses on the process catalogue."""
|
||||
base = _p900_shaped(tmp_path)
|
||||
bundle = okf.navigate_bundle(str(base))
|
||||
|
||||
assert okf.nearest_subdirectories(bundle, path) == ("R761",)
|
||||
assert okf.nearest_subdirectories(bundle, path) == ("P900",)
|
||||
|
|
|
|||
|
|
@ -37,7 +37,7 @@ def test_admits_in_dimension_measure_type_only() -> None:
|
|||
def test_admits_rejects_out_of_dimension_measure_type() -> None:
|
||||
"""Out-of-dimension measure_type is rejected regardless of codes."""
|
||||
dim = _energy_dim()
|
||||
assert admits(measure_type="paving", codes=frozenset({"07.1"}), dimension=dim) is False
|
||||
assert admits(measure_type="licence", codes=frozenset({"07.1"}), dimension=dim) is False
|
||||
|
||||
|
||||
def test_admits_prefix_branch_matching_code() -> None:
|
||||
|
|
|
|||
|
|
@ -44,9 +44,9 @@ _ENERGY_DIM = Dimension(
|
|||
)
|
||||
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (sim)"}
|
||||
|
||||
# A marker that appears ONLY in the asfalt-marked concept file's body, so it can reach the prompt
|
||||
# A marker that appears ONLY in the lisens-marked concept file's body, so it can reach the prompt
|
||||
# solely through un-filtered context — its presence/absence is the §4.1a leak probe.
|
||||
_ASFALT_SENTINEL = "ASFALT-LEAK-SENTINEL-x7y8z9"
|
||||
_LISENS_SENTINEL = "LISENS-LEAK-SENTINEL-x7y8z9"
|
||||
|
||||
|
||||
def _valid_reply(measure: str, code: str) -> str:
|
||||
|
|
@ -80,7 +80,7 @@ async def test_foreign_dimension_candidate_rejected_when_dimension_set() -> None
|
|||
that has nothing to do with the dimension, and this arm would have ridden on that instead of on
|
||||
``admits``. Only the MEASURE is foreign now, which is what §4.1b is about."""
|
||||
factory = _role_factory(
|
||||
_valid_reply("paving_renegotiation", "ENERGI-TOTAL-EL"), "VERDICT: APPROVE"
|
||||
_valid_reply("licence_renegotiation", "ENERGI-TOTAL-EL"), "VERDICT: APPROVE"
|
||||
)
|
||||
|
||||
result = await run_project(
|
||||
|
|
@ -125,19 +125,19 @@ async def test_in_dimension_candidate_passes() -> None:
|
|||
# --- Context scope (§4.1a) -----------------------------------------------------------------------
|
||||
|
||||
|
||||
def _bundle_with_asfalt_file(tmp_path: Path) -> str:
|
||||
"""A throwaway copy of the shared bundle with one extra asfalt-marked concept file carrying the
|
||||
def _bundle_with_lisens_file(tmp_path: Path) -> str:
|
||||
"""A throwaway copy of the shared bundle with one extra lisens-marked concept file carrying the
|
||||
sentinel in its body, linked from the index so ``navigate_bundle`` reaches it. The shared,
|
||||
framework-neutral fixture is never mutated (mirrors ``_copy_bundle``)."""
|
||||
dst = tmp_path / "bundle"
|
||||
shutil.copytree(BUNDLE_DIR, dst)
|
||||
(dst / "asfalt-note.md").write_text(
|
||||
f"---\ntype: methodology\ndimension: asfalt\n---\n\n{_ASFALT_SENTINEL} — paving method note\n",
|
||||
(dst / "lisens-note.md").write_text(
|
||||
f"---\ntype: methodology\ndimension: lisens\n---\n\n{_LISENS_SENTINEL} — licence method note\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
index = dst / "index.md"
|
||||
index.write_text(
|
||||
index.read_text(encoding="utf-8") + "\n- [asfalt](asfalt-note.md)\n", encoding="utf-8"
|
||||
index.read_text(encoding="utf-8") + "\n- [lisens](lisens-note.md)\n", encoding="utf-8"
|
||||
)
|
||||
return str(dst)
|
||||
|
||||
|
|
@ -194,20 +194,20 @@ async def test_dimension_scopes_the_agent_context(tmp_path) -> None:
|
|||
Asserts on the FILE NAME rather than the body sentinel, because a listing carries names and
|
||||
sizes, never bodies — an assert on the sentinel would be green against every implementation
|
||||
and would prove nothing (the vacuous form this repo keeps measuring)."""
|
||||
bundle_dir = _bundle_with_asfalt_file(tmp_path)
|
||||
bundle_dir = _bundle_with_lisens_file(tmp_path)
|
||||
|
||||
scoped = await _listing_the_debate_can_see(bundle_dir, _ENERGY_DIM)
|
||||
assert "asfalt-note.md" not in scoped, (
|
||||
assert "lisens-note.md" not in scoped, (
|
||||
"another dimension's document is listed to the agents — the tools are not dimension-scoped"
|
||||
)
|
||||
|
||||
|
||||
async def test_no_dimension_leaves_context_unscoped(tmp_path) -> None:
|
||||
"""CAUSALITY CONTROL: with ``dimension=None`` the asfalt document IS listed — proving its
|
||||
"""CAUSALITY CONTROL: with ``dimension=None`` the lisens document IS listed — proving its
|
||||
absence above is caused by the dimension scope, not by the file being unreachable."""
|
||||
bundle_dir = _bundle_with_asfalt_file(tmp_path)
|
||||
bundle_dir = _bundle_with_lisens_file(tmp_path)
|
||||
|
||||
unscoped = await _listing_the_debate_can_see(bundle_dir, None)
|
||||
assert "asfalt-note.md" in unscoped, (
|
||||
"the asfalt file is unreachable even without a filter — the control does not prove causality"
|
||||
assert "lisens-note.md" in unscoped, (
|
||||
"the lisens file is unreachable even without a filter — the control does not prove causality"
|
||||
)
|
||||
|
|
|
|||
|
|
@ -3,13 +3,13 @@
|
|||
P16 FUNN 2: the documented stress command names the same directory twice
|
||||
(``--docs-dir <base> --bundle-dir <base>``), because single-project mode demanded ``--docs-dir``
|
||||
even on the bundle path — where it is never read. Retrieval, the chunk tool and the "no citable
|
||||
content" check all live in the ROAD branch (``run.py``); the bundle branch builds its citations
|
||||
content" check all live in the REFERENCE branch (``run.py``); the bundle branch builds its citations
|
||||
from the navigated base. So the flag was required for a path that ignores it, and the published
|
||||
command had to satisfy the requirement by repeating itself.
|
||||
|
||||
**This is not the "--docs-dir omvei"** (feeding project documents through retrieval INSTEAD of
|
||||
ingesting them into a knowledge base), which STATE forbids and this order forbids again. No such
|
||||
path is opened: the road branch still refuses without a real ``--docs-dir``, and the value is
|
||||
path is opened: the reference branch still refuses without a real ``--docs-dir``, and the value is
|
||||
bound ONCE from ``--bundle-dir`` — byte-identically what the README already tells an operator to
|
||||
type by hand, so every existing invocation, the two-flag form included, is unchanged.
|
||||
"""
|
||||
|
|
@ -45,7 +45,7 @@ def _runnable(tmp_path: Path) -> str:
|
|||
def test_a_bundle_run_needs_no_docs_dir(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
|
||||
"""(a) The fix. A free dry run on ``--bundle-dir`` alone is ACCEPTED and reaches the bundle
|
||||
path — asserted on the run-config the dry run prints, not on rc alone, since rc 0 is also what
|
||||
a run that silently took the road path would return."""
|
||||
a run that silently took the reference path would return."""
|
||||
rc = run.main(["K2", "--bundle-dir", _runnable(tmp_path), "--live-dry-run"])
|
||||
|
||||
assert rc == 0
|
||||
|
|
@ -65,7 +65,7 @@ def test_the_documented_two_flag_form_still_works(
|
|||
|
||||
|
||||
def test_neither_flag_is_still_refused_by_name(capsys: pytest.CaptureFixture[str]) -> None:
|
||||
"""(c) The half that must NOT be relaxed: the road path has no base to fall back to, so an
|
||||
"""(c) The half that must NOT be relaxed: the reference path has no base to fall back to, so an
|
||||
argv naming neither is refused, and the refusal names BOTH doors rather than only the one it
|
||||
used to name."""
|
||||
rc = run.main(["K2", "--live-dry-run"])
|
||||
|
|
@ -79,14 +79,14 @@ def test_the_road_path_still_requires_a_real_docs_dir(
|
|||
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||
) -> None:
|
||||
"""(d) The anti-omvei arm. With no bundle, ``--docs-dir`` is still the only door AND it is
|
||||
still read: a directory holding nothing citable is refused by the road branch's own check, so
|
||||
still read: a directory holding nothing citable is refused by the reference branch's own check, so
|
||||
nothing here turns retrieval into a substitute for ingestion."""
|
||||
empty = tmp_path / "tomt"
|
||||
empty.mkdir()
|
||||
rc = run.main(["P1", "--docs-dir", str(empty), "--live-dry-run"])
|
||||
|
||||
assert rc == 1
|
||||
assert capsys.readouterr().err.strip(), "the road path must say why, not fail silently"
|
||||
assert capsys.readouterr().err.strip(), "the reference path must say why, not fail silently"
|
||||
|
||||
|
||||
def test_the_bundle_value_is_bound_once_and_reaches_run_project(
|
||||
|
|
|
|||
|
|
@ -11,8 +11,9 @@ refusal is built from).
|
|||
The reason is structural: the nearest listable ancestor of a guessed DOCUMENT path often holds
|
||||
documents and no subdirectories, and then the neighbour clause was omitted — deliberately, because
|
||||
an empty list is a sentence with nothing in it. Measured over those six misses, THREE land on such
|
||||
an ancestor (``krav/N100`` with 445 documents; ``R761/1`` with exactly ONE, which two separate
|
||||
guesses — ``R761/1/1-1.md`` and ``R761/1/R761-1-1_id-...md`` — were both reaching for) and three
|
||||
an ancestor (a requirements level with 445 documents; a process-catalogue chapter, ``P900/1``,
|
||||
with exactly ONE, which two separate guesses — ``P900/1/1-1.md`` and ``P900/1/P900-1-1_id-...md``
|
||||
— were both reaching for) and three
|
||||
have subdirectories and were already answered.
|
||||
|
||||
Same source and same property as its sibling: built from ``context_files`` through ``in_dimension``,
|
||||
|
|
@ -91,7 +92,7 @@ def test_a_missing_document_is_answered_with_the_names_that_level_holds(tmp_path
|
|||
|
||||
|
||||
def test_a_level_with_exactly_one_document_names_it(tmp_path: Path) -> None:
|
||||
"""The measured ``R761/1`` case, which is the sharpest one in the round-5 trace: TWO separate
|
||||
"""The measured ``P900/1`` case, which is the sharpest one in the round-5 trace: TWO separate
|
||||
guesses at one document's name, in one run, at a level that holds exactly that one document.
|
||||
Naming it answers both in a single step."""
|
||||
base = _write_tree(tmp_path / "korpus", dirs=1, per_dir=1)
|
||||
|
|
@ -136,9 +137,9 @@ def test_every_named_document_can_then_be_read(tmp_path: Path) -> None:
|
|||
|
||||
|
||||
def test_the_document_list_is_bounded(tmp_path: Path) -> None:
|
||||
"""A level of a delivered corpus can hold hundreds — ``krav/N100`` holds 445. The list is the
|
||||
same fixed five its sibling uses, for the same reason: this is help text on a refusal, and its
|
||||
price must not be set by how much the level contains."""
|
||||
"""A level of a delivered corpus can hold hundreds — one requirements level held 445. The list
|
||||
is the same fixed five its sibling uses, for the same reason: this is help text on a refusal,
|
||||
and its price must not be set by how much the level contains."""
|
||||
base = _write_tree(tmp_path / "korpus", dirs=1, per_dir=40)
|
||||
refusal = _read_file(base, "kategori-00/dok-99-gjettet.md")
|
||||
named = refusal.split("it holds the documents ")[1].split(")")[0].split(", ")
|
||||
|
|
@ -150,7 +151,7 @@ def test_the_closest_name_comes_first(tmp_path: Path) -> None:
|
|||
|
||||
Replacing ``_shared_prefix`` with a plain reverse sort left the whole suite green: the bound,
|
||||
the source and the resolve property were all gated, and the ORDER was not. For the measured
|
||||
``R761/1`` case that costs nothing — one document, one answer — but a level of a delivered
|
||||
``P900/1`` case that costs nothing — one document, one answer — but a level of a delivered
|
||||
corpus can hold 445, and then which five it names is the whole value of the clause.
|
||||
|
||||
The rule is the sibling's, through the SAME ``_shared_prefix`` helper rather than a second copy
|
||||
|
|
@ -160,7 +161,7 @@ def test_the_closest_name_comes_first(tmp_path: Path) -> None:
|
|||
"""
|
||||
base = tmp_path / "korpus"
|
||||
base.mkdir(parents=True)
|
||||
names = ["asfalt-slitelag-2024.md", "asfalt-dekke.md", "betong.md", "grus.md"]
|
||||
names = ["lagring-sikkerhetskopi-2024.md", "lagring-disk.md", "nettverk.md", "server.md"]
|
||||
lines = [f"# {base.name}", ""]
|
||||
for name in names:
|
||||
lines.append(f"- [{name}]({name})")
|
||||
|
|
@ -171,12 +172,12 @@ def test_the_closest_name_comes_first(tmp_path: Path) -> None:
|
|||
"---\ntype: index\n---\n\n" + "\n".join(lines) + "\n", encoding="utf-8"
|
||||
)
|
||||
|
||||
refusal = _read_file(base, "asfalt-slitelag.md")
|
||||
refusal = _read_file(base, "lagring-sikkerhetskopi.md")
|
||||
named = refusal.split("it holds the documents ")[1].split(")")[0].split(", ")
|
||||
# It shares the longest prefix AND is the longest name, so a length-only rule puts it last and
|
||||
# a reverse-alphabetical one puts ``grus.md`` first. Only the prefix rule puts it first.
|
||||
assert named[0] == "asfalt-slitelag-2024.md", named
|
||||
assert named[1] == "asfalt-dekke.md", named
|
||||
# a reverse-alphabetical one puts ``server.md`` first. Only the prefix rule puts it first.
|
||||
assert named[0] == "lagring-sikkerhetskopi-2024.md", named
|
||||
assert named[1] == "lagring-disk.md", named
|
||||
|
||||
|
||||
# --- (e) the other call site ---------------------------------------------------------------------
|
||||
|
|
@ -219,7 +220,7 @@ def test_a_foreign_dimension_document_is_never_named(tmp_path: Path) -> None:
|
|||
base = _write_tree(tmp_path / "korpus", dirs=1, per_dir=1)
|
||||
foreign = base / "kategori-00" / "fremmed.md"
|
||||
foreign.write_text(
|
||||
"---\ntype: concept\ndimension: asfalt\ntitle: Fremmed\n---\n\nkropp.\n", encoding="utf-8"
|
||||
"---\ntype: concept\ndimension: lisens\ntitle: Fremmed\n---\n\nkropp.\n", encoding="utf-8"
|
||||
)
|
||||
index = base / "kategori-00" / "index.md"
|
||||
index.write_text(
|
||||
|
|
|
|||
|
|
@ -40,8 +40,8 @@ from portfolio_optimiser import okf
|
|||
_TWO_SOURCES = (
|
||||
"---\n"
|
||||
"type: concept\n"
|
||||
"title: Tunnelbelysning\n"
|
||||
"sources: [{ id: n100, resource: vegnormal }, { id: ipmvp, resource: efficiency-valuation }]\n"
|
||||
"title: Serverromskjøling\n"
|
||||
"sources: [{ id: d100, resource: driftskrav }, { id: ipmvp, resource: efficiency-valuation }]\n"
|
||||
"verified: { by: human:kjell, at: 2026-09-02 }\n"
|
||||
"---\n"
|
||||
"\nInnhold.\n"
|
||||
|
|
@ -64,7 +64,7 @@ def test_a_non_verified_key_is_read_instead_of_raising(tmp_path: Path) -> None:
|
|||
|
||||
assert evidence.state == "present"
|
||||
assert evidence.items_seen == 2
|
||||
assert [e["id"] for e in evidence.entries] == ["n100", "ipmvp"]
|
||||
assert [e["id"] for e in evidence.entries] == ["d100", "ipmvp"]
|
||||
|
||||
|
||||
def test_a_non_verified_key_carries_no_tier(tmp_path: Path) -> None:
|
||||
|
|
@ -139,7 +139,7 @@ def test_the_three_non_present_states_are_unchanged_for_any_key(tmp_path: Path)
|
|||
# keeps asserting what it always asserted: an unreadable value is never tiered.
|
||||
blocked = _concept(
|
||||
tmp_path,
|
||||
"---\ntype: concept\nsources:\n - { id n100, resource vegnormal }\n---\n\nInnhold.\n",
|
||||
"---\ntype: concept\nsources:\n - { id d100, resource driftskrav }\n---\n\nInnhold.\n",
|
||||
)
|
||||
unreadable = okf.evidence_for(blocked, key="sources")
|
||||
assert unreadable.state == "unreadable"
|
||||
|
|
@ -149,7 +149,7 @@ def test_the_three_non_present_states_are_unchanged_for_any_key(tmp_path: Path)
|
|||
# The reverse direction, pinned in the same arm so the widening cannot regress silently.
|
||||
readable = _concept(
|
||||
tmp_path,
|
||||
"---\ntype: concept\nsources:\n - id: n100\n resource: vegnormal\n---\n\nInnhold.\n",
|
||||
"---\ntype: concept\nsources:\n - id: d100\n resource: driftskrav\n---\n\nInnhold.\n",
|
||||
)
|
||||
present = okf.evidence_for(readable, key="sources")
|
||||
assert present.state == "present"
|
||||
|
|
|
|||
|
|
@ -121,18 +121,18 @@ def test_the_dimension_refusal_carries_the_reason_and_never_the_content(tmp_path
|
|||
base = _marked_copy(tmp_path, _CONCEPT_FILE)
|
||||
foreign = base / "fremmed-dimensjon.md"
|
||||
foreign.write_text(
|
||||
f"---\ntype: methodology\ntitle: Asfalt\ndimension: asfalt\n---\n\n{_SENTINEL}\n",
|
||||
f"---\ntype: methodology\ntitle: Lisens\ndimension: lisens\n---\n\n{_SENTINEL}\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
index = base / "index.md"
|
||||
index.write_text(
|
||||
index.read_text(encoding="utf-8") + "\n- [Asfalt](fremmed-dimensjon.md)\n", encoding="utf-8"
|
||||
index.read_text(encoding="utf-8") + "\n- [Lisens](fremmed-dimensjon.md)\n", encoding="utf-8"
|
||||
)
|
||||
answer = _tools(base, dimension="tunnel")["read_file"].func(
|
||||
answer = _tools(base, dimension="kjoling")["read_file"].func(
|
||||
bundle_id=base.name, path="fremmed-dimensjon.md"
|
||||
)
|
||||
assert isinstance(answer, str) and answer.startswith("REFUSED")
|
||||
assert "tunnel" in answer
|
||||
assert "kjoling" in answer
|
||||
assert _SENTINEL not in answer, "the refusal leaked the out-of-scope document"
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -136,8 +136,8 @@ def test_the_worked_example_round_trips_through_the_real_readers(tmp_path: Path)
|
|||
# between this repo and the commons-owned artefact rather than a defect on either side.
|
||||
# The example declares `state: unreadable, reason: block-sequence, items_seen: 2` for a
|
||||
# SPEC §5.1 block sequence; po widened its reader to that carrier because all four
|
||||
# delivered knowledge bases write it and nothing else (measured 2026-09-12: n100 446/446,
|
||||
# n200 1133/1133, n500 270/270, r761 2756/2756 = 4605/4605, 0 in flow form). `shared/` is a
|
||||
# knowledge bases delivered at the time wrote it and nothing else (measured 2026-09-12:
|
||||
# 446/446, 1133/1133, 270/270, 2756/2756 = 4605/4605, 0 in flow form). `shared/` is a
|
||||
# PULL-ONLY subtree, so the declaration cannot be corrected from here: closing this needs a
|
||||
# commons amendment, and the divergence is asserted rather than skipped so it cannot sit
|
||||
# unnoticed until someone reads the prose.
|
||||
|
|
|
|||
|
|
@ -1,11 +1,11 @@
|
|||
"""The measurements read a FROZEN copy of the vegnormal bases, pinned by sha256 — never another
|
||||
"""The measurements read a FROZEN copy of the knowledge bases, pinned by sha256 — never another
|
||||
repository's live build directory.
|
||||
|
||||
Measured 2026-09-17 17:43: ``vegnormal-okf`` rebuilt ``build/ferdig/r761-2025`` while this
|
||||
repository's gate pointed straight at it. Rows 6-7 went "IKKE MÅLT" and five tests fell, for a
|
||||
change no one here made. The failure mode was never falsehood — the gate says IKKE MÅLT and exits
|
||||
non-zero, never green — it was that two projects shared a directory neither owns, so what this
|
||||
repository MEASURES could change without a commit here.
|
||||
Measured 2026-09-17: another project rebuilt the corpus directory this repository's gate pointed
|
||||
straight at. Rows 6-7 went "IKKE MÅLT" and five tests fell, for a change no one here made. The
|
||||
failure mode was never falsehood — the gate says IKKE MÅLT and exits non-zero, never green — it was
|
||||
that two projects shared a directory neither owns, so what this repository MEASURES could change
|
||||
without a commit here.
|
||||
|
||||
The fix is a copy outside both repositories plus a pin this repository tracks. The pin is the whole
|
||||
point: a copy with no pin is the same shared directory one move further away. So the three states
|
||||
|
|
@ -13,9 +13,10 @@ are separated by construction, and each has its own arm below:
|
|||
|
||||
* the copy matches the pin -> it resolves, and that is the only green path;
|
||||
* the copy is GONE -> ``FrozenBundleMissing`` (an ``OSError``): the gate says
|
||||
IKKE MÅLT and fails the exit code exactly as it did before this change; the delivered-corpus
|
||||
tests SKIP, which is MAJOR-3's ceiling rule (a hard error would break ``uv run pytest`` in the
|
||||
handover archive, where no corpus is mounted);
|
||||
IKKE MÅLT and fails the exit code exactly as it did before this change; the corpus tests SKIP.
|
||||
Since the bases became the package's own fictional examples (``data/kunnskapsbaser/``) the
|
||||
default store is always present, so this state is reached only through a user's own store
|
||||
(``PORTFOLIO_FROZEN_BUNDLES``) — and it is still pinned below, because that door stays open;
|
||||
* the copy DIFFERS from the pin -> ``FrozenBundleDrift`` (a ``ValueError``): loud, named, and
|
||||
NEVER a skip. A drifted copy is not an unreadable measurement, it is a measurement of the wrong
|
||||
corpus, which is the one thing that produces a silently wrong number.
|
||||
|
|
@ -42,18 +43,20 @@ _REPO = Path(__file__).resolve().parent.parent
|
|||
#: Joined at run time on purpose: this file NAMES the forbidden path spellings, and a literal
|
||||
#: would make the gate below red against its own source. (Implicit concatenation is not enough —
|
||||
#: ``ruff format`` folds ``"a" "b"`` back into one literal, measured here.)
|
||||
_OTHER_REPO = "-".join(("vegnormal", "okf"))
|
||||
#: The two spellings a path to that repository's build directory takes in this codebase. Both were
|
||||
#: present before this change (3 + 4 hits, the known positives recorded in the order's evidence).
|
||||
#: Each is paired with the line that PROVES it can match: a pattern that matches nothing makes a
|
||||
#: gate that can only be green, and the two spellings do not match each other's sample.
|
||||
_OTHER_REPO = "-".join(("another", "project"))
|
||||
#: The two spellings a path into ANOTHER repository's build directory takes. Both were present
|
||||
#: before this change (3 + 4 hits against the corpus project's build directory, the known
|
||||
#: positives recorded in the order's evidence); the patterns are now written for ANY sibling
|
||||
#: repository rather than for that one name, which is the rule the order stated. Each is paired
|
||||
#: with the line that PROVES it can match: a pattern that matches nothing makes a gate that can
|
||||
#: only be green, and the two spellings do not match each other's sample.
|
||||
_FORBIDDEN = (
|
||||
(
|
||||
re.compile(re.escape(_OTHER_REPO + "/build")),
|
||||
re.compile(r"~/repos/[\w.-]+/build"),
|
||||
f'ROOT = Path("~/repos/{_OTHER_REPO}/build/ferdig")',
|
||||
),
|
||||
(
|
||||
re.compile("[\"']" + re.escape(_OTHER_REPO) + "[\"']"),
|
||||
re.compile(r'home\(\) / "repos" / "[\w.-]+" / "build"'),
|
||||
'ROOT = Path.home() / "repos" / "' + _OTHER_REPO + '" / "build" / "ferdig"',
|
||||
),
|
||||
)
|
||||
|
|
@ -76,7 +79,7 @@ def _bundle(root: Path, body: str = "one") -> Path:
|
|||
return root
|
||||
|
||||
|
||||
def _store(tmp_path: Path, name: str = "n500-2024", body: str = "one") -> tuple[Path, Path]:
|
||||
def _store(tmp_path: Path, name: str = "driftskrav-2027", body: str = "one") -> tuple[Path, Path]:
|
||||
"""A frozen store holding ONE pinned bundle, and the pin file that names it."""
|
||||
store = tmp_path / "store"
|
||||
digest, files = fb.digest_bundle(_bundle(tmp_path / "src-of-truth", body))
|
||||
|
|
@ -134,18 +137,18 @@ def test_an_added_or_removed_file_is_drift_even_when_no_kept_byte_changes(
|
|||
only the SET of files moved, and the set is what a copy is."""
|
||||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
base = fb.bundle_dir("n500-2024") # control: green before either edit
|
||||
base = fb.bundle_dir("driftskrav-2027") # control: green before either edit
|
||||
|
||||
extra = base / "krav" / "b.md"
|
||||
extra.write_text("---\ntype: Krav\ntitle: B\n---\n", encoding="utf-8")
|
||||
with pytest.raises(fb.FrozenBundleDrift):
|
||||
fb.bundle_dir("n500-2024")
|
||||
fb.bundle_dir("driftskrav-2027")
|
||||
extra.unlink()
|
||||
assert fb.bundle_dir("n500-2024") == base # the edit, not the fixture, was the cause
|
||||
assert fb.bundle_dir("driftskrav-2027") == base # the edit, not the fixture, was the cause
|
||||
|
||||
(base / "krav" / "a.md").unlink()
|
||||
with pytest.raises(fb.FrozenBundleDrift) as exc:
|
||||
fb.bundle_dir("n500-2024")
|
||||
fb.bundle_dir("driftskrav-2027")
|
||||
assert "1 filer" in str(exc.value) # what disk holds now: one file of the pinned two
|
||||
|
||||
|
||||
|
|
@ -173,7 +176,7 @@ def test_a_matching_copy_resolves_to_the_pinned_directory(
|
|||
) -> None:
|
||||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
resolved = fb.bundle_dir("n500-2024")
|
||||
resolved = fb.bundle_dir("driftskrav-2027")
|
||||
assert resolved.parent == store
|
||||
assert (resolved / "index.md").is_file()
|
||||
|
||||
|
|
@ -188,15 +191,15 @@ def test_one_changed_byte_is_drift_named_with_both_digests(
|
|||
) -> None:
|
||||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
base = fb.bundle_dir("n500-2024") # green before the edit — the control
|
||||
expected = json.loads(pin.read_text(encoding="utf-8"))["bundles"]["n500-2024"]["sha256"]
|
||||
base = fb.bundle_dir("driftskrav-2027") # green before the edit — the control
|
||||
expected = json.loads(pin.read_text(encoding="utf-8"))["bundles"]["driftskrav-2027"]["sha256"]
|
||||
_touch_one_byte(base)
|
||||
|
||||
with pytest.raises(fb.FrozenBundleDrift) as exc:
|
||||
fb.bundle_dir("n500-2024")
|
||||
fb.bundle_dir("driftskrav-2027")
|
||||
message = str(exc.value)
|
||||
assert "avviker fra pin" in message
|
||||
assert "n500-2024" in message
|
||||
assert "driftskrav-2027" in message
|
||||
assert expected[: fb.SHORT] in message # what was pinned
|
||||
assert fb.digest_bundle(base)[0][: fb.SHORT] in message # what is on disk
|
||||
|
||||
|
|
@ -207,12 +210,12 @@ def test_a_missing_copy_is_missing_and_stays_an_oserror(
|
|||
"""``measure_stress`` already catches ``OSError`` -> IKKE MÅLT + exit 1; that path is unchanged."""
|
||||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
shutil.rmtree(store / fb.load_pins()["n500-2024"].directory)
|
||||
shutil.rmtree(store / fb.load_pins()["driftskrav-2027"].directory)
|
||||
|
||||
with pytest.raises(fb.FrozenBundleMissing) as exc:
|
||||
fb.bundle_dir("n500-2024")
|
||||
fb.bundle_dir("driftskrav-2027")
|
||||
assert isinstance(exc.value, OSError)
|
||||
assert "n500-2024" in str(exc.value)
|
||||
assert "driftskrav-2027" in str(exc.value)
|
||||
|
||||
|
||||
def test_drift_is_not_a_missing_copy(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
|
|
@ -224,10 +227,10 @@ def test_drift_is_not_a_missing_copy(tmp_path: Path, monkeypatch: pytest.MonkeyP
|
|||
|
||||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
_touch_one_byte(fb.bundle_dir("n500-2024"))
|
||||
_touch_one_byte(fb.bundle_dir("driftskrav-2027"))
|
||||
with pytest.raises(fb.FrozenBundleDrift):
|
||||
try:
|
||||
fb.bundle_dir("n500-2024")
|
||||
fb.bundle_dir("driftskrav-2027")
|
||||
except fb.FrozenBundleMissing: # pragma: no cover - the defect this arm forbids
|
||||
pytest.fail("drift was answered as a missing copy")
|
||||
|
||||
|
|
@ -238,22 +241,22 @@ def test_an_unknown_name_is_refused_by_name(
|
|||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
with pytest.raises(fb.FrozenBundleMissing) as exc:
|
||||
fb.bundle_dir("n100-2023")
|
||||
assert "n100-2023" in str(exc.value) and "n500-2024" in str(exc.value)
|
||||
fb.bundle_dir("prosesskatalog-2027")
|
||||
assert "prosesskatalog-2027" in str(exc.value) and "driftskrav-2027" in str(exc.value)
|
||||
|
||||
|
||||
def test_an_explicit_override_is_unpinned_and_the_operator_named_it(
|
||||
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""``--bundle-root`` / ``PORTFOLIO_VEGNORMAL_ROOT`` stays an escape hatch: the operator who
|
||||
"""``--bundle-root`` / ``PORTFOLIO_BUNDLE_ROOT`` stays an escape hatch: the operator who
|
||||
names a live mount gets it, pin or no pin. Absent it, the frozen store answers."""
|
||||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
live = tmp_path / "live"
|
||||
_bundle(live / "n500-2024", body="something else entirely")
|
||||
assert fb.bundle_dir("n500-2024", override=live) == live / "n500-2024"
|
||||
_bundle(live / "driftskrav-2027", body="something else entirely")
|
||||
assert fb.bundle_dir("driftskrav-2027", override=live) == live / "driftskrav-2027"
|
||||
monkeypatch.setenv(fb.OVERRIDE_ENV, str(live))
|
||||
assert fb.bundle_dir("n500-2024") == live / "n500-2024"
|
||||
assert fb.bundle_dir("driftskrav-2027") == live / "driftskrav-2027"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
|
@ -287,8 +290,31 @@ def test_the_pin_file_carries_no_key_nobody_reads() -> None:
|
|||
assert keys == {"source", "renewal", "bundles"}
|
||||
|
||||
|
||||
def test_the_bundle_itself_is_never_tracked_here() -> None:
|
||||
"""Vegnormal corpora must not reach a public remote: only the pin is tracked."""
|
||||
def test_the_default_store_is_the_packaged_example_bases(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""The pinned bases are the package's OWN fictional examples, so the default store sits next
|
||||
to the module — a wheel finds it without a checkout — and every pin resolves there, matched,
|
||||
with nothing in the environment helping.
|
||||
|
||||
Three halves, each a defect of its own: a default pointing at a machine-local path (the old
|
||||
``~/corpora`` store) makes every corpus arm skip on every other machine; a pin that does not
|
||||
resolve from the default is a pin nothing checks; and a copy left in the store with no pin is
|
||||
a stale base that ``ls`` shows and nothing measures. The pin file itself still carries only
|
||||
the pin, never bundle bytes.
|
||||
"""
|
||||
monkeypatch.delenv(fb.STORE_ENV, raising=False)
|
||||
monkeypatch.delenv(fb.OVERRIDE_ENV, raising=False)
|
||||
packaged = Path(fb.__file__).resolve().parent / "data" / "kunnskapsbaser"
|
||||
assert Path(fb.DEFAULT_STORE).resolve() == packaged
|
||||
pins = fb.load_pins()
|
||||
assert set(pins) == {"driftskrav-2027", "prosesskatalog-2027"}
|
||||
for name, pin in pins.items():
|
||||
base = fb.bundle_dir(name)
|
||||
assert base == Path(fb.DEFAULT_STORE) / pin.directory
|
||||
assert fb.digest_bundle(base)[1] == pin.files
|
||||
on_disk = sorted(p.name for p in packaged.iterdir() if p.is_dir())
|
||||
assert on_disk == sorted(pin.directory for pin in pins.values()), on_disk
|
||||
tracked = (_REPO / "src" / "portfolio_optimiser" / "frozen_bundles.json").read_text("utf-8")
|
||||
assert "index.md" not in tracked
|
||||
|
||||
|
|
@ -357,7 +383,7 @@ def _evidence(tmp_path: Path) -> tuple[dict[str, Any], Path]:
|
|||
"root": "scratchpad",
|
||||
"runs": [
|
||||
{
|
||||
"context": "contexts/tunnel-hauglia-2027",
|
||||
"context": "contexts/serverrom-2027",
|
||||
"outbox": "o",
|
||||
"run_id": "r",
|
||||
"bundle": None,
|
||||
|
|
@ -371,7 +397,7 @@ def test_a_drifted_copy_fails_the_gate_with_the_reason_said(
|
|||
) -> None:
|
||||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
_touch_one_byte(fb.bundle_dir("n500-2024"))
|
||||
_touch_one_byte(fb.bundle_dir("driftskrav-2027"))
|
||||
evidence, stress_root = _evidence(tmp_path)
|
||||
|
||||
m = gate.measure_stress(evidence, _REPO, stress_root, None)
|
||||
|
|
@ -422,14 +448,14 @@ def test_the_corpus_tests_skip_when_absent_but_fail_on_drift(
|
|||
store, pin = _store(tmp_path)
|
||||
_use(monkeypatch, store, pin)
|
||||
|
||||
assert resolve("n500-2024").is_dir() # control: the matching copy resolves
|
||||
_touch_one_byte(fb.bundle_dir("n500-2024"))
|
||||
assert resolve("driftskrav-2027").is_dir() # control: the matching copy resolves
|
||||
_touch_one_byte(fb.bundle_dir("driftskrav-2027"))
|
||||
# A skip is caught EXPLICITLY, never left to ``pytest.raises``: measured against the mutation
|
||||
# that makes drift a subclass of missing, these four arms SKIPPED instead of failing (5 -> 9
|
||||
# skipped over the whole suite) and stayed green — a gate that cannot see the one defect it
|
||||
# exists for.
|
||||
try:
|
||||
resolve("n500-2024")
|
||||
resolve("driftskrav-2027")
|
||||
except fb.FrozenBundleDrift:
|
||||
pass
|
||||
except pytest.skip.Exception as exc:
|
||||
|
|
@ -439,4 +465,4 @@ def test_the_corpus_tests_skip_when_absent_but_fail_on_drift(
|
|||
|
||||
monkeypatch.setenv(fb.STORE_ENV, str(tmp_path / "gone"))
|
||||
with pytest.raises(pytest.skip.Exception):
|
||||
resolve("n500-2024")
|
||||
resolve("driftskrav-2027")
|
||||
|
|
|
|||
|
|
@ -14,9 +14,9 @@ from portfolio_optimiser.ir import SavingsProposal
|
|||
from portfolio_optimiser.reference_domain import load_reference_projects
|
||||
from portfolio_optimiser.validator import Rejection, ValidatedProposal, proposal_for
|
||||
|
||||
# A feasible proposal for FV42-GSV-E1: 200k <= ~445k (30% of the affected total).
|
||||
# A feasible proposal for KONTOR-IT-E1: 200k <= ~445k (30% of the affected total).
|
||||
_VALID = (
|
||||
'{"project_id":"FV42-GSV-E1","measure":"Reduce scope",'
|
||||
'{"project_id":"KONTOR-IT-E1","measure":"Reduce scope",'
|
||||
'"affected_items":[{"code":"05.2","quantity":4300,"unit_cost":215},'
|
||||
'{"code":"03.1","quantity":1800,"unit_cost":310}],"claimed_saving_nok":200000}'
|
||||
)
|
||||
|
|
@ -24,7 +24,7 @@ _VALID = (
|
|||
|
||||
@pytest.fixture(scope="module")
|
||||
def project():
|
||||
return load_reference_projects()[0] # FV42-GSV-E1
|
||||
return load_reference_projects()[0] # KONTOR-IT-E1
|
||||
|
||||
|
||||
def _meter() -> TokenMeter:
|
||||
|
|
|
|||
|
|
@ -3,12 +3,12 @@
|
|||
P4 pkt. 0 anchored the demo against a repo-local reserve whose ``cost-baseline.json`` was DERIVED in
|
||||
code from the scripted register (``baseline_from_scripted_candidate``): the reserve's numbers are
|
||||
synthetic, so the script was the only ground truth there was. On GO day the direction **reverses** —
|
||||
a domain team delivered ``shared/examples/veglys-fv-soer/`` (commons ``002f000``) with its own
|
||||
``cost-baseline.json``, so the register is written FROM that file and the derivation helper is not
|
||||
used on this path.
|
||||
the delivered bundle (today the packaged, fictional ``data/bundles/klientpark-energi/``) ships its
|
||||
own ``cost-baseline.json``, so the register is written FROM that file and the derivation helper is
|
||||
not used on this path.
|
||||
|
||||
That reversal is what these tests guard. The risk it models is drift: two copies of the same numbers
|
||||
— one in commons' delivered file, one in our register — that stop agreeing without anyone noticing.
|
||||
— one in the delivered file, one in our register — that stop agreeing without anyone noticing.
|
||||
A drifted register does not crash; it makes the deterministic gate reject the demo's own hypothesis
|
||||
at stage 0 instead of at the P90 stage, which on screen is the SAME ``REJECTED`` line telling a
|
||||
different story. So the agreement is asserted directly (below), and the 10 % test proves the gate
|
||||
|
|
@ -22,13 +22,14 @@ mechanism could be perfect while the delivered file and the register disagree.
|
|||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from portfolio_optimiser import okf
|
||||
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
|
||||
from portfolio_optimiser.simulation import (
|
||||
_CANDIDATES,
|
||||
_INBOX_MARKER,
|
||||
_VEGLYS_PROJECT_ID,
|
||||
_KLIENTPARK_PROJECT_ID,
|
||||
ScriptedCandidate,
|
||||
_delivered_bundle_dir,
|
||||
materialize_anchored_bundle,
|
||||
|
|
@ -38,10 +39,10 @@ from portfolio_optimiser.persona import load_persona_example
|
|||
from portfolio_optimiser.validator import Rejection, ValidatedProposal
|
||||
|
||||
|
||||
def _veglys() -> ScriptedCandidate:
|
||||
def _klientpark() -> ScriptedCandidate:
|
||||
"""The registry entry for the delivered project — selected by id, never by index (the registry
|
||||
is a set of DATA entries whose order carries no meaning)."""
|
||||
(candidate,) = [c for c in _CANDIDATES if c.project_id == _VEGLYS_PROJECT_ID]
|
||||
(candidate,) = [c for c in _CANDIDATES if c.project_id == _KLIENTPARK_PROJECT_ID]
|
||||
return candidate
|
||||
|
||||
|
||||
|
|
@ -53,12 +54,12 @@ def _delivered_baseline() -> CostBaseline:
|
|||
|
||||
def test_the_delivered_bundle_ships_its_own_baseline() -> None:
|
||||
"""The delivered bundle is anchored WITHOUT ``materialize_anchored_bundle`` — which is exactly
|
||||
what let GO day be a call-site change rather than a new seam. RED if commons ever drops the
|
||||
what let GO day be a call-site change rather than a new seam. RED if the bundle ever drops the
|
||||
file: the demo would then run un-anchored while every other line stayed byte-identical."""
|
||||
baseline = _delivered_baseline()
|
||||
|
||||
assert baseline.project_id == _VEGLYS_PROJECT_ID
|
||||
assert set(baseline.items) == {"ENERGI-VEGLYS-EL"}
|
||||
assert baseline.project_id == _KLIENTPARK_PROJECT_ID
|
||||
assert set(baseline.items) == {"ENERGI-KLIENTPARK-EL"}
|
||||
|
||||
|
||||
def test_the_register_states_the_delivered_numbers_verbatim() -> None:
|
||||
|
|
@ -66,13 +67,13 @@ def test_the_register_states_the_delivered_numbers_verbatim() -> None:
|
|||
written FROM the delivered ``cost-baseline.json``, so they must equal it exactly — in BOTH
|
||||
scripted replies.
|
||||
|
||||
RED when the two drift: if commons re-states the baseline (or someone re-types it here), the
|
||||
RED when the two drift: if the bundle re-states the baseline (or someone re-types it here), the
|
||||
demo's own hypothesis starts being felled by stage 0's reconciliation instead of by the P90
|
||||
stage it narrates. Same ``REJECTED`` line on screen, different mechanism behind it.
|
||||
|
||||
Both replies are checked, not just the corrected one: the overclaimed hypothesis must differ
|
||||
from it in the CLAIM alone, never in the cost lines."""
|
||||
candidate = _veglys()
|
||||
candidate = _klientpark()
|
||||
delivered = _delivered_baseline().items
|
||||
|
||||
for reply in (candidate.overclaimed, candidate.corrected):
|
||||
|
|
@ -103,7 +104,7 @@ def test_the_overclaim_and_the_markers_are_absent_from_the_delivered_bundle() ->
|
|||
each appear to work while carrying nothing.
|
||||
|
||||
This is a CONTENT check, and delivered content is the reason it exists: the register's figure was
|
||||
chosen against commons' number inventory, and this is what keeps that choice honest if either
|
||||
chosen against the bundle's number inventory, and this is what keeps that choice honest if either
|
||||
side changes."""
|
||||
text = "\n".join(
|
||||
path.read_text(encoding="utf-8")
|
||||
|
|
@ -111,7 +112,7 @@ def test_the_overclaim_and_the_markers_are_absent_from_the_delivered_bundle() ->
|
|||
if path.is_file()
|
||||
)
|
||||
|
||||
for token in (_veglys().flip_key, load_persona_example().marker, _INBOX_MARKER):
|
||||
for token in (_klientpark().flip_key, load_persona_example().marker, _INBOX_MARKER):
|
||||
assert token not in text, (
|
||||
f"{token!r} already occurs in the delivered bundle — the walkthrough would trace a "
|
||||
"token the content supplied, not one the loop carried"
|
||||
|
|
@ -126,7 +127,7 @@ async def test_the_delivered_bundle_runs_the_whole_demo(tmp_path) -> None:
|
|||
meaning. It also pins WHICH stage rejects #1 — were stage 0 to start rejecting it, the screen
|
||||
would still show a REJECTED and a VALIDATED line while demonstrating a different mechanism."""
|
||||
result = await simulate_learning_loop(
|
||||
str(_delivered_bundle_dir()), str(tmp_path), project_id=_VEGLYS_PROJECT_ID
|
||||
str(_delivered_bundle_dir()), str(tmp_path), project_id=_KLIENTPARK_PROJECT_ID
|
||||
)
|
||||
|
||||
assert isinstance(result.run_a.outcome, ValidatedProposal)
|
||||
|
|
@ -149,8 +150,8 @@ async def test_a_deviating_delivered_baseline_forkaster_the_run_before_the_solve
|
|||
agreement is only worth something if a DISagreement would be caught. The script is byte-identical
|
||||
to the control above; only the declared baseline moves.
|
||||
|
||||
The deviated copy is materialized outside ``shared/`` — the delivered bundle is a pull-only
|
||||
subtree and criterion 8 requires it byte-unchanged."""
|
||||
The deviated copy is materialized in ``tmp_path`` — criterion 8 requires the delivered bundle
|
||||
byte-unchanged."""
|
||||
delivered = _delivered_baseline()
|
||||
deviated = CostBaseline(
|
||||
project_id=delivered.project_id,
|
||||
|
|
@ -162,7 +163,9 @@ async def test_a_deviating_delivered_baseline_forkaster_the_run_before_the_solve
|
|||
bundle = materialize_anchored_bundle(
|
||||
tmp_path / "avvikende", source=_delivered_bundle_dir(), baseline=deviated
|
||||
)
|
||||
result = await simulate_learning_loop(str(bundle), str(tmp_path), project_id=_VEGLYS_PROJECT_ID)
|
||||
result = await simulate_learning_loop(
|
||||
str(bundle), str(tmp_path), project_id=_KLIENTPARK_PROJECT_ID
|
||||
)
|
||||
|
||||
outcome = result.run_a.outcome
|
||||
assert isinstance(outcome, Rejection), (
|
||||
|
|
@ -170,24 +173,34 @@ async def test_a_deviating_delivered_baseline_forkaster_the_run_before_the_solve
|
|||
"deterministic gate is not anchored to the delivered baseline"
|
||||
)
|
||||
assert "outside the 5.0% tolerance" in outcome.reason
|
||||
assert "ENERGI-VEGLYS-EL" in outcome.reason
|
||||
assert "ENERGI-KLIENTPARK-EL" in outcome.reason
|
||||
assert "P90" not in outcome.reason, (
|
||||
"rejected by the solver stage, not by the reconciliation stage 0 that must run BEFORE it"
|
||||
)
|
||||
|
||||
|
||||
def test_the_delivered_bundle_is_never_mutated_by_a_run() -> None:
|
||||
"""Criterion 8's day-to-day half: the demo copies the bundle, so the commons-owned files under
|
||||
``shared/`` are byte-unchanged after everything above has run."""
|
||||
import subprocess
|
||||
def _snapshot(bundle: Path) -> dict[str, bytes]:
|
||||
return {
|
||||
path.relative_to(bundle).as_posix(): path.read_bytes()
|
||||
for path in sorted(bundle.rglob("*"))
|
||||
if path.is_file()
|
||||
}
|
||||
|
||||
proc = subprocess.run(
|
||||
["git", "status", "--porcelain", "--", "shared/examples/veglys-fv-soer"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=True,
|
||||
)
|
||||
assert proc.stdout == "", f"the delivered bundle was modified in place: {proc.stdout!r}"
|
||||
|
||||
async def test_the_delivered_bundle_is_never_mutated_by_a_run(tmp_path) -> None:
|
||||
"""Criterion 8's day-to-day half: the demo copies the bundle, so the delivered files are
|
||||
byte-unchanged after a whole two-run walkthrough over them.
|
||||
|
||||
Measured around a run of its own rather than read off the version-control status: the bundle's
|
||||
bytes before and after are compared directly, so the check holds whether or not the bundle is
|
||||
committed yet, and it does not depend on which other tests happened to run first."""
|
||||
bundle = _delivered_bundle_dir()
|
||||
before = _snapshot(bundle)
|
||||
assert before, "the delivered bundle has no files -- the comparison below would be vacuous"
|
||||
|
||||
await simulate_learning_loop(str(bundle), str(tmp_path), project_id=_KLIENTPARK_PROJECT_ID)
|
||||
|
||||
assert _snapshot(bundle) == before, "the delivered bundle was modified in place by a run"
|
||||
|
||||
|
||||
def test_the_reserve_entry_and_the_delivered_entry_are_both_registered() -> None:
|
||||
|
|
@ -195,7 +208,7 @@ def test_the_reserve_entry_and_the_delivered_entry_are_both_registered() -> None
|
|||
three-line revert, which it only is while the reserve's registry entry still exists."""
|
||||
ids = {c.project_id for c in _CANDIDATES}
|
||||
|
||||
assert {"BYGG-KONTOR-NORD", _VEGLYS_PROJECT_ID} <= ids
|
||||
assert {"BYGG-KONTOR-NORD", _KLIENTPARK_PROJECT_ID} <= ids
|
||||
|
||||
|
||||
def test_the_projection_names_the_project_the_register_keys_on() -> None:
|
||||
|
|
@ -206,5 +219,5 @@ def test_the_projection_names_the_project_the_register_keys_on() -> None:
|
|||
(_delivered_bundle_dir() / "validator-input.json").read_text(encoding="utf-8")
|
||||
)
|
||||
|
||||
assert projection["project_id"] == _VEGLYS_PROJECT_ID
|
||||
assert projection["project_id"] == _KLIENTPARK_PROJECT_ID
|
||||
assert _delivered_baseline().project_id == projection["project_id"]
|
||||
|
|
|
|||
|
|
@ -21,7 +21,8 @@ Arms, each with a named detach point:
|
|||
* **(a) the null offer is stated.** A delivered text with no identifier of any measured form and no
|
||||
baseline reports ``identifiers=0, cost_lines=0``. This is the K2 CONTROL — without it a green
|
||||
known positive proves nothing.
|
||||
* **(b) the positive offer is stated.** N100's delivered input carries the BINDING known positive
|
||||
* **(b) the positive offer is stated.** A requirements base's delivered input carries the BINDING
|
||||
known positive
|
||||
``Krav 3.3.1—13`` (P7 premiss (vii): it stands in 6 of 6 prompts, EM-DASH U+2014), and the offer
|
||||
is positive there. Paired with a control that the search CAN find the string at all.
|
||||
* **(c) the offer is measured on the SAME text the gate will see.** ``grounding_offer`` composes
|
||||
|
|
@ -42,11 +43,11 @@ only thing a mutation table is for.
|
|||
|
||||
The fixtures are P7's own tracked generation prompts. Tracked test data on purpose (``scratchpad/``
|
||||
is absent from ``git archive HEAD``) — and REUSED rather than extended with the delivered cuts,
|
||||
because those cuts are 10 kB and 96 kB of a real Norwegian tender and a public road standard, and
|
||||
because those cuts were 10 kB and 96 kB of a real Norwegian tender and a published standard, and
|
||||
committing corpus text into a repo published on ``open/`` is a publication decision that belongs to
|
||||
the operator, not to this order. The 8-distinct / 435-distinct figures for the delivered cut and the
|
||||
full grounding are MEASURED and reported in ``docs/2026-09-09-p8-forankringstilbudet.md``; what a
|
||||
test asserts here is the property, on text this repo already ships.
|
||||
full grounding were MEASURED at the time (the ledger is ``docs/invarianter.md``); what a test
|
||||
asserts here is the property, on text this repo already ships.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -71,7 +72,8 @@ KNOWN_POSITIVE = "Krav 3.3.1—13"
|
|||
|
||||
#: The base with a ``cost-baseline.json`` (one line) and the one without — the pair that makes the
|
||||
#: renderer's omission arm reachable rather than asserted.
|
||||
ANCHORED = SHARED / "veglys-fv-soer"
|
||||
BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
ANCHORED = BUNDLES / "klientpark-energi"
|
||||
UNANCHORED = SHARED / "bygg-energi-mikro"
|
||||
_UNANCHORED_PID = "BYGG-KONTOR-NORD"
|
||||
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (sim)"}
|
||||
|
|
@ -138,11 +140,11 @@ def test_the_known_positive_input_reports_a_positive_offer() -> None:
|
|||
"""LOAD-BEARING (b) — the BINDING known positive. Paired with the control that the string is
|
||||
actually there: a report that finds nothing because it searched for nothing would otherwise
|
||||
pass (a) and look measured."""
|
||||
text = _prompt("p4-n100-generation-prompt.txt")
|
||||
text = _prompt("p4-driftskrav-generation-prompt.txt")
|
||||
assert KNOWN_POSITIVE in text, "the fixture no longer carries the known positive"
|
||||
offer = grounding_offer(_project(), None, Grounding.of(text))
|
||||
assert offer.identifiers > 0, offer
|
||||
assert offer.cost_lines == 0, "a road standard carries no cost lines (F4's own finding)"
|
||||
assert offer.cost_lines == 0, "a requirements base carries no cost lines (F4's own finding)"
|
||||
|
||||
|
||||
def test_a_bare_number_is_not_counted_as_an_offer() -> None:
|
||||
|
|
@ -151,8 +153,8 @@ def test_a_bare_number_is_not_counted_as_an_offer() -> None:
|
|||
measurement inert — the repo's cardinal class, a gate that can only come out green.
|
||||
|
||||
**NARROWED in P19/B1, and the narrowing is a measured honesty limit rather than a weakening.**
|
||||
The list used to carry ``42.5`` as well. R761's requirement numbers — all six ``ref`` values in
|
||||
``contexts/kontrakt-sorasen-2027/fasit.json`` — are ``12.1`` / ``52.11`` / ``22.1``, which is
|
||||
The list used to carry ``42.5`` as well. A process catalogue's numbers — all six ``ref`` values
|
||||
in ``contexts/driftsavtale-2027/fasit.json`` — are ``12.1`` / ``52.11`` / ``22.1``, which is
|
||||
the SAME typography as a decimal. There is no rule that separates them, so the form counts
|
||||
both; the arm below states that in the open instead of leaving it in the docstring. What this
|
||||
arm still refuses is the class the K2 number was measured on: bare INTEGERS."""
|
||||
|
|
@ -161,9 +163,10 @@ def test_a_bare_number_is_not_counted_as_an_offer() -> None:
|
|||
|
||||
|
||||
def test_a_dotted_number_is_counted_and_that_is_an_honesty_limit() -> None:
|
||||
"""P19/B1: a decimal and an R761 process number are typographically the SAME token.
|
||||
"""P19/B1: a decimal and a catalogue's process number are typographically the SAME token.
|
||||
|
||||
Counted, therefore — which is the direction that keeps r761 measurable (its whole offer was 3
|
||||
Counted, therefore — which is the direction that keeps a process catalogue measurable (the
|
||||
one measured then had a whole offer of 3
|
||||
identifiers over 6.5 MB before this form) at the price of a decimal in prose being counted as
|
||||
one. Stated here rather than hidden: for the REPORT it inflates the count, and for P19/B3's
|
||||
gate it can turn the guard on in a base that carries decimals and no real identifiers — where
|
||||
|
|
@ -257,7 +260,7 @@ async def test_an_anchored_run_reports_its_cost_lines() -> None:
|
|||
"""LOAD-BEARING (d), the CONTROL that ``cost_lines`` is not a constant zero: the anchored base
|
||||
ships exactly one line, and the run reports it."""
|
||||
report = await run_project(
|
||||
"VEGLYS-FV-SOER",
|
||||
"KLIENTPARK-ENERGI",
|
||||
"local",
|
||||
docs_dir=str(ANCHORED),
|
||||
bundle_dir=str(ANCHORED),
|
||||
|
|
@ -321,7 +324,7 @@ def test_the_cli_says_nothing_when_the_run_can_anchor(capsys) -> None:
|
|||
all. Without it, a notice that always fired would pass the arm above."""
|
||||
rc = run_mod.main(
|
||||
[
|
||||
"VEGLYS-FV-SOER",
|
||||
"KLIENTPARK-ENERGI",
|
||||
"--docs-dir",
|
||||
str(ANCHORED),
|
||||
"--bundle-dir",
|
||||
|
|
|
|||
|
|
@ -62,8 +62,9 @@ from portfolio_optimiser import explore, okf
|
|||
from portfolio_optimiser.explore import navigator_tools
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
#: The largest FLAT base shipped - the one the MAJOR-3 gate bounds.
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
#: The largest FLAT example base the package ships - the one the MAJOR-3 gate bounds.
|
||||
_KJOLING = _BUNDLES / "driftssenter-kjoling"
|
||||
#: The only NESTED base shipped, and it is tiny: three levels, six concepts. It proves the shape is
|
||||
#: read correctly; it cannot prove the shape pays for itself, which is what arm (b) is for.
|
||||
_NESTED = _EXAMPLES / "nav-golden-hierarchy" / "bundle"
|
||||
|
|
@ -140,8 +141,8 @@ def test_a_flat_base_lists_exactly_what_it_listed_before() -> None:
|
|||
"""(a) The order names this one explicitly. Every base shipped in this repo except one is flat,
|
||||
and the whole MAJOR-3 measurement was taken over them: if the new rung changed what a flat base
|
||||
answers, this change would be a rewrite of that result rather than a level above it."""
|
||||
bundle = okf.navigate_bundle(str(_TUNNEL))
|
||||
listing = _read_bundle(_TUNNEL)
|
||||
bundle = okf.navigate_bundle(str(_KJOLING))
|
||||
listing = _read_bundle(_KJOLING)
|
||||
|
||||
assert listing["documents"] == [_entry(f, f.name) for f in bundle.context_files]
|
||||
assert listing["directories"] == [], "a flat base has no subdirectories to report"
|
||||
|
|
@ -255,10 +256,10 @@ def test_the_whole_document_is_still_one_call_away() -> None:
|
|||
"""(f) The bound is a disclosure LEVEL, not data loss - MAJOR-3's arm (e), re-asserted over the
|
||||
new shape because a listing that no longer names documents the way ``read_file`` takes them
|
||||
would have broken the ladder while every cost arm stayed green."""
|
||||
listing = _read_bundle(_TUNNEL)
|
||||
listing = _read_bundle(_KJOLING)
|
||||
biggest = max(listing["documents"], key=lambda e: int(e["chars"]))
|
||||
|
||||
whole = _tools(_TUNNEL)["read_file"].func(bundle_id=_TUNNEL.name, path=str(biggest["name"]))
|
||||
whole = _tools(_KJOLING)["read_file"].func(bundle_id=_KJOLING.name, path=str(biggest["name"]))
|
||||
|
||||
assert len(whole) > _CEILING_CHARS
|
||||
assert int(biggest["chars"]) <= len(whole)
|
||||
|
|
|
|||
|
|
@ -33,7 +33,7 @@ _PROVENANCE = ProvenanceStamp(
|
|||
role="proposer",
|
||||
validator_decision="validated",
|
||||
token_usage=8,
|
||||
# These fixtures stand in for an ordinary complete run; the road path is anchored by
|
||||
# These fixtures stand in for an ordinary complete run; the reference path is anchored by
|
||||
# construction, so ``True`` is the honest value here. The un-anchored case has its own file.
|
||||
cost_baseline_anchored=True,
|
||||
bundle_id_source=None,
|
||||
|
|
|
|||
|
|
@ -31,7 +31,7 @@ _PROVENANCE = ProvenanceStamp(
|
|||
role="proposer",
|
||||
validator_decision="validated",
|
||||
token_usage=8,
|
||||
# These fixtures stand in for an ordinary complete run; the road path is anchored by
|
||||
# These fixtures stand in for an ordinary complete run; the reference path is anchored by
|
||||
# construction, so ``True`` is the honest value here. The un-anchored case has its own file.
|
||||
cost_baseline_anchored=True,
|
||||
bundle_id_source=None,
|
||||
|
|
|
|||
|
|
@ -98,7 +98,7 @@ _PROVENANCE = ProvenanceStamp(
|
|||
role="proposer",
|
||||
validator_decision="validated",
|
||||
token_usage=8,
|
||||
# These fixtures stand in for an ordinary complete run; the road path is anchored by
|
||||
# These fixtures stand in for an ordinary complete run; the reference path is anchored by
|
||||
# construction, so ``True`` is the honest value here. The un-anchored case has its own file.
|
||||
cost_baseline_anchored=True,
|
||||
bundle_id_source=None,
|
||||
|
|
|
|||
|
|
@ -9,9 +9,11 @@ cost codes — ``M-04-01`` / ``M-04-03`` — that appear in NO prompt of that ru
|
|||
tied to whether a cost baseline happened to exist. **The input always exists; the baseline does
|
||||
not.**
|
||||
|
||||
Every fixture under ``tests/fixtures/p7-grounding/`` is the VERBATIM generation prompt of a free
|
||||
recording — the entire input the proposer saw on the attempt that produced its candidate. They are
|
||||
tracked test data on purpose: the recordings live in ``scratchpad/``, which ``git archive HEAD``
|
||||
Every fixture under ``tests/fixtures/p7-grounding/`` is a generation prompt in the exact form a
|
||||
free recording captured — the entire input the proposer saw on the attempt that produced its
|
||||
candidate. The two K2 fixtures are verbatim; the ``p4-driftskrav`` fixture is a fictional rewrite of
|
||||
a recorded prompt that keeps its structure and the two facts its tests read (the known positive is
|
||||
present verbatim, the control identifier is absent). They are tracked test data on purpose: the recordings live in ``scratchpad/``, which ``git archive HEAD``
|
||||
does not carry, so a test reading them would pass here and fail the handover gate.
|
||||
"""
|
||||
|
||||
|
|
@ -36,7 +38,7 @@ from portfolio_optimiser.validator import (
|
|||
|
||||
FIXTURES = Path(__file__).parent / "fixtures" / "p7-grounding"
|
||||
|
||||
#: The known positive (premiss (vii), `docs/2026-09-08-n-bundlene-hypoteseform.md` § 11.4): the
|
||||
#: The known positive (premiss (vii), measured on a requirements corpus during development): the
|
||||
#: model quoted this requirement number VERBATIM in 6 of 6 replies, and it stands in 6 of 6
|
||||
#: prompts. Note the EM-DASH (U+2014): the hyphen variant scores 0 of 6 in both.
|
||||
KNOWN_POSITIVE = "Krav 3.3.1—13"
|
||||
|
|
@ -51,7 +53,7 @@ def _proposal(*codes: str, unit_cost: float = 100.0) -> SavingsProposal:
|
|||
so a rejection can only have come from the grounding stage."""
|
||||
return SavingsProposal(
|
||||
project_id="K2",
|
||||
measure="reduce marking scope",
|
||||
measure="reduce cabling scope",
|
||||
affected_items=[AffectedItem(code=c, quantity=10.0, unit_cost=unit_cost) for c in codes],
|
||||
claimed_saving_nok=1.0,
|
||||
)
|
||||
|
|
@ -112,10 +114,10 @@ def test_b_the_reason_names_every_ungrounded_identifier_in_the_proposal_s_own_or
|
|||
|
||||
|
||||
def test_c_the_known_positive_is_not_flagged() -> None:
|
||||
"""``Krav 3.3.1—13`` stands VERBATIM in the N100 generation prompt, so a proposal that cites it
|
||||
is grounded and the grounding stage must stay silent. A rule that flags everything code-shaped
|
||||
"""``Krav 3.3.1—13`` stands VERBATIM in the driftskrav generation prompt, so a proposal that
|
||||
cites it is grounded and the grounding stage must stay silent. A rule that flags everything code-shaped
|
||||
is red here."""
|
||||
text = _prompt("p4-n100-generation-prompt.txt")
|
||||
text = _prompt("p4-driftskrav-generation-prompt.txt")
|
||||
assert KNOWN_POSITIVE in text, "fixture drifted: the known positive is not in the prompt"
|
||||
ruling = validate_proposal(
|
||||
_proposal(KNOWN_POSITIVE), baseline=None, grounding=Grounding.of(text)
|
||||
|
|
@ -125,9 +127,9 @@ def test_c_the_known_positive_is_not_flagged() -> None:
|
|||
|
||||
def test_c_control_the_rule_can_still_flag_on_that_same_recording() -> None:
|
||||
"""THE CONTROL that keeps the arm above from proving nothing. ``CRS-01`` is the one identifier
|
||||
the model produced on the N100 recording that no prompt carries (1 of 2 ungrounded). Same
|
||||
the model produced on the requirements recording that no prompt carries (1 of 2 ungrounded). Same
|
||||
fixture, same call — only the identifier differs, and this one must fall."""
|
||||
text = _prompt("p4-n100-generation-prompt.txt")
|
||||
text = _prompt("p4-driftskrav-generation-prompt.txt")
|
||||
assert "CRS-01" not in text
|
||||
ruling = validate_proposal(_proposal("CRS-01"), baseline=None, grounding=Grounding.of(text))
|
||||
assert isinstance(ruling, Rejection)
|
||||
|
|
@ -142,7 +144,7 @@ def test_c_control_the_rule_can_still_flag_on_that_same_recording() -> None:
|
|||
def test_d_a_code_quoted_verbatim_from_the_input_is_not_flagged() -> None:
|
||||
"""The discriminator between this rule and "flag anything that looks like a code". The token is
|
||||
deliberately shaped like the fabrications above; the ONLY difference is that the input says it."""
|
||||
text = "Context:\nPrice schedule line ZZZ-999-01 covers technical marking.\n"
|
||||
text = "Context:\nPrice schedule line ZZZ-999-01 covers network cabling.\n"
|
||||
ruling = validate_proposal(_proposal("ZZZ-999-01"), baseline=None, grounding=Grounding.of(text))
|
||||
assert isinstance(ruling, ValidatedProposal), getattr(ruling, "reason", "")
|
||||
|
||||
|
|
@ -211,7 +213,7 @@ async def test_f_generate_via_llm_grounds_the_candidate_in_the_prompt_it_sent()
|
|||
result = await generate_via_llm(
|
||||
ScriptedChatClient(reply=reply, tokens_per_reply=8),
|
||||
project,
|
||||
"The price schedule carries line REAL-11-02 for technical marking.",
|
||||
"The price schedule carries line REAL-11-02 for network cabling.",
|
||||
meter,
|
||||
max_attempts=1,
|
||||
)
|
||||
|
|
@ -233,7 +235,7 @@ async def test_f_control_a_code_the_context_names_survives_the_same_call() -> No
|
|||
result = await generate_via_llm(
|
||||
ScriptedChatClient(reply=reply, tokens_per_reply=8),
|
||||
project,
|
||||
"The price schedule carries line REAL-11-02 for technical marking.",
|
||||
"The price schedule carries line REAL-11-02 for network cabling.",
|
||||
meter,
|
||||
max_attempts=1,
|
||||
)
|
||||
|
|
@ -367,7 +369,7 @@ async def test_h_the_refusal_does_not_ground_the_next_attempt_that_repeats_the_c
|
|||
result = await generate_via_llm(
|
||||
client,
|
||||
_project(),
|
||||
"The price schedule carries line REAL-11-02 for technical marking.",
|
||||
"The price schedule carries line REAL-11-02 for network cabling.",
|
||||
TokenMeter(Budget(max_tokens=10**9, max_rounds=20)),
|
||||
max_attempts=2,
|
||||
)
|
||||
|
|
|
|||
|
|
@ -3,31 +3,33 @@
|
|||
P7 made stage 0b ``item.code in grounding``: plain containment over ONE concatenated string. P16
|
||||
then ran it against a delivered corpus and measured what containment cannot tell apart. The
|
||||
falsification arm ``a4-indeksregulering`` proposed a 250 000 NOK saving on a single cost line whose
|
||||
code was ``R761`` — the knowledge base's OWN NAME, which every one of its 2 756 concept documents
|
||||
carries — and the whole gate said ``validated``: stage 0 was skipped (un-anchored run), stage 0b
|
||||
was satisfied by the letterhead, and the checker approved.
|
||||
code was the knowledge base's OWN NAME — the catalogue designation every one of its 2 756 concept
|
||||
documents carries — and the whole gate said ``validated``: stage 0 was skipped (un-anchored run),
|
||||
stage 0b was satisfied by the letterhead, and the checker approved.
|
||||
|
||||
The rule this file measures: a code grounds only if it is at least ``_GROUNDING_MIN_LENGTH``
|
||||
characters AND appears in fewer than ``_GROUNDING_MAX_DOCUMENT_SHARE`` of the grounding's
|
||||
DOCUMENTS — with an absolute floor, because a share over a handful of documents is not a
|
||||
measurement (one of three is 33 % and says nothing).
|
||||
|
||||
**N and A are MEASURED, not chosen** (14.09, the four mounted vegnormal bases):
|
||||
**N and A are MEASURED, not chosen** (14.09, over the four corpora delivered at the time):
|
||||
|
||||
* every ``must_cite`` reference and every mandate ``affected_code`` in the four context sets: the
|
||||
shortest real identifier is FOUR characters (``12.1``, ``52.1``), so ``N = 3`` sits one below the
|
||||
measurement and cannot refuse anything measured;
|
||||
* document frequency of every code-shaped token (``generate._IDENTIFIER_FORMS``) in each base:
|
||||
1 692 distinct tokens and NOT ONE reaches 5 % of its base's documents. Highest anywhere 6 of 446
|
||||
(1.35 %); highest that a fasit names 3 of 446 (0.67 %); ``R761`` 2 756 of 2 756 (100 %). ``A =
|
||||
(1.35 %); highest that a fasit names 3 of 446 (0.67 %); the base's own name 2 756 of 2 756
|
||||
(100 %). ``A =
|
||||
0.05`` therefore sits 3.7x above the highest real token and 20x below the defect.
|
||||
|
||||
**The denominator is NAMED in the refusal**, because Step 5 feeds that reason verbatim into the
|
||||
next attempt's prompt: a proposer told only "ungrounded" answers with another token of the same
|
||||
kind, while one told "it is in 2 756 of 2 756 documents" has been told what is wrong with it.
|
||||
|
||||
The arms that need the delivered bases SKIP with the root named; the rule's own algebra, the
|
||||
floor, and the composition seam run over synthetic input and are UNCONDITIONAL.
|
||||
The arms that need a knowledge base read the package's pinned example bases (and SKIP, with the
|
||||
store named, only when a user's own store lacks them); the rule's own algebra, the floor, and the
|
||||
composition seam run over synthetic input and are UNCONDITIONAL.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -49,17 +51,12 @@ from portfolio_optimiser.validator import (
|
|||
validate_proposal,
|
||||
)
|
||||
|
||||
_A4 = Path(
|
||||
"scratchpad/p14-stress/kontrakt-sorasen-2027/"
|
||||
"kontrakt-sorasen-2027-01-a4-indeksregulering-proposal.json"
|
||||
)
|
||||
|
||||
|
||||
def _base(name: str) -> Path:
|
||||
"""The FROZEN copy this repository pins, resolved at call time.
|
||||
|
||||
Absence SKIPS (MAJOR-3's ceiling: no corpus is mounted in the handover archive), drift is
|
||||
allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
Absence SKIPS (a user's own store, named by ``PORTFOLIO_FROZEN_BUNDLES``, may not hold it),
|
||||
drift is allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
"""
|
||||
try:
|
||||
return frozen_bundles.bundle_dir(name)
|
||||
|
|
@ -98,22 +95,22 @@ def _corpus(*, documents: int, everywhere: str, once: str) -> Grounding:
|
|||
# --- the measured defect --------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_the_a4_proposal_p16_validated_is_now_refused() -> None:
|
||||
"""(a) THE KNOWN POSITIVE, and it is a recording rather than a construction: the proposal is
|
||||
the artefact P16's paid run wrote, replayed against the base that run was given, with
|
||||
``baseline=None`` — exactly the configuration under which it said ``validated``."""
|
||||
if not _A4.is_file():
|
||||
pytest.skip("P16's a4 artefact is not present in this checkout")
|
||||
proposal = SavingsProposal.model_validate(
|
||||
json.loads(_A4.read_text(encoding="utf-8"))["proposal"]
|
||||
)
|
||||
assert [i.code for i in proposal.affected_items] == ["R761"], "the artefact drifted"
|
||||
def test_the_catalogues_own_name_as_a_cost_code_is_refused() -> None:
|
||||
"""(a) THE KNOWN POSITIVE. P16's paid run proposed the catalogue's own designation as a cost
|
||||
code and the gate said ``validated``. That recording was made against a corpus this repository
|
||||
no longer carries, so the proposal is rebuilt in the same shape against the example catalogue,
|
||||
whose designation ``P900`` stands in every one of its concept documents — with
|
||||
``baseline=None``, exactly the configuration under which the original said ``validated``.
|
||||
CONTROL: the same base's own process number ``52.11`` still grounds (arm (b) covers every
|
||||
one)."""
|
||||
grounding = _grounding_over("prosesskatalog-2027")
|
||||
assert _inert_in(grounding, "52.11") is None, "control: a real process number must ground"
|
||||
|
||||
ruling = validate_proposal(proposal, baseline=None, grounding=_grounding_over("r761-2025"))
|
||||
ruling = validate_proposal(_proposal("P900"), baseline=None, grounding=grounding)
|
||||
|
||||
assert isinstance(ruling, Rejection)
|
||||
assert "'R761'" in ruling.reason
|
||||
assert "2756 of the 2756" in ruling.reason, (
|
||||
assert "'P900'" in ruling.reason
|
||||
assert "301 of the 301" in ruling.reason, (
|
||||
"the refusal must name the denominator: Step 5 feeds this reason verbatim into the next "
|
||||
f"attempt's prompt — got {ruling.reason!r}"
|
||||
)
|
||||
|
|
@ -123,10 +120,8 @@ def test_every_fasit_reference_still_grounds() -> None:
|
|||
"""(b) THE KNOWN NEGATIVE over the same corpora, with its denominator stated. A rule that made
|
||||
the defect inert by making real references inert too would pass (a) perfectly."""
|
||||
sets = {
|
||||
"gate-nordvik-2027": "n100-2023",
|
||||
"fv412-dekkefornyelse-2027": "n200-2024",
|
||||
"tunnel-hauglia-2027": "n500-2024",
|
||||
"kontrakt-sorasen-2027": "r761-2025",
|
||||
"serverrom-2027": "driftskrav-2027",
|
||||
"driftsavtale-2027": "prosesskatalog-2027",
|
||||
}
|
||||
checked = 0
|
||||
for context, base in sets.items():
|
||||
|
|
@ -138,7 +133,7 @@ def test_every_fasit_reference_still_grounds() -> None:
|
|||
f"{reference!r} is a real requirement of {base} and the rule made it inert"
|
||||
)
|
||||
checked += 1
|
||||
assert checked == 26, f"population moved: {checked} references, expected 26"
|
||||
assert checked == 12, f"population moved: {checked} references, expected 12"
|
||||
|
||||
|
||||
# --- the rule's own algebra, unconditional ---------------------------------------------------
|
||||
|
|
@ -159,7 +154,7 @@ def test_a_token_in_one_document_grounds_and_one_in_all_of_them_does_not() -> No
|
|||
|
||||
|
||||
def test_a_token_too_short_to_identify_anything_is_inert() -> None:
|
||||
"""(d) The length conjunct, which the SHARE does not cover: ``R761`` is four characters, so
|
||||
"""(d) The length conjunct, which the SHARE does not cover: ``P900`` is four characters, so
|
||||
length is not what made the measured defect inert. This is the coincidence class the
|
||||
measurement did not happen to contain — a one- or two-character token is in any prose."""
|
||||
grounding = Grounding(documents=("the line A is here", *("filler" for _ in range(50))))
|
||||
|
|
@ -181,8 +176,9 @@ def test_a_share_is_not_taken_over_a_handful_of_documents() -> None:
|
|||
|
||||
def test_a_caller_that_declares_no_boundaries_is_byte_for_byte_the_old_gate() -> None:
|
||||
"""(f) ``Grounding.of`` is the honest reading of a caller with nothing to declare, and it can
|
||||
never trip the share: one document cannot reach the floor. This is what keeps every road-path
|
||||
run and every pre-P18 test unchanged BY CONSTRUCTION rather than by exemption."""
|
||||
never trip the share: one document cannot reach the floor. This is what keeps every
|
||||
single-document run and every pre-P18 test unchanged BY CONSTRUCTION rather than by
|
||||
exemption."""
|
||||
text = "en tekst som nevner KODE-99 og ellers ingenting"
|
||||
single = Grounding.of(text)
|
||||
|
||||
|
|
|
|||
|
|
@ -126,7 +126,7 @@ def test_save_is_byte_identical_regardless_of_order(tmp_path) -> None:
|
|||
|
||||
# --- Step 8: hard/soft goal-stop in run_portfolio on the accumulated ledger (SC6/SC8) ------------
|
||||
|
||||
_PORTFOLIO_IDS = ["FV42-GSV-E1", "RV13-RAS-TP", "BRU-LAKS-REHAB"]
|
||||
_PORTFOLIO_IDS = ["KONTOR-IT-E1", "NETT-SIKR-TP", "ARKIV-LAGR-MIGR"]
|
||||
|
||||
|
||||
def _prefilled(project_id: str, amount_ore: int) -> SavingsLedger:
|
||||
|
|
@ -147,7 +147,7 @@ def _prefilled(project_id: str, amount_ore: int) -> SavingsLedger:
|
|||
|
||||
|
||||
async def test_portfolio_hard_goal_stops_the_whole_pass(make_portfolio_client_factory) -> None:
|
||||
ledger = _prefilled("FV42-GSV-E1", 100) # portfolio_total == 100
|
||||
ledger = _prefilled("KONTOR-IT-E1", 100) # portfolio_total == 100
|
||||
goals = GoalConfig(portfolio=GoalContract(absolute_ore=100)) # met at 100 (>=)
|
||||
result = await run_portfolio(
|
||||
_PORTFOLIO_IDS,
|
||||
|
|
@ -167,7 +167,7 @@ async def test_portfolio_hard_goal_stops_the_whole_pass(make_portfolio_client_fa
|
|||
|
||||
async def test_boundary_exact_equal_stops(make_portfolio_client_factory) -> None:
|
||||
""">= boundary: accumulated EXACTLY equal to the goal stops (reached, not strictly exceeded)."""
|
||||
ledger = _prefilled("FV42-GSV-E1", 500)
|
||||
ledger = _prefilled("KONTOR-IT-E1", 500)
|
||||
goals = GoalConfig(portfolio=GoalContract(absolute_ore=500))
|
||||
result = await run_portfolio(
|
||||
_PORTFOLIO_IDS,
|
||||
|
|
@ -181,10 +181,10 @@ async def test_boundary_exact_equal_stops(make_portfolio_client_factory) -> None
|
|||
|
||||
|
||||
async def test_per_project_hard_goal_skips_only_that_pid(make_portfolio_client_factory) -> None:
|
||||
ledger = _prefilled("FV42-GSV-E1", 1000)
|
||||
goals = GoalConfig(per_project={"FV42-GSV-E1": GoalContract(absolute_ore=1000)})
|
||||
ledger = _prefilled("KONTOR-IT-E1", 1000)
|
||||
goals = GoalConfig(per_project={"KONTOR-IT-E1": GoalContract(absolute_ore=1000)})
|
||||
result = await run_portfolio(
|
||||
["FV42-GSV-E1", "RV13-RAS-TP"],
|
||||
["KONTOR-IT-E1", "NETT-SIKR-TP"],
|
||||
"local",
|
||||
ledger=ledger,
|
||||
goals=goals,
|
||||
|
|
@ -192,16 +192,16 @@ async def test_per_project_hard_goal_skips_only_that_pid(make_portfolio_client_f
|
|||
max_rounds=1,
|
||||
)
|
||||
ran = [r.outcome.proposal.project_id for r in result.runs]
|
||||
assert "FV42-GSV-E1" not in ran # its goal is reached -> skipped
|
||||
assert ran == ["RV13-RAS-TP"] # the rest of the pass proceeds
|
||||
assert "KONTOR-IT-E1" not in ran # its goal is reached -> skipped
|
||||
assert ran == ["NETT-SIKR-TP"] # the rest of the pass proceeds
|
||||
assert result.stopped_early is False # a per-project skip is NOT a pass-stop
|
||||
|
||||
|
||||
async def test_soft_goal_flags_but_continues(make_portfolio_client_factory) -> None:
|
||||
ledger = _prefilled("FV42-GSV-E1", 100)
|
||||
ledger = _prefilled("KONTOR-IT-E1", 100)
|
||||
goals = GoalConfig(portfolio=GoalContract(absolute_ore=100, mode="soft"))
|
||||
result = await run_portfolio(
|
||||
["FV42-GSV-E1", "RV13-RAS-TP"],
|
||||
["KONTOR-IT-E1", "NETT-SIKR-TP"],
|
||||
"local",
|
||||
ledger=ledger,
|
||||
goals=goals,
|
||||
|
|
@ -219,7 +219,7 @@ async def test_stop_decision_is_deterministic(make_portfolio_client_factory) ->
|
|||
r1 = await run_portfolio(
|
||||
_PORTFOLIO_IDS,
|
||||
"local",
|
||||
ledger=_prefilled("FV42-GSV-E1", 100),
|
||||
ledger=_prefilled("KONTOR-IT-E1", 100),
|
||||
goals=goals,
|
||||
client_factory=make_portfolio_client_factory({}),
|
||||
max_rounds=1,
|
||||
|
|
@ -227,7 +227,7 @@ async def test_stop_decision_is_deterministic(make_portfolio_client_factory) ->
|
|||
r2 = await run_portfolio(
|
||||
_PORTFOLIO_IDS,
|
||||
"local",
|
||||
ledger=_prefilled("FV42-GSV-E1", 100),
|
||||
ledger=_prefilled("KONTOR-IT-E1", 100),
|
||||
goals=goals,
|
||||
client_factory=make_portfolio_client_factory({}),
|
||||
max_rounds=1,
|
||||
|
|
|
|||
|
|
@ -57,7 +57,7 @@ def test_same_candidate_two_dimensions_counts_once_and_flags_overlap() -> None:
|
|||
assert ledger.add_realized(_entry("energi", amount_ore=5000)) is True
|
||||
# Same candidate identity (same codes+measure+amount), a DIFFERENT dimension -> distinct full
|
||||
# key -> stored, so the overlap is visible; the sum must still count it once.
|
||||
assert ledger.add_realized(_entry("asfalt", amount_ore=5000)) is True
|
||||
assert ledger.add_realized(_entry("lisens", amount_ore=5000)) is True
|
||||
|
||||
assert ledger.portfolio_total() == 5000, (
|
||||
"double-counted the same candidate across dimensions (SC4)"
|
||||
|
|
@ -186,14 +186,14 @@ def test_realize_ore_conversion_is_exact() -> None:
|
|||
|
||||
# --- Step 8: goal-stop is load-bearing (SC6) -----------------------------------------------------
|
||||
|
||||
_PORTFOLIO_IDS = ["FV42-GSV-E1", "RV13-RAS-TP", "BRU-LAKS-REHAB"]
|
||||
_PORTFOLIO_IDS = ["KONTOR-IT-E1", "NETT-SIKR-TP", "ARKIV-LAGR-MIGR"]
|
||||
|
||||
|
||||
def _portfolio_ledger(amount_ore: int) -> SavingsLedger:
|
||||
led = SavingsLedger()
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="FV42-GSV-E1",
|
||||
project_id="KONTOR-IT-E1",
|
||||
dimension="energi",
|
||||
candidate_identity="prior-hitl",
|
||||
amount_ore=amount_ore,
|
||||
|
|
|
|||
|
|
@ -26,14 +26,14 @@ from portfolio_optimiser.reference_domain import load_reference_projects
|
|||
from portfolio_optimiser.validator import Rejection, ValidatedProposal
|
||||
|
||||
_VALID = (
|
||||
'{"project_id":"FV42-GSV-E1","measure":"Reduce scope",'
|
||||
'{"project_id":"KONTOR-IT-E1","measure":"Reduce scope",'
|
||||
'"affected_items":[{"code":"05.2","quantity":4300,"unit_cost":215},'
|
||||
'{"code":"03.1","quantity":1800,"unit_cost":310}],"claimed_saving_nok":200000}'
|
||||
)
|
||||
#: Same shape and a value the IR itself accepts, but above the method cap (30% of the ~1.48M
|
||||
#: affected total = ~445k) -> the DETERMINISTIC VALIDATOR is what must reject it, not the parser.
|
||||
_INFEASIBLE = (
|
||||
'{"project_id":"FV42-GSV-E1","measure":"Reduce scope",'
|
||||
'{"project_id":"KONTOR-IT-E1","measure":"Reduce scope",'
|
||||
'"affected_items":[{"code":"05.2","quantity":4300,"unit_cost":215},'
|
||||
'{"code":"03.1","quantity":1800,"unit_cost":310}],"claimed_saving_nok":800000}'
|
||||
)
|
||||
|
|
@ -47,7 +47,7 @@ _APPROACH = Approach(
|
|||
|
||||
@pytest.fixture(scope="module")
|
||||
def project():
|
||||
return load_reference_projects()[0] # FV42-GSV-E1
|
||||
return load_reference_projects()[0] # KONTOR-IT-E1
|
||||
|
||||
|
||||
def _meter() -> TokenMeter:
|
||||
|
|
|
|||
|
|
@ -167,7 +167,7 @@ async def test_portfolio_mode_gives_every_project_the_configured_tools(
|
|||
flag would be accepted and silently dropped in one of the two modes — the defect class this
|
||||
repo refuses by name ('refused, never ignored')."""
|
||||
result = await run_module.run_portfolio(
|
||||
["FV42-GSV-E1", "RV13-RAS-TP"],
|
||||
["KONTOR-IT-E1", "NETT-SIKR-TP"],
|
||||
"local",
|
||||
store=VerdictStore(verdicts=[]),
|
||||
client_factory=_factory,
|
||||
|
|
|
|||
|
|
@ -49,7 +49,7 @@ _ALIGNED_REPLY = (
|
|||
|
||||
|
||||
def _make_docs(tmp_path) -> str:
|
||||
"""A tmp docs folder with citable content, so the road path retrieval is non-empty."""
|
||||
"""A tmp docs folder with citable content, so the reference path retrieval is non-empty."""
|
||||
d = tmp_path / "ore-docs"
|
||||
d.mkdir()
|
||||
(d / "cost.txt").write_text(
|
||||
|
|
|
|||
|
|
@ -51,11 +51,12 @@ from portfolio_optimiser.simulation import ScriptedChatClient
|
|||
from portfolio_optimiser.verdicts import VerdictStore
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
#: Three bases with three DISTINCT project ids — which is what makes "the project comes from the
|
||||
#: base, not from the caller" a claim a test can actually falsify.
|
||||
_BYGG = _EXAMPLES / "bygg-energi-mikro" # BYGG-KONTOR-NORD
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia" # TUNNEL-HAUGLIA
|
||||
_VEGLYS = _EXAMPLES / "veglys-fv-soer" # VEGLYS-FV-SOER
|
||||
_KJOLING = _BUNDLES / "driftssenter-kjoling" # DRIFTSSENTER-KJOLING
|
||||
_KLIENTPARK = _BUNDLES / "klientpark-energi" # KLIENTPARK-ENERGI
|
||||
|
||||
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (sim)"}
|
||||
|
||||
|
|
@ -74,7 +75,10 @@ def test_an_approach_records_which_knowledge_base_it_belongs_to() -> None:
|
|||
one base to name.
|
||||
"""
|
||||
assert Approach(id="a", label="A").bundle_id == ""
|
||||
assert Approach(id="a", label="A", bundle_id="tunnel-hauglia").bundle_id == "tunnel-hauglia"
|
||||
assert (
|
||||
Approach(id="a", label="A", bundle_id="driftssenter-kjoling").bundle_id
|
||||
== "driftssenter-kjoling"
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
|
@ -93,13 +97,13 @@ def test_the_router_partitions_the_mandate_one_sub_mandate_per_named_base() -> N
|
|||
approaches, so the dispatch's spend order is a property of how the run was configured and not
|
||||
of how a model happened to sequence its hypotheses.
|
||||
"""
|
||||
a = Approach(id="a", label="A", bundle_id="tunnel-hauglia")
|
||||
b = Approach(id="b", label="B", bundle_id="veglys-fv-soer")
|
||||
c = Approach(id="c", label="C", bundle_id="tunnel-hauglia")
|
||||
a = Approach(id="a", label="A", bundle_id="driftssenter-kjoling")
|
||||
b = Approach(id="b", label="B", bundle_id="klientpark-energi")
|
||||
c = Approach(id="c", label="C", bundle_id="driftssenter-kjoling")
|
||||
|
||||
routed = route_by_bundle(_mandate(a, b, c), ("veglys-fv-soer", "tunnel-hauglia"))
|
||||
routed = route_by_bundle(_mandate(a, b, c), ("klientpark-energi", "driftssenter-kjoling"))
|
||||
|
||||
assert [bundle_id for bundle_id, _ in routed] == ["veglys-fv-soer", "tunnel-hauglia"]
|
||||
assert [bundle_id for bundle_id, _ in routed] == ["klientpark-energi", "driftssenter-kjoling"]
|
||||
assert [ap.id for ap in routed[0][1].approaches] == ["b"]
|
||||
assert [ap.id for ap in routed[1][1].approaches] == ["a", "c"]
|
||||
# The commission's own fields travel with every partition: each sub-run is still working on the
|
||||
|
|
@ -115,9 +119,9 @@ def test_a_single_base_absorbs_every_unassigned_approach() -> None:
|
|||
``_bundle_index`` already guarantees the id is unique. This is what keeps a single-base
|
||||
mandate (every mandate that exists today) dispatchable unchanged.
|
||||
"""
|
||||
routed = route_by_bundle(_mandate(Approach(id="a", label="A")), ("tunnel-hauglia",))
|
||||
routed = route_by_bundle(_mandate(Approach(id="a", label="A")), ("driftssenter-kjoling",))
|
||||
|
||||
assert [bundle_id for bundle_id, _ in routed] == ["tunnel-hauglia"]
|
||||
assert [bundle_id for bundle_id, _ in routed] == ["driftssenter-kjoling"]
|
||||
assert [ap.id for ap in routed[0][1].approaches] == ["a"]
|
||||
|
||||
|
||||
|
|
@ -130,7 +134,9 @@ def test_an_unassigned_approach_among_several_bases_is_refused_not_guessed() ->
|
|||
report would then describe work nobody ordered.
|
||||
"""
|
||||
with pytest.raises(MandateRoutingError) as exc:
|
||||
route_by_bundle(_mandate(Approach(id="a", label="A")), ("tunnel-hauglia", "veglys-fv-soer"))
|
||||
route_by_bundle(
|
||||
_mandate(Approach(id="a", label="A")), ("driftssenter-kjoling", "klientpark-energi")
|
||||
)
|
||||
|
||||
# The REPR, never the bare letter: "a" is a substring of almost any English sentence, so an
|
||||
# assertion on it would hold against a refusal raised for an entirely different reason.
|
||||
|
|
@ -145,11 +151,11 @@ def test_an_approach_naming_an_unconfigured_base_is_refused_by_name() -> None:
|
|||
"""
|
||||
approach = Approach(id="a", label="A", bundle_id="does-not-exist")
|
||||
with pytest.raises(MandateRoutingError) as exc:
|
||||
route_by_bundle(_mandate(approach), ("tunnel-hauglia", "veglys-fv-soer"))
|
||||
route_by_bundle(_mandate(approach), ("driftssenter-kjoling", "klientpark-energi"))
|
||||
|
||||
message = str(exc.value)
|
||||
assert "does-not-exist" in message
|
||||
assert "tunnel-hauglia" in message
|
||||
assert "driftssenter-kjoling" in message
|
||||
|
||||
|
||||
def test_the_routing_refusal_is_a_value_error() -> None:
|
||||
|
|
@ -279,14 +285,14 @@ async def test_a_marked_hypothesis_may_name_its_base_and_the_mandate_carries_it(
|
|||
result = await explore.explore(
|
||||
_PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(_TUNNEL), str(_VEGLYS)),
|
||||
bundle_dirs=(str(_KJOLING), str(_KLIENTPARK)),
|
||||
client_factory=_factory(
|
||||
ledgers=_two_round_ledgers(),
|
||||
hypothesiser=[_hypothesis_line("LED retrofit", "old fixtures", "veglys-fv-soer")],
|
||||
hypothesiser=[_hypothesis_line("PC-utskifting", "old desktops", "klientpark-energi")],
|
||||
),
|
||||
)
|
||||
|
||||
assert [a.bundle_id for a in result.mandate.approaches] == ["veglys-fv-soer"]
|
||||
assert [a.bundle_id for a in result.mandate.approaches] == ["klientpark-energi"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
|
|
@ -300,14 +306,14 @@ async def test_with_one_base_a_marker_that_names_none_still_yields_an_assigned_a
|
|||
result = await explore.explore(
|
||||
_PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(_TUNNEL),),
|
||||
bundle_dirs=(str(_KJOLING),),
|
||||
client_factory=_factory(
|
||||
ledgers=_two_round_ledgers(),
|
||||
hypothesiser=[_hypothesis_line("LED retrofit", "old fixtures")],
|
||||
hypothesiser=[_hypothesis_line("PC-utskifting", "old desktops")],
|
||||
),
|
||||
)
|
||||
|
||||
assert [a.bundle_id for a in result.mandate.approaches] == ["tunnel-hauglia"]
|
||||
assert [a.bundle_id for a in result.mandate.approaches] == ["driftssenter-kjoling"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
|
|
@ -322,10 +328,10 @@ async def test_with_several_bases_a_marker_that_names_none_is_refused() -> None:
|
|||
await explore.explore(
|
||||
_PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(_TUNNEL), str(_VEGLYS)),
|
||||
bundle_dirs=(str(_KJOLING), str(_KLIENTPARK)),
|
||||
client_factory=_factory(
|
||||
ledgers=_two_round_ledgers(),
|
||||
hypothesiser=[_hypothesis_line("LED retrofit", "old fixtures")],
|
||||
hypothesiser=[_hypothesis_line("PC-utskifting", "old desktops")],
|
||||
),
|
||||
)
|
||||
|
||||
|
|
@ -341,7 +347,7 @@ async def test_a_marker_naming_an_unconfigured_base_is_refused() -> None:
|
|||
await explore.explore(
|
||||
_PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(_TUNNEL), str(_VEGLYS)),
|
||||
bundle_dirs=(str(_KJOLING), str(_KLIENTPARK)),
|
||||
client_factory=_factory(
|
||||
ledgers=_two_round_ledgers(),
|
||||
hypothesiser=[_hypothesis_line("LED", "old fixtures", "no-such-base")],
|
||||
|
|
@ -363,7 +369,7 @@ async def test_a_seed_naming_an_unknown_base_is_refused_before_the_first_model_c
|
|||
await explore.explore(
|
||||
_PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(_TUNNEL), str(_VEGLYS)),
|
||||
bundle_dirs=(str(_KJOLING), str(_KLIENTPARK)),
|
||||
seed_approaches=(Approach(id="s1", label="Seed", bundle_id="no-such-base"),),
|
||||
client_factory=_factory(ledgers=_two_round_ledgers(), hypothesiser=[], sink=sink),
|
||||
)
|
||||
|
|
@ -384,7 +390,7 @@ async def test_an_unassigned_seed_with_several_bases_is_refused_before_the_first
|
|||
await explore.explore(
|
||||
_PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(_TUNNEL), str(_VEGLYS)),
|
||||
bundle_dirs=(str(_KJOLING), str(_KLIENTPARK)),
|
||||
seed_approaches=(Approach(id="s1", label="Seed"),),
|
||||
client_factory=_factory(ledgers=_two_round_ledgers(), hypothesiser=[], sink=sink),
|
||||
)
|
||||
|
|
@ -406,7 +412,7 @@ async def test_a_seed_is_never_rewritten_only_validated() -> None:
|
|||
result = await explore.explore(
|
||||
_PROMPT,
|
||||
contract=_CONTRACT,
|
||||
bundle_dirs=(str(_TUNNEL),),
|
||||
bundle_dirs=(str(_KJOLING),),
|
||||
seed_approaches=(seed,),
|
||||
client_factory=_factory(ledgers=_two_round_ledgers(), hypothesiser=[]),
|
||||
)
|
||||
|
|
@ -465,18 +471,18 @@ async def test_the_dispatch_runs_one_pipeline_per_base_with_that_bases_approache
|
|||
monkeypatch.setattr(run_module, "run_project", _recorder(calls))
|
||||
|
||||
mandate = _mandate(
|
||||
Approach(id="a", label="A", bundle_id="tunnel-hauglia"),
|
||||
Approach(id="b", label="B", bundle_id="veglys-fv-soer"),
|
||||
Approach(id="a", label="A", bundle_id="driftssenter-kjoling"),
|
||||
Approach(id="b", label="B", bundle_id="klientpark-energi"),
|
||||
)
|
||||
await run_mandate_across_bundles(
|
||||
mandate,
|
||||
(str(_TUNNEL), str(_VEGLYS)),
|
||||
(str(_KJOLING), str(_KLIENTPARK)),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
store=VerdictStore(verdicts=[]),
|
||||
)
|
||||
|
||||
assert len(calls) == 2
|
||||
assert [c["bundle_dir"] for c in calls] == [str(_TUNNEL), str(_VEGLYS)]
|
||||
assert [c["bundle_dir"] for c in calls] == [str(_KJOLING), str(_KLIENTPARK)]
|
||||
assert [[ap.id for ap in c["mandate"].approaches] for c in calls] == [["a"], ["b"]]
|
||||
|
||||
|
||||
|
|
@ -495,17 +501,17 @@ async def test_each_bases_project_id_comes_from_that_base_not_from_the_caller(
|
|||
monkeypatch.setattr(run_module, "run_project", _recorder(calls))
|
||||
|
||||
mandate = _mandate(
|
||||
Approach(id="a", label="A", bundle_id="tunnel-hauglia"),
|
||||
Approach(id="b", label="B", bundle_id="veglys-fv-soer"),
|
||||
Approach(id="a", label="A", bundle_id="driftssenter-kjoling"),
|
||||
Approach(id="b", label="B", bundle_id="klientpark-energi"),
|
||||
)
|
||||
await run_mandate_across_bundles(
|
||||
mandate,
|
||||
(str(_TUNNEL), str(_VEGLYS)),
|
||||
(str(_KJOLING), str(_KLIENTPARK)),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
store=VerdictStore(verdicts=[]),
|
||||
)
|
||||
|
||||
assert [c["project_id"] for c in calls] == ["TUNNEL-HAUGLIA", "VEGLYS-FV-SOER"]
|
||||
assert [c["project_id"] for c in calls] == ["DRIFTSSENTER-KJOLING", "KLIENTPARK-ENERGI"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
|
|
@ -523,7 +529,7 @@ async def test_one_base_is_dispatched_exactly_as_a_single_run(
|
|||
|
||||
await run_mandate_across_bundles(
|
||||
_mandate(Approach(id="a", label="A"), Approach(id="b", label="B")),
|
||||
(str(_TUNNEL),),
|
||||
(str(_KJOLING),),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
store=VerdictStore(verdicts=[]),
|
||||
)
|
||||
|
|
@ -549,10 +555,10 @@ async def test_a_base_that_cannot_be_funded_is_never_started_and_its_approaches_
|
|||
|
||||
result = await run_mandate_across_bundles(
|
||||
_mandate(
|
||||
Approach(id="a", label="A", bundle_id="tunnel-hauglia"),
|
||||
Approach(id="b", label="B", bundle_id="veglys-fv-soer"),
|
||||
Approach(id="a", label="A", bundle_id="driftssenter-kjoling"),
|
||||
Approach(id="b", label="B", bundle_id="klientpark-energi"),
|
||||
),
|
||||
(str(_TUNNEL), str(_VEGLYS)),
|
||||
(str(_KJOLING), str(_KLIENTPARK)),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
store=VerdictStore(verdicts=[]),
|
||||
portfolio_meter=meter,
|
||||
|
|
@ -560,7 +566,7 @@ async def test_a_base_that_cannot_be_funded_is_never_started_and_its_approaches_
|
|||
|
||||
# Base 1 is funded (1000 left, 500 required) and spends 600; base 2 then has 400 left against a
|
||||
# 500 reserve. The SECOND base is the one that must never start.
|
||||
assert [c["bundle_dir"] for c in calls] == [str(_TUNNEL)]
|
||||
assert [c["bundle_dir"] for c in calls] == [str(_KJOLING)]
|
||||
assert result.stopped_early is True
|
||||
assert result.budget_stop is not None
|
||||
assert {row.id for row in result.unreached} == {"b"}
|
||||
|
|
@ -579,7 +585,7 @@ async def test_a_pass_that_can_fund_nothing_at_all_is_refused_at_startup() -> No
|
|||
with pytest.raises(BudgetRefused):
|
||||
await run_mandate_across_bundles(
|
||||
_mandate(Approach(id="a", label="A")),
|
||||
(str(_TUNNEL),),
|
||||
(str(_KJOLING),),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
store=VerdictStore(verdicts=[]),
|
||||
portfolio_meter=meter,
|
||||
|
|
@ -606,7 +612,7 @@ async def test_the_dispatch_composes_with_the_real_run_project(tmp_path: Path) -
|
|||
Driven against two real bases end to end, offline, through the scripted client seam.
|
||||
"""
|
||||
bases = []
|
||||
for src in (_BYGG, _TUNNEL):
|
||||
for src in (_BYGG, _KJOLING):
|
||||
dst = tmp_path / src.name
|
||||
shutil.copytree(src, dst)
|
||||
bases.append(str(dst))
|
||||
|
|
@ -620,7 +626,7 @@ async def test_the_dispatch_composes_with_the_real_run_project(tmp_path: Path) -
|
|||
result = await run_mandate_across_bundles(
|
||||
_mandate(
|
||||
Approach(id="a", label="A", bundle_id="bygg-energi-mikro"),
|
||||
Approach(id="b", label="B", bundle_id="tunnel-hauglia"),
|
||||
Approach(id="b", label="B", bundle_id="driftssenter-kjoling"),
|
||||
),
|
||||
tuple(bases),
|
||||
"local",
|
||||
|
|
@ -630,8 +636,8 @@ async def test_the_dispatch_composes_with_the_real_run_project(tmp_path: Path) -
|
|||
max_rounds=1,
|
||||
)
|
||||
|
||||
assert [r.bundle_id for r in result.runs] == ["bygg-energi-mikro", "tunnel-hauglia"]
|
||||
assert [r.project_id for r in result.runs] == ["BYGG-KONTOR-NORD", "TUNNEL-HAUGLIA"]
|
||||
assert [r.bundle_id for r in result.runs] == ["bygg-energi-mikro", "driftssenter-kjoling"]
|
||||
assert [r.project_id for r in result.runs] == ["BYGG-KONTOR-NORD", "DRIFTSSENTER-KJOLING"]
|
||||
assert all(isinstance(r.result, RunResult) for r in result.runs)
|
||||
# Each run answered for ITS OWN commissioned approach, and for nobody else's.
|
||||
assert [{row.id for row in r.result.coverage} for r in result.runs] == [
|
||||
|
|
@ -662,10 +668,10 @@ async def test_one_store_is_threaded_across_every_base(
|
|||
store = VerdictStore(verdicts=[])
|
||||
await run_mandate_across_bundles(
|
||||
_mandate(
|
||||
Approach(id="a", label="A", bundle_id="tunnel-hauglia"),
|
||||
Approach(id="b", label="B", bundle_id="veglys-fv-soer"),
|
||||
Approach(id="a", label="A", bundle_id="driftssenter-kjoling"),
|
||||
Approach(id="b", label="B", bundle_id="klientpark-energi"),
|
||||
),
|
||||
(str(_TUNNEL), str(_VEGLYS)),
|
||||
(str(_KJOLING), str(_KLIENTPARK)),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
store=store,
|
||||
)
|
||||
|
|
@ -688,7 +694,7 @@ async def test_exploration_and_dispatch_close_the_loop_over_two_bases(tmp_path:
|
|||
expert actually reads: an approach that reached the wrong base would still appear evaluated.
|
||||
"""
|
||||
bases = []
|
||||
for src in (_BYGG, _TUNNEL):
|
||||
for src in (_BYGG, _KJOLING):
|
||||
dst = tmp_path / src.name
|
||||
shutil.copytree(src, dst)
|
||||
bases.append(str(dst))
|
||||
|
|
@ -705,14 +711,18 @@ async def test_exploration_and_dispatch_close_the_loop_over_two_bases(tmp_path:
|
|||
hypothesiser=[
|
||||
_hypothesis_line("Behovsstyrt lys", "fixtures are 1990s", "bygg-energi-mikro")
|
||||
+ "\n"
|
||||
+ _hypothesis_line("Nattsenking", "the tunnel runs lit all night", "tunnel-hauglia")
|
||||
+ _hypothesis_line(
|
||||
"Nattsenking",
|
||||
"the cooling runs at full level all night",
|
||||
"driftssenter-kjoling",
|
||||
)
|
||||
],
|
||||
),
|
||||
)
|
||||
|
||||
assert [a.bundle_id for a in exploration.mandate.approaches] == [
|
||||
"bygg-energi-mikro",
|
||||
"tunnel-hauglia",
|
||||
"driftssenter-kjoling",
|
||||
]
|
||||
|
||||
def factory(role: str) -> BaseChatClient:
|
||||
|
|
@ -731,7 +741,7 @@ async def test_exploration_and_dispatch_close_the_loop_over_two_bases(tmp_path:
|
|||
max_rounds=1,
|
||||
)
|
||||
|
||||
assert [r.bundle_id for r in result.runs] == ["bygg-energi-mikro", "tunnel-hauglia"]
|
||||
assert [r.bundle_id for r in result.runs] == ["bygg-energi-mikro", "driftssenter-kjoling"]
|
||||
assert [{row.id for row in r.result.coverage} for r in result.runs] == [
|
||||
{"hypothesis-1", "own-proposal"},
|
||||
{"hypothesis-2", "own-proposal"},
|
||||
|
|
|
|||
|
|
@ -2,8 +2,8 @@
|
|||
|
||||
MEASURED (P21 stress round 5, re-measured at the head of økt 126): all 26 rejections across the
|
||||
six paid runs read ``unknown cost code '<invention>': not in project P's cost baseline (5 known
|
||||
codes)``, and **0 of 20** approaches validated (round 4: 4 of 20). The model invented
|
||||
``signalregulering_konstruksjon``, ``VENTIL_IMP``, ``RIGG01``, ``bærelag_asfalt`` and 22 more. It
|
||||
codes)``, and **0 of 20** approaches validated (round 4: 4 of 20). The model invented 26 names
|
||||
such as ``RIGG01`` and ``VENTIL_IMP``. It
|
||||
could not have done otherwise: the price schedule reaches the VALIDATOR and never the proposer,
|
||||
and the refusal states the COUNT of known codes, not one code's name. Step 5 feeds that sentence
|
||||
verbatim into the next attempt's prompt — and "you guessed wrong, there are five right answers"
|
||||
|
|
@ -20,11 +20,12 @@ sever a code mid-name and hand the proposer an identifier that exists nowhere
|
|||
window can only ever emit whole codes.
|
||||
|
||||
Measured sizes, with denominators (økt 126): every cost baseline anywhere in this repo or its
|
||||
measured corpora is at most SIX codes (the five context sets: 5 · 5 · 5 · 5 · 6; the two shipped
|
||||
measured corpora was at most SIX codes (the five context sets of the time: 5 · 5 · 5 · 5 · 6; the
|
||||
two shipped
|
||||
``shared/examples`` baselines: 1 each; MAJOR-4's derivation of the synthetic K2 price schedule: 3),
|
||||
and the largest REAL delivered price schedule measured is K2's ``prissammenstilling-sheet-1.md`` at
|
||||
14 priced rows of 118 lines. Nothing measured reaches the window; it exists for the unmeasured
|
||||
R761-style mengdebeskrivelse, where a contract is priced BY prosesskode and the corpus declares
|
||||
schedule of an agreement priced BY process number, where the catalogue measured then declared
|
||||
2 727 of them.
|
||||
"""
|
||||
|
||||
|
|
@ -48,7 +49,7 @@ from portfolio_optimiser.validator import (
|
|||
validate_proposal,
|
||||
)
|
||||
|
||||
_SMALL: Final = ("RIGG", "ASFALT", "MASSE", "FROST", "SKILT")
|
||||
_SMALL: Final = ("RIGG", "LISENS", "MASSE", "FROST", "SKILT")
|
||||
|
||||
|
||||
def _baseline(codes: tuple[str, ...], project_id: str = "proj") -> CostBaseline:
|
||||
|
|
@ -112,7 +113,7 @@ def test_the_magnitude_half_does_not_grow_a_code_list() -> None:
|
|||
|
||||
|
||||
def test_a_large_schedule_is_bound_by_a_fixed_window() -> None:
|
||||
"""A mengdebeskrivelse priced by prosesskode can carry thousands of lines. The list is bound at
|
||||
"""A schedule priced by process number can carry thousands of lines. The list is bound at
|
||||
``_KNOWN_CODE_WINDOW`` codes, the denominator still states how many exist, and the cut is
|
||||
ANNOUNCED rather than left for the reader to subtract (``index_truncated``'s rule)."""
|
||||
many = tuple(f"P{i:04d}" for i in range(500))
|
||||
|
|
|
|||
|
|
@ -34,7 +34,7 @@ Two teeth:
|
|||
the honest reading is the opposite: an EMPTY tuple is a positive statement — "every cross-link was
|
||||
followed" — in the same class as ``ProvenanceStamp.external_calls`` ("nothing outside this process
|
||||
was contacted"). A caller that constructs a ``Bundle`` without a trace is not withholding a fact; it
|
||||
is stating one. So ``skipped`` defaults to ``()``, and the road path (which navigates nothing) is
|
||||
is stating one. So ``skipped`` defaults to ``()``, and the reference path (which navigates nothing) is
|
||||
honestly empty rather than dishonestly required to invent a value.
|
||||
|
||||
Arms:
|
||||
|
|
@ -250,10 +250,10 @@ async def test_run_result_carries_the_trace(tmp_path, fresh_store) -> None:
|
|||
|
||||
|
||||
async def test_road_path_has_an_empty_trace(docs_dir, fresh_store) -> None:
|
||||
"""The road path navigates no bundle, so "nothing was skipped" is literally true there — which
|
||||
"""The reference path navigates no bundle, so "nothing was skipped" is literally true there — which
|
||||
is exactly why the empty tuple is an honest DEFAULT rather than a withheld fact."""
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
|
|
|
|||
|
|
@ -1,12 +1,13 @@
|
|||
"""P18/A — a listing is a WINDOW, and a path the caller invented is refused by name.
|
||||
|
||||
P16 (``docs/2026-09-14-p16-stressrunde-1.md``) measured the ladder S7a-3 built against a DELIVERED
|
||||
corpus for the first time, and found two things the fixture bases could not show:
|
||||
P16 (the first stress round; the ledger is ``docs/invarianter.md``) measured the ladder S7a-3
|
||||
built against a DELIVERED corpus for the first time, and found two things the fixture bases could
|
||||
not show:
|
||||
|
||||
* **one level is not bounded by being one level.** ``okf.directory_listing`` on ``krav/N100`` was
|
||||
69 250 characters over 445 documents, ``krav/N200`` 169 974 over 1 132, and R761's own root
|
||||
110 912 over 2 728 SUBDIRECTORIES — 27-113x the 1 500-character ceiling S7a-3 set, riding in
|
||||
every later prompt. Three of seven paid runs died on the token cap.
|
||||
* **one level is not bounded by being one level.** ``okf.directory_listing`` on one requirements
|
||||
level was 69 250 characters over 445 documents, on another 169 974 over 1 132, and a process
|
||||
catalogue's own root 110 912 over 2 728 SUBDIRECTORIES — 27-113x the 1 500-character ceiling
|
||||
S7a-3 set, riding in every later prompt. Three of seven paid runs died on the token cap.
|
||||
* **0 of 26 fasit concepts were opened** in 24 ``read_file`` calls, and 10 of those calls named a
|
||||
path the base does not hold. Each reached the model as MAF's opaque ``"Error: Function failed."``
|
||||
(``agent_framework/_tools.py:1410-1432``) while counting toward the three consecutive tool errors
|
||||
|
|
@ -24,11 +25,11 @@ Three seams, each with its own arms below:
|
|||
nearest directory that actually holds documents. Narrowness is arm (h) and lives in
|
||||
``test_read_file_directory_refusal_loadbearing.py`` alongside the tripwire it replaced.
|
||||
|
||||
The bases this measures are the delivered vegnormal corpora (``PORTFOLIO_VEGNORMAL_ROOT``): arms
|
||||
that need them SKIP with the root named when it is not mounted, exactly as MAJOR-3's ceiling arm
|
||||
does — a hard failure would break ``uv run pytest`` inside the handover package. The arms that do
|
||||
NOT need them (the window's own algebra, the filter's negative, the refusal) are UNCONDITIONAL and
|
||||
run over a synthetic base, so this file can never be silently absent in full.
|
||||
The bases this measures are the package's two example bases (``data/kunnskapsbaser/``, pinned
|
||||
in ``frozen_bundles.json``), so the arms that need them run wherever the package does; they SKIP,
|
||||
with the store named, only when a user's own store (``PORTFOLIO_FROZEN_BUNDLES``) lacks them. The
|
||||
arms that do NOT need them (the window's own algebra, the filter's negative, the refusal) are
|
||||
UNCONDITIONAL and run over a synthetic base, so this file can never be silently absent in full.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -41,6 +42,8 @@ import pytest
|
|||
|
||||
from portfolio_optimiser import frozen_bundles, okf
|
||||
from portfolio_optimiser.explore import navigator_tools
|
||||
from portfolio_optimiser.mandate import load_mandate
|
||||
from portfolio_optimiser.stress import read_bundle_declarations
|
||||
|
||||
#: S7a-3's ceiling, restated here rather than imported: a gate that imported the implementation's
|
||||
#: own budget would move with it, and raising the budget is exactly the regression it guards.
|
||||
|
|
@ -50,8 +53,8 @@ _CEILING_CHARS = 1_500
|
|||
def _delivered(name: str) -> Path:
|
||||
"""The FROZEN copy this repository pins, resolved at call time.
|
||||
|
||||
Absence SKIPS (MAJOR-3's ceiling: no corpus is mounted in the handover archive), drift is
|
||||
allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
Absence SKIPS (a user's own store, named by ``PORTFOLIO_FROZEN_BUNDLES``, may not hold it),
|
||||
drift is allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
"""
|
||||
try:
|
||||
return frozen_bundles.bundle_dir(name)
|
||||
|
|
@ -105,36 +108,35 @@ def _synthetic(tmp_path: Path, *, dirs: int, per_dir: int) -> Path:
|
|||
@pytest.mark.parametrize(
|
||||
("name", "level", "before"),
|
||||
[
|
||||
("n100-2023", "krav/N100", 69_250),
|
||||
("n200-2024", "krav/N200", 169_974),
|
||||
("n500-2024", "krav/N500", 39_853),
|
||||
("r761-2025", "R761", 110_912),
|
||||
("driftskrav-2027", "krav/D200", 29_453),
|
||||
("prosesskatalog-2027", "P900", 12_345),
|
||||
],
|
||||
)
|
||||
def test_a_default_listing_of_a_delivered_level_is_bounded(
|
||||
name: str, level: str, before: int
|
||||
) -> None:
|
||||
"""(a) The order binds this by name on ``krav/N100`` (445 blades); the other three are the
|
||||
other levels P16 measured, including R761's root, whose 110 912 characters were DIRECTORIES —
|
||||
which is why the window covers both kinds and not only documents.
|
||||
"""(a) The two example levels that are far wider than one window: the requirements level
|
||||
``krav/D200`` (220 documents) and the process catalogue's root ``P900``, whose 300 entries are
|
||||
DIRECTORIES — which is why the window covers both kinds and not only documents.
|
||||
|
||||
``before`` is P16's measured cost of the SAME call, carried so the arm cannot pass by the level
|
||||
``before`` is what reading the WHOLE level costs through the window (every page at the maximum
|
||||
limit, summed), measured on the pinned base and carried so the arm cannot pass by the level
|
||||
having shrunk. The control below proves the fixture is the thing that did not fit."""
|
||||
base = _delivered(name)
|
||||
listing = _read_dir(base, level)
|
||||
|
||||
assert _chars(listing) <= _CEILING_CHARS or name == "n200-2024", (
|
||||
assert _chars(listing) <= _CEILING_CHARS, (
|
||||
f"{name}/{level} default listing is {_chars(listing)} chars, over {_CEILING_CHARS}"
|
||||
)
|
||||
# n200 carries the longest titles measured (entries up to 209 chars), so ten of them land 2.5 %
|
||||
# over. Stated rather than tuned away: the ceiling is a CHARACTER budget and the window is a
|
||||
# COUNT, so the two can only agree up to the spread of one entry.
|
||||
assert _chars(listing) <= 1_600
|
||||
assert _chars(listing) * 20 < before, "the bound must be a fall, not a rounding"
|
||||
assert int(listing["total"]) > 10 * int(listing["limit"]), (
|
||||
"the CONTROL is inert: this level must hold far more than one window, or the bound above "
|
||||
"is measuring a small level rather than the window"
|
||||
)
|
||||
whole = 0
|
||||
for offset in range(0, int(listing["total"]), okf._DIRECTORY_PAGE_MAX):
|
||||
whole += _chars(_read_dir(base, level, offset=offset, limit=okf._DIRECTORY_PAGE_MAX))
|
||||
assert whole == before, f"{name}/{level}: the whole level now costs {whole}, pinned {before}"
|
||||
|
||||
|
||||
def test_the_window_pages_the_whole_level_exactly_once(tmp_path: Path) -> None:
|
||||
|
|
@ -184,19 +186,19 @@ def test_an_offset_past_the_end_is_an_empty_window_over_an_honest_total(tmp_path
|
|||
|
||||
|
||||
def test_the_known_positive_filter_finds_the_fasit_concepts_and_not_the_level() -> None:
|
||||
"""(e) The order's own known positive: ``filter="rundkjoring"`` on n100 must answer with the
|
||||
two concepts gate-nordvik/a1 must cite, and NOT with 445 rows."""
|
||||
base = _delivered("n100-2023")
|
||||
fasit = json.loads(Path("contexts/gate-nordvik-2027/fasit.json").read_text(encoding="utf-8"))
|
||||
"""(e) The known positive: ``filter="autonomitid"`` on ``krav/D200`` must answer with the
|
||||
two concepts serverrom/a2 must cite, and NOT with 220 rows."""
|
||||
base = _delivered("driftskrav-2027")
|
||||
fasit = json.loads(Path("contexts/serverrom-2027/fasit.json").read_text(encoding="utf-8"))
|
||||
wanted = {
|
||||
c["path"]
|
||||
for m in fasit["must_cite"]
|
||||
if m["approach_id"] == "a1-rundkjoring-forenklet"
|
||||
if m["approach_id"] == "a2-ups-autonomi"
|
||||
for c in m["concepts"]
|
||||
}
|
||||
assert wanted, "the fixture must name concepts, or this arm proves nothing"
|
||||
|
||||
hits = _read_dir(base, "krav/N100", filter="rundkjøring", limit=50)
|
||||
hits = _read_dir(base, "krav/D200", filter="autonomitid", limit=50)
|
||||
|
||||
assert wanted <= {str(d["name"]) for d in hits["documents"]}
|
||||
assert int(hits["total_matches"]) < int(hits["total"]) / 50
|
||||
|
|
@ -216,8 +218,9 @@ def test_a_filter_that_matches_nothing_is_an_answer_and_not_a_refusal(tmp_path:
|
|||
|
||||
|
||||
def test_the_filter_reads_the_reference_number_and_not_only_the_title(tmp_path: Path) -> None:
|
||||
"""(g) The second field is load-bearing: on R761 the thing a navigator knows is the process
|
||||
number, which is not in the title. Written over a synthetic base so it is unconditional."""
|
||||
"""(g) The second field is load-bearing: on a process catalogue the thing a navigator knows
|
||||
is the process number, which is not in the title. Written over a synthetic base so it is
|
||||
unconditional."""
|
||||
base = _synthetic(tmp_path, dirs=2, per_dir=3)
|
||||
by_ref = _read_dir(base, "seksjon-01", filter="krav 1.2")
|
||||
|
||||
|
|
@ -226,8 +229,9 @@ def test_the_filter_reads_the_reference_number_and_not_only_the_title(tmp_path:
|
|||
|
||||
|
||||
def test_a_filter_narrows_directories_too(tmp_path: Path) -> None:
|
||||
"""(h) R761's root is 2 728 DIRECTORIES: a filter that only narrowed documents would leave the
|
||||
biggest measured level unnarrowable."""
|
||||
"""(h) A process catalogue's root is all DIRECTORIES (P900's: 300; the largest ever measured
|
||||
here: 2 728): a filter that only narrowed documents would leave the biggest level
|
||||
unnarrowable."""
|
||||
base = _synthetic(tmp_path, dirs=12, per_dir=2)
|
||||
narrowed = _read_dir(base, "", filter="seksjon-0")
|
||||
|
||||
|
|
@ -242,35 +246,47 @@ def test_an_invented_path_is_refused_by_name_over_a_delivered_base() -> None:
|
|||
"""(i) The measured live shape: a one-character slip in a UUID. Before, this reached the model
|
||||
as ``"Error: Function failed."``; the refusal now names the path AND the level that holds
|
||||
documents, which is the only thing the caller can act on."""
|
||||
base = _delivered("n100-2023")
|
||||
slip = "krav/N100/id-d2ebe771-5216-4d7f-92d2-95a31f2b2702.md"
|
||||
base = _delivered("driftskrav-2027")
|
||||
real = "krav/D200/id-b1e2ba25-9825-57b0-a358-569be457d2c8.md"
|
||||
slip = _slip(real)
|
||||
assert (base / real).is_file() and not (base / slip).exists(), "control: one real, one invented"
|
||||
answer = _read_file(base, slip)
|
||||
|
||||
assert answer.startswith(f"REFUSED ({okf.BundlePathNotFound.__name__})")
|
||||
assert slip in answer and "krav/N100" in answer and "read_dir" in answer
|
||||
assert slip in answer and "krav/D200" in answer and "read_dir" in answer
|
||||
# The named level resolves, and it is the one the caller was already in.
|
||||
assert int(_read_dir(base, "krav/N100")["total"]) > 0
|
||||
assert int(_read_dir(base, "krav/D200")["total"]) > 0
|
||||
|
||||
|
||||
def test_all_ten_of_p16s_unresolvable_calls_now_answer_instead_of_failing() -> None:
|
||||
"""(j) The nevner arm (ansikt 4). P16's four ``-debate.json`` artefacts ARE the population, and
|
||||
the denominator is re-measured here rather than quoted: 32 ``read_file`` calls (the order says
|
||||
24 — measured 14.09, that number is the four runs' DISTINCT paths, not their calls), of which
|
||||
10 named a path the base does not hold.
|
||||
def _slip(path: str) -> str:
|
||||
"""The measured live shape of an invented path: ONE character of the document id changed."""
|
||||
stem, suffix = path[:-3], path[-3:]
|
||||
return stem[:-1] + ("b" if stem[-1] == "a" else "a") + suffix
|
||||
|
||||
Every one is replayed. Each of the 10 must now come back as a NAMED refusal, and the other 22
|
||||
must still return the document — a gate that refused everything would pass the first half on
|
||||
its own, which is the failure this file's own A2 negative arm is written against."""
|
||||
artefacts = sorted(Path("scratchpad/p14-stress").glob("*/*-debate.json"))
|
||||
if not artefacts:
|
||||
pytest.skip("P16's stress artefacts are not present in this checkout")
|
||||
|
||||
calls: list[tuple[str, str]] = [
|
||||
(call["bundle_id"].removeprefix("vegnormal-"), call["path"])
|
||||
for artefact in artefacts
|
||||
for call in json.loads(artefact.read_text(encoding="utf-8"))["tool_calls"]
|
||||
if call["name"] == "read_file"
|
||||
]
|
||||
def test_every_invented_path_answers_by_name_and_every_real_one_is_served() -> None:
|
||||
"""(j) The nevner arm (ansikt 4). P16 replayed its own population — 32 ``read_file`` calls from
|
||||
four paid runs, 10 of them naming a path the base does not hold — but those runs read corpora
|
||||
this repository no longer carries, so the population here is CONSTRUCTED from the three context
|
||||
sets' own fasit: every cited path (routed at the base its approach names), and the same path
|
||||
with one id character changed. The denominator is re-measured from the fasit, not quoted.
|
||||
|
||||
Every invented path must come back as a NAMED refusal, and every real one must still return
|
||||
the document — a gate that refused everything would pass the first half on its own, which is
|
||||
the failure this file's own A2 negative arm is written against."""
|
||||
calls: list[tuple[str, str]] = []
|
||||
for fasit_path in sorted(Path("contexts").glob("*/fasit.json")):
|
||||
set_dir = fasit_path.parent
|
||||
names = {
|
||||
b["bundle_id"]: b["name"] for b in read_bundle_declarations(set_dir / "bundle.txt")
|
||||
}
|
||||
routed = {
|
||||
a.id: names[a.bundle_id] for a in load_mandate(set_dir / "mandate.json").approaches
|
||||
}
|
||||
for row in json.loads(fasit_path.read_text(encoding="utf-8"))["must_cite"]:
|
||||
for concept in row["concepts"]:
|
||||
calls.append((routed[row["approach_id"]], concept["path"]))
|
||||
calls.append((routed[row["approach_id"]], _slip(concept["path"])))
|
||||
refused = served = 0
|
||||
cache: dict[str, Any] = {}
|
||||
for name, path in calls:
|
||||
|
|
@ -285,4 +301,4 @@ def test_all_ten_of_p16s_unresolvable_calls_now_answer_instead_of_failing() -> N
|
|||
else:
|
||||
served += 1
|
||||
|
||||
assert (refused, served) == (10, 22), f"population moved: {refused} refused, {served} served"
|
||||
assert (refused, served) == (18, 18), f"population moved: {refused} refused, {served} served"
|
||||
|
|
|
|||
|
|
@ -209,7 +209,7 @@ def _dimension_bundle(tmp_path) -> str:
|
|||
(tmp_path / "index.md").write_text(
|
||||
"---\ntype: index\n---\n\n# Bundle\n\n"
|
||||
"- [energi](energi-method.md)\n"
|
||||
"- [asfalt](asfalt-method.md)\n"
|
||||
"- [lisens](lisens-method.md)\n"
|
||||
"- [shared](shared-note.md)\n"
|
||||
"- [verdict](verdict-x.md)\n",
|
||||
encoding="utf-8",
|
||||
|
|
@ -217,8 +217,8 @@ def _dimension_bundle(tmp_path) -> str:
|
|||
(tmp_path / "energi-method.md").write_text(
|
||||
"---\ntype: methodology\ndimension: energi\n---\n\nENERGI-SENTINEL body\n", encoding="utf-8"
|
||||
)
|
||||
(tmp_path / "asfalt-method.md").write_text(
|
||||
"---\ntype: methodology\ndimension: asfalt\n---\n\nASFALT-SENTINEL body\n", encoding="utf-8"
|
||||
(tmp_path / "lisens-method.md").write_text(
|
||||
"---\ntype: methodology\ndimension: lisens\n---\n\nLISENS-SENTINEL body\n", encoding="utf-8"
|
||||
)
|
||||
(tmp_path / "shared-note.md").write_text(
|
||||
"---\ntype: reference\n---\n\nSHARED-SENTINEL body\n", encoding="utf-8"
|
||||
|
|
@ -231,7 +231,7 @@ def _dimension_bundle(tmp_path) -> str:
|
|||
|
||||
def test_bundle_context_dimension_filter(tmp_path) -> None:
|
||||
"""SC7 forutsetning: with ``dimension="energi"`` only energi-marked + unmarked concept files
|
||||
render; an asfalt-marked file is omitted. ``dimension=None`` is byte-identical to the no-arg
|
||||
render; a lisens-marked file is omitted. ``dimension=None`` is byte-identical to the no-arg
|
||||
call (backward compat — protects the verdict-exclusion + step7/8 load-bearing tests).
|
||||
``type: verdict`` stays excluded in every case."""
|
||||
bundle = okf.navigate_bundle(_dimension_bundle(tmp_path))
|
||||
|
|
@ -239,12 +239,12 @@ def test_bundle_context_dimension_filter(tmp_path) -> None:
|
|||
scoped = okf.bundle_context(bundle, dimension="energi")
|
||||
assert "ENERGI-SENTINEL" in scoped # energi-marked concept file rendered
|
||||
assert "SHARED-SENTINEL" in scoped # unmarked knowledge is never dropped
|
||||
assert "ASFALT-SENTINEL" not in scoped # other-dimension file filtered out
|
||||
assert "LISENS-SENTINEL" not in scoped # other-dimension file filtered out
|
||||
assert "VERDICT-SENTINEL" not in scoped # verdict layer still excluded
|
||||
|
||||
default = okf.bundle_context(bundle)
|
||||
assert okf.bundle_context(bundle, dimension=None) == default # None == today, byte-identical
|
||||
assert "ASFALT-SENTINEL" in default # no filter -> asfalt present
|
||||
assert "LISENS-SENTINEL" in default # no filter -> lisens present
|
||||
assert "VERDICT-SENTINEL" not in default # verdict still excluded
|
||||
|
||||
|
||||
|
|
@ -317,13 +317,13 @@ def test_parse_frontmatter_reads_scalar_fields() -> None:
|
|||
|
||||
def test_parse_frontmatter_top_level_title_survives_nested_sources_title(tmp_path) -> None:
|
||||
"""A concept's OWN ``title`` sits at top level; ``sources:`` is a block sequence whose nested
|
||||
``title:`` names the SOURCE document, not the concept (measured on a real vegnormal-okf
|
||||
concept, ``krav/N500/id-bfb0edb4-…``). Top-level keys carry no indentation and must win —
|
||||
a nested line must never overwrite a top-level key of the same name, however late it appears
|
||||
in the scan. Without this, ``directory_listing`` on ``krav/N500`` returns 269 documents that
|
||||
all share the one nested title, ``N500:2024`` — rung 2/3 of the navigation ladder collapse to
|
||||
an opaque UUID filename and a character count (P14 finding,
|
||||
``docs/2026-09-12-p14-kontekstsett.md`` § 5).
|
||||
``title:`` names the SOURCE document, not the concept (measured on a real concept of a
|
||||
delivered requirements corpus). Top-level keys carry no indentation and must win — a nested
|
||||
line must never overwrite a top-level key of the same name, however late it appears in the
|
||||
scan. Without this, ``directory_listing`` on that corpus's level returned 269 documents that
|
||||
all shared the one nested title, the source standard's own name — rung 2/3 of the navigation
|
||||
ladder collapse to an opaque UUID filename and a character count (P14 finding; the ledger is
|
||||
``docs/invarianter.md``).
|
||||
|
||||
Known-negatives that must stay green: ``type`` and ``req_number`` are untouched top-level
|
||||
scalars either way, and the ``verified:`` block-form decoder (SPEC §5.2) reads
|
||||
|
|
@ -331,18 +331,18 @@ def test_parse_frontmatter_top_level_title_survives_nested_sources_title(tmp_pat
|
|||
text = (
|
||||
"---\n"
|
||||
"type: Krav\n"
|
||||
"title: Krav 10.4.3—1 Mekanisk ventilasjon (impulsventilator)\n"
|
||||
"title: Krav 10.4.3—1 Reservestrøm (nødstrømsaggregat)\n"
|
||||
"req_number: Krav 10.4.3—1\n"
|
||||
"sources:\n"
|
||||
" - resource: https://example.invalid/859990\n"
|
||||
" title: N500:2024\n"
|
||||
" title: D500:2027\n"
|
||||
"---\n\n"
|
||||
"## Krav\nbody\n"
|
||||
)
|
||||
path = tmp_path / "concept.md"
|
||||
path.write_text(text, encoding="utf-8")
|
||||
fm = okf.parse_frontmatter(path)
|
||||
assert fm["title"] == "Krav 10.4.3—1 Mekanisk ventilasjon (impulsventilator)"
|
||||
assert fm["title"] == "Krav 10.4.3—1 Reservestrøm (nødstrømsaggregat)"
|
||||
assert fm["type"] == "Krav"
|
||||
assert fm["req_number"] == "Krav 10.4.3—1"
|
||||
|
||||
|
|
|
|||
|
|
@ -21,7 +21,7 @@ whole fail-closed suite green. ``_claims_ingest_ownership`` now recognises BOTH
|
|||
further lift must re-measure that predicate against what the new release actually writes, because
|
||||
nothing in the suite would go red if a third spelling appeared. The refusal messages below name it,
|
||||
because a pin whose reason lives only in prose is a pin the next session lifts without
|
||||
re-measuring. Full numbers: ``docs/2026-09-12-p13-okf-pin-r761.md`` and ``…-p13b-okf-bump.md``.
|
||||
re-measuring. The ledger row is in ``docs/invarianter.md``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -47,7 +47,7 @@ _OKF_LIFT_COST = (
|
|||
"knew only the boolean left write_concept_file's forgery refusal INERT with the whole "
|
||||
"fail-closed suite green (P13 § 2c'). A third spelling would do the same. Then expect the four "
|
||||
"examples/ingest-golden-* byte goldens (seven concept files) and conftest."
|
||||
"expected_generated_stamp to move with it. See docs/2026-09-12-p13b-okf-bump.md."
|
||||
"expected_generated_stamp to move with it. See docs/invarianter.md."
|
||||
)
|
||||
|
||||
_GUARD_LIFT_COST = (
|
||||
|
|
|
|||
|
|
@ -62,7 +62,7 @@ _PROVENANCE = ProvenanceStamp(
|
|||
role="proposer",
|
||||
validator_decision="validated",
|
||||
token_usage=8,
|
||||
# These fixtures stand in for an ordinary complete run; the road path is anchored by
|
||||
# These fixtures stand in for an ordinary complete run; the reference path is anchored by
|
||||
# construction, so ``True`` is the honest value here. The un-anchored case has its own file.
|
||||
cost_baseline_anchored=True,
|
||||
bundle_id_source=None,
|
||||
|
|
|
|||
|
|
@ -1,14 +1,14 @@
|
|||
"""P4 pkt. 4 — the two honesty sentences the demo says out loud, and why one of them is DERIVED.
|
||||
|
||||
The plan pre-wrote two sentences for the stage. The first is the provenance claim about the cost
|
||||
numbers, already live as ``_VEGLYS_PROVENANCE`` since the GO (P3): the baseline the validator's
|
||||
stage 0 reconciles against was written by a domain team, not by the demo script.
|
||||
numbers, already live as ``_KLIENTPARK_PROVENANCE`` since the GO (P3): the baseline the validator's
|
||||
stage 0 reconciles against ships with the knowledge base, not with the demo script.
|
||||
|
||||
The second is the seed sentence, about where Run B's "previous verdicts" come from. **Its pre-written
|
||||
wording is wrong against the delivered content and was corrected against measurement.** The plan said
|
||||
"one of the TWO previous verdicts in Run B came with the example — the other is the one the demo
|
||||
learned"; measured against the delivered VEGLYS bundle, Run B retrieves THREE: one seeded
|
||||
(``verdict-veglys-fro.md``), and TWO the demo produced itself — one per time scale (the Step-8
|
||||
learned"; measured against the delivered KLIENTPARK bundle, Run B retrieves THREE: one seeded
|
||||
(``verdict-klientpark-fro.md``), and TWO the demo produced itself — one per time scale (the Step-8
|
||||
promotion and the Step-7 inbox note). A sentence that says "two" would be a false claim made on
|
||||
stage about a number printed one line above it.
|
||||
|
||||
|
|
@ -31,7 +31,7 @@ from pathlib import Path
|
|||
from portfolio_optimiser.persona import load_persona_example
|
||||
from portfolio_optimiser.simulation import (
|
||||
_INBOX_MARKER,
|
||||
_VEGLYS_PROVENANCE,
|
||||
_KLIENTPARK_PROVENANCE,
|
||||
_delivered_bundle_dir,
|
||||
_verdict_origin_line,
|
||||
)
|
||||
|
|
@ -46,10 +46,10 @@ def _verdict(vid: str, rationale: str) -> Verdict:
|
|||
return Verdict(
|
||||
id=vid,
|
||||
proposal_features=ProposalFeatures(
|
||||
affected_codes=frozenset({"ENERGI-VEGLYS-EL"}),
|
||||
measure_type="LED-utskifting",
|
||||
affected_codes=frozenset({"ENERGI-KLIENTPARK-EL"}),
|
||||
measure_type="PC-utskifting",
|
||||
claimed_saving_nok=445500.0,
|
||||
description="LED-utskifting",
|
||||
description="PC-utskifting",
|
||||
),
|
||||
decision="approved",
|
||||
rationale=rationale,
|
||||
|
|
@ -96,12 +96,12 @@ def test_both_sentences_are_in_the_pinned_transcript() -> None:
|
|||
fasit is byte-compared to a live run in ``test_golden_transcript_loadbearing``, so a sentence
|
||||
that is in the file but no longer printed fails there, and a sentence dropped from BOTH fails
|
||||
here. The provenance half needs this: the anchored-reserve test only rules the reserve's
|
||||
sentence OUT, so emptying ``_VEGLYS_PROVENANCE`` would leave it green.
|
||||
sentence OUT, so emptying ``_KLIENTPARK_PROVENANCE`` would leave it green.
|
||||
"""
|
||||
fasit = (Path(__file__).resolve().parent / "golden" / "demo-transcript.stdout").read_text(
|
||||
encoding="utf-8"
|
||||
)
|
||||
assert _VEGLYS_PROVENANCE in fasit
|
||||
assert _KLIENTPARK_PROVENANCE in fasit
|
||||
assert "med kunnskapsbasen; de øvrige" in fasit
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
names what it is about.
|
||||
|
||||
**C1, what was measured (P19 F4).** ``_fetch_parsed`` retried a malformed reply with the
|
||||
BYTE-IDENTICAL prompt. One round-3 run, ``kontrakt-sorasen-04``, left a
|
||||
BYTE-IDENTICAL prompt. One round-3 run, on the set priced by process number, left a
|
||||
``{run_id}-parse-failures.json`` with ELEVEN rows, every one of them the same failure
|
||||
(``claimed_saving_nok`` ≤ 0) — eleven of the run's twelve rounds, spent re-asking a question the
|
||||
model had already answered the same wrong way, because nothing ever told it what was wrong. Step
|
||||
|
|
@ -40,10 +40,10 @@ from portfolio_optimiser.reference_domain import Project
|
|||
from portfolio_optimiser.run import announced_subject
|
||||
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
_BASE = _BUNDLES / "driftssenter-kjoling"
|
||||
|
||||
#: The exact failure ``kontrakt-sorasen-04`` produced eleven times: a well-formed JSON object
|
||||
#: The exact failure that round-3 run produced eleven times: a well-formed JSON object
|
||||
#: whose ``claimed_saving_nok`` is 0, refused by pydantic before any validator sees it.
|
||||
_UNPARSEABLE = (
|
||||
'{"measure":"m","affected_items":[{"code":"C-1","quantity":1000,"unit_cost":100}],'
|
||||
|
|
@ -116,7 +116,7 @@ def test_the_verbatim_evidence_is_still_captured() -> None:
|
|||
|
||||
def test_the_announcement_names_the_routed_bases() -> None:
|
||||
"""(e) Two bases, two declared ids — not "the portfolio"."""
|
||||
assert announced_subject(None, (str(_TUNNEL),)) == "tunnel-hauglia"
|
||||
assert announced_subject(None, (str(_BASE),)) == "driftssenter-kjoling"
|
||||
|
||||
|
||||
def test_an_unresolvable_base_falls_back_to_its_directory_name(tmp_path: Path) -> None:
|
||||
|
|
@ -145,7 +145,7 @@ def test_the_cli_prints_the_routed_bases(
|
|||
rc = main(
|
||||
[
|
||||
"--across-bundle",
|
||||
str(_TUNNEL),
|
||||
str(_BASE),
|
||||
"--mandate",
|
||||
str(mandate),
|
||||
"--run-id",
|
||||
|
|
@ -157,7 +157,7 @@ def test_the_cli_prints_the_routed_bases(
|
|||
)
|
||||
out = capsys.readouterr().out
|
||||
assert rc == 0, out
|
||||
assert "Run mandate for tunnel-hauglia" in out
|
||||
assert "Run mandate for driftssenter-kjoling" in out
|
||||
assert "the portfolio" not in out
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -32,34 +32,34 @@ _OFFLINE_MODEL_MAP = {
|
|||
# therefore OMIT ``assumptions`` and carry explicit magnitudes so each ``claimed_saving_nok`` <= P90.
|
||||
#
|
||||
# S4.0: every line below quotes a cost line the project ACTUALLY has, verbatim from
|
||||
# reference_projects.json — the road path now anchors the validator to ``project.cost_items``, so a
|
||||
# reference_projects.json — the reference path now anchors the validator to ``project.cost_items``, so a
|
||||
# reply quoting another project's code (which these fixtures used to do) is rejected as a fabricated
|
||||
# cost line. Verified against reference_projects.json + validator + ir:
|
||||
# FV42-GSV-E1 01.1 1x850,000 + 05.2 4300x215 Σ=1,774,500 P90=532,350 claimed 200,000
|
||||
# RV13-RAS-TP 22.4 610x3,850 Σ=2,348,500 P90=704,550 claimed 130,000 (decoy)
|
||||
# BRU-LAKS-REHAB 01.1 1x620,000 + 87.3 640x980 Σ=1,247,200 P90=374,160 claimed 210,000
|
||||
# ``measure`` is byte-identical "Reduce scope" for FV42+BRU (measure-match is exact string
|
||||
# equality, verdicts.py:68) and "Material substitution" for the decoy, so the BRU<->FV42 pair still
|
||||
# KONTOR-IT-E1 01.1 1x850,000 + 05.2 4300x215 Σ=1,774,500 P90=532,350 claimed 200,000
|
||||
# NETT-SIKR-TP 22.4 610x3,850 Σ=2,348,500 P90=704,550 claimed 130,000 (decoy)
|
||||
# ARKIV-LAGR-MIGR 01.1 1x620,000 + 87.3 640x980 Σ=1,247,200 P90=374,160 claimed 210,000
|
||||
# ``measure`` is byte-identical "Reduce scope" for KONTOR+ARKIV (measure-match is exact string
|
||||
# equality, verdicts.py:68) and "Material substitution" for the decoy, so the ARKIV<->KONTOR pair still
|
||||
# overlaps (shared code 01.1 + measure + magnitude bucket) while the decoy does not. 01.1 replaces
|
||||
# 05.2 as the shared code because it is the only code both projects genuinely carry.
|
||||
REPLIES = {
|
||||
"FV42-GSV-E1": (
|
||||
"KONTOR-IT-E1": (
|
||||
'{"measure":"Reduce scope","affected_items":['
|
||||
'{"code":"01.1","quantity":1,"unit_cost":850000},'
|
||||
'{"code":"05.2","quantity":4300,"unit_cost":215}],"claimed_saving_nok":200000}'
|
||||
),
|
||||
"RV13-RAS-TP": (
|
||||
"NETT-SIKR-TP": (
|
||||
'{"measure":"Material substitution","affected_items":['
|
||||
'{"code":"22.4","quantity":610,"unit_cost":3850}],"claimed_saving_nok":130000}'
|
||||
),
|
||||
"BRU-LAKS-REHAB": (
|
||||
"ARKIV-LAGR-MIGR": (
|
||||
'{"measure":"Reduce scope","affected_items":['
|
||||
'{"code":"01.1","quantity":1,"unit_cost":620000},'
|
||||
'{"code":"87.3","quantity":640,"unit_cost":980}],"claimed_saving_nok":210000}'
|
||||
),
|
||||
}
|
||||
|
||||
_PORTFOLIO_IDS = ["FV42-GSV-E1", "RV13-RAS-TP", "BRU-LAKS-REHAB"]
|
||||
_PORTFOLIO_IDS = ["KONTOR-IT-E1", "NETT-SIKR-TP", "ARKIV-LAGR-MIGR"]
|
||||
|
||||
|
||||
async def test_a_fanout_returns_one_runresult_per_project(
|
||||
|
|
@ -122,10 +122,10 @@ async def test_a3_rejected_proposal_excluded_from_aggregate(
|
|||
make_portfolio_client_factory, fresh_store
|
||||
) -> None:
|
||||
"""F2 (rejection arm, previously unexercised): one project's claim exceeds its P90, so the
|
||||
validator REJECTS it (FV42 800000 > P90 444750, <= Σ 1,482,500), while RV13+BRU validate.
|
||||
validator REJECTS it (KONTOR 800000 > P90 444750, <= Σ 1,482,500), while NETT+ARKIV validate.
|
||||
rejected_count counts it and the validated-only sum (run.py:241) EXCLUDES its claim. Also the
|
||||
F1 portfolio seam: provenance.model mirrors the injected client across all N records."""
|
||||
replies = {**REPLIES, "FV42-GSV-E1": REPLIES["FV42-GSV-E1"].replace("200000", "800000")}
|
||||
replies = {**REPLIES, "KONTOR-IT-E1": REPLIES["KONTOR-IT-E1"].replace("200000", "800000")}
|
||||
result = await run_portfolio(
|
||||
_PORTFOLIO_IDS,
|
||||
"local",
|
||||
|
|
@ -149,9 +149,9 @@ async def test_b_shared_store_accumulates_and_surfaces_prior_verdict(
|
|||
) -> None:
|
||||
"""SC4 (load-bearing, not mere non-emptiness): the ONE shared store accumulates a distinct
|
||||
verdict per project (3 -> 3 pairwise-distinct ids), and the cross-project ExpeL retrieval
|
||||
surfaces the STRUCTURAL match — BRU (runs[2]) retrieves FV42's verdict (runs[0]), since
|
||||
sim(BRU,FV42)=0.60 (shared code 05.2 + identical "Reduce scope" measure + same magnitude
|
||||
bucket) ranks above the RV13 decoy at sim(BRU,decoy)=0.15. If retrieval ranking breaks, or
|
||||
surfaces the STRUCTURAL match — ARKIV (runs[2]) retrieves KONTOR's verdict (runs[0]), since
|
||||
sim(ARKIV,KONTOR)=0.60 (shared code 05.2 + identical "Reduce scope" measure + same magnitude
|
||||
bucket) ranks above the NETT decoy at sim(ARKIV,decoy)=0.15. If retrieval ranking breaks, or
|
||||
two proposals collide to a single minted id, this fails — it asserts the specific match, not
|
||||
store-non-empty."""
|
||||
result = await run_portfolio(
|
||||
|
|
@ -163,7 +163,7 @@ async def test_b_shared_store_accumulates_and_surfaces_prior_verdict(
|
|||
assert len(result.store.verdicts) == 3
|
||||
ids = [v.id for v in result.store.verdicts]
|
||||
assert len(set(ids)) == 3 # pairwise distinct minted ids (no collision)
|
||||
# BRU surfaces FV42 (the structural match), ranked above the decoy.
|
||||
# ARKIV surfaces KONTOR (the structural match), ranked above the decoy.
|
||||
assert result.runs[2].retrieved[0].id == result.runs[0].verdict.id
|
||||
|
||||
|
||||
|
|
@ -173,28 +173,28 @@ async def test_c_execution_state_isolation_is_load_bearing(
|
|||
"""SC3 (cap-independent automatic detach guard): each project's execution state (the budget
|
||||
meter) is built FRESH per run, so per-project token_usage is its OWN only. Built on the
|
||||
Step-1 ``meter=`` seam + Step-3 ``meter_factory``. The factory emits a UNIFORM but VALID
|
||||
proposal (REPLIES["FV42-GSV-E1"] for every call) so both projects complete — a bare
|
||||
proposal (REPLIES["KONTOR-IT-E1"] for every call) so both projects complete — a bare
|
||||
unparseable default would loop the generate fetch to BudgetExceeded instead. Both arms run
|
||||
every CI, so the detach is encoded automatically (no manual reviewer-revert)."""
|
||||
from portfolio_optimiser.budget import Budget, TokenMeter
|
||||
from portfolio_optimiser.reference_domain import load_reference_projects
|
||||
|
||||
f = make_client_factory(REPLIES["FV42-GSV-E1"])
|
||||
rv13 = {p.id: p for p in load_reference_projects()}["RV13-RAS-TP"]
|
||||
f = make_client_factory(REPLIES["KONTOR-IT-E1"])
|
||||
nett = {p.id: p for p in load_reference_projects()}["NETT-SIKR-TP"]
|
||||
|
||||
# 1. Baseline: RV13 run standalone (its own fresh meter).
|
||||
# 1. Baseline: NETT run standalone (its own fresh meter).
|
||||
standalone = await run_project(
|
||||
"RV13-RAS-TP",
|
||||
"NETT-SIKR-TP",
|
||||
"local",
|
||||
docs_dir=rv13.docs_dir,
|
||||
verdict_input=rv13.verdict_input,
|
||||
docs_dir=nett.docs_dir,
|
||||
verdict_input=nett.verdict_input,
|
||||
client_factory=f,
|
||||
)
|
||||
# 2. Isolated (default, no meter_factory): RV13 as project 1 in the portfolio. Its usage
|
||||
# 2. Isolated (default, no meter_factory): NETT as project 1 in the portfolio. Its usage
|
||||
# equals the standalone baseline — independent of project 0. THE load-bearing assertion:
|
||||
# if run_portfolio shared a meter by default, runs[1] would be cumulative and this breaks.
|
||||
iso = await run_portfolio(
|
||||
["FV42-GSV-E1", "RV13-RAS-TP"],
|
||||
["KONTOR-IT-E1", "NETT-SIKR-TP"],
|
||||
"local",
|
||||
store=fresh_store,
|
||||
client_factory=f,
|
||||
|
|
@ -204,7 +204,7 @@ async def test_c_execution_state_isolation_is_load_bearing(
|
|||
# (and that the == arm above would redden under sharing).
|
||||
shared = TokenMeter(Budget(max_tokens=1_000_000, max_rounds=1000))
|
||||
sh = await run_portfolio(
|
||||
["FV42-GSV-E1", "RV13-RAS-TP"],
|
||||
["KONTOR-IT-E1", "NETT-SIKR-TP"],
|
||||
"local",
|
||||
client_factory=f,
|
||||
meter_factory=lambda: shared,
|
||||
|
|
@ -269,7 +269,7 @@ def test_f_no_hardcoded_project_ids_in_src() -> None:
|
|||
backing the 'config-only' claim. Pattern: test_budget.py:88-99 (src anti-pattern grep)."""
|
||||
from pathlib import Path
|
||||
|
||||
ids = ("FV42-GSV-E1", "RV13-RAS-TP", "BRU-LAKS-REHAB", "SKOLE-VVS-OPPGR")
|
||||
ids = ("KONTOR-IT-E1", "NETT-SIKR-TP", "ARKIV-LAGR-MIGR", "SKOLE-VVS-OPPGR")
|
||||
offenders = [
|
||||
f"{py.name}: {pid}"
|
||||
for py in Path("src/portfolio_optimiser").rglob("*.py")
|
||||
|
|
|
|||
|
|
@ -41,7 +41,7 @@ from portfolio_optimiser.budget import (
|
|||
)
|
||||
from portfolio_optimiser.run import run_portfolio
|
||||
|
||||
_PORTFOLIO_IDS = ["FV42-GSV-E1", "RV13-RAS-TP", "BRU-LAKS-REHAB"]
|
||||
_PORTFOLIO_IDS = ["KONTOR-IT-E1", "NETT-SIKR-TP", "ARKIV-LAGR-MIGR"]
|
||||
|
||||
# Per-project replies: IMPORTED from tests/test_portfolio.py rather than copied. The copy claimed to
|
||||
# be "the tested constants ... unchanged" and then drifted — S4.0's baseline anchoring caught it,
|
||||
|
|
|
|||
|
|
@ -53,9 +53,9 @@ from portfolio_optimiser.ledger import LedgerEntry, SavingsLedger
|
|||
from portfolio_optimiser.run import BudgetStop, PortfolioResult, RunFailure
|
||||
from portfolio_optimiser.verdicts import VerdictStore
|
||||
|
||||
# The same caller-supplied answers the single-project door documents. On the shipped ROAD
|
||||
# The same caller-supplied answers the single-project door documents. On the shipped REFERENCE
|
||||
# portfolio these are rejected by S4.0's baseline anchoring (the cost code belongs to the bygg
|
||||
# bundle, not to a road project's estimate) — which is the correct outcome and is beside the
|
||||
# bundle, not to a reference project's estimate) — which is the correct outcome and is beside the
|
||||
# point here: the blades below assert that the pass RUNS offline, not that it validates.
|
||||
_VALID_PROPOSAL = (
|
||||
'{"measure":"LED-retrofit av kontorbelysning","affected_items":'
|
||||
|
|
@ -142,7 +142,7 @@ def test_no_banner_in_portfolio_mode_without_the_flag(tmp_path, capsys) -> None:
|
|||
led = SavingsLedger()
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="FV42-GSV-E1",
|
||||
project_id="KONTOR-IT-E1",
|
||||
dimension="energi",
|
||||
candidate_identity="c1",
|
||||
amount_ore=1,
|
||||
|
|
@ -204,23 +204,23 @@ def test_failures_are_visible_with_their_own_values(stub_portfolio, capsys) -> N
|
|||
stub_portfolio(
|
||||
_result(
|
||||
failures=(
|
||||
RunFailure("FV42-GSV-E1", "Connection error.", "APIConnectionError"),
|
||||
RunFailure("RV13-RAS-TP", "budget blew up", "BudgetExceeded"),
|
||||
RunFailure("KONTOR-IT-E1", "Connection error.", "APIConnectionError"),
|
||||
RunFailure("NETT-SIKR-TP", "budget blew up", "BudgetExceeded"),
|
||||
)
|
||||
)
|
||||
)
|
||||
rc = run.main(["--portfolio"])
|
||||
err = capsys.readouterr().err
|
||||
assert rc == 1
|
||||
assert "FV42-GSV-E1" in err and "Connection error." in err
|
||||
assert "RV13-RAS-TP" in err and "budget blew up" in err
|
||||
assert "KONTOR-IT-E1" in err and "Connection error." in err
|
||||
assert "NETT-SIKR-TP" in err and "budget blew up" in err
|
||||
assert "APIConnectionError" in err and "BudgetExceeded" in err
|
||||
|
||||
|
||||
def test_a_pass_with_failures_does_not_exit_zero(stub_portfolio, capsys) -> None:
|
||||
"""Blade 6 — silence-as-success was the worse half of the defect: a scripted caller checking
|
||||
only rc learned nothing. MEASURED before the fix: four dead projects, rc 0."""
|
||||
stub_portfolio(_result(failures=(RunFailure("FV42-GSV-E1", "boom", "RuntimeError"),)))
|
||||
stub_portfolio(_result(failures=(RunFailure("KONTOR-IT-E1", "boom", "RuntimeError"),)))
|
||||
assert run.main(["--portfolio"]) == 1
|
||||
capsys.readouterr()
|
||||
|
||||
|
|
@ -246,7 +246,7 @@ def test_completed_runs_still_print_alongside_a_failure(replies_file, monkeypatc
|
|||
genuine = await real_run_portfolio(*args, **kwargs)
|
||||
return _result(
|
||||
runs=genuine.runs,
|
||||
failures=(RunFailure("RV13-RAS-TP", "boom", "RuntimeError"),),
|
||||
failures=(RunFailure("NETT-SIKR-TP", "boom", "RuntimeError"),),
|
||||
)
|
||||
|
||||
monkeypatch.setattr(run, "run_portfolio", _fake)
|
||||
|
|
@ -254,7 +254,7 @@ def test_completed_runs_still_print_alongside_a_failure(replies_file, monkeypatc
|
|||
captured = capsys.readouterr()
|
||||
assert rc == 1
|
||||
assert captured.out.count("verdict id=") == 4, captured.out
|
||||
assert "RV13-RAS-TP" in captured.err
|
||||
assert "NETT-SIKR-TP" in captured.err
|
||||
|
||||
|
||||
def test_budget_stop_is_visible_with_its_own_numbers(stub_portfolio, capsys) -> None:
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@ exactly — nothing wider is needed, and nothing narrower suffices.
|
|||
|
||||
**The plan said "sort each wave's new verdicts on ``project_id``"; that was measured wrong and the
|
||||
correction is load-bearing, so it is recorded here rather than only in git.** Lexicographic order
|
||||
is BRU/FV42/RV13 while the sequential pass yields FV42/RV13/BRU — a ``project_id`` sort is
|
||||
is ARKIV/KONTOR/NETT while the sequential pass yields KONTOR/NETT/ARKIV — a ``project_id`` sort is
|
||||
deterministic yet NOT identical to ``concurrency=1``, and the second is the actual contract. The
|
||||
plan also named that sort as the detach point; it is not. With per-project snapshots the wave list
|
||||
is never reordered by completion, so a ``sorted(...)`` in the barrier would re-sort an already
|
||||
|
|
@ -60,7 +60,7 @@ from portfolio_optimiser.verdicts import VerdictStore
|
|||
# The shipped 3-project fixture, in submission order (mirrors tests/test_portfolio.py:57).
|
||||
# ``REPLIES`` is reused from there deliberately: its three proposals are already verified to
|
||||
# validate AND to mint three DISTINCT verdict ids, which is the precondition self-check 3 asserts.
|
||||
_PORTFOLIO_IDS = ["FV42-GSV-E1", "RV13-RAS-TP", "BRU-LAKS-REHAB"]
|
||||
_PORTFOLIO_IDS = ["KONTOR-IT-E1", "NETT-SIKR-TP", "ARKIV-LAGR-MIGR"]
|
||||
|
||||
# Scheduler yields per project, DESCENDING by submission index: the FIRST-submitted project yields
|
||||
# most and therefore tends to finish LAST. This is what makes completion order depart from
|
||||
|
|
@ -378,7 +378,7 @@ async def test_runs_follow_project_ids_not_completion_order() -> None:
|
|||
"completion order matched submission order, so this assertion would hold trivially"
|
||||
)
|
||||
# ``claimed_saving_nok`` is the per-project discriminator in REPLIES (200k / 130k / 210k) —
|
||||
# ``measure_type`` is NOT, since FV42 and BRU deliberately share "Reduce scope".
|
||||
# ``measure_type`` is NOT, since KONTOR and ARKIV deliberately share "Reduce scope".
|
||||
assert [r.verdict.proposal_features.claimed_saving_nok for r in result.runs] == [
|
||||
200_000.0,
|
||||
130_000.0,
|
||||
|
|
@ -397,7 +397,7 @@ async def test_runs_follow_project_ids_not_completion_order() -> None:
|
|||
# The MIDDLE project of the wave. Middle is deliberate: a failure at either end can be dropped by
|
||||
# a truncation bug and still leave the survivors in the right relative order, so an end position
|
||||
# would let the ordering assertion below pass for the wrong reason.
|
||||
_FAILING_PID = "RV13-RAS-TP"
|
||||
_FAILING_PID = "NETT-SIKR-TP"
|
||||
|
||||
|
||||
class _FailingProbeClient(_OrderProbeClient):
|
||||
|
|
@ -511,9 +511,9 @@ async def test_one_project_failure_does_not_cancel_its_siblings() -> None:
|
|||
assert failure.error_type == "RuntimeError"
|
||||
assert "synthetic backend failure" in failure.error
|
||||
|
||||
# The survivors are the two healthy projects, in SUBMISSION order — 200k is FV42 (submitted
|
||||
# first), 210k is BRU (submitted last). ``measure_type`` is not a discriminator here: FV42 and
|
||||
# BRU deliberately share "Reduce scope".
|
||||
# The survivors are the two healthy projects, in SUBMISSION order — 200k is KONTOR (submitted
|
||||
# first), 210k is ARKIV (submitted last). ``measure_type`` is not a discriminator here: KONTOR and
|
||||
# ARKIV deliberately share "Reduce scope".
|
||||
assert [r.verdict.proposal_features.claimed_saving_nok for r in result.runs] == [
|
||||
200_000.0,
|
||||
210_000.0,
|
||||
|
|
@ -638,20 +638,20 @@ async def test_goal_stop_membership_is_identical_across_concurrency() -> None:
|
|||
empty ``runs`` and the equality is nearly free. It is asserted anyway because a regression that
|
||||
let a hard stop leak one wave's worth of runs would show up here first."""
|
||||
# Strong half — a per-project HARD goal skips exactly that pid, at every k.
|
||||
per_project = GoalConfig(per_project={"RV13-RAS-TP": GoalContract(absolute_ore=1000)})
|
||||
per_project = GoalConfig(per_project={"NETT-SIKR-TP": GoalContract(absolute_ore=1000)})
|
||||
probe = _Recorder()
|
||||
concurrent = await _goal_pass(
|
||||
3, ledger=_prefilled("RV13-RAS-TP", 1000), goals=per_project, recorder=probe
|
||||
3, ledger=_prefilled("NETT-SIKR-TP", 1000), goals=per_project, recorder=probe
|
||||
)
|
||||
sequential = await _goal_pass(
|
||||
1, ledger=_prefilled("RV13-RAS-TP", 1000), goals=per_project, recorder=_Recorder()
|
||||
1, ledger=_prefilled("NETT-SIKR-TP", 1000), goals=per_project, recorder=_Recorder()
|
||||
)
|
||||
|
||||
assert probe.max_in_flight > 1, (
|
||||
f"max in-flight was {probe.max_in_flight}: the surviving two projects never overlapped, so "
|
||||
"membership under concurrency is untested here"
|
||||
)
|
||||
assert _ran(concurrent) == _ran(sequential) == ["FV42-GSV-E1", "BRU-LAKS-REHAB"], (
|
||||
assert _ran(concurrent) == _ran(sequential) == ["KONTOR-IT-E1", "ARKIV-LAGR-MIGR"], (
|
||||
f"membership diverged with k: k=3 ran {_ran(concurrent)}, k=1 ran {_ran(sequential)}"
|
||||
)
|
||||
assert concurrent.stopped_early is False and sequential.stopped_early is False, (
|
||||
|
|
@ -661,10 +661,10 @@ async def test_goal_stop_membership_is_identical_across_concurrency() -> None:
|
|||
# Weak half — a HARD portfolio goal already reached stops the pass identically at every k.
|
||||
portfolio = GoalConfig(portfolio=GoalContract(absolute_ore=100))
|
||||
stopped_concurrent = await _goal_pass(
|
||||
3, ledger=_prefilled("FV42-GSV-E1", 100), goals=portfolio, recorder=_Recorder()
|
||||
3, ledger=_prefilled("KONTOR-IT-E1", 100), goals=portfolio, recorder=_Recorder()
|
||||
)
|
||||
stopped_sequential = await _goal_pass(
|
||||
1, ledger=_prefilled("FV42-GSV-E1", 100), goals=portfolio, recorder=_Recorder()
|
||||
1, ledger=_prefilled("KONTOR-IT-E1", 100), goals=portfolio, recorder=_Recorder()
|
||||
)
|
||||
assert _ran(stopped_concurrent) == _ran(stopped_sequential) == []
|
||||
assert stopped_concurrent.stopped_early is stopped_sequential.stopped_early is True
|
||||
|
|
@ -773,7 +773,7 @@ async def test_intra_wave_visibility_is_the_documented_semantic_difference(
|
|||
**Why the shipped fixture cannot show it, and why this test needs its own.** The Step-1 ExpeL
|
||||
fold is ``bundle_dir``-gated, and no project in ``reference_projects.json`` sets ``bundle_dir``,
|
||||
so on that fixture the cross-project chain is live for store CONTENT but inert for OUTCOMES —
|
||||
the difference exists and is unobservable. The road-*k* + bundle-*(k+1)* pair from
|
||||
the difference exists and is unobservable. The ref-*k* + bundle-*(k+1)* pair from
|
||||
``test_portfolio_learning_loadbearing.py`` is the fixture where the fold actually fires, injected
|
||||
through the SAME ``load_reference_projects`` monkeypatch seam that file uses (``:163``). No
|
||||
``projects=`` parameter is added to production code for this: that would be a new public seam
|
||||
|
|
|
|||
|
|
@ -40,11 +40,11 @@ from portfolio_optimiser.budget import PortfolioBudget, PortfolioMeter
|
|||
from portfolio_optimiser.run import RunFailure, run_portfolio
|
||||
from portfolio_optimiser.simulation import ScriptedChatClient
|
||||
|
||||
_PORTFOLIO_IDS = ["FV42-GSV-E1", "RV13-RAS-TP", "BRU-LAKS-REHAB"]
|
||||
_PORTFOLIO_IDS = ["KONTOR-IT-E1", "NETT-SIKR-TP", "ARKIV-LAGR-MIGR"]
|
||||
|
||||
# The MIDDLE project. Middle is deliberate (mirroring the Step-4 test): a failure at either end can
|
||||
# be dropped by a truncation bug and still leave the survivors in the right relative order.
|
||||
_FAILING_PID = "RV13-RAS-TP"
|
||||
_FAILING_PID = "NETT-SIKR-TP"
|
||||
|
||||
_DEFAULT_REPLY = (
|
||||
'{"measure":"Reduce scope","affected_items":'
|
||||
|
|
@ -54,7 +54,7 @@ _DEFAULT_REPLY = (
|
|||
# ``REPLIES`` is IMPORTED from tests/test_portfolio.py (see the import above) rather than copied —
|
||||
# the local copy had drifted onto other projects' cost codes, which S4.0's baseline anchoring
|
||||
# rejects. ``_DEFAULT_REPLY`` above is reached only by a prompt naming none of the three mapped
|
||||
# projects; on the anchored road path such a reply is rejected as a fabricated cost line, which is
|
||||
# projects; on the anchored reference path such a reply is rejected as a fabricated cost line, which is
|
||||
# the correct outcome for a project this fixture never described.
|
||||
|
||||
# Measured in tests/test_portfolio_budget_loadbearing.py: 4 chat calls x ``tokens`` per reply, so a
|
||||
|
|
|
|||
|
|
@ -6,7 +6,7 @@ the ``bundle_dir``-gated Step-1 ExpeL fold never fired on the portfolio path and
|
|||
verdict never reached the next hypothesis — an honesty breach, since ``run_portfolio``'s docstring
|
||||
already claimed a "cross-project learning loop".
|
||||
|
||||
Topology (road-*k* + bundle-*(k+1)*, ONE bundle): project *k* is road-backed (``docs_dir``,
|
||||
Topology (ref-*k* + bundle-*(k+1)*, ONE bundle): project *k* is reference-backed (``docs_dir``,
|
||||
``bundle_dir=None``); project *k+1* is bundle-backed (``bundle_dir`` = the repo-local mini-bundle,
|
||||
``id`` = its IR ``project_id``). *k*'s scripted proposal shares *k+1*'s candidate features (code
|
||||
``ENERGI-TOTAL-EL``, magnitude ~18000), so *k*'s captured verdict is what *k+1*'s ExpeL query
|
||||
|
|
@ -47,11 +47,11 @@ def _generation_prompts(sink: list[str]) -> list[str]:
|
|||
|
||||
|
||||
def _make_docs(tmp_path, name: str) -> str:
|
||||
"""A tmp docs folder with citable content, so the road path retrieval is non-empty."""
|
||||
"""A tmp docs folder with citable content, so the reference path retrieval is non-empty."""
|
||||
d = tmp_path / name
|
||||
d.mkdir()
|
||||
(d / "cost.txt").write_text(
|
||||
"Asphalt Ab11 unit rate renegotiation reduced the paving cost on the school stretch.",
|
||||
"Licence unit price renegotiation reduced the office-suite cost for the head office.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return str(d)
|
||||
|
|
@ -59,9 +59,9 @@ def _make_docs(tmp_path, name: str) -> str:
|
|||
|
||||
def _road_k(tmp_path, *, rationale: str) -> Project:
|
||||
return Project(
|
||||
id="ROAD-K",
|
||||
name="Road k",
|
||||
description="road-backed project k",
|
||||
id="REF-K",
|
||||
name="Reference k",
|
||||
description="reference-backed project k",
|
||||
currency="NOK",
|
||||
cost_items=(),
|
||||
docs_dir=_make_docs(tmp_path, "k-docs"),
|
||||
|
|
@ -88,9 +88,9 @@ def _bundle_kplus1(tmp_path, *, verdict_dir: str | None = None) -> Project:
|
|||
async def test_verdict_on_k_reaches_kplus1_hypothesis_prompt(
|
||||
tmp_path, monkeypatch, make_recording_client_factory
|
||||
) -> None:
|
||||
"""T-2.0a LOAD-BEARING: a verdict captured on road-backed project *k* reaches bundle-backed
|
||||
"""T-2.0a LOAD-BEARING: a verdict captured on reference-backed project *k* reaches bundle-backed
|
||||
*k+1*'s hypothesis-generation prompt (both the sentinel rationale and *k*'s verdict id, via the
|
||||
ExpeL few-shot). Detach ``bundle_dir=project.bundle_dir`` at run.py:496 → *k+1* runs the road
|
||||
ExpeL few-shot). Detach ``bundle_dir=project.bundle_dir`` at run.py:496 → *k+1* runs the reference
|
||||
path → the Step-1 fold is skipped → the sentinel never arrives → RED."""
|
||||
k = _road_k(tmp_path, rationale=_SENTINEL)
|
||||
kplus1 = _bundle_kplus1(tmp_path)
|
||||
|
|
|
|||
|
|
@ -35,6 +35,6 @@ _NO_FOUNDRY = not (_ENDPOINT and _DEPLOYMENT and _OPTED_IN)
|
|||
)
|
||||
async def test_portfolio_live_azure_fanout() -> None:
|
||||
# No client_factory -> the real AZURE backend is used; hard-capped per D6.
|
||||
result = await run_portfolio(["FV42-GSV-E1"], "azure", max_tokens=2000)
|
||||
result = await run_portfolio(["KONTOR-IT-E1"], "azure", max_tokens=2000)
|
||||
assert isinstance(result, PortfolioResult)
|
||||
assert len(result.runs) == 1
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
"""Load-bearing gate for order 20260908T141941Z (P3): the excerpt's new fields reach the prompt.
|
||||
|
||||
P2 (`docs/2026-09-08-n-bundlene-hypoteseform.md` SS 4) measured that llm-ingestion-okf's producer
|
||||
P2 (its ledger row is in `docs/invarianter.md`) measured that llm-ingestion-okf's producer
|
||||
now emits ``title``, ``req_number``, ``sources`` and per-excerpt ``source_*`` locators on every
|
||||
excerpt (14 members, was 9) -- and that ``po`` drops all five TWICE: ``PrepassExcerpt`` is
|
||||
``extra="ignore"`` so parsing throws them away, and ``_data_blocks`` renders only ``concept_id``,
|
||||
|
|
@ -8,11 +8,11 @@ excerpt (14 members, was 9) -- and that ``po`` drops all five TWICE: ``PrepassEx
|
|||
DATA block. (b') -- "did the model name the concept" -- was therefore a ``po`` verdict, never a
|
||||
model verdict: no implementation of the model could have passed it while the field never arrived.
|
||||
|
||||
**Field-form decision, measured, not guessed.** The producer's OWN message (cited in P2 SS 4,
|
||||
"differansen er at K2 baerer fem source_*-lokatorer der N-bundlene baerer to -- prefiks-regelen,
|
||||
ikke en allowlist") states that ``source_*`` is an open-ended FAMILY, not a fixed two-member
|
||||
allowlist. Two forms were compared against "does a future producer's new ``source_foo`` key reach
|
||||
the prompt without a code change here":
|
||||
**Field-form decision, measured, not guessed.** The producer's OWN message (cited in P2: K2
|
||||
carries five ``source_*`` locators where the requirement bundles carry two, and the rule is a
|
||||
prefix rule, not an allowlist) states that ``source_*`` is an open-ended FAMILY, not a fixed
|
||||
two-member allowlist. Two forms were compared against "does a future producer's new ``source_foo``
|
||||
key reach the prompt without a code change here":
|
||||
|
||||
- **Named fields** (``source_element_id: str | None``, ``source_sha256: str | None``, ...): a
|
||||
producer adding a THIRD ``source_*`` member requires a new field declared on ``PrepassExcerpt``
|
||||
|
|
@ -53,7 +53,7 @@ def _raw_with_new_fields() -> dict[str, Any]:
|
|||
excerpt["title"] = "Kontorbygg Nord (BYGG-KONTOR-NORD)"
|
||||
excerpt["req_number"] = "Krav 3.3.1–13"
|
||||
excerpt["sources"] = [
|
||||
{"resource": "https://viewers.vegnorm.vegvesen.no/api/nisosts/859984", "title": "N100:2023"}
|
||||
{"resource": "https://example.invalid/driftskrav/859984", "title": "D100:2027"}
|
||||
]
|
||||
excerpt["source_element_id"] = "id-4b61eee9-a149-42b3-863d-293b8320c15a"
|
||||
excerpt["source_sha256"] = "c58e8bbc5fa9a5400c111e51b04c05f2cfd9edabd884ef5352a486fdab2cb5ab"
|
||||
|
|
@ -69,7 +69,7 @@ def test_the_excerpt_carries_title_req_number_and_sources() -> None:
|
|||
assert excerpt.title == "Kontorbygg Nord (BYGG-KONTOR-NORD)"
|
||||
assert excerpt.req_number == "Krav 3.3.1–13"
|
||||
assert excerpt.sources is not None
|
||||
assert excerpt.sources[0].resource == "https://viewers.vegnorm.vegvesen.no/api/nisosts/859984"
|
||||
assert excerpt.sources[0].resource == "https://example.invalid/driftskrav/859984"
|
||||
|
||||
|
||||
def test_a_future_source_star_key_survives_without_a_code_change() -> None:
|
||||
|
|
@ -120,7 +120,7 @@ def test_the_data_block_renders_req_number_title_and_the_address() -> None:
|
|||
header = _header_for(prepass.render_context(payload), "bygg-kontor-nord")
|
||||
assert "Krav 3.3.1–13" in header
|
||||
assert "Kontorbygg Nord (BYGG-KONTOR-NORD)" in header
|
||||
assert "https://viewers.vegnorm.vegvesen.no/api/nisosts/859984" in header
|
||||
assert "https://example.invalid/driftskrav/859984" in header
|
||||
assert "id-4b61eee9-a149-42b3-863d-293b8320c15a" in header
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -349,7 +349,7 @@ def test_a_foreign_dimension_excerpt_is_refused_with_an_in_dimension_control(
|
|||
path = Path(bundle_dir) / (victim["concept_id"] + ".md")
|
||||
lines = path.read_text(encoding="utf-8").split("\n")
|
||||
assert lines[0].strip() == "---"
|
||||
lines.insert(1, "dimension: asfalt")
|
||||
lines.insert(1, "dimension: lisens")
|
||||
path.write_text("\n".join(lines), encoding="utf-8")
|
||||
victim["sha256"] = hashlib.sha256(path.read_bytes()).hexdigest()
|
||||
victim["text"] = prepass.concept_text(path)
|
||||
|
|
@ -357,10 +357,10 @@ def test_a_foreign_dimension_excerpt_is_refused_with_an_in_dimension_control(
|
|||
payload = prepass.PrepassPayload.model_validate(raw)
|
||||
|
||||
_verify(payload, bundle_dir, dimension=None) # control: no dimension admits everything
|
||||
_verify(payload, bundle_dir, dimension="asfalt") # control: its OWN dimension admits it
|
||||
_verify(payload, bundle_dir, dimension="lisens") # control: its OWN dimension admits it
|
||||
|
||||
with pytest.raises(prepass.PrepassRefused, match="dimension"):
|
||||
_verify(payload, bundle_dir, dimension="tunnel")
|
||||
_verify(payload, bundle_dir, dimension="kjoling")
|
||||
|
||||
|
||||
def test_concept_text_reproduces_the_producers_derivation(tmp_path: Path) -> None:
|
||||
|
|
|
|||
|
|
@ -123,7 +123,7 @@ def _base(tmp_path: Path, *, sentinel: bool = False) -> tuple[str, prepass.Prepa
|
|||
|
||||
|
||||
def _docs(tmp_path: Path) -> str:
|
||||
"""The road path's data source, unused on the bundle arm but required by the signature."""
|
||||
"""The reference path's data source, unused on the bundle arm but required by the signature."""
|
||||
d = tmp_path / "docs"
|
||||
d.mkdir(exist_ok=True)
|
||||
(d / "cost.txt").write_text("Energitiltak i kontorbygg.", encoding="utf-8")
|
||||
|
|
@ -359,7 +359,7 @@ async def test_a_payload_that_delivers_nothing_is_refused_by_name(tmp_path: Path
|
|||
|
||||
|
||||
async def test_a_payload_without_a_bundle_dir_is_refused(tmp_path: Path) -> None:
|
||||
"""The road path has no base for the payload to agree with."""
|
||||
"""The reference path has no base for the payload to agree with."""
|
||||
_, payload = _base(tmp_path)
|
||||
with pytest.raises(prepass.PrepassRefused, match="bundle"):
|
||||
await run_project(
|
||||
|
|
|
|||
|
|
@ -52,6 +52,7 @@ from portfolio_optimiser.verdicts import VerdictStore
|
|||
_REPO = Path(__file__).resolve().parents[1]
|
||||
_EXAMPLES = _REPO / "shared" / "examples"
|
||||
_BUNDLE_DIR = _EXAMPLES / "bygg-energi-mikro"
|
||||
_DELIVERED = _REPO / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
_PID = "BYGG-KONTOR-NORD"
|
||||
_RUN_ID = "proposal-review-door"
|
||||
|
||||
|
|
@ -229,7 +230,7 @@ def test_the_payload_carries_every_field_and_takes_its_key_from_the_injected_fun
|
|||
|
||||
@pytest.fixture(scope="module")
|
||||
def project():
|
||||
return load_reference_projects()[0] # FV42-GSV-E1
|
||||
return load_reference_projects()[0] # KONTOR-IT-E1
|
||||
|
||||
|
||||
def test_prior_feedback_none_keeps_the_base_prompt_byte_identical(project) -> None:
|
||||
|
|
@ -294,7 +295,7 @@ def test_the_feedback_sentinel_cannot_arrive_from_the_knowledge_base() -> None:
|
|||
# Group A (Step 3) — the reviewer INSIDE the attempt loop. Driven at ``generate_via_llm``.
|
||||
# ---------------------------------------------------------------------------------------------
|
||||
|
||||
# FV42-GSV-E1 cost codes 05.2 + 03.1 -> affected total 1_482_500, degenerate Monte Carlo
|
||||
# KONTOR-IT-E1 cost codes 05.2 + 03.1 -> affected total 1_482_500, degenerate Monte Carlo
|
||||
# P90 = 0.30 x 1_482_500 = 444_750 (the ``test_step5_refine_loadbearing`` fixture, reused).
|
||||
_CODES = ["05.2", "03.1"]
|
||||
_BAD_CLAIM = 800_000 # parseable, but above P90 -> the DETERMINISTIC validator rejects it
|
||||
|
|
@ -1733,11 +1734,11 @@ async def test_t22_the_dispatcher_threads_one_reviewer_into_every_base(
|
|||
Mandate(
|
||||
objective="o",
|
||||
approaches=(
|
||||
Approach(id="a", label="A", bundle_id="tunnel-hauglia"),
|
||||
Approach(id="b", label="B", bundle_id="veglys-fv-soer"),
|
||||
Approach(id="a", label="A", bundle_id="driftssenter-kjoling"),
|
||||
Approach(id="b", label="B", bundle_id="klientpark-energi"),
|
||||
),
|
||||
),
|
||||
(str(_EXAMPLES / "tunnel-hauglia"), str(_EXAMPLES / "veglys-fv-soer")),
|
||||
(str(_DELIVERED / "driftssenter-kjoling"), str(_DELIVERED / "klientpark-energi")),
|
||||
proposal_reviewer=reviewer,
|
||||
)
|
||||
assert calls, "the dispatcher must have reached run_project at least once"
|
||||
|
|
|
|||
|
|
@ -1,8 +1,9 @@
|
|||
"""P19 DEL B — a "cost code" must have a FORM when the input offers forms.
|
||||
|
||||
**The measured defect.** P18's round 2 ended with two ``validated`` proposals whose
|
||||
``affected_item`` codes were ordinary words out of a road standard's prose:
|
||||
``impulsventilator`` (in 4 of 270 N500 documents) and ``bituminøst bærelag`` (4 of 1 133 N200).
|
||||
``affected_item`` codes were ordinary words out of a standard's prose, in 4 of the 270 and 4 of
|
||||
the 1 133 documents of the two bases they came from. This file names them by two stand-ins of the
|
||||
same kind, ``nødstrømsaggregat`` and ``redundant kjøling``.
|
||||
Both are GROUNDED in P7's sense — they appear verbatim in the input, which is all that stage asks —
|
||||
and neither is INERT in P18/B1's sense, because neither is anywhere near the 5 % document share.
|
||||
They are simply not identifiers of a cost line. The gate had no stage that could say so.
|
||||
|
|
@ -17,14 +18,14 @@ be answered in a form it does not use, and refusing there would be a rule about
|
|||
about grounding — exactly what ``_ground_against_input``'s own docstring refuses for the stage it
|
||||
sits inside.
|
||||
|
||||
**B1 — the third and fourth forms, transcribed from a measurement.** R761's requirement numbers are
|
||||
bare dotted numbers (``12.1``, ``52.11``); ALL SIX ``ref`` values in
|
||||
``contexts/kontrakt-sorasen-2027/fasit.json`` are of that shape and NEITHER pre-P19 form matched
|
||||
one, so r761's whole offer was **3 identifiers over 6.5 MB**. Measured after: **2 332**. And the
|
||||
first form was WIDENED in the same pass, because P19/B2 made these forms decide ``prose`` as well
|
||||
as count an offer: this repo's own ``ENERGI-TOTAL-EL`` matched neither, so the classifier called a
|
||||
real cost code prose and the new gate refused it. A gate may only be wrong in the direction that
|
||||
admits too much.
|
||||
**B1 — the third and fourth forms, transcribed from a measurement.** A process catalogue's numbers
|
||||
are bare dotted numbers (``12.1``, ``52.11``); ALL SIX ``ref`` values in
|
||||
``contexts/driftsavtale-2027/fasit.json`` are of that shape, and NEITHER pre-P19 form matched one:
|
||||
the catalogue measured then offered **3 identifiers over 6.5 MB**. Measured after: **2 332**. And
|
||||
the first form was WIDENED in the same pass, because P19/B2 made these forms decide ``prose`` as
|
||||
well as count an offer: this repo's own ``ENERGI-TOTAL-EL`` matched neither, so the classifier
|
||||
called a real cost code prose and the new gate refused it. A gate may only be wrong in the direction
|
||||
that admits too much.
|
||||
|
||||
**Two exemptions, both load-bearing:**
|
||||
|
||||
|
|
@ -32,7 +33,7 @@ admits too much.
|
|||
this project, and the weaker stage must not overrule the stronger falsifier (the sentence
|
||||
``_grounding_text`` already carries about its third source). A derived schedule whose codes are
|
||||
bare section numbers would otherwise be refused wholesale by the gate meant to protect it;
|
||||
* ``has_identifier_form`` FULL-matches. ``impulsventilator 12.1`` carrying a process number does
|
||||
* ``has_identifier_form`` FULL-matches. ``nødstrømsaggregat 12.1`` carrying a process number does
|
||||
not make the word a cost code, and a substring rule would let any prose code smuggle one along.
|
||||
|
||||
What each arm pins:
|
||||
|
|
@ -41,8 +42,8 @@ What each arm pins:
|
|||
reason verbatim into the next attempt, and "ungrounded" alone teaches nothing);
|
||||
(b) the generality guard: a base offering no identifier form leaves the gate OFF;
|
||||
(c) an anchored code is exempt — the stronger falsifier wins;
|
||||
(d) the 26 known negatives: every ``ref`` in all four fasit files still classifies as an identifier,
|
||||
and so does this repo's own ``ENERGI-TOTAL-EL``;
|
||||
(d) the 18 known negatives: every ``ref`` in all three fasit files still classifies as an
|
||||
identifier, and so does this repo's own ``ENERGI-TOTAL-EL``;
|
||||
(e) B1's known positives and negatives, one by one;
|
||||
(f) B2: the classification is reported, and it is the SAME classifier the gate uses.
|
||||
"""
|
||||
|
|
@ -68,12 +69,12 @@ from portfolio_optimiser.validator import (
|
|||
#: A base that OFFERS identifier forms: eleven documents carrying requirement numbers, one of which
|
||||
#: also mentions the prose word. Eleven because ``_GROUNDING_MIN_INERT_DOCUMENTS`` is 10 — with
|
||||
#: fewer, P18/B1's share rule cannot fire at all and this file would be measuring that instead.
|
||||
_OFFERING = tuple(f"Krav 10.4.3—{i} om ventilasjon i tunnelen." for i in range(1, 12)) + (
|
||||
"Krav 8.4.2—1 sier at impulsventilator skal dimensjoneres for brannlast.",
|
||||
_OFFERING = tuple(f"Krav 10.4.3—{i} om ventilasjon i serverrommet." for i in range(1, 12)) + (
|
||||
"Krav 8.4.2—1 sier at nødstrømsaggregat skal dimensjoneres for full last.",
|
||||
)
|
||||
|
||||
#: The same corpus with every identifier removed — prose only.
|
||||
_FORMLESS = ("ingen koder her, bare tekst om ventilasjon", "og enda mer prosa om tunnelen")
|
||||
_FORMLESS = ("ingen koder her, bare tekst om ventilasjon", "og enda mer prosa om serverrommet")
|
||||
|
||||
|
||||
def _proposal(code: str, *, quantity: float = 4.0, unit_cost: float = 250_000.0) -> SavingsProposal:
|
||||
|
|
@ -91,8 +92,8 @@ def _proposal(code: str, *, quantity: float = 4.0, unit_cost: float = 250_000.0)
|
|||
|
||||
|
||||
def test_a_word_from_the_prose_is_not_a_cost_code() -> None:
|
||||
"""(a) P18's ``impulsventilator``, in the shape that reached ``validated``."""
|
||||
verdict = validate_proposal(_proposal("impulsventilator"), grounding=Grounding(_OFFERING))
|
||||
"""(a) P18's ``nødstrømsaggregat``, in the shape that reached ``validated``."""
|
||||
verdict = validate_proposal(_proposal("nødstrømsaggregat"), grounding=Grounding(_OFFERING))
|
||||
assert isinstance(verdict, Rejection)
|
||||
assert "has no identifier form" in verdict.reason
|
||||
# The DENOMINATOR, not just a complaint: Step 5 feeds this reason into the next attempt.
|
||||
|
|
@ -124,11 +125,11 @@ def test_a_code_the_baseline_carries_is_never_refused_for_its_shape() -> None:
|
|||
"""(c) Stage 0 has already ruled it a real line; the weaker stage must not overrule it."""
|
||||
baseline = CostBaseline(
|
||||
project_id="t",
|
||||
items={"impulsventilator": CostBaselineLine(quantity=4.0, unit_cost=250_000.0)},
|
||||
items={"nødstrømsaggregat": CostBaselineLine(quantity=4.0, unit_cost=250_000.0)},
|
||||
)
|
||||
assert isinstance(
|
||||
validate_proposal(
|
||||
_proposal("impulsventilator"), baseline=baseline, grounding=Grounding(_OFFERING)
|
||||
_proposal("nødstrømsaggregat"), baseline=baseline, grounding=Grounding(_OFFERING)
|
||||
),
|
||||
ValidatedProposal,
|
||||
)
|
||||
|
|
@ -142,15 +143,15 @@ def test_a_code_the_baseline_carries_is_never_refused_for_its_shape() -> None:
|
|||
def test_every_fasit_reference_in_every_context_set_is_an_identifier() -> None:
|
||||
"""(d) Every fasit reference across every context set, with the denominator.
|
||||
|
||||
One of them — ``Krav 3.3.2—1_1`` — is why the second form grew an optional ``_<n>`` suffix.
|
||||
Measured, not anticipated: before that it was the single reference the classifier called prose.
|
||||
One of them — ``Krav 5.2.2—1_1`` — carries the optional ``_<n>`` suffix the second form grew
|
||||
when a reference of that shape was measured as the single one the classifier called prose.
|
||||
|
||||
**The denominator MOVED 26 -> 32 with P17b's fifth context set**, and it is asserted rather
|
||||
than dropped for the reason it was written down in the first place: a list comprehension over
|
||||
``contexts/*/fasit.json`` that quietly found fewer rows would make this arm weaker without
|
||||
making it red. The six new ones are four ``Krav x.y.z—n`` from n200-2024 and TWO bare
|
||||
``prosessnr`` from r761-2025 (``12.11``, ``12.12``) — the punctuation-and-digits form B1 added,
|
||||
now exercised by a fasit and not only by a known-positive.
|
||||
**The denominator is 18: six references in each of the three example sets**, and it is asserted
|
||||
rather than dropped for the reason it was written down in the first place: a list comprehension
|
||||
over ``contexts/*/fasit.json`` that quietly found fewer rows would make this arm weaker without
|
||||
making it red. Ten are ``Krav x.y.z—n`` from driftskrav-2027 and eight are bare ``prosessnr``
|
||||
from prosesskatalog-2027 (``12.1``, ``12.11``, ``12.12``, …) — the punctuation-and-digits form
|
||||
B1 added, exercised by a fasit and not only by a known-positive.
|
||||
"""
|
||||
refs = [
|
||||
concept["ref"]
|
||||
|
|
@ -158,14 +159,24 @@ def test_every_fasit_reference_in_every_context_set_is_an_identifier() -> None:
|
|||
for row in json.loads(open(path, encoding="utf-8").read())["must_cite"]
|
||||
for concept in row["concepts"]
|
||||
]
|
||||
assert len(refs) == 32, f"denominator moved: {len(refs)}"
|
||||
assert len(refs) == 18, f"denominator moved: {len(refs)}"
|
||||
assert [r for r in refs if not has_identifier_form(r)] == []
|
||||
assert has_identifier_form("ENERGI-TOTAL-EL"), "this repo's own reference cost code"
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"token",
|
||||
["12.1", "12.12", "22.1", "52.1", "52.11", "51.1", "65 ASFALTDEKKER", "SHA-01", "B-20-00-00"],
|
||||
[
|
||||
"12.1",
|
||||
"12.12",
|
||||
"22.1",
|
||||
"52.1",
|
||||
"52.11",
|
||||
"51.1",
|
||||
"65 LAGRINGSSYSTEMER",
|
||||
"SHA-01",
|
||||
"B-20-00-00",
|
||||
],
|
||||
)
|
||||
def test_b1_known_positives(token: str) -> None:
|
||||
"""(e) The forms the delivered corpora carry."""
|
||||
|
|
@ -179,8 +190,8 @@ def test_b1_known_positives(token: str) -> None:
|
|||
"250000",
|
||||
"15.09.2026",
|
||||
"2026-09-15",
|
||||
"impulsventilator",
|
||||
"bituminøst bærelag",
|
||||
"nødstrømsaggregat",
|
||||
"redundant kjøling",
|
||||
"0.70",
|
||||
],
|
||||
)
|
||||
|
|
@ -197,9 +208,9 @@ def test_b1_known_negatives(token: str) -> None:
|
|||
def test_the_classification_is_reported_and_is_the_gates_own() -> None:
|
||||
"""(f) One classifier, two consumers (kø-(p)): a report that disagreed with the gate about one
|
||||
proposal would be evidence about nothing."""
|
||||
forms = classify_codes(["impulsventilator", "Krav 8.4.2—1", "12.1", "ENERGI-TOTAL-EL"])
|
||||
forms = classify_codes(["nødstrømsaggregat", "Krav 8.4.2—1", "12.1", "ENERGI-TOTAL-EL"])
|
||||
assert forms == {
|
||||
"impulsventilator": "prose",
|
||||
"nødstrømsaggregat": "prose",
|
||||
"Krav 8.4.2—1": "identifier",
|
||||
"12.1": "identifier",
|
||||
"ENERGI-TOTAL-EL": "identifier",
|
||||
|
|
|
|||
|
|
@ -14,7 +14,7 @@ from portfolio_optimiser.retrieval import TextSpan
|
|||
|
||||
def _stamp() -> ProvenanceStamp:
|
||||
return ProvenanceStamp(
|
||||
citations=[Citation(file="asphalt.txt", locator=TextSpan(0, 10), snippet="0123456789")],
|
||||
citations=[Citation(file="licence.txt", locator=TextSpan(0, 10), snippet="0123456789")],
|
||||
model="qwen3:4b",
|
||||
role="proposer",
|
||||
validator_decision="validated",
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
MAJOR-3 (``docs/2026-09-02-misjonsreview-v2.md`` § 7), measured before anything was changed in
|
||||
``docs/2026-09-02-read-bundle-kontekstkostnad.md``:
|
||||
|
||||
read_bundle payload bygg 3 861 / veglys 10 406 / tunnel 12 595 o200k_base tokens
|
||||
read_bundle payload 3 861 / 10 406 / 12 595 o200k_base tokens (three example bases)
|
||||
prompts one result rides in FIVE (navigator 1, manager 3, hypothesiser 1)
|
||||
share of every prompt-token 54 % / 59 % / 59 % of one CLI ``--explore`` run
|
||||
|
||||
|
|
@ -21,7 +21,7 @@ level, not data loss: every byte is still exactly one call away, and a navigator
|
|||
documents it chose to open instead of for the ones it did not.
|
||||
|
||||
**A premise felled before it was built on** (see the measurement doc § 2): "the index body is the
|
||||
base's own navigation prose, so it belongs here". The tunnel base's root index is **4 763 chars
|
||||
base's own navigation prose, so it belongs here". The largest base's root index was **4 763 chars
|
||||
alone ≈ 1 400 tokens** — nearly the entire ceiling, for a field ``list_bundles`` already excerpts
|
||||
and ``read_file(id, "index.md")`` still returns whole.
|
||||
|
||||
|
|
@ -33,7 +33,7 @@ raising the budget is precisely the regression this file exists to catch.
|
|||
file bounds CHARACTERS instead. ``tiktoken`` is not a project dependency (and adding one for a
|
||||
gate would be a bigger decision than the gate), and a gate that skips when an optional package is
|
||||
missing is a gate that can be silently absent. The character ceiling is a proxy whose conversion
|
||||
was MEASURED rather than assumed: the new payload over the real tunnel base is **748 chars / 259
|
||||
was MEASURED rather than assumed: the new payload over the largest base was **748 chars / 259
|
||||
o200k tokens** (2.89 chars/token for this Norwegian markdown), so 1 500 characters is ≈ 520 tokens
|
||||
— comfortably inside the order's criterion, and twice the measured payload, so ordinary field
|
||||
growth does not force a rewrite. The order's own criterion is verified directly, once, by the
|
||||
|
|
@ -49,7 +49,7 @@ for byte -- and the hierarchy's own gate is ``tests/test_hierarchical_navigation
|
|||
|
||||
What the arms pin, and what each one refuses:
|
||||
|
||||
(a) the bound itself, over the REAL base the order names — refuses the unbounded form. A synthetic
|
||||
(a) the bound itself, over the largest SHIPPED base — refuses the unbounded form. A synthetic
|
||||
fixture here would measure the fixture writer, not the base;
|
||||
(b) cost tracks DOCUMENT COUNT, not document SIZE — ten times the prose, the same price. That is
|
||||
the property stated directly rather than inferred from (a);
|
||||
|
|
@ -76,10 +76,11 @@ from typing import Any
|
|||
from portfolio_optimiser import okf
|
||||
from portfolio_optimiser.explore import navigator_tools
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
|
||||
#: The base the order names. Measured, not chosen: the largest of the three example bases.
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||
#: The largest example base the package ships -- the structural counterpart (same documents, same
|
||||
#: link graph, same frontmatter) of the base the order named, which has since been replaced.
|
||||
_KJOLING = _BUNDLES / "driftssenter-kjoling"
|
||||
|
||||
#: Characters one ``read_bundle`` listing may cost. Test-owned on purpose — see the docstring, and
|
||||
#: the measured chars/token conversion that ties it to the order's token criterion.
|
||||
|
|
@ -130,12 +131,12 @@ def _write_base(
|
|||
return base
|
||||
|
||||
|
||||
def test_read_bundle_over_the_real_tunnel_base_is_bounded() -> None:
|
||||
"""(a) The headline, over the base the order names — not a fixture of my own making."""
|
||||
blob = _blob(_listing(_TUNNEL))
|
||||
def test_read_bundle_over_the_largest_shipped_base_is_bounded() -> None:
|
||||
"""(a) The headline, over the largest shipped base — not a fixture of my own making."""
|
||||
blob = _blob(_listing(_KJOLING))
|
||||
|
||||
assert len(blob) <= _CEILING_CHARS, (
|
||||
f"read_bundle over {_TUNNEL.name} costs {len(blob)} chars, over the ceiling "
|
||||
f"read_bundle over {_KJOLING.name} costs {len(blob)} chars, over the ceiling "
|
||||
f"{_CEILING_CHARS}; it used to be 39 583 (12 595 o200k tokens), riding in five prompts"
|
||||
)
|
||||
|
||||
|
|
@ -164,8 +165,8 @@ def test_the_listing_still_identifies_every_document(tmp_path: Path) -> None:
|
|||
will cost (``chars``) — so all four are asserted, and the entry count is tied to the navigated
|
||||
context files rather than to a number written here.
|
||||
"""
|
||||
bundle = okf.navigate_bundle(str(_TUNNEL))
|
||||
entries = _documents(_TUNNEL)
|
||||
bundle = okf.navigate_bundle(str(_KJOLING))
|
||||
entries = _documents(_KJOLING)
|
||||
|
||||
assert len(entries) == len(bundle.context_files) > 0, (
|
||||
"a listing that omits documents is a base the navigator cannot fully see"
|
||||
|
|
@ -200,10 +201,10 @@ def test_the_verdict_layer_is_still_excluded(tmp_path: Path) -> None:
|
|||
|
||||
def test_the_whole_document_is_still_one_call_away() -> None:
|
||||
"""(e) The bound is a disclosure LEVEL, not data loss."""
|
||||
entries = _documents(_TUNNEL)
|
||||
entries = _documents(_KJOLING)
|
||||
biggest = max(entries, key=lambda e: int(e["chars"]))
|
||||
|
||||
whole = _tools(_TUNNEL)["read_file"].func(bundle_id=_TUNNEL.name, path=str(biggest["name"]))
|
||||
whole = _tools(_KJOLING)["read_file"].func(bundle_id=_KJOLING.name, path=str(biggest["name"]))
|
||||
|
||||
assert len(whole) > _CEILING_CHARS, "read_file must still return the document, not a summary"
|
||||
assert str(biggest["chars"]) != "0" and int(biggest["chars"]) <= len(whole)
|
||||
|
|
@ -211,11 +212,11 @@ def test_the_whole_document_is_still_one_call_away() -> None:
|
|||
|
||||
def test_control_one_document_alone_would_blow_the_ceiling() -> None:
|
||||
"""(f) The ceiling discriminates — proved, not assumed. The order names this control."""
|
||||
bundle = okf.navigate_bundle(str(_TUNNEL))
|
||||
bundle = okf.navigate_bundle(str(_KJOLING))
|
||||
|
||||
biggest = max(len(f.body) for f in bundle.context_files)
|
||||
|
||||
assert biggest > _CEILING_CHARS, (
|
||||
f"the largest document in {_TUNNEL.name} is {biggest} chars; a ceiling it does not exceed "
|
||||
f"the largest document in {_KJOLING.name} is {biggest} chars; a ceiling it does not exceed "
|
||||
"would be a ceiling this base could pass while carrying everything"
|
||||
)
|
||||
|
|
|
|||
|
|
@ -184,7 +184,7 @@ def test_a_foreign_dimension_document_is_never_named(tmp_path: Path) -> None:
|
|||
name only" S2c measured. ONE predicate (``in_dimension``) serves the listing and this."""
|
||||
base = _base(tmp_path)
|
||||
|
||||
scoped = _read_dir(base, "energi-notat.md", dimension="tunnel")
|
||||
scoped = _read_dir(base, "energi-notat.md", dimension="kjoling")
|
||||
|
||||
assert "read_file" not in scoped["refused"]
|
||||
assert scoped["refusal"] == okf.BundlePathNotFound.__name__
|
||||
|
|
|
|||
|
|
@ -47,9 +47,9 @@ from portfolio_optimiser.simulation import ScriptedChatClient
|
|||
|
||||
_FIXTURES = Path(__file__).parent / "fixtures"
|
||||
_PRICED = str(_FIXTURES / "k2-prisskjema-SYNTETISK")
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
#: A base that SHIPS a hand-written ``cost-baseline.json`` - the anchored control.
|
||||
_ANCHORED_SOURCE = _EXAMPLES / "tunnel-hauglia"
|
||||
_ANCHORED_SOURCE = _BUNDLES / "driftssenter-kjoling"
|
||||
|
||||
_IR_PROJECTION = {
|
||||
"project_id": "K2",
|
||||
|
|
@ -170,7 +170,7 @@ async def test_it_composes_with_the_derivation(tmp_path: Path) -> None:
|
|||
|
||||
|
||||
def test_cli_requires_bundle_dir(capsys: pytest.CaptureFixture[str]) -> None:
|
||||
"""(f) The road path is anchored by construction, so on it the flag could never fire. A flag
|
||||
"""(f) The reference path is anchored by construction, so on it the flag could never fire. A flag
|
||||
that cannot fire is a claim the surface makes about itself (the Fase-3 class)."""
|
||||
rc = run.main(["P1", "--docs-dir", "docs", "--require-cost-baseline"])
|
||||
|
||||
|
|
|
|||
|
|
@ -16,8 +16,8 @@ measure without sharing a word with the name someone gave it — which is precis
|
|||
alternative rule P21/C1 measured and rejected failed, one rung over.
|
||||
|
||||
**Measured before it was built**, offline against the six traces: the rule speaks on **10 of 12**
|
||||
declarations and stays quiet on 2 (both fv412, on ``materialer``). A rule that spoke on 12 of 12,
|
||||
or on 0 of 12, could not tell the two classes apart.
|
||||
declarations and stays quiet on 2 (both in one context set, on ``materialer``). A rule that spoke
|
||||
on 12 of 12, or on 0 of 12, could not tell the two classes apart.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -33,24 +33,25 @@ from portfolio_optimiser.explore import ToolCall, navigator_tools
|
|||
from portfolio_optimiser.mandate import Approach, Mandate
|
||||
from portfolio_optimiser.run import run_project
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||
_BASE_ID = "tunnel-hauglia"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
_BASE = _BUNDLES / "driftssenter-kjoling"
|
||||
_BASE_ID = "driftssenter-kjoling"
|
||||
|
||||
#: The fixture document the arms declare. Its title is ASCII-clean, which is what lets a label
|
||||
#: share a word with it without a marker carrying a multibyte character into a scripted run.
|
||||
_DOC = "tiltak-portalskjerming.md"
|
||||
#: The fixture document the arms declare. The words the arms match on are ASCII-clean, which is
|
||||
#: what lets a label share a word with it without a marker carrying a multibyte character into a
|
||||
#: scripted run.
|
||||
_DOC = "tiltak-kaldgangsinnkapsling.md"
|
||||
|
||||
#: A direction whose words are IN that title ("Portalskjerming: senke L20 ...").
|
||||
_MATCHING = "Billigere portalskjerming"
|
||||
#: A direction whose words are IN that title ("Kaldgangsinnkapsling: senke varmetilskuddet ...").
|
||||
_MATCHING = "Billigere kaldgangsinnkapsling"
|
||||
#: A direction that shares nothing with it. Checked against the document's own tokens, both ways.
|
||||
_FOREIGN = "Asfaltdekke gjenbruk"
|
||||
_FOREIGN = "Lagringsenhet gjenbruk"
|
||||
|
||||
|
||||
def _wired(labels: tuple[str, ...]) -> tuple[dict[str, Any], list[ToolCall], list[Any]]:
|
||||
opened: list[ToolCall] = []
|
||||
declared: list[Any] = []
|
||||
tools = navigator_tools((str(_TUNNEL),), opened=opened, requirements=declared, labels=labels)
|
||||
tools = navigator_tools((str(_BASE),), opened=opened, requirements=declared, labels=labels)
|
||||
return {t.name: t for t in tools}, opened, declared
|
||||
|
||||
|
||||
|
|
@ -58,7 +59,7 @@ def _declare(
|
|||
labels: tuple[str, ...], *, ref: str = "Krav 1.1-1"
|
||||
) -> tuple[dict[str, Any], list[Any]]:
|
||||
tools, opened, declared = _wired(labels)
|
||||
for name in [f.name for f in okf.navigate_bundle(str(_TUNNEL)).context_files][:3]:
|
||||
for name in [f.name for f in okf.navigate_bundle(str(_BASE)).context_files][:3]:
|
||||
opened.append(ToolCall(name="read_file", bundle_id=_BASE_ID, path=name))
|
||||
opened.append(ToolCall(name="read_file", bundle_id=_BASE_ID, path=_DOC))
|
||||
answer = tools["declare_requirement"].func(
|
||||
|
|
@ -74,8 +75,8 @@ def test_a_direction_that_shares_a_word_is_told_which_one() -> None:
|
|||
"""KNOWN-POSITIVE. The document's title carries the direction's own word, and the reply says
|
||||
so — the half that keeps the report from being one that only ever complains."""
|
||||
answer, _ = _declare((_MATCHING,))
|
||||
assert answer["overlap"] == ["portalskjerming"], answer["overlap"]
|
||||
assert "portalskjerming" in answer["compare"]
|
||||
assert answer["overlap"] == ["kaldgangsinnkapsling"], answer["overlap"]
|
||||
assert "kaldgangsinnkapsling" in answer["compare"]
|
||||
|
||||
|
||||
def test_a_direction_that_shares_nothing_is_told_that_too() -> None:
|
||||
|
|
@ -108,7 +109,7 @@ def test_the_comparison_never_reads_the_callers_own_ref() -> None:
|
|||
input can only ever agree — P20/A1's rule (read off the base, never off the arguments)
|
||||
applied to the half P20 did not reach. A ``ref`` stuffed with the direction's words must not
|
||||
manufacture an overlap."""
|
||||
answer, _ = _declare((_FOREIGN,), ref="Asfaltdekke gjenbruk krav")
|
||||
answer, _ = _declare((_FOREIGN,), ref="Lagringsenhet gjenbruk krav")
|
||||
assert answer["overlap"] == [], answer["overlap"]
|
||||
|
||||
|
||||
|
|
@ -139,11 +140,13 @@ def test_without_directions_the_reply_is_the_one_p20_shipped() -> None:
|
|||
|
||||
def test_matching_is_generous_in_both_directions() -> None:
|
||||
"""The failure direction chosen on purpose. A direction whose word is a PREFIX of the
|
||||
document's own word counts, and so does the reverse — Norwegian inflects ("rundkjoring" vs
|
||||
"Rundkjoringer") and a strict rule would report "no overlap" on a declaration that was right,
|
||||
document's own word counts, and so does the reverse — Norwegian inflects ("sikkerhetskopi" vs
|
||||
"Sikkerhetskopier") and a strict rule would report "no overlap" on a declaration that was right,
|
||||
which is the only one of the two errors that can push a model away from a correct answer."""
|
||||
# label word LONGER than the document's own ("portalskjerming" is a prefix of it)
|
||||
assert _declare(("Portalskjermingen paa nordsiden",))[0]["overlap"] == ["portalskjermingen"]
|
||||
# label word LONGER than the document's own ("kaldgangsinnkapsling" is a prefix of it)
|
||||
assert _declare(("Kaldgangsinnkapslingen paa nordsiden",))[0]["overlap"] == [
|
||||
"kaldgangsinnkapslingen"
|
||||
]
|
||||
# and SHORTER: the document says "senke", the direction "senkekostnader"
|
||||
assert _declare(("Senkekostnader",))[0]["overlap"] == ["senkekostnader"]
|
||||
|
||||
|
|
@ -158,7 +161,7 @@ def test_short_words_cannot_manufacture_an_overlap() -> None:
|
|||
|
||||
# --- (h) the commission's labels actually REACH the tool in a run --------------------------------
|
||||
|
||||
_PID = "TUNNEL-HAUGLIA"
|
||||
_PID = "DRIFTSSENTER-KJOLING"
|
||||
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (test)"}
|
||||
_VALID_REPLY = (
|
||||
'{"measure":"LED-retrofit","affected_items":'
|
||||
|
|
@ -174,7 +177,7 @@ async def test_a_commissioned_run_reaches_the_tool_with_its_own_directions() ->
|
|||
a lint; this drives the real debate with a step manuscript that declares, and reads the
|
||||
comparison back out of the tool's own answer. Without ``labels=`` in ``run.py`` the reply
|
||||
carries no ``compare`` at all and this arm falls."""
|
||||
concepts = [f.name for f in okf.navigate_bundle(str(_TUNNEL)).context_files]
|
||||
concepts = [f.name for f in okf.navigate_bundle(str(_BASE)).context_files]
|
||||
script = {
|
||||
"proposer": [
|
||||
*(
|
||||
|
|
@ -218,8 +221,8 @@ async def test_a_commissioned_run_reaches_the_tool_with_its_own_directions() ->
|
|||
await run_project(
|
||||
_PID,
|
||||
"local",
|
||||
docs_dir=str(_TUNNEL),
|
||||
bundle_dir=str(_TUNNEL),
|
||||
docs_dir=str(_BASE),
|
||||
bundle_dir=str(_BASE),
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
client_factory=factory,
|
||||
mandate=Mandate(
|
||||
|
|
|
|||
|
|
@ -1,45 +1,50 @@
|
|||
"""P20 DEL B — a clause number is not a price, and the base's own vocabulary is what says so.
|
||||
|
||||
**What was measured.** Three paid rounds and one multi-base pass carried FOUR ``validated``
|
||||
proposals whose cost code was a chapter number of a standard. Two survive in the recorded outboxes
|
||||
and are this arm's known positives:
|
||||
proposals whose cost code was a chapter number of a standard. Two survived in the recorded
|
||||
outboxes and were this arm's known positives:
|
||||
|
||||
* ``10.4`` — tunnel-hauglia round 3, base ``vegnormal-n500-2024``, ``validated``;
|
||||
* ``1.10.4`` — lindaas P17b, base ``vegnormal-r761-2025``, ``validated``.
|
||||
* ``10.4`` — a stress round on a requirements base, ``validated``;
|
||||
* ``1.10.4`` — the multi-base pass, on the process catalogue, ``validated``.
|
||||
|
||||
Both are GROUNDED in P7's sense (they occur verbatim in the input) and neither is INERT in
|
||||
Both were GROUNDED in P7's sense (they occur verbatim in the input) and neither was INERT in
|
||||
P18/B1's sense (``10.4`` in 12 of 274 documents, ``1.10.4`` in 1 of 2 756). Stage 0 never ran:
|
||||
no vegnormal base ships a cost baseline. Nothing in the gate could say what they are.
|
||||
no requirements base ships a cost baseline. Nothing in the gate could say what they are.
|
||||
|
||||
**THE ORDER'S OWN RULE WAS FELLED BY MEASUREMENT BEFORE ANYTHING WAS BUILT ON IT.** B1 reads: a
|
||||
code is a requirement when it has form 2 or 3 AND "står som ``req_number``/``prosessnr`` i
|
||||
toppnivå-frontmatter i minst ett av grunnlagets dokumenter" — refuse that. Measured 15.09:
|
||||
|
||||
* n500 declares ``seksjon: 10.4.1`` … ``10.4.4`` and ``req_number: Krav 10.4.3—2``. The bare
|
||||
``10.4`` is declared NOWHERE — it is a section PREFIX;
|
||||
* r761 declares 2 727 ``prosessnr`` and 2 753 ``seksjon``. ``1.10.4`` is NONE of them: it occurs
|
||||
once, as prose, in "Krav til materialer skal være iht. vegnormal N200 Vegbygging kap. 1.10.4".
|
||||
* the requirements base declared ``seksjon: 10.4.1`` … ``10.4.4`` and ``req_number: Krav
|
||||
10.4.3—2``. The bare ``10.4`` was declared NOWHERE — it is a section PREFIX;
|
||||
* the process catalogue declared 2 727 ``prosessnr`` and 2 753 ``seksjon``. ``1.10.4`` was NONE of
|
||||
them: it occurred once, as prose, a chapter reference into ANOTHER standard.
|
||||
|
||||
The ordered rule therefore fires on NEITHER of its own known positives. The COMPLEMENT fires on
|
||||
BOTH, and it closes a hole ``_ground_against_input`` already admits in writing — "it fails OPEN …
|
||||
on a coincidental match". For one shape, a clause number, the base hands us the vocabulary needed
|
||||
to tell a real reference from a coincidence, and that is the rule built here.
|
||||
|
||||
The complement is also what SPARES the one context set built on real process codes: all five of
|
||||
``contexts/kontrakt-sorasen-2027``'s codes are declared ``prosessnr`` and pass. Under the ordered
|
||||
rule every one of them would have been refused on an unanchored r761 run, and the set's positive
|
||||
arms would have become unmeasurable — the R761 risk the order names, arriving through the door it
|
||||
was pointed away from.
|
||||
The complement is also what SPARES the context set built on the catalogue's own process codes:
|
||||
all five of ``contexts/driftsavtale-2027``'s codes are declared ``prosessnr`` and pass. Under the
|
||||
ordered rule every one of them would have been refused on an unanchored run on the catalogue, and
|
||||
the set's positive arms would have become unmeasurable — the risk the order names (a process
|
||||
number is both a clause and a settlement post), arriving through the door it was pointed away from.
|
||||
|
||||
Measured over EVERY code of round 3 and P17b (24 codes, 10 runs): exactly two are
|
||||
requirement-shaped, they are the two known positives, and the replay flips exactly those two.
|
||||
Measured over EVERY code of that round and the multi-base pass (24 codes, 10 runs): exactly two
|
||||
were requirement-shaped, they were the two known positives, and the replay flipped exactly those
|
||||
two. Those recordings were made against corpora this repository no longer carries, so (a) and (b)
|
||||
are now CONSTRUCTED in the same two shapes against the package's example bases: ``4.2.5`` — a
|
||||
section PREFIX in ``driftskrav-2027`` (it declares ``seksjon: 4.2.5.1`` and ``Krav 4.2.5.1—1`` …
|
||||
``—3``; the bare ``4.2.5`` stands in 3 of 306 documents and is no document's own number) — and
|
||||
``1.10.4`` — a clause of another document quoted once, as prose, in ``prosesskatalog-2027``.
|
||||
|
||||
What each arm pins:
|
||||
|
||||
(a) known positive — the ``10.4`` proposal, replayed against the base it actually ran on, is
|
||||
(a) known positive — the section prefix ``4.2.5``, unanchored against the requirements base, is
|
||||
``rejected``, and the reason names the denominator;
|
||||
(b) known positive — ditto ``1.10.4`` on r761;
|
||||
(c) known negative — a code the base DOES declare (sorasen's real ``12.1``) still validates;
|
||||
(b) known positive — ditto ``1.10.4`` on the process catalogue;
|
||||
(c) known negative — a code the base DOES declare (driftsavtale's ``12.1``) still validates;
|
||||
(d) known negative — the gate is OFF when the run is anchored, even for a clause-shaped code;
|
||||
(e) the generality guard — an input that declares no reference numbers at all cannot trip the rule,
|
||||
which is what leaves every pre-P20 fixture untouched rather than exempted;
|
||||
|
|
@ -74,15 +79,12 @@ from portfolio_optimiser.validator import (
|
|||
validate_proposal,
|
||||
)
|
||||
|
||||
_ROUND3 = Path("scratchpad/p19-stress/tunnel-hauglia-2027")
|
||||
_P17B = Path("scratchpad/p17b-multibase/lindaas")
|
||||
|
||||
|
||||
def _base(name: str) -> Path:
|
||||
"""The FROZEN copy this repository pins, resolved at call time.
|
||||
|
||||
Absence SKIPS (MAJOR-3's ceiling: no corpus is mounted in the handover archive), drift is
|
||||
allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
Absence SKIPS (a user's own store, named by ``PORTFOLIO_FROZEN_BUNDLES``, may not hold it),
|
||||
drift is allowed to propagate and FAIL — a measurement of the wrong corpus is not a missing one.
|
||||
"""
|
||||
try:
|
||||
return frozen_bundles.bundle_dir(name)
|
||||
|
|
@ -103,12 +105,6 @@ def _grounding_over(name: str) -> Grounding:
|
|||
)
|
||||
|
||||
|
||||
def _recorded(path: Path) -> SavingsProposal:
|
||||
if not path.is_file():
|
||||
pytest.skip(f"the recorded artefact {path} is not present in this checkout")
|
||||
return SavingsProposal.model_validate(json.loads(path.read_text(encoding="utf-8"))["proposal"])
|
||||
|
||||
|
||||
def _proposal(code: str, *, saving: float = 1000.0) -> SavingsProposal:
|
||||
return SavingsProposal(
|
||||
project_id="p",
|
||||
|
|
@ -121,23 +117,27 @@ def _proposal(code: str, *, saving: float = 1000.0) -> SavingsProposal:
|
|||
# ---------------------------------------------------------------------------------- known positives
|
||||
|
||||
|
||||
def test_the_section_number_that_reached_validated_on_n500_is_refused() -> None:
|
||||
"""(a) tunnel-04's ``10.4``, replayed against the base that run actually opened."""
|
||||
proposal = _recorded(_ROUND3 / "tunnel-hauglia-2027-04-a4-enhetspris-ventilator-proposal.json")
|
||||
assert [i.code for i in proposal.affected_items] == ["10.4"]
|
||||
outcome = validate_proposal(proposal, baseline=None, grounding=_grounding_over("n500-2024"))
|
||||
def test_a_section_prefix_of_the_requirements_base_is_refused() -> None:
|
||||
"""(a) The ``10.4`` shape: a section PREFIX that stands in the base, is not inert, and is no
|
||||
document's own number. The CONTROL is what makes the refusal this rule's and not another
|
||||
stage's: the prefix IS in the text, in too few documents to be inert, and not declared."""
|
||||
grounding = _grounding_over("driftskrav-2027")
|
||||
assert "4.2.5" in grounding.text and "4.2.5" not in grounding.reference_vocabulary
|
||||
assert grounding.document_frequency("4.2.5") == 3
|
||||
outcome = validate_proposal(_proposal("4.2.5"), baseline=None, grounding=grounding)
|
||||
assert isinstance(outcome, Rejection)
|
||||
assert "'10.4'" in outcome.reason
|
||||
assert "not one of the 365 this knowledge base declares" in outcome.reason
|
||||
assert "'4.2.5'" in outcome.reason
|
||||
assert "not one of the 279 this knowledge base declares" in outcome.reason
|
||||
|
||||
|
||||
def test_the_process_number_that_reached_validated_on_r761_is_refused() -> None:
|
||||
"""(b) lindaas a4's ``1.10.4`` — a chapter of ANOTHER standard, quoted in one r761 document."""
|
||||
proposal = _recorded(_P17B / "lindaas-01-vegnormal-r761-2025-a4-indeksregulering-proposal.json")
|
||||
assert [i.code for i in proposal.affected_items] == ["1.10.4"]
|
||||
outcome = validate_proposal(proposal, baseline=None, grounding=_grounding_over("r761-2025"))
|
||||
def test_a_clause_of_another_document_quoted_in_the_catalogue_is_refused() -> None:
|
||||
"""(b) The ``1.10.4`` shape: a clause of ANOTHER document, quoted once in the catalogue."""
|
||||
grounding = _grounding_over("prosesskatalog-2027")
|
||||
assert grounding.document_frequency("1.10.4") == 1
|
||||
assert "1.10.4" not in grounding.reference_vocabulary
|
||||
outcome = validate_proposal(_proposal("1.10.4"), baseline=None, grounding=grounding)
|
||||
assert isinstance(outcome, Rejection)
|
||||
assert "not one of the 2765 this knowledge base declares" in outcome.reason
|
||||
assert "not one of the 306 this knowledge base declares" in outcome.reason
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------- known negatives
|
||||
|
|
@ -146,25 +146,25 @@ def test_the_process_number_that_reached_validated_on_r761_is_refused() -> None:
|
|||
def test_a_process_code_the_base_declares_still_validates() -> None:
|
||||
"""(c) The arm that keeps this a rule about the corpus and not about shapes.
|
||||
|
||||
``12.1`` is ``contexts/kontrakt-sorasen-2027``'s own first code and a REAL declared
|
||||
``prosessnr`` of R761. Under the ordered rule it would have been refused; it must not be.
|
||||
``12.1`` is ``contexts/driftsavtale-2027``'s own first code and a declared ``prosessnr`` of
|
||||
the catalogue. Under the ordered rule it would have been refused; it must not be.
|
||||
"""
|
||||
grounding = _grounding_over("r761-2025")
|
||||
grounding = _grounding_over("prosesskatalog-2027")
|
||||
assert "12.1" in grounding.reference_vocabulary
|
||||
outcome = validate_proposal(_proposal("12.1"), baseline=None, grounding=grounding)
|
||||
assert isinstance(outcome, ValidatedProposal), getattr(outcome, "reason", "")
|
||||
|
||||
|
||||
def test_every_sorasen_code_is_in_the_bases_vocabulary() -> None:
|
||||
def test_every_driftsavtale_code_is_in_the_bases_vocabulary() -> None:
|
||||
"""(c) The whole context set, not one sample: five real codes, five declared numbers."""
|
||||
codes = [
|
||||
code
|
||||
for approach in json.loads(
|
||||
Path("contexts/kontrakt-sorasen-2027/mandate.json").read_text(encoding="utf-8")
|
||||
Path("contexts/driftsavtale-2027/mandate.json").read_text(encoding="utf-8")
|
||||
)["approaches"]
|
||||
for code in approach.get("affected_codes", [])
|
||||
]
|
||||
vocabulary = _grounding_over("r761-2025").reference_vocabulary
|
||||
vocabulary = _grounding_over("prosesskatalog-2027").reference_vocabulary
|
||||
shaped = [c for c in codes if has_requirement_form(c)]
|
||||
assert len(shaped) == 5, shaped
|
||||
assert [c for c in shaped if c not in vocabulary] == []
|
||||
|
|
@ -172,11 +172,9 @@ def test_every_sorasen_code_is_in_the_bases_vocabulary() -> None:
|
|||
|
||||
def test_the_fasit_references_are_classified_requirement() -> None:
|
||||
"""(g) known negative (c) of the order: a fasit reference IS a requirement, and says so."""
|
||||
fasit = json.loads(
|
||||
Path("contexts/kontrakt-sorasen-2027/fasit.json").read_text(encoding="utf-8")
|
||||
)
|
||||
fasit = json.loads(Path("contexts/driftsavtale-2027/fasit.json").read_text(encoding="utf-8"))
|
||||
refs = sorted({c["ref"] for entry in fasit["must_cite"] for c in entry["concepts"]})
|
||||
forms = classify_codes(refs, _grounding_over("r761-2025"))
|
||||
forms = classify_codes(refs, _grounding_over("prosesskatalog-2027"))
|
||||
assert set(forms.values()) == {"requirement"}, forms
|
||||
|
||||
|
||||
|
|
@ -203,16 +201,16 @@ def test_a_cost_line_identifier_is_not_requirement_shaped() -> None:
|
|||
"""(f) K2's 50 identifiers and this repo's own code: shape, measured."""
|
||||
assert not has_requirement_form("SHA-01")
|
||||
assert not has_requirement_form("ENERGI-TOTAL-EL")
|
||||
assert not has_requirement_form("65 ASFALTDEKKER")
|
||||
assert not has_requirement_form("65 LAGRINGSSYSTEMER")
|
||||
assert has_requirement_form("10.4") and has_requirement_form("Krav 4.1.2—1")
|
||||
|
||||
|
||||
def test_classify_codes_without_a_grounding_is_the_pre_p20_answer() -> None:
|
||||
"""(g) ``stress.py`` re-derives for runs written before the field existed."""
|
||||
assert classify_codes(["12.1", "SHA-01", "impulsventilator"]) == {
|
||||
assert classify_codes(["12.1", "SHA-01", "nødstrømsaggregat"]) == {
|
||||
"12.1": "identifier",
|
||||
"SHA-01": "identifier",
|
||||
"impulsventilator": "prose",
|
||||
"nødstrømsaggregat": "prose",
|
||||
}
|
||||
grounding = Grounding(documents=("x",), declared_references=("12.1",))
|
||||
assert classify_codes(["12.1"], grounding) == {"12.1": "requirement"}
|
||||
|
|
|
|||
|
|
@ -22,19 +22,19 @@ from portfolio_optimiser.retrieval import (
|
|||
def docs(tmp_path):
|
||||
d = tmp_path / "docs"
|
||||
d.mkdir()
|
||||
(d / "asphalt.txt").write_text(
|
||||
"Asphalt Ab11 unit rate renegotiation reduced the paving cost on the school stretch.",
|
||||
(d / "licence.txt").write_text(
|
||||
"Licence unit price renegotiation reduced the office-suite cost for the head office.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
(d / "drainage.txt").write_text(
|
||||
"Drainage length was reduced after a survey on the low-traffic section.",
|
||||
(d / "cabling.txt").write_text(
|
||||
"Cabling length was reduced after a survey on the low-use floor.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return d
|
||||
|
||||
|
||||
def test_locator_exactly_slices_snippet(docs) -> None:
|
||||
hits = retrieve("asphalt paving cost", str(docs), top_k=5)
|
||||
hits = retrieve("licence office-suite cost", str(docs), top_k=5)
|
||||
assert hits
|
||||
for h in hits:
|
||||
text = (docs / h.file).read_text(encoding="utf-8")
|
||||
|
|
@ -42,8 +42,8 @@ def test_locator_exactly_slices_snippet(docs) -> None:
|
|||
|
||||
|
||||
def test_retrieve_is_deterministic(docs) -> None:
|
||||
a = retrieve("drainage survey", str(docs), top_k=3)
|
||||
b = retrieve("drainage survey", str(docs), top_k=3)
|
||||
a = retrieve("cabling survey", str(docs), top_k=3)
|
||||
b = retrieve("cabling survey", str(docs), top_k=3)
|
||||
assert [(h.file, h.locator) for h in a] == [(h.file, h.locator) for h in b]
|
||||
|
||||
|
||||
|
|
@ -65,13 +65,13 @@ def test_parent_traversal_rejected(docs) -> None:
|
|||
def test_symlink_escape_is_not_read(tmp_path, docs) -> None:
|
||||
# A symlink INSIDE docs_dir pointing OUTSIDE must never be read (fail-closed).
|
||||
secret = tmp_path / "secret.txt"
|
||||
secret.write_text("TOPSECRET asphalt asphalt asphalt exfiltrated", encoding="utf-8")
|
||||
secret.write_text("TOPSECRET licence licence licence exfiltrated", encoding="utf-8")
|
||||
os.symlink(str(secret), str(docs / "link.txt"))
|
||||
# safe_resolve refuses the escaping symlink...
|
||||
with pytest.raises(PathSecurityError):
|
||||
safe_resolve(str(docs), "link.txt")
|
||||
# ...and retrieve never surfaces the outside content even on a matching query.
|
||||
hits = retrieve("asphalt", str(docs), top_k=10)
|
||||
hits = retrieve("licence", str(docs), top_k=10)
|
||||
assert all("TOPSECRET" not in h.snippet for h in hits)
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -49,7 +49,8 @@ from portfolio_optimiser.explore import DeclaredRequirement, ToolCall, navigator
|
|||
from portfolio_optimiser.mandate import Approach, Mandate, criteria_block
|
||||
|
||||
_EXAMPLES = Path(__file__).resolve().parents[1] / "shared" / "examples"
|
||||
_TUNNEL = _EXAMPLES / "tunnel-hauglia"
|
||||
_BUNDLES = Path(__file__).resolve().parents[1] / "src" / "portfolio_optimiser" / "data" / "bundles"
|
||||
_KJOLING = _BUNDLES / "driftssenter-kjoling"
|
||||
_BYGG = _EXAMPLES / "bygg-energi-mikro"
|
||||
_PID = "BYGG-KONTOR-NORD"
|
||||
|
||||
|
|
@ -84,7 +85,7 @@ def _declare(
|
|||
|
||||
|
||||
def _base_with_a_requirement(root: Path) -> Path:
|
||||
"""A base declaring its own ``req_number`` — the form the delivered N corpora carry.
|
||||
"""A base declaring its own ``req_number`` — the form the requirement bases carry.
|
||||
|
||||
Crafted rather than taken from ``shared/examples``: MEASURED, no example base declares a
|
||||
``req_number`` at all, so an arm built on one of them could not tell a reply that reads the
|
||||
|
|
@ -96,8 +97,8 @@ def _base_with_a_requirement(root: Path) -> Path:
|
|||
"---\nbundle_id: n-mini\n---\n\n- [Krav](krav.md) — one requirement.\n", encoding="utf-8"
|
||||
)
|
||||
(base / "krav.md").write_text(
|
||||
"---\ntype: Krav\ntitle: Krav 4.1.2-1 Rundkjoring\nreq_number: Krav 4.1.2-1\n"
|
||||
"seksjon: '4.1.2'\n---\n\nEn rundkjoring skal ha …\n",
|
||||
"---\ntype: Krav\ntitle: Krav 4.1.2-1 Sikkerhetskopi\nreq_number: Krav 4.1.2-1\n"
|
||||
"seksjon: '4.1.2'\n---\n\nEn sikkerhetskopi skal ha …\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return base
|
||||
|
|
@ -113,7 +114,7 @@ def test_the_reply_carries_the_documents_own_title_and_number(tmp_path: Path) ->
|
|||
base = _base_with_a_requirement(tmp_path)
|
||||
tools, opened, declared = _wired(base)
|
||||
answer = _declare(tools, opened, "n-mini", "krav.md")
|
||||
assert answer["title"] == "Krav 4.1.2-1 Rundkjoring"
|
||||
assert answer["title"] == "Krav 4.1.2-1 Sikkerhetskopi"
|
||||
assert answer["req_number"] == "Krav 4.1.2-1"
|
||||
assert "is the requirement the proposal rests on" in answer["binds"]
|
||||
assert declared == [
|
||||
|
|
@ -145,16 +146,16 @@ def test_a_verdict_document_is_never_named_back_by_title(tmp_path: Path) -> None
|
|||
assert answer["title"] == "", "the verdict layer must not be readable through this answer"
|
||||
# The control: the SAME base answers the concept's title, so the empty string above is the
|
||||
# layer being excluded rather than the reader being broken.
|
||||
assert _declare(tools, opened, "n-mini", "krav.md")["title"] == "Krav 4.1.2-1 Rundkjoring"
|
||||
assert _declare(tools, opened, "n-mini", "krav.md")["title"] == "Krav 4.1.2-1 Sikkerhetskopi"
|
||||
|
||||
|
||||
def test_the_tunnel_example_answers_its_own_title_and_no_number() -> None:
|
||||
def test_the_cooling_example_answers_its_own_title_and_no_number() -> None:
|
||||
"""(a) control on a REAL example base: title present, reference number honestly absent."""
|
||||
tools, opened, _declared = _wired(_TUNNEL)
|
||||
files = [f.name for f in okf.navigate_bundle(str(_TUNNEL)).context_files]
|
||||
tools, opened, _declared = _wired(_KJOLING)
|
||||
files = [f.name for f in okf.navigate_bundle(str(_KJOLING)).context_files]
|
||||
path = files[0]
|
||||
answer = _declare(tools, opened, "tunnel-hauglia", path, also=files[1:3])
|
||||
assert answer["title"].startswith("Kilder: tunnelbelysning")
|
||||
answer = _declare(tools, opened, "driftssenter-kjoling", path, also=files[1:3])
|
||||
assert answer["title"].startswith("Kilder: kjøling")
|
||||
assert answer["req_number"] == "", "this base declares none, and the reply must not invent one"
|
||||
|
||||
|
||||
|
|
@ -250,7 +251,7 @@ def test_both_descriptions_name_the_filter_route() -> None:
|
|||
instruction = explore._INSTRUCTIONS[explore.HYPOTHESISER_ROLE]
|
||||
assert "read_dir a 'filter' word" in instruction
|
||||
assert "Krav 4.1.2-1" in instruction, "the worked example is what makes the route concrete"
|
||||
tools = {t.name: t for t in navigator_tools((str(_TUNNEL),), opened=[], requirements=[])}
|
||||
tools = {t.name: t for t in navigator_tools((str(_KJOLING),), opened=[], requirements=[])}
|
||||
description = tools["declare_requirement"].description or ""
|
||||
assert "filter='rundkjoring'" in description
|
||||
assert "filter='sikkerhetskopi'" in description
|
||||
assert "own title and number" in description
|
||||
|
|
|
|||
|
|
@ -55,7 +55,8 @@ _ATTEST_FILE = "attestering.txt"
|
|||
#: carries. Everything this file asserts is counted from HERE, never from the builder's answer.
|
||||
#:
|
||||
#: The REJECTED row carries an amount ON PURPOSE, and no archived run does: measured 19.09 over
|
||||
#: the four real outboxes (``tunnel-hauglia-2027-04/-06/-07/-08``), every ``rejected`` row has
|
||||
#: the four real outboxes of one archived context set (run ids ``…-04/-06/-07/-08``), every
|
||||
#: ``rejected`` row has
|
||||
#: ``saving_nok = None``. The guard that only a ``validated`` row's figure may become
|
||||
#: ``validated_nok`` is therefore aimed at a coverage writer that does not exist yet — and a
|
||||
#: fixture that cannot produce the number cannot witness the guard at all, which is exactly why
|
||||
|
|
@ -66,40 +67,40 @@ _ATTEST_FILE = "attestering.txt"
|
|||
#: 137000001. One øre is the whole distance between the rule and its most plausible violation.
|
||||
_SPEC: tuple[tuple[str, str, str, float | None, float, str], ...] = (
|
||||
(
|
||||
"a1-kortere-rekkverk",
|
||||
"Kortere rekkverk langs fv. 12",
|
||||
"a1-faerre-rackskap",
|
||||
"Færre rackskap i rad 12",
|
||||
"validated",
|
||||
1_250_000.0,
|
||||
1_250_000.0,
|
||||
"",
|
||||
),
|
||||
(
|
||||
"a2-faerre-kummer",
|
||||
"Færre kummer i kryssene",
|
||||
"a2-faerre-svitsjer",
|
||||
"Færre svitsjer i etasjene",
|
||||
"rejected",
|
||||
900_000.0,
|
||||
900_000.0,
|
||||
"claimed saving 900000 exceeds P90 feasible 410000",
|
||||
),
|
||||
(
|
||||
"a3-tynnere-dekke",
|
||||
"Tynnere dekke på gang- og sykkelvegen",
|
||||
"a3-mindre-testlagring",
|
||||
"Mindre lagring i testmiljøet",
|
||||
"validated",
|
||||
60_000.005,
|
||||
60_000.005,
|
||||
"",
|
||||
),
|
||||
(
|
||||
"a4-smalere-skulder",
|
||||
"Smalere skulder på strekningen",
|
||||
"a4-kortere-reservetid",
|
||||
"Kortere reservetid på strømforsyningen",
|
||||
"unsupported",
|
||||
None,
|
||||
400_000.0,
|
||||
UNSUPPORTED_REASON,
|
||||
),
|
||||
(
|
||||
"a5-enklere-rekkverksender",
|
||||
"Enklere rekkverksender",
|
||||
"a5-enklere-kabelgater",
|
||||
"Enklere kabelgater",
|
||||
"validated",
|
||||
60_000.005,
|
||||
60_000.005,
|
||||
|
|
@ -107,7 +108,7 @@ _SPEC: tuple[tuple[str, str, str, float | None, float, str], ...] = (
|
|||
),
|
||||
(
|
||||
"a6-enklere-belysning",
|
||||
"Enklere belysning i krysset",
|
||||
"Enklere belysning i serverrommet",
|
||||
"not_evaluated",
|
||||
None,
|
||||
0.0,
|
||||
|
|
@ -147,7 +148,7 @@ _STATUS_WORDS: dict[str, str] = {
|
|||
#:
|
||||
#: Measured 19.09: swapping the sentences for ``stage4-p90`` and ``stage0b-grounding`` in
|
||||
#: ``round_builder.STAGE_PROSE`` passed all 43 arms of this file. The report then told the expert
|
||||
#: that «Færre kummer i kryssene» fell on the GROUNDING when it fell on the uncertainty
|
||||
#: that «Færre svitsjer i etasjene» fell on the GROUNDING when it fell on the uncertainty
|
||||
#: calculation — the one thing that section exists to say, said wrong, with nothing to catch it.
|
||||
#:
|
||||
#: Both halves are independent for a reason. The opening clause is what the expert reads as the
|
||||
|
|
@ -168,8 +169,8 @@ _STAGE_SENTENCES: dict[str, tuple[str, str]] = {
|
|||
#: Every artefact type OF THE RUNS MEASURED — the denominator, counted here rather than
|
||||
#: remembered, and deliberately not called «every type a real run leaves». METHOD: each file
|
||||
#: ``<run_id>-<rest>.json`` is typed as ``proposal``/``outcome`` when ``<rest>`` ends there and
|
||||
#: as ``<rest>`` itself otherwise. Counted 19.09 over the four archived runs
|
||||
#: (``scratchpad/p19..p22-stress/tunnel-hauglia-2027/``, run ids ``…-04/-06/-07/-08``): ``-06``,
|
||||
#: as ``<rest>`` itself otherwise. Counted 19.09 over the four archived runs of one context set
|
||||
#: (``scratchpad/p19..p22-stress/<set>/``, run ids ``…-04/-06/-07/-08``): ``-06``,
|
||||
#: ``-07`` and ``-08`` hold all SEVEN (15 files each); ``-04`` holds six (12 files, no
|
||||
#: ``parse-failures`` — that file is written only when something failed to parse, so its ABSENCE
|
||||
#: is the signal). Counted again over EVERY outbox in this repo holding exactly one coverage
|
||||
|
|
@ -226,7 +227,7 @@ _SHARED_CITED = 9
|
|||
|
||||
|
||||
def _snippet(aid: str, k: int) -> str:
|
||||
return f"Kravteksten bak {aid}, sted {k}: restriktiv bruk av kryss anbefales."
|
||||
return f"Kravteksten bak {aid}, sted {k}: restriktiv bruk av reservekapasitet anbefales."
|
||||
|
||||
|
||||
def _stamp(decision: str, aid: str, *, shared: bool = False) -> ProvenanceStamp:
|
||||
|
|
@ -241,7 +242,7 @@ def _stamp(decision: str, aid: str, *, shared: bool = False) -> ProvenanceStamp:
|
|||
return ProvenanceStamp(
|
||||
citations=[
|
||||
Citation(
|
||||
file=f"krav/N100/id-{'hele-kjoringen' if shared else aid}-{k}.md",
|
||||
file=f"krav/D100/id-{'hele-kjoringen' if shared else aid}-{k}.md",
|
||||
locator=TextSpan(start_index=0, end_index=48),
|
||||
snippet=_snippet("hele-kjoringen" if shared else aid, k),
|
||||
)
|
||||
|
|
@ -319,7 +320,7 @@ def _outbox(
|
|||
str(outbox),
|
||||
run_id,
|
||||
tool_calls=[{"tool": "read_dir", "argument": "krav/", "round": 1}],
|
||||
requirements=[{"id": "N100-1", "title": "Kryss"}],
|
||||
requirements=[{"id": "D100-1", "title": "Klient"}],
|
||||
)
|
||||
write_parse_failures(
|
||||
str(outbox),
|
||||
|
|
@ -483,7 +484,7 @@ def test_a_refused_rows_amount_never_becomes_a_validated_saving(tmp_path: Path)
|
|||
refused_with_amount = {
|
||||
aid for aid, _l, status, nok, *_rest in _SPEC if status != "validated" and nok is not None
|
||||
}
|
||||
assert refused_with_amount == {"a2-faerre-kummer"}, "the table lost the row this arm needs"
|
||||
assert refused_with_amount == {"a2-faerre-svitsjer"}, "the table lost the row this arm needs"
|
||||
rows = {
|
||||
row["id"]: row
|
||||
for row in json.loads((built.round_dir / "outcome.json").read_text(encoding="utf-8"))[
|
||||
|
|
@ -524,7 +525,7 @@ def test_only_a_validated_row_counts_even_when_a_refused_one_carries_a_figure()
|
|||
for row in leaky["approaches"]
|
||||
if not row["validated"] and row["validated_nok"] is not None
|
||||
]
|
||||
assert leaked_rows == ["a2-faerre-kummer"], "the table lost the row this arm needs"
|
||||
assert leaked_rows == ["a2-faerre-svitsjer"], "the table lost the row this arm needs"
|
||||
assert rb.validated_ore(leaky) == _VALIDATED_ORE == 137_000_002
|
||||
|
||||
|
||||
|
|
@ -962,7 +963,7 @@ def test_the_report_shows_each_proposals_source_and_how_many_places_it_cited(
|
|||
for aid in _EVALUATED:
|
||||
assert f"Kilde (1 av {_CITED[aid]} siterte steder)" in text, aid
|
||||
assert _snippet(aid, 0) in text, aid
|
||||
assert f"krav/N100/id-{aid}-0.md" in text, aid
|
||||
assert f"krav/D100/id-{aid}-0.md" in text, aid
|
||||
assert text.count("Kilde (1 av ") == len(_EVALUATED)
|
||||
assert "Felles kilde" not in text
|
||||
|
||||
|
|
@ -1320,7 +1321,7 @@ def test_the_ledgers_claim_about_the_five_other_types_is_what_the_repo_measures(
|
|||
times, and the third writing was still untrue: «``exploration``, ``prepass``, ``multibase``,
|
||||
``plan-review`` og ``proposal-reviews`` … ingen av dem ligger i dag i en utboks som også har
|
||||
coverage». Measured 19.09 in this working tree: ``multibase`` lies in FOUR outboxes that also
|
||||
hold coverage (``scratchpad/{p17b-multibase,p20-stress,p21-stress,p22-stress}/lindaas``), and
|
||||
hold coverage (the multi-base set's outbox under ``scratchpad/`` in four stress rounds), and
|
||||
each of those outboxes holds TWO coverage files — which is why they fall outside the «exactly
|
||||
one coverage» rule the union of seven is counted over, and at the same time why seven is a
|
||||
FLOOR and not a ceiling.
|
||||
|
|
|
|||
|
|
@ -267,7 +267,7 @@ def test_the_judge_attributes_addressed_declarations_and_labels_legacy_ones() ->
|
|||
async def test_a_run_offered_no_declaration_rung_keeps_the_validators_ruling(
|
||||
tmp_path: Path,
|
||||
) -> None:
|
||||
"""The road path holds no knowledge base, so no ``declare_requirement`` exists there: nothing
|
||||
"""The reference path holds no knowledge base, so no ``declare_requirement`` exists there: nothing
|
||||
could have been declared, and the rule stays out of it. Drives the REAL ``run_project`` —
|
||||
the arm above only proves ``_evaluate_mandate`` honours ``declared=None``, not that the run
|
||||
passes it."""
|
||||
|
|
|
|||
|
|
@ -137,7 +137,7 @@ def _met_portfolio_goal(tmp_path: Path) -> tuple[Path, Path]:
|
|||
led = SavingsLedger()
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="FV42-GSV-E1",
|
||||
project_id="KONTOR-IT-E1",
|
||||
dimension="energi",
|
||||
candidate_identity="c1",
|
||||
amount_ore=1,
|
||||
|
|
@ -276,12 +276,12 @@ def test_verdict_dir_ingested_at_main_level_offline(tmp_path, capsys) -> None:
|
|||
|
||||
def _report_ledger(tmp_path: Path) -> Path:
|
||||
"""A saved ledger with >=2 projects and one cross-dimension overlap (``c-a`` under both
|
||||
``energi`` and ``asfalt`` in FV42 -> counted once, flagged), for the value-report arms.
|
||||
``energi`` and ``lisens`` in KONTOR -> counted once, flagged), for the value-report arms.
|
||||
portfolio_total = 1234567 (overlap once) + 500000 = 1734567 øre."""
|
||||
led = SavingsLedger()
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="FV42-GSV-E1",
|
||||
project_id="KONTOR-IT-E1",
|
||||
dimension="energi",
|
||||
candidate_identity="c-a",
|
||||
amount_ore=1234567,
|
||||
|
|
@ -291,17 +291,17 @@ def _report_ledger(tmp_path: Path) -> Path:
|
|||
)
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="FV42-GSV-E1",
|
||||
dimension="asfalt",
|
||||
project_id="KONTOR-IT-E1",
|
||||
dimension="lisens",
|
||||
candidate_identity="c-a",
|
||||
amount_ore=1234567,
|
||||
verdict_id="v2",
|
||||
provenance="p2", # cross-dimension overlap on (FV42-GSV-E1, c-a)
|
||||
provenance="p2", # cross-dimension overlap on (KONTOR-IT-E1, c-a)
|
||||
)
|
||||
)
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="RV13-RAS-TP",
|
||||
project_id="NETT-SIKR-TP",
|
||||
dimension="energi",
|
||||
candidate_identity="c-b",
|
||||
amount_ore=500000,
|
||||
|
|
@ -320,7 +320,7 @@ def test_report_prints_table_rc0(tmp_path, capsys) -> None:
|
|||
rc = run.main(["--report", "--ledger", str(_report_ledger(tmp_path))])
|
||||
out = capsys.readouterr().out
|
||||
assert rc == 0
|
||||
assert "FV42-GSV-E1" in out # a per-project row
|
||||
assert "KONTOR-IT-E1" in out # a per-project row
|
||||
assert "17\xa0345,67\xa0kr" in out # portfolio total 1734567 øre
|
||||
|
||||
|
||||
|
|
@ -332,9 +332,9 @@ def test_report_json_rc0_parses_rollup(tmp_path, capsys) -> None:
|
|||
assert rc == 0
|
||||
payload = json.loads(out)
|
||||
assert payload["portfolio_total_ore"] == 1734567
|
||||
assert payload["per_project"]["FV42-GSV-E1"] == 1234567
|
||||
assert payload["per_project"]["KONTOR-IT-E1"] == 1234567
|
||||
assert isinstance(payload["overlaps"], list)
|
||||
assert ["FV42-GSV-E1", "c-a"] in payload["overlaps"] # tuple serialized as a JSON array
|
||||
assert ["KONTOR-IT-E1", "c-a"] in payload["overlaps"] # tuple serialized as a JSON array
|
||||
assert isinstance(payload["provenance"], list)
|
||||
assert len(payload["provenance"]) == 3 # one ProvenanceLine per ledger entry
|
||||
|
||||
|
|
|
|||
|
|
@ -76,12 +76,12 @@ def _isolate_model_env(monkeypatch: pytest.MonkeyPatch) -> None:
|
|||
|
||||
|
||||
def _marker_ledger(tmp_path: Path) -> Path:
|
||||
"""A ledger with exactly one realized entry totalling the marker øre against FV42-GSV-E1 —
|
||||
reused by both arms (portfolio_total == per_project_total('FV42-GSV-E1') == 13731)."""
|
||||
"""A ledger with exactly one realized entry totalling the marker øre against KONTOR-IT-E1 —
|
||||
reused by both arms (portfolio_total == per_project_total('KONTOR-IT-E1') == 13731)."""
|
||||
led = SavingsLedger()
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="FV42-GSV-E1",
|
||||
project_id="KONTOR-IT-E1",
|
||||
dimension="energi",
|
||||
candidate_identity="c-marker",
|
||||
amount_ore=_MARKER_ORE,
|
||||
|
|
@ -118,10 +118,10 @@ def test_per_project_hard_goal_control_distinguishes_scope(tmp_path, capsys) ->
|
|||
no client built) → scope=project, stopped_early=False, and NO portfolio-scope line. Proves the
|
||||
printed fields are the actual run_portfolio values, not a canned string."""
|
||||
goals = _write_goals(
|
||||
tmp_path, {"per_project": {"FV42-GSV-E1": {"absolute_ore": _MARKER_ORE, "mode": "hard"}}}
|
||||
tmp_path, {"per_project": {"KONTOR-IT-E1": {"absolute_ore": _MARKER_ORE, "mode": "hard"}}}
|
||||
)
|
||||
ledger = _marker_ledger(tmp_path)
|
||||
rc = run.main(["FV42-GSV-E1", "--portfolio", "--goals", str(goals), "--ledger", str(ledger)])
|
||||
rc = run.main(["KONTOR-IT-E1", "--portfolio", "--goals", str(goals), "--ledger", str(ledger)])
|
||||
assert rc == 0
|
||||
out = capsys.readouterr().out
|
||||
assert "scope=project" in out
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@ from portfolio_optimiser.run import run_project
|
|||
from portfolio_optimiser.validator import ValidatedProposal
|
||||
|
||||
_VALID = (
|
||||
'{"project_id":"FV42-GSV-E1","measure":"Reduce scope",'
|
||||
'{"project_id":"KONTOR-IT-E1","measure":"Reduce scope",'
|
||||
'"affected_items":[{"code":"05.2","quantity":4300,"unit_cost":215},'
|
||||
'{"code":"03.1","quantity":1800,"unit_cost":310}],"claimed_saving_nok":200000}'
|
||||
)
|
||||
|
|
@ -21,7 +21,7 @@ async def test_run_project_smoke(tmp_path) -> None:
|
|||
docs = tmp_path / "docs"
|
||||
docs.mkdir()
|
||||
(docs / "cost.txt").write_text(
|
||||
"Asphalt Ab11 unit rate renegotiation reduced the paving cost on the school stretch.",
|
||||
"Licence unit price renegotiation reduced the office-suite cost for the head office.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
|
|
@ -29,7 +29,7 @@ async def test_run_project_smoke(tmp_path) -> None:
|
|||
return FakeChatClient(default_reply=_VALID)
|
||||
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=str(docs),
|
||||
verdict_input={"decision": "approved", "rationale": "feasible within range"},
|
||||
|
|
|
|||
|
|
@ -13,7 +13,7 @@ validates), and the no-baseline arm proves the argument stays OPTIONAL (``None``
|
|||
behaviour, which is why the existing suite stands).
|
||||
|
||||
Measured detach points (see the session log): the reconciliation stage · the magnitude tolerance ·
|
||||
the road-path wiring in ``run.py`` · the bundle-path wiring · the method-cap registry key (F8).
|
||||
the reference-path wiring in ``run.py`` · the bundle-path wiring · the method-cap registry key (F8).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -44,7 +44,7 @@ PRE_AMENDMENT_BUNDLE = _DATA / "bygg-energi-mikro-a"
|
|||
|
||||
_VERDICT_INPUT = {"decision": "approved", "rationale": "expert reviewed (sim)"}
|
||||
|
||||
# FV42-GSV-E1's real cost line 05.2 (Asfalt Ab11): 4300 m2 x 215 NOK. Affected total 924500 ->
|
||||
# KONTOR-IT-E1's real cost line 05.2 (Lisens kontorpakke): 4300 licences x 215 NOK. Affected total 924500 ->
|
||||
# degenerate P90 = 0.30 x 924500 = 277350, so claimed 200000 clears every pre-S4.0 stage.
|
||||
_REAL_CODE = "05.2"
|
||||
_REAL_QTY = 4300.0
|
||||
|
|
@ -61,13 +61,13 @@ _FAKE_CLAIM = 3_000_000.0
|
|||
|
||||
def _fv42_baseline() -> CostBaseline:
|
||||
return baseline_from_project(
|
||||
next(p for p in load_reference_projects() if p.id == "FV42-GSV-E1")
|
||||
next(p for p in load_reference_projects() if p.id == "KONTOR-IT-E1")
|
||||
)
|
||||
|
||||
|
||||
def _proposal(code: str, quantity: float, unit_cost: float, claimed: float) -> SavingsProposal:
|
||||
return SavingsProposal(
|
||||
project_id="FV42-GSV-E1",
|
||||
project_id="KONTOR-IT-E1",
|
||||
measure="Reduce scope",
|
||||
affected_items=[AffectedItem(code=code, quantity=quantity, unit_cost=unit_cost)],
|
||||
claimed_saving_nok=claimed,
|
||||
|
|
@ -184,7 +184,7 @@ def test_malformed_baseline_raises_even_on_the_optional_path(tmp_path) -> None:
|
|||
okf.load_optional_cost_baseline(str(tmp_path))
|
||||
|
||||
|
||||
# --- Arm 4: the run-path wiring (road + bundle) ---------------------------------------------------
|
||||
# --- Arm 4: the run-path wiring (reference + bundle) ---------------------------------------------------
|
||||
|
||||
|
||||
def _reply(code: str, quantity: float, unit_cost: float, claimed: float) -> str:
|
||||
|
|
@ -205,11 +205,11 @@ def _factory(reply: str):
|
|||
|
||||
|
||||
async def test_road_path_anchors_the_gate_to_the_reference_baseline(docs_dir, fresh_store) -> None:
|
||||
"""RED (road wiring): a numerically-impeccable fabricated cost line is REJECTED end-to-end
|
||||
through ``run_project``. Detach the road-path baseline (stop passing it) and the same run
|
||||
"""RED (reference wiring): a numerically-impeccable fabricated cost line is REJECTED end-to-end
|
||||
through ``run_project``. Detach the reference-path baseline (stop passing it) and the same run
|
||||
returns a ValidatedProposal."""
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
|
|
@ -221,9 +221,9 @@ async def test_road_path_anchors_the_gate_to_the_reference_baseline(docs_dir, fr
|
|||
|
||||
|
||||
async def test_road_path_control_real_line_validates(docs_dir, fresh_store) -> None:
|
||||
"""Causality control for the road wiring: the real 05.2 line validates through the same path."""
|
||||
"""Causality control for the reference wiring: the real 05.2 line validates through the same path."""
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VERDICT_INPUT,
|
||||
|
|
@ -280,13 +280,13 @@ def test_method_cap_is_keyed_by_config_not_a_hardcoded_string() -> None:
|
|||
cap for a measure with no built-in entry rejects a proposal the generic P90 stage passes — so
|
||||
the rule is keyed by configuration, not by the ``energy_efficiency`` literal."""
|
||||
proposal = SavingsProposal(
|
||||
project_id="P-ASFALT",
|
||||
measure="asfalt_reduction",
|
||||
project_id="P-LISENS",
|
||||
measure="lisens_reduction",
|
||||
affected_items=[AffectedItem(code="05.2", quantity=100000, unit_cost=1.0)],
|
||||
claimed_saving_nok=20000,
|
||||
assumptions={},
|
||||
)
|
||||
assert isinstance(validate_proposal(proposal), ValidatedProposal) # generic P90 = 30000
|
||||
capped = validate_proposal(proposal, method_caps={"asfalt_reduction": 0.10})
|
||||
capped = validate_proposal(proposal, method_caps={"lisens_reduction": 0.10})
|
||||
assert isinstance(capped, Rejection)
|
||||
assert "method cap" in capped.reason
|
||||
|
|
|
|||
|
|
@ -1,10 +1,10 @@
|
|||
"""MAJOR-2 (docs/2026-08-25-syretest-vei-ab.md) — ``--explore --scripted-replies`` must not
|
||||
"""MAJOR-2 (see docs/invarianter.md) — ``--explore --scripted-replies`` must not
|
||||
crash with a raw ``KeyError: 'navigator'``.
|
||||
|
||||
``_SCRIPTED_ROLES = ("proposer", "checker")`` (``run.py:1531``) is the debate's two roles.
|
||||
``explore()`` asks the SAME ``client_factory`` for three more: ``manager``, ``navigator``,
|
||||
``hypothesiser`` (``explore.py:576``). ``_load_scripted_replies`` was fail-fast for the two roles
|
||||
it knew about — measured (docs/2026-08-25-syretest-vei-ab.md § MAJOR-2) to let the three it did not
|
||||
it knew about — measured (MAJOR-2) to let the three it did not
|
||||
know about surface exactly the ``KeyError`` deep inside ``scripted_factory``'s lookup its own
|
||||
docstring warns against, mid-run, after the banner had already printed.
|
||||
|
||||
|
|
|
|||
|
|
@ -50,7 +50,7 @@ def _features(
|
|||
codes: frozenset[str] = frozenset({"05.2", "03.1"}),
|
||||
measure_type: str = "scope_reduction",
|
||||
saving: float = 220_000.0,
|
||||
description: str = "asphalt base course reduction near school",
|
||||
description: str = "licence scope reduction near head office",
|
||||
) -> types.SimpleNamespace:
|
||||
"""Duck-typed stand-in for ``ProposalFeatures`` — semretrieval never imports ``verdicts``
|
||||
at runtime, so the unit tests do not need the real (MAF-bound) type either."""
|
||||
|
|
@ -134,7 +134,7 @@ def test_fake_embedder_is_bit_stable_across_processes(tmp_path: Path) -> None:
|
|||
" affected_codes=frozenset({'05.2', '03.1'}),\n"
|
||||
" measure_type='scope_reduction',\n"
|
||||
" claimed_saving_nok=220000.0,\n"
|
||||
" description='asphalt base course reduction near school',\n"
|
||||
" description='licence scope reduction near head office',\n"
|
||||
")\n"
|
||||
f"np.save({str(there)!r}, FakeEmbedder()(f))\n"
|
||||
)
|
||||
|
|
@ -197,7 +197,7 @@ _QUERY = ProposalFeatures(
|
|||
affected_codes=frozenset({"05.2", "03.1"}),
|
||||
measure_type="scope_reduction",
|
||||
claimed_saving_nok=220_000, # bucket [100k, 500k)
|
||||
description="asphalt base course reduction near school",
|
||||
description="licence scope reduction near head office",
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -222,7 +222,7 @@ def _store_with_true_match_and_decoys() -> tuple[VerdictStore, str]:
|
|||
affected_codes=frozenset({"09.1"}),
|
||||
measure_type="rate_renegotiation",
|
||||
claimed_saving_nok=50_000,
|
||||
description="asphalt base course reduction near school",
|
||||
description="licence scope reduction near head office",
|
||||
),
|
||||
decision="rejected",
|
||||
rationale="surface-text decoy",
|
||||
|
|
@ -233,7 +233,7 @@ def _store_with_true_match_and_decoys() -> tuple[VerdictStore, str]:
|
|||
affected_codes=frozenset({"21.2"}),
|
||||
measure_type="material_substitution",
|
||||
claimed_saving_nok=700_000,
|
||||
description="asphalt base course reduction extra words",
|
||||
description="licence scope reduction extra words",
|
||||
),
|
||||
decision="rejected",
|
||||
rationale="surface-text decoy",
|
||||
|
|
|
|||
|
|
@ -252,7 +252,7 @@ def test_semretrieval_imports_no_network_modules() -> None:
|
|||
# minting paths emit (``run._features_of`` sets ``description=proposal.measure``;
|
||||
# ``verdicts.features_from_ir`` sets ``description=ir["measure"]``).
|
||||
|
||||
_MEASURE = "asfalt"
|
||||
_MEASURE = "lisens"
|
||||
_QUERY_CODES = frozenset({"05.1", "05.2"})
|
||||
_CORRECT_CODES = frozenset({"05.1", "07.4"})
|
||||
_DISTRACTOR_CODES = frozenset({"05.2", "09.8"})
|
||||
|
|
@ -381,10 +381,10 @@ def test_minted_and_authored_verdicts_embed_identically_on_a_structural_tie() ->
|
|||
authored = verdict_from_dict(
|
||||
{
|
||||
**verdict_to_dict(minted),
|
||||
"rationale": "the base course reduction was accepted on the school approach",
|
||||
"rationale": "the licence scope reduction was accepted at the head office",
|
||||
"proposal_features": {
|
||||
**verdict_to_dict(minted)["proposal_features"],
|
||||
"description": "shallower asphalt base layer along the school approach",
|
||||
"description": "smaller licence tier across the head office",
|
||||
},
|
||||
}
|
||||
)
|
||||
|
|
|
|||
|
|
@ -105,9 +105,9 @@ def _bundle_with_patched_verdict(tmp_path: Path, old: str, new: str) -> str:
|
|||
frontmatter rewritten — the packaged fixture is never mutated."""
|
||||
dst = tmp_path / "bundle"
|
||||
shutil.copytree(_MULTI_BUNDLE, dst)
|
||||
target = dst / "verdict-b-asfalt.md"
|
||||
target = dst / "verdict-b-lisens.md"
|
||||
text = target.read_text(encoding="utf-8")
|
||||
assert old in text, f"fixture drift: {old!r} not in verdict-b-asfalt.md"
|
||||
assert old in text, f"fixture drift: {old!r} not in verdict-b-lisens.md"
|
||||
target.write_text(text.replace(old, new, 1), encoding="utf-8")
|
||||
return str(dst)
|
||||
|
||||
|
|
|
|||
|
|
@ -52,7 +52,7 @@ from portfolio_optimiser.validator import (
|
|||
validate_proposal,
|
||||
)
|
||||
|
||||
# Same fixture arithmetic as test_step5_refine_loadbearing: FV42-GSV-E1 codes 05.2 + 03.1 ->
|
||||
# Same fixture arithmetic as test_step5_refine_loadbearing: KONTOR-IT-E1 codes 05.2 + 03.1 ->
|
||||
# affected total 1_482_500 -> degenerate Monte Carlo P90 = 444_750. 800_000 parses (< affected
|
||||
# total) but exceeds P90 -> rejected; 200_000 validates. Both proposals are built via proposal_for,
|
||||
# so their quantity/unit_cost ARE the project's cost lines and the S4.0 baseline stage passes.
|
||||
|
|
@ -73,7 +73,7 @@ def _fixture() -> tuple[Project, str, str, str, Rejection]:
|
|||
flip key (the rejected claim value, which appears ONLY once the validator's reason is fed back),
|
||||
and the rejection the validator itself produces for the bad proposal — computed with the SAME
|
||||
validator the SUT uses, so the reason is byte-identical to the one the loop feeds back."""
|
||||
project = load_reference_projects()[0] # FV42-GSV-E1
|
||||
project = load_reference_projects()[0] # KONTOR-IT-E1
|
||||
bad = proposal_for(project, _CODES, claimed_saving_nok=_BAD_CLAIM)
|
||||
corrected = proposal_for(project, _CODES, claimed_saving_nok=_CORRECTED_CLAIM)
|
||||
rej = validate_proposal(bad)
|
||||
|
|
@ -140,7 +140,7 @@ async def test_run_result_carries_the_falsification_history(tmp_path: Path) -> N
|
|||
"""RUN-LEVEL WIRING: what generation returns must survive to ``RunResult`` — otherwise the seam
|
||||
exists but Step 5 is still invisible to the demo. RED if run.py drops ``refinements``.
|
||||
|
||||
Drives the ROAD path (a reference project, so the S4.0 baseline is always anchored) with the
|
||||
Drives the REFERENCE path (a reference project, so the S4.0 baseline is always anchored) with the
|
||||
same content-keyed proposer the simulation uses."""
|
||||
project, bad_json, corrected_json, flip_key, _rej = _fixture()
|
||||
|
||||
|
|
|
|||
|
|
@ -40,7 +40,7 @@ from portfolio_optimiser.validator import (
|
|||
validate_proposal,
|
||||
)
|
||||
|
||||
# FV42-GSV-E1 cost codes 05.2 + 03.1 -> affected total 1_482_500 -> degenerate Monte Carlo
|
||||
# KONTOR-IT-E1 cost codes 05.2 + 03.1 -> affected total 1_482_500 -> degenerate Monte Carlo
|
||||
# P90 = 0.30 x 1_482_500 = 444_750. A claim of 800_000 is Pydantic-parseable (< affected total)
|
||||
# but > P90 -> the validator REJECTS it (exercising the OUTER max_attempts path). A claim of
|
||||
# 200_000 (<= P90) validates. Both proposals are built via proposal_for + model_dump_json, so
|
||||
|
|
@ -96,7 +96,7 @@ async def test_refinement_feeds_prior_falsification_into_next_prompt() -> None:
|
|||
"""LOAD-BEARING: the validator's rejection from attempt 1 reaches attempt 2's prompt, so the
|
||||
proposer can correct. Goes RED if ``_build_messages`` stops injecting ``prior_rejection.reason``
|
||||
(the flip token never arrives -> no validated outcome AND the verbatim assertion fails)."""
|
||||
project = load_reference_projects()[0] # FV42-GSV-E1
|
||||
project = load_reference_projects()[0] # KONTOR-IT-E1
|
||||
bad = proposal_for(project, _CODES, claimed_saving_nok=_BAD_CLAIM)
|
||||
corrected = proposal_for(project, _CODES, claimed_saving_nok=_CORRECTED_CLAIM)
|
||||
|
||||
|
|
|
|||
|
|
@ -193,11 +193,11 @@ def test_promoted_verdict_about_another_candidate_keeps_its_own_key(tmp_path) ->
|
|||
id="STEG8-OTHER-CANDIDATE",
|
||||
proposal_features=ProposalFeatures(
|
||||
affected_codes=frozenset({"05.2", "03.1"}),
|
||||
measure_type="Redusert asfalttykkelse",
|
||||
measure_type="Redusert lisensomfang",
|
||||
claimed_saving_nok=900000,
|
||||
),
|
||||
decision="approved",
|
||||
rationale=f"asfalttiltak godkjent med kraftig realiseringskorreksjon ({other_marker})",
|
||||
rationale=f"lisenstiltak godkjent med kraftig realiseringskorreksjon ({other_marker})",
|
||||
)
|
||||
|
||||
path = promote_verdict(
|
||||
|
|
@ -206,7 +206,7 @@ def test_promoted_verdict_about_another_candidate_keeps_its_own_key(tmp_path) ->
|
|||
|
||||
fm = okf.parse_frontmatter(path)
|
||||
assert fm["affected_codes"] == "[03.1, 05.2]" # sorted -> deterministic bytes
|
||||
assert fm["measure_type"] == "Redusert asfalttykkelse"
|
||||
assert fm["measure_type"] == "Redusert lisensomfang"
|
||||
assert fm["claimed_saving_nok"] == "900000"
|
||||
|
||||
# The round trip: the next run's seed keys it on ITS candidate, so it does not surface for the
|
||||
|
|
@ -224,7 +224,7 @@ def test_promoted_verdict_about_another_candidate_keeps_its_own_key(tmp_path) ->
|
|||
(v.proposal_features.affected_codes, v.proposal_features.measure_type)
|
||||
for v in store.verdicts
|
||||
}
|
||||
assert (frozenset({"05.2", "03.1"}), "Redusert asfalttykkelse") in keys, (
|
||||
assert (frozenset({"05.2", "03.1"}), "Redusert lisensomfang") in keys, (
|
||||
f"the promoted verdict was seeded with the wrong structural key; got {keys}"
|
||||
)
|
||||
|
||||
|
|
|
|||
|
|
@ -14,17 +14,18 @@ tests`` hit only unrelated files).
|
|||
**THE ORDER'S (a) WAS VACUOUS AS WRITTEN, AND THAT IS MEASURED.** The order defines grounded as
|
||||
"a must_cite path was OPENED *or* CITED". But on the S2c navigation path ``run_project`` stamps
|
||||
``citations = bundle_citations(bundle)``, which is ONE CITATION PER CONTEXT FILE - the whole corpus.
|
||||
Measured on n100-2023: 446 context files, 446 citations, and **6 of 6 fasit paths already "cited"
|
||||
before a single model call**. A judge honouring the order literally would be a gate that can only
|
||||
be green, which is the repo's own vacuous-gate class, inside the gate built to stop it. So
|
||||
Measured on a requirements base during development: 446 context files, 446 citations, and **6
|
||||
of 6 fasit paths already "cited" before a single model call**. A judge honouring the order
|
||||
literally would be a gate that can only be green, which is the repo's own vacuous-gate class,
|
||||
inside the gate built to stop it. So
|
||||
``grounded`` counts a CITATION only when the citation list is NARROWER than the base (a declared
|
||||
pre-pass cut); a whole-base list is reported as such and carries nothing. Both halves are reported
|
||||
either way, so the operator can read which one fired - the deviation is stated, never silent.
|
||||
|
||||
**(b') was checked for the same vacuity and is CLEAN.** ``bundle_citations`` snippets are concept
|
||||
BODIES, and the ``ref``/``title`` live in FRONTMATTER: measured 0 of 446 n100 bodies contain
|
||||
``Krav 4.1.2-1``. So the snippet arm can carry (b') without being satisfied by construction, and
|
||||
the order's definition is kept. Which half fired is still reported.
|
||||
BODIES, and the ``ref``/``title`` live in FRONTMATTER: measured 0 of 446 bodies of that base
|
||||
contain its ``Krav 4.1.2-1``. So the snippet arm can carry (b') without being satisfied by
|
||||
construction, and the order's definition is kept. Which half fired is still reported.
|
||||
|
||||
**A denominator, always** (Verifiseringsloven ansikt 4): every verdict names how many tool calls,
|
||||
citations and approach rows it saw. An outbox with no artefacts RAISES rather than reporting
|
||||
|
|
@ -60,7 +61,7 @@ from portfolio_optimiser import stress
|
|||
_GOOD = "krav/N1/id-good.md"
|
||||
_OTHER = "krav/N1/id-other.md"
|
||||
_REF = "Krav 1.2.3-4"
|
||||
_TITLE = "Krav 1.2.3-4 Rundkjoringer"
|
||||
_TITLE = "Krav 1.2.3-4 Sikkerhetskopier"
|
||||
|
||||
|
||||
def _minibase(root: Path) -> Path:
|
||||
|
|
@ -272,8 +273,9 @@ def test_a_opening_some_other_document_does_not_ground_it(tmp_path: Path) -> Non
|
|||
|
||||
|
||||
def test_b_a_whole_base_citation_list_cannot_ground_an_approach(tmp_path: Path) -> None:
|
||||
"""Measured on n100-2023: 446 context files, 446 citations, 6/6 fasit paths 'cited' before
|
||||
any model call. Honouring the order literally would make (a) green by construction."""
|
||||
"""Measured on a requirements base: 446 context files, 446 citations, 6/6 fasit paths
|
||||
'cited' before any model call. Honouring the order literally would make (a) green by
|
||||
construction."""
|
||||
verdict = _judge(tmp_path, tool_calls=[])
|
||||
row = verdict.approaches[0]
|
||||
assert row.cited == (_GOOD,), "the path IS in the citation list"
|
||||
|
|
@ -331,8 +333,8 @@ def test_e_a_measure_that_names_nothing_is_carried_only_by_a_narrowed_snippet(
|
|||
def test_e_a_whole_base_snippet_does_not_name_the_concept(tmp_path: Path) -> None:
|
||||
"""P18/C2 (PM decision, P16 § 6.2). ``bundle_citations`` stamps EVERY context file before any
|
||||
model call, so a mark found in a whole-base snippet is evidence about what the base contains,
|
||||
not about what this run said. Measured on r761: the process number ``12.1`` stands in the
|
||||
bodies themselves, so that row came back ``named`` for a run that never named it.
|
||||
not about what this run said. Measured on a process catalogue: the process number ``12.1``
|
||||
stands in the bodies themselves, so that row came back ``named`` for a run that never named it.
|
||||
|
||||
The pair with the arm above is the discriminator: the SAME snippet, the same mark, and the only
|
||||
difference is the scope."""
|
||||
|
|
@ -660,10 +662,11 @@ def test_the_cli_refuses_to_guess_which_base_a_multi_base_outbox_is_for(
|
|||
# P21 B3: the judge reports what the run was ANCHORED on, which codes the project actually PRICES,
|
||||
# and WHICH falsifier caught the falsification arm.
|
||||
#
|
||||
# The measured reason. Rounds 1-4 all ran un-anchored — a vegnormal ships no ``cost-baseline.json``
|
||||
# and the only file loader read one out of the bundle — so stage 0 never spoke and the ``a4`` arm
|
||||
# fell, when it fell, on P7's grounding check. "It was refused" and "the stage that knows what this
|
||||
# project buys refused it" are different facts, and only the second is what anchoring bought.
|
||||
# The measured reason. Rounds 1-4 all ran un-anchored — a requirements base ships no
|
||||
# ``cost-baseline.json`` and the only file loader read one out of the bundle — so stage 0 never
|
||||
# spoke and the ``a4`` arm fell, when it fell, on P7's grounding check. "It was refused" and "the
|
||||
# stage that knows what this project buys refused it" are different facts, and only the second is
|
||||
# what anchoring bought.
|
||||
# --------------------------------------------------------------------------------------------
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -112,7 +112,7 @@ def _wire_reply(*, with_band: bool) -> str:
|
|||
)
|
||||
return json.dumps(
|
||||
{
|
||||
"project_id": "FV42-GSV-E1",
|
||||
"project_id": "KONTOR-IT-E1",
|
||||
"measure": _MEASURE,
|
||||
"affected_items": [{"code": _CODE, "quantity": _QUANTITY, "unit_cost": _UNIT_COST}],
|
||||
"claimed_saving_nok": _CLAIM,
|
||||
|
|
@ -326,7 +326,7 @@ async def test_band_round_trip_is_verbatim_and_the_ir_map_form_still_parses() ->
|
|||
|
||||
map_form = json.dumps(
|
||||
{
|
||||
"project_id": "FV42-GSV-E1",
|
||||
"project_id": "KONTOR-IT-E1",
|
||||
"measure": _MEASURE,
|
||||
"affected_items": [{"code": _CODE, "quantity": _QUANTITY, "unit_cost": _UNIT_COST}],
|
||||
"claimed_saving_nok": _CLAIM,
|
||||
|
|
|
|||
|
|
@ -3,10 +3,10 @@
|
|||
**DEL C, the measured silence.** P18 gave ``read_dir`` a window — ``filter`` / ``offset`` /
|
||||
``limit`` — and then measured its own paid round without being able to see it used: five of 31
|
||||
documents read lay OUTSIDE the default window, so the window had been widened, and the trace could
|
||||
not say with which knob (``docs/2026-09-14-p18-stressrunde-2.md`` § 1, finding 1). The recorder
|
||||
kept ``name`` / ``bundle_id`` / ``path`` and nothing else. "Did the model narrow the level, or page
|
||||
through it?" is the operative question about a corpus of 2 756 documents, and it was unanswerable
|
||||
from the artefact the run leaves behind.
|
||||
not say with which knob (stress round 2, finding 1; the ledger is ``docs/invarianter.md``). The
|
||||
recorder kept ``name`` / ``bundle_id`` / ``path`` and nothing else. "Did the model narrow the level,
|
||||
or page through it?" is the operative question about a corpus of 2 756 documents, and it was
|
||||
unanswerable from the artefact the run leaves behind.
|
||||
|
||||
**DEL D, and one of the two findings it fixes was WRONG AS WRITTEN.** P18's finding 4 said a
|
||||
successful run does not say what it used. Measured 14.09: ``provenance.token_usage`` has been
|
||||
|
|
@ -104,9 +104,12 @@ def _record(*calls: tuple[str, Any]) -> list[ex.ToolCall]:
|
|||
def test_the_window_arguments_are_recorded() -> None:
|
||||
"""(a) HOW the level was asked for, not only which one."""
|
||||
(call,) = _record(
|
||||
("read_dir", {"bundle_id": "k2", "path": "krav/N100", "filter": "rundkjoring", "limit": 25})
|
||||
(
|
||||
"read_dir",
|
||||
{"bundle_id": "k2", "path": "krav/D100", "filter": "sikkerhetskopi", "limit": 25},
|
||||
)
|
||||
)
|
||||
assert (call.filter, call.offset, call.limit) == ("rundkjoring", 0, 25)
|
||||
assert (call.filter, call.offset, call.limit) == ("sikkerhetskopi", 0, 25)
|
||||
|
||||
|
||||
def test_a_call_that_passes_none_of_them_yields_empty_fields() -> None:
|
||||
|
|
@ -316,7 +319,7 @@ def test_the_judge_reads_the_stop_reason_onto_every_unevaluated_row(tmp_path: Pa
|
|||
def test_the_judge_reports_what_the_run_spent(tmp_path: Path) -> None:
|
||||
"""(g) D1. P18's finding 4 was wrong as written: the field was always there, the READER was not.
|
||||
|
||||
The figure is the one measured on round 2's own artefacts (tunnel-hauglia-2027-02).
|
||||
The figure is the one measured on round 2's own artefacts (one context set's ``…-02`` run).
|
||||
"""
|
||||
verdict = stress.score_context_set(
|
||||
_context_dir(tmp_path),
|
||||
|
|
|
|||
|
|
@ -235,7 +235,7 @@ def _make_docs(tmp_path: Path, name: str) -> str:
|
|||
d = tmp_path / name
|
||||
d.mkdir()
|
||||
(d / "cost.txt").write_text(
|
||||
"Asphalt Ab11 unit rate renegotiation reduced the paving cost on the school stretch.",
|
||||
"Licence unit price renegotiation reduced the office-suite cost for the head office.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return str(d)
|
||||
|
|
@ -243,9 +243,9 @@ def _make_docs(tmp_path: Path, name: str) -> str:
|
|||
|
||||
def _road_k(tmp_path: Path, *, verdict_input: dict[str, str] | None) -> Project:
|
||||
return Project(
|
||||
id="ROAD-K",
|
||||
name="Road k",
|
||||
description="road-backed project k",
|
||||
id="REF-K",
|
||||
name="Reference k",
|
||||
description="reference-backed project k",
|
||||
currency="NOK",
|
||||
cost_items=(),
|
||||
docs_dir=_make_docs(tmp_path, "k-docs"),
|
||||
|
|
|
|||
|
|
@ -2,13 +2,13 @@
|
|||
|
||||
**The measured hole.** P9 ran ``okf check`` (0.7.0, 15 rules) over this repository's payloads and
|
||||
(P11 measured the count again on 0.8.1: still 15 -- the 16th rule, ``bundle_mismatch``, is on okf
|
||||
main ``7cca9e0`` and in no tag yet; ``docs/2026-09-11-p11-okf-081.md`` section 5)
|
||||
got rc 0 on ``payload-n100.json`` and rc 1 on BOTH K2 payloads -- **20 findings, every one of them
|
||||
``excerpt_unnamed``**: "an excerpt a reader cannot name is one an answer cannot cite, whatever its
|
||||
rank (SS 8)". That diagnosis is right and it is not a solution: on po's side the same absence is
|
||||
SILENT. ``_excerpt_header`` appends the ``title:`` field only ``if excerpt.title is not None`` and
|
||||
says nothing otherwise, so a payload whose every excerpt is unnamed renders exactly like one whose
|
||||
producer simply had nothing to add, and no surface of a run carries the difference.
|
||||
main ``7cca9e0`` and in no tag yet; the ledger is ``docs/invarianter.md``)
|
||||
got rc 0 on the requirements-base payload and rc 1 on BOTH K2 payloads -- **20 findings, every one
|
||||
of them ``excerpt_unnamed``**: "an excerpt a reader cannot name is one an answer cannot cite,
|
||||
whatever its rank (SS 8)". That diagnosis is right and it is not a solution: on po's side the same
|
||||
absence is SILENT. ``_excerpt_header`` appends the ``title:`` field only ``if excerpt.title is not
|
||||
None`` and says nothing otherwise, so a payload whose every excerpt is unnamed renders exactly like
|
||||
one whose producer simply had nothing to add, and no surface of a run carries the difference.
|
||||
|
||||
**F1 (reporting), not F2 (refusing) -- and the price of F2 is a NUMBER.** Making ``title``
|
||||
required and refusing in ``check_payload_shape`` would reverse a written decision:
|
||||
|
|
|
|||
|
|
@ -1048,7 +1048,7 @@ def test_m8_an_attestation_that_is_not_utf8_is_red_because_it_is_unreadable(
|
|||
|
||||
|
||||
def test_m8_the_output_says_the_stress_row_needs_the_untracked_scratchpad() -> None:
|
||||
"""Finding 7: row 7's 1 of 20 stands on ``scratchpad/`` being present. Without it the row is
|
||||
"""Finding 7: rows 6-7 stand on ``scratchpad/`` being present. Without it the row is
|
||||
NOT MEASURED, and that dependency is a fact about the checkout, not about the product."""
|
||||
assert "scratchpad" in gate.STRESS_DEPENDENCY and "IKKE MÅLT" in gate.STRESS_DEPENDENCY
|
||||
rows = gate.evaluate(
|
||||
|
|
@ -1354,31 +1354,31 @@ def test_row7_not_measured_is_not_green_either() -> None:
|
|||
|
||||
|
||||
def test_row6_measures_the_stress_outboxes_when_they_exist(tmp_path: Path) -> None:
|
||||
"""Against the real artefacts when this machine has them; otherwise the absence is named."""
|
||||
"""Against the real artefacts when this machine has them; otherwise the absence is named.
|
||||
|
||||
The numbers this arm used to pin (15 validated, 15 undeclared, 5 own proposals, 1 named, 20
|
||||
commissioned) were stress round 6's, on context sets this repository no longer carries. No paid
|
||||
round has been run on the three example sets, so the ABSENT half is what runs today — and it
|
||||
runs everywhere, because the knowledge bases now ship with the package and only the outboxes
|
||||
can be missing. The PRESENT half pins the denominator a round on the example sets will be
|
||||
judged against, never a result nobody has measured.
|
||||
"""
|
||||
evidence = _CONFIG["stress_evidence"]
|
||||
root = _REPO / evidence["root"]
|
||||
bundles = frozen_bundles.store_root()
|
||||
if not root.is_dir() or not bundles.is_dir():
|
||||
m = gate.measure_stress(evidence, _REPO, tmp_path / "absent", None)
|
||||
assert m.missing and m.validated == 0
|
||||
pytest.skip(f"stress artefacts not mounted ({root}, {bundles})")
|
||||
assert frozen_bundles.store_root().is_dir(), "the example bases ship with the package"
|
||||
absent = gate.measure_stress(evidence, _REPO, tmp_path / "absent", None)
|
||||
assert absent.missing and absent.validated == 0 and not absent.drift
|
||||
assert "serverrom-2027" in absent.missing, absent.missing
|
||||
row = gate.score_undeclared(_PROBES, _all_pass(_PROBES), absent, "s")
|
||||
assert (row.k, row.status) == (None, gate.NOT_MEASURED)
|
||||
if not root.is_dir():
|
||||
pytest.skip(f"no stress round on the example sets on this machine ({root})")
|
||||
m = gate.measure_stress(evidence, _REPO, root, None)
|
||||
if m.missing:
|
||||
# The mount belongs to another repository and can be mid-rebuild; the gate then says
|
||||
# "ikke målt", which test_row6_missing_artefacts_are_never_zero already pins.
|
||||
assert m.validated == 0
|
||||
pytest.skip(f"stress artefacts not judgeable right now: {m.missing}")
|
||||
# 10 commissioned + 5 of the runs' own proposals (M-3).
|
||||
assert (m.validated, m.undeclared, m.own_validated, m.named, m.commissioned) == (
|
||||
15,
|
||||
15,
|
||||
5,
|
||||
1,
|
||||
20,
|
||||
)
|
||||
assert m.unaddressed > 0 # stress round 6 predates approach-addressed declarations
|
||||
row = gate.score_undeclared(_PROBES, _all_pass(_PROBES), m, "s")
|
||||
assert (row.k, row.status) == (None, gate.NOT_MEASURED)
|
||||
# Three sets of four approaches; the two-base set is judged per base, two rows each.
|
||||
assert (m.commissioned, m.rows) == (12, 12)
|
||||
|
||||
|
||||
def test_row7_is_a_diagnosis_and_never_moves_the_exit_code() -> None:
|
||||
|
|
@ -1467,7 +1467,7 @@ def test_the_command_is_red_today_with_every_row_in_its_output(tmp_path: Path) -
|
|||
assert (rows["types"]["k"], rows["types"]["n"]) == (3, 8)
|
||||
assert rows["kept"]["status"] == gate.RED
|
||||
assert (rows["maf"]["k"], rows["maf"]["n"], rows["maf"]["status"]) == (3, 8, gate.RED)
|
||||
# Probes green since row 6; stress round 6 predates the rule (or is absent) -> never green.
|
||||
# Probes green since row 6; no stress round on the example sets exists -> never green.
|
||||
assert rows["undeclared"]["status"] == gate.NOT_MEASURED
|
||||
assert rows["undeclared"]["k"] is None
|
||||
assert rows["named"]["failing"] is False
|
||||
|
|
|
|||
|
|
@ -23,7 +23,7 @@ _ASSUMPTIONS = {"05.2": (200.0, 230.0), "03.1": (290.0, 330.0)}
|
|||
|
||||
@pytest.fixture(scope="module")
|
||||
def project():
|
||||
return load_reference_projects()[0] # FV42-GSV-E1
|
||||
return load_reference_projects()[0] # KONTOR-IT-E1
|
||||
|
||||
|
||||
def _valid(project) -> SavingsProposal:
|
||||
|
|
|
|||
|
|
@ -41,10 +41,10 @@ def _entry(
|
|||
|
||||
def _mixed_ledger() -> SavingsLedger:
|
||||
""">=2 projects, >=2 dimensions, one cross-dimension overlap (the same candidate ``c-a`` realized
|
||||
under both ``energi`` and ``asfalt`` in P1 -> counted ONCE by the deduped total, FLAGGED)."""
|
||||
under both ``energi`` and ``lisens`` in P1 -> counted ONCE by the deduped total, FLAGGED)."""
|
||||
led = SavingsLedger()
|
||||
led.add_realized(_entry("P1", "c-a", 1000, dimension="energi"))
|
||||
led.add_realized(_entry("P1", "c-a", 1000, dimension="asfalt")) # overlap: same candidate id
|
||||
led.add_realized(_entry("P1", "c-a", 1000, dimension="lisens")) # overlap: same candidate id
|
||||
led.add_realized(_entry("P1", "c-b", 2500, dimension="energi"))
|
||||
led.add_realized(_entry("P2", "c-c", 4000, dimension="energi"))
|
||||
return led
|
||||
|
|
@ -109,7 +109,7 @@ def _sc6_entries() -> list[LedgerEntry]:
|
|||
direct ``SavingsLedger(entries=[...])`` (``add_realized`` would dedup the second). This exercises
|
||||
the TOTAL provenance sort from Step 1: a triple-only sort key would TIE the pair and let
|
||||
insertion order flip the bytes. Also carries a genuine cross-dimension overlap (``c-ov`` under
|
||||
both ``energi`` and ``asfalt``) so ``overlaps`` is non-empty."""
|
||||
both ``energi`` and ``lisens``) so ``overlaps`` is non-empty."""
|
||||
return [
|
||||
LedgerEntry(
|
||||
project_id="P1",
|
||||
|
|
@ -137,7 +137,7 @@ def _sc6_entries() -> list[LedgerEntry]:
|
|||
),
|
||||
LedgerEntry(
|
||||
project_id="P1",
|
||||
dimension="asfalt",
|
||||
dimension="lisens",
|
||||
candidate_identity="c-ov",
|
||||
amount_ore=3000,
|
||||
verdict_id="v-d",
|
||||
|
|
|
|||
|
|
@ -31,11 +31,11 @@ _MARKER_ID = "c-marker-overlap-s54" # marker: appears nowhere else in the codeb
|
|||
|
||||
|
||||
def _overlap_ledger() -> SavingsLedger:
|
||||
"""One candidate (``_MARKER_ID``) realized under TWO dimensions (energi + asfalt) in P1 — a
|
||||
"""One candidate (``_MARKER_ID``) realized under TWO dimensions (energi + lisens) in P1 — a
|
||||
distinct FULL storage key each, so both are stored (the overlap is visible), while the
|
||||
dimension-free dedup counts the ``_MARKER_ORE`` amount ONCE."""
|
||||
led = SavingsLedger()
|
||||
for dimension in ("energi", "asfalt"):
|
||||
for dimension in ("energi", "lisens"):
|
||||
led.add_realized(
|
||||
LedgerEntry(
|
||||
project_id="P1",
|
||||
|
|
|
|||
|
|
@ -32,7 +32,7 @@ _QUERY = ProposalFeatures(
|
|||
affected_codes=frozenset({"05.2", "03.1"}),
|
||||
measure_type="scope_reduction",
|
||||
claimed_saving_nok=220_000, # bucket [100k, 500k)
|
||||
description="asphalt base course reduction near school",
|
||||
description="licence scope reduction near head office",
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -54,7 +54,7 @@ def _store_with_true_match_and_decoys() -> tuple[VerdictStore, str]:
|
|||
affected_codes=frozenset({"09.1"}),
|
||||
measure_type="rate_renegotiation",
|
||||
claimed_saving_nok=50_000,
|
||||
description="asphalt base course reduction near school", # same words as query
|
||||
description="licence scope reduction near head office", # same words as query
|
||||
),
|
||||
decision="rejected",
|
||||
rationale="surface-text decoy",
|
||||
|
|
@ -65,7 +65,7 @@ def _store_with_true_match_and_decoys() -> tuple[VerdictStore, str]:
|
|||
affected_codes=frozenset({"21.2"}),
|
||||
measure_type="material_substitution",
|
||||
claimed_saving_nok=700_000,
|
||||
description="asphalt base course reduction extra words",
|
||||
description="licence scope reduction extra words",
|
||||
),
|
||||
decision="rejected",
|
||||
rationale="surface-text decoy",
|
||||
|
|
|
|||
|
|
@ -16,7 +16,7 @@ from portfolio_optimiser.run import run_project
|
|||
from portfolio_optimiser.validator import Rejection, ValidatedProposal
|
||||
|
||||
_VALID = (
|
||||
'{"project_id":"FV42-GSV-E1","measure":"Reduce scope",'
|
||||
'{"project_id":"KONTOR-IT-E1","measure":"Reduce scope",'
|
||||
'"affected_items":[{"code":"05.2","quantity":4300,"unit_cost":215},'
|
||||
'{"code":"03.1","quantity":1800,"unit_cost":310}],"claimed_saving_nok":200000}'
|
||||
)
|
||||
|
|
@ -27,7 +27,7 @@ _VI = {"decision": "approved", "rationale": "feasible within range"}
|
|||
|
||||
async def test_a_valid_proposal_end_to_end(docs_dir, make_client_factory, fresh_store) -> None:
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -48,7 +48,7 @@ async def test_a1_provenance_stamps_injected_client_model_not_sentinel(
|
|||
(SyntheticUsageChatClient.model == 'synthetic'), never the 'fake-model' literal — the leak
|
||||
falsified provenance on the public deployer seam (run.py:197)."""
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -61,7 +61,7 @@ async def test_a1_provenance_stamps_injected_client_model_not_sentinel(
|
|||
|
||||
async def test_b_out_of_range_is_rejected(docs_dir, make_client_factory, fresh_store) -> None:
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input={"decision": "rejected", "rationale": "claim too high"},
|
||||
|
|
@ -74,7 +74,7 @@ async def test_b_out_of_range_is_rejected(docs_dir, make_client_factory, fresh_s
|
|||
|
||||
async def test_c_layer2_verdict_is_persisted(docs_dir, make_client_factory, fresh_store) -> None:
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -89,7 +89,7 @@ async def test_d_second_run_retrieves_prior_verdict(
|
|||
) -> None:
|
||||
factory = make_client_factory(_VALID)
|
||||
r1 = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -97,7 +97,7 @@ async def test_d_second_run_retrieves_prior_verdict(
|
|||
store=fresh_store,
|
||||
)
|
||||
r2 = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -113,7 +113,7 @@ async def test_e_tiny_budget_halts_without_exceeding(
|
|||
) -> None:
|
||||
with pytest.raises(BudgetExceeded):
|
||||
await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -143,7 +143,7 @@ async def test_wiring_budget_middleware_and_retrieval_tool(
|
|||
|
||||
monkeypatch.setattr(run_mod, "fresh_workflow", spy_fresh)
|
||||
await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -174,7 +174,7 @@ async def test_g_validated_proposal_derives_from_debate(
|
|||
|
||||
monkeypatch.setattr(run_mod, "generate_via_llm", spy_generate)
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -205,7 +205,7 @@ async def test_h_tiny_budget_halts_via_debate_middleware(
|
|||
monkeypatch.setattr(run_mod, "generate_via_llm", _forbidden)
|
||||
with pytest.raises(BudgetExceeded):
|
||||
await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -227,7 +227,7 @@ async def test_i_injected_meter_is_used(docs_dir, make_client_factory) -> None:
|
|||
|
||||
factory = make_client_factory(_VALID)
|
||||
baseline = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -237,7 +237,7 @@ async def test_i_injected_meter_is_used(docs_dir, make_client_factory) -> None:
|
|||
injected = TokenMeter(Budget(max_tokens=100_000, max_rounds=max(3 * 4, 4)))
|
||||
injected.charge(pre_charged)
|
||||
result = await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
@ -260,7 +260,7 @@ async def test_f_malformed_contract_raises_before_any_chat(docs_dir, make_client
|
|||
|
||||
with pytest.raises(ValidationError):
|
||||
await run_project(
|
||||
"FV42-GSV-E1",
|
||||
"KONTOR-IT-E1",
|
||||
"local",
|
||||
docs_dir=docs_dir,
|
||||
verdict_input=_VI,
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Add a link
Reference in a new issue