feat(p18): an identifier that stands everywhere identifies nothing

P18 parts B and C (order 20260914T105139Z).

B1 -- stage 0b. P7 made it `item.code in grounding`: plain containment over
ONE concatenated string. P16 ran it against a delivered corpus and measured
what containment cannot tell apart: the falsification arm a4-indeksregulering
put 250 000 NOK on a single cost line coded R761 -- the knowledge base's OWN
NAME, carried by all 2 756 of its concept documents -- and the whole gate
said validated (stage 0 skipped, un-anchored run; checker approve).

The grounding is now carried as the DOCUMENTS it is made of (validator.
Grounding), not as a blob. A structure and not a second argument beside the
text: the boundaries and the text are one fact, and .text is derived, so the
gate and P8's report measure the same characters. run.py composes one
document per concept file where the base is already walked; generate.
_grounding_text folds each cost line in as a one-line document.

N and A are MEASURED, not chosen (14.09, four mounted vegnormal bases):
- every must_cite ref and mandate affected_code in the four context sets --
  shortest real identifier is FOUR characters (12.1, 52.1), so N = 3 sits one
  below the measurement and cannot refuse anything measured;
- document frequency of every code-shaped token per base -- 1 692 distinct
  and NOT ONE reaches 5 %. Highest anywhere 6/446 (1.35 %), highest a fasit
  names 3/446 (0.67 %), R761 2 756/2 756 (100 %). A = 0.05 therefore sits
  3.7x above the highest real token and 20x below the defect.
Length is NOT what makes the defect inert (R761 is four characters); the
share is. And a share is not a measurement without a denominator big enough
to take one (ansikt 4): one of three is 33 %, so an ABSOLUTE floor of 10
documents gates it. Highest absolute count any real identifier reaches is 6,
and every fixture in the repo is far below 10 -- which is why every pre-P18
gate is UNTOUCHED by this rule rather than exempted from it. Grounding.of
(one document) can never reach the floor by construction.

The refusal NAMES the denominator ("appears in 2756 of the 2756 documents
this run was given"), because Step 5 feeds that reason verbatim into the next
attempt's prompt: a proposer told only "ungrounded" answers with another
token of the same kind.

B2 SPIKE (measured, NOT built) FELLED the order's own alternative: option (b)
"ground in what the run OPENED" was run over P16's 16 code rows -- R761
stands in every OPENED document too, so (b) would NOT have caught the defect,
while B1 makes it inert and still grounds the real process line 65
ASFALTDEKKER (29/2756 = 1.05 %). (b) is not a substitute for B1.

C1 -- --docs-dir is optional once --bundle-dir is given (P16 FUNN 2). On the
bundle path docs_dir is never read: retrieval, the chunk tool and the "no
citable content" check all live in the road branch. Bound ONCE from
--bundle-dir, which is byte-identically what the README already tells an
operator to type by hand. NOT the "--docs-dir omvei": no such path is opened
and the road branch still refuses without a real --docs-dir (own arm).

C2 -- the judge's snippet arm counts only under citation_scope == "narrowed",
as (a) already does (PM decision, P16 s 6.2). P16's reason for (b') being
clean -- snippets are bodies while ref/title live in frontmatter, 0 of 446
n100 bodies -- holds for "Krav 4.1.2-1" but NOT for R761, where a process
number like 12.1 stands in the bodies. Under a whole-base citation list that
mark was "cited" before any model call.

tests: test_inert_identifier_loadbearing.py (7 arms; known positive is P16's
OWN artefact replayed against the base that run was given, known negative is
26 of 26 fasit references still grounding), test_docs_dir_optional_
loadbearing.py (5 arms). test_stress_judge_loadbearing.py's snippet arm split
into narrowed/whole-base -- the pair is the discriminator, same snippet, same
mark, only the scope differs. The grounding tests migrate from str to
Grounding.of (the honest reading of a caller that declared no boundaries).

Verification: uv run pytest -q 1698 passed / 5 skipped (1685 after part A,
strict superset, 0 removed). ruff check + format clean, mypy clean (38
files). Golden demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-14 22:28:40 +02:00
commit 7a7c988253
11 changed files with 628 additions and 51 deletions

View file

@ -2521,6 +2521,57 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
kalles fortsatt per verktøykall (I/O, ikke tokens); A3-nekten navigerer basen ÉN gang ekstra på
feilstien for å finne den nærmeste listbare katalogen; at en LEVENDE modell BRUKER `filter` er
ikke bevist (structured-output-grensens klasse) — DEL D er målingen.
- **En identifikator som står OVERALT identifiserer INGENTING — og grunnlaget bæres som DOKUMENTER,
ikke som én blob (P18/B1, 14.09):** P7 gjorde stadium 0b til `item.code in grounding`, ren
delstreng over én sammenslått streng. P16 kjørte den mot et LEVERT korpus og målte hva
containment ikke kan skille: falsifiseringsarmen `a4-indeksregulering` la 250 000 NOK på ÉN
kostlinje kodet **`R761`** — basens EGET NAVN, som alle 2 756 konseptdokumenter bærer — og hele
gaten sa `validated` (stage 0 hoppet over, uforankret kjøring; checker `approve`).
**`Grounding` er en STRUKTUR, ikke et andre argument ved siden av teksten:** grensene og teksten
er ÉTT faktum, og to bærere for ett faktum står fritt til å være uenige (kø-(p)); `.text` utledes,
så gaten og P8-rapporten måler nøyaktig de samme tegnene. **N og A er MÅLT, ikke valgt:** over
hver `must_cite`-referanse og hver mandat-`affected_code` i de fire kontekstsettene er den
KORTESTE ekte identifikatoren FIRE tegn (`12.1`/`52.1`), så `_GROUNDING_MIN_LENGTH = 3` ligger ETT
under målingen og kan ikke nekte noe som er målt; og over dokumentfrekvensen til hvert
kodeformet token i hver base (`_IDENTIFIER_FORMS`) når **ingen av 1 692 distinkte** 5 % —
høyeste noe sted er 6 av 446 (**1,35 %**), høyeste en fasit faktisk navngir 3 av 446 (**0,67 %**),
mens `R761` er **2 756 av 2 756 (100 %)**. `_GROUNDING_MAX_DOCUMENT_SHARE = 0.05` ligger altså
3,7× over det høyeste ekte tokenet og 20× under defekten. **Lengden er IKKE dét som gjør defekten
inert** (`R761` er fire tegn) — andelen er; N dekker sammentreffs-klassen målingen ikke tilfeldigvis
inneholdt. **Og en andel er ingen måling uten en nevner stor nok til å ta den** (ansikt 4): ett av
tre dokumenter er 33 % og sier ingenting, så `_GROUNDING_MIN_INERT_DOCUMENTS = 10` er et ABSOLUTT
gulv — høyeste absolutte dokumenttall noen ekte identifikator når i de fire korpusene er 6, og
hver fixtur i repoet ligger langt under 10, hvilket er dét som gjør at HVER pre-P18-gate er
URØRT av regelen i stedet for UNNTATT fra den. `Grounding.of` (ett dokument) er den ærlige
lesningen av en kaller uten grenser å erklære og kan ved konstruksjon aldri utløse andelen.
**Nevneren NAVNGIS i nekten** («appears in 2756 of the 2756 documents this run was given»), fordi
Steg 5 mater den grunnen ORDRETT inn i neste forsøks prompt: en proposer som bare får «ungrounded»
svarer med et nytt token av samme slag. **B2-SPIKEN FELTE ORDRENS EGEN ALTERNATIV-HYPOTESE:** (b)
«grunnlag = det kjøringen ÅPNET» ble MÅLT over P16s 16 kode-rader — `R761` står i hvert ÅPNET
dokument også, så (b) ville **ikke** fanget defekten (grunner fortsatt), mens B1 gjør den inert og
lar den ekte prosesslinja `65 ASFALTDEKKER` (29/2756 = 1,05 %) grunne. (b) er altså ikke et
substitutt for B1; målt, ikke bygget. Load-bearing MÅLT
(`tests/test_inert_identifier_loadbearing.py`, 7 armer): kjent-positiv er P16s EGET artefakt
spilt av mot basen den fikk, kjent-negativ er **26 av 26** fasit-referanser som fortsatt grunner.
**Ærlighets-grenser, uttalt:** regelen er per-DOKUMENT, så en base med ett gigantisk dokument er
upåvirket; `assumptions`-nøkler er fortsatt utenfor (økt 109s grense); og ingen betalt kjøring
bekrefter at den endrer utfallet levende før DEL D.
- **`--docs-dir` er VALGFRI når `--bundle-dir` er gitt, og det er ikke omveien (P18/C1):** på
bundle-stien LESES `docs_dir` aldri — retrieval, chunk-verktøyet og «no citable content»-sjekken
bor alle i VEG-grenen — likevel KREVDE gaten den, så den publiserte kommandoen måtte navngi samme
katalog to ganger (P16 FUNN 2). Verdien bindes nå ÉN gang fra `--bundle-dir` når `--docs-dir`
mangler, hvilket er byte-identisk med dét READMEen alt ber operatøren skrive for hånd; hver
eksisterende invokasjon, to-flagg-formen inkludert, er uendret. **Dette er IKKE
«`--docs-dir`-omveien»** (å mate prosjektdokumenter gjennom retrieval I STEDET FOR å ingeste dem
inn i en kunnskapsbase): ingen slik sti åpnes, og veg-grenen nekter fortsatt uten en EKTE
`--docs-dir` (egen arm). Nekten navngir nå BEGGE dører, ikke bare den ene den pleide å navngi.
- **Dommerens snippet-arm teller KUN under `citation_scope == "narrowed"` (P18/C2, PM-valg P16
§ 6.2):** (a) hadde alt den korreksjonen; (b) hadde den ikke. P16s egen begrunnelse for at (b)
var ren — «snippetene er konsept-BODYER mens `ref`/`title` bor i FRONTMATTER, målt 0 av 446
n100-bodyer» — holder for `Krav 4.1.2—1`, men **IKKE for R761**, der et prosessnummer som `12.1`
står i bodyene selv. Under en helbase-siteringsliste var det merket «sitert» før noe modellkall,
så raden kom tilbake `named` for en kjøring der modellen ikke hadde sagt noe slikt. Diskriminatoren
er PARET av to armer med SAMME snippet og SAMME merke, der eneste forskjell er scopet.
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.

View file

@ -647,14 +647,35 @@ when the seam is detached, so the loop cannot silently degrade into theater.
One caveat is worth stating, because it decides what the verdict means: when the debate navigates
the base, the provenance stamp cites *every* concept file, so "the fasit path is cited" is true
before any model call. A citation therefore only counts as grounding under a declared pre-pass
cut; otherwise grounding must come from a document the run actually opened. Both halves are
reported either way.
cut; otherwise grounding must come from a document the run actually opened. The same correction
now applies to *naming* (P18): a requirement found in a whole-base citation snippet was "named"
before the model said anything, so the snippet half counts only under a narrowed scope too. Both
halves are reported either way.
```bash
uv run python -m portfolio_optimiser.stress contexts/<set> \
--outbox-dir <outbox> --run-id <run-id>
```
- **An identifier that stands everywhere identifies nothing (P18).** The validator's stage 0b
grounds each proposed cost code in the run's non-model-authored input. Plain containment was not
enough: measured on a live run, a proposal put 250 000 NOK on a line coded `R761` — the knowledge
base's own *name*, carried by all 2 756 of its documents — and the whole gate said `validated`.
A code now grounds only if it is at least 3 characters long **and** appears in fewer than 5 % of
the documents the input is made of, with an absolute floor of 10 documents so the share is never
taken over a handful. Both numbers are measured, not chosen: across the four delivered corpora no
code-shaped token of 1 692 distinct ones reaches 5 % (highest 1.35 %), and the shortest real
identifier is 4 characters. The refusal names the denominator — "appears in 2 756 of the 2 756
documents this run was given" — because that reason is fed verbatim into the next attempt's
prompt. A caller that declares no document boundaries is unchanged by construction: one document
can never reach the floor.
- **`--docs-dir` is optional when `--bundle-dir` is given (P18).** On the bundle path `docs_dir` is
never read — retrieval, the chunk tool and the citation check all live in the road branch — so
naming the same directory twice was a requirement for a path that ignores it. Both forms work;
neither turns retrieval into a substitute for ingesting documents into a knowledge base, and the
road path still requires a real `--docs-dir`.
- **Requiring the run to be anchored**`--require-cost-baseline` (opt-in, requires
`--bundle-dir`). Without a baseline the validator's stage 0 is skipped, and the run says so on
stdout — but it still finishes and still stamps `validator_decision: validated` over cost lines

View file

@ -40,6 +40,7 @@ from portfolio_optimiser.proposal_review import (
)
from portfolio_optimiser.reference_domain import Project
from portfolio_optimiser.validator import (
Grounding,
Rejection,
ValidatedProposal,
self_repair,
@ -421,12 +422,20 @@ def generate_with_validation(
return self_repair(_attempt, max_attempts=max_attempts)
def _grounding_text(project: Project, baseline: CostBaseline | None, delivered: str) -> str:
def _grounding_text(
project: Project, baseline: CostBaseline | None, delivered: Grounding
) -> Grounding:
"""P7: compose the ONE text a candidate's identifiers must be grounded in — the run's
non-model-authored input, and nothing else.
Three sources, each of which the run can point at without asking the model:
**P18/B1: the result carries the DOCUMENT BOUNDARIES, not only the text.** ``code in text``
cannot tell "this project has such a line" from "this word is in every letterhead" P16
measured a base's own name, ``R761``, carrying a fabricated 250 000 NOK line to ``validated``.
The two later sources join as ONE-LINE documents rather than being appended to a blob: each IS
one cost line, and the share rule then reads them exactly as it reads a concept file.
* ``delivered`` what the CALLER can prove this run was GIVEN. ``run_project`` fills it from
the delivered rendered context (the pre-pass cut, the bundle pointer, or the road path's
retrieved chunks) PLUS the navigated base's ``context_files`` — never ``files``, because that
@ -448,12 +457,12 @@ def _grounding_text(project: Project, baseline: CostBaseline | None, delivered:
very identifier it just refused. Grounding in the prompt would therefore let the gate's own
refusal ground the next attempt: a falsifier that disarms itself on its second round.
"""
return "\n".join(
[
delivered,
return Grounding(
documents=(
*delivered.documents,
*(item.code for item in project.cost_items),
*(() if baseline is None else baseline.items),
]
)
)
@ -508,7 +517,7 @@ class GroundingOffer:
def grounding_offer(
project: Project, baseline: CostBaseline | None, delivered: str
project: Project, baseline: CostBaseline | None, delivered: Grounding
) -> GroundingOffer:
"""Measure what ``delivered`` can ground, on the EXACT text the gate will see.
@ -519,7 +528,7 @@ def grounding_offer(
Deterministic and free: no model call, no network, and no second walk of the bundle.
"""
text = _grounding_text(project, baseline, delivered)
text = _grounding_text(project, baseline, delivered).text
found: set[str] = set()
for form in _IDENTIFIER_FORMS:
found |= set(form.findall(text))
@ -544,7 +553,7 @@ async def generate_via_llm(
reviews: list[ProposalReview] | None = None,
review_key: tuple[str | None, str | None] = (None, None),
checker_verdict: str = "absent",
grounding: str | None = None,
grounding: Grounding | None = None,
) -> GenerationResult:
"""Async LLM path: non-streaming chat -> parse -> validate, with TWO bounded retry kinds,
the meter checked in this loop:
@ -703,7 +712,7 @@ async def generate_via_llm(
# caller that declared nothing else IS the input it declared. ``run_project``
# always passes it EXPLICITLY, because on the debate path ``context`` has been
# replaced by the model's OWN summary of what it read.
context if grounding is None else grounding,
Grounding.of(context) if grounding is None else grounding,
),
)
last_ruling = result

View file

@ -120,6 +120,7 @@ from portfolio_optimiser.provenance import ProvenanceStamp
from portfolio_optimiser.reference_domain import Project, load_reference_projects
from portfolio_optimiser.tracing import TracingConfigError, configure_tracing, tracing_notice
from portfolio_optimiser.validator import (
Grounding,
Rejection,
ValidatedProposal,
baseline_from_project,
@ -1107,10 +1108,13 @@ async def run_project(
# P7: the delivered base is the run's own evidence for what identifiers EXIST. Built from
# ``context_files`` (MAJOR-3/S7a-3's rule), so the ``type: verdict`` layer stays out — a
# proposal grounded in a prior verdict would reach the ExpeL fold's material around its gate.
bundle_grounding = ""
# P18/B1: ONE document per concept file, not one blob. The boundaries ARE the denominator the
# gate's share rule needs, and composing them here — where the base is already walked — is what
# keeps them from being a second, drifting reconstruction (kø-(p)).
bundle_grounding: tuple[str, ...] = ()
if bundle_dir is not None:
bundle = okf.navigate_bundle(bundle_dir)
bundle_grounding = "\n".join(
bundle_grounding = tuple(
"\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files
)
# ONE bundle-id rule (Step 10, slackened S7a-3 pkt. 1): the DECLARED id is the identity and
@ -1224,7 +1228,7 @@ async def run_project(
# ``_fetch_parsed`` has returned, so a report from there could only ever speak once an attempt
# had been paid for. ONE binding feeding both the report and the gate: two compositions of one
# text are free to disagree, which is exactly what a report must not be able to do (kø-(p)).
delivered = "\n".join([context, bundle_grounding])
delivered = Grounding(documents=(context, *bundle_grounding))
offer = grounding_offer(project, baseline, delivered)
# Trekk B2 (krav 3): configured MCP servers become tools the AGENTS can call during the debate.
@ -2950,14 +2954,26 @@ def main(argv: list[str] | None = None) -> int:
# HOISTED above the scripted door (below) so an incomplete argv is refused BEFORE the honesty
# banner could claim a scripted run happened; the refusal ORDER within single-project mode
# (required args -> semantic-retrieval -> scripted) is unchanged.
if not args.portfolio and (args.project_id is None or args.docs_dir is None):
if not args.portfolio and (
args.project_id is None or (args.docs_dir is None and args.bundle_dir is None)
):
print(
"run refused: single-project mode requires PROJECT_ID and --docs-dir "
"(use --portfolio for portfolio mode)",
"run refused: single-project mode requires PROJECT_ID and either --docs-dir or "
"--bundle-dir (use --portfolio for portfolio mode)",
file=sys.stderr,
)
return 1
# P18/C1 (P16 FUNN 2): on the BUNDLE path ``docs_dir`` is never read — retrieval, the chunk
# tool and the citation check all live in the road branch — yet the guard above demanded it,
# so the documented command had to name the same directory twice. It is now bound ONCE, from
# ``--bundle-dir`` when ``--docs-dir`` is absent, which is byte-identically what the README
# tells an operator to type by hand; every existing invocation, including the two-flag form,
# is unchanged. This is NOT the "--docs-dir omvei" (feeding project documents through
# retrieval INSTEAD of ingesting them into a knowledge base): no such path is opened, and the
# road branch still refuses without a real ``--docs-dir``.
docs_dir = args.docs_dir if args.docs_dir is not None else args.bundle_dir
# The third projection reads a table INSIDE a bundle, so without one there is nothing to derive
# from: the road path's baseline comes from ``Project.cost_items`` and is anchored by
# construction. Refused by NAME here rather than left to surface later as a project-lookup
@ -3885,7 +3901,7 @@ def main(argv: list[str] | None = None) -> int:
run_project(
args.project_id,
args.profile,
docs_dir=args.docs_dir,
docs_dir=docs_dir,
bundle_dir=args.bundle_dir,
verdict_dir=args.verdict_dir,
dimension=(
@ -3963,7 +3979,7 @@ def main(argv: list[str] | None = None) -> int:
run_project(
args.project_id,
args.profile,
docs_dir=args.docs_dir,
docs_dir=docs_dir,
bundle_dir=args.bundle_dir,
verdict_dir=args.verdict_dir,
dimension=(

View file

@ -24,11 +24,17 @@ it. So a CITATION grounds an approach only when the citation list is NARROWER th
declared pre-pass cut, where the stamp really does name what was read). Both halves are reported
either way (``opened`` / ``cited`` / ``citation_scope``), so which one fired stays readable.
**(b') was checked for the same vacuity and is CLEAN, so the order stands.** ``bundle_citations``
snippets are concept BODIES while ``ref``/``title`` live in FRONTMATTER: measured 0 of 446
n100 bodies contain ``Krav 4.1.2-1``. The snippet arm can therefore carry (b') without being
satisfied by construction. ``named_in_measure`` / ``named_in_snippet`` are still reported apart,
because the measure is the model's own prose and a snippet is the base's.
**(b') was checked for the same vacuity and is CLEAN for the N corpora, so the order stands.**
``bundle_citations`` snippets are concept BODIES while ``ref``/``title`` live in FRONTMATTER:
measured 0 of 446 n100 bodies contain ``Krav 4.1.2-1``. ``named_in_measure`` / ``named_in_snippet``
are still reported apart, because the measure is the model's own prose and a snippet is the base's.
**P18/C2 (PM decision, P16 § 6.2): the snippet arm counts only under a NARROWED scope, as (a)
does.** The paragraph above holds for a reference like ``Krav 4.1.2-1``, which no body repeats it
does NOT hold for R761, where a process number such as ``12.1`` stands in the bodies themselves.
Under a whole-base citation list that mark is "cited" before any model call, so the row was
``named`` for a run in which the model had said nothing of the kind. The scope gate is the same
correction (a) already carries, applied to the half that was still exposed.
**A DENOMINATOR, ALWAYS** (Verifiseringsloven ansikt 4). Every verdict names how many tool calls,
citations, approach rows and base concepts it saw, and an outbox with no proposal artefact - or a
@ -274,7 +280,13 @@ def score_context_set(
snippets = " ".join(str(c.get("snippet", "")) for c in citations)
marks = [m for c in concepts for m in (c.get("ref", ""), c.get("title", "")) if m]
named_in_measure = any(m in measure for m in marks)
named_in_snippet = any(m in snippets for m in marks)
# P18/C2 (PM decision, P16 § 6.2): the snippet arm counts ONLY under a narrowed citation
# scope, exactly as (a) does. A whole-base citation list is stamped by ``bundle_citations``
# before a single model call — measured on n100, 446 context files, 446 citations, 6 of 6
# fasit paths "cited" for free — so a mark found in THOSE snippets is evidence about the
# base's contents, not about this run. Measured on r761: ``12.1`` appears in whole-base
# snippets and gave this row ``named`` without the model having said anything.
named_in_snippet = scope == "narrowed" and any(m in snippets for m in marks)
halluc = [f"citation:{f}" for f in sorted(cited_files - concept_names)]
allowed = set(approach.affected_codes) | baseline_codes

View file

@ -27,6 +27,7 @@ import warnings
from collections.abc import Callable, Mapping
from contextlib import contextmanager
from dataclasses import dataclass
from typing import Final
import pulp
@ -208,7 +209,95 @@ def _reconcile_against_baseline(
return Rejection(proposal=proposal, reason="; ".join(violations))
def _ground_against_input(proposal: SavingsProposal, grounding: str) -> Rejection | None:
#: P18/B1 — the shortest identifier the gate will let ground anything.
#:
#: MEASURED (14.09) over every ``must_cite`` reference and every mandate ``affected_code`` in the
#: four context sets: the shortest real identifier is FOUR characters (R761's ``12.1`` / ``52.1``).
#: Set one BELOW that, so the rule cannot refuse anything that has been measured, while a one- or
#: two-character token — which matches by coincidence in any prose — grounds nothing. The honest
#: failure direction for a gate that speaks about a model's invention.
#:
#: Length is NOT what makes the measured defect inert: P16's fabricated ``R761`` is four characters
#: long. The share below is what does. This covers the coincidence class the measurement did not
#: happen to contain.
_GROUNDING_MIN_LENGTH: Final = 3
#: The share of the grounding's DOCUMENTS above which a token identifies nothing.
#:
#: MEASURED over the four delivered corpora, counting document frequency for every code-shaped
#: token (``generate._IDENTIFIER_FORMS``): 1 692 distinct tokens, and NOT ONE reaches 5 % of its
#: base's documents. The highest anywhere is 6 of 446 (1.35 %); the highest that a fasit or mandate
#: actually names is 3 of 446 (0.67 %). P16's fabricated ``R761`` is 2 756 of 2 756 — 100 %.
#: 5 % therefore sits 3.7x above the highest real token measured and 20x below the defect.
_GROUNDING_MAX_DOCUMENT_SHARE: Final = 0.05
#: …and a share is not a measurement without a denominator big enough to take one (ansikt 4).
#: One document of three is 33 % and says nothing at all, so the share only fires once a token is
#: in at least this many documents. MEASURED: the highest ABSOLUTE document count any real
#: identifier reaches in the four corpora is 6, and every test fixture in this repo is far below
#: 10 — which is why every pre-P18 gate is untouched by this rule rather than exempted from it.
_GROUNDING_MIN_INERT_DOCUMENTS: Final = 10
@dataclass(frozen=True)
class Grounding:
"""The run's non-model-authored input, carried as the DOCUMENTS it is made of.
P7 carried it as ONE string, and P16 measured what that costs: ``R761`` the base's own NAME,
which every one of its 2 756 concept documents carries satisfied ``code in grounding`` and
carried a fabricated 250 000 NOK line through the whole gate to ``validated``. Containment in a
concatenation cannot tell "this project has such a line" from "this word is in the letterhead".
A structure rather than a second argument beside the text: the boundaries and the text are ONE
fact, and two carriers for one fact are free to disagree about it (-(p)). ``text`` is derived
here, so the gate and P8's report measure the very same characters.
A caller with no boundaries to declare the road path, a test builds ONE document from its
text and is unchanged by construction: one document can never reach
``_GROUNDING_MIN_INERT_DOCUMENTS``, so the share cannot fire on it.
"""
documents: tuple[str, ...]
@classmethod
def of(cls, text: str) -> Grounding:
"""One document. The honest reading of a caller that declared no boundaries."""
return cls(documents=(text,))
@property
def text(self) -> str:
"""The ONE composition. Byte-identical to P7's ``"\n".join`` of the same parts."""
return "\n".join(self.documents)
def document_frequency(self, token: str) -> int:
"""How many of the documents contain ``token`` — the numerator, in the unit of the rule."""
return sum(1 for document in self.documents if token in document)
def _inert_in(grounding: Grounding, code: str) -> str | None:
"""Why ``code`` identifies nothing in this input, or ``None`` when it identifies something.
The message NAMES THE DENOMINATOR, because "it is everywhere" and "it is not here" are
different findings and Step 5 feeds this reason verbatim into the next attempt's prompt: a
proposer told only "ungrounded" will re-answer with another token of the same kind.
"""
if len(code) < _GROUNDING_MIN_LENGTH:
return (
f"is {len(code)} characters long, too short to identify a cost line — any prose "
"contains it by coincidence"
)
hits = grounding.document_frequency(code)
total = len(grounding.documents)
floor = max(_GROUNDING_MIN_INERT_DOCUMENTS, total * _GROUNDING_MAX_DOCUMENT_SHARE)
if hits >= floor:
return (
f"appears in {hits} of the {total} documents this run was given — a token that is in "
"every document identifies none of them; name a cost line, not the corpus"
)
return None
def _ground_against_input(proposal: SavingsProposal, grounding: Grounding) -> Rejection | None:
"""P7: every identifier the proposal builds on must appear VERBATIM in the input it was built
from, or the verdict falls.
@ -244,12 +333,19 @@ def _ground_against_input(proposal: SavingsProposal, grounding: str) -> Rejectio
deliberately out of scope: the Monte Carlo never samples such a band (``SavingsProposal.
_assumption_bands_enclose_unit_cost`` says so in the same words), so it cannot move the verdict,
and a check on it would be a branch no recording exercises."""
violations = [
f"ungrounded identifier {item.code!r}: it appears nowhere in the input this proposal "
f"was built from ({len(grounding)} chars)"
for item in proposal.affected_items
if item.code not in grounding
]
violations = []
for item in proposal.affected_items:
if item.code not in grounding.text:
violations.append(
f"ungrounded identifier {item.code!r}: it appears nowhere in the input this "
f"proposal was built from ({len(grounding.text)} chars)"
)
continue
# P18/B1: present is not the same as identifying. An identifier that stands everywhere
# identifies nothing, and one too short to be an identifier is matched by coincidence.
inert = _inert_in(grounding, item.code)
if inert is not None:
violations.append(f"ungrounded identifier {item.code!r}: it {inert}")
if not violations:
return None
return Rejection(proposal=proposal, reason="; ".join(violations))
@ -259,7 +355,7 @@ def validate_proposal(
proposal: SavingsProposal,
*,
baseline: CostBaseline | None = None,
grounding: str | None = None,
grounding: Grounding | None = None,
tolerance: float = BASELINE_TOLERANCE_DEFAULT,
method_caps: Mapping[str, float] | None = None,
) -> ValidatedProposal | Rejection:

View file

@ -0,0 +1,108 @@
"""P18/C1 — ``--docs-dir`` is optional once ``--bundle-dir`` is given.
P16 FUNN 2: the documented stress command names the same directory twice
(``--docs-dir <base> --bundle-dir <base>``), because single-project mode demanded ``--docs-dir``
even on the bundle path where it is never read. Retrieval, the chunk tool and the "no citable
content" check all live in the ROAD branch (``run.py``); the bundle branch builds its citations
from the navigated base. So the flag was required for a path that ignores it, and the published
command had to satisfy the requirement by repeating itself.
**This is not the "--docs-dir omvei"** (feeding project documents through retrieval INSTEAD of
ingesting them into a knowledge base), which STATE forbids and this order forbids again. No such
path is opened: the road branch still refuses without a real ``--docs-dir``, and the value is
bound ONCE from ``--bundle-dir`` byte-identically what the README already tells an operator to
type by hand, so every existing invocation, the two-flag form included, is unchanged.
"""
from __future__ import annotations
import json
import shutil
from pathlib import Path
from typing import Any
import pytest
from portfolio_optimiser import run
_FIXTURES = Path(__file__).parent / "fixtures"
_PRICED = _FIXTURES / "k2-prisskjema-SYNTETISK"
_IR_PROJECTION = {
"project_id": "K2",
"measure": "energy_efficiency",
"claimed_saving_nok": 1000.0,
"affected_codes": ["01.1"],
}
def _runnable(tmp_path: Path) -> str:
root = tmp_path / "base"
shutil.copytree(_PRICED, root)
(root / "validator-input.json").write_text(json.dumps(_IR_PROJECTION), encoding="utf-8")
return str(root)
def test_a_bundle_run_needs_no_docs_dir(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
"""(a) The fix. A free dry run on ``--bundle-dir`` alone is ACCEPTED and reaches the bundle
path asserted on the run-config the dry run prints, not on rc alone, since rc 0 is also what
a run that silently took the road path would return."""
rc = run.main(["K2", "--bundle-dir", _runnable(tmp_path), "--live-dry-run"])
assert rc == 0
assert "LIVE-DRY-RUN OK" in capsys.readouterr().out
def test_the_documented_two_flag_form_still_works(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""(b) The CONTROL that this is a widening, not a change. The README's form — the same
directory twice must be byte-for-byte as accepted as it was before."""
base = _runnable(tmp_path)
rc = run.main(["K2", "--docs-dir", base, "--bundle-dir", base, "--live-dry-run"])
assert rc == 0
assert "LIVE-DRY-RUN OK" in capsys.readouterr().out
def test_neither_flag_is_still_refused_by_name(capsys: pytest.CaptureFixture[str]) -> None:
"""(c) The half that must NOT be relaxed: the road path has no base to fall back to, so an
argv naming neither is refused, and the refusal names BOTH doors rather than only the one it
used to name."""
rc = run.main(["K2", "--live-dry-run"])
assert rc == 1
err = capsys.readouterr().err
assert "run refused" in err and "--docs-dir" in err and "--bundle-dir" in err
def test_the_road_path_still_requires_a_real_docs_dir(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
"""(d) The anti-omvei arm. With no bundle, ``--docs-dir`` is still the only door AND it is
still read: a directory holding nothing citable is refused by the road branch's own check, so
nothing here turns retrieval into a substitute for ingestion."""
empty = tmp_path / "tomt"
empty.mkdir()
rc = run.main(["P1", "--docs-dir", str(empty), "--live-dry-run"])
assert rc == 1
assert capsys.readouterr().err.strip(), "the road path must say why, not fail silently"
def test_the_bundle_value_is_bound_once_and_reaches_run_project(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""(e) The SEAM. rc 0 above would also be satisfied by a CLI that never forwarded the value, so
the argument ``run_project`` actually receives is recorded one binding, both dispatch sites."""
base = _runnable(tmp_path)
seen: dict[str, Any] = {}
async def _record(project_id: str, profile: Any, **kw: Any) -> Any:
seen.update(kw)
raise SystemExit(0)
monkeypatch.setattr("portfolio_optimiser.run.run_project", _record)
with pytest.raises(SystemExit):
run.main(["K2", "--bundle-dir", base, "--live-dry-run"])
assert seen["docs_dir"] == base == seen["bundle_dir"]

View file

@ -60,7 +60,7 @@ from portfolio_optimiser.reference_domain import CostItem, Project
from portfolio_optimiser import run as run_mod
from portfolio_optimiser.run import grounding_offer_notice, run_project
from portfolio_optimiser.simulation import scripted_factory
from portfolio_optimiser.validator import Rejection
from portfolio_optimiser.validator import Grounding, Rejection
FIXTURES = Path(__file__).parent / "fixtures" / "p7-grounding"
SHARED = Path(__file__).resolve().parents[1] / "shared" / "examples"
@ -125,7 +125,7 @@ def test_a_delivered_input_with_no_cost_line_reports_zero() -> None:
measured form and the run is un-anchored, so the offer is null in BOTH numbers. Without this
control a green (b) would be satisfied by a report that always counts positive."""
for name in ("p6-k2-generation-prompt.txt", "s7c-k2-generation-prompt.txt"):
offer = grounding_offer(_project(), None, _prompt(name))
offer = grounding_offer(_project(), None, Grounding.of(_prompt(name)))
assert offer.identifiers == 0, (name, offer)
assert offer.cost_lines == 0, (name, offer)
assert offer.chars > 0, "the measurement must have had text to measure"
@ -140,7 +140,7 @@ def test_the_known_positive_input_reports_a_positive_offer() -> None:
pass (a) and look measured."""
text = _prompt("p4-n100-generation-prompt.txt")
assert KNOWN_POSITIVE in text, "the fixture no longer carries the known positive"
offer = grounding_offer(_project(), None, text)
offer = grounding_offer(_project(), None, Grounding.of(text))
assert offer.identifiers > 0, offer
assert offer.cost_lines == 0, "a road standard carries no cost lines (F4's own finding)"
@ -149,7 +149,7 @@ def test_a_bare_number_is_not_counted_as_an_offer() -> None:
"""LOAD-BEARING (b), the other half. K2 carries 46 394 bare-number occurrences over 2 117
distinct values (P7 § 2), so counting them would make every report positive and the whole
measurement inert the repo's cardinal class, a gate that can only come out green."""
offer = grounding_offer(_project(), None, "1234 5678 90 42.5 1000000")
offer = grounding_offer(_project(), None, Grounding.of("1234 5678 90 42.5 1000000"))
assert offer.identifiers == 0, offer
@ -161,12 +161,12 @@ def test_the_offer_is_measured_on_the_text_the_gate_will_see() -> None:
composer P7's gate uses, so the report and the gate cannot describe different texts. Both of
the composer's OTHER two sources are exercised: a project cost line and a baseline code each
raise the count, which a report built from ``delivered`` alone cannot do."""
delivered = "nothing citable here"
delivered = Grounding.of("nothing citable here")
project, baseline = _project("PRJ-77"), _baseline("BAS-88")
offer = grounding_offer(project, baseline, delivered)
assert offer.chars == len(_grounding_text(project, baseline, delivered))
assert offer.chars == len(_grounding_text(project, baseline, delivered).text)
assert offer.identifiers == 2, offer
assert offer.cost_lines == 1, offer

View file

@ -27,7 +27,12 @@ from portfolio_optimiser.ir import (
CostBaselineLine,
SavingsProposal,
)
from portfolio_optimiser.validator import Rejection, ValidatedProposal, validate_proposal
from portfolio_optimiser.validator import (
Grounding,
Rejection,
ValidatedProposal,
validate_proposal,
)
FIXTURES = Path(__file__).parent / "fixtures" / "p7-grounding"
@ -81,7 +86,7 @@ def test_a_fabricated_code_from_a_real_recording_falls(
text = _prompt(fixture)
for code in codes:
assert code not in text, f"fixture drifted: {code!r} is IN the prompt"
ruling = validate_proposal(_proposal(*codes), baseline=None, grounding=text)
ruling = validate_proposal(_proposal(*codes), baseline=None, grounding=Grounding.of(text))
assert isinstance(ruling, Rejection), f"{codes} cleared the gate on an input that names neither"
for code in codes:
assert repr(code) in ruling.reason
@ -92,7 +97,9 @@ def test_b_the_reason_names_every_ungrounded_identifier_in_the_proposal_s_own_or
message naming only the FIRST violation reads as an instruction to fix that one field, and the
proposer fixes one and rebreaks the other. Same ``"; "`` joiner, same PROPOSAL order."""
text = _prompt("p6-k2-generation-prompt.txt")
ruling = validate_proposal(_proposal("M-04-03", "M-04-01"), baseline=None, grounding=text)
ruling = validate_proposal(
_proposal("M-04-03", "M-04-01"), baseline=None, grounding=Grounding.of(text)
)
assert isinstance(ruling, Rejection)
parts = ruling.reason.split("; ")
assert len(parts) == 2, f"expected one sentence per violation, got {ruling.reason!r}"
@ -110,7 +117,9 @@ def test_c_the_known_positive_is_not_flagged() -> None:
is red here."""
text = _prompt("p4-n100-generation-prompt.txt")
assert KNOWN_POSITIVE in text, "fixture drifted: the known positive is not in the prompt"
ruling = validate_proposal(_proposal(KNOWN_POSITIVE), baseline=None, grounding=text)
ruling = validate_proposal(
_proposal(KNOWN_POSITIVE), baseline=None, grounding=Grounding.of(text)
)
assert isinstance(ruling, ValidatedProposal), getattr(ruling, "reason", "")
@ -120,7 +129,7 @@ def test_c_control_the_rule_can_still_flag_on_that_same_recording() -> None:
fixture, same call only the identifier differs, and this one must fall."""
text = _prompt("p4-n100-generation-prompt.txt")
assert "CRS-01" not in text
ruling = validate_proposal(_proposal("CRS-01"), baseline=None, grounding=text)
ruling = validate_proposal(_proposal("CRS-01"), baseline=None, grounding=Grounding.of(text))
assert isinstance(ruling, Rejection)
assert "'CRS-01'" in ruling.reason
@ -134,7 +143,7 @@ def test_d_a_code_quoted_verbatim_from_the_input_is_not_flagged() -> None:
"""The discriminator between this rule and "flag anything that looks like a code". The token is
deliberately shaped like the fabrications above; the ONLY difference is that the input says it."""
text = "Context:\nPrice schedule line ZZZ-999-01 covers technical marking.\n"
ruling = validate_proposal(_proposal("ZZZ-999-01"), baseline=None, grounding=text)
ruling = validate_proposal(_proposal("ZZZ-999-01"), baseline=None, grounding=Grounding.of(text))
assert isinstance(ruling, ValidatedProposal), getattr(ruling, "reason", "")
@ -148,7 +157,8 @@ def test_e_the_stage_fires_with_no_baseline_at_all() -> None:
proposal because stage 0 sits behind ``if baseline is not None``."""
text = _prompt("p6-k2-generation-prompt.txt")
assert isinstance(
validate_proposal(_proposal("M-04-01"), baseline=None, grounding=text), Rejection
validate_proposal(_proposal("M-04-01"), baseline=None, grounding=Grounding.of(text)),
Rejection,
)
@ -165,7 +175,8 @@ def test_e_the_stage_is_not_gated_on_the_absence_of_a_baseline() -> None:
project_id="K2", items={"M-04-01": CostBaselineLine(quantity=10.0, unit_cost=100.0)}
)
assert isinstance(
validate_proposal(_proposal("M-04-01"), baseline=baseline, grounding=text), Rejection
validate_proposal(_proposal("M-04-01"), baseline=baseline, grounding=Grounding.of(text)),
Rejection,
)
@ -245,7 +256,9 @@ def test_g_the_seam_grounds_a_code_the_baseline_proves_even_when_no_prompt_repea
baseline = CostBaseline(
project_id="K2", items={"M-04-01": CostBaselineLine(quantity=10.0, unit_cost=100.0)}
)
text = _grounding_text(project, baseline, "the debate summarised this in prose, naming no code")
text = _grounding_text(
project, baseline, Grounding.of("the debate summarised this in prose, naming no code")
)
assert isinstance(
validate_proposal(_proposal("M-04-01"), baseline=baseline, grounding=text),
ValidatedProposal,
@ -259,7 +272,9 @@ def test_g_the_seam_grounds_a_code_the_delivered_base_carries() -> None:
from portfolio_optimiser.generate import _grounding_text
project = _project()
text = _grounding_text(project, None, "the price schedule line BASE-77-01 is real")
text = _grounding_text(
project, None, Grounding.of("the price schedule line BASE-77-01 is real")
)
assert isinstance(
validate_proposal(_proposal("BASE-77-01"), baseline=None, grounding=text),
ValidatedProposal,

View file

@ -0,0 +1,221 @@
"""P18/B1 — an identifier that stands in every document identifies none of them.
P7 made stage 0b ``item.code in grounding``: plain containment over ONE concatenated string. P16
then ran it against a delivered corpus and measured what containment cannot tell apart. The
falsification arm ``a4-indeksregulering`` proposed a 250 000 NOK saving on a single cost line whose
code was ``R761`` the knowledge base's OWN NAME, which every one of its 2 756 concept documents
carries and the whole gate said ``validated``: stage 0 was skipped (un-anchored run), stage 0b
was satisfied by the letterhead, and the checker approved.
The rule this file measures: a code grounds only if it is at least ``_GROUNDING_MIN_LENGTH``
characters AND appears in fewer than ``_GROUNDING_MAX_DOCUMENT_SHARE`` of the grounding's
DOCUMENTS with an absolute floor, because a share over a handful of documents is not a
measurement (one of three is 33 % and says nothing).
**N and A are MEASURED, not chosen** (14.09, the four mounted vegnormal bases):
* every ``must_cite`` reference and every mandate ``affected_code`` in the four context sets: the
shortest real identifier is FOUR characters (``12.1``, ``52.1``), so ``N = 3`` sits one below the
measurement and cannot refuse anything measured;
* document frequency of every code-shaped token (``generate._IDENTIFIER_FORMS``) in each base:
1 692 distinct tokens and NOT ONE reaches 5 % of its base's documents. Highest anywhere 6 of 446
(1.35 %); highest that a fasit names 3 of 446 (0.67 %); ``R761`` 2 756 of 2 756 (100 %). ``A =
0.05`` therefore sits 3.7x above the highest real token and 20x below the defect.
**The denominator is NAMED in the refusal**, because Step 5 feeds that reason verbatim into the
next attempt's prompt: a proposer told only "ungrounded" answers with another token of the same
kind, while one told "it is in 2 756 of 2 756 documents" has been told what is wrong with it.
The arms that need the delivered bases SKIP with the root named; the rule's own algebra, the
floor, and the composition seam run over synthetic input and are UNCONDITIONAL.
"""
from __future__ import annotations
import json
import os
from pathlib import Path
import pytest
from portfolio_optimiser import okf
from portfolio_optimiser.ir import AffectedItem, SavingsProposal
from portfolio_optimiser.validator import (
Grounding,
Rejection,
ValidatedProposal,
_inert_in,
validate_proposal,
)
_DEFAULT_BUNDLE_ROOT = Path.home() / "repos" / "vegnormal-okf" / "build" / "ferdig"
_A4 = Path(
"scratchpad/p14-stress/kontrakt-sorasen-2027/"
"kontrakt-sorasen-2027-01-a4-indeksregulering-proposal.json"
)
def _base(name: str) -> Path:
root = Path(os.environ.get("PORTFOLIO_VEGNORMAL_ROOT", str(_DEFAULT_BUNDLE_ROOT)))
if not (root / name).is_dir():
pytest.skip(f"knowledge base {name!r} is not mounted under {root}")
return root / name
def _grounding_over(name: str) -> Grounding:
"""The delivered base as ``run_project`` composes it: ONE document per concept file."""
bundle = okf.navigate_bundle(str(_base(name)))
return Grounding(
documents=tuple(
"\n".join([f.name, *f.frontmatter.values(), f.body]) for f in bundle.context_files
)
)
def _proposal(code: str, *, saving: float = 1000.0) -> SavingsProposal:
return SavingsProposal(
project_id="p",
measure="m",
affected_items=[AffectedItem(code=code, quantity=1.0, unit_cost=100_000.0)],
claimed_saving_nok=saving,
)
def _corpus(*, documents: int, everywhere: str, once: str) -> Grounding:
"""A synthetic grounding: one token in every document, one in exactly one."""
return Grounding(
documents=tuple(
f"{everywhere} paragraf {n}" + (f" {once}" if n == 0 else "") for n in range(documents)
)
)
# --- the measured defect --------------------------------------------------------------------
def test_the_a4_proposal_p16_validated_is_now_refused() -> None:
"""(a) THE KNOWN POSITIVE, and it is a recording rather than a construction: the proposal is
the artefact P16's paid run wrote, replayed against the base that run was given, with
``baseline=None`` exactly the configuration under which it said ``validated``."""
if not _A4.is_file():
pytest.skip("P16's a4 artefact is not present in this checkout")
proposal = SavingsProposal.model_validate(
json.loads(_A4.read_text(encoding="utf-8"))["proposal"]
)
assert [i.code for i in proposal.affected_items] == ["R761"], "the artefact drifted"
ruling = validate_proposal(proposal, baseline=None, grounding=_grounding_over("r761-2025"))
assert isinstance(ruling, Rejection)
assert "'R761'" in ruling.reason
assert "2756 of the 2756" in ruling.reason, (
"the refusal must name the denominator: Step 5 feeds this reason verbatim into the next "
f"attempt's prompt — got {ruling.reason!r}"
)
def test_every_fasit_reference_still_grounds() -> None:
"""(b) THE KNOWN NEGATIVE over the same corpora, with its denominator stated. A rule that made
the defect inert by making real references inert too would pass (a) perfectly."""
sets = {
"gate-nordvik-2027": "n100-2023",
"fv412-dekkefornyelse-2027": "n200-2024",
"tunnel-hauglia-2027": "n500-2024",
"kontrakt-sorasen-2027": "r761-2025",
}
checked = 0
for context, base in sets.items():
fasit = json.loads(Path(f"contexts/{context}/fasit.json").read_text(encoding="utf-8"))
grounding = _grounding_over(base)
for reference in sorted({c["ref"] for m in fasit["must_cite"] for c in m["concepts"]}):
assert reference in grounding.text, f"{reference!r} is absent from {base}"
assert _inert_in(grounding, reference) is None, (
f"{reference!r} is a real requirement of {base} and the rule made it inert"
)
checked += 1
assert checked == 26, f"population moved: {checked} references, expected 26"
# --- the rule's own algebra, unconditional ---------------------------------------------------
def test_a_token_in_one_document_grounds_and_one_in_all_of_them_does_not() -> None:
"""(c) The discriminator, over synthetic input so it can never be absent. Both halves in one
arm on the SAME corpus: a rule that flagged everything and one that flagged nothing each fail
exactly one of them."""
grounding = _corpus(documents=100, everywhere="KORPUS-01", once="LINJE-77-01")
assert _inert_in(grounding, "LINJE-77-01") is None
assert _inert_in(grounding, "KORPUS-01") is not None
assert isinstance(
validate_proposal(_proposal("LINJE-77-01"), grounding=grounding), ValidatedProposal
)
assert isinstance(validate_proposal(_proposal("KORPUS-01"), grounding=grounding), Rejection)
def test_a_token_too_short_to_identify_anything_is_inert() -> None:
"""(d) The length conjunct, which the SHARE does not cover: ``R761`` is four characters, so
length is not what made the measured defect inert. This is the coincidence class the
measurement did not happen to contain a one- or two-character token is in any prose."""
grounding = Grounding(documents=("the line A is here", *("filler" for _ in range(50))))
assert _inert_in(grounding, "A") is not None
assert "too short" in str(_inert_in(grounding, "A"))
assert _inert_in(grounding, "A-1") is None, "three characters is the measured floor, not four"
def test_a_share_is_not_taken_over_a_handful_of_documents() -> None:
"""(e) The absolute floor, and the reason every pre-P18 fixture is untouched by this rule
rather than exempted from it: one document of three is 33 % and says nothing at all. Measured,
the highest ABSOLUTE document count any real identifier reaches in the four corpora is 6."""
tiny = Grounding(documents=("KODE-01 her", "KODE-01 og her", "KODE-01 og her"))
assert _inert_in(tiny, "KODE-01") is None, "3 of 3 is 100 %, and it is not a measurement"
assert isinstance(validate_proposal(_proposal("KODE-01"), grounding=tiny), ValidatedProposal)
def test_a_caller_that_declares_no_boundaries_is_byte_for_byte_the_old_gate() -> None:
"""(f) ``Grounding.of`` is the honest reading of a caller with nothing to declare, and it can
never trip the share: one document cannot reach the floor. This is what keeps every road-path
run and every pre-P18 test unchanged BY CONSTRUCTION rather than by exemption."""
text = "en tekst som nevner KODE-99 og ellers ingenting"
single = Grounding.of(text)
assert single.text == text, "the one-document form must not reshape the text"
assert single.document_frequency("KODE-99") == 1
assert _inert_in(single, "KODE-99") is None
# --- the seam: the boundaries reach the gate from the run ------------------------------------
def test_the_run_hands_the_gate_one_document_per_concept_file() -> None:
"""(g) The COMPOSITION arm. The rule is only as good as the boundaries it is given: a run that
still composed one blob would satisfy every arm above (which builds its own ``Grounding``) and
reproduce the measured defect exactly. Driven through ``_grounding_text``, the one composer the
run passes to the gate, and asserted on the COUNT of documents rather than on the text."""
from portfolio_optimiser.generate import _grounding_text
from portfolio_optimiser.ir import CostBaseline, CostBaselineLine
from portfolio_optimiser.reference_domain import CostItem, Project
delivered = Grounding(documents=("dokument A", "dokument B", "dokument C"))
project = Project(
id="p",
name="P",
description="d",
currency="NOK",
cost_items=(
CostItem(code="PRJ-01", description="d", unit="stk", quantity=1.0, unit_cost=1.0),
),
docs_dir="/nonexistent",
)
baseline = CostBaseline(
project_id="p", items={"BAS-01": CostBaselineLine(quantity=1.0, unit_cost=1.0)}
)
composed = _grounding_text(project, baseline, delivered)
assert len(composed.documents) == 5, "each later source is ONE document, never appended to one"
assert composed.text == "\n".join(
["dokument A", "dokument B", "dokument C", "PRJ-01", "BAS-01"]
)

View file

@ -286,15 +286,43 @@ def test_e_the_ref_in_the_measure_names_the_concept(tmp_path: Path) -> None:
assert row.named_in_measure is True
def test_e_a_measure_that_names_nothing_is_carried_only_by_a_snippet(tmp_path: Path) -> None:
def test_e_a_measure_that_names_nothing_is_carried_only_by_a_narrowed_snippet(
tmp_path: Path,
) -> None:
"""The snippet arm still carries (b') — but only under a NARROWED scope (P18/C2)."""
row = _judge(
tmp_path, measure="Do it cheaper", citation_snippet=f"see {_REF}", tool_calls=_opened(_GOOD)
tmp_path,
measure="Do it cheaper",
citation_files=[_GOOD],
citation_snippet=f"see {_REF}",
tool_calls=_opened(_GOOD),
).approaches[0]
assert row.citation_scope == "narrowed"
assert row.named_in_measure is False
assert row.named_in_snippet is True
assert row.named is True
def test_e_a_whole_base_snippet_does_not_name_the_concept(tmp_path: Path) -> None:
"""P18/C2 (PM decision, P16 § 6.2). ``bundle_citations`` stamps EVERY context file before any
model call, so a mark found in a whole-base snippet is evidence about what the base contains,
not about what this run said. Measured on r761: the process number ``12.1`` stands in the
bodies themselves, so that row came back ``named`` for a run that never named it.
The pair with the arm above is the discriminator: the SAME snippet, the same mark, and the only
difference is the scope."""
row = _judge(
tmp_path,
measure="Do it cheaper",
citation_files=[_GOOD, _OTHER],
citation_snippet=f"see {_REF}",
tool_calls=_opened(_GOOD),
).approaches[0]
assert row.citation_scope == "whole-base"
assert row.named_in_snippet is False
assert row.named is False
def test_e_naming_neither_way_fails_b_prime(tmp_path: Path) -> None:
row = _judge(tmp_path, measure="Do it cheaper", tool_calls=_opened(_GOOD)).approaches[0]
assert row.named is False