feat(p18): an identifier that stands everywhere identifies nothing
P18 parts B and C (order 20260914T105139Z).
B1 -- stage 0b. P7 made it `item.code in grounding`: plain containment over
ONE concatenated string. P16 ran it against a delivered corpus and measured
what containment cannot tell apart: the falsification arm a4-indeksregulering
put 250 000 NOK on a single cost line coded R761 -- the knowledge base's OWN
NAME, carried by all 2 756 of its concept documents -- and the whole gate
said validated (stage 0 skipped, un-anchored run; checker approve).
The grounding is now carried as the DOCUMENTS it is made of (validator.
Grounding), not as a blob. A structure and not a second argument beside the
text: the boundaries and the text are one fact, and .text is derived, so the
gate and P8's report measure the same characters. run.py composes one
document per concept file where the base is already walked; generate.
_grounding_text folds each cost line in as a one-line document.
N and A are MEASURED, not chosen (14.09, four mounted vegnormal bases):
- every must_cite ref and mandate affected_code in the four context sets --
shortest real identifier is FOUR characters (12.1, 52.1), so N = 3 sits one
below the measurement and cannot refuse anything measured;
- document frequency of every code-shaped token per base -- 1 692 distinct
and NOT ONE reaches 5 %. Highest anywhere 6/446 (1.35 %), highest a fasit
names 3/446 (0.67 %), R761 2 756/2 756 (100 %). A = 0.05 therefore sits
3.7x above the highest real token and 20x below the defect.
Length is NOT what makes the defect inert (R761 is four characters); the
share is. And a share is not a measurement without a denominator big enough
to take one (ansikt 4): one of three is 33 %, so an ABSOLUTE floor of 10
documents gates it. Highest absolute count any real identifier reaches is 6,
and every fixture in the repo is far below 10 -- which is why every pre-P18
gate is UNTOUCHED by this rule rather than exempted from it. Grounding.of
(one document) can never reach the floor by construction.
The refusal NAMES the denominator ("appears in 2756 of the 2756 documents
this run was given"), because Step 5 feeds that reason verbatim into the next
attempt's prompt: a proposer told only "ungrounded" answers with another
token of the same kind.
B2 SPIKE (measured, NOT built) FELLED the order's own alternative: option (b)
"ground in what the run OPENED" was run over P16's 16 code rows -- R761
stands in every OPENED document too, so (b) would NOT have caught the defect,
while B1 makes it inert and still grounds the real process line 65
ASFALTDEKKER (29/2756 = 1.05 %). (b) is not a substitute for B1.
C1 -- --docs-dir is optional once --bundle-dir is given (P16 FUNN 2). On the
bundle path docs_dir is never read: retrieval, the chunk tool and the "no
citable content" check all live in the road branch. Bound ONCE from
--bundle-dir, which is byte-identically what the README already tells an
operator to type by hand. NOT the "--docs-dir omvei": no such path is opened
and the road branch still refuses without a real --docs-dir (own arm).
C2 -- the judge's snippet arm counts only under citation_scope == "narrowed",
as (a) already does (PM decision, P16 s 6.2). P16's reason for (b') being
clean -- snippets are bodies while ref/title live in frontmatter, 0 of 446
n100 bodies -- holds for "Krav 4.1.2-1" but NOT for R761, where a process
number like 12.1 stands in the bodies. Under a whole-base citation list that
mark was "cited" before any model call.
tests: test_inert_identifier_loadbearing.py (7 arms; known positive is P16's
OWN artefact replayed against the base that run was given, known negative is
26 of 26 fasit references still grounding), test_docs_dir_optional_
loadbearing.py (5 arms). test_stress_judge_loadbearing.py's snippet arm split
into narrowed/whole-base -- the pair is the discriminator, same snippet, same
mark, only the scope differs. The grounding tests migrate from str to
Grounding.of (the honest reading of a caller that declared no boundaries).
Verification: uv run pytest -q 1698 passed / 5 skipped (1685 after part A,
strict superset, 0 removed). ruff check + format clean, mypy clean (38
files). Golden demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
9b47e5aa62
commit
7a7c988253
11 changed files with 628 additions and 51 deletions
|
|
@ -40,6 +40,7 @@ from portfolio_optimiser.proposal_review import (
|
|||
)
|
||||
from portfolio_optimiser.reference_domain import Project
|
||||
from portfolio_optimiser.validator import (
|
||||
Grounding,
|
||||
Rejection,
|
||||
ValidatedProposal,
|
||||
self_repair,
|
||||
|
|
@ -421,12 +422,20 @@ def generate_with_validation(
|
|||
return self_repair(_attempt, max_attempts=max_attempts)
|
||||
|
||||
|
||||
def _grounding_text(project: Project, baseline: CostBaseline | None, delivered: str) -> str:
|
||||
def _grounding_text(
|
||||
project: Project, baseline: CostBaseline | None, delivered: Grounding
|
||||
) -> Grounding:
|
||||
"""P7: compose the ONE text a candidate's identifiers must be grounded in — the run's
|
||||
non-model-authored input, and nothing else.
|
||||
|
||||
Three sources, each of which the run can point at without asking the model:
|
||||
|
||||
**P18/B1: the result carries the DOCUMENT BOUNDARIES, not only the text.** ``code in text``
|
||||
cannot tell "this project has such a line" from "this word is in every letterhead" — P16
|
||||
measured a base's own name, ``R761``, carrying a fabricated 250 000 NOK line to ``validated``.
|
||||
The two later sources join as ONE-LINE documents rather than being appended to a blob: each IS
|
||||
one cost line, and the share rule then reads them exactly as it reads a concept file.
|
||||
|
||||
* ``delivered`` — what the CALLER can prove this run was GIVEN. ``run_project`` fills it from
|
||||
the delivered rendered context (the pre-pass cut, the bundle pointer, or the road path's
|
||||
retrieved chunks) PLUS the navigated base's ``context_files`` — never ``files``, because that
|
||||
|
|
@ -448,12 +457,12 @@ def _grounding_text(project: Project, baseline: CostBaseline | None, delivered:
|
|||
very identifier it just refused. Grounding in the prompt would therefore let the gate's own
|
||||
refusal ground the next attempt: a falsifier that disarms itself on its second round.
|
||||
"""
|
||||
return "\n".join(
|
||||
[
|
||||
delivered,
|
||||
return Grounding(
|
||||
documents=(
|
||||
*delivered.documents,
|
||||
*(item.code for item in project.cost_items),
|
||||
*(() if baseline is None else baseline.items),
|
||||
]
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -508,7 +517,7 @@ class GroundingOffer:
|
|||
|
||||
|
||||
def grounding_offer(
|
||||
project: Project, baseline: CostBaseline | None, delivered: str
|
||||
project: Project, baseline: CostBaseline | None, delivered: Grounding
|
||||
) -> GroundingOffer:
|
||||
"""Measure what ``delivered`` can ground, on the EXACT text the gate will see.
|
||||
|
||||
|
|
@ -519,7 +528,7 @@ def grounding_offer(
|
|||
|
||||
Deterministic and free: no model call, no network, and no second walk of the bundle.
|
||||
"""
|
||||
text = _grounding_text(project, baseline, delivered)
|
||||
text = _grounding_text(project, baseline, delivered).text
|
||||
found: set[str] = set()
|
||||
for form in _IDENTIFIER_FORMS:
|
||||
found |= set(form.findall(text))
|
||||
|
|
@ -544,7 +553,7 @@ async def generate_via_llm(
|
|||
reviews: list[ProposalReview] | None = None,
|
||||
review_key: tuple[str | None, str | None] = (None, None),
|
||||
checker_verdict: str = "absent",
|
||||
grounding: str | None = None,
|
||||
grounding: Grounding | None = None,
|
||||
) -> GenerationResult:
|
||||
"""Async LLM path: non-streaming chat -> parse -> validate, with TWO bounded retry kinds,
|
||||
the meter checked in this loop:
|
||||
|
|
@ -703,7 +712,7 @@ async def generate_via_llm(
|
|||
# caller that declared nothing else IS the input it declared. ``run_project``
|
||||
# always passes it EXPLICITLY, because on the debate path ``context`` has been
|
||||
# replaced by the model's OWN summary of what it read.
|
||||
context if grounding is None else grounding,
|
||||
Grounding.of(context) if grounding is None else grounding,
|
||||
),
|
||||
)
|
||||
last_ruling = result
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue