feat(p18): an identifier that stands everywhere identifies nothing
P18 parts B and C (order 20260914T105139Z).
B1 -- stage 0b. P7 made it `item.code in grounding`: plain containment over
ONE concatenated string. P16 ran it against a delivered corpus and measured
what containment cannot tell apart: the falsification arm a4-indeksregulering
put 250 000 NOK on a single cost line coded R761 -- the knowledge base's OWN
NAME, carried by all 2 756 of its concept documents -- and the whole gate
said validated (stage 0 skipped, un-anchored run; checker approve).
The grounding is now carried as the DOCUMENTS it is made of (validator.
Grounding), not as a blob. A structure and not a second argument beside the
text: the boundaries and the text are one fact, and .text is derived, so the
gate and P8's report measure the same characters. run.py composes one
document per concept file where the base is already walked; generate.
_grounding_text folds each cost line in as a one-line document.
N and A are MEASURED, not chosen (14.09, four mounted vegnormal bases):
- every must_cite ref and mandate affected_code in the four context sets --
shortest real identifier is FOUR characters (12.1, 52.1), so N = 3 sits one
below the measurement and cannot refuse anything measured;
- document frequency of every code-shaped token per base -- 1 692 distinct
and NOT ONE reaches 5 %. Highest anywhere 6/446 (1.35 %), highest a fasit
names 3/446 (0.67 %), R761 2 756/2 756 (100 %). A = 0.05 therefore sits
3.7x above the highest real token and 20x below the defect.
Length is NOT what makes the defect inert (R761 is four characters); the
share is. And a share is not a measurement without a denominator big enough
to take one (ansikt 4): one of three is 33 %, so an ABSOLUTE floor of 10
documents gates it. Highest absolute count any real identifier reaches is 6,
and every fixture in the repo is far below 10 -- which is why every pre-P18
gate is UNTOUCHED by this rule rather than exempted from it. Grounding.of
(one document) can never reach the floor by construction.
The refusal NAMES the denominator ("appears in 2756 of the 2756 documents
this run was given"), because Step 5 feeds that reason verbatim into the next
attempt's prompt: a proposer told only "ungrounded" answers with another
token of the same kind.
B2 SPIKE (measured, NOT built) FELLED the order's own alternative: option (b)
"ground in what the run OPENED" was run over P16's 16 code rows -- R761
stands in every OPENED document too, so (b) would NOT have caught the defect,
while B1 makes it inert and still grounds the real process line 65
ASFALTDEKKER (29/2756 = 1.05 %). (b) is not a substitute for B1.
C1 -- --docs-dir is optional once --bundle-dir is given (P16 FUNN 2). On the
bundle path docs_dir is never read: retrieval, the chunk tool and the "no
citable content" check all live in the road branch. Bound ONCE from
--bundle-dir, which is byte-identically what the README already tells an
operator to type by hand. NOT the "--docs-dir omvei": no such path is opened
and the road branch still refuses without a real --docs-dir (own arm).
C2 -- the judge's snippet arm counts only under citation_scope == "narrowed",
as (a) already does (PM decision, P16 s 6.2). P16's reason for (b') being
clean -- snippets are bodies while ref/title live in frontmatter, 0 of 446
n100 bodies -- holds for "Krav 4.1.2-1" but NOT for R761, where a process
number like 12.1 stands in the bodies. Under a whole-base citation list that
mark was "cited" before any model call.
tests: test_inert_identifier_loadbearing.py (7 arms; known positive is P16's
OWN artefact replayed against the base that run was given, known negative is
26 of 26 fasit references still grounding), test_docs_dir_optional_
loadbearing.py (5 arms). test_stress_judge_loadbearing.py's snippet arm split
into narrowed/whole-base -- the pair is the discriminator, same snippet, same
mark, only the scope differs. The grounding tests migrate from str to
Grounding.of (the honest reading of a caller that declared no boundaries).
Verification: uv run pytest -q 1698 passed / 5 skipped (1685 after part A,
strict superset, 0 removed). ruff check + format clean, mypy clean (38
files). Golden demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
9b47e5aa62
commit
7a7c988253
11 changed files with 628 additions and 51 deletions
|
|
@ -27,6 +27,7 @@ import warnings
|
|||
from collections.abc import Callable, Mapping
|
||||
from contextlib import contextmanager
|
||||
from dataclasses import dataclass
|
||||
from typing import Final
|
||||
|
||||
import pulp
|
||||
|
||||
|
|
@ -208,7 +209,95 @@ def _reconcile_against_baseline(
|
|||
return Rejection(proposal=proposal, reason="; ".join(violations))
|
||||
|
||||
|
||||
def _ground_against_input(proposal: SavingsProposal, grounding: str) -> Rejection | None:
|
||||
#: P18/B1 — the shortest identifier the gate will let ground anything.
|
||||
#:
|
||||
#: MEASURED (14.09) over every ``must_cite`` reference and every mandate ``affected_code`` in the
|
||||
#: four context sets: the shortest real identifier is FOUR characters (R761's ``12.1`` / ``52.1``).
|
||||
#: Set one BELOW that, so the rule cannot refuse anything that has been measured, while a one- or
|
||||
#: two-character token — which matches by coincidence in any prose — grounds nothing. The honest
|
||||
#: failure direction for a gate that speaks about a model's invention.
|
||||
#:
|
||||
#: Length is NOT what makes the measured defect inert: P16's fabricated ``R761`` is four characters
|
||||
#: long. The share below is what does. This covers the coincidence class the measurement did not
|
||||
#: happen to contain.
|
||||
_GROUNDING_MIN_LENGTH: Final = 3
|
||||
|
||||
#: The share of the grounding's DOCUMENTS above which a token identifies nothing.
|
||||
#:
|
||||
#: MEASURED over the four delivered corpora, counting document frequency for every code-shaped
|
||||
#: token (``generate._IDENTIFIER_FORMS``): 1 692 distinct tokens, and NOT ONE reaches 5 % of its
|
||||
#: base's documents. The highest anywhere is 6 of 446 (1.35 %); the highest that a fasit or mandate
|
||||
#: actually names is 3 of 446 (0.67 %). P16's fabricated ``R761`` is 2 756 of 2 756 — 100 %.
|
||||
#: 5 % therefore sits 3.7x above the highest real token measured and 20x below the defect.
|
||||
_GROUNDING_MAX_DOCUMENT_SHARE: Final = 0.05
|
||||
|
||||
#: …and a share is not a measurement without a denominator big enough to take one (ansikt 4).
|
||||
#: One document of three is 33 % and says nothing at all, so the share only fires once a token is
|
||||
#: in at least this many documents. MEASURED: the highest ABSOLUTE document count any real
|
||||
#: identifier reaches in the four corpora is 6, and every test fixture in this repo is far below
|
||||
#: 10 — which is why every pre-P18 gate is untouched by this rule rather than exempted from it.
|
||||
_GROUNDING_MIN_INERT_DOCUMENTS: Final = 10
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Grounding:
|
||||
"""The run's non-model-authored input, carried as the DOCUMENTS it is made of.
|
||||
|
||||
P7 carried it as ONE string, and P16 measured what that costs: ``R761`` — the base's own NAME,
|
||||
which every one of its 2 756 concept documents carries — satisfied ``code in grounding`` and
|
||||
carried a fabricated 250 000 NOK line through the whole gate to ``validated``. Containment in a
|
||||
concatenation cannot tell "this project has such a line" from "this word is in the letterhead".
|
||||
|
||||
A structure rather than a second argument beside the text: the boundaries and the text are ONE
|
||||
fact, and two carriers for one fact are free to disagree about it (kø-(p)). ``text`` is derived
|
||||
here, so the gate and P8's report measure the very same characters.
|
||||
|
||||
A caller with no boundaries to declare — the road path, a test — builds ONE document from its
|
||||
text and is unchanged by construction: one document can never reach
|
||||
``_GROUNDING_MIN_INERT_DOCUMENTS``, so the share cannot fire on it.
|
||||
"""
|
||||
|
||||
documents: tuple[str, ...]
|
||||
|
||||
@classmethod
|
||||
def of(cls, text: str) -> Grounding:
|
||||
"""One document. The honest reading of a caller that declared no boundaries."""
|
||||
return cls(documents=(text,))
|
||||
|
||||
@property
|
||||
def text(self) -> str:
|
||||
"""The ONE composition. Byte-identical to P7's ``"\n".join`` of the same parts."""
|
||||
return "\n".join(self.documents)
|
||||
|
||||
def document_frequency(self, token: str) -> int:
|
||||
"""How many of the documents contain ``token`` — the numerator, in the unit of the rule."""
|
||||
return sum(1 for document in self.documents if token in document)
|
||||
|
||||
|
||||
def _inert_in(grounding: Grounding, code: str) -> str | None:
|
||||
"""Why ``code`` identifies nothing in this input, or ``None`` when it identifies something.
|
||||
|
||||
The message NAMES THE DENOMINATOR, because "it is everywhere" and "it is not here" are
|
||||
different findings and Step 5 feeds this reason verbatim into the next attempt's prompt: a
|
||||
proposer told only "ungrounded" will re-answer with another token of the same kind.
|
||||
"""
|
||||
if len(code) < _GROUNDING_MIN_LENGTH:
|
||||
return (
|
||||
f"is {len(code)} characters long, too short to identify a cost line — any prose "
|
||||
"contains it by coincidence"
|
||||
)
|
||||
hits = grounding.document_frequency(code)
|
||||
total = len(grounding.documents)
|
||||
floor = max(_GROUNDING_MIN_INERT_DOCUMENTS, total * _GROUNDING_MAX_DOCUMENT_SHARE)
|
||||
if hits >= floor:
|
||||
return (
|
||||
f"appears in {hits} of the {total} documents this run was given — a token that is in "
|
||||
"every document identifies none of them; name a cost line, not the corpus"
|
||||
)
|
||||
return None
|
||||
|
||||
|
||||
def _ground_against_input(proposal: SavingsProposal, grounding: Grounding) -> Rejection | None:
|
||||
"""P7: every identifier the proposal builds on must appear VERBATIM in the input it was built
|
||||
from, or the verdict falls.
|
||||
|
||||
|
|
@ -244,12 +333,19 @@ def _ground_against_input(proposal: SavingsProposal, grounding: str) -> Rejectio
|
|||
deliberately out of scope: the Monte Carlo never samples such a band (``SavingsProposal.
|
||||
_assumption_bands_enclose_unit_cost`` says so in the same words), so it cannot move the verdict,
|
||||
and a check on it would be a branch no recording exercises."""
|
||||
violations = [
|
||||
f"ungrounded identifier {item.code!r}: it appears nowhere in the input this proposal "
|
||||
f"was built from ({len(grounding)} chars)"
|
||||
for item in proposal.affected_items
|
||||
if item.code not in grounding
|
||||
]
|
||||
violations = []
|
||||
for item in proposal.affected_items:
|
||||
if item.code not in grounding.text:
|
||||
violations.append(
|
||||
f"ungrounded identifier {item.code!r}: it appears nowhere in the input this "
|
||||
f"proposal was built from ({len(grounding.text)} chars)"
|
||||
)
|
||||
continue
|
||||
# P18/B1: present is not the same as identifying. An identifier that stands everywhere
|
||||
# identifies nothing, and one too short to be an identifier is matched by coincidence.
|
||||
inert = _inert_in(grounding, item.code)
|
||||
if inert is not None:
|
||||
violations.append(f"ungrounded identifier {item.code!r}: it {inert}")
|
||||
if not violations:
|
||||
return None
|
||||
return Rejection(proposal=proposal, reason="; ".join(violations))
|
||||
|
|
@ -259,7 +355,7 @@ def validate_proposal(
|
|||
proposal: SavingsProposal,
|
||||
*,
|
||||
baseline: CostBaseline | None = None,
|
||||
grounding: str | None = None,
|
||||
grounding: Grounding | None = None,
|
||||
tolerance: float = BASELINE_TOLERANCE_DEFAULT,
|
||||
method_caps: Mapping[str, float] | None = None,
|
||||
) -> ValidatedProposal | Rejection:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue