feat(p18): an identifier that stands everywhere identifies nothing
P18 parts B and C (order 20260914T105139Z).
B1 -- stage 0b. P7 made it `item.code in grounding`: plain containment over
ONE concatenated string. P16 ran it against a delivered corpus and measured
what containment cannot tell apart: the falsification arm a4-indeksregulering
put 250 000 NOK on a single cost line coded R761 -- the knowledge base's OWN
NAME, carried by all 2 756 of its concept documents -- and the whole gate
said validated (stage 0 skipped, un-anchored run; checker approve).
The grounding is now carried as the DOCUMENTS it is made of (validator.
Grounding), not as a blob. A structure and not a second argument beside the
text: the boundaries and the text are one fact, and .text is derived, so the
gate and P8's report measure the same characters. run.py composes one
document per concept file where the base is already walked; generate.
_grounding_text folds each cost line in as a one-line document.
N and A are MEASURED, not chosen (14.09, four mounted vegnormal bases):
- every must_cite ref and mandate affected_code in the four context sets --
shortest real identifier is FOUR characters (12.1, 52.1), so N = 3 sits one
below the measurement and cannot refuse anything measured;
- document frequency of every code-shaped token per base -- 1 692 distinct
and NOT ONE reaches 5 %. Highest anywhere 6/446 (1.35 %), highest a fasit
names 3/446 (0.67 %), R761 2 756/2 756 (100 %). A = 0.05 therefore sits
3.7x above the highest real token and 20x below the defect.
Length is NOT what makes the defect inert (R761 is four characters); the
share is. And a share is not a measurement without a denominator big enough
to take one (ansikt 4): one of three is 33 %, so an ABSOLUTE floor of 10
documents gates it. Highest absolute count any real identifier reaches is 6,
and every fixture in the repo is far below 10 -- which is why every pre-P18
gate is UNTOUCHED by this rule rather than exempted from it. Grounding.of
(one document) can never reach the floor by construction.
The refusal NAMES the denominator ("appears in 2756 of the 2756 documents
this run was given"), because Step 5 feeds that reason verbatim into the next
attempt's prompt: a proposer told only "ungrounded" answers with another
token of the same kind.
B2 SPIKE (measured, NOT built) FELLED the order's own alternative: option (b)
"ground in what the run OPENED" was run over P16's 16 code rows -- R761
stands in every OPENED document too, so (b) would NOT have caught the defect,
while B1 makes it inert and still grounds the real process line 65
ASFALTDEKKER (29/2756 = 1.05 %). (b) is not a substitute for B1.
C1 -- --docs-dir is optional once --bundle-dir is given (P16 FUNN 2). On the
bundle path docs_dir is never read: retrieval, the chunk tool and the "no
citable content" check all live in the road branch. Bound ONCE from
--bundle-dir, which is byte-identically what the README already tells an
operator to type by hand. NOT the "--docs-dir omvei": no such path is opened
and the road branch still refuses without a real --docs-dir (own arm).
C2 -- the judge's snippet arm counts only under citation_scope == "narrowed",
as (a) already does (PM decision, P16 s 6.2). P16's reason for (b') being
clean -- snippets are bodies while ref/title live in frontmatter, 0 of 446
n100 bodies -- holds for "Krav 4.1.2-1" but NOT for R761, where a process
number like 12.1 stands in the bodies. Under a whole-base citation list that
mark was "cited" before any model call.
tests: test_inert_identifier_loadbearing.py (7 arms; known positive is P16's
OWN artefact replayed against the base that run was given, known negative is
26 of 26 fasit references still grounding), test_docs_dir_optional_
loadbearing.py (5 arms). test_stress_judge_loadbearing.py's snippet arm split
into narrowed/whole-base -- the pair is the discriminator, same snippet, same
mark, only the scope differs. The grounding tests migrate from str to
Grounding.of (the honest reading of a caller that declared no boundaries).
Verification: uv run pytest -q 1698 passed / 5 skipped (1685 after part A,
strict superset, 0 removed). ruff check + format clean, mypy clean (38
files). Golden demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
9b47e5aa62
commit
7a7c988253
11 changed files with 628 additions and 51 deletions
|
|
@ -24,11 +24,17 @@ it. So a CITATION grounds an approach only when the citation list is NARROWER th
|
|||
declared pre-pass cut, where the stamp really does name what was read). Both halves are reported
|
||||
either way (``opened`` / ``cited`` / ``citation_scope``), so which one fired stays readable.
|
||||
|
||||
**(b') was checked for the same vacuity and is CLEAN, so the order stands.** ``bundle_citations``
|
||||
snippets are concept BODIES while ``ref``/``title`` live in FRONTMATTER: measured 0 of 446
|
||||
n100 bodies contain ``Krav 4.1.2-1``. The snippet arm can therefore carry (b') without being
|
||||
satisfied by construction. ``named_in_measure`` / ``named_in_snippet`` are still reported apart,
|
||||
because the measure is the model's own prose and a snippet is the base's.
|
||||
**(b') was checked for the same vacuity and is CLEAN for the N corpora, so the order stands.**
|
||||
``bundle_citations`` snippets are concept BODIES while ``ref``/``title`` live in FRONTMATTER:
|
||||
measured 0 of 446 n100 bodies contain ``Krav 4.1.2-1``. ``named_in_measure`` / ``named_in_snippet``
|
||||
are still reported apart, because the measure is the model's own prose and a snippet is the base's.
|
||||
|
||||
**P18/C2 (PM decision, P16 § 6.2): the snippet arm counts only under a NARROWED scope, as (a)
|
||||
does.** The paragraph above holds for a reference like ``Krav 4.1.2-1``, which no body repeats — it
|
||||
does NOT hold for R761, where a process number such as ``12.1`` stands in the bodies themselves.
|
||||
Under a whole-base citation list that mark is "cited" before any model call, so the row was
|
||||
``named`` for a run in which the model had said nothing of the kind. The scope gate is the same
|
||||
correction (a) already carries, applied to the half that was still exposed.
|
||||
|
||||
**A DENOMINATOR, ALWAYS** (Verifiseringsloven ansikt 4). Every verdict names how many tool calls,
|
||||
citations, approach rows and base concepts it saw, and an outbox with no proposal artefact - or a
|
||||
|
|
@ -274,7 +280,13 @@ def score_context_set(
|
|||
snippets = " ".join(str(c.get("snippet", "")) for c in citations)
|
||||
marks = [m for c in concepts for m in (c.get("ref", ""), c.get("title", "")) if m]
|
||||
named_in_measure = any(m in measure for m in marks)
|
||||
named_in_snippet = any(m in snippets for m in marks)
|
||||
# P18/C2 (PM decision, P16 § 6.2): the snippet arm counts ONLY under a narrowed citation
|
||||
# scope, exactly as (a) does. A whole-base citation list is stamped by ``bundle_citations``
|
||||
# before a single model call — measured on n100, 446 context files, 446 citations, 6 of 6
|
||||
# fasit paths "cited" for free — so a mark found in THOSE snippets is evidence about the
|
||||
# base's contents, not about this run. Measured on r761: ``12.1`` appears in whole-base
|
||||
# snippets and gave this row ``named`` without the model having said anything.
|
||||
named_in_snippet = scope == "narrowed" and any(m in snippets for m in marks)
|
||||
|
||||
halluc = [f"citation:{f}" for f in sorted(cited_files - concept_names)]
|
||||
allowed = set(approach.affected_codes) | baseline_codes
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue