docs: general wording for the example-base document counts

Replace the exact document counts of earlier example bases (and the
per-base counts in the sources-format note) with general wording or
N-of-N in prose, comments and docstrings. Percentages and numerators
stay; no constant, assertion or test data changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-23 17:48:44 +02:00
commit 6b1046bc23
Signed by: ktg
SSH key fingerprint: SHA256:JakMjO6FTBBzN0Bhfj9saOoEjaFxlSdYuZQQpM/lF9Q
16 changed files with 56 additions and 56 deletions

View file

@ -880,12 +880,12 @@ when the seam is detached, so the loop cannot silently degrade into theater.
- **An identifier that stands everywhere identifies nothing (P18).** The validator's stage 0b
grounds each proposed cost code in the run's non-model-authored input. Plain containment was not
enough: measured on a live run, a proposal put 250 000 NOK on a line coded with the knowledge
base's own *name*, carried by all 2 756 of its documents — and the whole gate said `validated`.
base's own *name*, carried by every one of its documents — and the whole gate said `validated`.
A code now grounds only if it is at least 3 characters long **and** appears in fewer than 5 % of
the documents the input is made of, with an absolute floor of 10 documents so the share is never
taken over a handful. Both numbers are measured, not chosen: across the four delivered corpora no
code-shaped token of 1 692 distinct ones reaches 5 % (highest 1.35 %), and the shortest real
identifier is 4 characters. The refusal names the denominator — "appears in 2 756 of the 2 756
identifier is 4 characters. The refusal names the denominator — "appears in N of the N
documents this run was given" — because that reason is fed verbatim into the next attempt's
prompt. A caller that declares no document boundaries is unchanged by construction: one document
can never reach the floor.

View file

@ -2391,12 +2391,12 @@
alt finnes (`{run_id}[-{approach}]-proposal.json` · `-outcome.json` · `{run_id}-debate.json`) —
ingen kjoring far et nytt felt. **ORDRENS (a) VAR EN GATE SOM BARE KUNNE BLI GRONN:** pa
S2c-stien stempler `run_project` `citations = bundle_citations(bundle)`, altsa EN sitering per
konseptfil — MALT pa en kravbase i utviklingskorpuset: 446 kontekstfiler, 446 siteringer, **6 av 6 fasit-stier «sitert»
konseptfil — MALT pa en kravbase i utviklingskorpuset: noen hundre kontekstfiler, like mange siteringer, **6 av 6 fasit-stier «sitert»
for ett eneste modellkall**. En sitering grunner derfor KUN under en smalere liste (et erklaert
pre-pass-kutt); ellers ma grunningen komme av et dokument kjoringen faktisk APNET. Begge
halvdeler rapporteres uansett (`opened`/`cited`/`citation_scope`), sa avviket er uttalt, aldri
stille. **(b') ble sjekket for SAMME vakuitet og er REN:** snippetene er konsept-BODYER mens
`ref`/`title` bor i FRONTMATTER (MALT: 0 av 446 kravbase-bodyer inneholder `Krav 4.1.2—1`), sa
`ref`/`title` bor i FRONTMATTER (MALT: ingen av kravbasens bodyer inneholder `Krav 4.1.2—1`), sa
ordrens definisjon star. **NEVNER ALLTID** (ansikt 4): tool_calls, siteringer, approach-rader og
konsepter i basen; en utboks uten proposal-artefakt — eller en base som skannes til null
konsepter — REISER `EmptyMeasurement` i stedet for a rapportere «0 hallusinasjoner».
@ -2517,7 +2517,7 @@
ikke som én blob (P18/B1, 14.09):** P7 gjorde stadium 0b til `item.code in grounding`, ren
delstreng over én sammenslått streng. P16 kjørte den mot et LEVERT korpus og målte hva
containment ikke kan skille: falsifiseringsarmen `a4-indeksregulering` la 250 000 NOK på ÉN
kostlinje kodet med **basens EGET NAVN** (fire tegn), som alle 2 756 konseptdokumenter bærer — og hele
kostlinje kodet med **basens EGET NAVN** (fire tegn), som hvert eneste konseptdokument bærer — og hele
gaten sa `validated` (stage 0 hoppet over, uforankret kjøring; checker `approve`).
**`Grounding` er en STRUKTUR, ikke et andre argument ved siden av teksten:** grensene og teksten
er ÉTT faktum, og to bærere for ett faktum står fritt til å være uenige (kø-(p)); `.text` utledes,
@ -2526,8 +2526,8 @@
KORTESTE ekte identifikatoren FIRE tegn (`12.1`/`52.1`), så `_GROUNDING_MIN_LENGTH = 3` ligger ETT
under målingen og kan ikke nekte noe som er målt; og over dokumentfrekvensen til hvert
kodeformet token i hver base (`_IDENTIFIER_FORMS`) når **ingen av 1 692 distinkte** 5 % —
høyeste noe sted er 6 av 446 (**1,35 %**), høyeste en fasit faktisk navngir 3 av 446 (**0,67 %**),
mens basenavnet er **2 756 av 2 756 (100 %)**. `_GROUNDING_MAX_DOCUMENT_SHARE = 0.05` ligger altså
høyeste noe sted er **1,35 %** (6 dokumenter), høyeste en fasit faktisk navngir **0,67 %** (3 dokumenter),
mens basenavnet står i **alle dokumentene (100 %)**. `_GROUNDING_MAX_DOCUMENT_SHARE = 0.05` ligger altså
3,7× over det høyeste ekte tokenet og 20× under defekten. **Lengden er IKKE dét som gjør defekten
inert** (basenavnet er fire tegn) — andelen er; N dekker sammentreffs-klassen målingen ikke tilfeldigvis
inneholdt. **Og en andel er ingen måling uten en nevner stor nok til å ta den** (ansikt 4): ett av
@ -2536,12 +2536,12 @@
hver fixtur i repoet ligger langt under 10, hvilket er dét som gjør at HVER pre-P18-gate er
URØRT av regelen i stedet for UNNTATT fra den. `Grounding.of` (ett dokument) er den ærlige
lesningen av en kaller uten grenser å erklære og kan ved konstruksjon aldri utløse andelen.
**Nevneren NAVNGIS i nekten** («appears in 2756 of the 2756 documents this run was given»), fordi
**Nevneren NAVNGIS i nekten** («appears in N of the N documents this run was given»), fordi
Steg 5 mater den grunnen ORDRETT inn i neste forsøks prompt: en proposer som bare får «ungrounded»
svarer med et nytt token av samme slag. **B2-SPIKEN FELTE ORDRENS EGEN ALTERNATIV-HYPOTESE:** (b)
«grunnlag = det kjøringen ÅPNET» ble MÅLT over P16s 16 kode-rader — basenavnet står i hvert ÅPNET
dokument også, så (b) ville **ikke** fanget defekten (grunner fortsatt), mens B1 gjør den inert og
lar en ekte seksjonslinje (nummer + versaltittel, formen `65 LAGRINGSSYSTEMER`; 29/2756 = 1,05 %)
lar en ekte seksjonslinje (nummer + versaltittel, formen `65 LAGRINGSSYSTEMER`; 1,05 % av dokumentene)
grunne. (b) er altså ikke et
substitutt for B1; målt, ikke bygget. Load-bearing MÅLT
(`tests/test_inert_identifier_loadbearing.py`, **8 armer**), **sju mutasjoner mot HELE suiten** +
@ -2573,8 +2573,8 @@
5 armer): krev `--docs-dir` igjen (2) · la CLI-en aldri videresende basen (2).
- **Dommerens snippet-arm teller KUN under `citation_scope == "narrowed"` (P18/C2, PM-valg P16
§ 6.2):** (a) hadde alt den korreksjonen; (b′) hadde den ikke. P16s egen begrunnelse for at (b′)
var ren — «snippetene er konsept-BODYER mens `ref`/`title` bor i FRONTMATTER, målt 0 av 446
kravbase-bodyer» — holder for `Krav 4.1.2—1`, men **IKKE for prosesskatalogen**, der et prosessnummer som `12.1`
var ren — «snippetene er konsept-BODYER mens `ref`/`title` bor i FRONTMATTER, målt: ingen av
kravbase-bodyene» — holder for `Krav 4.1.2—1`, men **IKKE for prosesskatalogen**, der et prosessnummer som `12.1`
står i bodyene selv. Under en helbase-siteringsliste var det merket «sitert» før noe modellkall,
så raden kom tilbake `named` for en kjøring der modellen ikke hadde sagt noe slikt. Diskriminatoren
er PARET av to armer med SAMME snippet og SAMME merke, der eneste forskjell er scopet. Mutasjon
@ -2626,8 +2626,8 @@
ikke `ProvenanceStamp` (stempelet beskriver gaten som dømte ÉN kandidat).
- **En «kostkode» må ha en FORM når inputen tilbyr former — og formene bor ÉTT sted, i validatoren
(P19 DEL B, 15.09):** P18s runde 2 endte med TO `validated` forslag hvis `affected_item`-koder var
vanlige ord fra en kravstandards prosa (her gjengitt som `nødstrømsaggregat`, 4 av 270 dokumenter
i én kravbase, og `redundant kjøling`, 4 av 1 133 i en annen). Begge er GRUNNET i P7s forstand (de står ordrett i
vanlige ord fra en kravstandards prosa (her gjengitt som `nødstrømsaggregat`, 4 av noen hundre dokumenter
i én kravbase, og `redundant kjøling`, 4 av rundt tusen i en annen). Begge er GRUNNET i P7s forstand (de står ordrett i
inputen) og ingen av dem er INERT i P18/B1s forstand (langt under 5 %-andelen) — de er bare ikke
identifikatorer for en kostlinje, og gaten hadde ikke noe stadium som kunne si det. **KJENT-POSITIV
MÅLT, ikke påstått:** begge spilt av offline mot basene kjøringene faktisk fikk gir `Rejection`
@ -2757,7 +2757,7 @@
ett multi-base-pass bar `validated` forslag hvis kostkode var et kapittelnummer i en kravstandard.
To overlever i utboksene og er kjent-positivene: **`10.4`** (et kravbase-sett, runde 3, kravbase C) og
**`1.10.4`** (multi-base-passet P17b, prosesskatalogen). Begge er GRUNNET i P7s forstand og ingen er INERT i P18/B1s
(`10.4` i 12 av 274 dokumenter, `1.10.4` i **1 av 2 756**); stadium 0 kjørte aldri, fordi ingen
(`10.4` i 12 av noen hundre dokumenter, `1.10.4` i **1 av noen tusen**); stadium 0 kjørte aldri, fordi ingen
kravbase bærer en kostbaseline. **Ordrens B1 sier: form 2/3 OG «står som
`req_number`/`prosessnr` i toppnivå-frontmatter» → nekt.** MÅLT 15.09: kravbase C erklærer
`seksjon: 10.4.1`…`10.4.4` og `req_number: Krav 10.4.3—2`, men **aldri den bare `10.4`** — den er

View file

@ -1337,7 +1337,7 @@ Produce a REVISED SavingsProposal that resolves this.</code></pre>
<text class="lbl" x="452" y="424">classify_codes: identifier | prose | requirement</text>
<text class="tiny" x="452" y="448">En RAPPORT om hva kjøringen gjorde av koden, ikke porten. Havner i provenance-stemplet som code_forms.</text>
</svg>
<figcaption>Fire underkontroller per berørt post, første treff vinner. Målt i P18 runde 2: to alminnelige ord fra en standards prosa (4 av 270 dokumenter i én kravbase og 4 av 1 133 i en annen, i et kravkorpus brukt under utviklingen) passerte hele porten som kostnadskoder før kodeform-regelen fantes.</figcaption>
<figcaption>Fire underkontroller per berørt post, første treff vinner. Målt i P18 runde 2: to alminnelige ord fra en standards prosa (4 av noen hundre dokumenter i én kravbase og 4 av rundt tusen i en annen, i et kravkorpus brukt under utviklingen) passerte hele porten som kostnadskoder før kodeform-regelen fantes.</figcaption>
</figure>
<p class="src">Kilde: <code>src/portfolio_optimiser/validator.py:471-491</code> · <code>validator.py:494-571</code> · <code>validator.py:574-645</code> · <code>validator.py:388-408</code></p>
</div>
@ -2940,7 +2940,7 @@ Produce a REVISED SavingsProposal that resolves this.</code></pre>
<tbody>
<tr><td>To falsifiserere over én kandidat: validatoren gater tallene, checkeren gater resonnementet. De blandes aldri.</td><td>Metodespeken §2/§6, steg 3 og 4</td><td><code>test_checker_gate_loadbearing.py</code>, rød på begge frakoblingspunkter</td></tr>
<tr><td>Gaten er forankret i prosjektets egen kostnadsbasis. Stadium 0 avstemmer mot <code>CostBaseline</code> før solveren. Validering, aldri reparasjon.</td><td>Før S4.0 resonnerte hvert stadium bare om tall forslaget selv leverte, så en internt konsistent hallusinasjon klarerte hele gaten</td><td><code>test_s40_cost_baseline_loadbearing.py</code>, 6 mutasjoner røde</td></tr>
<tr><td>Et kravnummer er ikke en pris.</td><td>P20 del B, 15.09: <code>10.4</code> er erklært ingen steder og står i 12 av 274 dokumenter. Ordrens egen regel ble målt falsk mot begge sine kjent-positive</td><td><code>docs/invarianter.md:2750</code></td></tr>
<tr><td>Et kravnummer er ikke en pris.</td><td>P20 del B, 15.09: <code>10.4</code> er erklært ingen steder og står i 12 av noen hundre dokumenter. Ordrens egen regel ble målt falsk mot begge sine kjent-positive</td><td><code>docs/invarianter.md:2750</code></td></tr>
<tr><td>En ekspertdom kan ikke oppstå av stillhet. <code>RunResult.verdict</code> er <code>Verdict | None</code>.</td><td>F2, økt 66: hver flaggløs kjøring myntet en godkjenning ingen ga, den gikk i det delte lageret, og den ble båret inn i neste prosjekts prompt — på flaten som ble overlevert 14.08</td><td><code>test_ungiven_verdict_loadbearing.py</code>, 15 armer, 8 mutasjoner røde</td></tr>
<tr><td>Penger kvantiseres i én rekkefølge, fra én kilde: per beløp, så summeres heltallene.</td><td>Tre linjer à 60000,005 NOK er 18 000 003 øre kvantisert først, 18 000 001 summert først. En fiks på bare ett av de to kallstedene overlevde hele suiten</td><td><code>test_money_quantization_loadbearing.py</code>, 5 mutasjoner røde</td></tr>
<tr><td>Sporing er opt-in, og «av» betyr at MAF aldri kalles.</td><td>Et kall med tom exporter-liste ville installert providers og lest hver <code>OTEL_EXPORTER_OTLP_*</code> i omgivelsene. «Av» må være fravær av kallet</td><td><code>test_tracing_loadbearing.py</code>, 9 mutasjoner røde</td></tr>

View file

@ -319,7 +319,7 @@ def _citations_of(payload: Mapping[str, Any]) -> list[dict[str, Any]]:
def _citation_lines(citations: Sequence[Mapping[str, Any]], *, prefix: str) -> list[str]:
"""One quote with the COUNT of places it stands for, or nothing when there is no quote.
The count is part of the citation: a run that cited 446 places and one that cited a single
The count is part of the citation: a run that cited hundreds of places and one that cited a single
place both show one quote here, and a reader who cannot tell them apart cannot tell a
grounded proposal from a decorated one."""
if not citations:

View file

@ -309,10 +309,10 @@ def decode_block_mappings(
applies to the flow carrier. Two grammars would be two answers to one question, and a delivered
base would read differently depending on which spelling its producer chose. The colon-SPACE part
is load-bearing rather than stylistic: every delivered ``resource`` is a URL, so a reader
splitting on the FIRST colon would truncate all 4605 of them at ``https``.
splitting on the FIRST colon would truncate every one of them at ``https``.
Measured 2026-09-12: all four knowledge bases delivered at the time wrote ``sources`` in this
form and none in flow form (446/446, 1133/1133, 270/270 and 2756/2756). Reading it is not a
form and none in flow form (every concept file of every base). Reading it is not a
licence to WRITE it — ``write_concept_file``/``verified_field`` still refuse exactly what
``decode_flow_value`` refuses, so the emission rule and the round-trip gate are untouched.

View file

@ -17,8 +17,8 @@ and the set's own ``mandate.json`` / ``fasit.json`` / ``bundle.txt``.
**THE ORDER'S (a) WAS VACUOUS AS WRITTEN, AND THE DEVIATION IS MEASURED, NOT CHOSEN.** The order
defines grounded as "a ``must_cite`` path was OPENED *or* CITED". But on the S2c navigation path
``run_project`` stamps ``citations = bundle_citations(bundle)``, which is one citation PER CONTEXT
FILE - the whole corpus. Measured on a requirements base during development: 446 context files,
446 citations, and **6 of 6 fasit paths already "cited" before a single model call**. A judge
FILE - the whole corpus. Measured on a requirements base during development: a few hundred context files,
as many citations, and **6 of 6 fasit paths already "cited" before a single model call**. A judge
honouring that literally would be a gate that can only be green - the repo's own vacuous-gate
class, inside the gate built to catch it. So a CITATION grounds an approach only when the citation
list is NARROWER than the base (a declared pre-pass cut, where the stamp really does name what was
@ -27,7 +27,7 @@ one fired stays readable.
**(b') was checked for the same vacuity and is CLEAN for the requirement bases, so the order
stands.** ``bundle_citations`` snippets are concept BODIES while ``ref``/``title`` live in
FRONTMATTER: measured 0 of 446 bodies of that base contain its ``Krav 4.1.2-1``.
FRONTMATTER: measured, not one body of that base contains its ``Krav 4.1.2-1``.
``named_in_measure`` / ``named_in_snippet`` are still reported apart, because the measure is the
model's own prose and a snippet is the base's.
@ -476,8 +476,8 @@ def score_context_set(
named_in_measure = any(m in measure for m in marks)
# P18/C2 (PM decision, P16 § 6.2): the snippet arm counts ONLY under a narrowed citation
# scope, exactly as (a) does. A whole-base citation list is stamped by ``bundle_citations``
# before a single model call — measured on a requirements base, 446 context files, 446
# citations, 6 of 6 fasit paths "cited" for free — so a mark found in THOSE snippets is
# before a single model call — measured on a requirements base, a few hundred context files,
# as many citations, 6 of 6 fasit paths "cited" for free — so a mark found in THOSE snippets is
# evidence about the base's contents, not about this run. Measured on a process catalogue:
# ``12.1`` appears in whole-base snippets and gave this row ``named`` without the model
# having said anything.

View file

@ -296,8 +296,8 @@ _GROUNDING_MIN_LENGTH: Final = 3
#:
#: MEASURED over the four delivered corpora, counting document frequency for every code-shaped
#: token (``generate._IDENTIFIER_FORMS``): 1 692 distinct tokens, and NOT ONE reaches 5 % of its
#: base's documents. The highest anywhere is 6 of 446 (1.35 %); the highest that a fasit or mandate
#: actually names is 3 of 446 (0.67 %). P16's fabricated ``P900`` is 2 756 of 2 756 — 100 %.
#: base's documents. The highest anywhere is 1.35 % (6 documents); the highest that a fasit or mandate
#: actually names is 0.67 % (3 documents). P16's fabricated ``P900`` is in every one — 100 %.
#: 5 % therefore sits 3.7x above the highest real token measured and 20x below the defect.
_GROUNDING_MAX_DOCUMENT_SHARE: Final = 0.05
@ -348,7 +348,7 @@ IDENTIFIER_FORMS: Final = (
# its ``15.09`` prefix, and a date is not a requirement.
_FORM_PROCESS_NUMBER,
# ``65 LAGRINGSSYSTEMER`` — a process number and its heading, the form a price schedule's
# section rows carry (P18 § 2 measured it at 29 of 2 756 documents).
# section rows carry (P18 § 2 measured it at 1.05 % of the documents).
_FORM_PROCESS_HEADING,
)
@ -416,7 +416,7 @@ class Grounding:
"""The run's non-model-authored input, carried as the DOCUMENTS it is made of.
P7 carried it as ONE string, and P16 measured what that costs: ``P900`` — the base's own NAME,
which every one of its 2 756 concept documents carries — satisfied ``code in grounding`` and
which every one of its concept documents carries — satisfied ``code in grounding`` and
carried a fabricated 250 000 NOK line through the whole gate to ``validated``. Containment in a
concatenation cannot tell "this project has such a line" from "this word is in the letterhead".
@ -498,8 +498,8 @@ def _form_refusal(grounding: Grounding, code: str, anchored_codes: frozenset[str
"""Why ``code`` cannot be a cost code of THIS input, or ``None`` (P19/B3).
**The guard is what makes this a rule and not a preference.** MEASURED over P18's round 2: two
ordinary words from a standard's prose — each in 4 documents of its base (of 270 and of
1 133), words of the kind ``nødstrømsaggregat`` or ``redundant kjøling`` — passed the whole
ordinary words from a standard's prose — each in 4 documents of its base (of a few hundred
and of about a thousand), words of the kind ``nødstrømsaggregat`` or ``redundant kjøling`` — passed the whole
gate to ``validated`` as ``affected_item`` codes. Both are GROUNDED: they appear verbatim in the
input, which is all P7 asks. What they are not is an identifier of a cost line.
@ -537,7 +537,7 @@ def _reference_refusal(grounding: Grounding, code: str) -> str | None:
* ``10.4`` (a requirements base, serverrom-04, ``validated``) is declared NOWHERE in that
base's frontmatter. The base
declares ``seksjon: 10.4.1`` … ``10.4.4`` and ``req_number: Krav 10.4.3—2``; the bare ``10.4``
is a section PREFIX that occurs in 12 of 274 documents and is no document's own number;
is a section PREFIX that occurs in 12 of a few hundred documents and is no document's own number;
* ``1.10.4`` (the process catalogue, an across-bases run, ``validated``) is not one of the
catalogue's declared ``prosessnr`` or ``seksjon`` values. It occurs in ONE of the catalogue's
documents, as prose: a cross-reference to chapter 1.10.4 of a requirements standard.

View file

@ -47,7 +47,7 @@ What each arm pins, and what it refuses:
none is byte-identical to before. That second half is what keeps ``demo-transcript.stdout``
unchanged, and it is asserted here rather than left to the golden;
(h) A4 — the judge counts a hit against THIS APPROACH'S fasit concepts, never against the base. A
judge matching the whole base would mark every declaration a hit on a 2 756-document corpus,
judge matching the whole base would mark every declaration a hit on a corpus of a few thousand documents,
which is the P16 vacuity ``citation_scope`` already exists to refuse;
(i) the declaration reaches ``{run_id}-debate.json``, so a paid run's evidence survives the process
that produced it.
@ -260,7 +260,7 @@ def test_the_binding_requirement_reaches_the_proposer_verbatim() -> None:
def test_a_hit_is_counted_against_this_approachs_fasit_never_the_base() -> None:
"""(h) The P16 vacuity, one column over: on a 2 756-document base everything is 'in the base'."""
"""(h) The P16 vacuity, one column over: on a base of a few thousand documents everything is 'in the base'."""
wanted = {"krav/12-1/a.md", "krav/12-12/b.md"}
approach = Approach(
id="a1",
@ -392,7 +392,7 @@ def test_a_declaration_in_the_base_but_not_in_the_fasit_is_not_a_hit(tmp_path: P
own ``wanted`` set — left the WHOLE suite green (1710/5), because the arm above drives
``_attributable`` while the hit is computed at the call site in ``score_context_set``. The
declaration here names a document the base really does hold; what it is not is the one the
fasit asks for, which is the only distinction a 2 756-document corpus leaves standing.
fasit asks for, which is the only distinction a corpus of a few thousand documents leaves standing.
"""
base = _minibase(tmp_path)
ctx = _context_dir(tmp_path)

View file

@ -75,11 +75,11 @@ def own_frontmatter(path: Path) -> dict[str, str]:
title: D200:2027
and the indented ``title`` used to replace the concept's own. MEASURED on a delivered
requirements corpus during development, before the fix: ``okf.navigate_bundle`` yielded 270
concept files carrying **1 distinct title** (the sources title, 270 times).
requirements corpus during development, before the fix: ``okf.navigate_bundle`` yielded a few
hundred concept files carrying **1 distinct title** (the sources title, every time).
``okf.parse_frontmatter`` now makes indentation load-bearing — a top-level (unindented) key
always wins over a nested one of the same name — and re-measured AFTER the fix, the same base's
269 requirement documents carried **269 distinct titles**.
requirement documents carried **one distinct title each**.
**This helper still isn't a plain call to ``okf.parse_frontmatter``, and that remains
measured rather than assumed:** ``own_frontmatter`` also strips one layer of enclosing

View file

@ -137,7 +137,7 @@ def test_the_worked_example_round_trips_through_the_real_readers(tmp_path: Path)
# The example declares `state: unreadable, reason: block-sequence, items_seen: 2` for a
# SPEC §5.1 block sequence; po widened its reader to that carrier because all four
# knowledge bases delivered at the time wrote it and nothing else (measured 2026-09-12:
# 446/446, 1133/1133, 270/270, 2756/2756 = 4605/4605, 0 in flow form). `shared/` is a
# every concept file of every base, 0 in flow form). `shared/` is a
# PULL-ONLY subtree, so the declaration cannot be corrected from here: closing this needs a
# commons amendment, and the divergence is asserted rather than skipped so it cannot sit
# unnoticed until someone reads the prose.

View file

@ -3,7 +3,7 @@
P7 made stage 0b ``item.code in grounding``: plain containment over ONE concatenated string. P16
then ran it against a delivered corpus and measured what containment cannot tell apart. The
falsification arm ``a4-indeksregulering`` proposed a 250 000 NOK saving on a single cost line whose
code was the knowledge base's OWN NAME — the catalogue designation every one of its 2 756 concept
code was the knowledge base's OWN NAME — the catalogue designation every one of its concept
documents carries — and the whole gate said ``validated``: stage 0 was skipped (un-anchored run),
stage 0b was satisfied by the letterhead, and the checker approved.
@ -18,14 +18,14 @@ measurement (one of three is 33 % and says nothing).
shortest real identifier is FOUR characters (``12.1``, ``52.1``), so ``N = 3`` sits one below the
measurement and cannot refuse anything measured;
* document frequency of every code-shaped token (``generate._IDENTIFIER_FORMS``) in each base:
1 692 distinct tokens and NOT ONE reaches 5 % of its base's documents. Highest anywhere 6 of 446
(1.35 %); highest that a fasit names 3 of 446 (0.67 %); the base's own name 2 756 of 2 756
(100 %). ``A =
1 692 distinct tokens and NOT ONE reaches 5 % of its base's documents. Highest anywhere 1.35 %
(6 documents); highest that a fasit names 0.67 % (3 documents); the base's own name is in every
document (100 %). ``A =
0.05`` therefore sits 3.7x above the highest real token and 20x below the defect.
**The denominator is NAMED in the refusal**, because Step 5 feeds that reason verbatim into the
next attempt's prompt: a proposer told only "ungrounded" answers with another token of the same
kind, while one told "it is in 2 756 of 2 756 documents" has been told what is wrong with it.
kind, while one told "it is in N of N documents" has been told what is wrong with it.
The arms that need a knowledge base read the package's pinned example bases (and SKIP, with the
store named, only when a user's own store lacks them); the rule's own algebra, the floor, and the

View file

@ -1,8 +1,8 @@
"""P19 DEL B — a "cost code" must have a FORM when the input offers forms.
**The measured defect.** P18's round 2 ended with two ``validated`` proposals whose
``affected_item`` codes were ordinary words out of a standard's prose, in 4 of the 270 and 4 of
the 1 133 documents of the two bases they came from. This file names them by two stand-ins of the
``affected_item`` codes were ordinary words out of a standard's prose, in 4 of a few hundred and 4
of about a thousand documents of the two bases they came from. This file names them by two stand-ins of the
same kind, ``nødstrømsaggregat`` and ``redundant kjøling``.
Both are GROUNDED in P7's sense — they appear verbatim in the input, which is all that stage asks —
and neither is INERT in P18/B1's sense, because neither is anywhere near the 5 % document share.

View file

@ -8,7 +8,7 @@ outboxes and were this arm's known positives:
* ``1.10.4`` — the multi-base pass, on the process catalogue, ``validated``.
Both were GROUNDED in P7's sense (they occur verbatim in the input) and neither was INERT in
P18/B1's sense (``10.4`` in 12 of 274 documents, ``1.10.4`` in 1 of 2 756). Stage 0 never ran:
P18/B1's sense (``10.4`` in 12 of a few hundred documents, ``1.10.4`` in 1 of a few thousand). Stage 0 never ran:
no requirements base ships a cost baseline. Nothing in the gate could say what they are.
**THE ORDER'S OWN RULE WAS FELLED BY MEASUREMENT BEFORE ANYTHING WAS BUILT ON IT.** B1 reads: a

View file

@ -221,7 +221,7 @@ def _ir(aid: str, claimed: float, code: str | None = None) -> dict[str, Any]:
).model_dump()
#: The run-wide citation list every proposal carried in all four archived runs. 270 there, 9
#: The run-wide citation list every proposal carried in all four archived runs. Hundreds there, 9
#: here — the number is not the point, the IDENTITY across proposals is.
_SHARED_CITED = 9
@ -233,7 +233,7 @@ def _snippet(aid: str, k: int) -> str:
def _stamp(decision: str, aid: str, *, shared: bool = False) -> ProvenanceStamp:
"""The proposal's own provenance, with a citation list that is ITS OWN.
Measured 19.09 on all four archived runs: every proposal in a run carried the SAME 270
Measured 19.09 on all four archived runs: every proposal in a run carried the SAME
citations, byte for byte — the run's whole retrieved context, stamped once per proposal.
That is a property of the outbox, not of the report, and the report now states it once
instead of repeating it. A fixture that reproduced it everywhere could only witness the
@ -956,7 +956,7 @@ def test_the_report_shows_each_proposals_source_and_how_many_places_it_cited(
A proposal without its source cannot be checked against the knowledge base at all, and a
single quote without the count cannot tell a proposal grounded in one place from one that
swept 270. The counts in ``_CITED`` are DISTINCT per approach on purpose: a builder printing a
swept hundreds. The counts in ``_CITED`` are DISTINCT per approach on purpose: a builder printing a
constant would satisfy a fixture where every count was the same."""
text = _report(tmp_path)
assert len({_CITED[aid] for aid in _EVALUATED}) == len(_EVALUATED), "the counts must differ"
@ -986,7 +986,7 @@ def test_the_report_shows_the_cost_lines_each_proposal_touches(tmp_path: Path) -
def test_one_citation_list_shared_by_every_proposal_is_stated_once(tmp_path: Path) -> None:
"""Measured 19.09 on all four archived runs: every proposal carried the SAME citation list,
byte for byte (270 places, same order) — the run's whole retrieved context, stamped once per
byte for byte (hundreds of places, same order) — the run's whole retrieved context, stamped once per
proposal. The cause is in the OUTBOX, not in the builder reading a wrong field, so the report
cannot make the quote informative. What it can do is stop repeating it: say it once, say that
it is the run's list and not the measure's, and drop the per-proposal copies."""

View file

@ -14,8 +14,8 @@ tests`` hit only unrelated files).
**THE ORDER'S (a) WAS VACUOUS AS WRITTEN, AND THAT IS MEASURED.** The order defines grounded as
"a must_cite path was OPENED *or* CITED". But on the S2c navigation path ``run_project`` stamps
``citations = bundle_citations(bundle)``, which is ONE CITATION PER CONTEXT FILE - the whole corpus.
Measured on a requirements base during development: 446 context files, 446 citations, and **6
of 6 fasit paths already "cited" before a single model call**. A judge honouring the order
Measured on a requirements base during development: a few hundred context files, as many
citations, and **6 of 6 fasit paths already "cited" before a single model call**. A judge honouring the order
literally would be a gate that can only be green, which is the repo's own vacuous-gate class,
inside the gate built to stop it. So
``grounded`` counts a CITATION only when the citation list is NARROWER than the base (a declared
@ -23,8 +23,8 @@ pre-pass cut); a whole-base list is reported as such and carries nothing. Both h
either way, so the operator can read which one fired - the deviation is stated, never silent.
**(b') was checked for the same vacuity and is CLEAN.** ``bundle_citations`` snippets are concept
BODIES, and the ``ref``/``title`` live in FRONTMATTER: measured 0 of 446 bodies of that base
contain its ``Krav 4.1.2-1``. So the snippet arm can carry (b') without being satisfied by
BODIES, and the ``ref``/``title`` live in FRONTMATTER: measured, not one body of that base
contains its ``Krav 4.1.2-1``. So the snippet arm can carry (b') without being satisfied by
construction, and the order's definition is kept. Which half fired is still reported.
**A denominator, always** (Verifiseringsloven ansikt 4): every verdict names how many tool calls,
@ -273,7 +273,7 @@ def test_a_opening_some_other_document_does_not_ground_it(tmp_path: Path) -> Non
def test_b_a_whole_base_citation_list_cannot_ground_an_approach(tmp_path: Path) -> None:
"""Measured on a requirements base: 446 context files, 446 citations, 6/6 fasit paths
"""Measured on a requirements base: a few hundred context files, as many citations, 6/6 fasit paths
'cited' before any model call. Honouring the order literally would make (a) green by
construction."""
verdict = _judge(tmp_path, tool_calls=[])

View file

@ -5,7 +5,7 @@
documents read lay OUTSIDE the default window, so the window had been widened, and the trace could
not say with which knob (stress round 2, finding 1; the ledger is ``docs/invarianter.md``). The
recorder kept ``name`` / ``bundle_id`` / ``path`` and nothing else. "Did the model narrow the level,
or page through it?" is the operative question about a corpus of 2 756 documents, and it was
or page through it?" is the operative question about a corpus of a few thousand documents, and it was
unanswerable from the artefact the run leaves behind.
**DEL D, and one of the two findings it fixes was WRONG AS WRITTEN.** P18's finding 4 said a