docs: general wording for the example-base document counts
Replace the exact document counts of earlier example bases (and the per-base counts in the sources-format note) with general wording or N-of-N in prose, comments and docstrings. Percentages and numerators stay; no constant, assertion or test data changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
parent
0b7c853f3d
commit
6b1046bc23
16 changed files with 56 additions and 56 deletions
|
|
@ -47,7 +47,7 @@ What each arm pins, and what it refuses:
|
|||
none is byte-identical to before. That second half is what keeps ``demo-transcript.stdout``
|
||||
unchanged, and it is asserted here rather than left to the golden;
|
||||
(h) A4 — the judge counts a hit against THIS APPROACH'S fasit concepts, never against the base. A
|
||||
judge matching the whole base would mark every declaration a hit on a 2 756-document corpus,
|
||||
judge matching the whole base would mark every declaration a hit on a corpus of a few thousand documents,
|
||||
which is the P16 vacuity ``citation_scope`` already exists to refuse;
|
||||
(i) the declaration reaches ``{run_id}-debate.json``, so a paid run's evidence survives the process
|
||||
that produced it.
|
||||
|
|
@ -260,7 +260,7 @@ def test_the_binding_requirement_reaches_the_proposer_verbatim() -> None:
|
|||
|
||||
|
||||
def test_a_hit_is_counted_against_this_approachs_fasit_never_the_base() -> None:
|
||||
"""(h) The P16 vacuity, one column over: on a 2 756-document base everything is 'in the base'."""
|
||||
"""(h) The P16 vacuity, one column over: on a base of a few thousand documents everything is 'in the base'."""
|
||||
wanted = {"krav/12-1/a.md", "krav/12-12/b.md"}
|
||||
approach = Approach(
|
||||
id="a1",
|
||||
|
|
@ -392,7 +392,7 @@ def test_a_declaration_in_the_base_but_not_in_the_fasit_is_not_a_hit(tmp_path: P
|
|||
own ``wanted`` set — left the WHOLE suite green (1710/5), because the arm above drives
|
||||
``_attributable`` while the hit is computed at the call site in ``score_context_set``. The
|
||||
declaration here names a document the base really does hold; what it is not is the one the
|
||||
fasit asks for, which is the only distinction a 2 756-document corpus leaves standing.
|
||||
fasit asks for, which is the only distinction a corpus of a few thousand documents leaves standing.
|
||||
"""
|
||||
base = _minibase(tmp_path)
|
||||
ctx = _context_dir(tmp_path)
|
||||
|
|
|
|||
|
|
@ -75,11 +75,11 @@ def own_frontmatter(path: Path) -> dict[str, str]:
|
|||
title: D200:2027
|
||||
|
||||
and the indented ``title`` used to replace the concept's own. MEASURED on a delivered
|
||||
requirements corpus during development, before the fix: ``okf.navigate_bundle`` yielded 270
|
||||
concept files carrying **1 distinct title** (the sources title, 270 times).
|
||||
requirements corpus during development, before the fix: ``okf.navigate_bundle`` yielded a few
|
||||
hundred concept files carrying **1 distinct title** (the sources title, every time).
|
||||
``okf.parse_frontmatter`` now makes indentation load-bearing — a top-level (unindented) key
|
||||
always wins over a nested one of the same name — and re-measured AFTER the fix, the same base's
|
||||
269 requirement documents carried **269 distinct titles**.
|
||||
requirement documents carried **one distinct title each**.
|
||||
|
||||
**This helper still isn't a plain call to ``okf.parse_frontmatter``, and that remains
|
||||
measured rather than assumed:** ``own_frontmatter`` also strips one layer of enclosing
|
||||
|
|
|
|||
|
|
@ -137,7 +137,7 @@ def test_the_worked_example_round_trips_through_the_real_readers(tmp_path: Path)
|
|||
# The example declares `state: unreadable, reason: block-sequence, items_seen: 2` for a
|
||||
# SPEC §5.1 block sequence; po widened its reader to that carrier because all four
|
||||
# knowledge bases delivered at the time wrote it and nothing else (measured 2026-09-12:
|
||||
# 446/446, 1133/1133, 270/270, 2756/2756 = 4605/4605, 0 in flow form). `shared/` is a
|
||||
# every concept file of every base, 0 in flow form). `shared/` is a
|
||||
# PULL-ONLY subtree, so the declaration cannot be corrected from here: closing this needs a
|
||||
# commons amendment, and the divergence is asserted rather than skipped so it cannot sit
|
||||
# unnoticed until someone reads the prose.
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
P7 made stage 0b ``item.code in grounding``: plain containment over ONE concatenated string. P16
|
||||
then ran it against a delivered corpus and measured what containment cannot tell apart. The
|
||||
falsification arm ``a4-indeksregulering`` proposed a 250 000 NOK saving on a single cost line whose
|
||||
code was the knowledge base's OWN NAME — the catalogue designation every one of its 2 756 concept
|
||||
code was the knowledge base's OWN NAME — the catalogue designation every one of its concept
|
||||
documents carries — and the whole gate said ``validated``: stage 0 was skipped (un-anchored run),
|
||||
stage 0b was satisfied by the letterhead, and the checker approved.
|
||||
|
||||
|
|
@ -18,14 +18,14 @@ measurement (one of three is 33 % and says nothing).
|
|||
shortest real identifier is FOUR characters (``12.1``, ``52.1``), so ``N = 3`` sits one below the
|
||||
measurement and cannot refuse anything measured;
|
||||
* document frequency of every code-shaped token (``generate._IDENTIFIER_FORMS``) in each base:
|
||||
1 692 distinct tokens and NOT ONE reaches 5 % of its base's documents. Highest anywhere 6 of 446
|
||||
(1.35 %); highest that a fasit names 3 of 446 (0.67 %); the base's own name 2 756 of 2 756
|
||||
(100 %). ``A =
|
||||
1 692 distinct tokens and NOT ONE reaches 5 % of its base's documents. Highest anywhere 1.35 %
|
||||
(6 documents); highest that a fasit names 0.67 % (3 documents); the base's own name is in every
|
||||
document (100 %). ``A =
|
||||
0.05`` therefore sits 3.7x above the highest real token and 20x below the defect.
|
||||
|
||||
**The denominator is NAMED in the refusal**, because Step 5 feeds that reason verbatim into the
|
||||
next attempt's prompt: a proposer told only "ungrounded" answers with another token of the same
|
||||
kind, while one told "it is in 2 756 of 2 756 documents" has been told what is wrong with it.
|
||||
kind, while one told "it is in N of N documents" has been told what is wrong with it.
|
||||
|
||||
The arms that need a knowledge base read the package's pinned example bases (and SKIP, with the
|
||||
store named, only when a user's own store lacks them); the rule's own algebra, the floor, and the
|
||||
|
|
|
|||
|
|
@ -1,8 +1,8 @@
|
|||
"""P19 DEL B — a "cost code" must have a FORM when the input offers forms.
|
||||
|
||||
**The measured defect.** P18's round 2 ended with two ``validated`` proposals whose
|
||||
``affected_item`` codes were ordinary words out of a standard's prose, in 4 of the 270 and 4 of
|
||||
the 1 133 documents of the two bases they came from. This file names them by two stand-ins of the
|
||||
``affected_item`` codes were ordinary words out of a standard's prose, in 4 of a few hundred and 4
|
||||
of about a thousand documents of the two bases they came from. This file names them by two stand-ins of the
|
||||
same kind, ``nødstrømsaggregat`` and ``redundant kjøling``.
|
||||
Both are GROUNDED in P7's sense — they appear verbatim in the input, which is all that stage asks —
|
||||
and neither is INERT in P18/B1's sense, because neither is anywhere near the 5 % document share.
|
||||
|
|
|
|||
|
|
@ -8,7 +8,7 @@ outboxes and were this arm's known positives:
|
|||
* ``1.10.4`` — the multi-base pass, on the process catalogue, ``validated``.
|
||||
|
||||
Both were GROUNDED in P7's sense (they occur verbatim in the input) and neither was INERT in
|
||||
P18/B1's sense (``10.4`` in 12 of 274 documents, ``1.10.4`` in 1 of 2 756). Stage 0 never ran:
|
||||
P18/B1's sense (``10.4`` in 12 of a few hundred documents, ``1.10.4`` in 1 of a few thousand). Stage 0 never ran:
|
||||
no requirements base ships a cost baseline. Nothing in the gate could say what they are.
|
||||
|
||||
**THE ORDER'S OWN RULE WAS FELLED BY MEASUREMENT BEFORE ANYTHING WAS BUILT ON IT.** B1 reads: a
|
||||
|
|
|
|||
|
|
@ -221,7 +221,7 @@ def _ir(aid: str, claimed: float, code: str | None = None) -> dict[str, Any]:
|
|||
).model_dump()
|
||||
|
||||
|
||||
#: The run-wide citation list every proposal carried in all four archived runs. 270 there, 9
|
||||
#: The run-wide citation list every proposal carried in all four archived runs. Hundreds there, 9
|
||||
#: here — the number is not the point, the IDENTITY across proposals is.
|
||||
_SHARED_CITED = 9
|
||||
|
||||
|
|
@ -233,7 +233,7 @@ def _snippet(aid: str, k: int) -> str:
|
|||
def _stamp(decision: str, aid: str, *, shared: bool = False) -> ProvenanceStamp:
|
||||
"""The proposal's own provenance, with a citation list that is ITS OWN.
|
||||
|
||||
Measured 19.09 on all four archived runs: every proposal in a run carried the SAME 270
|
||||
Measured 19.09 on all four archived runs: every proposal in a run carried the SAME
|
||||
citations, byte for byte — the run's whole retrieved context, stamped once per proposal.
|
||||
That is a property of the outbox, not of the report, and the report now states it once
|
||||
instead of repeating it. A fixture that reproduced it everywhere could only witness the
|
||||
|
|
@ -956,7 +956,7 @@ def test_the_report_shows_each_proposals_source_and_how_many_places_it_cited(
|
|||
|
||||
A proposal without its source cannot be checked against the knowledge base at all, and a
|
||||
single quote without the count cannot tell a proposal grounded in one place from one that
|
||||
swept 270. The counts in ``_CITED`` are DISTINCT per approach on purpose: a builder printing a
|
||||
swept hundreds. The counts in ``_CITED`` are DISTINCT per approach on purpose: a builder printing a
|
||||
constant would satisfy a fixture where every count was the same."""
|
||||
text = _report(tmp_path)
|
||||
assert len({_CITED[aid] for aid in _EVALUATED}) == len(_EVALUATED), "the counts must differ"
|
||||
|
|
@ -986,7 +986,7 @@ def test_the_report_shows_the_cost_lines_each_proposal_touches(tmp_path: Path) -
|
|||
|
||||
def test_one_citation_list_shared_by_every_proposal_is_stated_once(tmp_path: Path) -> None:
|
||||
"""Measured 19.09 on all four archived runs: every proposal carried the SAME citation list,
|
||||
byte for byte (270 places, same order) — the run's whole retrieved context, stamped once per
|
||||
byte for byte (hundreds of places, same order) — the run's whole retrieved context, stamped once per
|
||||
proposal. The cause is in the OUTBOX, not in the builder reading a wrong field, so the report
|
||||
cannot make the quote informative. What it can do is stop repeating it: say it once, say that
|
||||
it is the run's list and not the measure's, and drop the per-proposal copies."""
|
||||
|
|
|
|||
|
|
@ -14,8 +14,8 @@ tests`` hit only unrelated files).
|
|||
**THE ORDER'S (a) WAS VACUOUS AS WRITTEN, AND THAT IS MEASURED.** The order defines grounded as
|
||||
"a must_cite path was OPENED *or* CITED". But on the S2c navigation path ``run_project`` stamps
|
||||
``citations = bundle_citations(bundle)``, which is ONE CITATION PER CONTEXT FILE - the whole corpus.
|
||||
Measured on a requirements base during development: 446 context files, 446 citations, and **6
|
||||
of 6 fasit paths already "cited" before a single model call**. A judge honouring the order
|
||||
Measured on a requirements base during development: a few hundred context files, as many
|
||||
citations, and **6 of 6 fasit paths already "cited" before a single model call**. A judge honouring the order
|
||||
literally would be a gate that can only be green, which is the repo's own vacuous-gate class,
|
||||
inside the gate built to stop it. So
|
||||
``grounded`` counts a CITATION only when the citation list is NARROWER than the base (a declared
|
||||
|
|
@ -23,8 +23,8 @@ pre-pass cut); a whole-base list is reported as such and carries nothing. Both h
|
|||
either way, so the operator can read which one fired - the deviation is stated, never silent.
|
||||
|
||||
**(b') was checked for the same vacuity and is CLEAN.** ``bundle_citations`` snippets are concept
|
||||
BODIES, and the ``ref``/``title`` live in FRONTMATTER: measured 0 of 446 bodies of that base
|
||||
contain its ``Krav 4.1.2-1``. So the snippet arm can carry (b') without being satisfied by
|
||||
BODIES, and the ``ref``/``title`` live in FRONTMATTER: measured, not one body of that base
|
||||
contains its ``Krav 4.1.2-1``. So the snippet arm can carry (b') without being satisfied by
|
||||
construction, and the order's definition is kept. Which half fired is still reported.
|
||||
|
||||
**A denominator, always** (Verifiseringsloven ansikt 4): every verdict names how many tool calls,
|
||||
|
|
@ -273,7 +273,7 @@ def test_a_opening_some_other_document_does_not_ground_it(tmp_path: Path) -> Non
|
|||
|
||||
|
||||
def test_b_a_whole_base_citation_list_cannot_ground_an_approach(tmp_path: Path) -> None:
|
||||
"""Measured on a requirements base: 446 context files, 446 citations, 6/6 fasit paths
|
||||
"""Measured on a requirements base: a few hundred context files, as many citations, 6/6 fasit paths
|
||||
'cited' before any model call. Honouring the order literally would make (a) green by
|
||||
construction."""
|
||||
verdict = _judge(tmp_path, tool_calls=[])
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@
|
|||
documents read lay OUTSIDE the default window, so the window had been widened, and the trace could
|
||||
not say with which knob (stress round 2, finding 1; the ledger is ``docs/invarianter.md``). The
|
||||
recorder kept ``name`` / ``bundle_id`` / ``path`` and nothing else. "Did the model narrow the level,
|
||||
or page through it?" is the operative question about a corpus of 2 756 documents, and it was
|
||||
or page through it?" is the operative question about a corpus of a few thousand documents, and it was
|
||||
unanswerable from the artefact the run leaves behind.
|
||||
|
||||
**DEL D, and one of the two findings it fixes was WRONG AS WRITTEN.** P18's finding 4 said a
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue