What the default was and is, the three signatures, the naming choice against
the module's other eight constants, the CLI-flag decision with its zero, both
gates with this round's numbers beside round 23's, the consumer list with its
denominator, and the honesty limits.
The deviation is stated first and it is the order's own acceptance row: `S1
spent 28 020 B at the default k` is round 23's X column -- the flagged bundle
with the line SCORED -- and the new default is round 23's Y, whose published
value for that cell is 31 031. Measured here in one process on one bundle:
`link_in_signal=True` gives 28 020 and the default gives 31 031, the unflagged
build gives 31 031 too. 16 of 16 cells of the Y column reproduce to the byte,
so neither failure the stop-rule guards against is present; the row was
transcribed from the column being retired.
Also corrects one label in round 23's own file, which is otherwise untouched
because a report is a measurement with a date. The row reading "distinct tokens
the path ever contributed" carried 31/22/8, which is the QUESTION-token count:
the path's segments are `r761` and `prosesskoden`, and `prosess` is a question
token that reaches the second through the stem prefix rule -- measured,
`consume.tokens_match('prosess', 'prosesskoden')` is True at
`MIN_SHARED_PREFIX = 4`. The label is what was wrong, and correcting it is what
makes the table agree with the paragraph under it, which already explains
`prosess` that way. The numbers stand, and no CHANGELOG figure is touched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
16 KiB
K3 round 23: the path in the body signal
Date: 2026-09-12 · Base: 2d4f56d · Commits: 18b3903 (red),
8e82da4 (the instrument) · Measured from: a frozen git archive export of
8e82da4 installed into a scratch virtualenv (__file__ under
/private/tmp/..., never /Users/ktg/repos, never the okf on PATH), Python
3.14, guard 1.4.0. One document: R761 Prosesskoden:2025, built twice in scratch
from the publisher's own NISO-STS source, once with --shell-parent and once
without. The consumer repository was read only: git status --porcelain empty
before and after, build/ferdig/ listing identical including mtimes.
Round 21 gave 675 of 710 heading-only sections a body line -- Enclosing section: [<title>](/<bundle-absolute path>) -- and reported that hit@k did not
move while the delivered SET did, on 2 of 8 questions at the default k and 3
of 8 at k 50. It attributed the cost only where a newcomer matched through
the link ALONE, and said so: "the rest of the delivered-set movement, and any
effect of the heavier excerpts on the knapsack, is not decomposed." This round
decomposes both.
0. Where this measurement differs from what it was given, first
- The link cost is 71 616 B = 4.45 %, not 72 265 B = 4.49 %. The order carried both figures and asked which one a fresh measurement reproduces: it reproduces the dispatch note's, not round 21's. Measured here as the byte difference between each concept's body and the same body with the door's line removed, over the 2 761 concepts of the flagged build: 71 616 B of 1 607 855 B body bytes, the line itself 70 941 B, median line 101 B, max 245 B, min 60 B, and 68.3 % of the 103 835 B those 675 bodies hold. Round 21's median and max are each exactly 2 B above these, which is what a per-line convention counting the newline and the blank line would give; that convention totals 72 291 B, still not its published 72 265 B. The rule used here is stated so the next round can disagree with a rule rather than with a number. Round 21's figure is left standing in its own file -- a report is a measurement with a date.
- Everything else round 21 published reproduces exactly. 675 of 710 shells
carry exactly one link, 0 without; the flagged and unflagged builds differ in
1 350 of 5 522 files (675 concepts + 675 index files, and
log.mdidentical here); hit@1/8/50 6/6 · 6/6 · 6/6 at bothkwith the known-positive at rank 1; delivered sets move on 2 of 8 questions at the defaultkand 3 of 8 atk50; S1'sspentat the defaultkis 28 020 B, to the byte. - No default moved. The instrument is a function parameter with no CLI flag, defaulting to today's behaviour.
1. The rig: one bundle, three readings
| reading | bundle | signal | excerpt bytes |
|---|---|---|---|
| X | flagged | link line scored | with the link |
| Y | flagged | link line NOT scored (link_in_signal=False) |
with the link |
| Z | unflagged | -- | without the link |
| W | flagged | link line scored | without the link (scratch rig only) |
X vs Y isolates RANKING (same bytes, same bundle). X vs W isolates the
BUDGET (same ranking, lighter excerpts). Y vs Z is the control that says the
instrument is honest, and it holds on 16 of 16 rows: Y's delivered list,
its order and its spent are identical to the unflagged build's, byte for
byte. The separation is therefore measured, not assumed.
W is the one configuration that does not exist in the library: it patches
delivered_text in the measuring script alone. Nothing in src/ knows about
it.
2. The base row, reproduced before anything else
| reading | hit@1 | hit@8 | hit@50 | KP rank, k 8 |
KP rank, k 50 |
denominator |
|---|---|---|---|---|---|---|
| X | 6/6 | 6/6 | 6/6 | 1 | 1 | 6 questions, 2 761 concepts |
| Y | 6/6 | 6/6 | 6/6 | 1 | 1 | 6 |
| Z | 6/6 | 6/6 | 6/6 | 1 | 1 | 6 |
The instrument moves no hit@k cell and no known-positive rank. That was the condition for reading anything else it produces.
3. The decomposition, per question, per k, per reading
pos counts positions where X and Y differ; new/out are set differences;
budget is X vs W, the displacement the ranking cannot explain.
Default k (8)
| id | delivered X / Y | spent X / Y | pos | new | out | gained a token | via PATH | via TITLE | budget |
|---|---|---|---|---|---|---|---|---|---|
| S1 | 7 / 7 | 28 020 / 31 031 | 3 of 7 | 3 | 3 | 3 of 3 | 3 | 0 | 0 |
| S2 | 8 / 8 | 14 949 / 14 949 | 0 | 0 | 0 | -- | 0 | 0 | 0 |
| S3 | 7 / 7 | 23 811 / 23 811 | 0 | 0 | 0 | -- | 0 | 0 | 0 |
| S4 | 8 / 8 | 54 025 / 54 025 | 0 | 0 | 0 | -- | 0 | 0 | 0 |
| S5 | 8 / 8 | 14 342 / 14 342 | 0 | 0 | 0 | -- | 0 | 0 | 0 |
| S6 | 8 / 8 | 24 424 / 24 424 | 0 | 0 | 0 | -- | 0 | 0 | 0 |
| KP | 7 / 7 | 35 050 / 35 050 | 0 | 0 | 0 | -- | 0 | 0 | 0 |
| KN | 7 / 7 | 10 514 / 10 151 | 5 of 7 | 2 | 2 | 2 of 2 | 2 | 0 | 0 |
S1, both lists in full (the six identical questions are identical in order as well as in membership):
| # | X | Y |
|---|---|---|
| 1 | 2-1/hovedprosesser |
2-1/hovedprosesser |
| 2 | hovedprosess-81-l-smasser |
hovedprosess-81-l-smasser |
| 3 | hovedprosess-83-konstruksjoner-i-grunnen-... |
same |
| 4 | hovedprosess-84-betong |
hovedprosess-84-betong |
| 5 | 32-113/delt-tverrsnitt-normal-salvelengde |
5/hierarkisk-oppbygging-av-prosesser |
| 6 | 32-114/delt-tverrsnitt-halv-salvelengde |
hovedprosess-82-berg |
| 7 | 36-111/hovedfordelinger |
hovedprosess-85-st-l |
KN, both lists in full:
| # | X | Y |
|---|---|---|
| 1 | 25-41/jordmasser-til-st-yvoll-... |
same |
| 2 | 1/bruksomr-der-for-prosesskoden |
same |
| 3 | 25-4/jordmasser-til-st-yvoll-ledevoll-steinfyllingsskr-ninger-mm |
26-4/sprengt-stein-... |
| 4 | 26-4/sprengt-stein-... |
32-225/steinmasser-fra-tunnelmunning-... |
| 5 | 31-51/injeksjons-og-kontrollhull-ved-sporadisk-injeksjon |
5/hierarkisk-oppbygging-av-prosesser |
| 6 | 32-225/steinmasser-... |
67-5/ledelinjer-i-gategrunn |
| 7 | 5/hierarkisk-oppbygging-av-prosesser |
88-1714/sporslitasje |
k 50
| id | delivered X / Y | spent X / Y | pos | new | out | gained a token | via PATH only | via BOTH | via TITLE only | budget |
|---|---|---|---|---|---|---|---|---|---|---|
| S1 | 41 / 40 | 107 803 / 106 610 | 38 of 41 | 7 | 6 | 6 of 7 | 6 | 0 | 0 | 0 |
| S2 | 43 / 43 | 109 618 / 109 618 | 0 | 0 | 0 | -- | 0 | 0 | 0 | 0 |
| S3 | 42 / 42 | 108 228 / 108 228 | 0 | 0 | 0 | -- | 0 | 0 | 0 | 0 |
| S4 | 39 / 39 | 108 158 / 108 158 | 0 | 0 | 0 | -- | 0 | 0 | 0 | 0 |
| S5 | 49 / 49 | 102 855 / 102 855 | 0 | 0 | 0 | -- | 0 | 0 | 0 | 0 |
| S6 | 48 / 48 | 102 211 / 102 211 | 0 | 0 | 0 | -- | 0 | 0 | 0 | 0 |
| KP | 48 / 44 | 96 965 / 108 749 | 46 of 48 | 23 | 19 | 22 of 23 | 20 | 2 | 0 | 0 |
| KN | 43 / 43 | 106 992 / 104 954 | 41 of 43 | 8 | 8 | 6 of 8 | 6 | 0 | 0 | 1 |
Where the newcomers enter, and what they push out. On KP at k 50, 22 of
the 23 newcomers are linked shells entering at positions 21, 22, 23, 24, 25,
26, 27, 28, 31, 32, 33, 34, 35, 38, 39, 40, 41, 42, 45, 46, 47, 48, and the 19
that leave held Y's positions 26 to 44 -- among them 84-2/forskaling,
84-3/armering, 87-1/fuktisolering-membran-... and
88-2/vedlikehold-beskyttelse-og-reparasjon-av-betong. On S1 at k 50 six
shells enter at positions 4, 5, 6, 8, 9, 10 -- near the top -- and six real
sections leave from Y's positions 35 to 40. The four newcomers carrying no
link of their own (1 on S1, 1 on KP, 2 on KN) gained nothing: they moved
because the concepts around them did.
The one number that decides everything below
| row | result | denominator |
|---|---|---|
| newcomers that gained a question token from the link | 39 | 39 link-bearing newcomers |
| of those, the gain came from the PATH | 37 path only + 2 path and title | 39 |
| of those, the gain came from the TITLE alone | 0 | 39 |
| distinct QUESTION tokens the path ever matched | prosesskoden (31), r761 (22), prosess (8) |
61 token hits |
Every token the link line ever added is a segment of the document's own
directory -- r761-prosesskoden -- and prosess reaches it by the stem
prefix rule. This is exactly the saturation shared_id_prefix (round 20) took
OUT of the id signal, arriving back through the body. The link's TITLE, which
is the part carrying meaning, contributed a hit on its own 0 times.
Round 21's hypothesis is therefore confirmed and sharpened: it is not the link that costs rank, it is the bundle-absolute PATH inside it. Only that second statement points at a fix.
The knapsack, which round 21 did not decompose
| row | result | denominator |
|---|---|---|
| rows where X and W deliver a different SET | 1 | 16 |
| the concept displaced | 12-11/tilrigging, KN at k 50: 43 delivered with the link bytes, 44 without |
1 |
rows where the budget binds at the default k |
0 (max spent 54 025 of 120 000) |
8 |
Rank movement and budget displacement are different sizes. At the default
k the budget is not binding at all, so 100 % of the movement there is
ranking. At k 50 the budget binds on every question, and the heavier excerpts
still displace one concept on one question -- the one with no fasit.
A single figure mixing the two would have read as "the link moves 5 of 8 rows";
it moves 5 by rank and 1 by weight, and the 1 is not on a scored row.
4. Can --shell-parent be on? No -- and the third exit is now measured
The acceptance the order set, answered with the numbers beside it:
| condition | result | verdict |
|---|---|---|
hit@1/8/50 and KP rank unchanged, both k |
6/6 · 6/6 · 6/6, KP 1 / 1, on all three readings | met |
| newcomers matching through the path = 0, or a stated number | 39 of 39 link-bearing newcomers gained through the path; 0 through the title | not met |
| delivered sets moved, per question | k 8: S1 3 of 7 positions, KN 5 of 7, six questions 0 · k 50: S1 38 of 41, KP 46 of 48, KN 41 of 43, five questions 0 |
stated, and it is movement |
--shell-parent stays OFF at its current link form. Two of three
conditions fail, and they fail for one reason with a name.
The third exit, measured with the same numbers. If the cost is the path, the question is no longer on-or-off but which of these:
| option | what it costs | what the numbers say |
|---|---|---|
| (a) leave the default off | the pointer round 21 built reaches no reader on any shipped bundle | 0 of 5 shipped bundles carry the line today, so this is the status quo |
| (b) change the link's FORM (relative, or title-only) | a file change: SPEC SS 6.1 calls the absolute form recommended, and inbox._link_enclosing's docstring gives a second reason (a relative link would count .. across a layout the next round may change) |
not measured here -- it needs a new build and a new form to measure |
(c) make link_in_signal=False the DEFAULT reading in consume |
a ranking change on a published payload form | measured: with (c), turning --shell-parent on moves nothing. Y equals Z on 16 of 16 rows -- list, order and spent -- so under (c) the flagged bundle delivers exactly what the unflagged one delivers |
Recommendation: (c), and (c) makes (a) unnecessary. The file keeps SS 6.1's
recommended form, the reader keeps the line in the excerpt, the checker keeps
parent_unfollowable, and the signal stops counting a path that says only which
document the concept was already known to be in.
The exposure of (c) is measured on bytes, not argued. body_without_link_line
is a no-op on any body that does not end in the door's exact form, and the door
writes that form only under --shell-parent:
| bundle | payload byte-identical under (c) | files carrying the door's line |
|---|---|---|
| N100 | yes | 0 |
| N200 | yes | 0 |
| N500 | yes | 0 |
| R761 as shipped | yes | 0 |
| R761 unflagged, built here | yes | 0 |
5 of 5, 0 of 5. Changing consume's default reading of the body is a rank
change on a published payload form, and it is stated here as one: it requires
the whole decomposition above behind it, which is what this report is. It moves
no byte of any bundle that exists today, and the day a bundle carries the line
is the day it would have started costing rank instead.
--shell-parent is a separate decision from (c) and is not taken here. With
(c) in place its acceptance would read: hit@k unchanged (already 6/6 on Y),
newcomers through the path 0, delivered sets moved 0 of 8 at both k.
All three hold on this document. What does not follow from one document is the
default.
5. What --follow-parent still lacks, and why it is not built here
DEFAULT_FOLLOW_PARENT = False because no fasit has a shell as its answer.
Measured here rather than quoted:
| row | result | denominator |
|---|---|---|
| questions whose fasit section is a heading-only concept | 0 | 7 with a fasit (8 questions, KN has none) |
| fasit sections present in the bundle at all | 7 | 7 |
| heading-only concepts in the document | 710 | 2 761 |
| of those, with an ancestor holding text (a parent to follow) | 675 | 710 |
| of those, with no such ancestor (nothing to inherit) | 35 | 710 |
What is missing is a question class, not a feature. Generically: a question whose answer is a section that STATES nothing itself and inherits everything from the section enclosing it -- so the correct answer can only be given by a reader who has the ancestor's text. Such a question can be asked of the 675 shells that have an ancestor holding text; it cannot be asked of the 35 without one, because there the inheritance does not exist, and asking it of the 2 051 sections holding their own text would not test the flag at all.
The denominator would be the number of such questions, and the acceptance would
have to separate two things the current instrument cannot: whether the shell is
DELIVERED (which --follow-parent does not change -- it delivers the same set
by construction and measured), and whether the answer is CORRECT, which needs a
judged reading and not a title match. Round 21 measured one probe question it
chose itself and said so.
The fasit is not written here. Choosing which sections become questions, and what counts as a correct answer for a section that states nothing, is the operator's decision; building it in a measurement session would make it cheap and would take the decision by making it.
6. Conformance
okf check on 48 of 48 payloads (X, Y and Z, eight questions, two k):
rc 0, 17 rules, 0 findings, the rule count read as a literal from
len(contract_check.RULES). The skill for each reading was generated from its
own bundle, so bundle_mismatch compared the identity it was meant to.
Conformance is the floor and never the proof: the known-negative question's
payloads are conformant too, and they answer nothing.
Honesty limits
- N = 1 document. Everything here is one 2 761-concept standard from one publisher. The mechanism -- a bundle-absolute path repeating the document directory in every linked body -- is a property of the FORM and would appear in any bundle, but its size depends on whether a question happens to name the document. Three of eight questions here do.
- The consumption half is not measured. This round measures delivery and rank only. Whether a reader ANSWERS better is a judged reading; round 21's own consumption rows were one non-deterministic draw per question.
- The instrument is someone else's and scores a title or a section-number
pair, not an answer.
hitk_sk2.pyat the consumer's HEADee4d7e1, copied to scratch with the hard-coded payload path changed, because a concurrent session writes the same/tmpfile. - Six of eight questions never move at all, which means the whole measurement rests on three rows (S1, KP, KN) -- and KN has no fasit, so the scored evidence is two.
- The four goldens and the K2 pin did not move, which is what says the instrument changed no default: the suite is 1 816 passed / 1 skipped against a baseline of 1 807 / 1, the nine new ones being this round's.
- Option (b) is unmeasured. It is listed because it is a real alternative, not because it was compared; a relative or title-only link needs its own build and its own row before anyone prefers it to (c).