feat(propose,consume,profiles,importer): recovery yields to declaration, and 9 % of the corpus that was in no segment

One rule explains every remaining `pdf` miss on the twelve-position reference:
where a document DECLARES headings, Arm D's RECOVERED headings are the whole of
the excess, and every declared one is a unit the reference wants. `--outline-gate`
admits recovery only where the document declares none of its own, plus any one
recovered heading covering OUTLINE_SHARE (0.20) of the text. It is `fold_units`
clause 2's own principle moved from voting to admission, and it filters at
ADMISSION so the text a removed mark opened is carried by the mark above it --
the post-filter form scores identically on all twelve positions and loses that
text, which is why only one of them shipped.

`--outline-gate` and `--drop-wrapped-outline` become the package default, one
decision because neither carries the reference alone: `pdf` 2 of 8 -> 5 of 8
alone, 7 of 8 together; the sheet 5 of 12 -> 10 of 12; `docx` unchanged at 3 of
3. Each keeps an explicit opt-out. The bar the move had to clear was not the
reference: hit@8 on a K2 bundle built with it holds 5 of 6 at ranks 1,1,1,1,1,-,
no row losing rank 1. `--sheet-section-rows --keep-table-heading` reaches 11 of
12 and does NOT ship, because on a bundle built with it row 1 falls rank 1 -> 2.
Cost to a consumer is a re-run: 492 concepts / 944 files -> 425 / 810.

DOCUMENT_PRIOR_EXPONENT makes the document prior sublinear (total/n**0.5). A sum
measures size and a density is diluted by every unit carrying none of the
question, so a document split 1 -> 12 lost its prior by 12. Swept over five
values on 18 rows it is at least as good as the delivered density everywhere and
strictly better on three. Stated plainly: end to end it moved NOT ONE hit@8 row
on any of four bundles, so it did not solve the knot it was adopted for -- what
did is that the `pdf` gain never needed `--sheet-section-rows`.

`--first-span-from-zero` is off and repairs a measured loss found while chasing
one position's 940 characters: 32 of the 32 documents that get a plan leave the
text above their first concept in no segment -- 159 704 characters, 9.18 % of
the corpus, 45 841 from one document. It changes nothing on the reference. Off
because it moves the first span of essentially every bundle with no hit@8 number
behind it yet.

vegnormal-okf FUNN 2: SPEC section 8's own star row parsed as prose, so every
concept behind one was unreachable to the section 9.2 walk. `IndexPolicy.also_reads`
carries it for the SEGMENTED profiles, read-only, after the emitted pattern
misses -- the asymmetry `sources` already has. DEFAULT and STRICT_V1 untouched (O2).

vegnormal-okf FUNN 1: Door C's own outcome was refused at exit 1,
`bundle_id_missing`. `import_bundle` now takes `root_frontmatter_values`,
keyword-only, rendered before any disk mutation, written only when the index is
created -- Door B's mechanism and ordering.

Report: docs/2026-09-09-k3-runde6-outline-gaten-og-prioren.md.
Suite 1478 passed (1449 before), ruff and mypy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-09 14:17:00 +02:00
commit 38104b7df5
16 changed files with 1301 additions and 42 deletions

View file

@ -949,6 +949,28 @@ def _overlap(
)
#: How a document's prior grows with its unit count. `1.0` is a DENSITY and
#: `0.0` is a SUM; this is the classical length normalisation between them, and
#: it is here rather than inline because the value is a decision a reader should
#: find where the decision was taken.
#:
#: WHY IT MOVED (2026-09-09). A sum measures size -- that is why the density
#: replaced it -- but a density is diluted by every unit carrying none of the
#: question, so a document the segmenter split from 1 concept into 12 lost its
#: prior by a factor of 12. That put the segmentation side and the retrieval
#: side in direct competition over one number, and it is what blocked a
#: reference-improving default from shipping.
#:
#: SWEPT, not chosen: the gold document's rank under this prior over 6 questions
#: x 3 bundles = 18 rows, at 0.0, 0.25, 0.5, 0.75 and 1.0. 0.5 is at least as
#: good as the delivered 1.0 on all 18 rows and strictly better on three; 0.25
#: loses one row and 0.0 and 0.75 are measured beside it. HONESTY LIMIT: four
#: alternatives on 18 rows, one gold set, one rater -- and the rank of the
#: PRIOR is not the rank of the excerpt, because RRF fuses it with two other
#: signals. The end-to-end hit@8 measurement is the one that decided it.
DOCUMENT_PRIOR_EXPONENT = 0.5
def document_scores(
bundle_root: Path,
question: str,
@ -966,7 +988,13 @@ def document_scores(
nor `bundle_id`; scoring them as members of some parent would put one bug in
three places.
**The score is a DENSITY, not a sum, and that is a correction rather than a
**The score grows SUBLINEARLY with the unit count** -- `total /
n**DOCUMENT_PRIOR_EXPONENT`, the exponent at 0.5. Both endpoints are wrong
and each is wrong in its own direction; the constant above carries the
measurement and the sweep. The original correction, from a sum to a
density, is kept here because it is still the reason a sum is not used:
**A sum is not a score, and that is a correction rather than a
preference.** A sum over a document's units grows with the number of units,
so a large document outscores a small one on size alone. Measured on K2 for
the price question: the competition document sums to 6.0 over 79 concepts
@ -1016,7 +1044,10 @@ def document_scores(
document,
_overlap(question_tokens, entry.label, cost_vocabulary=bridge, weights=weights),
)
return {document: totals[document] / units[document] for document in totals}
return {
document: totals[document] / units[document] ** DOCUMENT_PRIOR_EXPONENT
for document in totals
}
# --- Stage two: which concepts inside those documents -------------------------