feat(propose,consume,profiles,importer): recovery yields to declaration, and 9 % of the corpus that was in no segment
One rule explains every remaining `pdf` miss on the twelve-position reference: where a document DECLARES headings, Arm D's RECOVERED headings are the whole of the excess, and every declared one is a unit the reference wants. `--outline-gate` admits recovery only where the document declares none of its own, plus any one recovered heading covering OUTLINE_SHARE (0.20) of the text. It is `fold_units` clause 2's own principle moved from voting to admission, and it filters at ADMISSION so the text a removed mark opened is carried by the mark above it -- the post-filter form scores identically on all twelve positions and loses that text, which is why only one of them shipped. `--outline-gate` and `--drop-wrapped-outline` become the package default, one decision because neither carries the reference alone: `pdf` 2 of 8 -> 5 of 8 alone, 7 of 8 together; the sheet 5 of 12 -> 10 of 12; `docx` unchanged at 3 of 3. Each keeps an explicit opt-out. The bar the move had to clear was not the reference: hit@8 on a K2 bundle built with it holds 5 of 6 at ranks 1,1,1,1,1,-, no row losing rank 1. `--sheet-section-rows --keep-table-heading` reaches 11 of 12 and does NOT ship, because on a bundle built with it row 1 falls rank 1 -> 2. Cost to a consumer is a re-run: 492 concepts / 944 files -> 425 / 810. DOCUMENT_PRIOR_EXPONENT makes the document prior sublinear (total/n**0.5). A sum measures size and a density is diluted by every unit carrying none of the question, so a document split 1 -> 12 lost its prior by 12. Swept over five values on 18 rows it is at least as good as the delivered density everywhere and strictly better on three. Stated plainly: end to end it moved NOT ONE hit@8 row on any of four bundles, so it did not solve the knot it was adopted for -- what did is that the `pdf` gain never needed `--sheet-section-rows`. `--first-span-from-zero` is off and repairs a measured loss found while chasing one position's 940 characters: 32 of the 32 documents that get a plan leave the text above their first concept in no segment -- 159 704 characters, 9.18 % of the corpus, 45 841 from one document. It changes nothing on the reference. Off because it moves the first span of essentially every bundle with no hit@8 number behind it yet. vegnormal-okf FUNN 2: SPEC section 8's own star row parsed as prose, so every concept behind one was unreachable to the section 9.2 walk. `IndexPolicy.also_reads` carries it for the SEGMENTED profiles, read-only, after the emitted pattern misses -- the asymmetry `sources` already has. DEFAULT and STRICT_V1 untouched (O2). vegnormal-okf FUNN 1: Door C's own outcome was refused at exit 1, `bundle_id_missing`. `import_bundle` now takes `root_frontmatter_values`, keyword-only, rendered before any disk mutation, written only when the index is created -- Door B's mechanism and ordering. Report: docs/2026-09-09-k3-runde6-outline-gaten-og-prioren.md. Suite 1478 passed (1449 before), ruff and mypy clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
b01492b7f5
commit
38104b7df5
16 changed files with 1301 additions and 42 deletions
|
|
@ -949,6 +949,28 @@ def _overlap(
|
|||
)
|
||||
|
||||
|
||||
#: How a document's prior grows with its unit count. `1.0` is a DENSITY and
|
||||
#: `0.0` is a SUM; this is the classical length normalisation between them, and
|
||||
#: it is here rather than inline because the value is a decision a reader should
|
||||
#: find where the decision was taken.
|
||||
#:
|
||||
#: WHY IT MOVED (2026-09-09). A sum measures size -- that is why the density
|
||||
#: replaced it -- but a density is diluted by every unit carrying none of the
|
||||
#: question, so a document the segmenter split from 1 concept into 12 lost its
|
||||
#: prior by a factor of 12. That put the segmentation side and the retrieval
|
||||
#: side in direct competition over one number, and it is what blocked a
|
||||
#: reference-improving default from shipping.
|
||||
#:
|
||||
#: SWEPT, not chosen: the gold document's rank under this prior over 6 questions
|
||||
#: x 3 bundles = 18 rows, at 0.0, 0.25, 0.5, 0.75 and 1.0. 0.5 is at least as
|
||||
#: good as the delivered 1.0 on all 18 rows and strictly better on three; 0.25
|
||||
#: loses one row and 0.0 and 0.75 are measured beside it. HONESTY LIMIT: four
|
||||
#: alternatives on 18 rows, one gold set, one rater -- and the rank of the
|
||||
#: PRIOR is not the rank of the excerpt, because RRF fuses it with two other
|
||||
#: signals. The end-to-end hit@8 measurement is the one that decided it.
|
||||
DOCUMENT_PRIOR_EXPONENT = 0.5
|
||||
|
||||
|
||||
def document_scores(
|
||||
bundle_root: Path,
|
||||
question: str,
|
||||
|
|
@ -966,7 +988,13 @@ def document_scores(
|
|||
nor `bundle_id`; scoring them as members of some parent would put one bug in
|
||||
three places.
|
||||
|
||||
**The score is a DENSITY, not a sum, and that is a correction rather than a
|
||||
**The score grows SUBLINEARLY with the unit count** -- `total /
|
||||
n**DOCUMENT_PRIOR_EXPONENT`, the exponent at 0.5. Both endpoints are wrong
|
||||
and each is wrong in its own direction; the constant above carries the
|
||||
measurement and the sweep. The original correction, from a sum to a
|
||||
density, is kept here because it is still the reason a sum is not used:
|
||||
|
||||
**A sum is not a score, and that is a correction rather than a
|
||||
preference.** A sum over a document's units grows with the number of units,
|
||||
so a large document outscores a small one on size alone. Measured on K2 for
|
||||
the price question: the competition document sums to 6.0 over 79 concepts
|
||||
|
|
@ -1016,7 +1044,10 @@ def document_scores(
|
|||
document,
|
||||
_overlap(question_tokens, entry.label, cost_vocabulary=bridge, weights=weights),
|
||||
)
|
||||
return {document: totals[document] / units[document] for document in totals}
|
||||
return {
|
||||
document: totals[document] / units[document] ** DOCUMENT_PRIOR_EXPONENT
|
||||
for document in totals
|
||||
}
|
||||
|
||||
|
||||
# --- Stage two: which concepts inside those documents -------------------------
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue