feat(consume): the payload says what of the question it reached, and row 4 reads it
`coverage` carries three lists: the terms the pre-pass read the question as, the terms no concept in the bundle answers, and the terms no delivered excerpt answers. Without it a reader holding eight excerpts cannot tell a bundle that ANSWERED its question from one that merely ranked something -- the two payloads have the same shape. FACTS, AND NO VERDICT, which is a measurement and not caution. Two readings were built and both falsified over 81 questions (16 synthetic, 65 across the three real sets, 2026-09-20): the share of question terms a delivered excerpt answers separates the synthetic controls at 0.33 against 0.50 and REVERSES on real data (covered questions down to 0.27, one genuinely uncovered question at 0.71); the share of a bundle tying the best lexical match is ~0.00 for every real question either way. Question style dominates the first, corpus size the second. The one bar this repository declares is the gate's: `UNANSWERED_BAR = 2/3` over `unanswered_in_bundle`, swept and collapsing at both ends -- at 0.50 eleven real covered questions are marked, at 0.70 the row falls to 5 of 6, at 2/3 the row is 6 of 6 and 0 of 65 real questions are marked. The margin is thin (0.6087 against 0.6667) and is published that way, together with what it does not catch: r761-sk2's own known-negative sits at 0.2857. Row 4: 3 of 6 RED -> 6 of 6 GREEN, with the 10 answered synthetic questions held unmarked as the known-negative. The contract's SS 8 gains point 7, the consumption skill is told to read the block, and the SS 7.4 known-positive moves with the document (14 721/375 -> 16 389/417). Suite 2292 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
90394c383d
commit
05cb19087a
8 changed files with 216 additions and 42 deletions
|
|
@ -2661,14 +2661,14 @@ def test_the_payload_names_the_question_terms_the_bundle_answers_to_none_of() ->
|
|||
the ones no delivered excerpt answers.
|
||||
"""
|
||||
payload = okf_consume.build_payload(
|
||||
FIXTURE, question="Hva koster et doegn paa hytta for gjester?"
|
||||
FIXTURE, question="Hva koster et doegn for gjester i kravet?"
|
||||
)
|
||||
assert "coverage" in payload, "the payload says nothing about what it covered"
|
||||
coverage = payload["coverage"]
|
||||
assert isinstance(coverage, dict)
|
||||
# The terms are the ranker's own, counted here rather than read back.
|
||||
expected = list(
|
||||
dict.fromkeys(okf_consume.normalise("Hva koster et doegn paa hytta for gjester?"))
|
||||
dict.fromkeys(okf_consume.normalise("Hva koster et doegn for gjester i kravet?"))
|
||||
)
|
||||
assert coverage["question_terms"] == expected
|
||||
assert expected, "a question with no terms would make every assertion below vacuous"
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue