feat(propose,consume,tools): the type that declares nothing, and the prefix that is not a word
Three of round 9's four measured holes, each closed with a rule chosen on a measurement rather than named as a limit. `rtf` GIVES 0 SEGMENTS -> 6 of 6 AUTHORED TITLES over N = 4. The container has no heading style, so the author's title is bold text. The grammar is markdown, not `rtf`: the converter already writes that title as `**...**` in the same output every office row produces, so no `rtf`-only heading form exists. Three parameters were swept over 47 readable documents and ONE carried -- refusing a line that ends in terminal punctuation takes false-positive lines from 9-12 to 1-2. A maximum title length (unlimited/40/60/80/120) and a must-stand-between-blank-lines clause are both FLAT, so neither is in the rule. The last false positive is closed by G1, the principle `_gate_outline` already carries: recovery yields to declaration. False positives are then 0 of the 31 declaring documents by construction, and 0 of 27 on the corpus. Reach: 2 of 39 corpus documents, both `docx`, 0 of 33 `pdf` and 0 of 2 `xlsx`. Behind `--bold-title`, default OFF pending the hit@8 measurement; the default bundle is byte-identical without it. BOTH ALTERNATIVES THE ORDER NAMED WERE MEASURED AND FELLED. A fourth hand-laid fixture DECLARES heading styles in a stylesheet and the converter discards them, emitting the same bold line -- so "read the declared headings out of the markdown" has nothing to read. `rtf` -> `docx` -> markdown yields 0 ATX headings on that same document, because the loss is in the `rtf` READER before any writer sees the style. Fixtures are hand-laid in `make_k2_office.py` with the fasit written first; they live in their own directory because Door B walks a drop directory recursively and `k2-office/` reads its N off the listing. THE PREFIX OVER-MATCH: THREE CANDIDATES MEASURED, ALL THREE FAILED ON ONE ROW. Re-measured on the pinned 453-concept bundle with the control run first: `under` occurs 79 times by equality and matches 172 by prefix, `undersjoisk` 0 and 172, `bilateral` 0 and 400 of 453, `standhaftig` 0 and 219. The two extra known-negatives were FOUND, not chosen -- every 4-character prefix ranked by document frequency, then a real word taken from the widest. A longer floor (5-8), a coverage share (0.5-0.8) and a long-words-only floor (>= 8) each cost row 1 its rank on the default bundle and the whole row on Arm B. Decomposed: row 1's token `prisene` reaches its gold document through `pris|sammenstilling` on four characters -- 0.57 of one word and 0.22 of the other -- so the over-match and the wanted match are one mechanism. THE FOURTH CANDIDATE IS THE ANSWER: the shared prefix must be a WORD the bundle uses. `pris` is; `bila` and `stan` are not. `bilateral` 400 -> 0 and 512 -> 0, `standhaftig` 219 -> 56 and 235 -> 33, every hit@8 row keeping rank 1 on BOTH bundles. `undersjoisk` stops at 162 because `under` IS a word here -- a genuine Norwegian morpheme, so that residual is a different answer, not a ceiling. ON by default (`--no-stem-prefix`), pinned with its own known-negative on the shipped bytes. THE SHIM: a path importer holds the object `module_from_spec` made, and `sys.modules[__name__] = _impl` never reaches it. Measured under both counting methods -- 3 of 76 public names by `vars()`. One line copies the public names into this file's globals; the dunder filter is load-bearing, because an unfiltered copy overwrites `__name__` before the next line uses it as the alias key. It restores attribute ACCESS and not patch-through, which is why the alias stays. A CHANGELOG note under 0.7.0 and a shim docstring line say so, since what the consumer asked for was the note. Suite 1515 -> 1535. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
6ff84d71c8
commit
191de89f41
16 changed files with 1100 additions and 22 deletions
|
|
@ -183,6 +183,15 @@ DEFAULT_CLOSE_SPAN_GAPS = True
|
|||
#: them from a contents entry. It only ever REMOVES members from a run, so it
|
||||
#: can only add concepts, never take one away. Opt-out `--no-contents-name`.
|
||||
DEFAULT_CONTENTS_NAME = True
|
||||
#: Round 10's rule for the type that declares nothing. `rtf` came back at 0 of
|
||||
#: 0 declared headings, 0 concepts and 1368 of 1368 characters in no segment:
|
||||
#: the container has no heading style, so the author's title is bold text.
|
||||
#: The grammar is markdown, not `rtf` -- the converter already writes that
|
||||
#: title as `**...**` in the same output every office row produces -- and it is
|
||||
#: gated by the principle Arm D already carries, that recovery yields to
|
||||
#: declaration. Three parameters were swept over 47 readable documents and one
|
||||
#: carried (see `tests/test_bold_title.py`). Default set by measurement below.
|
||||
DEFAULT_BOLD_TITLE = False
|
||||
|
||||
#: Round 3's two spreadsheet rules (D1 and D3), held back through rounds 5 and
|
||||
#: 6 by a RETRIEVAL regression rather than by the reference: they take the
|
||||
|
|
@ -250,6 +259,7 @@ def _propose_plans(
|
|||
first_span_from_zero: bool = False,
|
||||
close_span_gaps: bool = False,
|
||||
contents_name: bool = False,
|
||||
bold_title: bool = False,
|
||||
pdf_headings: bool = False,
|
||||
pdf_headings_reserve: bool = False,
|
||||
ocr: bool = False,
|
||||
|
|
@ -287,6 +297,7 @@ def _propose_plans(
|
|||
first_span_from_zero=first_span_from_zero,
|
||||
close_span_gaps=close_span_gaps,
|
||||
contents_name=contents_name,
|
||||
bold_title=bold_title,
|
||||
pdf_headings=pdf_headings,
|
||||
pdf_headings_reserve=pdf_headings_reserve,
|
||||
ocr=ocr,
|
||||
|
|
@ -323,6 +334,7 @@ def build(
|
|||
first_span_from_zero: bool = DEFAULT_FIRST_SPAN_FROM_ZERO,
|
||||
close_span_gaps: bool = DEFAULT_CLOSE_SPAN_GAPS,
|
||||
contents_name: bool = DEFAULT_CONTENTS_NAME,
|
||||
bold_title: bool = DEFAULT_BOLD_TITLE,
|
||||
pdf_headings: bool = DEFAULT_PDF_HEADINGS,
|
||||
pdf_headings_reserve: bool = DEFAULT_PDF_HEADINGS_RESERVE,
|
||||
ocr: bool = DEFAULT_OCR,
|
||||
|
|
@ -394,6 +406,7 @@ def build(
|
|||
first_span_from_zero=first_span_from_zero,
|
||||
close_span_gaps=close_span_gaps,
|
||||
contents_name=contents_name,
|
||||
bold_title=bold_title,
|
||||
pdf_headings=pdf_headings,
|
||||
pdf_headings_reserve=pdf_headings_reserve,
|
||||
ocr=ocr,
|
||||
|
|
@ -725,6 +738,27 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
|||
"pre-2026-09-09 contents-run predicate byte for byte"
|
||||
),
|
||||
)
|
||||
build_parser.add_argument(
|
||||
"--bold-title",
|
||||
action="store_true",
|
||||
default=DEFAULT_BOLD_TITLE,
|
||||
help=(
|
||||
"Read a line that is ONE bold span as a title, in a document that "
|
||||
"declares no heading of its own. For the type whose container has "
|
||||
"no heading style at all: `rtf` reached round 9 at 0 of 0 declared "
|
||||
"headings, 0 concepts and 1368 of 1368 characters in no segment. "
|
||||
"The grammar is markdown, so it reaches every type the converter "
|
||||
"writes bold for, and it is inert for `pdf`, which never goes "
|
||||
"through the converter. Measured over 47 readable documents: 0 "
|
||||
"false positives on the 31 that declare, by the gate"
|
||||
),
|
||||
)
|
||||
build_parser.add_argument(
|
||||
"--no-bold-title",
|
||||
action="store_false",
|
||||
dest="bold_title",
|
||||
help="The rule's explicit opt-out",
|
||||
)
|
||||
build_parser.add_argument(
|
||||
"--pdf-headings",
|
||||
choices=("none", "font", "font-reserve"),
|
||||
|
|
@ -799,6 +833,7 @@ def main(argv: list[str] | None = None) -> int:
|
|||
first_span_from_zero=args.first_span_from_zero,
|
||||
close_span_gaps=args.close_span_gaps,
|
||||
contents_name=args.contents_name,
|
||||
bold_title=args.bold_title,
|
||||
pdf_headings=args.pdf_headings == "font",
|
||||
pdf_headings_reserve=args.pdf_headings == "font-reserve",
|
||||
ocr=args.ocr,
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue