llm-ingestion-okf/src/llm_ingestion_okf/propose.py
Kjell Tore Guttormsen b01492b7f5 feat(propose,cli): typography as a reserve, and the two of our own numbers it took to measure it
K3 round 5. Three questions, three answers, and two of them correct a figure
this repository published.

RETRIEVAL FIRST, because it could have reversed a default. hit@8 over the six
questions on BOTH K2 bundles -- Arm B at 629 concepts and the shipped default
at 492 -- is 5 of 6 with ranks 1,1,1,1,1,- on each, so 0 of 6 rows lost. The
order's rule reverses `--unit-fold` at >= 2 of 6; it does not fire, and the
default stands. The gold sets shrink (49->26, 20->17, 43->36, 11->18) while
every rank holds at 1, which is the fold merging concepts rather than removing
a document from the top.

TWO PUBLISHED NUMBERS CORRECTED, both ours. The S7 candidate ranks 96 of 629
and 159 of 492 were measured with the cost vocabulary passed to
`concept_scores` and NOT to `document_scores`, while `build_payload` passes it
to both; scored the way the shipped payload scores it, the same concept is 10
of 629 and 19 of 492. And round 4 attributed its non-delivery to the default
move -- measured here, it is not delivered on the Arm B bundle either, for a
different reason (knapsack eviction at 68 046 bytes of a 120 000 budget, versus
`below_k`). That column had been inherited from round 3's own build, never
re-measured.

`--pdf-headings font-reserve`, OFF, and the hypothesis behind it is falsified
by its own condition rather than by a score: position 7, the one position the
flag exists for, has THREE outline runs, so the reserve is silent there at
every minimum. It changes 0 of 12 cells on the reference and reaches 4 of 39
corpus documents, none of them rated. Built anyway because it was authorised
and because the condition is now measured rather than assumed. The predicate
lives in one place (`propose.heading_reserve_applies`) and the door receives it
as a callable, like `gate`: a plan indexes the exact string it was proposed
against, so a reserve firing on one side only would make every document it
touches a coded rejection.

The `xlsx` re-reading is confirmed on the artifact -- 11 `rule:sheet-section`
units plus 1 `rule:table-block` ingress -- but the number alone makes the cell
worse (distance 1 -> 2), because the criterion counts that ingress as a table
that should have been merged. A hit needs both halves ratified, and the
reference is the operator's.

`--sheet-section-rows` as a default: three cells better and none worse on the
twelve positions, but the K2 control moves -- row 1's gold document splits 1 ->
12 concepts and its best concept ranks 2 instead of 1. Condition not met,
default not moved.

Default build byte-identical before and after (`diff -r`, 30 md files).
Suite 1441 -> 1449; three of the eight were red first.

Report: docs/2026-09-08-k3-runde5-hitat8-og-skriftakse.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 00:25:51 +02:00

1409 lines
64 KiB
Python
Executable file

#!/usr/bin/env python3
"""Propose a segmentation plan for one document. A human adjudicates it.
Pipeline step 3. It lives in the package because `okf build` has to reach it
from an INSTALLED copy, where `tools/` does not exist -- it was outside `src/`
until then, and the reason it could move is that the reason it sat outside was
never about dependencies: every rule below is mechanical, so nothing here adds
a model call to a package that promises none.
What DID have to survive the move is the separation the old location expressed
physically: the split of a document into units of knowledge is a judgement, and
the run path replays a decision somebody already made. That separation is
carried by `adjudicated: false` and the `PROPOSED` marker, which is where it
belonged all along -- a directory boundary cannot enforce it, and a caller who
ingests a proposal unadjudicated could always do so.
## What the research says this tool may and may not claim
Topic 2 measured the OKF reference agent's granularity criteria against
`_okf-canonical`: it splits on **what a thing is**, not on layout, and makes
"multiple `write_concept_doc` calls ... rather than dumping everything into one
doc". Four of its gates are semantic and need a model. A handful of MECHANICAL
rules port today, and those are the ones below.
Topic 1b measured heading derivation on the K2 corpus: 11 of 11 prose headings
recovered -- from ONE document. 23 of 33 PDFs carry no outline at all and 95 %
of the outline entries that do exist are AutoCAD export metadata. The
denominator is 1. A rule validated on n=1 is not validated, and this tool says
so by marking every entry it emits `PROPOSED` rather than adjudicated.
Topic 1a measured that the best deterministic heading rule from poppler is a
CONJUNCTION -- `size AND bold`, via `-fontfullname` -- at recall 1.000 and
precision 0.846, and that adding weight as a DISJUNCT makes precision worse
(0.786 -> 0.524). That path is implemented here and nowhere else: poppler is a
SYSTEM binary the `[extract]` extra cannot express, so it may never be on the
run path or in a golden fixture.
## The one rule that is not a heuristic
**Nothing here is ever adjudicated.** `adjudicated: false` sits at the top of
every artifact and `PROPOSED` in every entry's `derived` list. A plan is
replayed deterministically and forever by the run path, so a proposal that
could pass for an adjudication would put a machine's guess where a human's
judgement is supposed to be, permanently and silently.
Stdlib only. No network: the model-backed path this tool deliberately does not
have would need the per-run network opt-in, and the socket-free test suite
proves the absence rather than assuming it.
"""
from __future__ import annotations
import argparse
import hashlib
import json
import re
import sys
import unicodedata
from dataclasses import dataclass, replace
from pathlib import Path
from typing import Any
from .errors import IngestError
from .extract import extract_text
from .materialize import reduce_to_id_grammar
from .segmentation import observed_extractor_version
#: Stamped into every entry's `derived` list. The marker is what keeps a
#: proposal from being mistaken for the judgement the run path replays.
PROPOSED_MARKER = "PROPOSED"
#: This tool's identity, written into the artifact so an operator reading a
#: plan six months later can tell what produced it.
PROPOSER_ID = "okf-propose-segments"
PROPOSER_VERSION = "1"
#: The rules that survived Topic 2's port test. Every entry names the rule that
#: OPENED its span, so a proposal an operator disagrees with is traceable to the
#: rule that made it rather than to the tool as a whole. An entry an arm later
#: reshaped names that arm too -- Arm C's cut, Arm E's join -- because the rule
#: that opened the span is still true of it. Exactly one name per entry held
#: until the arms existed; it is the ORIGIN that is single, not the list.
RULE_HEADING = "rule:heading"
RULE_TABLE_BLOCK = "rule:table-block"
RULE_POPPLER_SIZE_AND_BOLD = "rule:poppler-size-and-bold"
#: Arm C only. NOT one of Topic 2's ported rules and not a heading rule at
#: all: it names the fact that a span was cut because it was too long, which
#: is a judgement about SIZE and says nothing about where a unit of knowledge
#: begins. It is emitted ALONGSIDE the rule that proposed the origin span, so
#: an operator reading a part can still see what opened it.
RULE_SIZE_SPLIT = "rule:size-split"
#: Arm D only. Like Arm C it is NOT one of Topic 2's ported rules and NOT
#: defined upstream: `docs/2026-09-02-k3-k4-k5-metode.md` contains no
#: occurrence of the word "arm" at all, so this definition was written for the
#: brief of order 20260906T213322Z and is reported as the author's. Unlike Arm
#: C it says nothing about size -- it names the fact that the DOCUMENT ITSELF
#: declared a chapter there, by numbering it in an ascending run its own
#: outline sustains.
RULE_OUTLINE = "rule:outline"
#: Arm E only. Like Arm C and Arm D it is NOT one of Topic 2's ported rules and
#: NOT defined upstream -- `docs/2026-09-02-k3-k4-k5-metode.md` contains no
#: occurrence of the word "arm" at all -- so this definition was written for
#: order 20260907T075834Z-18584396-from-.claude and is reported as the author's.
#: Its axis is a third one. Arm C names SIZE and Arm D names what the DOCUMENT
#: declared; this names what the CONVERTER emitted: a table block that was
#: joined across a grid-table rule line. Emitted ALONGSIDE `rule:table-block`,
#: which is still what opened the span, and only on a block that was ACTUALLY
#: joined -- never on one whose span merely happens to contain a rule line, so
#: a single-row grid table stays byte-identical to Arm D.
RULE_TABLE_GRID = "rule:table-grid"
#: D3 round 3 only, and it is the first rule in this module that opens a span
#: INSIDE a table rather than at one. Like Arm C, D and E it is NOT one of Topic
#: 2's ported rules and NOT defined upstream -- `docs/2026-09-02-k3-k4-k5-metode.md`
#: contains no occurrence of the word "arm" at all -- so this definition was
#: written for order 20260908T170037Z-3622420612-from-.claude and is reported as
#: the author's. Its axis is a fourth one. Arm C names SIZE, Arm D what the
#: DOCUMENT declared, Arm E what the CONVERTER emitted; this names what the
#: SHEET labelled: a row whose first cell is a bare numeric label, in a run of
#: such rows. It is the opposite DIRECTION from Arm E -- that one stops a rule
#: line from closing a block so a grid table proposes one candidate instead of
#: many, this one cuts an open block at the rows that label its sections -- and
#: the two compose in one order: Arm E decides how far a block extends, this
#: decides where it is cut inside.
RULE_SHEET_SECTION = "rule:sheet-section"
RULE_NAMES = (
RULE_HEADING,
RULE_TABLE_BLOCK,
RULE_POPPLER_SIZE_AND_BOLD,
RULE_SIZE_SPLIT,
RULE_OUTLINE,
RULE_TABLE_GRID,
RULE_SHEET_SECTION,
)
#: How many characters of context each side of a quote anchor carries. Enough
#: to separate two occurrences of a repeated heading, short enough that an
#: edit NEAR a segment does not invalidate the anchor FOR it -- the anchor
#: exists to survive shifts, so making it fragile would defeat it.
ANCHOR_CONTEXT = 48
#: Norwegian and English function words. A heading made only of these names no
#: unit of knowledge -- it is a connective that happened to sit on its own line.
#: Topic 2's stop-word gate, and the only place this tool judges wording.
STOP_WORDS = frozenset(
{
"and",
"as",
"at",
"av",
"be",
"by",
"da",
"de",
"den",
"der",
"det",
"en",
"er",
"et",
"for",
"fra",
"i",
"in",
"is",
"it",
"med",
"of",
"og",
"om",
"on",
"or",
"over",
"paa",
"som",
"til",
"the",
"to",
"under",
"ved",
"with",
}
)
# An ATX heading, or a numbered section opening a line (`3.1 Brannkonsept`).
# A BARE integer is not a section number, for the same reason `structure.py`
# refuses one: `12 ting` is an ordinary line and admitting it would cut a
# document at every list item.
#
# That claim still holds, and Arm D does not weaken it. `_OUTLINE` below admits
# a bare integer ONLY inside an ascending run the document sustains for at
# least a declared length -- which is a property of the whole text, not of the
# line -- and the rule is off unless a caller asks for it. An UNGATED widening
# was measured and rejected: 1681 raw hits against 618 candidates, admitting
# list items, quantities and page furniture. The gate is what makes the signal
# a signal.
_ATX = re.compile(r"^(?P<hashes>#{1,6})\s+(?P<title>\S.*?)\s*$")
_NUMBERED = re.compile(r"^(?P<number>\d+(?:\.\d+)+)\s+(?P<title>\S.*?)\s*$")
_TABLE_ROW = re.compile(r"^\s*\|.*\|\s*$")
# Arm E's grammar: a pandoc GRID-table rule line. The converter separates a grid
# table's rows with `+---+---+`, and its header from its body with `+===+===+`.
# Neither matches `_TABLE_ROW`, so `in_table` is reset between every pair of rows
# and ONE table becomes one candidate per row group. Measured on the K2 corpus:
# three documents carry grid tables, and they account for 33 of the 709 entries
# Arm D proposes.
#
# The character class is measured rather than guessed. Across those three
# documents, 38 of 38 lines whose stripped form starts with `+` match this
# pattern, and `+`, `-`, `:`, `=` is the COMPLETE set of characters occurring on
# them. The `:` is pandoc's column-alignment marker, and it is not decoration: a
# first pass with `[-=+]` matched 37 of 38 and, through that one miss, read one
# document as having two tables where it has one.
#
# `\s*` on both ends mirrors `_TABLE_ROW` rather than tightening on it, because
# the loop iterates `splitlines(keepends=True)` -- every line carries its `\n`,
# and an indented rule line is a real shape that must still be admitted.
_GRID_RULE = re.compile(r"^\s*\+[-=:+]+\+\s*$")
# Arm D's grammar. Integer-only BY CONSTRUCTION: `\s+` after the optional
# separator is what keeps `1.1 Brannkonsept` out, because `_NUMBERED` requires
# a dot and this requires whitespace, so no line can match both. No exclusion
# clause is written for that: a filter with a measured effect of zero is dead
# code that reads like a guard.
_OUTLINE = re.compile(r"^\s{0,4}(?P<number>\d{1,2})[.)]?\s+(?P<title>\S.*?)\s*$")
# A contents line carries the page it points at (`Innledning 6`). Measured on
# the K2 corpus: stripping it changes 0 of the 144 outline counts and 9 emitted
# titles. It is load-bearing anyway, because titles become concept paths
# through `_segment_path` -- an unstripped page number would become part of a
# filename.
_TRAILING_PAGE_NUMBER = re.compile(r"[\s.]+\d{1,4}\s*$")
# D3's grammar, and it reads a CELL rather than a line. A sheet's section label
# is a bare number, optionally joined to another by a separator where two groups
# were merged (`11+12`), and nothing else: no letters, so a row opening with a
# word is not a section, and no word list, so the rule knows nothing about which
# numbers any real sheet uses.
_SHEET_SECTION_LABEL = re.compile(r"^\d+(?:[+./-]\d+)*$")
# A pipe that pandoc did not escape. Splitting a row on a bare `|` would cut a
# cell containing a literal pipe in half and misread the FIRST cell of the row
# after it, which is the only cell this rule judges.
_UNESCAPED_PIPE = re.compile(r"(?<!\\)\|")
class ProposerError(Exception):
"""The run failed. NOT 'nothing to propose' -- the two must stay distinct."""
@dataclass(frozen=True)
class Candidate:
"""One proposed boundary, before it becomes an entry."""
title: str
level: int
number: str | None
rule: str
start: int
end: int
#: True when this candidate is one PART of a longer span that Arm C cut.
#: Kept on the candidate rather than recomputed at write time so the entry
#: and the reason it exists cannot drift apart.
split: bool = False
#: True when Arm E JOINED this table block across at least one grid-rule
#: line. Same reasoning as `split`, and the same trap: both reconstruction
#: sites below rebuild a `Candidate` from an explicit keyword list, so a
#: field not copied there is silently defaulted back and the entry loses
#: the only trace of why it exists. NOT set merely because a span contains
#: a rule line -- a single-row grid table joins nothing and stays
#: byte-identical to Arm D.
grid: bool = False
#: True when this candidate belongs to a run of page-numbered siblings that
#: was a contents list BEFORE the orphan check thinned it. Set only when
#: Arm F is on, because it exists only for Arm F's clause 1 to read: a
#: contents list without dot leaders is a run of bodiless headings, so the
#: orphan check deletes all but the last and the run clause 1 looks for is
#: gone by the time `fold_units` sees the list. Computed where the whole
#: pre-orphan list is still in hand, and nowhere else -- no candidate
#: carrying it survives clause 1, so it never reaches an artifact.
contents: bool = False
def _is_stop_word_only(title: str) -> bool:
words = [word for word in re.split(r"[^\w]+", title.lower()) if word]
return bool(words) and all(word in STOP_WORDS for word in words)
def _strip_page_number(title: str) -> str:
"""Remove a trailing page number from a contents-listing title.
Deliberately NOT applied to a title that is only digits: `477` has no
separator before the number, so the pattern cannot match it and the title
survives for the stop-word and junk paths to see. Emptying it would fall
back to the `seksjon` stem and dress junk as a named section.
"""
return _TRAILING_PAGE_NUMBER.sub("", title)
def outline_lines(text: str) -> list[tuple[int, int, str]]:
"""Every line the outline grammar admits, as `(line index, integer, title)`.
Module level and importable on purpose: the reach instrument measures this
rule, and an instrument that re-implements the grammar it measures is
measuring a second definition that can silently drift from the shipped one.
"""
found: list[tuple[int, int, str]] = []
for index, line in enumerate(text.splitlines()):
match = _OUTLINE.match(line)
if match is None:
continue
title = _strip_page_number(match.group("title")).strip()
if not title or _is_stop_word_only(title):
continue
found.append((index, int(match.group("number")), title))
return found
def heading_reserve_applies(text: str, *, outline_run: int) -> bool:
"""Whether this text needs a SECOND heading source, having no run of its own.
The font reader's reserve condition, and the only place it is decided. The
proposer and the door both call this, because a plan indexes the exact
string it was proposed against: a reserve that fired on one side and not
the other would make every document it touched a coded rejection.
It reads the gate AS CONFIGURED rather than a fixed minimum -- raising the
arm's threshold widens the reserve, which is the same document property
seen through the same threshold. At `outline_run` 0 the gate admits nothing
at all, so the reserve is unconditional; that combination is round 4's
"font instead of Arm D", measured at 1 of 8, and a caller reaching it gets
it deliberately.
"""
if outline_run <= 0:
return True
return not outline_runs(outline_lines(text), outline_run)
def outline_runs(
entries: list[tuple[int, int, str]], minimum: int
) -> list[list[tuple[int, int, str]]]:
"""The maximal ascending runs among `entries`, each at least `minimum` long.
A run is anchored at `1` and every later member is its predecessor plus
one; a number that is neither is skipped without closing the run, so a
stray page number between two chapters does not truncate the outline. A new
`1` closes the current run and opens the next, which is what makes a
contents listing and the body it lists two runs rather than one.
Returned in document order. The CALLER chooses among them -- last-run
selection was measured against the alternatives and is stated where it is
applied, not hidden in here.
"""
runs: list[list[tuple[int, int, str]]] = []
current: list[tuple[int, int, str]] = []
for entry in entries:
number = entry[1]
if number == 1:
if current:
runs.append(current)
current = [entry]
elif current and number == current[-1][1] + 1:
current.append(entry)
if current:
runs.append(current)
return [run for run in runs if len(run) >= minimum]
#: D3. How many CONSECUTIVE numbered rows make a sectioned table. Three, and it
#: is the same bounding device -- and the same number -- `CONTENTS_RUN` uses,
#: for the same reason: a single numbered row is a stated quantity, not a
#: section, and a rule that read one would cut a sheet at every computation
#: basis. NOT swept on a reference, because no reference exists for its value;
#: the corpus SENSITIVITY of the choice is published instead, in
#: `docs/2026-09-08-k3-runde3-per-filtype.md`. Not a CLI knob for the reason
#: `CONTENTS_RUN` is not one.
SHEET_SECTION_RUN = 3
def _wraps_onto_next_line(lines: list[str], index: int) -> bool:
"""True when the line at `index` is a sentence that continues below it.
The test is the next line's first character being lower case. A heading is
a complete line -- the line under it opens a new sentence, is blank, or is
a bullet -- while a hard-wrapped paragraph carries its own continuation.
This is deliberately NOT a length test. Round 2 sorted every outline title
in the K3 sample and measured the two classes overlapping: a real chapter
heading of 88 characters against quoted sentences of 86, 91, 92 and 100, so
no threshold separates them. Measured on the same sample, this axis
separates 8 of 34 -- the four quoted regulation paragraphs and four risk
table rows, and none of the 26 headings the operator kept.
"""
if index + 1 >= len(lines):
return False
following = lines[index + 1].strip()
return bool(following) and following[0].islower()
def _sheet_section_rows(lines: list[str]) -> dict[int, tuple[str, str]]:
"""D3: line index -> (label, row text) for every section row, or empty.
A section row is a table row whose FIRST cell is a bare numeric label and
which carries at least one other non-empty cell -- the label alone names
nothing, and the first non-empty cell after it is what the section is
called. The rows must come in a run of at least :data:`SHEET_SECTION_RUN`
consecutive lines: that is what separates a labelled section column from a
quantity stated on its own row, and it is the module's existing guard
rather than a new one.
Computed over the whole line list before the marking loop, for the reason
the outline runs are: a run is a property of the text, and a forward scan
that decided one row at a time could not know whether the run it is inside
is long enough.
"""
labelled: dict[int, tuple[str, str]] = {}
for index, line in enumerate(lines):
if not _TABLE_ROW.match(line):
continue
cells = [cell.strip() for cell in _UNESCAPED_PIPE.split(line.strip().strip("|"))]
if not cells or not _SHEET_SECTION_LABEL.match(cells[0]):
continue
rest = [cell for cell in cells[1:] if cell]
if not rest:
continue
labelled[index] = (cells[0], rest[0])
sections: dict[int, tuple[str, str]] = {}
ordered = sorted(labelled)
start = 0
while start < len(ordered):
end = start + 1
while end < len(ordered) and ordered[end] == ordered[end - 1] + 1:
end += 1
if end - start >= SHEET_SECTION_RUN:
for position in range(start, end):
sections[ordered[position]] = labelled[ordered[position]]
start = end
return sections
def find_candidates(
text: str,
*,
outline_run: int = 0,
table_grid: bool = False,
unit_fold: bool = False,
keep_table_heading: bool = False,
sheet_section_rows: bool = False,
drop_wrapped_outline: bool = False,
) -> list[Candidate]:
"""Every boundary the mechanical rules propose, in document order.
Two gates from Topic 2 are applied here and both REMOVE candidates:
- the **stop-word gate**: a heading made only of function words is not a
unit of knowledge;
- the **orphan check**: a heading with no body under it proposes nothing,
because an empty concept is the silent skip this library refuses
everywhere else.
`outline_run` is Arm D's gate and it is OFF at 0: the function then behaves
exactly as it did before the rule existed. At `N >= 1` the document's own
numbered outline contributes boundaries where the integers sustain an
ascending run of at least `N`.
`unit_fold` is Arm F's gate and it is OFF at False, where `fold_units` is
not called at all. On, the candidate list is folded once, at the end: the
arm adds no boundary, so every plan it can produce is a subset of the one
the flags below produce.
`table_grid` is Arm E's gate and it is OFF at False, where the branch is not
even evaluated. On, a pandoc grid-table rule line no longer closes an open
table block, so one grid table proposes one candidate instead of one per row
group. It only ever REMOVES marks, which is what keeps every surviving
candidate's `start` fixed and the orphan check monotone.
`keep_table_heading` is D1's gate and it is OFF at False. On, a heading
whose body is empty ONLY because a table block opens under it keeps that
table instead of being dropped: the table is absorbed into the heading's
span rather than emitted, so the concept starts at the heading line. The
count does not move -- one candidate either way -- and the first byte does.
It is its own flag and not part of an arm because the orphan check is
reached by every file type, and moving it is a decision about all of them.
`sheet_section_rows` is D3's gate and it is OFF at False, where the scan is
not run at all. On, a RUN of numbered rows inside an open table block cuts
it: each such row opens a candidate that reaches the next section row, or
the end of the block. It is the only rule here that opens a span inside a
table, and it is its own flag for the same reason D1 is -- a sheet is the
one file type whose units are rows, and every other type reaches this scan
too.
"""
lines = text.splitlines(keepends=True)
offsets: list[int] = []
position = 0
for line in lines:
offsets.append(position)
position += len(line)
end_of_text = position
# Computed BEFORE the loop, and that is a correctness requirement rather
# than a style choice: run selection is a whole-text decision (the LAST
# maximal run wins, because a contents listing precedes the body it lists),
# and a forward scan cannot know which run is last. Deciding it up front is
# also what keeps `marked` sorted by construction -- appending outline
# candidates in a second pass would leave `end < start` on some spans, and
# `text[start:end]` is then `""`, so the orphan check DELETES them
# silently. Silent loss, not a raise: nothing would announce it.
admitted: dict[int, str] = {}
if outline_run > 0:
runs = outline_runs(outline_lines(text), outline_run)
if runs:
# LAST run, not longest and not first. Measured against both:
# first-run opens segments inside the table of contents on 14/39
# documents; longest-run differs on 5/39 with no measured reason to
# prefer it. "Later occurrence wins" states the document's own
# ordering rather than a property of this corpus.
admitted = {index: title for index, _, title in runs[-1]}
# D3's input, and the same whole-text reasoning as `admitted` above: a run
# is a property of the line list, not of a line.
sections = _sheet_section_rows(lines) if sheet_section_rows else {}
marked: list[tuple[int, Candidate]] = []
in_table = False
# Arm E's state, and all three of these are cleared together on the
# fall-through below. `rule_pending` is the one that matters: a grid table
# ends with a bottom rule, which sets it, and if the blank line after the
# table did not clear it the NEXT table's first row would be recorded as a
# join although nothing was joined. `joined` holds positions in `marked`
# rather than mutating a candidate, because `Candidate` is frozen and the
# orphan-check pass below already rebuilds every one of them.
rule_pending = False
open_block: int | None = None
joined: set[int] = set()
for index, line in enumerate(lines):
if _TABLE_ROW.match(line):
section = sections.get(index)
if section is not None:
label, name = section
# A section row never ALSO opens a table block, even when the
# run starts on the block's first line: two marks on one line
# would give the second an empty span, and the orphan check
# deletes an empty span silently. `open_block` is cleared with
# it, because a block cut here is no longer one span an Arm E
# join could extend.
in_table = True
open_block = None
rule_pending = False
marked.append(
(
index,
Candidate(
title=f"{label} {name}",
# The table sentinel, not a heading depth: a sheet
# row declares no level, and letting it vote in
# Arm F's clause 2 would make a sectioned sheet a
# one-level document.
level=9,
number=label,
rule=RULE_SHEET_SECTION,
start=offsets[index],
end=end_of_text,
),
)
)
continue
if not in_table:
in_table = True
open_block = len(marked)
marked.append(
(
index,
Candidate(
title=f"Tabell linje {index + 1}",
level=9,
number=None,
rule=RULE_TABLE_BLOCK,
start=offsets[index],
end=end_of_text,
),
)
)
elif rule_pending and open_block is not None:
joined.add(open_block)
rule_pending = False
continue
if table_grid and in_table and _GRID_RULE.match(line):
# The whole of Arm E: do NOT close the block. Reaching here with
# `in_table` False is impossible by construction, so a rule line can
# never OPEN a block -- a table's top border is not a boundary, and
# every surviving candidate keeps the `start` it had under Arm D.
rule_pending = True
continue
in_table = False
rule_pending = False
open_block = None
outline_title = admitted.get(index)
if outline_title is not None:
if drop_wrapped_outline and _wraps_onto_next_line(lines, index):
# Filtered at ADMISSION rather than at run selection: the run
# this document sustains is a property of its numbering, and
# re-selecting it from a thinned list would move boundaries on
# documents where nothing wraps. The rule declines candidates;
# it does not rewrite which run won.
continue
outline_match = _OUTLINE.match(line)
assert outline_match is not None, "an admitted index still matches the grammar"
marked.append(
(
index,
Candidate(
title=outline_title,
level=1,
number=outline_match.group("number"),
rule=RULE_OUTLINE,
start=offsets[index],
end=end_of_text,
),
)
)
continue
atx = _ATX.match(line)
numbered = _NUMBERED.match(line)
if atx is None and numbered is None:
continue
if atx is not None:
title = atx.group("title")
level = len(atx.group("hashes"))
inner = _NUMBERED.match(title)
number = inner.group("number") if inner else None
else:
assert numbered is not None
title = numbered.group("title")
number = numbered.group("number")
level = number.count(".") + 1
# The stop-word gate. Applied to the TITLE, after any section number
# has been split off, so `3.1 Og` is judged on `Og`.
if _is_stop_word_only(title):
continue
marked.append(
(
index,
Candidate(
title=title,
level=level,
number=number,
rule=RULE_HEADING,
start=offsets[index],
end=end_of_text,
),
)
)
candidates: list[Candidate] = []
# The name an orphaned heading leaves behind, and the ONE candidate allowed
# to pick it up.
#
# WHY THIS EXISTS. A table opening on the line below a heading gives that
# heading an empty body, so the orphan check drops it, and the table block
# keeps its mechanical `Tabell linje <n>`: the section's NAME is destroyed
# even though its content survives. Measured on a 629-concept bundle after
# a spreadsheet began rendering as pipe rows -- the concept fell from
# candidate rank 10 to 19 for a question naming its subject, and restoring
# the title alone put it back at 10. The name is not invented here, it is
# carried across: it moves to the segment that holds the heading's content.
#
# CONDITIONED ON THE DROP, deliberately. A heading that keeps its own body
# is still carried by a live candidate, so copying its title onto the table
# as well would put one name on two concepts and rescue none.
orphaned_name: tuple[str, str | None] | None = None
# D1. Which table blocks a heading ABSORBS, decided before the pass that
# consumes it: once a table is absorbed the heading's body is no longer
# empty, so the orphan check below stops firing on it by itself and no
# branch is needed there. Computed only when the caller asked, so every
# other arm's `marked` -> `candidates` mapping is untouched code.
absorbed = _absorbed_tables(text, marked, offsets, end_of_text) if keep_table_heading else set()
# Arm F clause 1's input, and it must be read HERE: the orphan pass below
# deletes every bodiless heading, which is every entry of a contents list
# but the last, and a run of one is below `CONTENTS_RUN`.
contents_run = _contents_run_positions(marked) if unit_fold else set()
for position_in_list, (_, candidate) in enumerate(marked):
if position_in_list in absorbed:
continue
following = [
entry
for position, entry in enumerate(marked)
if position > position_in_list and position not in absorbed
]
end = offsets[following[0][0]] if following else end_of_text
body = text[candidate.start : end]
# The orphan check: everything after the heading line itself.
#
# D3 IS EXEMPT, and the exception is stated rather than worked around:
# the check asks whether anything stands UNDER a candidate's first
# line, which is the right question for a heading and the wrong one for
# a row. A one-row section carries its content in its own cells, so
# every section but the last would be read as bodiless and deleted --
# the rule could not fire at all. It is the same shape as a contents
# list without dot leaders, and it is the reason that one needed
# `Candidate.contents`.
orphan = candidate.rule != RULE_SHEET_SECTION and (
not body.splitlines()[1:] or not "".join(body.splitlines()[1:]).strip()
)
if orphan:
orphaned_name = (candidate.title, candidate.number)
continue
inherited, orphaned_name = orphaned_name, None
title, number = candidate.title, candidate.number
if inherited is not None and candidate.rule == RULE_TABLE_BLOCK:
# Both members, because `_segment_path` reads both: the number
# becomes the directory AND is stripped from the stem, so carrying
# the title alone would emit a name the heading never had.
title, number = inherited
candidates.append(
Candidate(
title=title,
level=candidate.level,
number=number,
rule=candidate.rule,
start=candidate.start,
end=end,
grid=position_in_list in joined,
contents=position_in_list in contents_run,
)
)
return fold_units(candidates) if unit_fold else candidates
def _absorbed_tables(
text: str,
marked: list[tuple[int, Candidate]],
offsets: list[int],
end_of_text: int,
) -> set[int]:
"""D1: positions in `marked` of table blocks a bodiless heading keeps.
The condition is the orphan check's own, evaluated on the UNABSORBED
neighbour distance -- a heading is eligible only when the very next mark is
a table block and there is nothing between them but the heading line. That
is what keeps the variant from merging a section into a table it merely
contains: a heading with a paragraph of its own is not orphaned, so it
absorbs nothing.
"""
absorbed: set[int] = set()
for position, (_, candidate) in enumerate(marked):
if candidate.rule == RULE_TABLE_BLOCK or position + 1 >= len(marked):
continue
line_index, following = marked[position + 1]
if following.rule != RULE_TABLE_BLOCK:
continue
body = text[candidate.start : offsets[line_index]]
if body.splitlines()[1:] and "".join(body.splitlines()[1:]).strip():
continue
absorbed.add(position + 1)
return absorbed
def _contents_run_positions(marked: list[tuple[int, Candidate]]) -> set[int]:
"""Arm F clause 1, measured on the PRE-orphan list. Same predicate as there.
A run of at least `CONTENTS_RUN` consecutive page-numbered headings, table
blocks excluded and never a single line. Written once here and read by
`fold_units` through `Candidate.contents`, so the two cannot drift into two
definitions of what a contents list is.
LEVEL IS NOT PART OF THE PREDICATE, and that is the second round-2 change.
A numbered report's contents list interleaves `1`, `1.1`, `1.1.1`, `2`, so
requiring the members to be siblings breaks the run at every level change:
on K3 position 7 the level-2 entries formed runs long enough to discard and
`6.2.2 Tverrfaglig kontroll` did not, so one contents line was emitted as a
concept while its neighbours were not. What still bounds the rule is the run
LENGTH, which is what the `CONTENTS_RUN` sweep bought and is unchanged.
"""
inside: set[int] = set()
index = 0
while index < len(marked):
candidate = marked[index][1]
if candidate.rule == RULE_TABLE_BLOCK or not _TRAILING_PAGE_NUMBER.search(candidate.title):
index += 1
continue
end = index
while (
end < len(marked)
and marked[end][1].rule != RULE_TABLE_BLOCK
and _TRAILING_PAGE_NUMBER.search(marked[end][1].title)
):
end += 1
if end - index >= CONTENTS_RUN:
inside.update(range(index, end))
index = end
return inside
#: Arm F clause 1. How many CONSECUTIVE same-level page-numbered headings make
#: a contents list. Three, and it is swept rather than guessed: at 1 and 2 the
#: rule deletes body chapters (a heading like `... i henhold til TEK 17` ends
#: in a number and is not a contents line), and the sweep front is published in
#: `docs/2026-09-08-k3-arm-f-mot-enhetsarket.md`. Not a CLI knob: a number a
#: caller can turn without recording the sweep is a number nobody has measured.
CONTENTS_RUN = 3
def fold_units(candidates: list[Candidate]) -> list[Candidate]:
"""Arm F: ONE rule, three clauses, and it only MERGES or DISCARDS.
Derived from the three rules the operator wrote across the K3 unit
worksheet (2026-09-08), not from twelve special cases:
1. "innholdsfortegnelsen er ikke konsepter" -- a RUN of at least
`CONTENTS_RUN` consecutive same-level headings each ending in a page
number is a contents list and is discarded. A run of siblings, never a
single line: one body heading ending in a number is not a contents list.
The run is measured on this list AND on the list before the orphan check
(`Candidate.contents`), because a contents list whose entries carry no
dot leaders is bodiless and reaches here as one surviving line.
2. "hvert h2-kapittel med sine h3" -- the unit level is the SHALLOWEST
heading level occurring more than once; anything deeper folds into the
preceding candidate at or above that level, which EXTENDS the parent's
span rather than deleting the child's body.
3. "tabellen med innledningen" -- a table folds back into the heading
immediately before it when that heading's own span is shorter than the
table's. The surviving concept keeps the HEADING's name: merging the
right bytes under `Tabell linje 48` would produce a concept no reader
can look up.
It proposes no boundary of its own, so it can only ever reduce a plan. A
document with one heading level and no table comes out identical.
"""
if not candidates:
return candidates
# Clause 1. Two inputs, one predicate. `contents` carries the run measured
# on the list BEFORE the orphan check thinned it -- a contents list with no
# dot leaders is a run of bodiless headings, so only its last entry reaches
# here and a run of one is below `CONTENTS_RUN`. The scan below still has
# work to do: a contents list WITH dot leaders keeps every entry, and that
# run exists only in this list.
drop: set[int] = {position for position, c in enumerate(candidates) if c.contents}
index = 0
while index < len(candidates):
candidate = candidates[index]
if candidate.rule == RULE_TABLE_BLOCK or not _TRAILING_PAGE_NUMBER.search(candidate.title):
index += 1
continue
end = index
while (
end < len(candidates)
and candidates[end].rule != RULE_TABLE_BLOCK
and candidates[end].level == candidate.level
and _TRAILING_PAGE_NUMBER.search(candidates[end].title)
):
end += 1
if end - index >= CONTENTS_RUN:
drop.update(range(index, end))
index = end
kept = [c for position, c in enumerate(candidates) if position not in drop]
if not kept:
return kept
# Clause 2. The unit level is read from DECLARED headings only -- ATX and
# dotted-numbered, `RULE_HEADING`. Two exclusions, each measured:
#
# - a table block carries the sentinel level 9, and letting it vote would
# make every document with two tables a two-level document;
# - `RULE_OUTLINE` is Arm D's RECOVERY of an integer numbering run, not a
# level the document declares. Letting it vote made it the shallowest
# repeated level on every PDF that has both, folding the dotted headings
# the operator actually named into it: 23 -> 3, 48 -> 4 and 11 -> 7 on
# K3 positions 1, 7 and 9. The unit worksheet showed the operator ATX and
# dotted headings and nothing else, which is the same set.
#
# With no declared heading at all the fold has no level to work from, and
# `unit_level` is set below every candidate's, so clause 2 is inert and
# only clause 3 can fire. That is the honest behaviour: a document whose
# structure was recovered rather than declared has no unit level to read.
levels = [c.level for c in kept if c.rule == RULE_HEADING]
repeated = sorted({level for level in levels if levels.count(level) > 1})
if repeated:
unit_level = repeated[0]
elif levels:
unit_level = min(levels)
else:
unit_level = -1
folded: list[Candidate] = []
for candidate in kept:
table = candidate.rule == RULE_TABLE_BLOCK
deeper = candidate.rule == RULE_HEADING and candidate.level > unit_level
if table and folded:
previous = folded[-1]
introduces = (
previous.rule != RULE_TABLE_BLOCK
and previous.end - previous.start < candidate.end - candidate.start
)
else:
introduces = False
if (deeper or introduces) and folded:
previous = folded[-1]
folded[-1] = replace(previous, end=candidate.end)
continue
folded.append(candidate)
return folded
def _cut_points(text: str, start: int, end: int, cap: int) -> list[int]:
"""Where to cut `text[start:end]` so no part exceeds `cap` characters.
The cut prefers a PARAGRAPH boundary (a blank line) inside the window, then
a line boundary, and only then cuts mid-line. The order is the whole
content of the rule: a cut that lands mid-sentence splits one unit of
knowledge for no reason other than arithmetic, and the K3 categories count
that as `too fine`. The last resort exists anyway, because a document whose
body is one unbroken line is exactly where a cap that quietly stopped
binding would be least defensible.
"""
cuts: list[int] = []
position = start
while end - position > cap:
window_end = position + cap
paragraph = text.rfind("\n\n", position, window_end)
if paragraph != -1:
cut = paragraph + 2
else:
line = text.rfind("\n", position, window_end)
cut = line + 1 if line != -1 else window_end
# rfind can only return an index at or after `position`, so every
# branch advances. The assertion states that rather than trusting it:
# a cut that did not advance would loop forever on a corpus run.
assert cut > position, f"cut {cut} did not advance past {position}"
cuts.append(cut)
position = cut
return cuts
def subdivide(text: str, candidates: list[Candidate], cap: int) -> list[Candidate]:
"""Arm C. Arm B's candidates, with every over-long span cut down to `cap`.
ARM C IS NOT DEFINED IN `docs/2026-09-02-k3-k4-k5-metode.md`; that file
contains no occurrence of the word. This definition was written for order
20260904T145630Z and is reported as the author's, not as a ratified one.
Two callers' cases, one rule. When Arm B found boundaries but a span still
runs long (a PDF whose headings are its table of contents, so the trailing
segment absorbs the body), the span is cut. When Arm B found NO boundary at
all, the whole document is that span -- which is the `no declared
structure` case § 10 names, and 23 of 33 PDFs in the K2 corpus are in it.
A document with no boundaries that is already under the cap proposes
NOTHING, exactly as Arm B does. Arm C fires on size; where size is not the
problem it has nothing to say, and a one-entry plan would only dress a
single concept in a plan file.
"""
if cap <= 0:
return candidates
if not candidates:
if len(text) <= cap:
return []
# The synthetic span. Its rule is the size rule alone, because no
# heading rule proposed it -- there was no heading.
candidates = [
Candidate(
title="Del",
level=1,
number=None,
rule=RULE_SIZE_SPLIT,
start=0,
end=len(text),
split=False,
)
]
unnumbered_parts = True
else:
unnumbered_parts = False
out: list[Candidate] = []
for candidate in candidates:
cuts = _cut_points(text, candidate.start, candidate.end, cap)
if not cuts:
out.append(candidate)
continue
edges = [candidate.start, *cuts, candidate.end]
for part, (start, end) in enumerate(zip(edges, edges[1:]), start=1):
if unnumbered_parts:
title = f"Del {part}"
else:
title = candidate.title if part == 1 else f"{candidate.title} (del {part})"
out.append(
Candidate(
title=title,
level=candidate.level,
number=candidate.number,
rule=candidate.rule,
start=start,
end=end,
split=True,
grid=candidate.grid,
)
)
return out
def _derived_names(candidate: Candidate) -> list[str]:
"""The `derived` list for one candidate, in the order a test pins.
The marker, the rule that OPENED the span, then any rule that reshaped it:
Arm E's join before Arm C's cut, because a joined span is what Arm C would
then have been given to cut.
"""
names = [PROPOSED_MARKER, candidate.rule]
if candidate.grid:
names.append(RULE_TABLE_GRID)
if candidate.split and candidate.rule != RULE_SIZE_SPLIT:
names.append(RULE_SIZE_SPLIT)
return names
def _segment_path(candidate: Candidate, taken: set[str], prefix: str = "") -> str:
title = unicodedata.normalize("NFC", candidate.title)
# The section number becomes the DIRECTORY, so leaving it in the stem too
# yields `3-1/3-1-brannkonsept.md` -- correct and unreadable.
if candidate.number and title.startswith(candidate.number):
title = title[len(candidate.number) :]
stem = reduce_to_id_grammar(title)
if not stem:
stem = "seksjon"
directory = reduce_to_id_grammar(candidate.number or "") if candidate.number else ""
# The caller's scope comes FIRST and is never deduplicated against: it is
# the same for every entry in this document by construction, and that is
# the whole point -- one document's sections must not be able to claim
# another's path.
head = f"{prefix}/" if prefix else ""
path = f"{head}{directory}/{stem}.md" if directory else f"{head}{stem}.md"
suffix = 2
while path in taken:
path = f"{head}{directory}/{stem}-{suffix}.md" if directory else f"{head}{stem}-{suffix}.md"
suffix += 1
taken.add(path)
return path
def build_plan(
source: Path,
text: str,
source_bytes: bytes,
*,
okf_type: str,
proposed_at: str,
path_prefix: str = "",
max_segment_chars: int = 0,
outline_run: int = 0,
table_grid: bool = False,
unit_fold: bool = False,
keep_table_heading: bool = False,
sheet_section_rows: bool = False,
drop_wrapped_outline: bool = False,
) -> dict[str, Any]:
"""The artifact. Every entry PROPOSED, the plan itself never adjudicated."""
taken: set[str] = set()
extractor_id = source.suffix.lower().lstrip(".") or "none"
entries: list[dict[str, Any]] = []
candidates = find_candidates(
text,
outline_run=outline_run,
table_grid=table_grid,
unit_fold=unit_fold,
keep_table_heading=keep_table_heading,
sheet_section_rows=sheet_section_rows,
drop_wrapped_outline=drop_wrapped_outline,
)
for candidate in subdivide(text, candidates, max_segment_chars):
entries.append(
{
"segment_id": f"p{len(entries) + 1}",
"path": _segment_path(candidate, taken, path_prefix),
"title": candidate.title,
"okf_type": okf_type,
"span": [candidate.start, candidate.end],
"ingested_at": proposed_at,
# The offsets are a hint the anchor may correct. Written at
# proposal time because that is the only moment the text the
# adjudicator will judge and the offsets naming it are known
# to agree -- reconstructing it later would anchor to whatever
# the extraction had already become.
"anchor": {
"quote": text[candidate.start : candidate.end],
"prefix": text[max(0, candidate.start - ANCHOR_CONTEXT) : candidate.start],
"suffix": text[candidate.end : candidate.end + ANCHOR_CONTEXT],
},
# PROPOSED first, then the rule that proposed it. `derived` is
# this library's existing "which of these did we infer" marker,
# so a consumer that already distrusts derived fields
# distrusts these by construction.
# PROPOSED first, then the rule that OPENED the span, then --
# for an Arm E join -- the grid rule, then -- for an Arm C part
# only -- the size rule that cut it. More than one name on those
# entries because the rule that opened the span is still true of
# them, and dropping it would leave a part traceable to nothing
# but arithmetic or nothing but a join.
#
# Written as an ordered build rather than a conditional
# expression: four combinations exist now, and the order is
# itself a claim a test pins.
"derived": _derived_names(candidate),
}
)
return {
"version": "1",
"source_sha256": hashlib.sha256(source_bytes).hexdigest(),
# The hash the offsets actually depend on. Source bytes alone cannot
# see a converter reshaping its output, so the staleness signal this
# plan is supposed to carry did not exist until this line did.
"text_sha256": hashlib.sha256(text.encode("utf-8")).hexdigest(),
"extractor_id": extractor_id,
# The EXTRACTOR's version, not this tool's. `PROPOSER_VERSION` sat here
# and named the wrong thing: a converter bump left the field frozen at
# the proposer's own number, so the component could not move.
"extractor_version": observed_extractor_version(extractor_id),
"adjudicated_at": proposed_at,
# NOT a timestamp question. `adjudicated_at` records when this artifact
# was produced; this records whether a human has looked at it, and it is
# false until one replaces the file.
"adjudicated": False,
"proposed_by": f"{PROPOSER_ID}/{PROPOSER_VERSION}",
"entries": entries,
}
def run(
source: Path,
out: Path,
*,
okf_type: str,
proposed_at: str,
path_prefix: str = "",
max_segment_chars: int = 0,
outline_run: int = 0,
table_grid: bool = False,
unit_fold: bool = False,
keep_table_heading: bool = False,
sheet_section_rows: bool = False,
drop_wrapped_outline: bool = False,
pdf_headings: bool = False,
pdf_headings_reserve: bool = False,
ocr: bool = False,
) -> int:
if max_segment_chars < 0:
raise ProposerError(
f"--max-segment-chars {max_segment_chars} is negative; the cap is a "
"character count, and 0 means off (Arm B)"
)
if outline_run < 0:
raise ProposerError(
f"--outline-run {outline_run} is negative; the gate is a run LENGTH, "
"and 0 means off (Arm B)"
)
# Reduced HERE, before anything is read: a prefix that survives to the
# entries as an empty component would produce exactly the unscoped paths
# the caller asked to avoid, and would do it silently.
#
# PER COMPONENT, because the prefix carries a DIRECTORY now that Door B
# walks the inbox recursively and records a relative `source_file`.
# Reducing the whole string would fold `/` into a `-` and flatten
# `sub/sub2` into the single component `sub-sub2` -- a bundle shaped unlike
# the inbox it came from, and unlike what the caller wrote.
components = (
[reduce_to_id_grammar(part) for part in path_prefix.split("/")] if path_prefix else []
)
if path_prefix and not all(components):
raise ProposerError(
f"--path-prefix {path_prefix!r} has a component that reduces to nothing under "
"the id grammar ([a-z0-9][a-z0-9-]*); refusing to write unscoped paths under a "
"scope that was asked for"
)
scope = "/".join(components)
if not source.is_file():
raise ProposerError(f"source is not a file: {source}")
try:
source_bytes = source.read_bytes()
except OSError as exc:
raise ProposerError(f"cannot read {source}: {exc}") from exc
try:
# The two READER options, not arms: they change what the extraction
# says, and every arm below reads whatever it says. Passed here as well
# as to the run path because the plan's `text_sha256` indexes this
# exact string -- a plan proposed against one rendering and replayed
# against another is refused by `assert_plan_applies`, which is the
# right outcome and a confusing one to debug.
text = extract_text(source.name, source_bytes, pdf_headings=pdf_headings, ocr=ocr)
# The reserve, and the reason it re-extracts rather than post-processes:
# the font reader works on the PDF's glyph geometry, which the joined
# text no longer carries. Skipped outright when the font reader is
# already on -- `font` and `font-reserve` are two values of one option,
# never a pair to combine.
if pdf_headings_reserve and not pdf_headings:
if heading_reserve_applies(text, outline_run=outline_run):
text = extract_text(source.name, source_bytes, pdf_headings=True, ocr=ocr)
except IngestError as exc:
raise ProposerError(f"cannot extract text from {source.name}: {exc}") from exc
payload = build_plan(
source,
text,
source_bytes,
okf_type=okf_type,
proposed_at=proposed_at,
path_prefix=scope,
max_segment_chars=max_segment_chars,
outline_run=outline_run,
table_grid=table_grid,
unit_fold=unit_fold,
keep_table_heading=keep_table_heading,
sheet_section_rows=sheet_section_rows,
drop_wrapped_outline=drop_wrapped_outline,
)
# Nothing to propose is an OUTCOME, and it is not an artifact. An empty
# plan cannot be replayed -- `process_inbox` refuses one, because a plan
# naming no entry would persist nothing for a document that was dropped --
# so the only thing a zero-entry file can do is fail a run later. Its own
# exit status, distinct from 2, so a driver can tell "this document lands
# as one flat concept" from "stop".
if not payload["entries"]:
print(
f"{PROPOSER_ID}: nothing to propose for {source.name} — the mechanical "
"rules found no boundary. No artifact written; this document lands as "
"one concept unless someone segments it by hand.",
file=sys.stderr,
)
return 1
out.parent.mkdir(parents=True, exist_ok=True)
out.write_bytes((json.dumps(payload, indent=2, ensure_ascii=False) + "\n").encode("utf-8"))
print(
f"{PROPOSER_ID}: proposed {len(payload['entries'])} segment(s) -> {out}\n"
f"{PROPOSER_ID}: every entry is PROPOSED. Adjudicate before ingesting.",
file=sys.stderr,
)
return 0
def parse_args(argv: list[str] | None) -> argparse.Namespace:
parser = argparse.ArgumentParser(
prog=PROPOSER_ID,
description="Propose a segmentation plan. A human adjudicates it before use.",
)
parser.add_argument("source", type=Path, help="the document to segment")
parser.add_argument("--out", type=Path, required=True, help="where to write the artifact")
parser.add_argument("--okf-type", default="reference", help="okf_type for every entry")
parser.add_argument(
"--path-prefix",
default="",
help=(
"scope every entry's path under this directory, `/`-separated for a "
"nested one (each component is reduced on its own). Required for a corpus: "
"section numbering is document-local, so two documents propose the same "
"path and Door B refuses both. An argument rather than something this "
"tool derives -- it sees one document and cannot know what else is in "
"the bundle"
),
)
parser.add_argument(
"--max-segment-chars",
type=int,
default=0,
metavar="N",
help=(
"Arm C: cut any proposed span longer than N characters at the nearest "
"paragraph boundary, the whole document counting as one span when the "
"mechanical rules find no boundary at all. 0 (the default) is OFF and "
"leaves the artifact byte-identical to Arm B. Arm C is the author's "
"definition, written for order 20260904T145630Z; it is not defined in "
"the K3 method file"
),
)
parser.add_argument(
"--outline-run",
type=int,
default=0,
metavar="N",
help=(
"Arm D: also propose a boundary at each line of the document's own "
"numbered outline (the bare integers the heading grammar cannot "
"match, since it requires a dot), but only where those integers "
"sustain an ascending run of at least N entries, and only for the "
"LAST such run when the outline repeats, because a contents listing "
"precedes the body it lists. 0 (the default) is OFF and leaves the "
"artifact byte-identical to Arm B. Arm D is the author's definition, "
"written for order 20260906T213322Z; it is not defined upstream, and "
"the K3 method file does not name it either"
),
)
parser.add_argument(
"--unit-fold",
action="store_true",
help=(
"Arm F: fold the proposed candidates into the operator's units --"
" discard a run of contents-list headings, fold a deeper heading"
" into its parent, and fold a table back into the shorter heading"
" that introduces it. ONE rule with three clauses, derived from the"
" three rules the K3 unit worksheet records (2026-09-08). It adds"
" no boundary, so it can only reduce a plan. Absent (the default)"
" is OFF and leaves the artifact byte-identical to the arm below"
" it. A boolean: the rule's one number, CONTENTS_RUN, is a module"
" constant whose sweep is published, not a knob a caller can turn"
),
)
parser.add_argument(
"--table-grid",
action="store_true",
help=(
"Arm E: do not let a pandoc grid-table rule line (`+---+---+`, and "
"`+===+===+` under a header) close an open table block. The table "
"grammar cannot match a rule line, so without this one grid table "
"becomes one concept per row group. Absent (the default) is OFF and "
"leaves the artifact byte-identical to Arm D (Arm D rather than "
"Arm B, because Arm E is defined on top of it). A boolean and not a "
"number: the rule has no parameter to sweep. Arm E is the author's "
"definition, written for order 20260907T075834Z-18584396-from-.claude; "
"it is not defined upstream, and the K3 method file does not name it "
"either"
),
)
parser.add_argument(
"--keep-table-heading",
action="store_true",
help=(
"D1: keep a heading whose body is empty ONLY because a table block "
"opens under it, and absorb that table into the heading's span "
"instead of emitting it. Without this the orphan check drops the "
"heading, the table inherits its NAME, and the concept starts at "
"the first table row -- so the heading line is in no concept's "
"body. Measured on a spreadsheet: the concept count does not move "
"(1 -> 1), the first byte does. Absent (the default) is OFF and "
"leaves every artifact byte-identical. Its own flag rather than "
"part of an arm: the orphan check is reached by every file type"
),
)
parser.add_argument(
"--sheet-section-rows",
action="store_true",
help=(
"D3: cut an open table block at the rows that label its sections. "
"A section row is one of a RUN of at least SHEET_SECTION_RUN "
"consecutive rows whose first cell is a bare numeric label and "
"which carry at least one other non-empty cell; each opens a "
"candidate reaching the next section row or the end of the block. "
"The opposite direction from Arm E, which stops a grid rule line "
"from CLOSING a block: that arm decides how far a block extends, "
"this rule where it is cut inside, and they compose in that order. "
"Written for a spreadsheet whose whole body is one table block and "
"whose units are rows. Absent (the default) is OFF and leaves every "
"artifact byte-identical. A boolean: the run length is a module "
"constant, not a knob a caller can turn. D3 is the author's "
"definition, written for order "
"20260908T170037Z-3622420612-from-.claude; it is not defined "
"upstream, and the K3 method file does not name it either"
),
)
parser.add_argument(
"--drop-wrapped-outline",
action="store_true",
help=(
"D3: do not admit an outline candidate whose line continues onto "
"the next one, because a wrapped sentence is not a heading. Judges "
"RECOVERED candidates only, never a heading the document declares "
"for itself. Written for quoted regulation text, whose numbered "
"paragraphs match Arm D's grammar exactly; a TITLE LENGTH rule was "
"tried first and falsified, because a real 88-character heading "
"sits between the quoted sentences at 86 and 91. Absent (the "
"default) is OFF and leaves every artifact byte-identical. It is "
"the author's definition, written for order "
"20260908T170037Z-3622420612-from-.claude; it is not defined "
"upstream, and the K3 method file does not name it either"
),
)
parser.add_argument(
"--proposed-at",
default="1970-01-01T00:00:00Z",
help="the timestamp written into the artifact; explicit so a run is reproducible",
)
return parser.parse_args(argv)
def main(argv: list[str] | None = None) -> int:
args = parse_args(argv)
try:
return run(
args.source,
args.out,
okf_type=args.okf_type,
proposed_at=args.proposed_at,
path_prefix=args.path_prefix,
max_segment_chars=args.max_segment_chars,
outline_run=args.outline_run,
table_grid=args.table_grid,
unit_fold=args.unit_fold,
keep_table_heading=args.keep_table_heading,
sheet_section_rows=args.sheet_section_rows,
drop_wrapped_outline=args.drop_wrapped_outline,
)
except ProposerError as exc:
print(f"{PROPOSER_ID}: FAILED - {exc}", file=sys.stderr)
print(
f"{PROPOSER_ID}: this is NOT 'nothing to propose'. Nothing was written.",
file=sys.stderr,
)
return 2
if __name__ == "__main__":
raise SystemExit(main())