#!/usr/bin/env python3 """Propose a segmentation plan for one document. A human adjudicates it. Pipeline step 3. It lives in the package because `okf build` has to reach it from an INSTALLED copy, where `tools/` does not exist -- it was outside `src/` until then, and the reason it could move is that the reason it sat outside was never about dependencies: every rule below is mechanical, so nothing here adds a model call to a package that promises none. What DID have to survive the move is the separation the old location expressed physically: the split of a document into units of knowledge is a judgement, and the run path replays a decision somebody already made. That separation is carried by `adjudicated: false` and the `PROPOSED` marker, which is where it belonged all along -- a directory boundary cannot enforce it, and a caller who ingests a proposal unadjudicated could always do so. ## What the research says this tool may and may not claim Topic 2 measured the OKF reference agent's granularity criteria against `_okf-canonical`: it splits on **what a thing is**, not on layout, and makes "multiple `write_concept_doc` calls ... rather than dumping everything into one doc". Four of its gates are semantic and need a model. A handful of MECHANICAL rules port today, and those are the ones below. Topic 1b measured heading derivation on the K2 corpus: 11 of 11 prose headings recovered -- from ONE document. 23 of 33 PDFs carry no outline at all and 95 % of the outline entries that do exist are AutoCAD export metadata. The denominator is 1. A rule validated on n=1 is not validated, and this tool says so by marking every entry it emits `PROPOSED` rather than adjudicated. Topic 1a measured that the best deterministic heading rule from poppler is a CONJUNCTION -- `size AND bold`, via `-fontfullname` -- at recall 1.000 and precision 0.846, and that adding weight as a DISJUNCT makes precision worse (0.786 -> 0.524). That path is implemented here and nowhere else: poppler is a SYSTEM binary the `[extract]` extra cannot express, so it may never be on the run path or in a golden fixture. ## The one rule that is not a heuristic **Nothing here is ever adjudicated.** `adjudicated: false` sits at the top of every artifact and `PROPOSED` in every entry's `derived` list. A plan is replayed deterministically and forever by the run path, so a proposal that could pass for an adjudication would put a machine's guess where a human's judgement is supposed to be, permanently and silently. Stdlib only. No network: the model-backed path this tool deliberately does not have would need the per-run network opt-in, and the socket-free test suite proves the absence rather than assuming it. """ from __future__ import annotations import argparse import hashlib import json import re import sys import unicodedata from dataclasses import dataclass, replace from pathlib import Path from typing import Any from .errors import IngestError from .extract import extract_text from .materialize import reduce_to_id_grammar from .segmentation import observed_extractor_version #: Stamped into every entry's `derived` list. The marker is what keeps a #: proposal from being mistaken for the judgement the run path replays. PROPOSED_MARKER = "PROPOSED" #: This tool's identity, written into the artifact so an operator reading a #: plan six months later can tell what produced it. PROPOSER_ID = "okf-propose-segments" PROPOSER_VERSION = "1" #: The rules that survived Topic 2's port test. Every entry names the rule that #: OPENED its span, so a proposal an operator disagrees with is traceable to the #: rule that made it rather than to the tool as a whole. An entry an arm later #: reshaped names that arm too -- Arm C's cut, Arm E's join -- because the rule #: that opened the span is still true of it. Exactly one name per entry held #: until the arms existed; it is the ORIGIN that is single, not the list. RULE_HEADING = "rule:heading" RULE_TABLE_BLOCK = "rule:table-block" RULE_POPPLER_SIZE_AND_BOLD = "rule:poppler-size-and-bold" #: Arm C only. NOT one of Topic 2's ported rules and not a heading rule at #: all: it names the fact that a span was cut because it was too long, which #: is a judgement about SIZE and says nothing about where a unit of knowledge #: begins. It is emitted ALONGSIDE the rule that proposed the origin span, so #: an operator reading a part can still see what opened it. RULE_SIZE_SPLIT = "rule:size-split" #: Arm D only. Like Arm C it is NOT one of Topic 2's ported rules and NOT #: defined upstream: `docs/2026-09-02-k3-k4-k5-metode.md` contains no #: occurrence of the word "arm" at all, so this definition was written for the #: brief of order 20260906T213322Z and is reported as the author's. Unlike Arm #: C it says nothing about size -- it names the fact that the DOCUMENT ITSELF #: declared a chapter there, by numbering it in an ascending run its own #: outline sustains. RULE_OUTLINE = "rule:outline" #: Arm E only. Like Arm C and Arm D it is NOT one of Topic 2's ported rules and #: NOT defined upstream -- `docs/2026-09-02-k3-k4-k5-metode.md` contains no #: occurrence of the word "arm" at all -- so this definition was written for #: order 20260907T075834Z-18584396-from-.claude and is reported as the author's. #: Its axis is a third one. Arm C names SIZE and Arm D names what the DOCUMENT #: declared; this names what the CONVERTER emitted: a table block that was #: joined across a grid-table rule line. Emitted ALONGSIDE `rule:table-block`, #: which is still what opened the span, and only on a block that was ACTUALLY #: joined -- never on one whose span merely happens to contain a rule line, so #: a single-row grid table stays byte-identical to Arm D. RULE_TABLE_GRID = "rule:table-grid" #: D3 round 3 only, and it is the first rule in this module that opens a span #: INSIDE a table rather than at one. Like Arm C, D and E it is NOT one of Topic #: 2's ported rules and NOT defined upstream -- `docs/2026-09-02-k3-k4-k5-metode.md` #: contains no occurrence of the word "arm" at all -- so this definition was #: written for order 20260908T170037Z-3622420612-from-.claude and is reported as #: the author's. Its axis is a fourth one. Arm C names SIZE, Arm D what the #: DOCUMENT declared, Arm E what the CONVERTER emitted; this names what the #: SHEET labelled: a row whose first cell is a bare numeric label, in a run of #: such rows. It is the opposite DIRECTION from Arm E -- that one stops a rule #: line from closing a block so a grid table proposes one candidate instead of #: many, this one cuts an open block at the rows that label its sections -- and #: the two compose in one order: Arm E decides how far a block extends, this #: decides where it is cut inside. RULE_SHEET_SECTION = "rule:sheet-section" RULE_NAMES = ( RULE_HEADING, RULE_TABLE_BLOCK, RULE_POPPLER_SIZE_AND_BOLD, RULE_SIZE_SPLIT, RULE_OUTLINE, RULE_TABLE_GRID, RULE_SHEET_SECTION, ) #: How many characters of context each side of a quote anchor carries. Enough #: to separate two occurrences of a repeated heading, short enough that an #: edit NEAR a segment does not invalidate the anchor FOR it -- the anchor #: exists to survive shifts, so making it fragile would defeat it. ANCHOR_CONTEXT = 48 #: Norwegian and English function words. A heading made only of these names no #: unit of knowledge -- it is a connective that happened to sit on its own line. #: Topic 2's stop-word gate, and the only place this tool judges wording. STOP_WORDS = frozenset( { "and", "as", "at", "av", "be", "by", "da", "de", "den", "der", "det", "en", "er", "et", "for", "fra", "i", "in", "is", "it", "med", "of", "og", "om", "on", "or", "over", "paa", "som", "til", "the", "to", "under", "ved", "with", } ) # An ATX heading, or a numbered section opening a line (`3.1 Brannkonsept`). # A BARE integer is not a section number, for the same reason `structure.py` # refuses one: `12 ting` is an ordinary line and admitting it would cut a # document at every list item. # # That claim still holds, and Arm D does not weaken it. `_OUTLINE` below admits # a bare integer ONLY inside an ascending run the document sustains for at # least a declared length -- which is a property of the whole text, not of the # line -- and the rule is off unless a caller asks for it. An UNGATED widening # was measured and rejected: 1681 raw hits against 618 candidates, admitting # list items, quantities and page furniture. The gate is what makes the signal # a signal. _ATX = re.compile(r"^(?P#{1,6})\s+(?P\S.*?)\s*$") _NUMBERED = re.compile(r"^(?P<number>\d+(?:\.\d+)+)\s+(?P<title>\S.*?)\s*$") _TABLE_ROW = re.compile(r"^\s*\|.*\|\s*$") # Arm E's grammar: a pandoc GRID-table rule line. The converter separates a grid # table's rows with `+---+---+`, and its header from its body with `+===+===+`. # Neither matches `_TABLE_ROW`, so `in_table` is reset between every pair of rows # and ONE table becomes one candidate per row group. Measured on the K2 corpus: # three documents carry grid tables, and they account for 33 of the 709 entries # Arm D proposes. # # The character class is measured rather than guessed. Across those three # documents, 38 of 38 lines whose stripped form starts with `+` match this # pattern, and `+`, `-`, `:`, `=` is the COMPLETE set of characters occurring on # them. The `:` is pandoc's column-alignment marker, and it is not decoration: a # first pass with `[-=+]` matched 37 of 38 and, through that one miss, read one # document as having two tables where it has one. # # `\s*` on both ends mirrors `_TABLE_ROW` rather than tightening on it, because # the loop iterates `splitlines(keepends=True)` -- every line carries its `\n`, # and an indented rule line is a real shape that must still be admitted. _GRID_RULE = re.compile(r"^\s*\+[-=:+]+\+\s*$") # Arm D's grammar. Integer-only BY CONSTRUCTION: `\s+` after the optional # separator is what keeps `1.1 Brannkonsept` out, because `_NUMBERED` requires # a dot and this requires whitespace, so no line can match both. No exclusion # clause is written for that: a filter with a measured effect of zero is dead # code that reads like a guard. _OUTLINE = re.compile(r"^\s{0,4}(?P<number>\d{1,2})[.)]?\s+(?P<title>\S.*?)\s*$") # A contents line carries the page it points at (`Innledning 6`). Measured on # the K2 corpus: stripping it changes 0 of the 144 outline counts and 9 emitted # titles. It is load-bearing anyway, because titles become concept paths # through `_segment_path` -- an unstripped page number would become part of a # filename. _TRAILING_PAGE_NUMBER = re.compile(r"[\s.]+\d{1,4}\s*$") # D3's grammar, and it reads a CELL rather than a line. A sheet's section label # is a bare number, optionally joined to another by a separator where two groups # were merged (`11+12`), and nothing else: no letters, so a row opening with a # word is not a section, and no word list, so the rule knows nothing about which # numbers any real sheet uses. _SHEET_SECTION_LABEL = re.compile(r"^\d+(?:[+./-]\d+)*$") # A pipe that pandoc did not escape. Splitting a row on a bare `|` would cut a # cell containing a literal pipe in half and misread the FIRST cell of the row # after it, which is the only cell this rule judges. _UNESCAPED_PIPE = re.compile(r"(?<!\\)\|") class ProposerError(Exception): """The run failed. NOT 'nothing to propose' -- the two must stay distinct.""" @dataclass(frozen=True) class Candidate: """One proposed boundary, before it becomes an entry.""" title: str level: int number: str | None rule: str start: int end: int #: True when this candidate is one PART of a longer span that Arm C cut. #: Kept on the candidate rather than recomputed at write time so the entry #: and the reason it exists cannot drift apart. split: bool = False #: True when Arm E JOINED this table block across at least one grid-rule #: line. Same reasoning as `split`, and the same trap: both reconstruction #: sites below rebuild a `Candidate` from an explicit keyword list, so a #: field not copied there is silently defaulted back and the entry loses #: the only trace of why it exists. NOT set merely because a span contains #: a rule line -- a single-row grid table joins nothing and stays #: byte-identical to Arm D. grid: bool = False #: True when this candidate belongs to a run of page-numbered siblings that #: was a contents list BEFORE the orphan check thinned it. Set only when #: Arm F is on, because it exists only for Arm F's clause 1 to read: a #: contents list without dot leaders is a run of bodiless headings, so the #: orphan check deletes all but the last and the run clause 1 looks for is #: gone by the time `fold_units` sees the list. Computed where the whole #: pre-orphan list is still in hand, and nowhere else -- no candidate #: carrying it survives clause 1, so it never reaches an artifact. contents: bool = False def _is_stop_word_only(title: str) -> bool: words = [word for word in re.split(r"[^\w]+", title.lower()) if word] return bool(words) and all(word in STOP_WORDS for word in words) def _strip_page_number(title: str) -> str: """Remove a trailing page number from a contents-listing title. Deliberately NOT applied to a title that is only digits: `477` has no separator before the number, so the pattern cannot match it and the title survives for the stop-word and junk paths to see. Emptying it would fall back to the `seksjon` stem and dress junk as a named section. """ return _TRAILING_PAGE_NUMBER.sub("", title) def outline_lines(text: str) -> list[tuple[int, int, str]]: """Every line the outline grammar admits, as `(line index, integer, title)`. Module level and importable on purpose: the reach instrument measures this rule, and an instrument that re-implements the grammar it measures is measuring a second definition that can silently drift from the shipped one. """ found: list[tuple[int, int, str]] = [] for index, line in enumerate(text.splitlines()): match = _OUTLINE.match(line) if match is None: continue title = _strip_page_number(match.group("title")).strip() if not title or _is_stop_word_only(title): continue found.append((index, int(match.group("number")), title)) return found def heading_reserve_applies(text: str, *, outline_run: int) -> bool: """Whether this text needs a SECOND heading source, having no run of its own. The font reader's reserve condition, and the only place it is decided. The proposer and the door both call this, because a plan indexes the exact string it was proposed against: a reserve that fired on one side and not the other would make every document it touched a coded rejection. It reads the gate AS CONFIGURED rather than a fixed minimum -- raising the arm's threshold widens the reserve, which is the same document property seen through the same threshold. At `outline_run` 0 the gate admits nothing at all, so the reserve is unconditional; that combination is round 4's "font instead of Arm D", measured at 1 of 8, and a caller reaching it gets it deliberately. """ if outline_run <= 0: return True return not outline_runs(outline_lines(text), outline_run) def outline_runs( entries: list[tuple[int, int, str]], minimum: int ) -> list[list[tuple[int, int, str]]]: """The maximal ascending runs among `entries`, each at least `minimum` long. A run is anchored at `1` and every later member is its predecessor plus one; a number that is neither is skipped without closing the run, so a stray page number between two chapters does not truncate the outline. A new `1` closes the current run and opens the next, which is what makes a contents listing and the body it lists two runs rather than one. Returned in document order. The CALLER chooses among them -- last-run selection was measured against the alternatives and is stated where it is applied, not hidden in here. """ runs: list[list[tuple[int, int, str]]] = [] current: list[tuple[int, int, str]] = [] for entry in entries: number = entry[1] if number == 1: if current: runs.append(current) current = [entry] elif current and number == current[-1][1] + 1: current.append(entry) if current: runs.append(current) return [run for run in runs if len(run) >= minimum] #: D3. How many CONSECUTIVE numbered rows make a sectioned table. Three, and it #: is the same bounding device -- and the same number -- `CONTENTS_RUN` uses, #: for the same reason: a single numbered row is a stated quantity, not a #: section, and a rule that read one would cut a sheet at every computation #: basis. NOT swept on a reference, because no reference exists for its value; #: the corpus SENSITIVITY of the choice is published instead, in #: `docs/2026-09-08-k3-runde3-per-filtype.md`. Not a CLI knob for the reason #: `CONTENTS_RUN` is not one. SHEET_SECTION_RUN = 3 def _wraps_onto_next_line(lines: list[str], index: int) -> bool: """True when the line at `index` is a sentence that continues below it. The test is the next line's first character being lower case. A heading is a complete line -- the line under it opens a new sentence, is blank, or is a bullet -- while a hard-wrapped paragraph carries its own continuation. This is deliberately NOT a length test. Round 2 sorted every outline title in the K3 sample and measured the two classes overlapping: a real chapter heading of 88 characters against quoted sentences of 86, 91, 92 and 100, so no threshold separates them. Measured on the same sample, this axis separates 8 of 34 -- the four quoted regulation paragraphs and four risk table rows, and none of the 26 headings the operator kept. """ if index + 1 >= len(lines): return False following = lines[index + 1].strip() return bool(following) and following[0].islower() def _sheet_section_rows(lines: list[str]) -> dict[int, tuple[str, str]]: """D3: line index -> (label, row text) for every section row, or empty. A section row is a table row whose FIRST cell is a bare numeric label and which carries at least one other non-empty cell -- the label alone names nothing, and the first non-empty cell after it is what the section is called. The rows must come in a run of at least :data:`SHEET_SECTION_RUN` consecutive lines: that is what separates a labelled section column from a quantity stated on its own row, and it is the module's existing guard rather than a new one. Computed over the whole line list before the marking loop, for the reason the outline runs are: a run is a property of the text, and a forward scan that decided one row at a time could not know whether the run it is inside is long enough. """ labelled: dict[int, tuple[str, str]] = {} for index, line in enumerate(lines): if not _TABLE_ROW.match(line): continue cells = [cell.strip() for cell in _UNESCAPED_PIPE.split(line.strip().strip("|"))] if not cells or not _SHEET_SECTION_LABEL.match(cells[0]): continue rest = [cell for cell in cells[1:] if cell] if not rest: continue labelled[index] = (cells[0], rest[0]) sections: dict[int, tuple[str, str]] = {} ordered = sorted(labelled) start = 0 while start < len(ordered): end = start + 1 while end < len(ordered) and ordered[end] == ordered[end - 1] + 1: end += 1 if end - start >= SHEET_SECTION_RUN: for position in range(start, end): sections[ordered[position]] = labelled[ordered[position]] start = end return sections def find_candidates( text: str, *, outline_run: int = 0, table_grid: bool = False, unit_fold: bool = False, keep_table_heading: bool = False, sheet_section_rows: bool = False, drop_wrapped_outline: bool = False, ) -> list[Candidate]: """Every boundary the mechanical rules propose, in document order. Two gates from Topic 2 are applied here and both REMOVE candidates: - the **stop-word gate**: a heading made only of function words is not a unit of knowledge; - the **orphan check**: a heading with no body under it proposes nothing, because an empty concept is the silent skip this library refuses everywhere else. `outline_run` is Arm D's gate and it is OFF at 0: the function then behaves exactly as it did before the rule existed. At `N >= 1` the document's own numbered outline contributes boundaries where the integers sustain an ascending run of at least `N`. `unit_fold` is Arm F's gate and it is OFF at False, where `fold_units` is not called at all. On, the candidate list is folded once, at the end: the arm adds no boundary, so every plan it can produce is a subset of the one the flags below produce. `table_grid` is Arm E's gate and it is OFF at False, where the branch is not even evaluated. On, a pandoc grid-table rule line no longer closes an open table block, so one grid table proposes one candidate instead of one per row group. It only ever REMOVES marks, which is what keeps every surviving candidate's `start` fixed and the orphan check monotone. `keep_table_heading` is D1's gate and it is OFF at False. On, a heading whose body is empty ONLY because a table block opens under it keeps that table instead of being dropped: the table is absorbed into the heading's span rather than emitted, so the concept starts at the heading line. The count does not move -- one candidate either way -- and the first byte does. It is its own flag and not part of an arm because the orphan check is reached by every file type, and moving it is a decision about all of them. `sheet_section_rows` is D3's gate and it is OFF at False, where the scan is not run at all. On, a RUN of numbered rows inside an open table block cuts it: each such row opens a candidate that reaches the next section row, or the end of the block. It is the only rule here that opens a span inside a table, and it is its own flag for the same reason D1 is -- a sheet is the one file type whose units are rows, and every other type reaches this scan too. """ lines = text.splitlines(keepends=True) offsets: list[int] = [] position = 0 for line in lines: offsets.append(position) position += len(line) end_of_text = position # Computed BEFORE the loop, and that is a correctness requirement rather # than a style choice: run selection is a whole-text decision (the LAST # maximal run wins, because a contents listing precedes the body it lists), # and a forward scan cannot know which run is last. Deciding it up front is # also what keeps `marked` sorted by construction -- appending outline # candidates in a second pass would leave `end < start` on some spans, and # `text[start:end]` is then `""`, so the orphan check DELETES them # silently. Silent loss, not a raise: nothing would announce it. admitted: dict[int, str] = {} if outline_run > 0: runs = outline_runs(outline_lines(text), outline_run) if runs: # LAST run, not longest and not first. Measured against both: # first-run opens segments inside the table of contents on 14/39 # documents; longest-run differs on 5/39 with no measured reason to # prefer it. "Later occurrence wins" states the document's own # ordering rather than a property of this corpus. admitted = {index: title for index, _, title in runs[-1]} # D3's input, and the same whole-text reasoning as `admitted` above: a run # is a property of the line list, not of a line. sections = _sheet_section_rows(lines) if sheet_section_rows else {} marked: list[tuple[int, Candidate]] = [] in_table = False # Arm E's state, and all three of these are cleared together on the # fall-through below. `rule_pending` is the one that matters: a grid table # ends with a bottom rule, which sets it, and if the blank line after the # table did not clear it the NEXT table's first row would be recorded as a # join although nothing was joined. `joined` holds positions in `marked` # rather than mutating a candidate, because `Candidate` is frozen and the # orphan-check pass below already rebuilds every one of them. rule_pending = False open_block: int | None = None joined: set[int] = set() for index, line in enumerate(lines): if _TABLE_ROW.match(line): section = sections.get(index) if section is not None: label, name = section # A section row never ALSO opens a table block, even when the # run starts on the block's first line: two marks on one line # would give the second an empty span, and the orphan check # deletes an empty span silently. `open_block` is cleared with # it, because a block cut here is no longer one span an Arm E # join could extend. in_table = True open_block = None rule_pending = False marked.append( ( index, Candidate( title=f"{label} {name}", # The table sentinel, not a heading depth: a sheet # row declares no level, and letting it vote in # Arm F's clause 2 would make a sectioned sheet a # one-level document. level=9, number=label, rule=RULE_SHEET_SECTION, start=offsets[index], end=end_of_text, ), ) ) continue if not in_table: in_table = True open_block = len(marked) marked.append( ( index, Candidate( title=f"Tabell linje {index + 1}", level=9, number=None, rule=RULE_TABLE_BLOCK, start=offsets[index], end=end_of_text, ), ) ) elif rule_pending and open_block is not None: joined.add(open_block) rule_pending = False continue if table_grid and in_table and _GRID_RULE.match(line): # The whole of Arm E: do NOT close the block. Reaching here with # `in_table` False is impossible by construction, so a rule line can # never OPEN a block -- a table's top border is not a boundary, and # every surviving candidate keeps the `start` it had under Arm D. rule_pending = True continue in_table = False rule_pending = False open_block = None outline_title = admitted.get(index) if outline_title is not None: if drop_wrapped_outline and _wraps_onto_next_line(lines, index): # Filtered at ADMISSION rather than at run selection: the run # this document sustains is a property of its numbering, and # re-selecting it from a thinned list would move boundaries on # documents where nothing wraps. The rule declines candidates; # it does not rewrite which run won. continue outline_match = _OUTLINE.match(line) assert outline_match is not None, "an admitted index still matches the grammar" marked.append( ( index, Candidate( title=outline_title, level=1, number=outline_match.group("number"), rule=RULE_OUTLINE, start=offsets[index], end=end_of_text, ), ) ) continue atx = _ATX.match(line) numbered = _NUMBERED.match(line) if atx is None and numbered is None: continue if atx is not None: title = atx.group("title") level = len(atx.group("hashes")) inner = _NUMBERED.match(title) number = inner.group("number") if inner else None else: assert numbered is not None title = numbered.group("title") number = numbered.group("number") level = number.count(".") + 1 # The stop-word gate. Applied to the TITLE, after any section number # has been split off, so `3.1 Og` is judged on `Og`. if _is_stop_word_only(title): continue marked.append( ( index, Candidate( title=title, level=level, number=number, rule=RULE_HEADING, start=offsets[index], end=end_of_text, ), ) ) candidates: list[Candidate] = [] # The name an orphaned heading leaves behind, and the ONE candidate allowed # to pick it up. # # WHY THIS EXISTS. A table opening on the line below a heading gives that # heading an empty body, so the orphan check drops it, and the table block # keeps its mechanical `Tabell linje <n>`: the section's NAME is destroyed # even though its content survives. Measured on a 629-concept bundle after # a spreadsheet began rendering as pipe rows -- the concept fell from # candidate rank 10 to 19 for a question naming its subject, and restoring # the title alone put it back at 10. The name is not invented here, it is # carried across: it moves to the segment that holds the heading's content. # # CONDITIONED ON THE DROP, deliberately. A heading that keeps its own body # is still carried by a live candidate, so copying its title onto the table # as well would put one name on two concepts and rescue none. orphaned_name: tuple[str, str | None] | None = None # D1. Which table blocks a heading ABSORBS, decided before the pass that # consumes it: once a table is absorbed the heading's body is no longer # empty, so the orphan check below stops firing on it by itself and no # branch is needed there. Computed only when the caller asked, so every # other arm's `marked` -> `candidates` mapping is untouched code. absorbed = _absorbed_tables(text, marked, offsets, end_of_text) if keep_table_heading else set() # Arm F clause 1's input, and it must be read HERE: the orphan pass below # deletes every bodiless heading, which is every entry of a contents list # but the last, and a run of one is below `CONTENTS_RUN`. contents_run = _contents_run_positions(marked) if unit_fold else set() for position_in_list, (_, candidate) in enumerate(marked): if position_in_list in absorbed: continue following = [ entry for position, entry in enumerate(marked) if position > position_in_list and position not in absorbed ] end = offsets[following[0][0]] if following else end_of_text body = text[candidate.start : end] # The orphan check: everything after the heading line itself. # # D3 IS EXEMPT, and the exception is stated rather than worked around: # the check asks whether anything stands UNDER a candidate's first # line, which is the right question for a heading and the wrong one for # a row. A one-row section carries its content in its own cells, so # every section but the last would be read as bodiless and deleted -- # the rule could not fire at all. It is the same shape as a contents # list without dot leaders, and it is the reason that one needed # `Candidate.contents`. orphan = candidate.rule != RULE_SHEET_SECTION and ( not body.splitlines()[1:] or not "".join(body.splitlines()[1:]).strip() ) if orphan: orphaned_name = (candidate.title, candidate.number) continue inherited, orphaned_name = orphaned_name, None title, number = candidate.title, candidate.number if inherited is not None and candidate.rule == RULE_TABLE_BLOCK: # Both members, because `_segment_path` reads both: the number # becomes the directory AND is stripped from the stem, so carrying # the title alone would emit a name the heading never had. title, number = inherited candidates.append( Candidate( title=title, level=candidate.level, number=number, rule=candidate.rule, start=candidate.start, end=end, grid=position_in_list in joined, contents=position_in_list in contents_run, ) ) return fold_units(candidates) if unit_fold else candidates def _absorbed_tables( text: str, marked: list[tuple[int, Candidate]], offsets: list[int], end_of_text: int, ) -> set[int]: """D1: positions in `marked` of table blocks a bodiless heading keeps. The condition is the orphan check's own, evaluated on the UNABSORBED neighbour distance -- a heading is eligible only when the very next mark is a table block and there is nothing between them but the heading line. That is what keeps the variant from merging a section into a table it merely contains: a heading with a paragraph of its own is not orphaned, so it absorbs nothing. """ absorbed: set[int] = set() for position, (_, candidate) in enumerate(marked): if candidate.rule == RULE_TABLE_BLOCK or position + 1 >= len(marked): continue line_index, following = marked[position + 1] if following.rule != RULE_TABLE_BLOCK: continue body = text[candidate.start : offsets[line_index]] if body.splitlines()[1:] and "".join(body.splitlines()[1:]).strip(): continue absorbed.add(position + 1) return absorbed def _contents_run_positions(marked: list[tuple[int, Candidate]]) -> set[int]: """Arm F clause 1, measured on the PRE-orphan list. Same predicate as there. A run of at least `CONTENTS_RUN` consecutive page-numbered headings, table blocks excluded and never a single line. Written once here and read by `fold_units` through `Candidate.contents`, so the two cannot drift into two definitions of what a contents list is. LEVEL IS NOT PART OF THE PREDICATE, and that is the second round-2 change. A numbered report's contents list interleaves `1`, `1.1`, `1.1.1`, `2`, so requiring the members to be siblings breaks the run at every level change: on K3 position 7 the level-2 entries formed runs long enough to discard and `6.2.2 Tverrfaglig kontroll` did not, so one contents line was emitted as a concept while its neighbours were not. What still bounds the rule is the run LENGTH, which is what the `CONTENTS_RUN` sweep bought and is unchanged. """ inside: set[int] = set() index = 0 while index < len(marked): candidate = marked[index][1] if candidate.rule == RULE_TABLE_BLOCK or not _TRAILING_PAGE_NUMBER.search(candidate.title): index += 1 continue end = index while ( end < len(marked) and marked[end][1].rule != RULE_TABLE_BLOCK and _TRAILING_PAGE_NUMBER.search(marked[end][1].title) ): end += 1 if end - index >= CONTENTS_RUN: inside.update(range(index, end)) index = end return inside #: Arm F clause 1. How many CONSECUTIVE same-level page-numbered headings make #: a contents list. Three, and it is swept rather than guessed: at 1 and 2 the #: rule deletes body chapters (a heading like `... i henhold til TEK 17` ends #: in a number and is not a contents line), and the sweep front is published in #: `docs/2026-09-08-k3-arm-f-mot-enhetsarket.md`. Not a CLI knob: a number a #: caller can turn without recording the sweep is a number nobody has measured. CONTENTS_RUN = 3 def fold_units(candidates: list[Candidate]) -> list[Candidate]: """Arm F: ONE rule, three clauses, and it only MERGES or DISCARDS. Derived from the three rules the operator wrote across the K3 unit worksheet (2026-09-08), not from twelve special cases: 1. "innholdsfortegnelsen er ikke konsepter" -- a RUN of at least `CONTENTS_RUN` consecutive same-level headings each ending in a page number is a contents list and is discarded. A run of siblings, never a single line: one body heading ending in a number is not a contents list. The run is measured on this list AND on the list before the orphan check (`Candidate.contents`), because a contents list whose entries carry no dot leaders is bodiless and reaches here as one surviving line. 2. "hvert h2-kapittel med sine h3" -- the unit level is the SHALLOWEST heading level occurring more than once; anything deeper folds into the preceding candidate at or above that level, which EXTENDS the parent's span rather than deleting the child's body. 3. "tabellen med innledningen" -- a table folds back into the heading immediately before it when that heading's own span is shorter than the table's. The surviving concept keeps the HEADING's name: merging the right bytes under `Tabell linje 48` would produce a concept no reader can look up. It proposes no boundary of its own, so it can only ever reduce a plan. A document with one heading level and no table comes out identical. """ if not candidates: return candidates # Clause 1. Two inputs, one predicate. `contents` carries the run measured # on the list BEFORE the orphan check thinned it -- a contents list with no # dot leaders is a run of bodiless headings, so only its last entry reaches # here and a run of one is below `CONTENTS_RUN`. The scan below still has # work to do: a contents list WITH dot leaders keeps every entry, and that # run exists only in this list. drop: set[int] = {position for position, c in enumerate(candidates) if c.contents} index = 0 while index < len(candidates): candidate = candidates[index] if candidate.rule == RULE_TABLE_BLOCK or not _TRAILING_PAGE_NUMBER.search(candidate.title): index += 1 continue end = index while ( end < len(candidates) and candidates[end].rule != RULE_TABLE_BLOCK and candidates[end].level == candidate.level and _TRAILING_PAGE_NUMBER.search(candidates[end].title) ): end += 1 if end - index >= CONTENTS_RUN: drop.update(range(index, end)) index = end kept = [c for position, c in enumerate(candidates) if position not in drop] if not kept: return kept # Clause 2. The unit level is read from DECLARED headings only -- ATX and # dotted-numbered, `RULE_HEADING`. Two exclusions, each measured: # # - a table block carries the sentinel level 9, and letting it vote would # make every document with two tables a two-level document; # - `RULE_OUTLINE` is Arm D's RECOVERY of an integer numbering run, not a # level the document declares. Letting it vote made it the shallowest # repeated level on every PDF that has both, folding the dotted headings # the operator actually named into it: 23 -> 3, 48 -> 4 and 11 -> 7 on # K3 positions 1, 7 and 9. The unit worksheet showed the operator ATX and # dotted headings and nothing else, which is the same set. # # With no declared heading at all the fold has no level to work from, and # `unit_level` is set below every candidate's, so clause 2 is inert and # only clause 3 can fire. That is the honest behaviour: a document whose # structure was recovered rather than declared has no unit level to read. levels = [c.level for c in kept if c.rule == RULE_HEADING] repeated = sorted({level for level in levels if levels.count(level) > 1}) if repeated: unit_level = repeated[0] elif levels: unit_level = min(levels) else: unit_level = -1 folded: list[Candidate] = [] for candidate in kept: table = candidate.rule == RULE_TABLE_BLOCK deeper = candidate.rule == RULE_HEADING and candidate.level > unit_level if table and folded: previous = folded[-1] introduces = ( previous.rule != RULE_TABLE_BLOCK and previous.end - previous.start < candidate.end - candidate.start ) else: introduces = False if (deeper or introduces) and folded: previous = folded[-1] folded[-1] = replace(previous, end=candidate.end) continue folded.append(candidate) return folded def _cut_points(text: str, start: int, end: int, cap: int) -> list[int]: """Where to cut `text[start:end]` so no part exceeds `cap` characters. The cut prefers a PARAGRAPH boundary (a blank line) inside the window, then a line boundary, and only then cuts mid-line. The order is the whole content of the rule: a cut that lands mid-sentence splits one unit of knowledge for no reason other than arithmetic, and the K3 categories count that as `too fine`. The last resort exists anyway, because a document whose body is one unbroken line is exactly where a cap that quietly stopped binding would be least defensible. """ cuts: list[int] = [] position = start while end - position > cap: window_end = position + cap paragraph = text.rfind("\n\n", position, window_end) if paragraph != -1: cut = paragraph + 2 else: line = text.rfind("\n", position, window_end) cut = line + 1 if line != -1 else window_end # rfind can only return an index at or after `position`, so every # branch advances. The assertion states that rather than trusting it: # a cut that did not advance would loop forever on a corpus run. assert cut > position, f"cut {cut} did not advance past {position}" cuts.append(cut) position = cut return cuts def subdivide(text: str, candidates: list[Candidate], cap: int) -> list[Candidate]: """Arm C. Arm B's candidates, with every over-long span cut down to `cap`. ARM C IS NOT DEFINED IN `docs/2026-09-02-k3-k4-k5-metode.md`; that file contains no occurrence of the word. This definition was written for order 20260904T145630Z and is reported as the author's, not as a ratified one. Two callers' cases, one rule. When Arm B found boundaries but a span still runs long (a PDF whose headings are its table of contents, so the trailing segment absorbs the body), the span is cut. When Arm B found NO boundary at all, the whole document is that span -- which is the `no declared structure` case § 10 names, and 23 of 33 PDFs in the K2 corpus are in it. A document with no boundaries that is already under the cap proposes NOTHING, exactly as Arm B does. Arm C fires on size; where size is not the problem it has nothing to say, and a one-entry plan would only dress a single concept in a plan file. """ if cap <= 0: return candidates if not candidates: if len(text) <= cap: return [] # The synthetic span. Its rule is the size rule alone, because no # heading rule proposed it -- there was no heading. candidates = [ Candidate( title="Del", level=1, number=None, rule=RULE_SIZE_SPLIT, start=0, end=len(text), split=False, ) ] unnumbered_parts = True else: unnumbered_parts = False out: list[Candidate] = [] for candidate in candidates: cuts = _cut_points(text, candidate.start, candidate.end, cap) if not cuts: out.append(candidate) continue edges = [candidate.start, *cuts, candidate.end] for part, (start, end) in enumerate(zip(edges, edges[1:]), start=1): if unnumbered_parts: title = f"Del {part}" else: title = candidate.title if part == 1 else f"{candidate.title} (del {part})" out.append( Candidate( title=title, level=candidate.level, number=candidate.number, rule=candidate.rule, start=start, end=end, split=True, grid=candidate.grid, ) ) return out def _derived_names(candidate: Candidate) -> list[str]: """The `derived` list for one candidate, in the order a test pins. The marker, the rule that OPENED the span, then any rule that reshaped it: Arm E's join before Arm C's cut, because a joined span is what Arm C would then have been given to cut. """ names = [PROPOSED_MARKER, candidate.rule] if candidate.grid: names.append(RULE_TABLE_GRID) if candidate.split and candidate.rule != RULE_SIZE_SPLIT: names.append(RULE_SIZE_SPLIT) return names def _segment_path(candidate: Candidate, taken: set[str], prefix: str = "") -> str: title = unicodedata.normalize("NFC", candidate.title) # The section number becomes the DIRECTORY, so leaving it in the stem too # yields `3-1/3-1-brannkonsept.md` -- correct and unreadable. if candidate.number and title.startswith(candidate.number): title = title[len(candidate.number) :] stem = reduce_to_id_grammar(title) if not stem: stem = "seksjon" directory = reduce_to_id_grammar(candidate.number or "") if candidate.number else "" # The caller's scope comes FIRST and is never deduplicated against: it is # the same for every entry in this document by construction, and that is # the whole point -- one document's sections must not be able to claim # another's path. head = f"{prefix}/" if prefix else "" path = f"{head}{directory}/{stem}.md" if directory else f"{head}{stem}.md" suffix = 2 while path in taken: path = f"{head}{directory}/{stem}-{suffix}.md" if directory else f"{head}{stem}-{suffix}.md" suffix += 1 taken.add(path) return path def build_plan( source: Path, text: str, source_bytes: bytes, *, okf_type: str, proposed_at: str, path_prefix: str = "", max_segment_chars: int = 0, outline_run: int = 0, table_grid: bool = False, unit_fold: bool = False, keep_table_heading: bool = False, sheet_section_rows: bool = False, drop_wrapped_outline: bool = False, ) -> dict[str, Any]: """The artifact. Every entry PROPOSED, the plan itself never adjudicated.""" taken: set[str] = set() extractor_id = source.suffix.lower().lstrip(".") or "none" entries: list[dict[str, Any]] = [] candidates = find_candidates( text, outline_run=outline_run, table_grid=table_grid, unit_fold=unit_fold, keep_table_heading=keep_table_heading, sheet_section_rows=sheet_section_rows, drop_wrapped_outline=drop_wrapped_outline, ) for candidate in subdivide(text, candidates, max_segment_chars): entries.append( { "segment_id": f"p{len(entries) + 1}", "path": _segment_path(candidate, taken, path_prefix), "title": candidate.title, "okf_type": okf_type, "span": [candidate.start, candidate.end], "ingested_at": proposed_at, # The offsets are a hint the anchor may correct. Written at # proposal time because that is the only moment the text the # adjudicator will judge and the offsets naming it are known # to agree -- reconstructing it later would anchor to whatever # the extraction had already become. "anchor": { "quote": text[candidate.start : candidate.end], "prefix": text[max(0, candidate.start - ANCHOR_CONTEXT) : candidate.start], "suffix": text[candidate.end : candidate.end + ANCHOR_CONTEXT], }, # PROPOSED first, then the rule that proposed it. `derived` is # this library's existing "which of these did we infer" marker, # so a consumer that already distrusts derived fields # distrusts these by construction. # PROPOSED first, then the rule that OPENED the span, then -- # for an Arm E join -- the grid rule, then -- for an Arm C part # only -- the size rule that cut it. More than one name on those # entries because the rule that opened the span is still true of # them, and dropping it would leave a part traceable to nothing # but arithmetic or nothing but a join. # # Written as an ordered build rather than a conditional # expression: four combinations exist now, and the order is # itself a claim a test pins. "derived": _derived_names(candidate), } ) return { "version": "1", "source_sha256": hashlib.sha256(source_bytes).hexdigest(), # The hash the offsets actually depend on. Source bytes alone cannot # see a converter reshaping its output, so the staleness signal this # plan is supposed to carry did not exist until this line did. "text_sha256": hashlib.sha256(text.encode("utf-8")).hexdigest(), "extractor_id": extractor_id, # The EXTRACTOR's version, not this tool's. `PROPOSER_VERSION` sat here # and named the wrong thing: a converter bump left the field frozen at # the proposer's own number, so the component could not move. "extractor_version": observed_extractor_version(extractor_id), "adjudicated_at": proposed_at, # NOT a timestamp question. `adjudicated_at` records when this artifact # was produced; this records whether a human has looked at it, and it is # false until one replaces the file. "adjudicated": False, "proposed_by": f"{PROPOSER_ID}/{PROPOSER_VERSION}", "entries": entries, } def run( source: Path, out: Path, *, okf_type: str, proposed_at: str, path_prefix: str = "", max_segment_chars: int = 0, outline_run: int = 0, table_grid: bool = False, unit_fold: bool = False, keep_table_heading: bool = False, sheet_section_rows: bool = False, drop_wrapped_outline: bool = False, pdf_headings: bool = False, pdf_headings_reserve: bool = False, ocr: bool = False, ) -> int: if max_segment_chars < 0: raise ProposerError( f"--max-segment-chars {max_segment_chars} is negative; the cap is a " "character count, and 0 means off (Arm B)" ) if outline_run < 0: raise ProposerError( f"--outline-run {outline_run} is negative; the gate is a run LENGTH, " "and 0 means off (Arm B)" ) # Reduced HERE, before anything is read: a prefix that survives to the # entries as an empty component would produce exactly the unscoped paths # the caller asked to avoid, and would do it silently. # # PER COMPONENT, because the prefix carries a DIRECTORY now that Door B # walks the inbox recursively and records a relative `source_file`. # Reducing the whole string would fold `/` into a `-` and flatten # `sub/sub2` into the single component `sub-sub2` -- a bundle shaped unlike # the inbox it came from, and unlike what the caller wrote. components = ( [reduce_to_id_grammar(part) for part in path_prefix.split("/")] if path_prefix else [] ) if path_prefix and not all(components): raise ProposerError( f"--path-prefix {path_prefix!r} has a component that reduces to nothing under " "the id grammar ([a-z0-9][a-z0-9-]*); refusing to write unscoped paths under a " "scope that was asked for" ) scope = "/".join(components) if not source.is_file(): raise ProposerError(f"source is not a file: {source}") try: source_bytes = source.read_bytes() except OSError as exc: raise ProposerError(f"cannot read {source}: {exc}") from exc try: # The two READER options, not arms: they change what the extraction # says, and every arm below reads whatever it says. Passed here as well # as to the run path because the plan's `text_sha256` indexes this # exact string -- a plan proposed against one rendering and replayed # against another is refused by `assert_plan_applies`, which is the # right outcome and a confusing one to debug. text = extract_text(source.name, source_bytes, pdf_headings=pdf_headings, ocr=ocr) # The reserve, and the reason it re-extracts rather than post-processes: # the font reader works on the PDF's glyph geometry, which the joined # text no longer carries. Skipped outright when the font reader is # already on -- `font` and `font-reserve` are two values of one option, # never a pair to combine. if pdf_headings_reserve and not pdf_headings: if heading_reserve_applies(text, outline_run=outline_run): text = extract_text(source.name, source_bytes, pdf_headings=True, ocr=ocr) except IngestError as exc: raise ProposerError(f"cannot extract text from {source.name}: {exc}") from exc payload = build_plan( source, text, source_bytes, okf_type=okf_type, proposed_at=proposed_at, path_prefix=scope, max_segment_chars=max_segment_chars, outline_run=outline_run, table_grid=table_grid, unit_fold=unit_fold, keep_table_heading=keep_table_heading, sheet_section_rows=sheet_section_rows, drop_wrapped_outline=drop_wrapped_outline, ) # Nothing to propose is an OUTCOME, and it is not an artifact. An empty # plan cannot be replayed -- `process_inbox` refuses one, because a plan # naming no entry would persist nothing for a document that was dropped -- # so the only thing a zero-entry file can do is fail a run later. Its own # exit status, distinct from 2, so a driver can tell "this document lands # as one flat concept" from "stop". if not payload["entries"]: print( f"{PROPOSER_ID}: nothing to propose for {source.name} — the mechanical " "rules found no boundary. No artifact written; this document lands as " "one concept unless someone segments it by hand.", file=sys.stderr, ) return 1 out.parent.mkdir(parents=True, exist_ok=True) out.write_bytes((json.dumps(payload, indent=2, ensure_ascii=False) + "\n").encode("utf-8")) print( f"{PROPOSER_ID}: proposed {len(payload['entries'])} segment(s) -> {out}\n" f"{PROPOSER_ID}: every entry is PROPOSED. Adjudicate before ingesting.", file=sys.stderr, ) return 0 def parse_args(argv: list[str] | None) -> argparse.Namespace: parser = argparse.ArgumentParser( prog=PROPOSER_ID, description="Propose a segmentation plan. A human adjudicates it before use.", ) parser.add_argument("source", type=Path, help="the document to segment") parser.add_argument("--out", type=Path, required=True, help="where to write the artifact") parser.add_argument("--okf-type", default="reference", help="okf_type for every entry") parser.add_argument( "--path-prefix", default="", help=( "scope every entry's path under this directory, `/`-separated for a " "nested one (each component is reduced on its own). Required for a corpus: " "section numbering is document-local, so two documents propose the same " "path and Door B refuses both. An argument rather than something this " "tool derives -- it sees one document and cannot know what else is in " "the bundle" ), ) parser.add_argument( "--max-segment-chars", type=int, default=0, metavar="N", help=( "Arm C: cut any proposed span longer than N characters at the nearest " "paragraph boundary, the whole document counting as one span when the " "mechanical rules find no boundary at all. 0 (the default) is OFF and " "leaves the artifact byte-identical to Arm B. Arm C is the author's " "definition, written for order 20260904T145630Z; it is not defined in " "the K3 method file" ), ) parser.add_argument( "--outline-run", type=int, default=0, metavar="N", help=( "Arm D: also propose a boundary at each line of the document's own " "numbered outline (the bare integers the heading grammar cannot " "match, since it requires a dot), but only where those integers " "sustain an ascending run of at least N entries, and only for the " "LAST such run when the outline repeats, because a contents listing " "precedes the body it lists. 0 (the default) is OFF and leaves the " "artifact byte-identical to Arm B. Arm D is the author's definition, " "written for order 20260906T213322Z; it is not defined upstream, and " "the K3 method file does not name it either" ), ) parser.add_argument( "--unit-fold", action="store_true", help=( "Arm F: fold the proposed candidates into the operator's units --" " discard a run of contents-list headings, fold a deeper heading" " into its parent, and fold a table back into the shorter heading" " that introduces it. ONE rule with three clauses, derived from the" " three rules the K3 unit worksheet records (2026-09-08). It adds" " no boundary, so it can only reduce a plan. Absent (the default)" " is OFF and leaves the artifact byte-identical to the arm below" " it. A boolean: the rule's one number, CONTENTS_RUN, is a module" " constant whose sweep is published, not a knob a caller can turn" ), ) parser.add_argument( "--table-grid", action="store_true", help=( "Arm E: do not let a pandoc grid-table rule line (`+---+---+`, and " "`+===+===+` under a header) close an open table block. The table " "grammar cannot match a rule line, so without this one grid table " "becomes one concept per row group. Absent (the default) is OFF and " "leaves the artifact byte-identical to Arm D (Arm D rather than " "Arm B, because Arm E is defined on top of it). A boolean and not a " "number: the rule has no parameter to sweep. Arm E is the author's " "definition, written for order 20260907T075834Z-18584396-from-.claude; " "it is not defined upstream, and the K3 method file does not name it " "either" ), ) parser.add_argument( "--keep-table-heading", action="store_true", help=( "D1: keep a heading whose body is empty ONLY because a table block " "opens under it, and absorb that table into the heading's span " "instead of emitting it. Without this the orphan check drops the " "heading, the table inherits its NAME, and the concept starts at " "the first table row -- so the heading line is in no concept's " "body. Measured on a spreadsheet: the concept count does not move " "(1 -> 1), the first byte does. Absent (the default) is OFF and " "leaves every artifact byte-identical. Its own flag rather than " "part of an arm: the orphan check is reached by every file type" ), ) parser.add_argument( "--sheet-section-rows", action="store_true", help=( "D3: cut an open table block at the rows that label its sections. " "A section row is one of a RUN of at least SHEET_SECTION_RUN " "consecutive rows whose first cell is a bare numeric label and " "which carry at least one other non-empty cell; each opens a " "candidate reaching the next section row or the end of the block. " "The opposite direction from Arm E, which stops a grid rule line " "from CLOSING a block: that arm decides how far a block extends, " "this rule where it is cut inside, and they compose in that order. " "Written for a spreadsheet whose whole body is one table block and " "whose units are rows. Absent (the default) is OFF and leaves every " "artifact byte-identical. A boolean: the run length is a module " "constant, not a knob a caller can turn. D3 is the author's " "definition, written for order " "20260908T170037Z-3622420612-from-.claude; it is not defined " "upstream, and the K3 method file does not name it either" ), ) parser.add_argument( "--drop-wrapped-outline", action="store_true", help=( "D3: do not admit an outline candidate whose line continues onto " "the next one, because a wrapped sentence is not a heading. Judges " "RECOVERED candidates only, never a heading the document declares " "for itself. Written for quoted regulation text, whose numbered " "paragraphs match Arm D's grammar exactly; a TITLE LENGTH rule was " "tried first and falsified, because a real 88-character heading " "sits between the quoted sentences at 86 and 91. Absent (the " "default) is OFF and leaves every artifact byte-identical. It is " "the author's definition, written for order " "20260908T170037Z-3622420612-from-.claude; it is not defined " "upstream, and the K3 method file does not name it either" ), ) parser.add_argument( "--proposed-at", default="1970-01-01T00:00:00Z", help="the timestamp written into the artifact; explicit so a run is reproducible", ) return parser.parse_args(argv) def main(argv: list[str] | None = None) -> int: args = parse_args(argv) try: return run( args.source, args.out, okf_type=args.okf_type, proposed_at=args.proposed_at, path_prefix=args.path_prefix, max_segment_chars=args.max_segment_chars, outline_run=args.outline_run, table_grid=args.table_grid, unit_fold=args.unit_fold, keep_table_heading=args.keep_table_heading, sheet_section_rows=args.sheet_section_rows, drop_wrapped_outline=args.drop_wrapped_outline, ) except ProposerError as exc: print(f"{PROPOSER_ID}: FAILED - {exc}", file=sys.stderr) print( f"{PROPOSER_ID}: this is NOT 'nothing to propose'. Nothing was written.", file=sys.stderr, ) return 2 if __name__ == "__main__": raise SystemExit(main())