feat(description): an STS section's description is its own first spec point

K3-19 c. The NISO-STS reader records, per titled <sec>, the FIRST <p> of its
FIRST direct-child <sec sec-type="spec"> as `OutlineMark.description`. The
plan entry carries it (`description`, only where the source has one, so every
other row's plan keeps its bytes), `parse_segmentation_plan` refuses an empty,
multi-line or non-string value, and the door writes it as the concept's
`description` after the gate has seen it: it is document text persisted
outside the body the gate screens, so it is kept only on the non-blocking
floor and as the sanitized text.

SPEC SS 4.1 makes `description` RECOMMENDED and sets no length, in SS 4.1,
SS 8 or SS 11. The limit is ours and structural -- one paragraph, whole --
because a cut inside it writes a sentence the source never wrote. Measured on
R761: 2 026 of 2 761 titled sections carry a direct-child spec point; the
first <p> runs 17 / 109 / 273 / 521 / 942 characters (min / median / p90 /
p99 / max). A section with none gets no key; nothing is derived from the
title. A stated `--frontmatter description=...` replaces it.

The extracted text does not move: the description is read beside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-11 03:22:44 +02:00
commit de7849e35b
6 changed files with 96 additions and 4 deletions

View file

@ -130,7 +130,11 @@ one boundary rule:
one", which a name read from inside a document cannot obey. The `sources` one", which a name read from inside a document cannot obey. The `sources`
title is `<doc-number>` + `<year>`, then `<title-wrap>`, then the file name: title is `<doc-number>` + `<year>`, then `<title-wrap>`, then the file name:
R761's `<full>` carries a COMMA, a flow terminator, so it is never written R761's `<full>` carries a COMMA, a flow terminator, so it is never written
and never cleaned up. The registries are COUPLED: a row in and never cleaned up. A titled section's `description` is its own FIRST
spec point (first `<p>` of the first DIRECT-child `sec-type="spec"`, whole),
carried by the plan entry and screened by the gate: 2 026 of 2 761 on R761,
none invented for the rest. SS 4.1 sets no length, so the one-paragraph
limit is ours. The registries are COUPLED: a row in
`_CORE_EXTRACTORS` and not in `segmentation._STDLIB_EXTRACTOR_IDS` refuses `_CORE_EXTRACTORS` and not in `segmentation._STDLIB_EXTRACTOR_IDS` refuses
every proposal for the type, two layers away from the extractor. every proposal for the type, two layers away from the extractor.
`pdf`/`docx`/`xlsx` only via `pdf`/`docx`/`xlsx` only via

View file

@ -640,7 +640,10 @@ bundle:
the `sources` title from `<doc-number>` + `<year>`, then `<title-wrap>`, the `sources` title from `<doc-number>` + `<year>`, then `<title-wrap>`,
then the file name, whichever is the first that can be written verbatim. then the file name, whichever is the first that can be written verbatim.
Stated more than once, or claimed by a second document in the same run, a Stated more than once, or claimed by a second document in the same run, a
declared name is not used and the file name stays. declared name is not used and the file name stays. Each titled section's
`description` is its own first spec point — the first `<p>` of its first
direct-child `<sec sec-type="spec">`, whole — and a section with none gets
no `description` at all; nothing is derived from the title.
The drop directory is walked **recursively**, in sorted relative-path order: The drop directory is walked **recursively**, in sorted relative-path order:
a file at any depth is ingested and records its path relative to the inbox a file at any depth is ingested and records its path relative to the inbox
root as its `source_file`, while dot-directories and a bundle directory root as its `source_file`, while dot-directories and a bundle directory

View file

@ -524,6 +524,30 @@ class _XmlTextExtractor:
"""An element's whole text, whitespace collapsed.""" """An element's whole text, whitespace collapsed."""
return " ".join("".join(element.itertext()).split()) return " ".join("".join(element.itertext()).split())
def _spec_point(self, section: Element) -> str | None:
"""The FIRST `<p>` of the FIRST direct-child `sec-type="spec"`, whole.
The limit is this package's, not the spec's: SPEC SS 4.1 asks for "a
single sentence" and sets no length anywhere. It is STRUCTURAL rather
than a character count, because a cut inside a paragraph writes a
sentence the source never wrote. Measured on the one STS document this
row has: 2 026 of 2 761 titled sections carry a direct-child spec
point; 264 of those points hold more than one `<p>` and 2 hold none;
the first `<p>` runs 17 / 109 / 273 / 521 / 942 characters at min /
median / p90 / p99 / max.
A DIRECT child only: a spec point belongs to the section it opens under,
and a container borrowing its first child's would describe a section by
a sentence about another one.
"""
for child in section:
if _local_name(child.tag) == "sec" and child.get("sec-type") == "spec":
for paragraph in child:
if _local_name(paragraph.tag) == "p":
return self._text_of(paragraph) or None
return None
return None
def _table(self, element: Element) -> bool: def _table(self, element: Element) -> bool:
"""A `<table-wrap>`: its label on a line, its rows as ONE table block. """A `<table-wrap>`: its label on a line, its rows as ONE table block.
@ -564,7 +588,12 @@ class _XmlTextExtractor:
heading = " ".join(part for part in parts if part) heading = " ".join(part for part in parts if part)
self._emit("#" * level + " " + heading) self._emit("#" * level + " " + heading)
self.marks.append( self.marks.append(
OutlineMark(line=len(self._lines) - 1, level=level, title=heading) OutlineMark(
line=len(self._lines) - 1,
level=level,
title=heading,
description=self._spec_point(element),
)
) )
skip = {id(title)} | ({id(label)} if label is not None else set()) skip = {id(title)} | ({id(label)} if label is not None else set())
elif label is not None: elif label is not None:
@ -1036,6 +1065,10 @@ class OutlineMark:
line: int line: int
level: int level: int
title: str title: str
#: The section's own first spec point, where the SOURCE declares one. Set
#: only by the NISO-STS reader; `None` on every bookmark mark, because a
#: bookmark declares a place and never a summary.
description: str | None = None
@dataclass(frozen=True) @dataclass(frozen=True)

View file

@ -213,6 +213,10 @@ def render_inbox_concept(
frontmatter[policy.offset_key] = _render_flow_list([str(offset) for offset in segment.span]) frontmatter[policy.offset_key] = _render_flow_list([str(offset) for offset in segment.span])
if segment.parent_id is not None: if segment.parent_id is not None:
frontmatter["parent"] = segment.parent_id frontmatter["parent"] = segment.parent_id
if segment.description is not None:
# The SOURCE's words, carried by the plan. Absent is the source
# saying nothing -- never a summary derived from the title.
frontmatter["description"] = segment.description
if policy.adjudication_key is not None: if policy.adjudication_key is not None:
# The per-entry verdict IS the discriminator. A plan-level # The per-entry verdict IS the discriminator. A plan-level
# `adjudicated: true` with no record for THIS entry leaves it # `adjudicated: true` with no record for THIS entry leaves it
@ -834,7 +838,13 @@ def _render_segments(
ingested_at=entry.ingested_at, ingested_at=entry.ingested_at,
profile=profile, profile=profile,
structure=structure, structure=structure,
segment=entry, # The description is document text persisted OUTSIDE the
# body this gate just screened, so it is screened too.
segment=(
replace(entry, description=_screened(gate, entry.description))
if entry.description is not None
else entry
),
bundle_id=bundle_id, bundle_id=bundle_id,
units=units, units=units,
source_title=source_title, source_title=source_title,

View file

@ -1514,6 +1514,22 @@ def build_plan(
outline_marks=outline_marks, outline_marks=outline_marks,
outline_rule=outline_rule, outline_rule=outline_rule,
) )
# A mark's description travels to the candidate whose span holds the mark's
# line -- the FIRST such mark, so a span that opens at zero (the front
# matter above the first section) still takes its own section's and no
# other. Only the NISO-STS reader sets one, so every other row's plan keeps
# its bytes: an absent key is the source saying nothing.
starts: list[int] = []
notes: list[str | None] = []
if outline_marks:
offsets = [0]
for line in text.splitlines(keepends=True):
offsets.append(offsets[-1] + len(line))
for mark in sorted(outline_marks, key=lambda item: item.line):
if mark.line < len(offsets) - 1:
starts.append(offsets[mark.line])
notes.append(mark.description)
cursor = 0
for candidate in subdivide(text, candidates, max_segment_chars): for candidate in subdivide(text, candidates, max_segment_chars):
entries.append( entries.append(
{ {
@ -1550,6 +1566,12 @@ def build_plan(
"derived": _derived_names(candidate), "derived": _derived_names(candidate),
} }
) )
while cursor < len(starts) and starts[cursor] < candidate.start:
cursor += 1
if cursor < len(starts) and starts[cursor] < candidate.end:
description = notes[cursor]
if description is not None:
entries[-1]["description"] = description
return { return {
"version": "1", "version": "1",
"source_sha256": hashlib.sha256(source_bytes).hexdigest(), "source_sha256": hashlib.sha256(source_bytes).hexdigest(),

View file

@ -157,6 +157,11 @@ class SegmentEntry:
derived: frozenset[str] = field(default_factory=frozenset) derived: frozenset[str] = field(default_factory=frozenset)
anchor: SegmentAnchor | None = None anchor: SegmentAnchor | None = None
adjudication: SegmentVerdict | None = None adjudication: SegmentVerdict | None = None
#: The SOURCE's own one-line summary of this segment, where it declares one
#: (NISO-STS: the section's first spec point). Carried in the plan beside
#: `title` for the same reason `title` is: a plan is the record a rebuild
#: replays, and the door writes it as the concept's `description`.
description: str | None = None
@dataclass(frozen=True) @dataclass(frozen=True)
@ -363,6 +368,20 @@ def _parse_entry(payload: Any, *, position: int) -> SegmentEntry:
f"{where} field 'parent_id' must be a non-empty string or absent, got {parent_id!r}", f"{where} field 'parent_id' must be a non-empty string or absent, got {parent_id!r}",
code="segmentation_plan_invalid", code="segmentation_plan_invalid",
) )
description = payload.get("description")
if description is not None and (
not isinstance(description, str)
or not description
or "\n" in description
or "\r" in description
):
# Written into line-oriented frontmatter, so a line break would inject
# a key and an empty value would state a summary that says nothing.
raise SegmentationError(
f"{where} field 'description' must be a non-empty single-line string or "
f"absent, got {description!r}",
code="segmentation_plan_invalid",
)
return SegmentEntry( return SegmentEntry(
segment_id=segment_id, segment_id=segment_id,
path=normalize_segment_path(payload["path"], where=where), path=normalize_segment_path(payload["path"], where=where),
@ -374,6 +393,7 @@ def _parse_entry(payload: Any, *, position: int) -> SegmentEntry:
derived=_parse_derived(payload.get("derived", ()), where=where), derived=_parse_derived(payload.get("derived", ()), where=where),
anchor=_parse_anchor(payload.get("anchor"), where=where), anchor=_parse_anchor(payload.get("anchor"), where=where),
adjudication=_parse_verdict(payload.get("adjudication"), where=where), adjudication=_parse_verdict(payload.get("adjudication"), where=where),
description=description,
) )