feat(identity): an STS document's doc-number names its directory and its title the address

K3-19 a. `extract.declared_identity` reads what a NISO-STS document states
about itself -- exactly one <std-ident> (<doc-number>, <year>) and exactly one
<title-wrap> (<full>, else <main>) -- and returns None for every other row,
for XML that is not STS, for an unparseable file, and for a document that
states neither. A value stated more than once is not read: an adopted
standard carries one <std-ident> per issuing body, and picking one is a guess.

`okf build` names a document's directory from its <doc-number> through the id
grammar, replacing only the file's stem. A declared name another document in
the run also claims falls back to the file name for both, said on stderr: the
existing collision gate would refuse both with "rename one", and a name read
from inside a document is not one a rename can change.

`sources[0].title` becomes <doc-number> + <year>, then the <title-wrap>
title, then the file name -- the first that survives the gate and can be
written into the flow mapping verbatim. Measured on R761, <full> carries a
comma, which ends a flow mapping, so it is never the title there; it is never
cleaned up either. `resource` stays the inbox-relative file.

Every other row, and every profile without an address, is untouched: the
identity is asked for only where `sources` is written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-11 03:08:30 +02:00
commit ee8d5b5776
6 changed files with 230 additions and 13 deletions

View file

@ -120,7 +120,17 @@ one boundary rule:
reachable without inventing an id; round 13's 14 such directories were false reachable without inventing an id; round 13's 14 such directories were false
positives of the TEXT route reading the document's own contents listing and positives of the TEXT route reading the document's own contents listing and
are gone. Report: are gone. Report:
`docs/2026-09-10-k3-runde14-deklarert-struktur-tar-ruten.md`. The registries are COUPLED: a row in `docs/2026-09-10-k3-runde14-deklarert-struktur-tar-ruten.md`.
**Since K3-19 an STS document's own identity names its directory**
(`extract.declared_identity`, read by `cli._document_prefixes` and the door):
the directory was the delivery file's stem, a UUID occurring **0 times** in
the document, while its one `<std-ident>` carried `<doc-number>`. Only the
stem is replaced, and a declared name two documents in one run claim is
used by NEITHER -- the `slug_owners` gate would refuse both with "rename
one", which a name read from inside a document cannot obey. The `sources`
title is `<doc-number>` + `<year>`, then `<title-wrap>`, then the file name:
R761's `<full>` carries a COMMA, a flow terminator, so it is never written
and never cleaned up. The registries are COUPLED: a row in
`_CORE_EXTRACTORS` and not in `segmentation._STDLIB_EXTRACTOR_IDS` refuses `_CORE_EXTRACTORS` and not in `segmentation._STDLIB_EXTRACTOR_IDS` refuses
every proposal for the type, two layers away from the extractor. every proposal for the type, two layers away from the extractor.
`pdf`/`docx`/`xlsx` only via `pdf`/`docx`/`xlsx` only via

View file

@ -626,6 +626,13 @@ bundle:
`<table-wrap>` becomes one markdown table. Any other XML keeps its text in `<table-wrap>` becomes one markdown table. Any other XML keeps its text in
document order and gets no invented structure. XML carrying a document order and gets no invented structure. XML carrying a
`<!DOCTYPE` is refused unparsed. `<!DOCTYPE` is refused unparsed.
A NISO-STS document that **states who it is** names its own directory:
`okf build` takes the directory from the document's one `<std-ident>`
`<doc-number>` (reduced to the id grammar) instead of the file's stem, and
the `sources` title from `<doc-number>` + `<year>`, then `<title-wrap>`,
then the file name, whichever is the first that can be written verbatim.
Stated more than once, or claimed by a second document in the same run, a
declared name is not used and the file name stays.
The drop directory is walked **recursively**, in sorted relative-path order: The drop directory is walked **recursively**, in sorted relative-path order:
a file at any depth is ingested and records its path relative to the inbox a file at any depth is ingested and records its path relative to the inbox
root as its `source_file`, while dot-directories and a bundle directory root as its `source_file`, while dot-directories and a bundle directory

View file

@ -70,12 +70,15 @@ from __future__ import annotations
import argparse import argparse
import sys import sys
import tempfile import tempfile
from collections.abc import Sequence
from functools import partial from functools import partial
from pathlib import Path from pathlib import Path
from .corpus import LOG_NAME, CorpusReport, load_plans, measure from .corpus import LOG_NAME, CorpusReport, load_plans, measure
from .errors import IngestError from .errors import IngestError
from .extract import declared_identity
from .inbox import walk_inbox from .inbox import walk_inbox
from .materialize import reduce_to_id_grammar
from .profiles import SEGMENTED_OKF_V0_2, STRUCTURED_V1, BundleProfile from .profiles import SEGMENTED_OKF_V0_2, STRUCTURED_V1, BundleProfile
from .propose import ProposerError, heading_reserve_applies from .propose import ProposerError, heading_reserve_applies
from .propose import run as propose_run from .propose import run as propose_run
@ -248,6 +251,60 @@ DEFAULT_PDF_OUTLINE = False
DEFAULT_STAMP = "1970-01-01T00:00:00Z" DEFAULT_STAMP = "1970-01-01T00:00:00Z"
def _document_prefixes(inbox: Path, walked: Sequence[Path]) -> dict[Path, str]:
"""Each document's directory: the name it declares, else its file name.
MEASURED: a NISO-STS delivery landed every one of its 2 761 concepts under
a directory named for the delivery path's file name, a UUID occurring 0
times in the document, while the document's own `<doc-number>` said what it
was. Only the file's STEM is replaced; the folders above it are the
operator's arrangement and stay.
A declared name another document in this run also claims -- by declaring
it, or by its file name reducing to it -- is not taken by either, and both
keep their file name. The gate Door B already has would refuse both and
tell the operator to rename one, and a name read from inside a document is
not one a rename can change. Said on stderr rather than silently, because a
directory that stays a UUID is otherwise indistinguishable from this rule
never having run.
"""
def scope(prefix: str) -> str:
return "/".join(reduce_to_id_grammar(part) for part in prefix.split("/"))
named = {source: source.relative_to(inbox).with_suffix("").as_posix() for source in walked}
declared: dict[Path, tuple[str, str]] = {}
for source in walked:
try:
identity = declared_identity(source.name, source.read_bytes())
except OSError:
# The proposer reads the same file next and reports it per file.
continue
if identity is None or identity.doc_number is None:
continue
slug = reduce_to_id_grammar(identity.doc_number)
if slug:
parent = source.relative_to(inbox).parent
declared[source] = ((parent / slug).as_posix(), identity.doc_number)
claims: dict[str, set[Path]] = {}
for source in walked:
claims.setdefault(scope(named[source]), set()).add(source)
for source, (prefix, _) in declared.items():
claims.setdefault(scope(prefix), set()).add(source)
prefixes = dict(named)
for source, (prefix, doc_number) in declared.items():
others = sorted(named[other] for other in claims[scope(prefix)] if other != source)
if others:
print(
f"{CLI_ID}: {named[source]}: <doc-number> {doc_number!r} names {prefix!r}, "
f"which {', '.join(others)} also claims; both keep their file name",
file=sys.stderr,
)
continue
prefixes[source] = prefix
return prefixes
def _propose_plans( def _propose_plans(
inbox: Path, inbox: Path,
bundle: Path, bundle: Path,
@ -284,6 +341,7 @@ def _propose_plans(
corpora the two-script path completes. corpora the two-script path completes.
""" """
walked, _ = walk_inbox(inbox, exclude=bundle) walked, _ = walk_inbox(inbox, exclude=bundle)
prefixes = _document_prefixes(inbox, walked)
written = nothing = failed = 0 written = nothing = failed = 0
for position, source in enumerate(walked, start=1): for position, source in enumerate(walked, start=1):
relative = source.relative_to(inbox) relative = source.relative_to(inbox)
@ -293,7 +351,7 @@ def _propose_plans(
plans_dir / f"{position:02d}.json", plans_dir / f"{position:02d}.json",
okf_type=okf_type, okf_type=okf_type,
proposed_at=proposed_at, proposed_at=proposed_at,
path_prefix=relative.with_suffix("").as_posix(), path_prefix=prefixes[source],
outline_run=outline_run, outline_run=outline_run,
table_grid=table_grid, table_grid=table_grid,
unit_fold=unit_fold, unit_fold=unit_fold,

View file

@ -610,6 +610,13 @@ def _xml_document(data: bytes) -> tuple[str, tuple[OutlineMark, ...]]:
document type declaration costs nothing here (0 of 1 file carries one) and document type declaration costs nothing here (0 of 1 file carries one) and
holds on every interpreter. holds on every interpreter.
""" """
root = _parse_xml(data)
reader = _XmlTextExtractor(sts=_is_sts(root))
return reader.text(root), tuple(reader.marks)
def _parse_xml(data: bytes) -> Element:
"""The one parse, with the DTD refusal in front of it (see `_xml_document`)."""
text = decode_text(data) text = decode_text(data)
prologue = text[: text.find("<", text.find("<") + 1) + 1] if "<" in text else text prologue = text[: text.find("<", text.find("<") + 1) + 1] if "<" in text else text
if "<!DOCTYPE" in prologue or "<!DOCTYPE" in text[:4096]: if "<!DOCTYPE" in prologue or "<!DOCTYPE" in text[:4096]:
@ -620,14 +627,81 @@ def _xml_document(data: bytes) -> tuple[str, tuple[OutlineMark, ...]]:
code="extractor_xml_doctype", code="extractor_xml_doctype",
) )
try: try:
root = ElementTree.fromstring(text) return ElementTree.fromstring(text)
except ElementTree.ParseError as exc: except ElementTree.ParseError as exc:
raise ExtractionError( raise ExtractionError(
f"the XML parser failed on this file: {exc}", code="extractor_xml_parse_error" f"the XML parser failed on this file: {exc}", code="extractor_xml_parse_error"
) from exc ) from exc
sts = _local_name(root.tag) == _STS_ROOT or next(root.iter("sec"), None) is not None
reader = _XmlTextExtractor(sts=sts)
return reader.text(root), tuple(reader.marks) def _is_sts(root: Element) -> bool:
"""The NAMED schema test: a `<standard>` root, or any `<sec>`."""
return _local_name(root.tag) == _STS_ROOT or next(root.iter("sec"), None) is not None
@dataclass(frozen=True)
class DeclaredIdentity:
"""What a document states about itself, read from its own elements.
Each field is `None` when the document does not state it, and ALSO when it
states it more than once: an adopted standard carries one `<std-ident>` per
body that issued it, and taking the first would be a guess dressed as a
reading. The caller falls back to the file name for whatever is `None`.
"""
doc_number: str | None
year: str | None
title: str | None
def declared_identity(name: str, data: bytes) -> DeclaredIdentity | None:
"""`xml`: the identity a NISO-STS document declares, or `None`.
MEASURED ON THE ONE STS DOCUMENT THIS ROW HAS: exactly one `<std-ident>`
(`<doc-number>R761 Prosesskoden</doc-number>` beside `<year>2025</year>`)
and one `<title-wrap>` whose `<full>` is the document's title -- while the
file carrying it was named for a delivery path, a UUID occurring 0 times in
the document. `<doc-type>` is read by nobody: it said `Innledning` there,
which is the name of a chapter and not a kind of document.
`None` for every other row and for XML that is not STS: a declaration is a
property of a schema, and a text that merely LOOKS like one declares
nothing. An unparseable file is `None` too, never an exception -- extracting
the same bytes refuses it with its own code, and an identity is not the
place a document is refused.
"""
if Path(name).suffix.lower() != ".xml":
return None
try:
root = _parse_xml(data)
except ExtractionError:
return None
if not _is_sts(root):
return None
declared = [
(_child_text(element, "doc-number"), _child_text(element, "year"))
for element in root.iter()
if _local_name(element.tag) == "std-ident"
]
declared = [pair for pair in declared if pair[0]]
doc_number, year = declared[0] if len(declared) == 1 else (None, None)
wraps = [element for element in root.iter() if _local_name(element.tag) == "title-wrap"]
title = (
(_child_text(wraps[0], "full") or _child_text(wraps[0], "main"))
if len(wraps) == 1
else None
)
if doc_number is None and title is None:
return None
return DeclaredIdentity(doc_number=doc_number, year=year, title=title)
def _child_text(element: Element, name: str) -> str | None:
"""A direct child's whole text, whitespace collapsed; `None` when absent or empty."""
for child in element:
if _local_name(child.tag) == name:
return " ".join("".join(child.itertext()).split()) or None
return None
def _extract_xml(data: bytes) -> str: def _extract_xml(data: bytes) -> str:

View file

@ -27,7 +27,7 @@ from dataclasses import dataclass, replace
from pathlib import Path, PurePosixPath from pathlib import Path, PurePosixPath
from .errors import IngestError, MaterializationError, SegmentationError, SourceError from .errors import IngestError, MaterializationError, SegmentationError, SourceError
from .extract import SourceUnits, extract_text, source_units from .extract import DeclaredIdentity, SourceUnits, declared_identity, extract_text, source_units
from .materialize import ( from .materialize import (
_render_root_frontmatter, _render_root_frontmatter,
check_filename_length, check_filename_length,
@ -113,6 +113,7 @@ def render_inbox_concept(
bundle_id: str | None = None, bundle_id: str | None = None,
units: SourceUnits | None = None, units: SourceUnits | None = None,
span: tuple[int, int] | None = None, span: tuple[int, int] | None = None,
source_title: str | None = None,
) -> str: ) -> str:
"""Frame extracted text as an inbox concept file with its provenance layer. """Frame extracted text as an inbox concept file with its provenance layer.
@ -127,6 +128,9 @@ def render_inbox_concept(
the text arriving here is the SANITIZED text and its length is not the text arriving here is the SANITIZED text and its length is not
necessarily the extracted text's. necessarily the extracted text's.
`source_title` is what the document calls itself, for the `sources`
entry's `title`; `None` keeps the file name there, as before it existed.
`segment` and `bundle_id` carry the 1-to-N identity layer and are read ONLY `segment` and `bundle_id` carry the 1-to-N identity layer and are read ONLY
when the profile declares the segmentation capability. A concept the plan when the profile declares the segmentation capability. A concept the plan
does not cover keeps today's rule verbatim, and the four shipped profiles does not cover keeps today's rule verbatim, and the four shipped profiles
@ -231,6 +235,7 @@ def render_inbox_concept(
source_file=source_file, source_file=source_file,
units=units, units=units,
span=located, span=located,
title=source_title,
) )
) )
return f"---\n{profile.frontmatter.emit(frontmatter)}\n---\n\n{_normalize_body(text)}" return f"---\n{profile.frontmatter.emit(frontmatter)}\n---\n\n{_normalize_body(text)}"
@ -249,6 +254,7 @@ def _provenance_frontmatter(
source_file: str, source_file: str,
units: SourceUnits | None, units: SourceUnits | None,
span: tuple[int, int] | None, span: tuple[int, int] | None,
title: str | None = None,
) -> dict[str, str]: ) -> dict[str, str]:
"""The address, and the locator when one is available. """The address, and the locator when one is available.
@ -265,11 +271,14 @@ def _provenance_frontmatter(
"than mangled", "than mangled",
code="inbox_source_file_unaddressable", code="inbox_source_file_unaddressable",
) )
values = { if title is not None and not _flow_expressible(title):
policy.sources_key: ( raise MaterializationError(
f"[{{ resource: {source_file}, title: {PurePosixPath(source_file).name} }}]" f"source title {title!r} cannot be written into the `sources` flow "
"mapping verbatim; refused rather than mangled",
code="inbox_source_title_unaddressable",
) )
} shown = title if title is not None else PurePosixPath(source_file).name
values = {policy.sources_key: f"[{{ resource: {source_file}, title: {shown} }}]"}
if units is None or span is None: if units is None or span is None:
return values return values
first, last = units.covering(*span) first, last = units.covering(*span)
@ -288,6 +297,54 @@ def _provenance_frontmatter(
return values return values
def _flow_expressible(value: str) -> bool:
"""Whether `value` survives as a plain scalar inside a flow mapping."""
return bool(value) and not any(char in value for char in f"{_FLOW_TERMINATORS}\n\r")
def _screened(gate: Gate, value: str | None) -> str | None:
"""A value read from the DOCUMENT and persisted outside its screened body.
The body goes through the gate before anything is written; a frontmatter
value taken from the same bytes would otherwise be the one route around
it. Kept only on the gate's non-blocking floor, as the SANITIZED text, and
dropped rather than refused otherwise: the body carrying the same words is
judged on its own, and a document is never lost over an optional key.
"""
if value is None:
return None
decision = gate(value)
if decision.disposition != _DISPOSITION_PERSIST:
return None
return " ".join(decision.sanitized_text.split()) or None
def _declared_sources_title(identity: DeclaredIdentity | None, gate: Gate) -> str | None:
"""The `sources` title a document declares, or `None` for the file name.
`<doc-number>` + `<year>` first, then the `<title-wrap>` title: measured on
the one STS document this row has, the `<full>` title carries a COMMA,
which ends a flow mapping, and the guard refuses the quoted scalar that
could have carried it. A declared value that cannot be written verbatim
falls to the next layer -- never cleaned up, because a title with its comma
removed is a title the document does not carry.
"""
if identity is None:
return None
candidates: list[str] = []
if identity.doc_number is not None:
candidates.append(
f"{identity.doc_number} {identity.year}" if identity.year else identity.doc_number
)
if identity.title is not None:
candidates.append(identity.title)
for candidate in candidates:
kept = _screened(gate, candidate)
if kept is not None and _flow_expressible(kept):
return kept
return None
# --- the guard seam ------------------------------------------------------- # --- the guard seam -------------------------------------------------------
# The guard's non-blocking floor. `Disposition` is a `str, Enum` in # The guard's non-blocking floor. `Disposition` is a `str, Enum` in
@ -620,6 +677,7 @@ def _render_segments(
bundle_id: str, bundle_id: str,
source_file: str, source_file: str,
units: SourceUnits | None, units: SourceUnits | None,
source_title: str | None = None,
) -> BlockedFile | None: ) -> BlockedFile | None:
"""Render every segment, or refuse the WHOLE document. """Render every segment, or refuse the WHOLE document.
@ -686,6 +744,7 @@ def _render_segments(
segment=entry, segment=entry,
bundle_id=bundle_id, bundle_id=bundle_id,
units=units, units=units,
source_title=source_title,
), ),
decision.reasons, decision.reasons,
) )
@ -996,6 +1055,14 @@ def process_inbox(
if profile.provenance is not None if profile.provenance is not None
else None else None
) )
# What the document says it is, for the address's title. Asked only
# where an address is written, so the four profiles without one do
# not parse anything they would never emit.
source_title = (
_declared_sources_title(declared_identity(source_name(path), source_bytes), gate)
if profile.provenance is not None
else None
)
covering = _plan_covering(plans, source_bytes) covering = _plan_covering(plans, source_bytes)
if covering is not None: if covering is not None:
blocked = _render_segments( blocked = _render_segments(
@ -1009,6 +1076,7 @@ def process_inbox(
bundle_id=(root_frontmatter_values or {})[_bundle_id_key(profile)], bundle_id=(root_frontmatter_values or {})[_bundle_id_key(profile)],
source_file=source_name(path), source_file=source_name(path),
units=units, units=units,
source_title=source_title,
) )
if blocked is not None: if blocked is not None:
if blocked.disposition == _DISPOSITION_QUARANTINE: if blocked.disposition == _DISPOSITION_QUARANTINE:
@ -1061,6 +1129,7 @@ def process_inbox(
# gate that removed a character would shift every # gate that removed a character would shift every
# unit boundary after it. # unit boundary after it.
span=(0, len(text)), span=(0, len(text)),
source_title=source_title,
), ),
decision.reasons, decision.reasons,
) )

View file

@ -113,8 +113,7 @@ def test_two_doc_numbers_are_no_doc_number() -> None:
""" """
data = IDENTITY.read_bytes().replace( data = IDENTITY.read_bytes().replace(
b"</std-doc-meta>", b"</std-doc-meta>",
b"<std-ident><doc-number>NS 9000</doc-number><year>2020</year></std-ident>" b"<std-ident><doc-number>NS 9000</doc-number><year>2020</year></std-ident></std-doc-meta>",
b"</std-doc-meta>",
) )
identity = extract.declared_identity(IDENTITY.name, data) identity = extract.declared_identity(IDENTITY.name, data)