feat(import): Door C flow against an injected import gate (Phase 2 step 5)

Reads an external OKF bundle as {bundle-relative path -> document text},
hands it WHOLE to an injected gate over the guard's okf.import_bundle (a
bundle-level call: it resolves the cross-link graph across concepts), and
merges only concepts clearing the non-blocking floor. Same injection pattern
as Door B, so the core stays dependency-free while the CI channel for the
real guard is settled.

Two constraints shaped the design and are pinned by tests:

- A merged concept is written VERBATIM. Stamping provenance into it would
  require round-tripping its frontmatter through this library's line-oriented
  parser, which cannot represent the block lists the guard's parser accepts --
  silent data loss -- and would persist bytes the guard never screened.
- Ownership is therefore proven by content identity: identical bytes at the
  target name are a no-op re-merge (re-import of an unchanged bundle is
  idempotent), and anything else at the name is refused. Curated content and
  an updated concept are refused alike; refusing is what never destroys.

The floor is fail-closed beyond the plan's "no error" wording: an error, an
unrecognised disposition, and a concept the gate returned no verdict for are
all refusals. quarantine_review stays its own bucket, as at Door B.
origin/channel are validated against the guard's pinned vocabulary -- it
derives trust from origin by enum identity, so an unrecognised string would be
silently downgraded rather than caught.

Three primitives promoted for reuse rather than duplicated:
reduce_to_id_grammar and check_filename_length to materialize.py, and
extract.decode_text. Door C slugs the WHOLE concept path, so tables/users.md
and views/users.md stay distinct. Concept discovery folds case explicitly
rather than globbing *.md, whose case-sensitivity follows the filesystem and
would import the same bundle differently on APFS and ext4.

README's "what is gated today" section corrected: it claimed nothing is gated,
which is no longer true, but the honest statement is narrower than "the doors
are gated" -- the library cannot verify that an injected adapter is a real
guard, and a permissive stub is believed.

405 tests green; ruff, ruff format and mypy --strict clean.
This commit is contained in:
Kjell Tore Guttormsen 2026-07-25 06:57:25 +02:00
commit f10fc60de2
11 changed files with 1250 additions and 61 deletions

View file

@ -12,6 +12,7 @@ from __future__ import annotations
import hashlib
import logging
import re
import unicodedata
from dataclasses import dataclass
from pathlib import Path
@ -64,6 +65,61 @@ def validate_ingested_at(ingested_at: str) -> None:
)
# --- generated-name primitives (shared by Doors B and C) ------------------
# Every character outside the Phase 1 id grammar (`[a-z0-9][a-z0-9-]*`) is a
# separator. Deliberately NOT a transliteration: mapping non-ASCII letters to
# ASCII ones would be a semantic claim the slugger cannot make — Norwegian
# `møte` (meeting) would become `mote` (fashion). The readable name survives
# verbatim in the concept's title or index label; the slug is an identifier,
# not a label.
_SEPARATOR_RUN_RE = re.compile(r"[^a-z0-9]+")
# NAME_MAX: the per-component limit on every filesystem this library targets
# (APFS, ext4, NTFS all cap at 255). Checked before the write rather than
# caught at it, because the OS signals it as an OSError whose errno differs per
# platform (63 on macOS, 36 on Linux) — an untyped, unportable failure at the
# very moment the caller needs a typed per-file outcome. Verified empirically
# on APFS 2026-07-25: a 255-byte name writes, a 258-byte one raises errno 63.
NAME_MAX_BYTES = 255
def reduce_to_id_grammar(text: str) -> str:
"""Reduce arbitrary text to the Phase 1 id grammar, or to the empty string.
Lowercased, with every run of non-grammar characters collapsed to a single
`-` and both ends stripped. The caller decides what an empty result means:
both doors refuse it rather than invent a fallback name.
"""
# NFC first: macOS (APFS/HFS+) hands filenames over DECOMPOSED, so an `é`
# arrives as `e` + combining acute. Without normalising, the same visual
# name reduces differently depending on where it came from — the combining
# mark alone becomes a separator and the base letter survives (`cafe`),
# where a composed `é` is one non-grammar character (`caf`). Composing
# first makes the whole letter one unit, so non-ASCII is uniformly a
# separator and the result is stable across both forms.
return _SEPARATOR_RUN_RE.sub("-", unicodedata.normalize("NFC", text).lower()).strip("-")
def check_filename_length(name: str, *, code: str) -> str:
"""Refuse a generated filename the filesystem cannot hold.
Never truncated: truncation is lossy AND collision-prone (two long names
sharing a prefix would reduce to one filename, and the second write would
silently claim the first file). The message carries both the actual size
and the limit, because the operator's fix is to shorten the source name.
"""
size = len(name.encode("utf-8"))
if size > NAME_MAX_BYTES:
raise MaterializationError(
f"the generated filename would be {size} bytes, over the "
f"{NAME_MAX_BYTES}-byte filesystem limit — shorten the source name; "
"refusing to truncate (lossy and collision-prone)",
code=code,
)
return name
def _collapse_whitespace(value: str) -> str:
# §5 mandates whitespace-run collapse for ONE field only: `source_query`
# (ingest-spec.md:140-141), where a legitimately multi-line SQL SELECT