feat(import): Door C flow against an injected import gate (Phase 2 step 5)
Reads an external OKF bundle as {bundle-relative path -> document text},
hands it WHOLE to an injected gate over the guard's okf.import_bundle (a
bundle-level call: it resolves the cross-link graph across concepts), and
merges only concepts clearing the non-blocking floor. Same injection pattern
as Door B, so the core stays dependency-free while the CI channel for the
real guard is settled.
Two constraints shaped the design and are pinned by tests:
- A merged concept is written VERBATIM. Stamping provenance into it would
require round-tripping its frontmatter through this library's line-oriented
parser, which cannot represent the block lists the guard's parser accepts --
silent data loss -- and would persist bytes the guard never screened.
- Ownership is therefore proven by content identity: identical bytes at the
target name are a no-op re-merge (re-import of an unchanged bundle is
idempotent), and anything else at the name is refused. Curated content and
an updated concept are refused alike; refusing is what never destroys.
The floor is fail-closed beyond the plan's "no error" wording: an error, an
unrecognised disposition, and a concept the gate returned no verdict for are
all refusals. quarantine_review stays its own bucket, as at Door B.
origin/channel are validated against the guard's pinned vocabulary -- it
derives trust from origin by enum identity, so an unrecognised string would be
silently downgraded rather than caught.
Three primitives promoted for reuse rather than duplicated:
reduce_to_id_grammar and check_filename_length to materialize.py, and
extract.decode_text. Door C slugs the WHOLE concept path, so tables/users.md
and views/users.md stay distinct. Concept discovery folds case explicitly
rather than globbing *.md, whose case-sensitivity follows the filesystem and
would import the same bundle differently on APFS and ext4.
README's "what is gated today" section corrected: it claimed nothing is gated,
which is no longer true, but the honest statement is narrower than "the doors
are gated" -- the library cannot verify that an injected adapter is a real
guard, and a permissive stub is believed.
405 tests green; ruff, ruff format and mypy --strict clean.
This commit is contained in:
parent
d812a839be
commit
f10fc60de2
11 changed files with 1250 additions and 61 deletions
|
|
@ -12,6 +12,7 @@ from __future__ import annotations
|
|||
import hashlib
|
||||
import logging
|
||||
import re
|
||||
import unicodedata
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
|
|
@ -64,6 +65,61 @@ def validate_ingested_at(ingested_at: str) -> None:
|
|||
)
|
||||
|
||||
|
||||
# --- generated-name primitives (shared by Doors B and C) ------------------
|
||||
|
||||
# Every character outside the Phase 1 id grammar (`[a-z0-9][a-z0-9-]*`) is a
|
||||
# separator. Deliberately NOT a transliteration: mapping non-ASCII letters to
|
||||
# ASCII ones would be a semantic claim the slugger cannot make — Norwegian
|
||||
# `møte` (meeting) would become `mote` (fashion). The readable name survives
|
||||
# verbatim in the concept's title or index label; the slug is an identifier,
|
||||
# not a label.
|
||||
_SEPARATOR_RUN_RE = re.compile(r"[^a-z0-9]+")
|
||||
|
||||
# NAME_MAX: the per-component limit on every filesystem this library targets
|
||||
# (APFS, ext4, NTFS all cap at 255). Checked before the write rather than
|
||||
# caught at it, because the OS signals it as an OSError whose errno differs per
|
||||
# platform (63 on macOS, 36 on Linux) — an untyped, unportable failure at the
|
||||
# very moment the caller needs a typed per-file outcome. Verified empirically
|
||||
# on APFS 2026-07-25: a 255-byte name writes, a 258-byte one raises errno 63.
|
||||
NAME_MAX_BYTES = 255
|
||||
|
||||
|
||||
def reduce_to_id_grammar(text: str) -> str:
|
||||
"""Reduce arbitrary text to the Phase 1 id grammar, or to the empty string.
|
||||
|
||||
Lowercased, with every run of non-grammar characters collapsed to a single
|
||||
`-` and both ends stripped. The caller decides what an empty result means:
|
||||
both doors refuse it rather than invent a fallback name.
|
||||
"""
|
||||
# NFC first: macOS (APFS/HFS+) hands filenames over DECOMPOSED, so an `é`
|
||||
# arrives as `e` + combining acute. Without normalising, the same visual
|
||||
# name reduces differently depending on where it came from — the combining
|
||||
# mark alone becomes a separator and the base letter survives (`cafe`),
|
||||
# where a composed `é` is one non-grammar character (`caf`). Composing
|
||||
# first makes the whole letter one unit, so non-ASCII is uniformly a
|
||||
# separator and the result is stable across both forms.
|
||||
return _SEPARATOR_RUN_RE.sub("-", unicodedata.normalize("NFC", text).lower()).strip("-")
|
||||
|
||||
|
||||
def check_filename_length(name: str, *, code: str) -> str:
|
||||
"""Refuse a generated filename the filesystem cannot hold.
|
||||
|
||||
Never truncated: truncation is lossy AND collision-prone (two long names
|
||||
sharing a prefix would reduce to one filename, and the second write would
|
||||
silently claim the first file). The message carries both the actual size
|
||||
and the limit, because the operator's fix is to shorten the source name.
|
||||
"""
|
||||
size = len(name.encode("utf-8"))
|
||||
if size > NAME_MAX_BYTES:
|
||||
raise MaterializationError(
|
||||
f"the generated filename would be {size} bytes, over the "
|
||||
f"{NAME_MAX_BYTES}-byte filesystem limit — shorten the source name; "
|
||||
"refusing to truncate (lossy and collision-prone)",
|
||||
code=code,
|
||||
)
|
||||
return name
|
||||
|
||||
|
||||
def _collapse_whitespace(value: str) -> str:
|
||||
# §5 mandates whitespace-run collapse for ONE field only: `source_query`
|
||||
# (ingest-spec.md:140-141), where a legitimately multi-line SQL SELECT
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue