# Intake boundaries: what will never enter through the inbox Two boundaries on Door B's intake are decided, not open. Both were left unstated for long enough to be mistaken for gaps someone would eventually close. They will not be closed, so they are written here instead — an unstated consequence is the one that bites later. ## 1. OCR is permanently out of scope **Decision: operator, 2026-09-01.** The library does no optical character recognition, and no work assumes it will. A PDF whose pages are raster images carries no extractable text layer, `pdfplumber` returns nothing for it, and the file is refused with `extractor_empty_pdf`. That refusal is correct and final, not a placeholder. **The consequence, stated plainly.** The `vegnormal-okf` PDFs are rastered 11 of 11 (measured, `docs/2026-08-29-konverteringsarmen-arm-a.md`). They will never enter through the inbox. Any statement of the form "the Vegnormalene PDFs in ⇒ bundle out" is not a goal this library holds, and the phrasing should not survive in planning documents. For the record, that premise never described anything that happened: `vegnormal-okf` builds its bundle from seven JSON files through **Door C**, not through the inbox. The inbox path has never been the route those concepts took. A consumer with rastered source documents needs a text layer produced outside this library before the inbox is involved. Where that text comes from is the consumer's decision; this library's answer is a coded refusal, not a best-effort guess. ## 2. Door B lowercases every path component `materialize.reduce_to_id_grammar` reduces each path component to `[a-z0-9][a-z0-9-]*`, and `segmentation.normalize_segment_path` applies the same grammar to segment paths. A source path of `krav/N500/id-` therefore becomes `krav/n500/id-`. **This is the library's Phase 1 grammar, not the specification's.** OKF v0.2 § 2 defines a concept ID as the file path minus `.md` and imposes no case rule. The grammar is stricter than the spec requires. **Decision: keep it, 2026-09-01.** The grammar is shared with Door A, which implements a spec `portfolio-optimiser-commons` owns, and it is pinned by byte-exact golden fixtures. Softening it on the segmentation path alone would put two grammars in one codebase — the same drift class that a single ordering helper was introduced to remove. **The consequence, stated plainly.** A source whose identities are case-sensitive cannot be reproduced verbatim through Door B. This is measured, not predicted: in `vegnormal-okf`'s B3 run against N500:2024, 267 of 270 concept identities matched case-folded and 0 matched verbatim — the entire difference across all 267 was case alone. A consumer that needs verbatim case-sensitive identities should use Door C, which writes external concepts verbatim, or accept the folded form. ## 3. Structured sources are not segmentable as they stand Measured by `vegnormal-okf` 2026-09-01 (B3, commit `0f8a27f`), with a known-positive control that held at 270 of 270 through their domain path: - The segment proposer returned **0 segments** for N500:2024. The raw file is 1 633 250 characters on a single line, and the proposer's rules are line-oriented. The same proposer returns 3 segments for a markdown document with headings, so the zero is a measurement rather than a broken query. - A hand-authored plan reproduced **267 of 270 identities but 0 of 267 bodies byte-identically**, because a span cuts the canonical extracted text — and for `.json` that text is the raw data inside a fenced block, not prose. The span model assumes the extracted text *is* the concept text. For a hierarchical element tree it is not. This is a structural limit of the current extraction layer, and unlike §§ 1 and 2 above it is **open work**, not a closed boundary. What is decided is only the shape of the eventual fix: the extraction layer gains a renderer a profile can name, so a structured source can be rendered rather than handed to segmentation as fenced raw bytes. A renderer that reproduces a specific domain's prose needs that domain's vocabulary and is not this library's to write without a commission.