Operator decision 2026-09-18, and the whole of it: okf removes the soft hyphen, the guard is untouched, and the number is BOOKED. `normalise_extracted(text) -> (text, removed)` in `extract`, applied once at the end of `extract_document`, so both entry points and both sides of a plan's `text_sha256` see one string. `_pdf_units` measures its page offsets through the same door, because that table is CHARACTER offsets rebuilt from the pages while the text it indexes comes back normalised -- two readings of one document, and a table built against the other names the wrong page with full confidence. Removing a character removes no newline, so every line-indexed rule downstream is unmoved by construction. `ExtractedDocument.soft_hyphens` carries the count out; `InboxResult.normalised` carries it per document; `account_run` reads it off the RUN rather than recounting the source, because a second count would be a second reader. It reaches the accounting JSON as `normalised_soft_hyphen` at both levels and `log.md` as a `**Normalisation**` bullet naming the count, the documents touched and that no other character is moved. EXPOSURE, measured and not assumed: 0 of the 78 readable documents of the K2 reference corpus carry U+00AD or any of the four real zero-width characters (the 8 unreadable ones raise before extraction), 0 in the pinned K2 bundle's concept bodies, and 0 across `tests/fixtures`, `examples`, `skills`, `docs`, `src`, README and CHANGELOG. The door cannot have moved a byte anyone here has measured. Suite 2179 passed, 1 skipped; ruff and `mypy --strict src/` clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2834 lines
126 KiB
Python
2834 lines
126 KiB
Python
"""Door B extraction registry: dropped file bytes -> text, per file type.
|
|
|
|
All file-type -> text extraction lives here (the guard is text-only). The core
|
|
registry is stdlib-only and deterministic: `md`/`txt` pass through, `csv` renders
|
|
the Phase 1 markdown table, `json` is fenced verbatim, and `html`/`htm` are
|
|
reduced to text with `html.parser`. Binary types are `[extract]`-gated and each
|
|
now has a reader: `pdf` through `pdfplumber`, and the five office rows
|
|
(`docx`/`xlsx`/`pptx`/`odt`/`rtf`) through a table-driven converter seam over
|
|
the vendored binary. Every one of those gates is an IMPORT PROBE rather than a
|
|
membership test, so an absent extra is rejected with the same typed error
|
|
whatever the type. Never a silent skip and never a bundled parser in core.
|
|
|
|
Two of the five office rows are `measured` and three are `unmeasured` -- the
|
|
corpus this arm was built on contains zero `pptx`, `odt` or `rtf` files, so
|
|
those rows work by construction and have never met a document anyone wrote.
|
|
`_EVIDENCE` carries that per row and the suite asserts it, because an
|
|
unmeasured row must not read as a supported one.
|
|
|
|
`extract_text` returns the extracted text *content*; final LF framing and the
|
|
concept frontmatter are the materializer's concern (Phase 2 step 2), not this
|
|
registry's. No guard call and no model call anywhere in this module.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import base64
|
|
import binascii
|
|
import collections
|
|
import csv
|
|
import functools
|
|
import io
|
|
import math
|
|
import re
|
|
import statistics
|
|
import tempfile
|
|
import urllib.parse
|
|
import warnings
|
|
import zipfile
|
|
from collections.abc import Callable, Sequence
|
|
from dataclasses import dataclass
|
|
from html.parser import HTMLParser
|
|
from pathlib import Path
|
|
from xml.etree import ElementTree
|
|
from xml.etree.ElementTree import Element
|
|
|
|
from .assets import (
|
|
PDF_FILTER_OUTPUT_RATIO,
|
|
AssetRejection,
|
|
ExtractedImage,
|
|
check_filter_cost,
|
|
check_payload,
|
|
check_size,
|
|
encode_png,
|
|
inflate_bounded,
|
|
inflate_limit_for,
|
|
inflated_size,
|
|
read_image,
|
|
render_block,
|
|
render_missing,
|
|
sniff,
|
|
)
|
|
from .connectors import safe_resolve
|
|
from .errors import ExtractionError, ExtractionWarning, SourceError
|
|
from .render import render_fenced_block, render_table
|
|
|
|
# Binary types gated behind the optional `[extract]` extra that it ships no
|
|
# parser for. EMPTY, and kept rather than deleted: the dispatch branch it feeds
|
|
# still raises `extractor_extra_missing`, and a later type that arrives before
|
|
# its reader belongs here rather than in a new mechanism. Every type the extra
|
|
# names now has a reader -- `.pdf` through the import probe in `_extract_pdf`,
|
|
# the five office rows through the converter seam below -- so the gate for all
|
|
# of them is an import probe, which is why the two tests for that code reach it
|
|
# that way.
|
|
_UNPARSED_OPTIONAL_EXTENSIONS: frozenset[str] = frozenset()
|
|
|
|
# The office rows: suffix -> the converter's reader name. THESE ROWS AND NO
|
|
# OTHERS. `.html` is excluded although the converter can read it: it already
|
|
# has a stdlib extractor here, so routing it through the converter would buy
|
|
# nothing and would add CVE-2025-51591 (SSRF via an iframe in HTML input),
|
|
# unpatched in every converter version. `.epub` is excluded on the "no gain"
|
|
# half of the same reason.
|
|
_PANDOC_FORMATS: dict[str, str] = {
|
|
".docx": "docx",
|
|
".xlsx": "xlsx",
|
|
".pptx": "pptx",
|
|
".odt": "odt",
|
|
".rtf": "rtf",
|
|
}
|
|
|
|
# What each row's behaviour actually rests on, asserted in the suite rather
|
|
# than written in a comment that rots. THREE classes, and the third exists
|
|
# because the first two could not tell the truth about these rows:
|
|
#
|
|
# measured real CORPUS files and a hand-counted fasit. Someone wrote the
|
|
# document for their own purposes and we counted what we got.
|
|
# constructed hand-built or generator-built documents with a hand-written
|
|
# fasit, and no corpus file at all. The row has now met a
|
|
# document end to end -- it is not `unmeasured` -- but a
|
|
# document written to exercise it is not a document anyone
|
|
# wrote, so it is not `measured` either.
|
|
# unmeasured the corpus contains ZERO files of the type AND no document
|
|
# has ever been put through the row. It works by construction.
|
|
#
|
|
# An unmeasured row must not read as a supported one, and neither must a
|
|
# constructed one.
|
|
#
|
|
# THE THREE OFFICE ROWS MOVED unmeasured -> constructed ON 2026-09-09, each on
|
|
# its own hand-built document (N = 1, except `pptx` at N = 2):
|
|
#
|
|
# .odt 1 of 1 declared headings recovered, 1 concept, 0 characters in no
|
|
# segment.
|
|
# .pptx 2 of 2 declared slide titles recovered on a deck that DECLARES them
|
|
# (`<p:ph type="title"/>`), 0 of 2 on one that does not -- the latter
|
|
# lands as `Slide 1`/`Slide 2`, which is the converter naming a slide
|
|
# it has no title for, not a segmentation failure.
|
|
# .rtf 0 declared headings, because the container has no heading style and
|
|
# the author's title is bold text. The proposer therefore proposes
|
|
# nothing and the document reaches Door B's INBOX as one concept --
|
|
# content preserved, structure zero. That is the row's honest result
|
|
# and it is the one open finding of the three.
|
|
#
|
|
# `.html` JOINED THE TABLE 2026-09-09, as `measured`, and the class was chosen
|
|
# against the definitions above rather than assumed: the 828 files are a
|
|
# consumer's own export of a real published handbook, produced for their
|
|
# ingestion and not to exercise this row, with a fasit written before any
|
|
# lookup -- which is `measured`'s test, "someone wrote the document for their
|
|
# own purposes and we counted what we got". What that class does NOT claim, and
|
|
# the honesty limit that travels with it: the 828 files are ONE product in ONE
|
|
# format from ONE publisher, and the file boundaries and `<h1>`s are a
|
|
# generator's cut of that document, not 828 documents anyone wrote.
|
|
#
|
|
# `.xml` JOINED THE TABLE 2026-09-11, as `measured`, and the class was read off
|
|
# the definitions above rather than inherited: the one file is a publisher's own
|
|
# NISO-STS delivery of R761, written for their purposes years before any lookup
|
|
# of ours, and its 2 761 titled `<sec>` are a fasit nobody here authored. The
|
|
# honesty limit that travels with it and does NOT move when the build reaches
|
|
# the reader's ceiling: the denominator is ONE file, ONE publisher, ONE schema.
|
|
# `.xml` as a file type is far wider than NISO-STS, and a document in any other
|
|
# schema keeps its text in document order and gets no structure at all -- which
|
|
# is measured on fixtures, not on a corpus.
|
|
#
|
|
# `.pdf` JOINED THE TABLE 2026-09-10, as `measured`, and it enters on the
|
|
# strongest evidence of any row here: eight real corpus PDFs with a fasit the
|
|
# operator hand-counted document by document, plus a 701-page process code
|
|
# whose PUBLISHER also ships a NISO-STS structure for it -- 2 761 titled
|
|
# sections, written for their own purposes and long before any lookup of ours.
|
|
# The honesty limit that travels with it: those 2 761 rows are ONE product in
|
|
# ONE format from ONE publisher, its structure is a strict numbered hierarchy
|
|
# on 2 739 of 2 761 titles, and a running prose document would measure
|
|
# something else entirely.
|
|
_EVIDENCE: dict[str, str] = {
|
|
".pdf": "measured",
|
|
".docx": "measured",
|
|
".xlsx": "measured",
|
|
".pptx": "constructed",
|
|
".odt": "constructed",
|
|
".rtf": "constructed",
|
|
".html": "measured",
|
|
".xml": "measured",
|
|
}
|
|
|
|
# Load-bearing, all three, and none of them hygiene:
|
|
#
|
|
# --eol=lf the defaults produce DIFFERENT BYTES (maximum line length 75
|
|
# --wrap=none against 447), which a byte-pinned golden registers as a change
|
|
# nobody made.
|
|
# -t markdown never `-t plain`: plain destroys the headings the segment
|
|
# proposer reads. Measured -- a document yielding 15 entries
|
|
# including two real headings yields 13 with none under `plain`,
|
|
# so the writer choice silently sets the ceiling for the arm
|
|
# downstream of it.
|
|
_PANDOC_WRITER = "markdown"
|
|
_PANDOC_ARGS = ("--eol=lf", "--wrap=none")
|
|
|
|
# The spreadsheet row writes PIPE tables, and it is the only row that does.
|
|
#
|
|
# The default writer prefers simple tables, which pad every cell out to the
|
|
# width of the widest cell in its column. Measured on the K2 price sheet: one
|
|
# 594-character prose cell turned every other row in that column into a run of
|
|
# up to 887 spaces between a label and its amount, 100 795 characters in all,
|
|
# and the header row named ONE column because only the first cell of the source
|
|
# row 1 is filled. The bytes reached the reader and the structure did not. The
|
|
# same sheet through this writer is 11 221 characters with no whitespace run
|
|
# longer than two, one row per line, each source column its own cell.
|
|
#
|
|
# `--columns=1` is load-bearing rather than cosmetic: the pipe writer pads cells
|
|
# out to the column width it computes from that setting, so at the default 72 a
|
|
# NARROW table gains runs of up to 45 spaces -- the same defect at a smaller
|
|
# scale. Measured across every office fixture and every K2 office file, the
|
|
# longest whitespace run with it is 2.
|
|
#
|
|
# SPREADSHEET-ONLY, deliberately. The other four rows have the same defect
|
|
# available to the same one-line fix (measured: the odt fixture 1366 -> 1105
|
|
# characters), but a spreadsheet IS a grid with no prose fallback, while moving
|
|
# the prose rows would move a corpus denominator that nothing has measured.
|
|
# `tests/test_extract.py` pins that scoping with three digests.
|
|
_SPREADSHEET_WRITER = "markdown-simple_tables-multiline_tables-grid_tables"
|
|
_SPREADSHEET_ARGS = (*_PANDOC_ARGS, "--columns=1")
|
|
|
|
# SpreadsheetML's namespace, needed to read the workbook's shared string table.
|
|
_SSML = "http://schemas.openxmlformats.org/spreadsheetml/2006/main"
|
|
|
|
# A table cell whose whole content is an integer with the converter's trailing
|
|
# `.0`. Bounded by unescaped pipes on both sides so a cell containing an escaped
|
|
# `\|` can never be split in the middle.
|
|
_INTEGRAL_CELL = re.compile(r"(?<!\\)\|(\s*)(-?\d+)\.0(\s*)(?=(?<!\\)\|)")
|
|
|
|
# Conversion recovers text, on the same terms as PDF extraction: a drawing has
|
|
# no text to recover. Said out loud on every conversion rather than detected
|
|
# per document, for the same reason.
|
|
_OFFICE_LOSSY_WARNING = (
|
|
"office-file conversion recovers text only: figures, diagrams, images and "
|
|
"drawn shapes are not represented in the output (their captions are). A "
|
|
"bundle built from drawn documents is incomplete by construction."
|
|
)
|
|
|
|
# Text extraction recovers text. A figure is a vector drawing with no text to
|
|
# recover — only its caption survives — so any bundle built from drawn
|
|
# documents is incomplete by construction. Said out loud on every PDF rather
|
|
# than detected per document: deciding "is there a figure here" is a layout
|
|
# heuristic this library does not own.
|
|
_PDF_LOSSY_WARNING = (
|
|
"PDF extraction recovers text only: figures, diagrams and images are not "
|
|
"represented in the output (their captions are). A bundle built from "
|
|
"drawn documents is incomplete by construction."
|
|
)
|
|
|
|
# Tags whose text content is never document prose.
|
|
#
|
|
# TWO MEMBERS, and it stays two in round 11. Dropping `nav`/`header`/`footer`
|
|
# as well is a DIFFERENT change with a different guarantee: the exact
|
|
# text-preservation invariant below holds only while nothing is dropped, and a
|
|
# quiet widening here would hide exactly how many characters left the document.
|
|
_SKIP_TAGS = frozenset({"script", "style"})
|
|
|
|
# The heading tags, and the ATX level each becomes. The level is the TAG's:
|
|
# a flat `#` for every heading would hand `propose._ATX` three top-level
|
|
# boundaries where the document declares one section and two subsections.
|
|
_HEADING_TAGS: dict[str, int] = {f"h{level}": level for level in range(1, 7)}
|
|
|
|
# Tags that open a line of their own. Everything NOT here is inline and stays a
|
|
# word boundary inside the current line, which is what `b`/`em`/`a`/`span` were
|
|
# already treated as.
|
|
#
|
|
# WHY THIS IS A SET AND NOT THE FIVE TAGS THE CORPUS EXERCISES. Block versus
|
|
# inline is a property of HTML, not of one corpus. The measured corpus writes
|
|
# its prose in `p`, `li` and `tr`; a `div`-structured page -- the ordinary shape
|
|
# of hand-written and exported HTML -- carries the same prose in containers this
|
|
# corpus never uses, and restricting the set to what was measured would leave
|
|
# that page collapsing into one line, which IS the defect. Adding a line break
|
|
# never removes text and never invents a boundary on its own: a boundary needs a
|
|
# line that MATCHES a grammar.
|
|
_BLOCK_TAGS = frozenset(
|
|
{
|
|
"p",
|
|
"li",
|
|
"tr",
|
|
"pre",
|
|
"div",
|
|
"section",
|
|
"article",
|
|
"header",
|
|
"footer",
|
|
"nav",
|
|
"main",
|
|
"aside",
|
|
"table",
|
|
"thead",
|
|
"tbody",
|
|
"tfoot",
|
|
"caption",
|
|
"ul",
|
|
"ol",
|
|
"dl",
|
|
"dt",
|
|
"dd",
|
|
"blockquote",
|
|
"figure",
|
|
"figcaption",
|
|
"hr",
|
|
"address",
|
|
"form",
|
|
"fieldset",
|
|
"legend",
|
|
"title",
|
|
"body",
|
|
}
|
|
)
|
|
|
|
|
|
# --- NISO-STS, and the two facts the whole reader turns on ------------------
|
|
#
|
|
# ONE: `<label>` carries the number and `<title>` carries the text. Measured on
|
|
# the 701-page process code, 2 of its 2 761 `<title>` strings begin with a
|
|
# digit -- the number is a sibling element, never glued on. A fasit that shows
|
|
# `"2.1Hovedprosesser"` is its BUILDER joining the two. Emitting `<title>`
|
|
# alone therefore scores 0 of 2 761 while every line of this file looks right,
|
|
# because the number is what okf reduces to a directory name.
|
|
#
|
|
# TWO: a `<sec>` with a `<label>` and no `<title>` is not a section heading. It
|
|
# is a lettered point (`a)`, `c)`, `sec-type="spec"`) inside a process
|
|
# description, and there are 4 954 of them against the document's own 2 761.
|
|
# One heading each and the document's structure is the minority of its own
|
|
# outline.
|
|
_STS_ROOT = "standard"
|
|
|
|
# Inline by allowlist, block by default -- the INVERSE of the HTML reader, and
|
|
# for the reason that reader gives for its own direction. Block versus inline
|
|
# is a property of HTML; XML has no such universal, so an unknown element
|
|
# cannot be assumed inline without fusing two paragraphs into one line. It can
|
|
# safely be assumed block: an extra line break never removes text and never
|
|
# invents a boundary, because a boundary needs a line that MATCHES a grammar.
|
|
#
|
|
# The members are NISO-STS's own inline set, and they are load-bearing rather
|
|
# than decorative: that document carries 1 701 `<italic>` and 1 396 `<bold>`
|
|
# inside its prose, so breaking on them would shred a paragraph into fragments
|
|
# that are individually true and collectively unreadable.
|
|
_XML_INLINE_TAGS = frozenset(
|
|
{
|
|
"italic",
|
|
"bold",
|
|
"underline",
|
|
"sup",
|
|
"sub",
|
|
"sc",
|
|
"monospace",
|
|
"roman",
|
|
"sans-serif",
|
|
"overline",
|
|
"strike",
|
|
"xref",
|
|
"ext-link",
|
|
"uri",
|
|
"std-ref",
|
|
"inline-formula",
|
|
"styled-content",
|
|
"named-content",
|
|
"break",
|
|
}
|
|
)
|
|
|
|
# The maximum ATX level `propose._ATX` can read (`#{1,6}`), which `_HEADING_TAGS`
|
|
# stops at for the same reason. STS nesting goes DEEPER: 9 of the 2 761 titled
|
|
# sections in that document sit at depth 7, and `#######` matches nothing at
|
|
# all. The depth is CLIPPED rather than dropped -- a clipped heading still sets
|
|
# its boundary, where a dropped one loses the section entirely. The clip is the
|
|
# HEADING's alone (K3-21): the `OutlineMark` beside it carries the declared
|
|
# depth, because the declared route builds its plan from the mark and a plan
|
|
# reading 7 as 6 gave a depth-7 section the ancestor one level too high.
|
|
_ATX_MAX_LEVEL = 6
|
|
|
|
|
|
# --- assets: the images a document carries beside its text ------------------
|
|
#
|
|
# ADDED IN 0.10.0, and off unless the caller asks. Until then no reader here
|
|
# fetched, named or copied an image: `<img>`'s attributes were never read, an
|
|
# STS `<graphic>` was walked past, a PDF was opened for `extract_text` alone,
|
|
# and the converter's markdown writer dropped every picture. The two lossiness
|
|
# warnings above said so on every run, which made the loss honest and did not
|
|
# make it smaller.
|
|
#
|
|
# THE READER PLACES, THE COLLECTOR DECIDES. Each reader knows where in its own
|
|
# document an image stands and what the document calls it; nothing else. What
|
|
# an image IS, what it is named in the bundle and how it is pointed at is
|
|
# `assets.py`'s, so the four readers cannot drift into four grammars, and
|
|
# "carried N of M" means one thing across all of them.
|
|
|
|
#: How a reader asks for bytes the document only POINTS at. The inbox supplies
|
|
#: one rooted at the dropped file's own directory and refuses to leave the drop
|
|
#: -- an `<img src="../../../etc/passwd">` is a path traversal written by
|
|
#: whoever wrote the document, which is exactly the class of input this package
|
|
#: treats as untrusted. `None` means the caller gave no resolver, and every
|
|
#: pointer then resolves to nothing rather than to a guess.
|
|
Resolver = Callable[[str], bytes | None]
|
|
|
|
#: A scheme this package will not open. Extraction opens no socket at all --
|
|
#: network access here is an explicit per-run opt-in and extraction is not on
|
|
#: that path -- so a remote image is carried as a POINTER and never as bytes.
|
|
_REMOTE_SOURCE = re.compile(r"^(?:[a-zA-Z][a-zA-Z0-9+.-]*:|//)")
|
|
|
|
_DATA_URI = re.compile(r"^data:(?P<media>[^;,]*)(?P<base64>;base64)?,(?P<payload>.*)$", re.DOTALL)
|
|
|
|
|
|
#: The ONE character the normalisation door removes, and the reason it is one
|
|
#: character and not a class. `llm-ingestion-guard` 1.4.0 puts U+00AD in
|
|
#: `_ZERO_WIDTH_CPS` beside U+200B, U+200C, U+200D and U+FEFF, and
|
|
#: `output:zero-width-present` is an any-tier carrier: `fail_secure` at every
|
|
#: trust level, with no sanitisation and no exception. Measured by PM
|
|
#: 2026-09-18, R761 Prosesskoden:2025 carries 71 U+00AD and 0 of the other
|
|
#: four, so a 701-page process code is unreadable for the whole chain over
|
|
#: Norwegian hyphenation points inside words -- `ar[SHY]beider`,
|
|
#: `bitu[SHY]men`, `asfalt[SHY]betong`. The verdict is formally right and
|
|
#: materially a false positive, and the operator's answer (2026-09-18) is that
|
|
#: okf removes this character before the guard and SAYS SO in the accounting.
|
|
#:
|
|
#: The other four stay. They carry no typographic job in running text, so
|
|
#: removing one would be a decision about what the guard screens for, taken in
|
|
#: the wrong repository. U+00A0 NBSP is not in the guard's set at all and is
|
|
#: not touched either -- R761 ships 6 633 of them.
|
|
SOFT_HYPHEN = "\u00ad"
|
|
|
|
|
|
def normalise_extracted(text: str) -> tuple[str, int]:
|
|
"""The normalisation door: the text without U+00AD, and how many were removed.
|
|
|
|
Applied ONCE, at the end of :func:`extract_document`, so every caller of
|
|
either entry point gets the same string and `propose` and Door B cannot
|
|
disagree about the text a plan's `text_sha256` indexes. The only other
|
|
place that has to know about it is :func:`_pdf_units`, which rebuilds a
|
|
table of CHARACTER offsets from the pages rather than from the returned
|
|
text: two readings of one document, and a table built against the other
|
|
one would name the wrong page with full confidence.
|
|
|
|
Removing a character never removes a newline, so every LINE-indexed rule
|
|
downstream -- the proposer's grammars, `xml_outline`, `pdf_outline`'s
|
|
per-page line check -- is unmoved by construction.
|
|
"""
|
|
removed = text.count(SOFT_HYPHEN)
|
|
if not removed:
|
|
return text, 0
|
|
return text.replace(SOFT_HYPHEN, ""), removed
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class ExtractedDocument:
|
|
"""One dropped file's text, and the images that stand inside that text.
|
|
|
|
The text is what it always was when `images` and `rejected` are empty --
|
|
which is what the `assets=False` default guarantees byte for byte, so every
|
|
golden, every pinned bundle and every published digest is unmoved by this
|
|
capability existing.
|
|
"""
|
|
|
|
text: str
|
|
images: tuple[ExtractedImage, ...] = ()
|
|
rejected: tuple[AssetRejection, ...] = ()
|
|
#: The references, relative to the document's own directory, whose bytes
|
|
#: the resolver returned and that were CARRIED as images. Recorded where
|
|
#: the resolution happened rather than inferred from bytes afterwards: an
|
|
#: unpointed file with the same bytes as a carried one was carried through
|
|
#: nothing (R761 ships eight such duplicates).
|
|
files: tuple[str, ...] = ()
|
|
#: How many U+00AD the normalisation door removed from this text. Zero for
|
|
#: every document that carried none, which is 0 of the 78 readable
|
|
#: documents of the K2 reference corpus (measured 2026-09-19).
|
|
soft_hyphens: int = 0
|
|
|
|
|
|
class _AssetCollector:
|
|
"""What a reader hands an image to, and the only place a rejection is made.
|
|
|
|
A rejection is a ROW, never a failed document: one unreadable picture must
|
|
not cost the text around it. Every rejection also leaves a line in the
|
|
concept saying what was there and why it is not here -- an absence a reader
|
|
is never shown is an absence they cannot weigh, which is the defect this
|
|
whole capability closes one level up.
|
|
"""
|
|
|
|
def __init__(self, resolve: Resolver | None = None) -> None:
|
|
self._resolve = resolve
|
|
self.images: list[ExtractedImage] = []
|
|
self.rejected: list[AssetRejection] = []
|
|
self.files: list[str] = []
|
|
|
|
def carry(self, data: bytes, *, name: str, label: str | None = None) -> str:
|
|
"""Bytes the reader already holds, as the block that stands in the text."""
|
|
try:
|
|
image = read_image(data, name=name, label=label)
|
|
except ExtractionError as exc:
|
|
return self.reject(name, code=exc.code, reason=str(exc), label=label)
|
|
self.images.append(image)
|
|
return render_block(image)
|
|
|
|
def reject(
|
|
self,
|
|
name: str,
|
|
*,
|
|
code: str,
|
|
reason: str,
|
|
label: str | None = None,
|
|
href: str | None = None,
|
|
) -> str:
|
|
self.rejected.append(AssetRejection(name=name, code=code, reason=reason))
|
|
return render_missing(name, reason=reason, label=label, href=href)
|
|
|
|
def local(self, source: str, *, label: str | None = None, sibling: str | None = None) -> str:
|
|
"""An href the DOCUMENT points at: a data URI, a local path, or remote.
|
|
|
|
`sibling` is a SECOND path to try, and it exists because one format's
|
|
convention is not another's: a NISO-STS delivery writes a bare file
|
|
name and ships the files in a `graphics/` directory beside the
|
|
document. Passing the candidate rather than teaching this method about
|
|
STS keeps the convention with the reader that has it, and the
|
|
containment rule with the resolver that owns it.
|
|
"""
|
|
data_uri = _DATA_URI.match(source)
|
|
if data_uri is not None:
|
|
return self._data_uri(data_uri, label=label)
|
|
if _REMOTE_SOURCE.match(source):
|
|
return self.reject(
|
|
source,
|
|
code="asset_remote",
|
|
reason="the source is off this machine and extraction opens no socket",
|
|
label=label,
|
|
href=source,
|
|
)
|
|
found = source
|
|
data = self._resolve(source) if self._resolve is not None else None
|
|
if data is None and sibling is not None and sibling != source and self._resolve is not None:
|
|
found = sibling
|
|
data = self._resolve(sibling)
|
|
if data is None:
|
|
return self.reject(
|
|
source,
|
|
code="asset_unresolved",
|
|
reason="the file the document points at was not found beside it",
|
|
label=label,
|
|
)
|
|
carried = len(self.images)
|
|
block = self.carry(data, name=source, label=label)
|
|
if len(self.images) > carried:
|
|
self.files.append(found)
|
|
return block
|
|
|
|
def _data_uri(self, match: re.Match[str], *, label: str | None) -> str:
|
|
payload = match.group("payload")
|
|
name = f"data-uri{Path(match.group('media').split('/')[-1] or 'bin').suffix or ''}"
|
|
# Before decoding: base64 expands by 3/4, and a percent-encoded payload
|
|
# by at most 1. An inline picture is small by nature, so a payload over
|
|
# the bound is refused unread -- the same bomb class as a PDF image
|
|
# declaring an enormous size. A refusal is a ROW, like every other.
|
|
try:
|
|
check_payload(
|
|
len(payload) * 3 // 4 if match.group("base64") else len(payload), name=name
|
|
)
|
|
except ExtractionError as exc:
|
|
return self.reject(name, code=exc.code, reason=str(exc), label=label)
|
|
try:
|
|
raw = (
|
|
base64.b64decode(payload, validate=True)
|
|
if match.group("base64")
|
|
else urllib.parse.unquote_to_bytes(payload)
|
|
)
|
|
except (binascii.Error, ValueError) as exc:
|
|
return self.reject(
|
|
name,
|
|
code="asset_unresolved",
|
|
reason=f"the inline data URI could not be decoded: {exc}",
|
|
label=label,
|
|
)
|
|
return self.carry(raw, name=name, label=label)
|
|
|
|
|
|
def directory_resolver(root: Path) -> Resolver:
|
|
"""A resolver rooted at ONE directory, fail-closed, reading nothing else.
|
|
|
|
Containment is against the DOCUMENT'S OWN directory rather than against the
|
|
whole drop, and that is not caution for its own sake: the proposer reads a
|
|
file straight off disk and the door reads it out of the inbox, and a plan
|
|
indexes the exact string it was proposed against. One root both sides can
|
|
compute from the document alone is what makes the two renderings identical
|
|
without threading a second path through either. A reference above the
|
|
document's directory is refused (`asset_unresolved`) rather than followed;
|
|
the limit is stated in the concept, like every other rejection.
|
|
|
|
`safe_resolve` is Door A's own rule, reused verbatim: `..` traversal, an
|
|
absolute path, a symlink escape and a prefix-collision sibling all fail
|
|
closed. An `<img src="../../../etc/passwd">` is a path written by whoever
|
|
wrote the document, which is exactly the class of input this package treats
|
|
as untrusted.
|
|
"""
|
|
|
|
def resolve(relative: str) -> bytes | None:
|
|
try:
|
|
target = safe_resolve(root, relative)
|
|
except SourceError:
|
|
return None
|
|
try:
|
|
if not target.is_file():
|
|
return None
|
|
return target.read_bytes()
|
|
except OSError:
|
|
return None
|
|
|
|
return resolve
|
|
|
|
|
|
def decode_text(data: bytes) -> str:
|
|
"""Decode file bytes as UTF-8 (BOM-stripping), typed on failure.
|
|
|
|
utf-8-sig so a byte-order mark never leaks into the first character
|
|
(baseline parity with Door A's read_csv). A non-UTF-8 file is a corrupt
|
|
input: fail fast with a typed error rather than leaking UnicodeDecodeError.
|
|
"""
|
|
try:
|
|
return data.decode("utf-8-sig")
|
|
except UnicodeDecodeError as exc:
|
|
raise ExtractionError(
|
|
f"file bytes are not valid UTF-8: {exc}", code="extractor_decode_error"
|
|
) from exc
|
|
|
|
|
|
def _extract_passthrough(data: bytes) -> str:
|
|
"""`md`/`txt`: the decoded text verbatim."""
|
|
return decode_text(data)
|
|
|
|
|
|
def _extract_csv(data: bytes) -> str:
|
|
"""`csv`: parse with the stdlib reader, render the Phase 1 markdown table."""
|
|
reader = csv.reader(io.StringIO(decode_text(data)))
|
|
header = next(reader, None)
|
|
if header is None:
|
|
raise ExtractionError("CSV has no header row", code="extractor_empty_csv")
|
|
rows = list(reader)
|
|
return render_table(header, rows)
|
|
|
|
|
|
def _extract_json(data: bytes) -> str:
|
|
"""`json`: the decoded text verbatim inside a fenced block (Phase 1 renderer)."""
|
|
return render_fenced_block(decode_text(data))
|
|
|
|
|
|
class _HTMLTextExtractor(HTMLParser):
|
|
"""Collect document text as LINES, skipping `script`/`style`.
|
|
|
|
A block tag opens a line of its own, a heading tag opens one carrying the
|
|
ATX marker for its level, `br` breaks the current line, and every other tag
|
|
stays what it always was: a word boundary inside the line, so adjacent
|
|
inline text (``<b>a</b>b``) does not fuse. Runs of whitespace inside a line
|
|
collapse to single spaces.
|
|
|
|
THE OUTPUT GRAMMAR IS MARKDOWN, and deliberately the same markdown the
|
|
office rows reach the proposer through. `_ATX` and every other boundary
|
|
grammar is line-anchored, so this class decides -- alone -- whether an HTML
|
|
document can be segmented at all. It emitted one line for any input until
|
|
2026-09-09, which is why 828 of 828 real sections produced zero boundaries.
|
|
|
|
TEXT IS PRESERVED EXACTLY. The only characters this adds are the ATX
|
|
markers; strip those and the non-whitespace sequence is identical to the
|
|
one-line form. Nothing is ever dropped here beyond `_SKIP_TAGS`.
|
|
"""
|
|
|
|
def __init__(self, collector: _AssetCollector | None = None) -> None:
|
|
super().__init__(convert_charrefs=True)
|
|
self._lines: list[str] = []
|
|
self._current: list[str] = []
|
|
self._prefix = ""
|
|
self._skip_depth = 0
|
|
self._collector = collector
|
|
|
|
def _break(self, prefix: str = "") -> None:
|
|
"""Close the line being accumulated and open the next one."""
|
|
line = " ".join("".join(self._current).split())
|
|
self._current = []
|
|
if line:
|
|
self._lines.append(f"{self._prefix}{line}")
|
|
self._prefix = prefix
|
|
|
|
def _image(self, attrs: list[tuple[str, str | None]]) -> None:
|
|
"""An `<img>`, in the place it stands.
|
|
|
|
`attrs` was read by nothing here until 0.10.0, so `src` and `alt` were
|
|
both dropped -- a document's figures left no trace in the concept at
|
|
all, not even their alt text. The block goes on its own lines because
|
|
markdown wants a paragraph and because every boundary grammar in
|
|
`propose.py` is line-anchored.
|
|
"""
|
|
if self._collector is None or self._skip_depth:
|
|
return
|
|
values = {key: value or "" for key, value in attrs}
|
|
source = values.get("src", "").strip()
|
|
if not source:
|
|
return
|
|
self._break()
|
|
for line in self._collector.local(source, label=values.get("alt") or None).split("\n"):
|
|
self._lines.append(line)
|
|
self._break()
|
|
|
|
def _open(self, tag: str) -> bool:
|
|
"""Break for a block or heading tag; report whether it was one."""
|
|
level = _HEADING_TAGS.get(tag)
|
|
if level is not None:
|
|
self._break("#" * level + " ")
|
|
return True
|
|
if tag in _BLOCK_TAGS or tag == "br":
|
|
self._break()
|
|
return True
|
|
return False
|
|
|
|
def handle_starttag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None:
|
|
if not self._open(tag):
|
|
self._current.append(" ")
|
|
if tag in _SKIP_TAGS:
|
|
self._skip_depth += 1
|
|
if tag == "img":
|
|
self._image(attrs)
|
|
|
|
def handle_startendtag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None:
|
|
if not self._open(tag):
|
|
self._current.append(" ")
|
|
if tag == "img":
|
|
self._image(attrs)
|
|
|
|
def handle_endtag(self, tag: str) -> None:
|
|
if tag in _SKIP_TAGS and self._skip_depth > 0:
|
|
self._skip_depth -= 1
|
|
if not self._open(tag):
|
|
self._current.append(" ")
|
|
|
|
def handle_data(self, data: str) -> None:
|
|
if self._skip_depth == 0:
|
|
self._current.append(data)
|
|
|
|
def text(self) -> str:
|
|
self._break()
|
|
return "\n".join(self._lines)
|
|
|
|
|
|
def _extract_html(data: bytes, collector: _AssetCollector | None = None) -> str:
|
|
"""`html`/`htm`: text via `html.parser`, script/style stripped (spec B3)."""
|
|
parser = _HTMLTextExtractor(collector)
|
|
parser.feed(decode_text(data))
|
|
parser.close()
|
|
return parser.text()
|
|
|
|
|
|
class _XmlTextExtractor:
|
|
"""Collect an XML document's text as LINES of markdown.
|
|
|
|
THE OUTPUT GRAMMAR IS MARKDOWN, and deliberately the same markdown the
|
|
office rows and the HTML row reach the proposer through. Nothing in
|
|
`propose.py` knows this format exists: `_ATX` reads the heading lines and
|
|
`_TABLE_ROW` reads the table ones, exactly as they read a converted
|
|
`.docx`. That is the whole reason this is an extractor and not a
|
|
segmentation arm.
|
|
|
|
TWO PATHS, and which one runs is NAMED rather than guessed:
|
|
|
|
STS the root is `<standard>` or the document contains a `<sec>`.
|
|
A `<sec>` with a `<title>` becomes one ATX line whose level is
|
|
its `<sec>`-nesting depth and whose text is `<label>` + space +
|
|
`<title>`; a `<sec>` with only a `<label>` becomes a body line
|
|
with the label in front of it, the way `li` is treated in HTML;
|
|
a `<table-wrap>` becomes its label on a line and its rows as one
|
|
markdown table block.
|
|
GENERIC everything else. Text content in document order, one line per
|
|
block-like element, and NO element name is ever promoted to a
|
|
heading. An RSS feed is not a numbered standard, and reading it
|
|
as one would state a structure its author did not.
|
|
|
|
TEXT IS PRESERVED EXACTLY. The only characters added are the ATX markers
|
|
and the table pipes; strip those and the non-whitespace sequence is
|
|
identical to `"".join(root.itertext())`.
|
|
"""
|
|
|
|
def __init__(self, *, sts: bool, collector: _AssetCollector | None = None) -> None:
|
|
self._sts = sts
|
|
self._collector = collector
|
|
self._lines: list[str] = []
|
|
self._current: list[str] = []
|
|
self._prefix = ""
|
|
# The declared structure, recorded WHERE it is written rather than
|
|
# recovered from the finished string. The PDF arm has to bridge from
|
|
# (page, `/XYZ` top) onto a line index and was wrong on 1 840 of 2 762
|
|
# nodes under the naive rule; here the reader appended the line, so the
|
|
# index is not a guess and carries no tolerance. Empty for a document
|
|
# that is not STS -- that is "this schema declares no section", and it
|
|
# must not collapse into "this document has no structure to state".
|
|
self.marks: list[OutlineMark] = []
|
|
|
|
def _break(self, prefix: str = "") -> None:
|
|
"""Close the line being accumulated and open the next one.
|
|
|
|
A pending prefix SURVIVES a break that emitted nothing -- a label-only
|
|
`<sec>` holds its `a)` until the first line that has words in it, which
|
|
may be several empty elements later.
|
|
"""
|
|
line = " ".join("".join(self._current).split())
|
|
self._current = []
|
|
if line:
|
|
self._lines.append(f"{self._prefix}{line}")
|
|
self._prefix = ""
|
|
if prefix:
|
|
# A prefix still pending here belongs to a section that turned out
|
|
# to have no body line at all, and REPLACING it would drop it from
|
|
# the document. Measured on R761: exactly one `x)`, two characters,
|
|
# which is the whole distance between 0.999998 and exact.
|
|
if self._prefix:
|
|
self._lines.append(self._prefix.rstrip())
|
|
self._prefix = prefix
|
|
|
|
def _emit(self, line: str) -> None:
|
|
"""Put a whole line out, ahead of whatever is being accumulated.
|
|
|
|
A prefix still pending is FLUSHED first rather than carried past a
|
|
heading: the section it belongs to is above this one, and holding it
|
|
would either attach it to the wrong body or lose it outright.
|
|
"""
|
|
self._break()
|
|
if self._prefix:
|
|
self._lines.append(self._prefix.rstrip())
|
|
self._prefix = ""
|
|
self._lines.append(line)
|
|
|
|
def _text_of(self, element: Element) -> str:
|
|
"""An element's whole text, whitespace collapsed."""
|
|
return " ".join("".join(element.itertext()).split())
|
|
|
|
def _spec_point(self, section: Element) -> str | None:
|
|
"""The FIRST `<p>` of the FIRST direct-child `sec-type="spec"`, whole.
|
|
|
|
The limit is this package's, not the spec's: SPEC SS 4.1 asks for "a
|
|
single sentence" and sets no length anywhere. It is STRUCTURAL rather
|
|
than a character count, because a cut inside a paragraph writes a
|
|
sentence the source never wrote. Measured on the one STS document this
|
|
row has: 2 026 of 2 761 titled sections carry a direct-child spec
|
|
point; 264 of those points hold more than one `<p>` and 2 hold none;
|
|
the first `<p>` runs 17 / 109 / 273 / 521 / 942 characters at min /
|
|
median / p90 / p99 / max.
|
|
|
|
A DIRECT child only: a spec point belongs to the section it opens under,
|
|
and a container borrowing its first child's would describe a section by
|
|
a sentence about another one.
|
|
"""
|
|
for child in section:
|
|
if _local_name(child.tag) == "sec" and child.get("sec-type") == "spec":
|
|
for paragraph in child:
|
|
if _local_name(paragraph.tag) == "p":
|
|
return self._text_of(paragraph) or None
|
|
return None
|
|
return None
|
|
|
|
def _table(self, element: Element) -> bool:
|
|
"""A `<table-wrap>`: its label on a line, its rows as ONE table block.
|
|
|
|
The separator line is what makes it a block rather than two pipe lines
|
|
-- `--table-grid` and `--keep-table-heading` read the block, and the
|
|
PDF path delivered 0 of this document's 10 tables as one.
|
|
"""
|
|
rows = [
|
|
[self._text_of(cell) for cell in row if cell.tag in ("td", "th")]
|
|
for row in element.iter("tr")
|
|
]
|
|
rows = [row for row in rows if row]
|
|
if not rows:
|
|
return False
|
|
label = element.find("label")
|
|
if label is not None:
|
|
self._emit(self._text_of(label))
|
|
caption = element.find("caption")
|
|
if caption is not None:
|
|
self._emit(self._text_of(caption))
|
|
self._break()
|
|
self._lines.extend(render_table(rows[0], rows[1:]).rstrip("\n").split("\n"))
|
|
return True
|
|
|
|
def _graphic(self, element: Element) -> bool:
|
|
"""A `<graphic>`, in the place it stands. Reports whether it was one.
|
|
|
|
Measured on the R761 delivery, 2026-09-16: 50 `<graphic>` elements, all
|
|
50 direct children of a `<sec>`, none inside a `<table-wrap>`, none
|
|
carrying a caption element of any kind -- the "Figur 11.1 ..." line a
|
|
human reads is a sibling `<p>` this reader already emits on its own
|
|
line. So the label falls back to the file name rather than being
|
|
guessed from the neighbourhood.
|
|
|
|
TWO RESOLUTION ROUTES, and the second is the delivery's own convention:
|
|
the href as written, and then `graphics/<name>`, because that publisher
|
|
writes a BARE file name and ships the files in a sibling directory.
|
|
Both are tried through the caller's resolver, which is what keeps the
|
|
containment rule in one place.
|
|
"""
|
|
if self._collector is None:
|
|
return False
|
|
href = next(
|
|
(value for key, value in element.attrib.items() if _local_name(key) == "href"),
|
|
None,
|
|
)
|
|
if not href:
|
|
return False
|
|
block = self._collector.local(href, sibling=f"graphics/{Path(href).name}")
|
|
self._emit_lines(block.split("\n"))
|
|
return True
|
|
|
|
def _emit_lines(self, lines: list[str]) -> None:
|
|
for line in lines:
|
|
self._emit(line)
|
|
|
|
def _walk(self, element: Element, depth: int) -> None:
|
|
tag = _local_name(element.tag)
|
|
if tag in ("graphic", "inline-graphic") and self._graphic(element):
|
|
return
|
|
if self._sts and tag == "table-wrap" and self._table(element):
|
|
return
|
|
skip: set[int] = set()
|
|
if self._sts and tag == "sec":
|
|
depth += 1
|
|
label = element.find("label")
|
|
title = element.find("title")
|
|
if title is not None:
|
|
level = min(depth, _ATX_MAX_LEVEL)
|
|
parts = [self._text_of(label)] if label is not None else []
|
|
parts.append(self._text_of(title))
|
|
heading = " ".join(part for part in parts if part)
|
|
self._emit("#" * level + " " + heading)
|
|
self.marks.append(
|
|
OutlineMark(
|
|
line=len(self._lines) - 1,
|
|
# The DECLARED depth, never the clipped one: the mark
|
|
# is not a markdown heading, and the plan the declared
|
|
# route builds from it reads nesting off this level.
|
|
level=depth,
|
|
title=heading,
|
|
description=self._spec_point(element),
|
|
)
|
|
)
|
|
skip = {id(title)} | ({id(label)} if label is not None else set())
|
|
elif label is not None:
|
|
# NEVER a heading. The label goes in FRONT of the body line the
|
|
# way `li` is treated in the HTML reader.
|
|
self._break(self._text_of(label) + " ")
|
|
skip = {id(label)}
|
|
inline = _local_name(element.tag) in _XML_INLINE_TAGS
|
|
if not inline:
|
|
self._break()
|
|
else:
|
|
self._current.append(" ")
|
|
if element.text:
|
|
self._current.append(element.text)
|
|
for child in element:
|
|
if id(child) not in skip:
|
|
self._walk(child, depth)
|
|
if child.tail:
|
|
self._current.append(child.tail)
|
|
if not inline:
|
|
self._break()
|
|
|
|
def text(self, root: Element) -> str:
|
|
self._walk(root, 0)
|
|
self._break()
|
|
if self._prefix:
|
|
self._lines.append(self._prefix.rstrip())
|
|
return "\n".join(self._lines)
|
|
|
|
|
|
def _xml_document(
|
|
data: bytes, collector: _AssetCollector | None = None
|
|
) -> tuple[str, tuple[OutlineMark, ...]]:
|
|
"""`xml`: NISO-STS structure as markdown, any other schema as its text.
|
|
|
|
A DTD IS REFUSED RATHER THAN PARSED, and that is a guarantee about this
|
|
code instead of one about the machine. Measured on this interpreter
|
|
(3.14.0, `pyexpat.version_info` 2.7.3): an external `SYSTEM` entity is
|
|
refused by the stdlib and never fetched, but the entity-amplification limit
|
|
that stops a billion-laughs comes from libexpat >= 2.4.0 and NOT from
|
|
Python -- five levels still expanded, six and seven were refused -- while
|
|
`pyproject.toml` requires only `>=3.10` and no lockfile pins an
|
|
interpreter. `XMLParser` exposes no `.parser` attribute on the C
|
|
accelerator either, so the handler route is not portable. Refusing every
|
|
document type declaration costs nothing here (0 of 1 file carries one) and
|
|
holds on every interpreter.
|
|
"""
|
|
root = _parse_xml(data)
|
|
reader = _XmlTextExtractor(sts=_is_sts(root), collector=collector)
|
|
return reader.text(root), tuple(reader.marks)
|
|
|
|
|
|
def _parse_xml(data: bytes) -> Element:
|
|
"""The one parse, with the DTD refusal in front of it (see `_xml_document`)."""
|
|
text = decode_text(data)
|
|
prologue = text[: text.find("<", text.find("<") + 1) + 1] if "<" in text else text
|
|
if "<!DOCTYPE" in prologue or "<!DOCTYPE" in text[:4096]:
|
|
raise ExtractionError(
|
|
"XML carrying a document type declaration is refused unparsed: a DTD can "
|
|
"define entities, and the parser's amplification limit is a property of "
|
|
"the installed libexpat rather than of this package",
|
|
code="extractor_xml_doctype",
|
|
)
|
|
try:
|
|
return ElementTree.fromstring(text)
|
|
except ElementTree.ParseError as exc:
|
|
raise ExtractionError(
|
|
f"the XML parser failed on this file: {exc}", code="extractor_xml_parse_error"
|
|
) from exc
|
|
|
|
|
|
def _is_sts(root: Element) -> bool:
|
|
"""The NAMED schema test: a `<standard>` root, or any `<sec>`."""
|
|
return _local_name(root.tag) == _STS_ROOT or next(root.iter("sec"), None) is not None
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class DeclaredIdentity:
|
|
"""What a document states about itself, read from its own elements.
|
|
|
|
Each field is `None` when the document does not state it, and ALSO when it
|
|
states it more than once: an adopted standard carries one `<std-ident>` per
|
|
body that issued it, and taking the first would be a guess dressed as a
|
|
reading. The caller falls back to the file name for whatever is `None`.
|
|
"""
|
|
|
|
doc_number: str | None
|
|
year: str | None
|
|
title: str | None
|
|
|
|
|
|
def declared_identity(name: str, data: bytes) -> DeclaredIdentity | None:
|
|
"""`xml`: the identity a NISO-STS document declares, or `None`.
|
|
|
|
MEASURED ON THE ONE STS DOCUMENT THIS ROW HAS: exactly one `<std-ident>`
|
|
(`<doc-number>R761 Prosesskoden</doc-number>` beside `<year>2025</year>`)
|
|
and one `<title-wrap>` whose `<full>` is the document's title -- while the
|
|
file carrying it was named for a delivery path, a UUID occurring 0 times in
|
|
the document. `<doc-type>` is read by nobody: it said `Innledning` there,
|
|
which is the name of a chapter and not a kind of document.
|
|
|
|
`None` for every other row and for XML that is not STS: a declaration is a
|
|
property of a schema, and a text that merely LOOKS like one declares
|
|
nothing. An unparseable file is `None` too, never an exception -- extracting
|
|
the same bytes refuses it with its own code, and an identity is not the
|
|
place a document is refused.
|
|
"""
|
|
if Path(name).suffix.lower() != ".xml":
|
|
return None
|
|
try:
|
|
root = _parse_xml(data)
|
|
except ExtractionError:
|
|
return None
|
|
if not _is_sts(root):
|
|
return None
|
|
declared = [
|
|
(_child_text(element, "doc-number"), _child_text(element, "year"))
|
|
for element in root.iter()
|
|
if _local_name(element.tag) == "std-ident"
|
|
]
|
|
declared = [pair for pair in declared if pair[0]]
|
|
doc_number, year = declared[0] if len(declared) == 1 else (None, None)
|
|
wraps = [element for element in root.iter() if _local_name(element.tag) == "title-wrap"]
|
|
title = (
|
|
(_child_text(wraps[0], "full") or _child_text(wraps[0], "main"))
|
|
if len(wraps) == 1
|
|
else None
|
|
)
|
|
if doc_number is None and title is None:
|
|
return None
|
|
return DeclaredIdentity(doc_number=doc_number, year=year, title=title)
|
|
|
|
|
|
def _child_text(element: Element, name: str) -> str | None:
|
|
"""A direct child's whole text, whitespace collapsed; `None` when absent or empty."""
|
|
for child in element:
|
|
if _local_name(child.tag) == name:
|
|
return " ".join("".join(child.itertext()).split()) or None
|
|
return None
|
|
|
|
|
|
def _extract_xml(data: bytes, collector: _AssetCollector | None = None) -> str:
|
|
return _xml_document(data, collector)[0]
|
|
|
|
|
|
def xml_outline(
|
|
name: str, data: bytes, *, assets: bool = False, resolve: Resolver | None = None
|
|
) -> tuple[OutlineMark, ...]:
|
|
"""`xml`: the sections the document DECLARES, as marks on the extracted text.
|
|
|
|
The counterpart of `pdf_outline`, and the difference between them is the
|
|
whole point of the row. A bookmark states a page and a y position, so that
|
|
arm has to BRIDGE onto a line and reports what did not bridge; an STS
|
|
`<sec><title>` is written into the output by this reader, so the line index
|
|
is the one it appended at -- nothing is recovered, nothing is unresolved,
|
|
and there is no tolerance constant to choose.
|
|
|
|
RE-READS the bytes rather than returning both from one call, for the same
|
|
reason `pdf_outline` does: `extract_text` has one signature that every
|
|
caller and every registry entry is keyed to, and a second return value
|
|
would change it for eight rows to serve one. The parse is stdlib and the
|
|
document is read twice; measured on a 2.4 MB NISO-STS file, that is the
|
|
smaller cost by a wide margin.
|
|
|
|
Empty for every schema that is not STS. That is a statement about the
|
|
document -- it declares no section -- and `find_candidates` reads an empty
|
|
list as "leave every rule untouched", never as a route.
|
|
"""
|
|
del name # the registry decides which reader runs; kept for `pdf_outline`'s shape
|
|
# `assets` and `resolve` are NOT options of this arm, exactly as
|
|
# `pdf_headings` is not one of `pdf_outline`'s: carrying an image inserts
|
|
# lines into the extracted text, so marks computed with the images off name
|
|
# the right sections at the wrong line numbers. They are threaded so both
|
|
# sides of the plan can be computed against ONE rendering.
|
|
return _xml_document(data, _AssetCollector(resolve) if assets else None)[1]
|
|
|
|
|
|
def _local_name(tag: str) -> str:
|
|
"""`{ns}sec` -> `sec`. A namespaced document names the same elements."""
|
|
return tag.rsplit("}", 1)[-1]
|
|
|
|
|
|
def _extra_missing(suffix: str) -> ExtractionError:
|
|
"""The one rejection for a `[extract]` type without the extra installed.
|
|
|
|
One constructor, one wording: the import probe and the still-unparsed
|
|
types must be indistinguishable to a consumer, because they are the same
|
|
fact — the extra is not installed.
|
|
"""
|
|
return ExtractionError(
|
|
f"extracting {suffix!r} requires the optional 'extract' extra "
|
|
f"(pip install 'llm-ingestion-okf[extract]'); it is not installed",
|
|
code="extractor_extra_missing",
|
|
)
|
|
|
|
|
|
def _ocr_group_missing() -> ExtractionError:
|
|
"""The one rejection for `--ocr` without the optional `ocr` group.
|
|
|
|
A DIFFERENT code from `extractor_extra_missing`, because it is a different
|
|
fact and a different remedy: the `[extract]` extra can be fully installed
|
|
-- the document parsed, the pages counted -- and the OCR engine still be
|
|
absent. One error naming both would send an operator to reinstall
|
|
something they already have.
|
|
"""
|
|
return ExtractionError(
|
|
"reading a PDF page with OCR requires the optional 'ocr' group "
|
|
"(pip install 'llm-ingestion-okf[extract,ocr]'), which ships rapidocr "
|
|
"on onnxruntime; it is not installed",
|
|
code="extractor_ocr_group_missing",
|
|
)
|
|
|
|
|
|
#: The literal placeholder `pdfminer.six` (behind `pdfplumber`) emits for a
|
|
#: glyph whose font carries no usable ToUnicode mapping. The text is present on
|
|
#: the page and unreadable in the extraction -- a failure that looks like
|
|
#: success, which is why it needs a measurement rather than an exception.
|
|
_CID_CODE = re.compile(r"\(cid:\d+\)")
|
|
|
|
#: The share of a page's extracted characters that must be `(cid:N)` codes
|
|
#: before `--ocr` reads the page as an image instead.
|
|
#:
|
|
#: MEASURED, not chosen: `docs/2026-09-08-k3-runde4-pdf-skrift-og-ocr.md`
|
|
#: reports the per-page distribution over the K2 corpus, and it is bimodal
|
|
#: with nothing in between -- one document's pages sit near 1.0 and every other
|
|
#: page in the corpus sits at 0.0. Any value in that gap selects the same
|
|
#: pages, which is what makes 0.10 defensible and also what makes it
|
|
#: uninformative about a corpus that has intermediate pages. Stated rather than
|
|
#: implied: this threshold is bounded by the corpus, not by a property of the
|
|
#: format.
|
|
OCR_CID_SHARE = 0.10
|
|
|
|
#: The resolution a page is rendered at before it is read as an image.
|
|
#: 200 dpi is what the round-4 measurement was taken at; the engine's own
|
|
#: preprocessing rescales from there, so this is a floor on how much of the
|
|
#: page's detail reaches it rather than a tuning knob. It is part of the output
|
|
#: contract in the same way the parser version is: OCR text is deterministic
|
|
#: within one resolution and one model version, and across neither.
|
|
OCR_DPI = 200
|
|
|
|
|
|
def cid_share(text: str) -> float:
|
|
"""The share of `text` made of `(cid:N)` placeholder codes, 0.0 for empty.
|
|
|
|
Module level and importable: `tools/okf_cid_measure.py` answers the same
|
|
question at DOCUMENT level, and two definitions of one metric drift.
|
|
"""
|
|
if not text:
|
|
return 0.0
|
|
return sum(len(match.group(0)) for match in _CID_CODE.finditer(text)) / len(text)
|
|
|
|
|
|
def needs_ocr(text: str) -> bool:
|
|
"""Whether a page's extracted text is unusable enough to read the image.
|
|
|
|
TWO conditions, because there are two ways a page's text never arrives and
|
|
they look nothing alike: a page with no text layer extracts as the empty
|
|
string, and a page whose fonts carry no ToUnicode mapping extracts as a
|
|
full page of `(cid:N)`. A trigger written for one of them would leave the
|
|
other exactly where it was.
|
|
"""
|
|
return not text.strip() or cid_share(text) >= OCR_CID_SHARE
|
|
|
|
|
|
def _ocr_reader() -> Callable[[object], list[str]]:
|
|
"""The OCR engine, or the typed refusal. The import IS the gate.
|
|
|
|
Same shape as `_extract_pdf`'s probe and for the same reason: membership in
|
|
a suffix set cannot tell whether a package is importable, and this group is
|
|
the one a consumer is most likely not to have.
|
|
"""
|
|
try:
|
|
import rapidocr
|
|
except ImportError as exc:
|
|
raise _ocr_group_missing() from exc
|
|
if rapidocr is None: # pragma: no cover - the sys.modules probe in tests
|
|
raise _ocr_group_missing()
|
|
engine = rapidocr.RapidOCR()
|
|
|
|
def read(image: object) -> list[str]:
|
|
result = engine(image)
|
|
# `txts` is None when the detector found nothing at all, which is a
|
|
# legitimate answer for a blank page and not an error.
|
|
return [str(line) for line in (getattr(result, "txts", None) or ())]
|
|
|
|
return read
|
|
|
|
|
|
#: Bold as a PDF says it: in the font's NAME (`Helvetica-Bold`,
|
|
#: `ABCDEF+Arial-BoldMT`). There is no weight attribute on a character, so the
|
|
#: name is the only place a text extractor can read it.
|
|
_PDF_BOLD_MARKER = "bold"
|
|
|
|
#: The deepest ATX level the emitted markdown may use. `_ATX` in `propose.py`
|
|
#: reads one to six hashes, and a document with seven distinct heading sizes
|
|
#: would otherwise emit a line the proposer reads as body.
|
|
_PDF_MAX_HEADING_LEVEL = 6
|
|
|
|
|
|
def _dominant(values: list[str]) -> str:
|
|
"""The most frequent value, ties broken by first occurrence.
|
|
|
|
`Counter.most_common(1)` reduces to `max` over the items in insertion
|
|
order, so the tie-break is document order and the result is deterministic
|
|
for identical bytes -- which is the property everything downstream is
|
|
pinned to.
|
|
"""
|
|
return collections.Counter(values).most_common(1)[0][0]
|
|
|
|
|
|
def _typography(line: dict[str, object]) -> tuple[float, str] | None:
|
|
"""One line's dominant font size and font name, or `None` if it is blank.
|
|
|
|
Blank characters are excluded from both: a space carries a size and a font
|
|
like any other character, and a heading padded with body-sized spaces would
|
|
read as body.
|
|
"""
|
|
chars = [char for char in line["chars"] if str(char["text"]).strip()] # type: ignore[attr-defined]
|
|
if not chars:
|
|
return None
|
|
sizes = [f"{float(char['size']):.1f}" for char in chars]
|
|
fonts = [str(char["fontname"]) for char in chars]
|
|
return float(_dominant(sizes)), _dominant(fonts)
|
|
|
|
|
|
def _heading_levels(lines: list[tuple[str, float, str]]) -> dict[float, int]:
|
|
"""Which font sizes are headings in this document, and at what ATX level.
|
|
|
|
The rule is the CONJUNCTION this repository already measured: larger than
|
|
the body AND bold. `docs/2026-09-07-k3-arm-d.md`'s predecessor measured
|
|
size-and-bold from poppler at recall 1.000 / precision 0.846, and measured
|
|
that adding weight as a DISJUNCT made precision worse (0.786 -> 0.524). A
|
|
disjunction here would mark every emphasised phrase in the body.
|
|
|
|
The body size is the CHARACTER-weighted median over the whole document, not
|
|
the page: a title page is 100 % heading by line count, and a per-page
|
|
median would compare it with itself and mark nothing. Weighted by
|
|
characters rather than lines for the same reason in miniature -- a document
|
|
front-loaded with short lines has a line median that no paragraph shares.
|
|
|
|
The ATX LEVEL is the size's rank among the heading sizes, largest first, so
|
|
a document's own typographic hierarchy survives into the markdown instead
|
|
of flattening to one level. Deeper than six is clamped, because `_ATX`
|
|
reads six.
|
|
"""
|
|
weighted: list[float] = []
|
|
for text, size, _ in lines:
|
|
weighted.extend([size] * len(text.replace(" ", "")))
|
|
if not weighted:
|
|
return {}
|
|
body = statistics.median(weighted)
|
|
sizes = {size for _, size, font in lines if size > body and _PDF_BOLD_MARKER in font.lower()}
|
|
return {
|
|
size: min(rank, _PDF_MAX_HEADING_LEVEL)
|
|
for rank, size in enumerate(sorted(sizes, reverse=True), start=1)
|
|
}
|
|
|
|
|
|
def _mark_headings(lines: list[tuple[str, float, str]], levels: dict[float, int]) -> str:
|
|
"""One page's lines as markdown, the heading sizes carrying their hashes.
|
|
|
|
BOLD is checked again here rather than folded into the size map: a document
|
|
can set a caption in the same size as a heading without setting it bold,
|
|
and a map keyed on size alone would promote it.
|
|
"""
|
|
out: list[str] = []
|
|
for text, size, font in lines:
|
|
level = levels.get(size) if _PDF_BOLD_MARKER in font.lower() else None
|
|
out.append(f"{'#' * level} {text}" if level is not None and text else text)
|
|
return "\n".join(out)
|
|
|
|
|
|
# How `_extract_pdf` joins its pages, named because the locator below has to
|
|
# reproduce the exact same arithmetic to turn a character offset back into a
|
|
# page number. Two constants that must agree, written once.
|
|
_PDF_PAGE_SEPARATOR = "\n\n"
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class _PdfPage:
|
|
"""One page's text, and the images drawn on it, kept APART on purpose.
|
|
|
|
The pointer blocks are appended after the page's own lines by
|
|
`_pdf_page_text`, and the body is kept separately because `pdf_outline`
|
|
compares the page's line splitting against `page.extract_text_lines()` --
|
|
a per-page check that ships and that decides whether the primary bridge
|
|
route may be used at all. Appended lines are not in that geometry, so a
|
|
joined string would fail the check on every page carrying an image and
|
|
silently demote 2 762 bookmarks to the fallback route.
|
|
"""
|
|
|
|
number: int
|
|
text: str
|
|
images: tuple[ExtractedImage, ...] = ()
|
|
rejected: tuple[AssetRejection, ...] = ()
|
|
|
|
|
|
def _pdf_page_text(page: _PdfPage) -> str:
|
|
"""A page as it reaches the extracted text: its lines, then its pointers.
|
|
|
|
END OF PAGE, not the image's y position, and the reason is stated rather
|
|
than hidden: a PDF image has a bounding box and no place in the reading
|
|
order, so "where it stands" is the page. Inserting by y would reorder the
|
|
page's own lines against the geometry `pdf_outline` checks itself against,
|
|
and would put a pointer inside a sentence. A caption printed above a figure
|
|
therefore keeps its own line where the document put it, and the pointer
|
|
follows the page it was drawn on.
|
|
"""
|
|
blocks = [render_block(image) for image in page.images]
|
|
blocks += [
|
|
render_missing(rejection.name, reason=rejection.reason) for rejection in page.rejected
|
|
]
|
|
if not blocks:
|
|
return page.text
|
|
joined = "\n\n".join(blocks)
|
|
return f"{page.text}\n\n{joined}" if page.text else joined
|
|
|
|
|
|
#: The bits-per-component this encoder expresses. A PDF may store 1, 2, 4, 8 or
|
|
#: 16, and everything but 8 is REFUSED with a code rather than rescaled --
|
|
#: rescaling a 1-bit stencil to 8 bits is a decision about what black means, and
|
|
#: a wrong one looks exactly like a right one.
|
|
_PDF_SAMPLE_BITS = 8
|
|
|
|
|
|
def _pdf_colour(space: object) -> tuple[int, bytes | None] | None:
|
|
"""A PDF colour space as `(channels, palette)`, or `None` if not expressible.
|
|
|
|
`None` is the honest answer for CMYK, for a separation space and for
|
|
anything with a transfer function: converting those needs a colour model
|
|
this package does not carry, and a guess would be a picture that is
|
|
plausibly the wrong colour. It is counted and stated, never approximated.
|
|
"""
|
|
from pdfminer.pdftypes import PDFStream, resolve1
|
|
|
|
space = resolve1(space)
|
|
name = getattr(space, "name", None)
|
|
if name in ("DeviceGray", "CalGray", "G"):
|
|
return 1, None
|
|
if name in ("DeviceRGB", "CalRGB", "RGB"):
|
|
return 3, None
|
|
if not isinstance(space, list) or not space:
|
|
return None
|
|
head = getattr(resolve1(space[0]), "name", None)
|
|
if head == "ICCBased" and len(space) > 1:
|
|
profile = resolve1(space[1])
|
|
components = resolve1(profile.attrs.get("N")) if isinstance(profile, PDFStream) else None
|
|
return (int(components), None) if components in (1, 3) else None
|
|
if head in ("CalGray",):
|
|
return 1, None
|
|
if head in ("CalRGB", "Lab"):
|
|
return 3, None
|
|
if head in ("Indexed", "I") and len(space) >= 4:
|
|
base = _pdf_colour(space[1])
|
|
if base is None:
|
|
return None
|
|
lookup = resolve1(space[3])
|
|
if isinstance(lookup, PDFStream):
|
|
lookup = lookup.get_data()
|
|
if not isinstance(lookup, bytes):
|
|
return None
|
|
if base[0] == 3:
|
|
palette = lookup[: (len(lookup) // 3) * 3]
|
|
else:
|
|
# PNG's PLTE is RGB triples only, so a grey palette is widened
|
|
# rather than refused. Widening a grey to r=g=b is exact, not an
|
|
# approximation -- which is why this branch exists and the CMYK one
|
|
# does not.
|
|
palette = b"".join(bytes([value, value, value]) for value in lookup)
|
|
return (1, palette) if palette else None
|
|
return None
|
|
|
|
|
|
def _pdf_alpha(attrs: dict[str, object], width: int, height: int) -> bytes | None | bool:
|
|
"""A soft mask as one alpha byte per pixel, `None` for none, `False` to refuse.
|
|
|
|
An `SMask` this encoder cannot express is a REFUSAL rather than a dropped
|
|
channel: an image whose transparency is thrown away is composited against
|
|
nothing and reads as a black or white rectangle over the page, which is a
|
|
picture that is wrong in a way no consumer can detect.
|
|
"""
|
|
from pdfminer.pdftypes import PDFStream, resolve1
|
|
|
|
mask = resolve1(attrs.get("SMask"))
|
|
if mask is None:
|
|
return None
|
|
if not isinstance(mask, PDFStream):
|
|
return False
|
|
shape = mask.attrs
|
|
if (
|
|
resolve1(shape.get("Width")) != width
|
|
or resolve1(shape.get("Height")) != height
|
|
or resolve1(shape.get("BitsPerComponent")) != _PDF_SAMPLE_BITS
|
|
):
|
|
return False
|
|
try:
|
|
alpha = mask.get_data()
|
|
except Exception:
|
|
return False
|
|
return alpha if len(alpha) >= width * height else False
|
|
|
|
|
|
def bounded_pdf_filters() -> frozenset[str]:
|
|
"""The PDF stream filters an image may be reached through, by NAME.
|
|
|
|
TWO CLASSES, and what separates them is HOW the cost of a link is bounded,
|
|
never whether the link is safe. Every one of them is bounded.
|
|
|
|
* `FlateDecode` is MEASURED as it is paid: inflated a chunk at a time,
|
|
refused the moment the running total crosses the bound, with the output
|
|
discarded unless a link behind it has to be measured from those bytes.
|
|
* `ASCII85Decode`, `ASCIIHexDecode`, `DCTDecode`, `JPXDecode` and
|
|
`JBIG2Decode` are PREDICTED before they are paid: each carries a measured
|
|
worst-case cost per byte of input (`assets.PDF_FILTER_COST_RATIO`), and a
|
|
link whose input times that ratio is over the budget is refused before
|
|
its decoder is called. Their decoders take a whole string and return a
|
|
whole string, so there is no moment between the two at which a cost could
|
|
be observed.
|
|
|
|
0.10.1 had a third class, and it was WRONG. `ASCII85Decode` and
|
|
`ASCIIHexDecode` were called bounded "by their own input because they
|
|
shrink". `z` is ASCII85's shorthand for four zero bytes, so that filter
|
|
QUADRUPLES its input, and `base64.a85decode` appends one 4-byte object per
|
|
group to a list, so it costs about a hundred bytes of memory per byte of
|
|
input. Measured on the pinned tree, its own interpreter, idle machine: a
|
|
33 475-byte PDF decoding through `[/FlateDecode /ASCII85Decode]` cost
|
|
3 261 599 744 bytes of peak RSS and the picture was CARRIED. Under the
|
|
ratios it is 42 070 016 bytes and `asset_too_large`, and at twice the run
|
|
of `z` -- which trebled the old cost to 6 461 558 784 -- it is 40 280 064:
|
|
the cost no longer follows the bomb.
|
|
|
|
EVERYTHING ELSE IS REFUSED with `asset_pdf_unbounded` before any of the
|
|
stream is decoded -- `LZWDecode`, `RunLengthDecode`, `CCITTFaxDecode`,
|
|
`/Crypt`, and any filter written after this one. A filter with no measured
|
|
ratio has no budget to be checked against, and decoding one to find out
|
|
what it costs IS the failure this bound exists to stop. Refusing an unknown
|
|
name rather than passing it through is the same decision
|
|
`corpus.resolve_gate` takes for an unknown gate name: a fallback reproduces
|
|
the defect with an extra step.
|
|
|
|
The reach is measured rather than assumed. Over the 9 668 image objects of
|
|
the 77 PDFs on this machine (2026-09-18, enumerated through pdfminer's own
|
|
page walk), the filter chains are 6 235 `[/FlateDecode]`, 2 459
|
|
`[/DCTDecode]`, 596 `[/FlateDecode /DCTDecode]`, 296 `[/Fl]`, 42
|
|
unfiltered, 16 `[/ASCII85Decode /FlateDecode]`, 16 `[/JPXDecode]` and 8
|
|
`[/CCITTFaxDecode]`. The refused class is those 8 objects, 1-bit stencil
|
|
masks (`/ImageMask true`, `/BitsPerComponent 1`) already refused one step
|
|
later by the encoder. The largest input any `ASCII85Decode` link is handed
|
|
is 450 739 bytes, more than ten times under the budget's cap, which is why
|
|
the cap costs no picture the corpora hold.
|
|
"""
|
|
return _BOUNDED_PDF_FILTERS
|
|
|
|
|
|
#: The names in `bounded_pdf_filters`, as a constant the test suite pins. The
|
|
#: docstring above is the published claim; this is what the code enforces, and
|
|
#: `_pdf_filter_names` maps every pdfminer spelling of these onto the canonical
|
|
#: name that `assets.PDF_FILTER_COST_RATIO` budgets.
|
|
_BOUNDED_PDF_FILTERS = frozenset(
|
|
{
|
|
"FlateDecode",
|
|
"ASCII85Decode",
|
|
"ASCIIHexDecode",
|
|
"DCTDecode",
|
|
"JPXDecode",
|
|
"JBIG2Decode",
|
|
}
|
|
)
|
|
|
|
|
|
def _pdf_filter_names() -> dict[object, str]:
|
|
"""Every pdfminer literal this package bounds, mapped to its CANONICAL name.
|
|
|
|
Read from pdfminer rather than written out here, because a filter has more
|
|
than one spelling (`/Fl` is `/FlateDecode`) and a set of names written by
|
|
hand would refuse the abbreviation a real document uses. The canonical name
|
|
is the key into `assets.PDF_FILTER_COST_RATIO`, so both spellings of a
|
|
filter are budgeted by one measured number.
|
|
"""
|
|
from pdfminer.pdftypes import (
|
|
LITERALS_ASCII85_DECODE,
|
|
LITERALS_ASCIIHEX_DECODE,
|
|
LITERALS_DCT_DECODE,
|
|
LITERALS_FLATE_DECODE,
|
|
LITERALS_JBIG2_DECODE,
|
|
LITERALS_JPX_DECODE,
|
|
)
|
|
|
|
families = {
|
|
"FlateDecode": LITERALS_FLATE_DECODE,
|
|
"ASCII85Decode": LITERALS_ASCII85_DECODE,
|
|
"ASCIIHexDecode": LITERALS_ASCIIHEX_DECODE,
|
|
"DCTDecode": LITERALS_DCT_DECODE,
|
|
"JPXDecode": LITERALS_JPX_DECODE,
|
|
"JBIG2Decode": LITERALS_JBIG2_DECODE,
|
|
}
|
|
return {
|
|
literal: canonical
|
|
for canonical, literals in families.items()
|
|
for literal in literals
|
|
if canonical in _BOUNDED_PDF_FILTERS
|
|
}
|
|
|
|
|
|
#: The two canonical names whose decoders produce fewer or more bytes than they
|
|
#: were given, and which this package therefore has to run to learn the size of
|
|
#: the link behind them. Everything else in the table is pass-through.
|
|
_SHRINKING_FILTER_NAMES = frozenset({"ASCII85Decode", "ASCIIHexDecode"})
|
|
|
|
|
|
def _pdf_stream_bytes(stream: object, name: str) -> bytes | None:
|
|
"""The stream's raw bytes, DECIPHERED where the document is encrypted.
|
|
|
|
`None` means there are none left to measure: pdfminer's `decode()` sets
|
|
`rawdata` to `None`, so a stream something else has already decoded was
|
|
paid for before this package was asked anything. That is the one path the
|
|
backstop after `get_data()` exists for.
|
|
|
|
Until this commit an encrypted stream RETURNED here without being measured,
|
|
which made "the document declares encryption" a way past the bound.
|
|
Deciphering does not change a stream's length, so this does what pdfminer's
|
|
own `decode()` does -- decipher first, then read the filters -- and the
|
|
bound applies to an encrypted document exactly as it does to any other.
|
|
"""
|
|
raw = stream.get_rawdata() if hasattr(stream, "get_rawdata") else None
|
|
if raw is None:
|
|
return None
|
|
decipher = getattr(stream, "decipher", None)
|
|
if decipher is None:
|
|
return bytes(raw)
|
|
objid = getattr(stream, "objid", None)
|
|
genno = getattr(stream, "genno", None)
|
|
if objid is None or genno is None:
|
|
raise ExtractionError(
|
|
f"the encrypted stream behind {name!r} carries no object number, so its "
|
|
"bytes cannot be deciphered before they are decoded; refused rather "
|
|
"than decoded to find out what it costs",
|
|
code="asset_pdf_unbounded",
|
|
)
|
|
try:
|
|
return bytes(decipher(objid, genno, raw, getattr(stream, "attrs", {})))
|
|
except Exception as exc:
|
|
raise ExtractionError(
|
|
f"the encrypted stream behind {name!r} could not be deciphered: {exc}",
|
|
code="asset_pdf_unbounded",
|
|
) from exc
|
|
|
|
|
|
def _check_stream_cost(stream: object, name: str) -> None:
|
|
"""Refuse a PDF image stream whose DECODED size is over the bound.
|
|
|
|
`check_size` reads `/Width` and `/Height`, which an untrusted document
|
|
writes and which say nothing about what `get_data()` will return: `/Length`
|
|
is the COMPRESSED length, and a dictionary declaring 1x1 may hang 400 MB of
|
|
deflated zeros off it. Measured by an independent review of 0.10.1 on
|
|
`230d1cb`: that document is 408 516 bytes, was carried without a rejection,
|
|
and cost 892 MB of peak RSS -- the same failure mode the 0.10.0 review
|
|
named, reached through the other number.
|
|
|
|
THE CHAIN, NOT ITS FIRST LINK. 0.10.1 measured `filters[0]` and returned,
|
|
which is not a bound: a PDF decodes a stream through a LIST of filters, and
|
|
`/Filter [/FlateDecode /FlateDecode]` puts the whole expansion in the
|
|
second one. Measured 2026-09-18 on `0f308c1`: 1 636 bytes of file became
|
|
886 554 624 bytes of peak RSS, and three links did the same from 1 070
|
|
bytes -- about 542 000x the file, and the picture WAS refused at the end,
|
|
by `check_payload`, after the memory had been spent. It also left the 16
|
|
corpus objects behind an `[/ASCII85Decode /FlateDecode]` chain unmeasured,
|
|
because `filters[0]` is not `FlateDecode` there.
|
|
|
|
AND THE COST OF A LINK, NOT THE SIZE OF ITS OUTPUT. Bounding every
|
|
`FlateDecode` was still not a bound, because the bomb moved into a link
|
|
0.10.1 had documented as safe: `ASCII85Decode`'s `z` is the shorthand for
|
|
four zero bytes, and its decoder holds about a hundred bytes per byte of
|
|
input. Measured on the pinned tree: 33 475 bytes of file cost
|
|
3 261 599 744 bytes of peak RSS and the picture was CARRIED. Three rounds
|
|
of this review each bound an OUTPUT and the bomb stepped one link along;
|
|
what they had in common is that a decoder's working set is not its output.
|
|
|
|
So every link is walked, in order, and each is given a BUDGET
|
|
(`assets.MAX_FILTER_DECODE_BYTES`) rather than a class:
|
|
|
|
* a filter with no measured cost ratio is refused BEFORE anything is
|
|
decoded (`bounded_pdf_filters` says which, and why);
|
|
* a `FlateDecode` is measured as it is paid, under a limit that is the
|
|
smaller of the picture's own bound and what the NEXT link's decoder may
|
|
be handed -- which is how the budget travels down the chain instead of
|
|
being applied to each link in isolation;
|
|
* the last `FlateDecode` in the chain has its output counted and thrown
|
|
away, which is the common case and costs exactly what 0.10.1 cost; an
|
|
earlier one is inflated under the same limit and handed on, so the link
|
|
behind it can be measured from real bytes;
|
|
* every other filter has its cost PREDICTED from its input size and its
|
|
measured ratio, and is refused before its decoder is called.
|
|
|
|
WHAT THIS STILL DOES NOT BOUND, stated rather than implied: a stream
|
|
something else has already decoded (`_pdf_stream_bytes` returns `None`),
|
|
where the memory is spent before this package is asked. That one is caught
|
|
by `check_payload` AFTER `get_data()`, which makes it a counted refusal
|
|
rather than a bounded one.
|
|
"""
|
|
names = _pdf_filter_names()
|
|
|
|
data = _pdf_stream_bytes(stream, name)
|
|
if data is None:
|
|
return
|
|
raw_filters = stream.get_filters() if hasattr(stream, "get_filters") else []
|
|
# THE WHOLE CHAIN IS READ BEFORE THE FIRST LINK IS DECODED. A filter this
|
|
# package cannot bound must be refused without having paid for the links in
|
|
# front of it, which is only possible if the refusal is decided up front.
|
|
chain: list[tuple[str, object]] = []
|
|
for literal, params in raw_filters:
|
|
canonical = names.get(literal)
|
|
if canonical is None:
|
|
raise ExtractionError(
|
|
f"the image {name!r} is decoded through {literal}, a filter whose "
|
|
"cost this package has no measured ratio for; refused unread "
|
|
"rather than decoded to find out what it costs",
|
|
code="asset_pdf_unbounded",
|
|
)
|
|
chain.append((canonical, params))
|
|
|
|
# The LAST deflate link is the one whose bytes nothing behind has to be
|
|
# measured from, so it is counted and thrown away; every earlier one is
|
|
# inflated under the same bound and handed on. Deciding this by index
|
|
# rather than by a running flag is what keeps "the bytes are gone" and "a
|
|
# link still needs them" from ever being true at once.
|
|
last_flate = max(
|
|
(index for index, (canonical, _) in enumerate(chain) if canonical == "FlateDecode"),
|
|
default=-1,
|
|
)
|
|
size = len(data)
|
|
for index, (canonical, params) in enumerate(chain):
|
|
behind = [name_behind for name_behind, _ in chain[index + 1 :]]
|
|
if canonical == "FlateDecode":
|
|
limit = inflate_limit_for(behind[0] if behind else None)
|
|
if index == last_flate:
|
|
# Nothing behind has to be measured, so the output is counted
|
|
# and thrown away: the cheap common case, and what 0.10.1 cost.
|
|
size = inflated_size(data, name=name, limit=limit)
|
|
data = b""
|
|
else:
|
|
if _has_predictor(params):
|
|
raise ExtractionError(
|
|
f"the image {name!r} applies a predictor to a link that is not "
|
|
"the last one, so the bytes this package would hand to the next "
|
|
"filter are not the bytes pdfminer decodes; refused unread",
|
|
code="asset_pdf_unbounded",
|
|
)
|
|
data = inflate_bounded(data, name=name, limit=limit)
|
|
size = len(data)
|
|
else:
|
|
# PREDICTED, not measured, and predicted BEFORE the decoder is
|
|
# called: these decoders take a whole string and return a whole
|
|
# string, so there is no moment between the two at which the cost
|
|
# could be observed.
|
|
check_filter_cost(size, canonical=canonical, name=name)
|
|
if canonical in _SHRINKING_FILTER_NAMES:
|
|
if index > last_flate >= 0:
|
|
# The bytes were discarded at the last deflate link, so the
|
|
# bound travels on as the WIDEST this link could produce.
|
|
size = _widest_output(canonical, size)
|
|
else:
|
|
try:
|
|
data = _shrink(canonical, data)
|
|
except Exception:
|
|
# Not this function's problem: a stream that is not
|
|
# valid input for its own filter is reported by the
|
|
# reader behind it, in that reader's vocabulary.
|
|
return
|
|
size = len(data)
|
|
# A pass-through filter leaves the bytes exactly as they are.
|
|
check_payload(size, name=name)
|
|
|
|
|
|
def _widest_output(canonical: str, size: int) -> int:
|
|
"""The most `canonical` can produce from `size` bytes, for a link whose
|
|
bytes were discarded and whose SIZE is all that is carried forward.
|
|
|
|
Rounded up rather than down, and never below one byte: a bound that is
|
|
optimistic by a byte is not a bound.
|
|
"""
|
|
ratio = PDF_FILTER_OUTPUT_RATIO.get(canonical)
|
|
if ratio is None: # pragma: no cover - only `FlateDecode`, handled above
|
|
return size
|
|
return max(1, math.ceil(size * ratio))
|
|
|
|
|
|
def _has_predictor(params: object) -> bool:
|
|
"""Whether a `DecodeParms` entry asks for a predictor other than `1`."""
|
|
if not isinstance(params, dict) or "Predictor" not in params:
|
|
return False
|
|
from pdfminer.pdftypes import resolve1
|
|
|
|
predictor = resolve1(params["Predictor"])
|
|
return isinstance(predictor, int) and predictor > 1
|
|
|
|
|
|
def _shrink(canonical: str, data: bytes) -> bytes:
|
|
"""The two ASCII filters, decoded with pdfminer's own readers so both sides
|
|
agree on what the bytes are.
|
|
|
|
Called only after `check_filter_cost` has allowed the input size, which is
|
|
what makes handing a whole string to a decoder that returns a whole string
|
|
a bounded thing to do. The name "shrink" is kept for the pair, but only
|
|
`ASCIIHexDecode` actually shrinks: `ASCII85Decode` can quadruple its input,
|
|
which is the defect this round was opened for.
|
|
"""
|
|
from pdfminer.ascii85 import ascii85decode, asciihexdecode
|
|
|
|
if canonical == "ASCII85Decode":
|
|
return ascii85decode(data)
|
|
return asciihexdecode(data)
|
|
|
|
|
|
def _pdf_image(stream: object, name: str) -> ExtractedImage:
|
|
"""One image XObject, carried verbatim where it already is a file.
|
|
|
|
TWO ROUTES, and which one runs is decided by the BYTES rather than by the
|
|
filter name. `get_data()` applies every filter pdfminer knows and stops at
|
|
the image codecs, so a `DCTDecode` stream comes back as a finished JPEG and
|
|
a `FlateDecode` one comes back as raw samples. Sniffing the result is what
|
|
makes the first route exact: an embedded JPEG is written to the bundle as
|
|
the publisher's own bytes, unre-encoded, and its content-addressed name is
|
|
therefore stable for as long as the document is.
|
|
|
|
Measured on R761 (2026-09-16): 29 of 50 image objects are `DCTDecode` and
|
|
take the verbatim route; 21 are `FlateDecode` and are encoded here. Over
|
|
the 33-document K2 reference corpus the population is 4 828 objects, and
|
|
the filters are mixed enough (`FlateDecode`, `DCTDecode`, `JPXDecode`,
|
|
`ASCII85Decode` chains, `CCITTFaxDecode`) that guessing from the filter
|
|
name would have been wrong on several hundred.
|
|
|
|
RENDERING THE PAGE REGION WAS THE ALTERNATIVE AND IT WAS NOT TAKEN. A
|
|
rasterised crop would be one code path and would handle every filter, but
|
|
its bytes -- and therefore the asset's name and the bundle's digest --
|
|
would depend on the version of the rasteriser installed, which is the one
|
|
property `OCR_DPI`'s docstring already admits OCR text cannot have. An
|
|
embedded stream has no such dependency.
|
|
"""
|
|
from pdfminer.pdftypes import resolve1
|
|
|
|
# THE DECLARED SIZE IS READ FIRST, and the stream is not touched until it
|
|
# is within the bound: `get_data()` decompresses, so a check after it has
|
|
# already paid for a picture of compressed zeros. `channels=1` because the
|
|
# colour space is not resolved until further down and one channel is the
|
|
# LOWEST estimate of the cost -- a bound that under-counts refuses nothing
|
|
# it should carry, and the pixel count beside it does not depend on it.
|
|
# What that leaves unbounded is the stream, which `_check_stream_cost` reads
|
|
# for what it actually costs rather than for what it claims.
|
|
declared = dict(getattr(stream, "attrs", {}))
|
|
declared_width = resolve1(declared.get("Width"))
|
|
declared_height = resolve1(declared.get("Height"))
|
|
check_size(
|
|
declared_width if isinstance(declared_width, int) else None,
|
|
declared_height if isinstance(declared_height, int) else None,
|
|
name=name,
|
|
channels=1,
|
|
)
|
|
_check_stream_cost(stream, name)
|
|
try:
|
|
data = stream.get_data() # type: ignore[attr-defined]
|
|
except Exception as exc:
|
|
raise ExtractionError(
|
|
f"the PDF image stream behind {name!r} could not be decoded: {exc}",
|
|
code="asset_pdf_unsupported",
|
|
) from exc
|
|
# THE BACKSTOP, and it is a weaker guarantee than the one above it: this
|
|
# one COUNTS a stream that `_check_stream_cost` could not bound in advance
|
|
# (see its docstring for which those are), after the memory has been spent.
|
|
check_payload(len(data), name=name)
|
|
if data and sniff(data) is not None:
|
|
return read_image(data, name=name)
|
|
|
|
attrs = dict(getattr(stream, "attrs", {}))
|
|
width = resolve1(attrs.get("Width"))
|
|
height = resolve1(attrs.get("Height"))
|
|
bits = resolve1(attrs.get("BitsPerComponent"))
|
|
if not isinstance(width, int) or not isinstance(height, int):
|
|
raise ExtractionError(
|
|
f"the PDF image {name!r} declares no usable size",
|
|
code="asset_pdf_unsupported",
|
|
)
|
|
if resolve1(attrs.get("ImageMask")):
|
|
raise ExtractionError(
|
|
f"the PDF image {name!r} is a stencil mask, which paints the current "
|
|
"fill colour rather than carrying one of its own",
|
|
code="asset_pdf_unsupported",
|
|
)
|
|
if bits != _PDF_SAMPLE_BITS:
|
|
raise ExtractionError(
|
|
f"the PDF image {name!r} stores {bits}-bit samples; this encoder writes "
|
|
f"{_PDF_SAMPLE_BITS}-bit ones and will not rescale, because rescaling a "
|
|
"stencil is a decision about what black means",
|
|
code="asset_pdf_unsupported",
|
|
)
|
|
if attrs.get("Decode") is not None:
|
|
raise ExtractionError(
|
|
f"the PDF image {name!r} carries a Decode array, which remaps every "
|
|
"sample; carrying it unmapped would invert the picture",
|
|
code="asset_pdf_unsupported",
|
|
)
|
|
colour = _pdf_colour(attrs.get("ColorSpace"))
|
|
if colour is None:
|
|
raise ExtractionError(
|
|
f"the PDF image {name!r} uses a colour space this encoder does not "
|
|
f"express ({attrs.get('ColorSpace')!r})",
|
|
code="asset_pdf_unsupported",
|
|
)
|
|
alpha = _pdf_alpha(attrs, width, height)
|
|
if alpha is False:
|
|
raise ExtractionError(
|
|
f"the PDF image {name!r} has a soft mask this encoder cannot express; "
|
|
"dropping transparency would composite the picture against nothing",
|
|
code="asset_pdf_unsupported",
|
|
)
|
|
channels, palette = colour
|
|
encoded = encode_png(
|
|
width,
|
|
height,
|
|
data,
|
|
channels=channels,
|
|
palette=palette,
|
|
alpha=alpha if isinstance(alpha, bytes) else None,
|
|
)
|
|
return read_image(encoded, name=name)
|
|
|
|
|
|
def _pdf_images(page: object) -> tuple[tuple[ExtractedImage, ...], tuple[AssetRejection, ...]]:
|
|
"""Every image drawn on one page, with the failures kept beside them."""
|
|
carried: list[ExtractedImage] = []
|
|
rejected: list[AssetRejection] = []
|
|
number = getattr(page, "page_number", 0)
|
|
for index, drawn in enumerate(getattr(page, "images", []) or [], start=1):
|
|
# The name a PDF image does NOT have. An XObject is reached through a
|
|
# resource name local to one page's dictionary, so it is not an
|
|
# identifier -- the page number in front of it is what makes the string
|
|
# readable, and the content-addressed digest is what makes it unique.
|
|
# A name pdfminer derived from `id()` is not a name. An inline image
|
|
# (`BI ... EI`) has no resource name, and pdfminer fills the field with
|
|
# the address of a Python object, so 0.10.0 wrote a pointer line that
|
|
# changed between two runs of one build -- two concept files of the
|
|
# reference corpus differed, which breaks the bit-exact rebuild.
|
|
# Measured 2026-09-17.
|
|
raw = str(drawn.get("name") or "").lstrip("/")
|
|
label = raw if raw and not raw.isdigit() else f"inline-{index}"
|
|
name = f"page-{number}-{label}"
|
|
stream = drawn.get("stream")
|
|
if stream is None:
|
|
rejected.append(
|
|
AssetRejection(name, "asset_pdf_unsupported", "the image object has no stream")
|
|
)
|
|
continue
|
|
try:
|
|
carried.append(_pdf_image(stream, name))
|
|
except ExtractionError as exc:
|
|
rejected.append(AssetRejection(name, exc.code, str(exc)))
|
|
return tuple(carried), tuple(rejected)
|
|
|
|
|
|
@functools.lru_cache(maxsize=1)
|
|
def _pdf_pages(
|
|
data: bytes, headings: bool = False, ocr: bool = False, assets: bool = False
|
|
) -> tuple[_PdfPage, ...]:
|
|
"""Every page that produced content, as a `_PdfPage`, in page order.
|
|
|
|
The page NUMBER is 1-based and comes from the document, so a page that
|
|
yielded nothing removes itself from the sequence without renumbering the
|
|
ones after it -- which is the difference between "the third page that
|
|
produced text" and "page 3", and the whole reason a locator is worth
|
|
writing down.
|
|
|
|
Memoised on the bytes AND on the two options, with room for exactly one
|
|
entry: extraction and location are two calls about the same file with the
|
|
same options, back to back, and parsing it twice would double the PDF cost
|
|
of every corpus run for nothing. The options are part of the key because
|
|
two renderings of one document are two different strings, and a locator
|
|
built against the wrong one points at the wrong place with full confidence.
|
|
|
|
`headings` and `ocr` are INDEPENDENT and compose. With both off this is the
|
|
path every byte-pinned golden was measured on, unchanged: the default
|
|
branch still calls `page.extract_text()` rather than reassembling the page
|
|
from its lines. Measured, the two agree on 11 of 11 pages of a real tender
|
|
PDF -- but "agree on the document I tried" is not a contract, so the
|
|
default does not depend on it.
|
|
|
|
`assets` is the third, and with it off not one line below it runs: no
|
|
stream is decoded, no sample buffer is allocated, and the emitted pages are
|
|
the objects they always were. A page that produced no TEXT is still
|
|
dropped even when it carries an image, because `_extract_pdf` refuses a
|
|
document with no text at all (`extractor_empty_pdf`) and an image-only
|
|
document is `--ocr`'s question, not this one's.
|
|
"""
|
|
try:
|
|
import pdfplumber
|
|
except ImportError as exc:
|
|
raise _extra_missing(".pdf") from exc
|
|
|
|
read = _ocr_reader() if ocr else None
|
|
try:
|
|
with pdfplumber.open(io.BytesIO(data)) as pdf:
|
|
# PASS ONE. Nothing is emitted here, because the heading rule needs
|
|
# a fact about the WHOLE document -- the body's size -- and a page
|
|
# cannot supply it. A title page is 100 % heading, and a per-page
|
|
# median would compare it with itself and mark nothing.
|
|
recovered: list[str | list[tuple[str, float, str]]] = []
|
|
numbers: list[int] = []
|
|
drawn: list[tuple[tuple[ExtractedImage, ...], tuple[AssetRejection, ...]]] = []
|
|
for page in pdf.pages:
|
|
flat = (page.extract_text() or "").rstrip()
|
|
numbers.append(page.page_number)
|
|
drawn.append(_pdf_images(page) if assets else ((), ()))
|
|
if read is not None and needs_ocr(flat):
|
|
# The page's own text is unusable, so it is replaced
|
|
# WHOLESALE rather than merged with: a page of `(cid:N)`
|
|
# has nothing worth keeping, and interleaving two readings
|
|
# of one page would put a guess and a fact in one paragraph
|
|
# with no way to tell them apart. An OCR'd page carries no
|
|
# typography either -- the engine reports text, not fonts --
|
|
# so it is a finished string and never a heading candidate.
|
|
recovered.append("\n".join(read(page.to_image(resolution=OCR_DPI).original)))
|
|
elif not headings:
|
|
recovered.append(flat)
|
|
else:
|
|
recovered.append(
|
|
[
|
|
(str(line["text"]), *found)
|
|
for line in page.extract_text_lines()
|
|
if (found := _typography(line)) is not None
|
|
]
|
|
)
|
|
levels = _heading_levels(
|
|
[
|
|
line
|
|
for page_lines in recovered
|
|
if not isinstance(page_lines, str)
|
|
for line in page_lines
|
|
]
|
|
)
|
|
# PASS TWO.
|
|
pages = [
|
|
page_lines
|
|
if isinstance(page_lines, str)
|
|
else _mark_headings(page_lines, levels).rstrip()
|
|
for page_lines in recovered
|
|
]
|
|
except ExtractionError:
|
|
raise
|
|
except Exception as exc:
|
|
raise ExtractionError(
|
|
f"the PDF parser failed on this file: {exc}", code="extractor_pdf_error"
|
|
) from exc
|
|
return tuple(
|
|
_PdfPage(number=number, text=text, images=images, rejected=rejected)
|
|
for number, text, (images, rejected) in zip(numbers, pages, drawn)
|
|
if text
|
|
)
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class OutlineMark:
|
|
"""One `/Outlines` node, placed on a LINE of the string `extract_text` returns.
|
|
|
|
`level` is what the TREE declares, not a distance normalised against
|
|
anything: a document whose outline carries its own root node puts its
|
|
chapters at level 2, and rewriting that here would state a structure the
|
|
publisher did not. Measured on a 701-page process code -- the tree's levels
|
|
2..8 hold 28/118/500/1141/872/93/9 nodes against the publisher's own
|
|
NISO-STS depths 1..7 at 28/118/500/1141/868/97/9, so the mapping is level
|
|
minus one on five rows and the publisher disagrees with the publisher on
|
|
four nodes. That disagreement is data, and it survives only if the level is
|
|
reported rather than fixed up.
|
|
"""
|
|
|
|
line: int
|
|
level: int
|
|
title: str
|
|
#: The section's own first spec point, where the SOURCE declares one. Set
|
|
#: only by the NISO-STS reader; `None` on every bookmark mark, because a
|
|
#: bookmark declares a place and never a summary.
|
|
description: str | None = None
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class PdfOutline:
|
|
"""The bookmark tree, bridged onto lines -- with what did not bridge counted.
|
|
|
|
`unresolved` is not decoration. A `/Dest` that names an object which is not
|
|
a page, or a page that produced no text, has to be DROPPED: fabricating a
|
|
boundary from it would put a heading somewhere the document never had one,
|
|
and raising would refuse a file over a defect in one of its bookmarks.
|
|
Dropping silently is the third option this library refuses everywhere else,
|
|
so the count is part of the return value.
|
|
|
|
`collided` is the same principle applied to the OTHER way a node leaves
|
|
without a boundary. Two bookmarks can resolve to one line -- measured on
|
|
the 701-page process code, its tree's root node and `SVV - Forside` both
|
|
land on line 0 -- and only the first can become a mark, because two
|
|
candidates at one offset give the first an empty span that the orphan check
|
|
then deletes without a word. That was measured on that document and is why
|
|
keeping both was felled rather than argued. What the count buys is the
|
|
identity: NODES IN == len(marks) + unresolved + collided, so a document
|
|
that loses several nodes this way says so instead of returning a shorter
|
|
list that looks complete.
|
|
"""
|
|
|
|
marks: tuple[OutlineMark, ...]
|
|
unresolved: int
|
|
collided: int = 0
|
|
|
|
|
|
def _outline_page_and_top(doc: object, dest: object, action: object) -> tuple[object, float | None]:
|
|
"""`(page reference, /XYZ top)` from a bookmark's destination, or `(None, None)`.
|
|
|
|
Four shapes reach here and all four are in the wild: an explicit array, a
|
|
NAMED destination resolved through the document's name tree, a `GoTo`
|
|
action carrying either, and an indirect reference to any of them.
|
|
"""
|
|
target = dest
|
|
if target is None and action is not None:
|
|
resolved = action.resolve() if hasattr(action, "resolve") else action
|
|
if isinstance(resolved, dict):
|
|
target = resolved.get("D")
|
|
if isinstance(target, (bytes, str)) or hasattr(target, "name"):
|
|
name = target.name if hasattr(target, "name") else target
|
|
try:
|
|
target = doc.get_dest(name) # type: ignore[attr-defined]
|
|
except Exception:
|
|
return (None, None)
|
|
if hasattr(target, "resolve"):
|
|
try:
|
|
target = target.resolve()
|
|
except Exception:
|
|
return (None, None)
|
|
if isinstance(target, dict):
|
|
target = target.get("D")
|
|
if not isinstance(target, list) or not target:
|
|
return (None, None)
|
|
top: float | None = None
|
|
if len(target) > 3 and getattr(target[1], "name", None) == "XYZ":
|
|
candidate = target[3]
|
|
if isinstance(candidate, (int, float)):
|
|
top = float(candidate)
|
|
return (target[0], top)
|
|
|
|
|
|
def pdf_outline(
|
|
name: str,
|
|
data: bytes,
|
|
*,
|
|
pdf_headings: bool = False,
|
|
ocr: bool = False,
|
|
assets: bool = False,
|
|
) -> PdfOutline:
|
|
"""`pdf`: the file's own `/Outlines` tree, as marks on the extracted text.
|
|
|
|
THE BRIDGE IS THE WHOLE PROBLEM, and both routes are measured rather than
|
|
argued. A bookmark states a PAGE and a y position; a candidate needs a LINE
|
|
index. On the 701-page document this was built against, 2 706 of 2 761
|
|
bookmarks share a destination page with another bookmark, so the page alone
|
|
is never a cut point.
|
|
|
|
Y ROUTE (primary). `page.extract_text_lines()` carries a `top` per line,
|
|
and the mark takes the FIRST line at or below the destination. It needs
|
|
the line splitting to be the one `page.extract_text()` produced -- an
|
|
assumption, so it is CHECKED per page and the route is used only where
|
|
the two strings are identical. Measured: 701 of 701 pages, and the
|
|
resulting index agrees with the title route on 2 762 of 2 762 nodes,
|
|
flat from a 0 pt tolerance to 8 pt and collapsing at 12 (the line
|
|
spacing). It therefore ships with NO tolerance constant at all.
|
|
|
|
TITLE ROUTE (fallback). The bookmark's title, normalised, searched in the
|
|
destination page's own lines. It resolved 2 762 of 2 763 on that
|
|
document, and its weakness is real: a title like `Armering` occurs nine
|
|
times in that structure, so it is scoped to the destination page and is
|
|
never asked a question the y route already answered.
|
|
|
|
`pdf_headings` and `ocr` are passed through so the line indices address the
|
|
SAME rendering the caller extracted. They are not options of this arm: a
|
|
plan indexes one exact string, and marks computed against another one point
|
|
at the right words in the wrong places.
|
|
"""
|
|
if Path(name).suffix.lower() != ".pdf":
|
|
return PdfOutline((), 0)
|
|
try:
|
|
import pdfplumber
|
|
except ImportError as exc:
|
|
raise _extra_missing(".pdf") from exc
|
|
from pdfminer.pdfdocument import PDFNoOutlines
|
|
from pdfminer.pdfpage import PDFPage
|
|
|
|
rendered = _pdf_pages(data, pdf_headings, ocr, assets)
|
|
starts: dict[int, int] = {}
|
|
page_lines: dict[int, list[str]] = {}
|
|
# The page's OWN lines, without the pointer blocks appended after them.
|
|
# `starts` has to count the appended lines (they are in the text a mark
|
|
# indexes) while the geometry check must not see them, because
|
|
# `extract_text_lines` reports the page and knows nothing about what this
|
|
# package added underneath it.
|
|
body_lines: dict[int, list[str]] = {}
|
|
offset = 0
|
|
for rendered_page in rendered:
|
|
starts[rendered_page.number] = offset
|
|
page_lines[rendered_page.number] = _pdf_page_text(rendered_page).split("\n")
|
|
body_lines[rendered_page.number] = rendered_page.text.split("\n")
|
|
offset += len(page_lines[rendered_page.number]) + 1
|
|
|
|
unresolved = 0
|
|
collided = 0
|
|
placed: dict[int, OutlineMark] = {}
|
|
with pdfplumber.open(io.BytesIO(data)) as pdf:
|
|
try:
|
|
nodes = list(pdf.doc.get_outlines())
|
|
except PDFNoOutlines:
|
|
# NOT an error, and not zero concepts either: this file simply
|
|
# carries no index, which is the common case and the one the
|
|
# byte-identical guarantee below rests on.
|
|
return PdfOutline((), 0)
|
|
except Exception as exc:
|
|
raise ExtractionError(
|
|
f"the PDF parser failed reading /Outlines: {exc}",
|
|
code="extractor_pdf_error",
|
|
) from exc
|
|
numbers = {
|
|
page.pageid: index + 1 for index, page in enumerate(PDFPage.create_pages(pdf.doc))
|
|
}
|
|
wanted: dict[int, list[tuple[int, str, float | None]]] = {}
|
|
for level, title, dest, action, _ in nodes:
|
|
reference, top = _outline_page_and_top(pdf.doc, dest, action)
|
|
page_number = numbers.get(getattr(reference, "objid", None))
|
|
if page_number is None or page_number not in starts:
|
|
unresolved += 1
|
|
continue
|
|
wanted.setdefault(page_number, []).append((int(level), str(title), top))
|
|
# Geometry is read only for the pages that carry a bookmark, because
|
|
# `extract_text_lines` costs a second render of every page it is asked
|
|
# about -- 78 s over 701 pages, and nothing at all over the pages no
|
|
# bookmark points at.
|
|
for page in pdf.pages:
|
|
number = page.page_number
|
|
group = wanted.get(number)
|
|
if not group:
|
|
continue
|
|
lines = page_lines[number]
|
|
tops: list[float] | None = None
|
|
geometry = page.extract_text_lines()
|
|
if [str(entry["text"]) for entry in geometry] == body_lines[number]:
|
|
tops = [float(entry["top"]) for entry in geometry]
|
|
height = float(page.height)
|
|
for level, title, top in group:
|
|
index: int | None = None
|
|
if tops is not None and top is not None:
|
|
want = height - top
|
|
index = next(
|
|
(position for position, value in enumerate(tops) if value >= want),
|
|
len(tops) - 1,
|
|
)
|
|
else:
|
|
target = _normalise_outline(title)
|
|
joined = ""
|
|
bounds: list[int] = []
|
|
for rendered_line in lines:
|
|
bounds.append(len(joined))
|
|
joined += _normalise_outline(rendered_line)
|
|
found = joined.find(target)
|
|
if found >= 0:
|
|
index = max(position for position, at in enumerate(bounds) if at <= found)
|
|
if index is None:
|
|
unresolved += 1
|
|
continue
|
|
at = starts[number] + index
|
|
# FIRST in tree order wins a shared line. Two marks on one line
|
|
# would give the second an empty span, and the orphan check
|
|
# deletes an empty span silently -- the same trap the proposer
|
|
# documents for a second-pass candidate list. The loser is
|
|
# COUNTED rather than dropped: `setdefault` alone made a lost
|
|
# node indistinguishable from a node that was never there.
|
|
if at in placed:
|
|
collided += 1
|
|
continue
|
|
placed[at] = OutlineMark(line=at, level=level, title=title)
|
|
return PdfOutline(tuple(placed[at] for at in sorted(placed)), unresolved, collided)
|
|
|
|
|
|
def _normalise_outline(value: str) -> str:
|
|
"""Whitespace out, case folded -- the form the title route compares on."""
|
|
return re.sub(r"\s+", "", value).lower()
|
|
|
|
|
|
def _extract_pdf(
|
|
data: bytes, *, headings: bool = False, ocr: bool = False, assets: bool = False
|
|
) -> str:
|
|
"""`pdf`: page text via `pdfplumber`, in page order, pages separated by a
|
|
blank line.
|
|
|
|
The gate is this import, not a membership test: without the `[extract]`
|
|
extra the very same typed rejection is raised as for the types that ship
|
|
no parser at all. Text is returned VERBATIM — no Unicode normalization,
|
|
matching `md`/`txt` passthrough; normalizing would edit source content,
|
|
and NFC folding belongs to filenames and titles, not to document bodies.
|
|
|
|
`pdfplumber` was chosen on ONE measured property (2026-08-21,
|
|
docs/2026-08-21-g2-pdf-extraction-measurement.md): on a real requirement
|
|
table it keeps label and value on the same line, where pypdf, pdfminer.six
|
|
and pymupdf each emit all labels then all values. Re-pairing those is
|
|
guesswork, and in a requirements document a wrong pairing looks right.
|
|
"""
|
|
pages = _pdf_pages(data, headings, ocr, assets)
|
|
text = _PDF_PAGE_SEPARATOR.join(_pdf_page_text(page) for page in pages)
|
|
if not text:
|
|
raise ExtractionError(
|
|
"the PDF yielded no text on any page; a scanned or image-only "
|
|
"document needs OCR, which this registry does only behind the "
|
|
"optional 'ocr' group and only when asked",
|
|
code="extractor_empty_pdf",
|
|
)
|
|
# After the parse, not before: a run that produced no text has nothing to
|
|
# be lossy about, and warning there would just add noise to a failure.
|
|
warnings.warn(_PDF_LOSSY_WARNING, ExtractionWarning, stacklevel=3)
|
|
return text
|
|
|
|
|
|
def _convert_bytes(source: bytes, to: str, format: str, extra_args: Sequence[str]) -> str:
|
|
"""The one converter call, isolated so the seam above it is testable.
|
|
|
|
Separated for a reason beyond tidiness: every test of the seam's behaviour
|
|
would otherwise need the binary present and a real office document, which
|
|
would make the seam's own logic untestable on a machine without the extra.
|
|
The conversion itself is covered by the frozen-text fixtures instead.
|
|
|
|
THE INPUT GOES THROUGH A FILE, NOT THROUGH THE TEXT ENTRY POINT. Every
|
|
format here is a binary container, and the converter's text entry point
|
|
takes an `encoding` because it treats its source as text -- which corrupts
|
|
a zip. Measured: a hand-laid `.xlsx` that pandoc reads correctly from disk
|
|
fails through the text path with `Failed to unpack XLSX archive: not enough
|
|
bytes`. A `.docx` of the same shape happened to survive, which is what
|
|
makes this worth writing down: the defect is SILENT for some inputs and
|
|
fatal for others, so "it worked on the file I tried" is not evidence here.
|
|
|
|
The temporary directory is removed on every path, including the failure
|
|
one, and nothing outside it is written.
|
|
"""
|
|
import pypandoc
|
|
|
|
from ._pandoc import converter_path
|
|
|
|
with tempfile.TemporaryDirectory() as staging:
|
|
staged = Path(staging) / f"input.{format}"
|
|
staged.write_bytes(source)
|
|
with converter_path():
|
|
return str(
|
|
pypandoc.convert_file(str(staged), to, format=format, extra_args=list(extra_args))
|
|
)
|
|
|
|
|
|
#: A markdown image as the converter's own writer emits it. Measured against
|
|
#: pandoc 3.10.2 on hand-laid fixtures: a `.docx` picture arrives as
|
|
#: `{width="..." height="..."}` and a `.pptx` one as
|
|
#: ``, so the title form and the attribute form are
|
|
#: both real and a regex written for one of them silently leaves the other's
|
|
#: link in the text.
|
|
_MEDIA_LINK = re.compile(
|
|
r"!\[(?P<alt>[^\]]*)\]\("
|
|
r"(?:<(?P<angle>[^>]*)>|(?P<plain>[^)\s]*))"
|
|
r'(?:\s+"(?P<title>[^"]*)")?\)'
|
|
r"(?P<attrs>\{[^}]*\})?"
|
|
)
|
|
|
|
|
|
def _convert_with_media(
|
|
source: bytes, to: str, format: str, extra_args: Sequence[str]
|
|
) -> tuple[str, dict[str, bytes]]:
|
|
"""The converter call again, with `--extract-media` and the files read back.
|
|
|
|
A separate function rather than a flag on `_convert_bytes` because the
|
|
media must be READ INSIDE the temporary directory's lifetime: the directory
|
|
is removed on every path, and a caller handed a rewritten markdown string
|
|
pointing into it would hold links to files that no longer exist. Returning
|
|
the bytes is what makes the seam closed.
|
|
|
|
The staging path is absolute, so the converter writes absolute links. That
|
|
is deliberate: every one of them is replaced below, and a link that somehow
|
|
survived would carry a temporary directory name into a concept -- a string
|
|
that differs on every run, which a byte-determinism rule would catch loudly
|
|
rather than never.
|
|
"""
|
|
import pypandoc
|
|
|
|
from ._pandoc import converter_path
|
|
|
|
with tempfile.TemporaryDirectory() as staging:
|
|
staged = Path(staging) / f"input.{format}"
|
|
staged.write_bytes(source)
|
|
media_root = Path(staging) / "extracted"
|
|
with converter_path():
|
|
text = str(
|
|
pypandoc.convert_file(
|
|
str(staged),
|
|
to,
|
|
format=format,
|
|
extra_args=[*extra_args, f"--extract-media={media_root}"],
|
|
)
|
|
)
|
|
media: dict[str, bytes] = {}
|
|
if media_root.is_dir():
|
|
for path in sorted(media_root.rglob("*")):
|
|
if path.is_file():
|
|
media[str(path)] = path.read_bytes()
|
|
return text, media
|
|
|
|
|
|
def _rewrite_media_links(text: str, media: dict[str, bytes], collector: _AssetCollector) -> str:
|
|
"""Every converter image link, replaced by this package's own pointer block.
|
|
|
|
UNCONDITIONAL, including the links that cannot be resolved. The converter
|
|
already emitted a markdown image before this existed -- measured on a
|
|
hand-laid `.docx`, today's output carries
|
|
`` with no such file anywhere, which
|
|
`structure._scan_references` reads as a cross-reference to a concept that
|
|
cannot exist. Leaving an unresolvable link in place would keep that defect
|
|
and add a temporary directory name to it.
|
|
"""
|
|
|
|
def replace(match: re.Match[str]) -> str:
|
|
target = match.group("angle") or match.group("plain") or ""
|
|
label = match.group("alt") or match.group("title") or None
|
|
data = media.get(target)
|
|
if data is not None:
|
|
# The name the CONTAINER gave it, not the staging path: pandoc
|
|
# preserves the part name under its own media directory, so
|
|
# `word/media/tabell-84-2.png` arrives as `media/tabell-84-2.png`.
|
|
inside = target.split("/extracted/", 1)[-1]
|
|
return collector.carry(data, name=inside, label=label)
|
|
if not target:
|
|
return collector.reject(
|
|
"image", code="asset_unresolved", reason="the converter emitted no target"
|
|
)
|
|
return collector.local(target, label=label)
|
|
|
|
return _MEDIA_LINK.sub(replace, text)
|
|
|
|
|
|
def _extract_office(suffix: str, data: bytes, collector: _AssetCollector | None = None) -> str:
|
|
"""The five office rows, converted through the vendored binary.
|
|
|
|
Shaped after `_extract_pdf`: the gate is an import probe rather than a
|
|
membership test, third-party failures are wrapped rather than leaked, empty
|
|
output is refused rather than persisted, and the lossiness is stated after
|
|
the parse rather than before it.
|
|
"""
|
|
try:
|
|
import pypandoc # noqa: F401
|
|
except ImportError as exc:
|
|
raise _extra_missing(suffix) from exc
|
|
|
|
spreadsheet = suffix == ".xlsx"
|
|
writer = _SPREADSHEET_WRITER if spreadsheet else _PANDOC_WRITER
|
|
args = _SPREADSHEET_ARGS if spreadsheet else _PANDOC_ARGS
|
|
try:
|
|
if collector is None:
|
|
text = _convert_bytes(data, writer, _PANDOC_FORMATS[suffix], args)
|
|
else:
|
|
text, media = _convert_with_media(data, writer, _PANDOC_FORMATS[suffix], args)
|
|
text = _rewrite_media_links(text, media, collector)
|
|
except ExtractionError:
|
|
raise
|
|
except Exception as exc:
|
|
raise ExtractionError(
|
|
f"the converter failed on this {suffix} file: {exc}",
|
|
code="extractor_convert_error",
|
|
) from exc
|
|
|
|
text = text.strip()
|
|
if not text:
|
|
raise ExtractionError(
|
|
f"the converter returned no text for this {suffix} file; refused "
|
|
"rather than persisted as an empty concept",
|
|
code="extractor_empty_conversion",
|
|
)
|
|
|
|
if spreadsheet:
|
|
text = _drop_converter_decimals(text, data)
|
|
|
|
# After the parse, not before: a run that produced no text has nothing to
|
|
# be lossy about, and warning there would just add noise to a failure.
|
|
warnings.warn(_OFFICE_LOSSY_WARNING, ExtractionWarning, stacklevel=3)
|
|
return text
|
|
|
|
|
|
def _shared_strings(data: bytes) -> frozenset[str]:
|
|
"""Every literal in a workbook's shared string table, or nothing.
|
|
|
|
Read for one purpose: to tell a NUMBER from TEXT THAT LOOKS LIKE ONE. The
|
|
converter renders a numeric cell as a double, so an integral value arrives
|
|
as `5647500.0` -- and a text cell reading `92.0` arrives as `92.0` too. The
|
|
output alone cannot separate them, and rewriting on the output alone would
|
|
silently edit somebody's authored text.
|
|
|
|
Shared strings are the only text the converter recovers from a sheet at
|
|
all: an inline string (`t="inlineStr"`) is read as an EMPTY cell, measured
|
|
while the first xlsx fixture was built (`tests/fixtures/README.md`). So a
|
|
`<digits>.0` that is not in this set did not come from text.
|
|
|
|
Every failure returns the empty set, which makes the rewrite a no-op rather
|
|
than a guess: a workbook this cannot read keeps its converter decimals.
|
|
"""
|
|
try:
|
|
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
|
raw = archive.read("xl/sharedStrings.xml")
|
|
root = ElementTree.fromstring(raw)
|
|
except (KeyError, OSError, zipfile.BadZipFile, ElementTree.ParseError):
|
|
return frozenset()
|
|
return frozenset(
|
|
"".join(node.text or "" for node in item.iter(f"{{{_SSML}}}t")) for item in root
|
|
)
|
|
|
|
|
|
def _drop_converter_decimals(text: str, data: bytes) -> str:
|
|
"""Undo the converter's `N.0` on cells the workbook stores as integers.
|
|
|
|
Cell-scoped and never applied to prose: the pattern is anchored between two
|
|
unescaped pipes, so only a cell whose ENTIRE content is an integer with a
|
|
trailing `.0` is rewritten, and only when that same literal is absent from
|
|
the shared string table.
|
|
"""
|
|
literals = _shared_strings(data)
|
|
|
|
def rewrite(match: re.Match[str]) -> str:
|
|
digits = match.group(2)
|
|
if f"{digits}.0" in literals:
|
|
return match.group(0)
|
|
return f"|{match.group(1)}{digits}{match.group(3)}"
|
|
|
|
return _INTEGRAL_CELL.sub(rewrite, text)
|
|
|
|
|
|
_CORE_EXTRACTORS: dict[str, Callable[[bytes], str]] = {
|
|
".md": _extract_passthrough,
|
|
".txt": _extract_passthrough,
|
|
".csv": _extract_csv,
|
|
".json": _extract_json,
|
|
".html": _extract_html,
|
|
".htm": _extract_html,
|
|
".xml": _extract_xml,
|
|
}
|
|
|
|
# Types the `[extract]` extra ships a parser for. Kept separate from the core
|
|
# registry so "adds no runtime dependency" stays readable at a glance.
|
|
_OPTIONAL_EXTRACTORS: dict[str, Callable[[bytes], str]] = {
|
|
".pdf": _extract_pdf,
|
|
**{suffix: functools.partial(_extract_office, suffix) for suffix in _PANDOC_FORMATS},
|
|
}
|
|
|
|
|
|
# --- provenance: a character range of the extracted text -> a place in the
|
|
# original document ---------------------------------------------------------
|
|
#
|
|
# `source_offset` alone is a position in OUR extraction, so following it back
|
|
# needs the corpus directory, the extractor and its exact version -- none of
|
|
# which a bundle carries. A unit table is that mapping, saved AT EXTRACTION
|
|
# where the two are known to agree, rather than guessed afterwards from text
|
|
# whose page breaks are gone.
|
|
#
|
|
# THE UNIT IS PER FORMAT AND IS NAMED, never assumed:
|
|
#
|
|
# pages a PDF page number, from the document itself.
|
|
# rows a spreadsheet row, within the sheet named by `scope_of`.
|
|
# lines a line of the EXTRACTED text. For `md`/`txt` that text is the
|
|
# dropped file, so the number is the original's own line; for the
|
|
# converted formats it is not, and the key says `lines` rather than
|
|
# `paragraphs` for exactly that reason. Measured on the five K2
|
|
# `.docx` documents: `<w:p>` counts 108/27/65/176/57 against
|
|
# converted-markdown line counts 75/33/67/144/63 -- not one pair
|
|
# agrees, so a `paragraphs` key would name a number the original does
|
|
# not have.
|
|
#
|
|
# The heading a spreadsheet's sheet becomes, as the converter writes it:
|
|
# `## <sheet name> {#sheet-<n>}`. Anchored to the line start so a pipe cell
|
|
# containing a `#` cannot be read as a sheet.
|
|
_SHEET_HEADING = re.compile(r"^#{1,6} (?P<name>.*?) \{#sheet-\d+\}$")
|
|
|
|
# Pandoc's ATTRIBUTE syntax at the end of a heading, which is what the sheet
|
|
# and slide anchors above are an instance of. Deliberately NARROW, because the
|
|
# known-negative is the whole point: an author writing `Mal for {kundenavn}` or
|
|
# `Feltet {"id": 4}` wrote a title, and stripping that would be this same
|
|
# defect pointed the other way.
|
|
#
|
|
# The three narrowings, each doing work: the block must be at the END of the
|
|
# title (`$`), it must OPEN with `#` (pandoc's identifier -- `{.class}` and
|
|
# `{key=val}` alone are not what any converter here emits, and matching them
|
|
# would reach further than measured), and the identifier is the restricted
|
|
# character set pandoc actually generates, so a brace holding a space, a quote
|
|
# or a colon is not an attribute.
|
|
_CONVERTER_ATTRIBUTE = re.compile(r"\s*\{#[A-Za-z0-9_.:-]+\}\s*$")
|
|
|
|
# A line the converter wrote as part of a pipe table. Whether one of them is
|
|
# the table's SEPARATOR is decided by POSITION, never by content: an empty
|
|
# spreadsheet row renders as `| | |` and a separator as `|----|----|`, and
|
|
# every content rule that tells those apart also swallows a data row that
|
|
# happens to hold only dashes. Measured on the K2 price sheet: a content rule
|
|
# ate 8 empty rows and reported the sheet's last row as 92 against a workbook
|
|
# that says 100.
|
|
_TABLE_LINE = "|"
|
|
|
|
|
|
def strip_converter_attribute(title: str) -> str:
|
|
"""Remove a trailing pandoc attribute anchor from a heading's title.
|
|
|
|
ONE definition, read by both title-forming sites: `propose` names a
|
|
segment from an ATX heading, `structure` derives a document title from its
|
|
leading heading, and a rule living in only one of them would strip the
|
|
attribute on one path and leave it on the other -- with the id and the
|
|
title then disagreeing about the same concept.
|
|
|
|
Lives HERE because the attribute is a CONVERTER artefact: `_SHEET_HEADING`
|
|
above is the same syntax read for a different purpose, and this module is
|
|
the one that knows what pandoc writes. That reading must keep its
|
|
attribute, which is why the strip is applied to a title downstream and
|
|
never to the extracted text.
|
|
|
|
RENAMES CONCEPT IDS, by design and with the operator's authorisation
|
|
(2026-09-09): a filename is reduced FROM the title, so the two move
|
|
together. Measured exposure at the time: 2 of 810 concepts on the default
|
|
K2 bundle and 2 of 1108 on Arm B.
|
|
"""
|
|
return _CONVERTER_ATTRIBUTE.sub("", title)
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class SourceUnits:
|
|
"""Where in the ORIGINAL each stretch of the extracted text came from.
|
|
|
|
`starts[i]` is the character offset in the extracted text at which unit
|
|
`numbers[i]` begins, and `scopes[i]` is the sheet that unit belongs to (or
|
|
`None` for a format that has no sheets). The three tuples are parallel and
|
|
`starts` ascends, which is what lets `covering` be a bisection rather than
|
|
a scan.
|
|
|
|
`numbers` is separate from the index on purpose. A PDF page that yielded no
|
|
text is not in this table, and a pipe table's separator line is a row of
|
|
nothing -- in both cases the position in the table and the number in the
|
|
original have already parted company, and an index standing in for a number
|
|
is the off-by-one this whole object exists to prevent.
|
|
"""
|
|
|
|
unit: str
|
|
starts: tuple[int, ...]
|
|
numbers: tuple[int, ...]
|
|
scopes: tuple[str | None, ...] = ()
|
|
|
|
def __post_init__(self) -> None:
|
|
if len(self.starts) != len(self.numbers):
|
|
raise ValueError("a unit table needs one number per start offset")
|
|
if self.scopes and len(self.scopes) != len(self.starts):
|
|
raise ValueError("a unit table needs one scope per start offset, or none at all")
|
|
|
|
def _index(self, offset: int) -> int:
|
|
"""The table row covering `offset`, clamped to the table's own ends."""
|
|
low, high = 0, len(self.starts) - 1
|
|
while low < high:
|
|
middle = (low + high + 1) // 2
|
|
if self.starts[middle] <= offset:
|
|
low = middle
|
|
else:
|
|
high = middle - 1
|
|
return low
|
|
|
|
def covering(self, start: int, end: int) -> tuple[int, int]:
|
|
"""The first and last original unit the half-open `[start, end)` touches.
|
|
|
|
`end` is exclusive, so a range ending exactly where the next unit
|
|
begins does not claim that unit -- a segment that stops at a page
|
|
boundary is on the page it was written on.
|
|
"""
|
|
if not self.starts:
|
|
raise ValueError("an empty unit table locates nothing")
|
|
first = self._index(start)
|
|
last = self._index(max(start, end - 1))
|
|
return self.numbers[first], self.numbers[last]
|
|
|
|
def scope_of(self, offset: int) -> str | None:
|
|
"""The sheet `offset` falls in, or `None` for a format without sheets."""
|
|
if not self.scopes:
|
|
return None
|
|
return self.scopes[self._index(offset)]
|
|
|
|
def scopes_covering(self, start: int, end: int) -> tuple[str | None, ...]:
|
|
"""Every distinct scope the range touches, in order, without repeats."""
|
|
if not self.scopes:
|
|
return ()
|
|
first = self._index(start)
|
|
last = self._index(max(start, end - 1))
|
|
seen: list[str | None] = []
|
|
for scope in self.scopes[first : last + 1]:
|
|
if not seen or seen[-1] != scope:
|
|
seen.append(scope)
|
|
return tuple(seen)
|
|
|
|
|
|
def _line_units(text: str) -> SourceUnits:
|
|
starts: list[int] = []
|
|
offset = 0
|
|
for line in text.split("\n"):
|
|
starts.append(offset)
|
|
offset += len(line) + 1
|
|
return SourceUnits("lines", tuple(starts), tuple(range(1, len(starts) + 1)))
|
|
|
|
|
|
def _pdf_units(data: bytes, headings: bool, ocr: bool, assets: bool = False) -> SourceUnits:
|
|
starts: list[int] = []
|
|
numbers: list[int] = []
|
|
offset = 0
|
|
for page in _pdf_pages(data, headings, ocr, assets):
|
|
starts.append(offset)
|
|
numbers.append(page.number)
|
|
# The page as it reaches the text, pointers included: a locator built
|
|
# from the body alone would drift by two lines per carried image and
|
|
# would name the wrong page from the first one onwards.
|
|
# Through the normalisation door for the same reason one step smaller:
|
|
# the text this table indexes has had its soft hyphens removed, so
|
|
# measuring the page before the door would drift by one character per
|
|
# hyphen and name the wrong page.
|
|
offset += len(normalise_extracted(_pdf_page_text(page))[0]) + len(_PDF_PAGE_SEPARATOR)
|
|
return SourceUnits("pages", tuple(starts), tuple(numbers))
|
|
|
|
|
|
def _spreadsheet_units(text: str) -> SourceUnits | None:
|
|
"""Sheet and row for a converted spreadsheet, or `None` if it is not one.
|
|
|
|
The converter writes one heading per sheet and then one pipe-table line per
|
|
source row, with a separator line after the first. Row numbering therefore
|
|
restarts at every heading and skips that one line by POSITION.
|
|
|
|
The row number is the ORIGINAL sheet's, and that holds exactly as far as
|
|
one converted line per `<row>` element holds. Measured on the two K2
|
|
spreadsheets and both fixtures: 39 rows for 39, 100 for 100, 4 for 4, 6 for
|
|
6 and 3 for 3 -- every one contiguous from row 1. A sheet whose XML omits a
|
|
row entirely would number from the converted table instead, and nothing
|
|
here can see that.
|
|
"""
|
|
starts: list[int] = []
|
|
numbers: list[int] = []
|
|
scopes: list[str | None] = []
|
|
sheet: str | None = None
|
|
seen = 0
|
|
offset = 0
|
|
for line in text.split("\n"):
|
|
heading = _SHEET_HEADING.match(line)
|
|
if heading is not None:
|
|
sheet = heading.group("name")
|
|
seen = 0
|
|
elif sheet is not None and line.startswith(_TABLE_LINE):
|
|
seen += 1
|
|
# The SECOND table line of a sheet is the separator the converter
|
|
# writes under the header, and it is a row of no spreadsheet. Every
|
|
# line after it is one row further on than its position suggests.
|
|
if seen != 2:
|
|
starts.append(offset)
|
|
numbers.append(seen if seen == 1 else seen - 1)
|
|
scopes.append(sheet)
|
|
offset += len(line) + 1
|
|
if not starts:
|
|
return None
|
|
return SourceUnits("rows", tuple(starts), tuple(numbers), tuple(scopes))
|
|
|
|
|
|
def source_units(
|
|
filename: str,
|
|
data: bytes,
|
|
text: str,
|
|
*,
|
|
pdf_headings: bool = False,
|
|
ocr: bool = False,
|
|
assets: bool = False,
|
|
) -> SourceUnits | None:
|
|
"""The unit table for one dropped file, or `None` when it has none.
|
|
|
|
`text` must be what `extract_text` returned for these exact bytes: the
|
|
table indexes that string, and a table built against a different rendering
|
|
would point a consumer at the wrong place with full confidence.
|
|
|
|
`None` is a measurement, not a failure -- a spreadsheet the converter wrote
|
|
no table for has no rows to name, and the caller writes the address without
|
|
a locator rather than inventing one.
|
|
"""
|
|
suffix = Path(filename).suffix.lower()
|
|
if suffix == ".pdf":
|
|
return _pdf_units(data, pdf_headings, ocr, assets)
|
|
if suffix == ".xlsx":
|
|
return _spreadsheet_units(text)
|
|
if suffix in _CORE_EXTRACTORS or suffix in _PANDOC_FORMATS:
|
|
return _line_units(text)
|
|
return None
|
|
|
|
|
|
def extract_text(
|
|
filename: str,
|
|
data: bytes,
|
|
*,
|
|
renderer: Callable[[str], str] | None = None,
|
|
pdf_headings: bool = False,
|
|
ocr: bool = False,
|
|
assets: bool = False,
|
|
resolve: Resolver | None = None,
|
|
) -> str:
|
|
"""Convert one dropped file's bytes to OKF concept text, dispatched by type.
|
|
|
|
`filename` supplies the extension (case-insensitive); `data` is the raw
|
|
bytes. A core stdlib type is extracted; a `[extract]`-gated binary type
|
|
without the extra, and any unregistered extension, fail fast with a typed
|
|
:class:`ExtractionError`. Extracting a `pdf` also emits an
|
|
:class:`ExtractionWarning`: drawn content has no text to recover.
|
|
|
|
`renderer`, when given, is applied to the EXTRACTED TEXT before it is
|
|
returned -- after extraction, never instead of it, so a renderer never has
|
|
to re-implement a reader and the two cannot drift. It is a plain callable
|
|
rather than anything profile-shaped ON PURPOSE: this module is the
|
|
extraction registry and must not import the contract layer, or the
|
|
dependency would run backwards and the registry would stop standing on its
|
|
own. Resolving a profile's NAMED renderer to a function is the caller's
|
|
job, in the layer that already holds the profile.
|
|
|
|
The default is identity, which is what keeps every existing byte-pinned
|
|
golden byte-pinned.
|
|
|
|
`pdf_headings` and `ocr` are PDF-only and both default to off. They are
|
|
branched on here rather than expressed as two more registry rows because
|
|
the registry's contract is `bytes -> str`: a row per option combination
|
|
would be four rows for one reader, and a reader chosen by a suffix lookup
|
|
that also has to consult two flags is not a lookup. A non-PDF caller
|
|
passing either argument gets today's behaviour, silently, which is correct
|
|
-- the options describe a reader, not a policy for the run.
|
|
"""
|
|
return extract_document(
|
|
filename,
|
|
data,
|
|
renderer=renderer,
|
|
pdf_headings=pdf_headings,
|
|
ocr=ocr,
|
|
assets=assets,
|
|
resolve=resolve,
|
|
).text
|
|
|
|
|
|
#: The types that can carry an image, and the reader that places it. Kept apart
|
|
#: from `_CORE_EXTRACTORS` and `_OPTIONAL_EXTRACTORS` rather than folded into
|
|
#: them, and that separation is the byte-identity guarantee: with `assets=False`
|
|
#: not one entry below is consulted and the dispatch is the one every golden,
|
|
#: every pinned bundle and every published digest was measured on. `.csv`,
|
|
#: `.json`, `.md` and `.txt` are absent because the formats carry no image;
|
|
#: `.xlsx` is absent because its converter writes one pipe table per sheet and
|
|
#: a two-line block inside one would break the row locator `_spreadsheet_units`
|
|
#: reads back out of it -- measured 2026-09-16, 0 of 4 K2 workbooks hold any
|
|
#: media at all, so the row is a limit stated rather than a loss taken.
|
|
_ASSET_READERS: dict[str, Callable[[bytes, _AssetCollector], str]] = {
|
|
".html": _extract_html,
|
|
".htm": _extract_html,
|
|
".xml": _extract_xml,
|
|
**{
|
|
suffix: functools.partial(_extract_office, suffix)
|
|
for suffix in _PANDOC_FORMATS
|
|
if suffix != ".xlsx"
|
|
},
|
|
}
|
|
|
|
|
|
def extract_document(
|
|
filename: str,
|
|
data: bytes,
|
|
*,
|
|
renderer: Callable[[str], str] | None = None,
|
|
pdf_headings: bool = False,
|
|
ocr: bool = False,
|
|
assets: bool = False,
|
|
resolve: Resolver | None = None,
|
|
) -> ExtractedDocument:
|
|
"""One dropped file as text PLUS the images that stand inside that text.
|
|
|
|
The entry point :func:`extract_text` keeps for the eight callers that want
|
|
a string, and the one Door B uses since 0.10.0. With `assets=False` -- the
|
|
default, everywhere -- this runs exactly the dispatch that existed before
|
|
the asset layer did, and returns an :class:`ExtractedDocument` whose text is
|
|
byte-identical and whose two image tuples are empty.
|
|
|
|
`resolve` answers for the formats that POINT at a file instead of embedding
|
|
it (`html`, `xml`). Without one every pointer resolves to nothing and is
|
|
stated as such; with one, containment is that resolver's rule and not this
|
|
module's. `pdf` and the office rows embed their images and never consult it.
|
|
"""
|
|
suffix = Path(filename).suffix.lower()
|
|
extractor = _CORE_EXTRACTORS.get(suffix) or _OPTIONAL_EXTRACTORS.get(suffix)
|
|
if extractor is None:
|
|
if suffix in _UNPARSED_OPTIONAL_EXTENSIONS:
|
|
raise _extra_missing(suffix)
|
|
raise ExtractionError(
|
|
f"no extractor is registered for file extension {suffix!r} ({filename!r})",
|
|
code="extractor_unknown",
|
|
)
|
|
collector = _AssetCollector(resolve) if assets else None
|
|
if suffix == ".pdf":
|
|
if pdf_headings or ocr or assets:
|
|
text = _extract_pdf(data, headings=pdf_headings, ocr=ocr, assets=assets)
|
|
else:
|
|
text = extractor(data)
|
|
if collector is not None:
|
|
for page in _pdf_pages(data, pdf_headings, ocr, True):
|
|
collector.images.extend(page.images)
|
|
collector.rejected.extend(page.rejected)
|
|
elif collector is not None and suffix in _ASSET_READERS:
|
|
text = _ASSET_READERS[suffix](data, collector)
|
|
else:
|
|
text = extractor(data)
|
|
rendered, soft_hyphens = normalise_extracted(renderer(text) if renderer is not None else text)
|
|
return ExtractedDocument(
|
|
text=rendered,
|
|
images=tuple(collector.images) if collector is not None else (),
|
|
rejected=tuple(collector.rejected) if collector is not None else (),
|
|
files=tuple(collector.files) if collector is not None else (),
|
|
soft_hyphens=soft_hyphens,
|
|
)
|