feat(assets): every carried image is one a model can be shown
Chosen: a stdlib BMP reader, because `read_image` is on the CORE path and an asset's name is its content digest. Measured first, as the order requires: Pillow 12.3.0 IS in this tree (transitively under `pdfplumber`) and it DOES decode RLE8 correctly -- a hand-written stdlib decoder and Pillow agree on 19 of 19 of R761's real files, RGB per pixel. So the choice does not rest on capability. It rests on two properties of this package: `.html` and `.xml` carry images with no `[extract]` extra installed, so a Pillow converter either makes a core path depend on an optional binary wheel or buys the second runtime dependency; and encoding through an installed library would make a bundle's identity move with that library's version, which is the property 0.10.0 felled page rasterisation over and `encode_png`'s docstring already defends. Pillow keeps the job it is good for: the INDEPENDENT decoder in the tests, on neither side of the conversion. The defect, measured over the frozen R761 delivery's `assets/`, denominator 50: 29 JPEG, 2 PNG and 19 RLE8 BMP. The 19 are byte-correct files nothing reads, so 19 figures were present and invisible while `images: N` reported that they had arrived. - `VIEWABLE_MEDIA_TYPES` is tested against every asset's SNIFFED type, so it is a property and not a list of formats we met. WebP is on it and `sniff` does not recognise one; the limit is stated, not implied. - `bmp_to_png`: 8-bit uncompressed, 8-bit RLE8, 24-bit uncompressed. All five RLE8 opcodes. 19 of 19 real files convert with RGB identical to Pillow's decoding of the source, 2 366 365 pixels compared. - `asset_not_viewable` and `asset_bmp_unsupported`, both published, both leaving the concept's "not carried" line. - Traceability on the pointer's second line, where the rest of the asset metadata already lives: original media type, original sha256 in full, new sha256 in full. A converted asset is ONE asset. - The ceiling is paid on the DECLARATION before a row is allocated, and an RLE run is one clipped slice -- painting pixel by pixel leaves the memory bounded and the CPU unbounded. Two repairs the change forced, each measured rather than assumed: - `tests/test_assets.py`'s "dimensions absent is absent" used a TIFF, which is now refused before `read_image` returns. The property still has a reachable case -- a JPEG whose frame header never arrives -- and uses it. - `asset_holds` in the accounting gate proved a carry by hashing the SOURCE file, which a converted image's bundle cannot satisfy. It now also reads the two digests the bundle states and HASHES THE ASSET ITSELF, so a bundle claiming a conversion it did not perform still fails. `tools/okf_asset_census.py` is the committed instrument for the known-positive: one row per image, from two pinned trees. It was caught by the rule it serves -- its first version handed `_pdf_images` the wrong page object and reported 0 images over 67 PDFs with exit 0. The attribute is asserted now and a known-positive runs before the sweep. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
955ec4b2ca
commit
b0b5e71658
10 changed files with 826 additions and 5 deletions
|
|
@ -361,6 +361,19 @@ def _sha12(path: Path) -> str:
|
|||
return _sha256(path)[:12]
|
||||
|
||||
|
||||
#: The conversion clause `assets.render_block` writes: the source's media type
|
||||
#: and sha256, and the media type and sha256 of what the run actually carried.
|
||||
#: One expression, so the judge has one definition of the claim it verifies.
|
||||
_CONVERSION = re.compile(
|
||||
r"converted from \S+ sha256:(?P<before>[0-9a-f]{64}) to \S+ sha256:(?P<after>[0-9a-f]{64})"
|
||||
)
|
||||
|
||||
|
||||
def _conversions(bundle_text: str) -> dict[str, str]:
|
||||
"""source digest -> the digest the bundle says it carried instead."""
|
||||
return {m.group("before"): m.group("after") for m in _CONVERSION.finditer(bundle_text)}
|
||||
|
||||
|
||||
def asset_holds(build: Build, source: Path) -> bool:
|
||||
"""Did the run carry THESE bytes, placed under their own content address?
|
||||
|
||||
|
|
@ -377,10 +390,26 @@ def asset_holds(build: Build, source: Path) -> bool:
|
|||
by construction -- and it would be wrong: measured 2026-09-18 on R761,
|
||||
whose own hrefs carry spaces, capitals and parentheses, a judge checking
|
||||
the full name reported 50 of 50 carried images as missing.
|
||||
|
||||
A SECOND ROUTE, for an image the build CONVERTS. Since the viewable-asset
|
||||
round a source in a format no model can be shown reaches the bundle as a
|
||||
PNG, so its own bytes are not in `assets/` and never will be -- measured,
|
||||
the day that landed R761 went from 0 to 19 claimed-and-not-found, which is
|
||||
exactly its RLE8 BMP count. The bundle states both digests on the pointer
|
||||
line, and this reads them and then HASHES THE ASSET ITSELF: the claim is
|
||||
accepted only when a file in `assets/` really holds the bytes the bundle
|
||||
says were written. A bundle claiming a conversion it did not perform still
|
||||
fails, which is the difference between reading the bundle and believing
|
||||
the report.
|
||||
"""
|
||||
digest = _sha256(source)
|
||||
return any(
|
||||
if any(
|
||||
found == digest and name.startswith(digest[:12]) for name, found in build.assets.items()
|
||||
):
|
||||
return True
|
||||
written = _conversions(build.bundle_text).get(digest)
|
||||
return written is not None and any(
|
||||
found == written and name.startswith(written[:12]) for name, found in build.assets.items()
|
||||
)
|
||||
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue