docs(readme): one file-type table, not two

`### Binary extraction` carried its own six-row Format/Reader/Evidence table
over the same rows the pinned table now holds. It was true when written and
reachable by exactly the failure this module exists for: three evidence classes
copied into prose no test reads. That section now points at the pinned table
and keeps its prose about the extra.

A fifth assertion in `tests/test_docs_promises.py` holds the duplicate gone.
Known-positive: the same query finds 8 table lines in that section on the
previous commit, so it can go red.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-12 21:36:28 +02:00
commit d310dbb34e
3 changed files with 26 additions and 11 deletions

View file

@ -1057,14 +1057,11 @@ rather than found on `PATH`, with its version asserted against a pin. A host
carrying a different converter is refused, not silently used: extraction is
deterministic within a converter version and not across one.
| Format | Reader | Evidence |
|---|---|---|
| `pdf` | `pdfplumber` | measured |
| `docx` | converter | measured |
| `xlsx` | converter | measured |
| `pptx` | converter | **constructed** |
| `odt` | converter | **constructed** |
| `rtf` | converter | **constructed** |
These six rows have their reader and their evidence class in the one table
this README carries: [Supported file types](#supported-file-types), which is
pinned to the extraction registry cell by cell. A second table here would be a
copy nothing checks, and a copy of an evidence class is exactly the thing that
goes false quietly.
**`constructed` means what it says, and it is a weaker word than `measured`
on purpose.** The corpus this work was measured on contains **zero** `pptx`,