test(accounting): content-accounting gate for okf build, written red
Capability loop step 3, no capability. tools/okf_accounting_gate.py asks, per README file type, how many of the elements a SOURCE holds okf build books as carried / pointer / coded rejection, with unaccounted and double-booked both required to be 0. Exit 1 today on rows 2, 3, 4 and 6. The fasit is tools/okf_witness.py (stdlib + pdfplumber + poppler, no package import; tested on the live import graph), committed as tests/fixtures/accounting/*inventory.json over one fixture per type. Measured: no source inventory (0 of 13); two graphics/ files carried through documents AND counted extractor_unknown (50 on R761 under --gate none); a refused document logged "0 carried of 0 found"; R761 refused whole because guard 1.4.0 treats its 71 U+00AD soft hyphens as an invisible carrier (asked of the security repo). The two R761 witnesses agree once STS labels are counted by role, not tag. Report: docs/2026-09-17-innholdsregnskapet-rod-gate.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
bc39e8091f
commit
0b00de4408
26 changed files with 2598 additions and 0 deletions
13
CLAUDE.md
13
CLAUDE.md
|
|
@ -890,6 +890,19 @@ and fixtures, never code.
|
|||
move (it is not delivered on the Arm B bundle either, by a different
|
||||
mechanism). hit@8 over the six published questions holds at **5 of 6 on both
|
||||
K2 bundles**, so the default move cost the retrieval side nothing.
|
||||
- **Account for content: `python3 tools/okf_accounting_gate.py` (2026-09-17), RED by
|
||||
design.** Capability loop step 3, no capability: per README file type, of M
|
||||
elements the SOURCE holds, how many `okf build` books as carried / pointer /
|
||||
coded rejection, with u (unaccounted) and d (double-booked) required to be 0.
|
||||
The fasit is `tools/okf_witness.py`, which imports no package module (tested on
|
||||
the live import graph), committed as `tests/fixtures/accounting/*inventory.json`.
|
||||
It reads a door that does not exist yet (`okf build --accounting PATH`, contract
|
||||
in the module docstring). Red today on rows 2, 3, 4 and 6: no inventory; a
|
||||
`graphics/` file carried through a document AND counted `extractor_unknown`
|
||||
(2 on the fixtures, 50 on R761 under `--gate none`); a refused document logged
|
||||
`0 carried of 0 found`; R761 refused whole because guard 1.4.0 treats its 71
|
||||
U+00AD soft hyphens as an invisible carrier (asked of the security repo).
|
||||
Report: `docs/2026-09-17-innholdsregnskapet-rod-gate.md`.
|
||||
- **Judge a bundle: `okf quality <bundle>` (G37, 2026-09-12).** A per-file-type
|
||||
verdict, with the denominator on every line, and it is a SEPARATE command
|
||||
from `okf check` on purpose: `check` is the contract check, and a green one
|
||||
|
|
|
|||
149
docs/2026-09-17-innholdsregnskapet-rod-gate.md
Normal file
149
docs/2026-09-17-innholdsregnskapet-rod-gate.md
Normal file
|
|
@ -0,0 +1,149 @@
|
|||
# Content accounting in `okf build`: the gate, written red
|
||||
|
||||
Capability loop, step 3. This session adds no capability. It adds the gate a
|
||||
later capability has to turn green, together with the fasit that gate judges
|
||||
against.
|
||||
|
||||
- Command: `python3 tools/okf_accounting_gate.py [--json] [--r761 DIR] [--consume]`
|
||||
- Exit codes: 0 when every failing row (1–6) is green, 1 otherwise, 2 when
|
||||
the gate did not run.
|
||||
- Fasit: `tools/okf_witness.py`, committed as
|
||||
`tests/fixtures/accounting/*inventory.json`.
|
||||
- Tests: `tests/test_accounting_gate.py`, 42 tests, green.
|
||||
|
||||
## Why a gate at all
|
||||
|
||||
`okf build`'s conservation identity, `merged + coded rejections == N`, counts
|
||||
FILES. Two things happen that the identity cannot see, and both were
|
||||
reproduced here on `bc39e80` (v0.10.0) before any code was written:
|
||||
|
||||
- **A file can be merged while content inside it is gone.** No inventory of
|
||||
the source exists, so nothing can state what a merged file lost.
|
||||
- **A file can be carried and rejected at once (finding B).** The fixture
|
||||
corpus has a `graphics/` directory next to its documents. Its two PNGs are
|
||||
carried into `assets/` through the documents that point at them, and the
|
||||
same two files are also counted as `extractor_unknown` rejections. Each has
|
||||
two fates.
|
||||
- **A rejected document is logged as if it held nothing (finding C).** The
|
||||
fixture document the guard refuses produces
|
||||
`Images: 0 carried of 0 found`, although its source declares one image.
|
||||
|
||||
## The fasit is independent
|
||||
|
||||
The witness imports no `llm_ingestion_okf` module. A test runs every witness
|
||||
in a fresh process and then reads `sys.modules`. A control run of the same
|
||||
check, with the package pre-loaded, shows the check fires.
|
||||
|
||||
The witness counts each format by the format's own rules:
|
||||
|
||||
- STS: ElementTree, and the publisher's JSON node tree.
|
||||
- docx, pptx, xlsx and odt: the zip members' own XML.
|
||||
- PDF: pdfplumber objects, and poppler (`pdfinfo`, `pdfimages -list`).
|
||||
- HTML: `html.parser`.
|
||||
- md, txt, csv, json and rtf: stdlib readers, and a control-word scan for rtf.
|
||||
|
||||
Four documents are also pinned to hand counts.
|
||||
|
||||
**Two witnesses disagreed by tag, and both disagreements were about placement
|
||||
rather than content** (R761, measured). A section's label is `sec/label` on
|
||||
7 714 sections in the XML. The JSON puts it at `sec/label` on 4 954 of them
|
||||
and at `sec/title/label` on the other 2 760. A table's label is
|
||||
`table-wrap/label` in the XML (10) and `table-wrap/table/caption` in the JSON
|
||||
(10). The witness therefore counts ROLES, and the mapping is written in
|
||||
`okf_witness._sts_role` together with these numbers. After the mapping, the
|
||||
two R761 witnesses agree on all ten roles.
|
||||
|
||||
The PM's figures were re-measured and all hold: `p` 12 528, `title` 2 761,
|
||||
`td` 197, `list-item` 111, `graphic` 50, 109 other files in the zip (50
|
||||
pointed at, 59 not). The PDF witnesses agree at 701 pages and 50 images;
|
||||
pdfplumber takes 50 s for that and poppler 1 s.
|
||||
|
||||
## The door the capability must open
|
||||
|
||||
The gate reads `okf build --accounting PATH`: one JSON object with a
|
||||
per-document `inventory` taken before extraction and before the gate,
|
||||
per-element `fates` (`carried` / `pointer` / `rejected{code}`), and exactly
|
||||
one `fate` per inbox file that is not a document. It also reads one line in
|
||||
`log.md` per rejected document:
|
||||
|
||||
<source_file>: <M> elements found in the source, 0 carried: document rejected `<code>`
|
||||
|
||||
The full contract is in the module docstring. Until the flag exists, row 2 is
|
||||
0 of 13, and row 3 books every element as unaccounted. For a file, it derives
|
||||
the file's fates from the bundle bytes (assets are named by content) and from
|
||||
K1b.
|
||||
|
||||
## Finding A: why the guard refuses R761, and the decision
|
||||
|
||||
The cause is measured. The R761 XML carries **71 U+00AD SOFT HYPHEN** and no
|
||||
other invisible codepoint, and the extracted text carries the same 71. Guard
|
||||
1.4.0 lists U+00AD in `_ZERO_WIDTH_CPS` (`output.py:250`). One soft hyphen is
|
||||
enough to fail_secure at both tiers; the same string without it is `clean`.
|
||||
Image pointers are not the cause: `--no-assets` gives the same refusal, as the
|
||||
PM also measured.
|
||||
|
||||
This is a question for the guard. It was sent to
|
||||
`llm-ingestion-pipeline-security` with the repro, and nothing was built there.
|
||||
|
||||
**Decision on exit 0 when nothing was persisted:** `okf build` should NOT exit
|
||||
0 when at least one document was extracted and none was persisted. The
|
||||
conservation identity holding is necessary, not sufficient.
|
||||
|
||||
- This repository already refuses the analogous case: a run with zero plans
|
||||
fails (exit 2) rather than emitting a flat bundle and calling it success.
|
||||
- `okf build` is the trusted-source path, an operator pointing at their own
|
||||
folder, where "everything refused" is a finding and not a normal outcome.
|
||||
- A script running `okf build && okf skill` would otherwise wire a skill to an
|
||||
empty bundle.
|
||||
|
||||
Door B's library function keeps its behaviour, because for a hostile inbox
|
||||
"all rejected" is correct. The change is an exit code, which this order
|
||||
forbids, so it belongs to the capability order and not here.
|
||||
|
||||
## Exceptions proposed, none approved
|
||||
|
||||
The gate lists three exceptions in every run. None of them lowers a
|
||||
denominator:
|
||||
|
||||
- **PDF headings, paragraphs and tables.** A PDF without a structure tree
|
||||
declares none of them, so no witness can count them.
|
||||
- **xlsx images.** The reader deliberately does not carry them.
|
||||
- **Images in md, txt, csv, json, odt and rtf.** No reader for these types
|
||||
carries image bytes.
|
||||
|
||||
Approving any of them is the operator's decision.
|
||||
|
||||
## Output at this commit
|
||||
|
||||
See the order's closing message for the full text. In short:
|
||||
|
||||
| Row | Result | Status |
|
||||
|---|---|---|
|
||||
| 1 | 13 of 13 | GREEN |
|
||||
| 2 | 0 of 13 | RED |
|
||||
| 3 | 0 of 15, u = 248, d = 2 | RED |
|
||||
| 4 | 0 of 1 | RED |
|
||||
| 5 | 4 of 4 | GREEN |
|
||||
| 6 | 59 of 110 (R761, two builds) | RED |
|
||||
|
||||
Row 6 runs two builds:
|
||||
|
||||
- **Default gate:** the document is refused, so every element is unaccounted.
|
||||
- **`--gate none`:** the document is persisted, and the 50 pictures it points at
|
||||
are both carried and rejected, which is finding B at scale (d = 50). A first
|
||||
run reported 57: 7 of the 59 unpointed files share bytes with pointed ones,
|
||||
and bytes in `assets/` had been credited to them. A file nothing points at
|
||||
is never carried; a test holds that.
|
||||
|
||||
Row 7 (`--consume`) is a diagnostic and never fails.
|
||||
|
||||
## Limits
|
||||
|
||||
- The witness vocabulary is ours. A format element the witness does not name
|
||||
is invisible to the gate.
|
||||
- The fixture corpus is constructed, one document per type. The real-corpus
|
||||
row is N = 1 (R761).
|
||||
- Row 6 needs the R761 source on this machine. Without it the row is RED
|
||||
locally and SKIPPED when `CI` is set, never green.
|
||||
- The `--r761` default points into another repository's data directory, and
|
||||
the gate only reads from it.
|
||||
28
tests/fixtures/README.md
vendored
28
tests/fixtures/README.md
vendored
|
|
@ -175,6 +175,34 @@ The version range that carries this lives in `pyproject.toml`'s
|
|||
`[project.optional-dependencies] extract`, with the same reasoning at the
|
||||
declaration site.
|
||||
|
||||
## The content-accounting fixtures (`accounting/`)
|
||||
|
||||
The fasit side of `tools/okf_accounting_gate.py`. `accounting/corpus/` holds
|
||||
one document per row of README's file-type table (13 of 13) plus a `graphics/`
|
||||
directory next to them that the HTML, STS and markdown documents point at --
|
||||
the layout under which a picture is carried through a document AND booked as a
|
||||
rejected file. `accounting/rejected/` holds one HTML document with a
|
||||
zero-width space in its prose, which the guard refuses at every tier, and the
|
||||
image it points at.
|
||||
|
||||
`inventory.json` and `rejected-inventory.json` are what `tools/okf_witness.py`
|
||||
counts in those two directories, committed as data and regenerated only with
|
||||
that tool:
|
||||
|
||||
```
|
||||
python3 tools/okf_witness.py tests/fixtures/accounting/corpus > tests/fixtures/accounting/inventory.json
|
||||
python3 tools/okf_witness.py tests/fixtures/accounting/rejected > tests/fixtures/accounting/rejected-inventory.json
|
||||
```
|
||||
|
||||
`witness/prosess-84-sts.twin.json` is the STS document written by hand in the
|
||||
publisher's JSON node form (`standardContent`, nodes with `e`/`t`/`x`), so the
|
||||
two STS witnesses can be compared on a fixture as well as on R761. Eight of
|
||||
the thirteen documents are byte copies of fixtures documented above
|
||||
(`image-inbox/`, `k2-office/`, `prisark.xlsx`); the other five
|
||||
(`notat.md`, `logg.txt`, `mengder.csv`, `parametre.json`, `side.htm`) are
|
||||
written here, and `notat.md` carries a fenced `# ...` line that is not a
|
||||
heading.
|
||||
|
||||
## What these fixtures do not cover
|
||||
|
||||
Structured table recovery. Measured on real Vegnormalene, only 45 of 196
|
||||
|
|
|
|||
BIN
tests/fixtures/accounting/corpus/graphics/figur-84-1.png
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/graphics/figur-84-1.png
vendored
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 94 B |
BIN
tests/fixtures/accounting/corpus/graphics/tabell-84-2.png
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/graphics/tabell-84-2.png
vendored
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 122 B |
52
tests/fixtures/accounting/corpus/krav-rikt-tekstformat.rtf
vendored
Normal file
52
tests/fixtures/accounting/corpus/krav-rikt-tekstformat.rtf
vendored
Normal file
|
|
@ -0,0 +1,52 @@
|
|||
{\rtf1\ansi\ansicpg1252\deff0{\fonttbl{\f0\froman Times New Roman;}}\pard\sa180\b\fs32 Kravspesifikasjon for tunnelbelysning\b0\fs24\par
|
||||
\pard\sa180 Dokumentet samler kravene til belysning i vegtunneler over 500 meter.\par
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Dokumentnummer:\cell \pard\intbl SVV-2026-0417\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Tittel:\cell \pard\intbl Tunnelbelysning i hovedl\u248 ?pet\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Ansvarlig avdeling:\cell \pard\intbl Utbyggingsdivisjonen\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Fagomr\u229 ?de:\cell \pard\intbl Elektro og belysning\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Versjon:\cell \pard\intbl 2.3 godkjent\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Gyldig fra:\cell \pard\intbl 01.03.2026\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Erstatter:\cell \pard\intbl SVV-2024-0188\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Tunnellengde:\cell \pard\intbl 1240 meter\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Dimensjonerende fart:\cell \pard\intbl 80 km/t\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl \u197 ?rsd\u248 ?gntrafikk:\cell \pard\intbl 12400 kj\u248 ?ret\u248 ?y\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Terskelluminans:\cell \pard\intbl 145 candela\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Overgangssone:\cell \pard\intbl Tre trinn nedtrapping\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Innerstrekning:\cell \pard\intbl 3,0 candela\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Utkj\u248 ?ringssone:\cell \pard\intbl Ingen forsterkning\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Fargetemperatur:\cell \pard\intbl 4000 kelvin\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Fargegjengivelse:\cell \pard\intbl Ra st\u248 ?rre enn 70\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl N\u248 ?dbelysning:\cell \pard\intbl 60 minutter drift\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Vedlikeholdsfaktor:\cell \pard\intbl 0,80 beregnet\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl M\u229 ?lemetode:\cell \pard\intbl M\u229 ?ling med luminanskamera\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000
|
||||
\pard\intbl Avviksh\u229 ?ndtering:\cell \pard\intbl S\u248 ?knad om fravik\cell \row
|
||||
\pard\sa180 Luminansmatrise per trafikklasse\par
|
||||
\trowd\trgaph108\cellx3000\cellx6000\cellx9000\cellx12000
|
||||
\pard\intbl Sone\cell \pard\intbl Klasse A\cell \pard\intbl Klasse B\cell \pard\intbl Klasse C\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000\cellx9000\cellx12000
|
||||
\pard\intbl Terskel\cell \pard\intbl 150 cd\cell \pard\intbl 120 cd\cell \pard\intbl 95 cd\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000\cellx9000\cellx12000
|
||||
\pard\intbl Overgang\cell \pard\intbl 45 cd\cell \pard\intbl 36 cd\cell \pard\intbl 28 cd\cell \row
|
||||
\trowd\trgaph108\cellx3000\cellx6000\cellx9000\cellx12000
|
||||
\pard\intbl Indre\cell \pard\intbl 3,5 cd\cell \pard\intbl 3,0 cd\cell \pard\intbl 2,5 cd\cell \row
|
||||
}
|
||||
BIN
tests/fixtures/accounting/corpus/krav-tekstdokument.odt
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/krav-tekstdokument.odt
vendored
Normal file
Binary file not shown.
6
tests/fixtures/accounting/corpus/logg.txt
vendored
Normal file
6
tests/fixtures/accounting/corpus/logg.txt
vendored
Normal file
|
|
@ -0,0 +1,6 @@
|
|||
Driftslogg for tunnelen.
|
||||
|
||||
Vask utført i mars.
|
||||
Lys kontrollert i april.
|
||||
|
||||
Ingen avvik registrert.
|
||||
3
tests/fixtures/accounting/corpus/mengder.csv
vendored
Normal file
3
tests/fixtures/accounting/corpus/mengder.csv
vendored
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
Post,Mengde,Enhet
|
||||
84.1,120,m3
|
||||
84.2,45,m2
|
||||
|
19
tests/fixtures/accounting/corpus/notat.md
vendored
Normal file
19
tests/fixtures/accounting/corpus/notat.md
vendored
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
# Drift av tunneler
|
||||
|
||||
Tunnelen driftes etter håndboka.
|
||||
|
||||
## Vask
|
||||
|
||||
Vask skjer to ganger i året.
|
||||
|
||||
| Sone | Intervall |
|
||||
| --- | --- |
|
||||
| Inngang | 6 mnd |
|
||||
| Hovedløp | 12 mnd |
|
||||
|
||||

|
||||
|
||||
```bash
|
||||
# ikke en overskrift
|
||||
okf build .
|
||||
```
|
||||
1
tests/fixtures/accounting/corpus/parametre.json
vendored
Normal file
1
tests/fixtures/accounting/corpus/parametre.json
vendored
Normal file
|
|
@ -0,0 +1 @@
|
|||
{"tunnel": "Eksempeltunnelen", "lengde_m": 1240, "soner": ["inngang", "hovedlop"], "belysning": {"terskel_cd": 145, "fart_kmt": 80}}
|
||||
BIN
tests/fixtures/accounting/corpus/prisark.xlsx
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/prisark.xlsx
vendored
Normal file
Binary file not shown.
BIN
tests/fixtures/accounting/corpus/prosess-84-notat.docx
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/prosess-84-notat.docx
vendored
Normal file
Binary file not shown.
BIN
tests/fixtures/accounting/corpus/prosess-84-presentasjon.pptx
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/prosess-84-presentasjon.pptx
vendored
Normal file
Binary file not shown.
18
tests/fixtures/accounting/corpus/prosess-84-sts.xml
vendored
Normal file
18
tests/fixtures/accounting/corpus/prosess-84-sts.xml
vendored
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
<?xml version="1.0" encoding="UTF-8"?>
|
||||
<standard xmlns:xlink="http://www.w3.org/1999/xlink">
|
||||
<front><std-ident><doc-number>R761</doc-number></std-ident></front>
|
||||
<body>
|
||||
<sec>
|
||||
<label>84</label>
|
||||
<title>Konstruksjoner av betong</title>
|
||||
<sec>
|
||||
<label>84.1</label>
|
||||
<title>Toleranser</title>
|
||||
<p>Toleranseklasse er gitt i tabell 84-2.</p>
|
||||
<graphic xlink:href="graphics/tabell-84-2.png"/>
|
||||
<p>Figur 84-1 viser prinsippet.</p>
|
||||
<graphic xlink:href="figur-84-1.png"/>
|
||||
</sec>
|
||||
</sec>
|
||||
</body>
|
||||
</standard>
|
||||
BIN
tests/fixtures/accounting/corpus/prosess-84-tabell.pdf
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/prosess-84-tabell.pdf
vendored
Normal file
Binary file not shown.
14
tests/fixtures/accounting/corpus/prosess-84-web.html
vendored
Normal file
14
tests/fixtures/accounting/corpus/prosess-84-web.html
vendored
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
<!doctype html>
|
||||
<html><head><title>Prosess 84</title></head>
|
||||
<body>
|
||||
<h1>84 Konstruksjoner av betong</h1>
|
||||
<p>Toleranseklasse for de enkelte konstruksjonsdeler er gitt i tabell 84-2.</p>
|
||||
<img src="graphics/tabell-84-2.png" alt="Tabell 84-2 Toleranseklasser">
|
||||
<p>Figuren under viser prinsippet.</p>
|
||||
<figure>
|
||||
<img src="graphics/figur-84-1.png" alt="Figur 84-1 Prinsippskisse">
|
||||
<figcaption>Figur 84-1 Prinsippskisse</figcaption>
|
||||
</figure>
|
||||
<p>Og en som ligger et annet sted:</p>
|
||||
<img src="https://example.invalid/ekstern.png" alt="Ekstern figur">
|
||||
</body></html>
|
||||
9
tests/fixtures/accounting/corpus/side.htm
vendored
Normal file
9
tests/fixtures/accounting/corpus/side.htm
vendored
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
<!doctype html>
|
||||
<html><head><title>Side</title></head>
|
||||
<body>
|
||||
<h1>Belysning</h1>
|
||||
<p>Terskelluminansen er 145 candela.</p>
|
||||
<h2>Soner</h2>
|
||||
<ul><li>Inngang</li><li>Hovedløp</li></ul>
|
||||
<table><tr><th>Sone</th><th>Lengde</th></tr><tr><td>Inngang</td><td>120 m</td></tr></table>
|
||||
</body></html>
|
||||
230
tests/fixtures/accounting/inventory.json
vendored
Normal file
230
tests/fixtures/accounting/inventory.json
vendored
Normal file
|
|
@ -0,0 +1,230 @@
|
|||
{
|
||||
"documents": {
|
||||
"krav-rikt-tekstformat.rtf": {
|
||||
"elements": {
|
||||
"cell": 56,
|
||||
"image": 0,
|
||||
"paragraph": 3,
|
||||
"table_row": 24
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".rtf",
|
||||
"witness": "rtf control words"
|
||||
},
|
||||
"krav-tekstdokument.odt": {
|
||||
"elements": {
|
||||
"cell": 56,
|
||||
"heading": 1,
|
||||
"image": 0,
|
||||
"list_item": 0,
|
||||
"paragraph": 2,
|
||||
"table": 2
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".odt",
|
||||
"witness": "odt zip xml"
|
||||
},
|
||||
"logg.txt": {
|
||||
"elements": {
|
||||
"line": 4,
|
||||
"paragraph": 3
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".txt",
|
||||
"witness": "text lines"
|
||||
},
|
||||
"mengder.csv": {
|
||||
"elements": {
|
||||
"cell": 6,
|
||||
"header_cell": 3,
|
||||
"row": 2
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".csv",
|
||||
"witness": "csv"
|
||||
},
|
||||
"notat.md": {
|
||||
"elements": {
|
||||
"code_block": 1,
|
||||
"heading": 2,
|
||||
"image": 1,
|
||||
"paragraph": 2,
|
||||
"table": 1,
|
||||
"table_row": 2
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "graphics/tabell-84-2.png",
|
||||
"target": "graphics/tabell-84-2.png"
|
||||
}
|
||||
],
|
||||
"suffix": ".md",
|
||||
"witness": "markdown lines"
|
||||
},
|
||||
"parametre.json": {
|
||||
"elements": {
|
||||
"key": 6,
|
||||
"value": 6
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".json",
|
||||
"witness": "json"
|
||||
},
|
||||
"prisark.xlsx": {
|
||||
"elements": {
|
||||
"cell": 18,
|
||||
"image": 0,
|
||||
"row": 9,
|
||||
"sheet": 2
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".xlsx",
|
||||
"witness": "xlsx zip xml"
|
||||
},
|
||||
"prosess-84-notat.docx": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"footnote": 0,
|
||||
"heading": 0,
|
||||
"image": 1,
|
||||
"paragraph": 2,
|
||||
"table": 0
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "embedded",
|
||||
"ref": "",
|
||||
"target": null
|
||||
}
|
||||
],
|
||||
"suffix": ".docx",
|
||||
"witness": "docx zip xml"
|
||||
},
|
||||
"prosess-84-presentasjon.pptx": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"image": 1,
|
||||
"paragraph": 0,
|
||||
"slide": 1,
|
||||
"table": 0,
|
||||
"title": 1
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "embedded",
|
||||
"ref": "",
|
||||
"target": null
|
||||
}
|
||||
],
|
||||
"suffix": ".pptx",
|
||||
"witness": "pptx zip xml"
|
||||
},
|
||||
"prosess-84-sts.xml": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"footnote": 0,
|
||||
"image": 2,
|
||||
"list_item": 0,
|
||||
"paragraph": 2,
|
||||
"section": 2,
|
||||
"section_label": 2,
|
||||
"table": 0,
|
||||
"table_label": 0,
|
||||
"title": 2
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "graphics/tabell-84-2.png",
|
||||
"target": "graphics/tabell-84-2.png"
|
||||
},
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "figur-84-1.png",
|
||||
"target": "graphics/figur-84-1.png"
|
||||
}
|
||||
],
|
||||
"suffix": ".xml",
|
||||
"witness": "xml.etree"
|
||||
},
|
||||
"prosess-84-tabell.pdf": {
|
||||
"elements": {
|
||||
"image": 2,
|
||||
"page": 1
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "embedded",
|
||||
"ref": "",
|
||||
"target": null
|
||||
},
|
||||
{
|
||||
"kind": "embedded",
|
||||
"ref": "",
|
||||
"target": null
|
||||
}
|
||||
],
|
||||
"suffix": ".pdf",
|
||||
"witness": "pdfplumber objects"
|
||||
},
|
||||
"prosess-84-web.html": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"heading": 1,
|
||||
"image": 3,
|
||||
"list_item": 0,
|
||||
"paragraph": 3,
|
||||
"table": 0
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "graphics/tabell-84-2.png",
|
||||
"target": "graphics/tabell-84-2.png"
|
||||
},
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "graphics/figur-84-1.png",
|
||||
"target": "graphics/figur-84-1.png"
|
||||
},
|
||||
{
|
||||
"kind": "remote",
|
||||
"ref": "https://example.invalid/ekstern.png",
|
||||
"target": null
|
||||
}
|
||||
],
|
||||
"suffix": ".html",
|
||||
"witness": "html.parser"
|
||||
},
|
||||
"side.htm": {
|
||||
"elements": {
|
||||
"cell": 4,
|
||||
"heading": 2,
|
||||
"image": 0,
|
||||
"list_item": 2,
|
||||
"paragraph": 1,
|
||||
"table": 1
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".htm",
|
||||
"witness": "html.parser"
|
||||
}
|
||||
},
|
||||
"files": {
|
||||
"graphics/figur-84-1.png": {
|
||||
"pointed_at_by": [
|
||||
"prosess-84-sts.xml",
|
||||
"prosess-84-web.html"
|
||||
]
|
||||
},
|
||||
"graphics/tabell-84-2.png": {
|
||||
"pointed_at_by": [
|
||||
"notat.md",
|
||||
"prosess-84-sts.xml",
|
||||
"prosess-84-web.html"
|
||||
]
|
||||
}
|
||||
},
|
||||
"witness_version": 1
|
||||
}
|
||||
31
tests/fixtures/accounting/rejected-inventory.json
vendored
Normal file
31
tests/fixtures/accounting/rejected-inventory.json
vendored
Normal file
|
|
@ -0,0 +1,31 @@
|
|||
{
|
||||
"documents": {
|
||||
"avvist.html": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"heading": 1,
|
||||
"image": 1,
|
||||
"list_item": 0,
|
||||
"paragraph": 1,
|
||||
"table": 0
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "graphics/figur.png",
|
||||
"target": "graphics/figur.png"
|
||||
}
|
||||
],
|
||||
"suffix": ".html",
|
||||
"witness": "html.parser"
|
||||
}
|
||||
},
|
||||
"files": {
|
||||
"graphics/figur.png": {
|
||||
"pointed_at_by": [
|
||||
"avvist.html"
|
||||
]
|
||||
}
|
||||
},
|
||||
"witness_version": 1
|
||||
}
|
||||
7
tests/fixtures/accounting/rejected/avvist.html
vendored
Normal file
7
tests/fixtures/accounting/rejected/avvist.html
vendored
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
<!doctype html>
|
||||
<html><head><title>Avvist</title></head>
|
||||
<body>
|
||||
<h1>Skjult</h1>
|
||||
<p>Denne teksten bærer et usynlig tegn.</p>
|
||||
<img src="graphics/figur.png" alt="Figur 1">
|
||||
</body></html>
|
||||
BIN
tests/fixtures/accounting/rejected/graphics/figur.png
vendored
Normal file
BIN
tests/fixtures/accounting/rejected/graphics/figur.png
vendored
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 94 B |
216
tests/fixtures/accounting/witness/prosess-84-sts.twin.json
vendored
Normal file
216
tests/fixtures/accounting/witness/prosess-84-sts.twin.json
vendored
Normal file
|
|
@ -0,0 +1,216 @@
|
|||
{
|
||||
"productInfo": {
|
||||
"meta": {
|
||||
"productTitle": "R761 fixture twin"
|
||||
}
|
||||
},
|
||||
"standardContent": {
|
||||
"id": null,
|
||||
"tag": "#document",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "standard",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "front",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "std-ident",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "doc-number",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 2,
|
||||
"t": "R761",
|
||||
"x": null
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "body",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "sec",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "label",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 2,
|
||||
"t": "84",
|
||||
"x": null
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "title",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 2,
|
||||
"t": "Konstruksjoner av betong",
|
||||
"x": null
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "sec",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "label",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 2,
|
||||
"t": "84.1",
|
||||
"x": null
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "title",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 2,
|
||||
"t": "Toleranser",
|
||||
"x": null
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "p",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 2,
|
||||
"t": "Toleranseklasse er gitt i tabell 84-2.",
|
||||
"x": null
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "graphic",
|
||||
"a": {
|
||||
"xlink:href": "graphic_0001.png"
|
||||
},
|
||||
"c": []
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "p",
|
||||
"a": {},
|
||||
"c": [
|
||||
{
|
||||
"e": 2,
|
||||
"t": "Figur 84-1 viser prinsippet.",
|
||||
"x": null
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
{
|
||||
"e": 1,
|
||||
"t": null,
|
||||
"x": {
|
||||
"id": null,
|
||||
"tag": "graphic",
|
||||
"a": {
|
||||
"xlink:href": "graphic_0002.png"
|
||||
},
|
||||
"c": []
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
425
tests/test_accounting_gate.py
Normal file
425
tests/test_accounting_gate.py
Normal file
|
|
@ -0,0 +1,425 @@
|
|||
"""The content-accounting gate's own tests (`tools/okf_accounting_gate.py`).
|
||||
|
||||
The gate is RED today by design: `okf build` has no source inventory and no
|
||||
per-element accounting. These tests are GREEN and prove three things.
|
||||
|
||||
1. The fasit is independent. `tools/okf_witness.py` imports no
|
||||
`llm_ingestion_okf` module -- checked on the live import graph of a process
|
||||
that ran every witness, with a control that shows the check fires -- and
|
||||
the committed inventories are exactly what the witness counts today, pinned
|
||||
again to hand counts on four documents.
|
||||
2. Every row CAN turn green and CAN turn red, each boundary driven from both
|
||||
sides with synthetic build output (the door the capability must open).
|
||||
3. Run against the real `okf build` at this commit, the gate sees the three
|
||||
defects it was ordered for: no inventory, a file carried AND rejected, and
|
||||
a rejected document logged as "0 carried of 0 found".
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
|
||||
TOOLS = Path(__file__).resolve().parents[1] / "tools"
|
||||
sys.path.insert(0, str(TOOLS))
|
||||
|
||||
import okf_accounting_gate as gate # noqa: E402
|
||||
import okf_witness as witness # noqa: E402
|
||||
|
||||
TABLE = [
|
||||
".csv",
|
||||
".docx",
|
||||
".htm",
|
||||
".html",
|
||||
".json",
|
||||
".md",
|
||||
".odt",
|
||||
".pdf",
|
||||
".pptx",
|
||||
".rtf",
|
||||
".txt",
|
||||
".xlsx",
|
||||
".xml",
|
||||
]
|
||||
|
||||
|
||||
# --- 1. the fasit ------------------------------------------------------------
|
||||
|
||||
_IMPORT_PROBE = """
|
||||
import sys
|
||||
sys.path.insert(0, {tools!r})
|
||||
{preload}
|
||||
import okf_witness as w
|
||||
from pathlib import Path
|
||||
fixtures = Path({fixtures!r})
|
||||
w.witness_inbox(fixtures / "corpus")
|
||||
w.witness_inbox(fixtures / "rejected")
|
||||
w.count_sts_json((fixtures / "witness" / "prosess-84-sts.twin.json").read_bytes())
|
||||
w.pdf_poppler(fixtures / "corpus" / "prosess-84-tabell.pdf")
|
||||
print(sorted(m for m in sys.modules if m.split(".")[0] == "llm_ingestion_okf"))
|
||||
"""
|
||||
|
||||
|
||||
def _package_modules_after_witness(preload: str) -> list[str]:
|
||||
script = _IMPORT_PROBE.format(tools=str(TOOLS), fixtures=str(gate.FIXTURES), preload=preload)
|
||||
out = subprocess.run(
|
||||
[sys.executable, "-c", script], capture_output=True, text=True, check=True
|
||||
).stdout
|
||||
result: list[str] = json.loads(out.strip().splitlines()[-1].replace("'", '"'))
|
||||
return result
|
||||
|
||||
|
||||
def test_the_witness_imports_no_module_of_the_package_it_judges() -> None:
|
||||
pytest.importorskip("pdfplumber")
|
||||
assert _package_modules_after_witness("") == []
|
||||
|
||||
|
||||
def test_the_import_check_fires_when_the_package_is_loaded() -> None:
|
||||
pytest.importorskip("pdfplumber")
|
||||
assert "llm_ingestion_okf" in _package_modules_after_witness("import llm_ingestion_okf")
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("inbox", "committed"),
|
||||
[(gate.CORPUS, gate.INVENTORY), (gate.REJECTED, gate.REJECTED_INVENTORY)],
|
||||
)
|
||||
def test_the_committed_fasit_is_what_the_witness_counts(inbox: Path, committed: Path) -> None:
|
||||
pytest.importorskip("pdfplumber")
|
||||
assert witness.witness_inbox(inbox) == gate.load_inventory(committed)
|
||||
|
||||
|
||||
def test_the_fixture_corpus_covers_every_readme_file_type() -> None:
|
||||
inventory = gate.load_inventory(gate.INVENTORY)
|
||||
assert gate.readme_types() == TABLE
|
||||
assert {entry["suffix"] for entry in inventory["documents"].values()} == set(TABLE)
|
||||
|
||||
|
||||
# Counted by hand from the fixture bytes, not by the witness.
|
||||
HAND_COUNTS = {
|
||||
"notat.md": {
|
||||
"code_block": 1,
|
||||
"heading": 2,
|
||||
"image": 1,
|
||||
"paragraph": 2,
|
||||
"table": 1,
|
||||
"table_row": 2,
|
||||
},
|
||||
"side.htm": {"cell": 4, "heading": 2, "image": 0, "list_item": 2, "paragraph": 1, "table": 1},
|
||||
"krav-rikt-tekstformat.rtf": {"cell": 56, "image": 0, "paragraph": 3, "table_row": 24},
|
||||
"prosess-84-sts.xml": {
|
||||
"cell": 0,
|
||||
"footnote": 0,
|
||||
"image": 2,
|
||||
"list_item": 0,
|
||||
"paragraph": 2,
|
||||
"section": 2,
|
||||
"section_label": 2,
|
||||
"table": 0,
|
||||
"table_label": 0,
|
||||
"title": 2,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(HAND_COUNTS))
|
||||
def test_the_witness_matches_a_hand_count(name: str) -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.CORPUS / name)
|
||||
assert inventory.elements == HAND_COUNTS[name]
|
||||
|
||||
|
||||
def test_a_fenced_heading_is_not_a_heading_to_the_witness() -> None:
|
||||
elements, _ = witness.count_markdown("# Real\n\n```bash\n# not one\n```\n")
|
||||
assert elements["heading"] == 1
|
||||
assert elements["code_block"] == 1
|
||||
|
||||
|
||||
def test_the_sts_image_reference_resolves_through_the_graphics_directory() -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.STS_FIXTURE)
|
||||
assert [ref.target for ref in inventory.images] == [
|
||||
"graphics/tabell-84-2.png",
|
||||
"graphics/figur-84-1.png",
|
||||
]
|
||||
|
||||
|
||||
def test_a_reference_above_the_document_resolves_to_nothing(tmp_path: Path) -> None:
|
||||
(tmp_path / "secret.png").write_bytes(b"x")
|
||||
(tmp_path / "docs").mkdir()
|
||||
document = tmp_path / "docs" / "a.html"
|
||||
assert witness.resolve_local(tmp_path, document, "../secret.png") is None
|
||||
|
||||
|
||||
def test_the_witness_refuses_a_doctype() -> None:
|
||||
with pytest.raises(witness.WitnessRefused):
|
||||
witness.count_sts_xml(b'<!DOCTYPE x [<!ENTITY a "b">]><standard/>')
|
||||
|
||||
|
||||
# --- 2. every row can go both ways -------------------------------------------
|
||||
|
||||
|
||||
def _inventory() -> dict[str, Any]:
|
||||
return {
|
||||
"documents": {
|
||||
"a.md": {"suffix": ".md", "elements": {"heading": 2, "image": 1}, "images": []},
|
||||
},
|
||||
"files": {"graphics/x.png": {"pointed_at_by": ["a.md"]}},
|
||||
}
|
||||
|
||||
|
||||
def _build(
|
||||
*,
|
||||
accounting: dict[str, Any] | None = None,
|
||||
sources: set[str] | None = None,
|
||||
assets: set[str] | None = None,
|
||||
log: str = "",
|
||||
exit_code: int = 0,
|
||||
) -> gate.Build:
|
||||
return gate.Build(
|
||||
exit_code=exit_code,
|
||||
log=log,
|
||||
accounting=accounting,
|
||||
source_files={"a.md"} if sources is None else sources,
|
||||
asset_prefixes=assets or set(),
|
||||
)
|
||||
|
||||
|
||||
def _declared(heading: int = 2, image: int = 1, fate: str = "rejected") -> dict[str, Any]:
|
||||
return {
|
||||
"accounting_version": 1,
|
||||
"documents": [
|
||||
{
|
||||
"source_file": "a.md",
|
||||
"status": "persisted",
|
||||
"code": None,
|
||||
"inventory": {"heading": 2, "image": 1},
|
||||
"fates": {
|
||||
"heading": {"carried": heading},
|
||||
"image": {"pointer": image},
|
||||
},
|
||||
}
|
||||
],
|
||||
"files": [{"source_file": "graphics/x.png", "fate": fate, "code": "extractor_unknown"}],
|
||||
}
|
||||
|
||||
|
||||
def _corpus(tmp_path: Path) -> Path:
|
||||
(tmp_path / "graphics").mkdir()
|
||||
(tmp_path / "graphics" / "x.png").write_bytes(b"png bytes")
|
||||
return tmp_path
|
||||
|
||||
|
||||
def test_row1_is_green_when_every_type_has_a_fasit() -> None:
|
||||
inventory = {
|
||||
"documents": {f"f{s}": {"suffix": s, "elements": {}} for s in TABLE},
|
||||
"files": {},
|
||||
}
|
||||
row = gate.row1(TABLE, inventory, inventory)
|
||||
assert (row.k, row.m, row.status) == (13, 13, gate.GREEN)
|
||||
|
||||
|
||||
def test_row1_is_red_when_one_type_lacks_a_fasit() -> None:
|
||||
inventory = {
|
||||
"documents": {f"f{s}": {"suffix": s, "elements": {}} for s in TABLE[:-1]},
|
||||
"files": {},
|
||||
}
|
||||
row = gate.row1(TABLE, inventory, inventory)
|
||||
assert (row.k, row.m, row.status) == (12, 13, gate.RED)
|
||||
|
||||
|
||||
def test_row1_is_red_when_the_committed_fasit_is_stale() -> None:
|
||||
inventory = {"documents": {"f.md": {"suffix": ".md", "elements": {"heading": 1}}}}
|
||||
fresh = {"documents": {"f.md": {"suffix": ".md", "elements": {"heading": 2}}}}
|
||||
row = gate.row1([".md"], inventory, fresh)
|
||||
assert (row.k, row.status) == (0, gate.RED)
|
||||
|
||||
|
||||
def test_row2_is_red_without_the_door() -> None:
|
||||
row = gate.row2([".md"], _inventory(), _build(), door=False)
|
||||
assert (row.k, row.m, row.status) == (0, 1, gate.RED)
|
||||
|
||||
|
||||
def test_row2_is_green_when_the_declared_inventory_equals_the_witness() -> None:
|
||||
row = gate.row2([".md"], _inventory(), _build(accounting=_declared()), door=True)
|
||||
assert (row.k, row.m, row.status) == (1, 1, gate.GREEN)
|
||||
|
||||
|
||||
def test_row2_is_red_when_the_declared_inventory_is_one_off() -> None:
|
||||
declared = _declared()
|
||||
declared["documents"][0]["inventory"]["heading"] = 3
|
||||
row = gate.row2([".md"], _inventory(), _build(accounting=declared), door=True)
|
||||
assert (row.k, row.status) == (0, gate.RED)
|
||||
|
||||
|
||||
def test_row3_is_green_when_every_element_and_file_has_one_fate(tmp_path: Path) -> None:
|
||||
units = gate.account(_inventory(), _build(accounting=_declared()), _corpus(tmp_path))
|
||||
row = gate.row3(units, door=True)
|
||||
assert (row.k, row.m, row.status) == (2, 2, gate.GREEN)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(("heading", "u", "d"), [(1, 1, 0), (3, 0, 1)])
|
||||
def test_row3_is_red_one_element_either_side(tmp_path: Path, heading: int, u: int, d: int) -> None:
|
||||
units = gate.account(
|
||||
_inventory(), _build(accounting=_declared(heading=heading)), _corpus(tmp_path)
|
||||
)
|
||||
document = units[0]
|
||||
assert (document.unaccounted, document.double) == (u, d)
|
||||
assert gate.row3(units, door=True).status == gate.RED
|
||||
|
||||
|
||||
def test_row3_is_red_when_no_fate_is_declared(tmp_path: Path) -> None:
|
||||
units = gate.account(_inventory(), _build(), _corpus(tmp_path))
|
||||
assert units[0].unaccounted == 3
|
||||
assert gate.row3(units, door=False).status == gate.RED
|
||||
|
||||
|
||||
def test_a_file_carried_through_a_document_and_rejected_is_double_booked(tmp_path: Path) -> None:
|
||||
corpus = _corpus(tmp_path)
|
||||
carried = {gate._sha12(corpus / "graphics" / "x.png")}
|
||||
units = gate.account(_inventory(), _build(accounting=_declared(), assets=carried), corpus)
|
||||
assert (units[1].unaccounted, units[1].double) == (0, 1)
|
||||
|
||||
|
||||
def test_a_file_carried_through_a_document_and_declared_carried_is_clean(tmp_path: Path) -> None:
|
||||
corpus = _corpus(tmp_path)
|
||||
carried = {gate._sha12(corpus / "graphics" / "x.png")}
|
||||
build = _build(accounting=_declared(fate="carried"), assets=carried)
|
||||
assert gate.account(_inventory(), build, corpus)[1].clean
|
||||
|
||||
|
||||
def test_a_file_declared_carried_without_its_bytes_is_unaccounted(tmp_path: Path) -> None:
|
||||
build = _build(accounting=_declared(fate="carried"))
|
||||
unit = gate.account(_inventory(), build, _corpus(tmp_path))[1]
|
||||
assert (unit.unaccounted, unit.double) == (1, 0)
|
||||
|
||||
|
||||
def test_an_unpointed_file_sharing_bytes_with_a_carried_one_is_not_carried(
|
||||
tmp_path: Path,
|
||||
) -> None:
|
||||
corpus = _corpus(tmp_path)
|
||||
(corpus / "graphics" / "twin.png").write_bytes(b"png bytes")
|
||||
inventory = _inventory()
|
||||
inventory["files"]["graphics/twin.png"] = {"pointed_at_by": []}
|
||||
carried = {gate._sha12(corpus / "graphics" / "x.png")}
|
||||
units = gate.account(inventory, _build(assets=carried), corpus)
|
||||
assert [(u.name, u.double) for u in units[1:]] == [
|
||||
("graphics/twin.png", 0),
|
||||
("graphics/x.png", 1),
|
||||
]
|
||||
|
||||
|
||||
def test_without_the_door_double_booking_is_derived_from_conservation(tmp_path: Path) -> None:
|
||||
corpus = _corpus(tmp_path)
|
||||
carried = {gate._sha12(corpus / "graphics" / "x.png")}
|
||||
assert gate.account(_inventory(), _build(assets=carried), corpus)[1].double == 1
|
||||
assert gate.account(_inventory(), _build(), corpus)[1].clean
|
||||
|
||||
|
||||
_HONEST_LOG = (
|
||||
"* **Images**: 0 carried of 1 found, written to `assets/`.\n"
|
||||
"* a.md: 3 elements found in the source, 0 carried: document rejected `fail_secure`\n"
|
||||
)
|
||||
|
||||
|
||||
def _rejected_inventory() -> dict[str, Any]:
|
||||
inventory = _inventory()
|
||||
inventory["documents"]["a.md"]["images"] = [{"kind": "local"}]
|
||||
return inventory
|
||||
|
||||
|
||||
def test_row4_is_green_when_the_log_names_what_the_rejected_document_held() -> None:
|
||||
row = gate.row4(_rejected_inventory(), _build(sources=set(), log=_HONEST_LOG))
|
||||
assert (row.k, row.m, row.status) == (1, 1, gate.GREEN)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"log",
|
||||
[
|
||||
_HONEST_LOG.replace("0 carried of 1 found", "0 carried of 0 found"),
|
||||
_HONEST_LOG.replace("3 elements", "2 elements"),
|
||||
_HONEST_LOG.replace(" `fail_secure`", ""),
|
||||
],
|
||||
)
|
||||
def test_row4_is_red_when_the_log_understates_the_rejected_document(log: str) -> None:
|
||||
row = gate.row4(_rejected_inventory(), _build(sources=set(), log=log))
|
||||
assert (row.k, row.status) == (0, gate.RED)
|
||||
|
||||
|
||||
def test_row4_cannot_be_green_when_nothing_was_rejected() -> None:
|
||||
row = gate.row4(_rejected_inventory(), _build(log=_HONEST_LOG))
|
||||
assert (row.m, row.status) == (0, gate.RED)
|
||||
|
||||
|
||||
def test_row5_is_green_when_the_witnesses_agree() -> None:
|
||||
row = gate.row5([("pair", gate.compare({"p": 3}, {"p": 3}))], [])
|
||||
assert (row.k, row.m, row.status) == (1, 1, gate.GREEN)
|
||||
|
||||
|
||||
def test_row5_is_red_with_both_numbers_when_they_disagree_by_one() -> None:
|
||||
row = gate.row5([("pair", gate.compare({"p": 3}, {"p": 4}))], [])
|
||||
assert row.status == gate.RED
|
||||
assert row.details == ["pair: p: 3 vs 4"]
|
||||
|
||||
|
||||
def test_row5_is_red_when_a_witness_is_missing() -> None:
|
||||
assert gate.row5([("pair", gate.compare({"p": 3}, None))], []).status == gate.RED
|
||||
|
||||
|
||||
def test_row6_without_its_source_is_red_locally_and_skipped_in_ci(tmp_path: Path) -> None:
|
||||
missing = tmp_path / "absent"
|
||||
local = gate.row6(missing, ci=False)
|
||||
ci = gate.row6(missing, ci=True)
|
||||
assert (local.status, local.fails) == (gate.RED, True)
|
||||
assert (ci.status, ci.fails) == (gate.SKIPPED, False)
|
||||
assert "source missing" in ci.reason
|
||||
|
||||
|
||||
def test_the_proposed_exceptions_are_not_applied() -> None:
|
||||
assert gate.APPROVED_EXCEPTIONS == frozenset()
|
||||
assert "NOT APPROVED" in gate.render([])
|
||||
|
||||
|
||||
def test_bad_usage_exits_two() -> None:
|
||||
with pytest.raises(SystemExit) as exc:
|
||||
gate.main(["--no-such-flag"])
|
||||
assert exc.value.code == 2
|
||||
|
||||
|
||||
# --- 3. the real build at this commit ----------------------------------------
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def real_rows() -> list[gate.Row]:
|
||||
pytest.importorskip("pdfplumber")
|
||||
pytest.importorskip("pypandoc")
|
||||
if gate.door_available():
|
||||
pytest.skip("the accounting door exists; these rows describe the build before it")
|
||||
return gate.evaluate(r761=None, ci=True, consume=False)
|
||||
|
||||
|
||||
def test_the_real_gate_is_red_on_rows_two_three_and_four(real_rows: list[gate.Row]) -> None:
|
||||
assert [r.number for r in real_rows if r.fails] == [2, 3, 4]
|
||||
assert [(r.number, r.status) for r in real_rows if not r.fails] == [
|
||||
(1, gate.GREEN),
|
||||
(5, gate.GREEN),
|
||||
(6, gate.SKIPPED),
|
||||
]
|
||||
|
||||
|
||||
def test_the_real_gate_sees_the_graphics_carried_and_rejected(real_rows: list[gate.Row]) -> None:
|
||||
row3 = real_rows[2]
|
||||
assert "d = 2 double-booked" in row3.reason
|
||||
doubled = [d for d in row3.details if "AND rejected" in d]
|
||||
assert [d.split(":")[0] for d in doubled] == [
|
||||
"file graphics/figur-84-1.png",
|
||||
"file graphics/tabell-84-2.png",
|
||||
]
|
||||
|
||||
|
||||
def test_the_real_gate_sees_the_rejected_document_logged_as_empty(
|
||||
real_rows: list[gate.Row],
|
||||
) -> None:
|
||||
assert "log says 0 carried of 0 found; the source declares 1" in "\n".join(real_rows[3].details)
|
||||
680
tools/okf_accounting_gate.py
Normal file
680
tools/okf_accounting_gate.py
Normal file
|
|
@ -0,0 +1,680 @@
|
|||
"""The content-accounting gate for `okf build` (capability loop, step 3).
|
||||
|
||||
One command, one exit code. For every supported file type it asks whether the
|
||||
bundle accounts for what the SOURCE holds: of M elements the source carries,
|
||||
how many does the build account for as carried, as a pointer or as a coded
|
||||
rejection -- and how many does it account for NOT AT ALL (u) or TWICE (d).
|
||||
|
||||
Written RED, before any capability. `okf build`'s conservation identity,
|
||||
`merged + coded rejections == N`, counts FILES: a file can be "merged" while
|
||||
content inside it is gone, and a file can be "rejected" while its bytes ride
|
||||
into the bundle through a document that points at it. Neither is visible to
|
||||
the identity, and both are measured here.
|
||||
|
||||
THE FASIT NEVER COMES FROM THE READER IT JUDGES. Element counts come from
|
||||
`tools/okf_witness.py`, which imports no `llm_ingestion_okf` module (a test
|
||||
proves it on the live import graph) and is committed as data in
|
||||
`tests/fixtures/accounting/*inventory.json`. This module imports the package
|
||||
only to RUN the build it judges.
|
||||
|
||||
THE DOOR THE CAPABILITY MUST OPEN (the contract this gate reads). `okf build`
|
||||
accepts `--accounting PATH` and writes one JSON object there:
|
||||
|
||||
{"accounting_version": 1,
|
||||
"documents": [
|
||||
{"source_file": "<inbox-relative path>",
|
||||
"status": "persisted" | "rejected", "code": "<rejection code>" | null,
|
||||
"inventory": {"<element>": <count>, ...},
|
||||
"fates": {"<element>": {"carried": n, "pointer": n,
|
||||
"rejected": {"<code>": n}}}}],
|
||||
"files": [
|
||||
{"source_file": "<inbox-relative path>",
|
||||
"fate": "carried" | "merged" | "rejected", "code": "<code>" | null}]}
|
||||
|
||||
`inventory` is taken BEFORE extraction and before the persist gate, in the
|
||||
witness's element vocabulary (per file type, defined in `okf_witness.py`), so
|
||||
a document the gate refuses still has one. `files` covers every inbox file
|
||||
that is not a document the build reads; `fate` is exactly one value, so a file
|
||||
whose bytes were carried through a document is `carried` and never also
|
||||
`rejected`. Until the flag exists, rows 2 and 3 say so and stay red.
|
||||
|
||||
A rejected document is reported in `log.md` as ONE line, and row 4 reads it:
|
||||
|
||||
<source_file>: <M> elements found in the source, 0 carried: document rejected `<code>`
|
||||
|
||||
with M the document's inventory total. The `Images` bullet's "found" count is
|
||||
the SOURCE's (every image the documents declare), never what a reader got to.
|
||||
|
||||
EXCEPTIONS to 100 % are listed in the output and are NOT APPROVED: none of
|
||||
them lowers a denominator until the operator approves it by name.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import contextlib
|
||||
import hashlib
|
||||
import io
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import tempfile
|
||||
import warnings
|
||||
import zipfile
|
||||
from collections.abc import Iterable, Mapping
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
TOOLS = Path(__file__).resolve().parent
|
||||
REPO = TOOLS.parent
|
||||
if str(TOOLS) not in sys.path:
|
||||
sys.path.insert(0, str(TOOLS))
|
||||
|
||||
import okf_witness as witness # noqa: E402
|
||||
|
||||
FIXTURES = REPO / "tests" / "fixtures" / "accounting"
|
||||
CORPUS = FIXTURES / "corpus"
|
||||
REJECTED = FIXTURES / "rejected"
|
||||
INVENTORY = FIXTURES / "inventory.json"
|
||||
REJECTED_INVENTORY = FIXTURES / "rejected-inventory.json"
|
||||
STS_FIXTURE = CORPUS / "prosess-84-sts.xml"
|
||||
STS_TWIN = FIXTURES / "witness" / "prosess-84-sts.twin.json"
|
||||
PDF_FIXTURE = CORPUS / "prosess-84-tabell.pdf"
|
||||
README = REPO / "README.md"
|
||||
|
||||
R761_DEFAULT = Path.home() / "repos" / "vegnormal-okf" / "data" / "raw" / "860019"
|
||||
R761_ZIP = "14ce59dc-2150-480b-b661-6ea605fe3b24.zip"
|
||||
R761_JSON = "R761-2025-860019.json"
|
||||
R761_PDF = "R761-prosesskoden-2025.pdf"
|
||||
|
||||
ACCOUNTING_FLAG = "--accounting"
|
||||
ACCOUNTING_VERSION = 1
|
||||
BUNDLE_ID = "accounting-gate"
|
||||
OKF_VERSION = "0.2"
|
||||
CONSUME_QUESTION = "Hvilken toleranseklasse gjelder for konstruksjoner av betong?"
|
||||
|
||||
GREEN = "GREEN"
|
||||
RED = "RED"
|
||||
SKIPPED = "SKIPPED"
|
||||
DIAGNOSTIC = "DIAGNOSTIC"
|
||||
|
||||
#: Exceptions the operator has approved, by (suffix, element). Empty: the
|
||||
#: first time an exception arises it is the operator's to approve.
|
||||
APPROVED_EXCEPTIONS: frozenset[tuple[str, str]] = frozenset()
|
||||
|
||||
#: Exceptions this gate PROPOSES. Listed in every run; none of them is applied.
|
||||
PROPOSED_EXCEPTIONS: tuple[dict[str, str], ...] = (
|
||||
{
|
||||
"suffix": ".pdf",
|
||||
"element": "heading, paragraph, table",
|
||||
"reason": "a PDF without a structure tree declares none of them, so no "
|
||||
"witness can count them; the witness counts pages and image placements",
|
||||
"carried_instead": "the page text, with headings recovered by rule",
|
||||
},
|
||||
{
|
||||
"suffix": ".xlsx",
|
||||
"element": "image",
|
||||
"reason": "the converter writes one pipe table per sheet and a pointer "
|
||||
"block inside it would break source_rows (README, 0.10.0)",
|
||||
"carried_instead": "nothing; the image is absent from the bundle",
|
||||
},
|
||||
{
|
||||
"suffix": ".md .txt .csv .json .odt .rtf",
|
||||
"element": "image",
|
||||
"reason": "no reader for these types carries image bytes (README, 0.10.0)",
|
||||
"carried_instead": "the reference text as written, if the format has one",
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Row:
|
||||
number: int
|
||||
name: str
|
||||
k: int
|
||||
m: int
|
||||
status: str
|
||||
reason: str
|
||||
details: list[str] = field(default_factory=list)
|
||||
|
||||
@property
|
||||
def fails(self) -> bool:
|
||||
return self.status == RED and self.number <= 6
|
||||
|
||||
def to_json(self) -> dict[str, Any]:
|
||||
return {
|
||||
"row": self.number,
|
||||
"name": self.name,
|
||||
"k": self.k,
|
||||
"m": self.m,
|
||||
"status": self.status,
|
||||
"reason": self.reason,
|
||||
"details": self.details,
|
||||
}
|
||||
|
||||
|
||||
def _row(number: int, name: str, k: int, m: int, reason: str, details: list[str]) -> Row:
|
||||
return Row(number, name, k, m, GREEN if m > 0 and k == m else RED, reason, details)
|
||||
|
||||
|
||||
# --- inputs ------------------------------------------------------------------
|
||||
|
||||
|
||||
def readme_types(readme: Path = README) -> list[str]:
|
||||
"""T: the rows of README's supported-file-types table."""
|
||||
text = readme.read_text(encoding="utf-8")
|
||||
section = text.split("## Supported file types", 1)[1].split("\n## ", 1)[0]
|
||||
return sorted(set(re.findall(r"^\| `(\.[a-z0-9]+)` \|", section, re.MULTILINE)))
|
||||
|
||||
|
||||
def load_inventory(path: Path) -> dict[str, Any]:
|
||||
data: dict[str, Any] = json.loads(path.read_text(encoding="utf-8"))
|
||||
return data
|
||||
|
||||
|
||||
def door_available() -> bool:
|
||||
"""Does `okf build` accept the accounting flag?"""
|
||||
from llm_ingestion_okf import cli
|
||||
|
||||
argv = ["build", "in", "--bundle", "out", "--bundle-id", "x", "--okf-version", "0.2"]
|
||||
with contextlib.redirect_stderr(io.StringIO()):
|
||||
try:
|
||||
cli.parse_args([*argv, ACCOUNTING_FLAG, "accounting.json"])
|
||||
except SystemExit:
|
||||
return False
|
||||
return True
|
||||
|
||||
|
||||
@dataclass
|
||||
class Build:
|
||||
"""What one `okf build` run left behind, read back from the artifacts."""
|
||||
|
||||
exit_code: int
|
||||
log: str
|
||||
accounting: dict[str, Any] | None
|
||||
source_files: set[str]
|
||||
asset_prefixes: set[str]
|
||||
|
||||
|
||||
def _frontmatter_source_file(text: str) -> str | None:
|
||||
if not text.startswith("---\n"):
|
||||
return None
|
||||
head = text[4:].split("\n---\n", 1)[0]
|
||||
match = re.search(r"^source_file:\s*(.+?)\s*$", head, re.MULTILINE)
|
||||
if match is None:
|
||||
return None
|
||||
value = match.group(1)
|
||||
if len(value) >= 2 and value[0] == value[-1] and value[0] in "\"'":
|
||||
value = value[1:-1]
|
||||
return value
|
||||
|
||||
|
||||
def read_bundle(bundle: Path, exit_code: int, accounting_path: Path | None) -> Build:
|
||||
sources: set[str] = set()
|
||||
for path in bundle.rglob("*.md"):
|
||||
if "assets" in path.relative_to(bundle).parts:
|
||||
continue
|
||||
found = _frontmatter_source_file(path.read_text(encoding="utf-8"))
|
||||
if found:
|
||||
sources.add(found)
|
||||
assets = bundle / "assets"
|
||||
prefixes = {p.name[:12] for p in assets.iterdir()} if assets.is_dir() else set()
|
||||
log_path = bundle / "log.md"
|
||||
accounting = None
|
||||
if accounting_path is not None and accounting_path.is_file():
|
||||
accounting = json.loads(accounting_path.read_text(encoding="utf-8"))
|
||||
return Build(
|
||||
exit_code=exit_code,
|
||||
log=log_path.read_text(encoding="utf-8") if log_path.is_file() else "",
|
||||
accounting=accounting,
|
||||
source_files=sources,
|
||||
asset_prefixes=prefixes,
|
||||
)
|
||||
|
||||
|
||||
def run_build(corpus: Path, workdir: Path, *, door: bool, gate: str | None = None) -> Build:
|
||||
"""Run the real `okf build` in-process and read back what it wrote."""
|
||||
from llm_ingestion_okf import cli
|
||||
|
||||
bundle = workdir / "bundle"
|
||||
accounting_path = workdir / "accounting.json" if door else None
|
||||
argv = [
|
||||
"build",
|
||||
str(corpus),
|
||||
"--bundle",
|
||||
str(bundle),
|
||||
"--bundle-id",
|
||||
BUNDLE_ID,
|
||||
"--okf-version",
|
||||
OKF_VERSION,
|
||||
]
|
||||
if gate is not None:
|
||||
argv += ["--gate", gate]
|
||||
if accounting_path is not None:
|
||||
argv += [ACCOUNTING_FLAG, str(accounting_path)]
|
||||
sink = io.StringIO()
|
||||
with (
|
||||
contextlib.redirect_stdout(sink),
|
||||
contextlib.redirect_stderr(sink),
|
||||
warnings.catch_warnings(),
|
||||
):
|
||||
warnings.simplefilter("ignore")
|
||||
try:
|
||||
code = cli.main(argv)
|
||||
except SystemExit as exc:
|
||||
code = exc.code if isinstance(exc.code, int) else 2
|
||||
return read_bundle(bundle, code, accounting_path)
|
||||
|
||||
|
||||
def _sha12(path: Path) -> str:
|
||||
return hashlib.sha256(path.read_bytes()).hexdigest()[:12]
|
||||
|
||||
|
||||
# --- accounting --------------------------------------------------------------
|
||||
|
||||
|
||||
@dataclass
|
||||
class Unit:
|
||||
"""One inventoried thing: a document, or an inbox file that is not one."""
|
||||
|
||||
name: str
|
||||
kind: str
|
||||
unaccounted: int
|
||||
double: int
|
||||
notes: list[str] = field(default_factory=list)
|
||||
|
||||
@property
|
||||
def clean(self) -> bool:
|
||||
return self.unaccounted == 0 and self.double == 0
|
||||
|
||||
|
||||
def _conservation_held(build: Build) -> bool:
|
||||
return build.exit_code == 0 and "K1b FAILED" not in build.log
|
||||
|
||||
|
||||
def account(inventory: Mapping[str, Any], build: Build, corpus: Path) -> list[Unit]:
|
||||
"""Give every inventoried element and file its fate, or say it has none."""
|
||||
declared_docs = {}
|
||||
declared_files = {}
|
||||
if build.accounting is not None:
|
||||
declared_docs = {d["source_file"]: d for d in build.accounting.get("documents", [])}
|
||||
declared_files = {f["source_file"]: f for f in build.accounting.get("files", [])}
|
||||
units: list[Unit] = []
|
||||
for name, entry in sorted(inventory["documents"].items()):
|
||||
elements: dict[str, int] = entry["elements"]
|
||||
declared = declared_docs.get(name)
|
||||
if declared is None:
|
||||
units.append(Unit(name, "document", sum(elements.values()), 0, ["no declared fates"]))
|
||||
continue
|
||||
unaccounted = double = 0
|
||||
notes: list[str] = []
|
||||
fates: dict[str, Any] = declared.get("fates", {})
|
||||
for element in sorted(set(elements) | set(fates)):
|
||||
fate = fates.get(element, {})
|
||||
booked = (
|
||||
int(fate.get("carried", 0))
|
||||
+ int(fate.get("pointer", 0))
|
||||
+ sum(int(v) for v in fate.get("rejected", {}).values())
|
||||
)
|
||||
have = elements.get(element, 0)
|
||||
if booked < have:
|
||||
unaccounted += have - booked
|
||||
notes.append(f"{element}: {booked} booked of {have}")
|
||||
elif booked > have:
|
||||
double += booked - have
|
||||
notes.append(f"{element}: {booked} booked, source has {have}")
|
||||
units.append(Unit(name, "document", unaccounted, double, notes))
|
||||
for name, entry in sorted(inventory["files"].items()):
|
||||
pointed_by = entry["pointed_at_by"]
|
||||
# Bytes in assets/ prove a carry only for a file a document points at:
|
||||
# an unpointed file with the same bytes (R761 ships 8 such duplicates)
|
||||
# was not carried through anything.
|
||||
carried = bool(pointed_by) and _sha12(corpus / name) in build.asset_prefixes
|
||||
merged = name in build.source_files
|
||||
declared = declared_files.get(name)
|
||||
notes = []
|
||||
false_claim = False
|
||||
if declared is not None:
|
||||
rejected = declared.get("fate") == "rejected"
|
||||
false_claim = declared.get("fate") == "carried" and not carried
|
||||
if false_claim:
|
||||
notes.append("declared carried, bytes absent from assets/")
|
||||
else:
|
||||
# K1b: every walked file is merged or a coded rejection, so a file
|
||||
# that is not merged was booked as a rejection.
|
||||
rejected = not merged and _conservation_held(build)
|
||||
fates = sum((carried, merged, rejected))
|
||||
if carried and rejected:
|
||||
notes.append(f"carried via {', '.join(pointed_by) or 'a document'} AND rejected")
|
||||
unaccounted = 1 if fates == 0 or false_claim else 0
|
||||
units.append(Unit(name, "file", unaccounted, max(0, fates - 1), notes))
|
||||
return units
|
||||
|
||||
|
||||
# --- rows --------------------------------------------------------------------
|
||||
|
||||
|
||||
def row1(table: Iterable[str], inventory: Mapping[str, Any], fresh: Mapping[str, Any]) -> Row:
|
||||
table = sorted(table)
|
||||
covered: set[str] = set()
|
||||
stale: list[str] = []
|
||||
for name, entry in inventory["documents"].items():
|
||||
if fresh["documents"].get(name) == entry:
|
||||
covered.add(entry["suffix"])
|
||||
else:
|
||||
stale.append(name)
|
||||
missing = [t for t in table if t not in covered]
|
||||
k = len(table) - len(missing)
|
||||
reason = "every README type has a fixture with a reproducible witness count"
|
||||
if missing:
|
||||
reason = f"no fixture fasit for {', '.join(missing)}"
|
||||
details = [f"committed fasit differs from a fresh witness count: {n}" for n in stale]
|
||||
return _row(1, "file types with a fasit fixture", k, len(table), reason, details)
|
||||
|
||||
|
||||
def row2(table: Iterable[str], inventory: Mapping[str, Any], build: Build, door: bool) -> Row:
|
||||
table = sorted(table)
|
||||
name = "source inventory before build"
|
||||
if not door:
|
||||
return _row(
|
||||
2, name, 0, len(table), f"`okf build` has no `{ACCOUNTING_FLAG}` door; no inventory", []
|
||||
)
|
||||
declared = {d["source_file"]: d for d in (build.accounting or {}).get("documents", [])}
|
||||
good: list[str] = []
|
||||
details: list[str] = []
|
||||
for suffix in table:
|
||||
docs = [n for n, e in inventory["documents"].items() if e["suffix"] == suffix]
|
||||
ok = bool(docs)
|
||||
for doc in docs:
|
||||
got = declared.get(doc, {}).get("inventory")
|
||||
want = inventory["documents"][doc]["elements"]
|
||||
if got != want:
|
||||
ok = False
|
||||
details.append(f"{doc}: declared {got}, witness {want}")
|
||||
if ok:
|
||||
good.append(suffix)
|
||||
missing = [t for t in table if t not in good]
|
||||
reason = (
|
||||
"every type's inventory equals the witness"
|
||||
if not missing
|
||||
else (f"inventory absent or wrong for {', '.join(missing)}")
|
||||
)
|
||||
return _row(2, name, len(good), len(table), reason, details)
|
||||
|
||||
|
||||
def row3(units: list[Unit], door: bool) -> Row:
|
||||
clean = sum(1 for u in units if u.clean)
|
||||
u_total = sum(u.unaccounted for u in units)
|
||||
d_total = sum(u.double for u in units)
|
||||
reason = f"u = {u_total} unaccounted, d = {d_total} double-booked"
|
||||
if not door:
|
||||
reason += f"; no `{ACCOUNTING_FLAG}` door, so no element has a declared fate"
|
||||
details = [
|
||||
f"{u.kind} {u.name}: u={u.unaccounted} d={u.double}"
|
||||
+ (f" ({'; '.join(u.notes)})" if u.notes else "")
|
||||
for u in units
|
||||
if not u.clean
|
||||
]
|
||||
return _row(3, "accounting after build (u = 0 and d = 0)", clean, len(units), reason, details)
|
||||
|
||||
|
||||
def row4(inventory: Mapping[str, Any], build: Build) -> Row:
|
||||
name = "a rejected document is reported honestly"
|
||||
rejected = [n for n in sorted(inventory["documents"]) if n not in build.source_files]
|
||||
if not rejected:
|
||||
return _row(4, name, 0, 0, "the fixture was not rejected; the row cannot judge", [])
|
||||
declared_images = sum(len(e["images"]) for e in inventory["documents"].values())
|
||||
found = re.search(r"\*\*Images\*\*: (\d+) carried of (\d+) found", build.log)
|
||||
good = 0
|
||||
details: list[str] = []
|
||||
for doc in rejected:
|
||||
total = sum(inventory["documents"][doc]["elements"].values())
|
||||
line = re.compile(
|
||||
rf"{re.escape(doc)}: {total} elements found in the source, 0 carried: "
|
||||
r"document rejected `[a-z_]+`"
|
||||
)
|
||||
problems = []
|
||||
if not line.search(build.log):
|
||||
problems.append(f"no line '{doc}: {total} elements found in the source, 0 carried'")
|
||||
if found is None:
|
||||
problems.append("no Images bullet")
|
||||
elif int(found.group(2)) != declared_images:
|
||||
problems.append(
|
||||
f"log says {found.group(1)} carried of {found.group(2)} found; "
|
||||
f"the source declares {declared_images}"
|
||||
)
|
||||
if problems:
|
||||
details.extend(f"{doc}: {p}" for p in problems)
|
||||
else:
|
||||
good += 1
|
||||
reason = (
|
||||
"every rejected document names what it held"
|
||||
if good == len(rejected)
|
||||
else (f"{len(rejected) - good} rejected document(s) reported as if they held less")
|
||||
)
|
||||
return _row(4, name, good, len(rejected), reason, details)
|
||||
|
||||
|
||||
def compare(left: Mapping[str, int] | None, right: Mapping[str, int] | None) -> list[str]:
|
||||
"""Disagreements between two witnesses, both numbers kept."""
|
||||
if left is None or right is None:
|
||||
return ["a witness is unavailable"]
|
||||
return [
|
||||
f"{e}: {left[e]} vs {right[e]}"
|
||||
for e in sorted(set(left) & set(right))
|
||||
if left[e] != right[e]
|
||||
]
|
||||
|
||||
|
||||
def row5(pairs: list[tuple[str, list[str]]], notes: list[str]) -> Row:
|
||||
good = sum(1 for _, problems in pairs if not problems)
|
||||
details = [f"{label}: {'; '.join(p)}" for label, p in pairs if p] + notes
|
||||
reason = (
|
||||
"both witnesses agree on every element"
|
||||
if good == len(pairs)
|
||||
else (f"{len(pairs) - good} pair(s) disagree")
|
||||
)
|
||||
return _row(5, "two witnesses agree", good, len(pairs), reason, details)
|
||||
|
||||
|
||||
def witness_pairs(r761: Path | None) -> tuple[list[tuple[str, list[str]]], list[str]]:
|
||||
sts_xml, _, _ = witness.count_sts_xml(STS_FIXTURE.read_bytes())
|
||||
pairs = [
|
||||
(
|
||||
"sts fixture (xml | json)",
|
||||
compare(sts_xml, witness.count_sts_json(STS_TWIN.read_bytes())),
|
||||
),
|
||||
(
|
||||
"pdf fixture (pdfplumber | poppler)",
|
||||
compare(witness.pdf_objects(PDF_FIXTURE), witness.pdf_poppler(PDF_FIXTURE)),
|
||||
),
|
||||
]
|
||||
notes: list[str] = []
|
||||
if r761 is None or not r761.is_dir():
|
||||
notes.append("R761 pairs not measured: source missing")
|
||||
return pairs, notes
|
||||
with zipfile.ZipFile(r761 / R761_ZIP) as archive:
|
||||
xml_name = next(n for n in archive.namelist() if n.endswith(".xml"))
|
||||
r761_xml, _, _ = witness.count_sts_xml(archive.read(xml_name))
|
||||
pairs.append(
|
||||
(
|
||||
"R761 sts (xml | json)",
|
||||
compare(r761_xml, witness.count_sts_json((r761 / R761_JSON).read_bytes())),
|
||||
)
|
||||
)
|
||||
pdf = r761 / R761_PDF
|
||||
pairs.append(
|
||||
(
|
||||
"R761 pdf (pdfplumber | poppler)",
|
||||
compare(witness.pdf_objects(pdf), witness.pdf_poppler(pdf)),
|
||||
)
|
||||
)
|
||||
return pairs, notes
|
||||
|
||||
|
||||
#: The two builds row 6 runs: the default gate is what a user gets; `none`
|
||||
#: persists the document, which is the only way its pictures are carried and
|
||||
#: the double booking of the files beside it becomes visible.
|
||||
R761_GATES: tuple[str | None, ...] = (None, "none")
|
||||
|
||||
|
||||
def row6(r761: Path | None, ci: bool) -> Row:
|
||||
"""R761 through two builds; a unit is clean only if it is clean in both."""
|
||||
name = "real corpus: R761 Prosesskoden:2025"
|
||||
if r761 is None or not r761.is_dir():
|
||||
status = SKIPPED if ci else RED
|
||||
return Row(6, name, 0, 0, status, f"not measured, source missing: {r761}")
|
||||
door = door_available()
|
||||
runs: list[tuple[str, Build, list[Unit]]] = []
|
||||
with zipfile.ZipFile(r761 / R761_ZIP) as archive, tempfile.TemporaryDirectory() as tmp:
|
||||
inbox = Path(tmp) / "inbox"
|
||||
archive.extractall(inbox)
|
||||
inventory = witness.witness_inbox(inbox)
|
||||
json_counts = witness.count_sts_json((r761 / R761_JSON).read_bytes())
|
||||
for index, gate_name in enumerate(R761_GATES):
|
||||
build = run_build(inbox, Path(tmp) / f"work{index}", door=door, gate=gate_name)
|
||||
runs.append((gate_name or "default", build, account(inventory, build, inbox)))
|
||||
files = inventory["files"]
|
||||
pointed = sum(1 for f in files.values() if f["pointed_at_by"])
|
||||
documents = inventory["documents"]
|
||||
details = [
|
||||
f"witness (json): {json.dumps(json_counts, sort_keys=True)}",
|
||||
f"zip: {len(documents)} document(s), {len(files)} other files "
|
||||
f"({pointed} pointed at, {len(files) - pointed} not)",
|
||||
]
|
||||
reasons = []
|
||||
for label, build, units in runs:
|
||||
persisted = sum(1 for d in documents if d in build.source_files)
|
||||
u_total = sum(u.unaccounted for u in units)
|
||||
d_total = sum(u.double for u in units)
|
||||
reasons.append(f"gate {label}: u = {u_total}, d = {d_total}")
|
||||
details.append(
|
||||
f"gate {label}: exit {build.exit_code}, {persisted} of {len(documents)} "
|
||||
f"document(s) persisted, {len(build.asset_prefixes)} asset file(s)"
|
||||
)
|
||||
for unit in units:
|
||||
if unit.kind == "document":
|
||||
details.append(
|
||||
f" {unit.name}: u={unit.unaccounted} d={unit.double} ({'; '.join(unit.notes)})"
|
||||
)
|
||||
doubled = [u for u in units if u.kind == "file" and u.double]
|
||||
if doubled:
|
||||
details.append(
|
||||
f" {len(doubled)} file(s) carried through the document AND rejected, "
|
||||
f"e.g. {doubled[0].name}"
|
||||
)
|
||||
clean = sum(1 for parts in zip(*(units for _, _, units in runs)) if all(u.clean for u in parts))
|
||||
total = len(runs[0][2])
|
||||
return _row(6, name, clean, total, f"{'; '.join(reasons)} over {total} units", details)
|
||||
|
||||
|
||||
def row7(workdir: Path) -> Row:
|
||||
"""Diagnostic only: a question through `okf consume` on the fixture bundle."""
|
||||
from llm_ingestion_okf import cli
|
||||
|
||||
bundle = workdir / "bundle"
|
||||
out = workdir / "payload.json"
|
||||
sink = io.StringIO()
|
||||
with contextlib.redirect_stdout(sink), contextlib.redirect_stderr(sink):
|
||||
try:
|
||||
code = cli.main(
|
||||
["consume", str(bundle), "--question", CONSUME_QUESTION, "--out", str(out)]
|
||||
)
|
||||
except SystemExit as exc:
|
||||
code = exc.code if isinstance(exc.code, int) else 2
|
||||
excerpts = 0
|
||||
if out.is_file():
|
||||
excerpts = len(json.loads(out.read_text(encoding="utf-8")).get("excerpts", []))
|
||||
return Row(
|
||||
7,
|
||||
"question set via okf consume (diagnostic)",
|
||||
excerpts,
|
||||
0,
|
||||
DIAGNOSTIC,
|
||||
f"exit {code}, {excerpts} excerpt(s) for {CONSUME_QUESTION!r}",
|
||||
)
|
||||
|
||||
|
||||
# --- the run -----------------------------------------------------------------
|
||||
|
||||
|
||||
def evaluate(*, r761: Path | None, ci: bool, consume: bool) -> list[Row]:
|
||||
table = readme_types()
|
||||
inventory = load_inventory(INVENTORY)
|
||||
rejected_inventory = load_inventory(REJECTED_INVENTORY)
|
||||
fresh = witness.witness_inbox(CORPUS)
|
||||
door = door_available()
|
||||
rows = [row1(table, inventory, fresh)]
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
work = Path(tmp) / "corpus"
|
||||
build = run_build(CORPUS, work, door=door)
|
||||
rows.append(row2(table, inventory, build, door))
|
||||
rows.append(row3(account(inventory, build, CORPUS), door))
|
||||
rejected_build = run_build(REJECTED, Path(tmp) / "rejected", door=door)
|
||||
rows.append(row4(rejected_inventory, rejected_build))
|
||||
rows.append(row5(*witness_pairs(r761)))
|
||||
rows.append(row6(r761, ci))
|
||||
if consume:
|
||||
rows.append(row7(work))
|
||||
return rows
|
||||
|
||||
|
||||
def render(rows: list[Row]) -> str:
|
||||
lines = ["row | k of M | status | reason"]
|
||||
for row in rows:
|
||||
count = "n/a" if row.status == DIAGNOSTIC else f"{row.k} of {row.m}"
|
||||
lines.append(f"{row.number} {row.name} | {count} | {row.status} | {row.reason}")
|
||||
lines.extend(f" - {d}" for d in row.details)
|
||||
lines += ["", "exceptions (PROPOSED, NOT APPROVED -- none lowers a denominator):"]
|
||||
for item in PROPOSED_EXCEPTIONS:
|
||||
lines.append(
|
||||
f" - {item['suffix']} {item['element']}: {item['reason']}; "
|
||||
f"carried instead: {item['carried_instead']}"
|
||||
)
|
||||
lines.append(f"approved exceptions: {len(APPROVED_EXCEPTIONS)}")
|
||||
failing = [str(r.number) for r in rows if r.fails]
|
||||
lines.append("")
|
||||
lines.append(
|
||||
f"GATE {'RED' if failing else 'GREEN'}"
|
||||
+ (f": rows {', '.join(failing)}" if failing else "")
|
||||
)
|
||||
return "\n".join(lines) + "\n"
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0])
|
||||
parser.add_argument("--json", action="store_true", help="emit the rows as JSON")
|
||||
parser.add_argument(
|
||||
"--r761",
|
||||
type=Path,
|
||||
default=R761_DEFAULT,
|
||||
help="directory holding the R761 zip, JSON and PDF (read only)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--consume", action="store_true", help="also run row 7 (diagnostic, never fails)"
|
||||
)
|
||||
args = parser.parse_args(argv)
|
||||
try:
|
||||
rows = evaluate(r761=args.r761, ci=bool(os.environ.get("CI")), consume=args.consume)
|
||||
except (OSError, ValueError, KeyError) as exc:
|
||||
print(f"okf-accounting-gate: did not run: {exc}", file=sys.stderr)
|
||||
return 2
|
||||
if args.json:
|
||||
payload = {
|
||||
"rows": [r.to_json() for r in rows],
|
||||
"exceptions": {
|
||||
"approved": sorted(APPROVED_EXCEPTIONS),
|
||||
"proposed": PROPOSED_EXCEPTIONS,
|
||||
},
|
||||
"gate": RED if any(r.fails for r in rows) else GREEN,
|
||||
}
|
||||
print(json.dumps(payload, indent=2, ensure_ascii=False))
|
||||
else:
|
||||
print(render(rows), end="")
|
||||
return 1 if any(r.fails for r in rows) else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
697
tools/okf_witness.py
Normal file
697
tools/okf_witness.py
Normal file
|
|
@ -0,0 +1,697 @@
|
|||
"""The independent witness: what a SOURCE file holds, counted by the format's rules.
|
||||
|
||||
This module is the fasit side of the content-accounting gate
|
||||
(`tools/okf_accounting_gate.py`). It answers one question per file -- "how many
|
||||
of each element does this source carry?" -- and it answers it WITHOUT this
|
||||
package: no `llm_ingestion_okf` module is imported, directly or through a
|
||||
helper, and `tests/test_accounting_gate.py` proves that on the live import
|
||||
graph rather than by searching the text. A fasit computed by the reader it is
|
||||
meant to judge is the reader agreeing with itself.
|
||||
|
||||
Each counter reads the container the way the format defines it:
|
||||
|
||||
- XML (NISO-STS): `xml.etree.ElementTree` straight on the bytes.
|
||||
- STS JSON twin: the publisher's own node tree (`standardContent`, nodes with
|
||||
`e`/`t`/`x`), walked with the same element roles as the XML.
|
||||
- docx / pptx / xlsx / odt: the zip members' own XML.
|
||||
- PDF: pdfplumber OBJECTS (pages, image placements) and, as a second witness,
|
||||
poppler (`pdfinfo`, `pdfimages -list`). pdfplumber is also what the reader
|
||||
extracts text with, which is why the gate never trusts a PDF count that the
|
||||
poppler side does not repeat.
|
||||
- HTML: `html.parser` from the stdlib.
|
||||
- md / txt / csv / json / rtf: stdlib line, csv and json readers, and a
|
||||
control-word scan for rtf.
|
||||
|
||||
The ELEMENT VOCABULARY is part of the gate's contract: a build that declares
|
||||
an inventory must use these names, per file type, or it is not comparable.
|
||||
Each name is defined where it is counted, and a type counts only what that
|
||||
format actually carries.
|
||||
|
||||
Image REFERENCES are resolved here too, because the accounting has to know
|
||||
which inbox files a document points at: a relative reference is taken against
|
||||
the document's own directory, never above it, and an STS reference that is
|
||||
not found there is looked for as `graphics/<basename>` -- the layout the
|
||||
publisher's STS delivery ships in. That is a fact about the delivery format,
|
||||
written down here, not borrowed from the reader.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import csv
|
||||
import io
|
||||
import json
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import zipfile
|
||||
from collections.abc import Iterator
|
||||
from dataclasses import dataclass, field
|
||||
from html.parser import HTMLParser
|
||||
from pathlib import Path, PurePosixPath
|
||||
from xml.etree import ElementTree as ET
|
||||
|
||||
WITNESS_VERSION = 1
|
||||
|
||||
#: Pointer kinds a document can hold for an image.
|
||||
LOCAL = "local"
|
||||
REMOTE = "remote"
|
||||
EMBEDDED = "embedded"
|
||||
|
||||
|
||||
class WitnessRefused(Exception):
|
||||
"""The witness will not read this file (for example a DOCTYPE)."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ImageRef:
|
||||
"""One image a document declares.
|
||||
|
||||
`target` is the inbox-relative POSIX path of the file a LOCAL reference
|
||||
resolves to, or None when it resolves to nothing inside the document's
|
||||
directory.
|
||||
"""
|
||||
|
||||
kind: str
|
||||
ref: str
|
||||
target: str | None = None
|
||||
|
||||
|
||||
@dataclass
|
||||
class Inventory:
|
||||
"""What one source file holds, element type by element type."""
|
||||
|
||||
source_file: str
|
||||
suffix: str
|
||||
witness: str
|
||||
elements: dict[str, int] = field(default_factory=dict)
|
||||
images: list[ImageRef] = field(default_factory=list)
|
||||
|
||||
@property
|
||||
def total(self) -> int:
|
||||
return sum(self.elements.values())
|
||||
|
||||
def to_json(self) -> dict[str, object]:
|
||||
return {
|
||||
"suffix": self.suffix,
|
||||
"witness": self.witness,
|
||||
"elements": dict(sorted(self.elements.items())),
|
||||
"images": [
|
||||
{"kind": ref.kind, "ref": ref.ref, "target": ref.target} for ref in self.images
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
def _local(tag: str) -> str:
|
||||
return tag.rsplit("}", 1)[-1] if "}" in tag else tag.split(":")[-1]
|
||||
|
||||
|
||||
def _is_remote(ref: str) -> bool:
|
||||
return bool(re.match(r"^[a-zA-Z][a-zA-Z0-9+.-]*:", ref)) or ref.startswith("//")
|
||||
|
||||
|
||||
def resolve_local(inbox: Path, document: Path, ref: str, *, sts: bool = False) -> str | None:
|
||||
"""The inbox-relative path a LOCAL reference names, or None.
|
||||
|
||||
Contained in the document's own directory: an absolute path or one that
|
||||
climbs above that directory resolves to nothing.
|
||||
"""
|
||||
base = document.parent
|
||||
candidates = [ref]
|
||||
if sts:
|
||||
candidates.append(f"graphics/{PurePosixPath(ref).name}")
|
||||
for candidate in candidates:
|
||||
pure = PurePosixPath(candidate)
|
||||
if pure.is_absolute() or ".." in pure.parts:
|
||||
continue
|
||||
target = base.joinpath(*pure.parts)
|
||||
if target.is_file():
|
||||
return target.relative_to(inbox).as_posix()
|
||||
return None
|
||||
|
||||
|
||||
# --- markdown / text ---------------------------------------------------------
|
||||
|
||||
_FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})(.*)$")
|
||||
_ATX = re.compile(r"^ {0,3}#{1,6}(\s|$)")
|
||||
_DELIMITER_ROW = re.compile(r"^\s*\|?\s*:?-{3,}:?\s*(\|\s*:?-{3,}:?\s*)*\|?\s*$")
|
||||
_MD_IMAGE = re.compile(r"!\[[^\]]*\]\(\s*<?([^)\s>]+)>?[^)]*\)")
|
||||
|
||||
|
||||
def _unfenced(lines: list[str]) -> tuple[list[str | None], int]:
|
||||
"""Lines with fenced ones replaced by None, and the number of fences.
|
||||
|
||||
CommonMark SS 4.5 in the parts that decide which lines are fenced: up to
|
||||
three leading spaces, a backtick info string may not hold a backtick, the
|
||||
closing fence is the same character and at least as long, and an unclosed
|
||||
fence runs to the end of the text.
|
||||
"""
|
||||
out: list[str | None] = []
|
||||
fences = 0
|
||||
opener: str | None = None
|
||||
for line in lines:
|
||||
if opener is None:
|
||||
match = _FENCE_OPEN.match(line)
|
||||
if match and not (match.group(1)[0] == "`" and "`" in match.group(2)):
|
||||
opener = match.group(1)
|
||||
fences += 1
|
||||
out.append(None)
|
||||
continue
|
||||
out.append(line)
|
||||
continue
|
||||
out.append(None)
|
||||
stripped = line.strip()
|
||||
if (
|
||||
stripped
|
||||
and set(stripped) == {opener[0]}
|
||||
and len(stripped) >= len(opener)
|
||||
and len(line) - len(line.lstrip(" ")) <= 3
|
||||
):
|
||||
opener = None
|
||||
return out, fences
|
||||
|
||||
|
||||
def count_markdown(text: str) -> tuple[dict[str, int], list[str]]:
|
||||
"""heading: ATX lines outside a fence. table: a pipe row followed by a
|
||||
delimiter row. table_row: the body rows under it. image: ``
|
||||
outside a fence. code_block: a fence. paragraph: a run of non-blank lines
|
||||
outside a fence that holds none of the above."""
|
||||
lines, fences = _unfenced(text.split("\n"))
|
||||
elements = {
|
||||
"heading": 0,
|
||||
"paragraph": 0,
|
||||
"table": 0,
|
||||
"table_row": 0,
|
||||
"image": 0,
|
||||
"code_block": fences,
|
||||
}
|
||||
refs: list[str] = []
|
||||
in_table = False
|
||||
in_paragraph = False
|
||||
delimiter_rows: set[int] = set()
|
||||
for index, line in enumerate(lines):
|
||||
if index in delimiter_rows:
|
||||
continue
|
||||
if line is None or not line.strip():
|
||||
in_table = False
|
||||
in_paragraph = False
|
||||
continue
|
||||
if in_table:
|
||||
if "|" in line:
|
||||
elements["table_row"] += 1
|
||||
continue
|
||||
in_table = False
|
||||
following = lines[index + 1] if index + 1 < len(lines) else None
|
||||
if "|" in line and following is not None and _DELIMITER_ROW.match(following):
|
||||
elements["table"] += 1
|
||||
in_table = True
|
||||
delimiter_rows.add(index + 1) # the delimiter row is not a body row
|
||||
in_paragraph = False
|
||||
continue
|
||||
if _ATX.match(line):
|
||||
elements["heading"] += 1
|
||||
in_paragraph = False
|
||||
continue
|
||||
found = _MD_IMAGE.findall(line)
|
||||
if found:
|
||||
elements["image"] += len(found)
|
||||
refs.extend(found)
|
||||
if not _MD_IMAGE.sub("", line).strip():
|
||||
in_paragraph = False
|
||||
continue
|
||||
if not in_paragraph:
|
||||
elements["paragraph"] += 1
|
||||
in_paragraph = True
|
||||
return elements, refs
|
||||
|
||||
|
||||
def count_text(text: str) -> dict[str, int]:
|
||||
"""paragraph: a run of non-blank lines. line: a non-blank line."""
|
||||
paragraphs = 0
|
||||
lines = 0
|
||||
previous_blank = True
|
||||
for line in text.split("\n"):
|
||||
if line.strip():
|
||||
lines += 1
|
||||
if previous_blank:
|
||||
paragraphs += 1
|
||||
previous_blank = False
|
||||
else:
|
||||
previous_blank = True
|
||||
return {"paragraph": paragraphs, "line": lines}
|
||||
|
||||
|
||||
def count_csv(text: str) -> dict[str, int]:
|
||||
"""header_cell: cells of the first row. row / cell: every row after it."""
|
||||
rows = [row for row in csv.reader(io.StringIO(text)) if row]
|
||||
if not rows:
|
||||
return {"header_cell": 0, "row": 0, "cell": 0}
|
||||
return {
|
||||
"header_cell": len(rows[0]),
|
||||
"row": len(rows) - 1,
|
||||
"cell": sum(len(row) for row in rows[1:]),
|
||||
}
|
||||
|
||||
|
||||
def count_json(text: str) -> dict[str, int]:
|
||||
"""key: an object member. value: a leaf (string, number, boolean, null)."""
|
||||
counts = {"key": 0, "value": 0}
|
||||
|
||||
def walk(node: object) -> None:
|
||||
if isinstance(node, dict):
|
||||
counts["key"] += len(node)
|
||||
for child in node.values():
|
||||
walk(child)
|
||||
elif isinstance(node, list):
|
||||
for child in node:
|
||||
walk(child)
|
||||
else:
|
||||
counts["value"] += 1
|
||||
|
||||
walk(json.loads(text))
|
||||
return counts
|
||||
|
||||
|
||||
# --- html --------------------------------------------------------------------
|
||||
|
||||
|
||||
class _HtmlCounter(HTMLParser):
|
||||
"""heading: h1-h6. paragraph: p. list_item: li. table: table. cell: td,
|
||||
th. image: img."""
|
||||
|
||||
_ROLES = {
|
||||
**{f"h{level}": "heading" for level in range(1, 7)},
|
||||
"p": "paragraph",
|
||||
"li": "list_item",
|
||||
"table": "table",
|
||||
"td": "cell",
|
||||
"th": "cell",
|
||||
"img": "image",
|
||||
}
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__(convert_charrefs=True)
|
||||
self.elements = {role: 0 for role in sorted(set(self._ROLES.values()))}
|
||||
self.refs: list[str] = []
|
||||
|
||||
def handle_starttag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None:
|
||||
role = self._ROLES.get(tag)
|
||||
if role is None:
|
||||
return
|
||||
self.elements[role] += 1
|
||||
if tag == "img":
|
||||
self.refs.append(dict(attrs).get("src") or "")
|
||||
|
||||
|
||||
# --- xml / sts ---------------------------------------------------------------
|
||||
|
||||
STS_ROLES = (
|
||||
"section",
|
||||
"title",
|
||||
"section_label",
|
||||
"paragraph",
|
||||
"table",
|
||||
"table_label",
|
||||
"cell",
|
||||
"list_item",
|
||||
"image",
|
||||
"footnote",
|
||||
)
|
||||
|
||||
|
||||
def _sts_role(tag: str, parent: str | None, grandparent: str | None) -> str | None:
|
||||
"""The one mapping from an STS element to its accounting role.
|
||||
|
||||
Used by BOTH STS witnesses, and the role is the unit, not the tag, because
|
||||
the publisher's two deliveries of one document place the same text
|
||||
differently (measured on R761 Prosesskoden:2025, 2026-09-17):
|
||||
|
||||
- a section's label: XML `sec/label` on 7 714 sections; JSON `sec/label`
|
||||
on 4 954 and `sec/title/label` on the 2 760 that carry a title.
|
||||
- a table's label: XML `table-wrap/label` (10); JSON
|
||||
`table-wrap/table/caption` (10).
|
||||
|
||||
Counted by tag, the two witnesses disagree by 2 760 and by 10 on text
|
||||
both of them carry.
|
||||
"""
|
||||
if tag == "sec":
|
||||
return "section"
|
||||
if tag == "title" and parent == "sec":
|
||||
return "title"
|
||||
if tag == "label" and (parent == "sec" or (parent == "title" and grandparent == "sec")):
|
||||
return "section_label"
|
||||
if tag == "label" and parent == "table-wrap":
|
||||
return "table_label"
|
||||
if tag == "caption" and parent == "table" and grandparent == "table-wrap":
|
||||
return "table_label"
|
||||
if tag == "p":
|
||||
return "paragraph"
|
||||
if tag == "table-wrap":
|
||||
return "table"
|
||||
if tag in ("td", "th"):
|
||||
return "cell"
|
||||
if tag == "list-item":
|
||||
return "list_item"
|
||||
if tag in ("graphic", "inline-graphic"):
|
||||
return "image"
|
||||
if tag == "fn":
|
||||
return "footnote"
|
||||
return None
|
||||
|
||||
|
||||
def count_sts_xml(data: bytes) -> tuple[dict[str, int], list[str], bool]:
|
||||
"""Element roles of an STS document; `element` alone for other XML."""
|
||||
if b"<!DOCTYPE" in data:
|
||||
raise WitnessRefused("a DOCTYPE is not parsed")
|
||||
root = ET.fromstring(data)
|
||||
sts = _local(root.tag) == "standard" or any(_local(el.tag) == "sec" for el in root.iter())
|
||||
if not sts:
|
||||
return {"element": sum(1 for _ in root.iter())}, [], False
|
||||
elements = {role: 0 for role in STS_ROLES}
|
||||
refs: list[str] = []
|
||||
|
||||
def walk(node: ET.Element, parent: str | None, grandparent: str | None) -> None:
|
||||
tag = _local(node.tag)
|
||||
role = _sts_role(tag, parent, grandparent)
|
||||
if role is not None:
|
||||
elements[role] += 1
|
||||
if role == "image":
|
||||
href = next((value for key, value in node.attrib.items() if _local(key) == "href"), "")
|
||||
refs.append(href)
|
||||
for child in node:
|
||||
walk(child, tag, parent)
|
||||
|
||||
walk(root, None, None)
|
||||
return elements, refs, True
|
||||
|
||||
|
||||
def count_sts_json(data: bytes) -> dict[str, int]:
|
||||
"""The same roles, read from the publisher's JSON node tree."""
|
||||
document = json.loads(data)
|
||||
elements = {role: 0 for role in STS_ROLES}
|
||||
|
||||
def walk(node: dict[str, object], parent: str | None, grandparent: str | None) -> None:
|
||||
body = node.get("x")
|
||||
if not isinstance(body, dict):
|
||||
return
|
||||
tag = str(body.get("tag"))
|
||||
role = _sts_role(tag, parent, grandparent)
|
||||
if role is not None:
|
||||
elements[role] += 1
|
||||
for child in body.get("c") or []:
|
||||
walk(child, tag, parent)
|
||||
|
||||
for child in document["standardContent"]["c"]:
|
||||
walk(child, None, None)
|
||||
return elements
|
||||
|
||||
|
||||
# --- office zips -------------------------------------------------------------
|
||||
|
||||
_W = "{http://schemas.openxmlformats.org/wordprocessingml/2006/main}"
|
||||
_A = "{http://schemas.openxmlformats.org/drawingml/2006/main}"
|
||||
_P = "{http://schemas.openxmlformats.org/presentationml/2006/main}"
|
||||
_S = "{http://schemas.openxmlformats.org/spreadsheetml/2006/main}"
|
||||
_XDR = "{http://schemas.openxmlformats.org/drawingml/2006/spreadsheetDrawing}"
|
||||
_TEXT = "{urn:oasis:names:tc:opendocument:xmlns:text:1.0}"
|
||||
_TABLE = "{urn:oasis:names:tc:opendocument:xmlns:table:1.0}"
|
||||
_DRAW = "{urn:oasis:names:tc:opendocument:xmlns:drawing:1.0}"
|
||||
_HEADING_STYLE = re.compile(r"^(heading|overskrift|title|tittel)\s*\d*$", re.IGNORECASE)
|
||||
|
||||
|
||||
def _text_of(node: ET.Element, tag: str) -> str:
|
||||
return "".join(t.text or "" for t in node.iter(tag))
|
||||
|
||||
|
||||
def count_docx(data: bytes) -> dict[str, int]:
|
||||
"""heading: a w:p whose style is a heading or title style. paragraph: any
|
||||
other w:p with text. table: w:tbl. cell: w:tc. image: a:blip.
|
||||
footnote: a w:footnote with a positive id."""
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
root = ET.fromstring(archive.read("word/document.xml"))
|
||||
footnotes = 0
|
||||
if "word/footnotes.xml" in archive.namelist():
|
||||
notes = ET.fromstring(archive.read("word/footnotes.xml"))
|
||||
footnotes = sum(
|
||||
1 for note in notes.iter(f"{_W}footnote") if int(note.get(f"{_W}id", "0")) > 0
|
||||
)
|
||||
headings = paragraphs = 0
|
||||
for para in root.iter(f"{_W}p"):
|
||||
style = para.find(f"{_W}pPr/{_W}pStyle")
|
||||
if style is not None and _HEADING_STYLE.match(style.get(f"{_W}val", "")):
|
||||
headings += 1
|
||||
elif _text_of(para, f"{_W}t").strip():
|
||||
paragraphs += 1
|
||||
return {
|
||||
"heading": headings,
|
||||
"paragraph": paragraphs,
|
||||
"table": sum(1 for _ in root.iter(f"{_W}tbl")),
|
||||
"cell": sum(1 for _ in root.iter(f"{_W}tc")),
|
||||
"image": sum(1 for _ in root.iter(f"{_A}blip")),
|
||||
"footnote": footnotes,
|
||||
}
|
||||
|
||||
|
||||
def count_pptx(data: bytes) -> dict[str, int]:
|
||||
"""slide: ppt/slides/slideN.xml. title: a shape whose placeholder is a
|
||||
title. paragraph: an a:p with text outside a table and outside a title.
|
||||
table: a:tbl. cell: a:tc. image: p:pic."""
|
||||
counts = {"slide": 0, "title": 0, "paragraph": 0, "table": 0, "cell": 0, "image": 0}
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
slides = [n for n in archive.namelist() if re.fullmatch(r"ppt/slides/slide\d+\.xml", n)]
|
||||
for name in slides:
|
||||
counts["slide"] += 1
|
||||
root = ET.fromstring(archive.read(name))
|
||||
counts["table"] += sum(1 for _ in root.iter(f"{_A}tbl"))
|
||||
counts["cell"] += sum(1 for _ in root.iter(f"{_A}tc"))
|
||||
counts["image"] += sum(1 for _ in root.iter(f"{_P}pic"))
|
||||
for shape in root.iter(f"{_P}sp"):
|
||||
placeholder = shape.find(f"{_P}nvSpPr/{_P}nvPr/{_P}ph")
|
||||
is_title = placeholder is not None and placeholder.get("type") in (
|
||||
"title",
|
||||
"ctrTitle",
|
||||
)
|
||||
texts = [p for p in shape.iter(f"{_A}p") if _text_of(p, f"{_A}t").strip()]
|
||||
if is_title and texts:
|
||||
counts["title"] += 1
|
||||
else:
|
||||
counts["paragraph"] += len(texts)
|
||||
return counts
|
||||
|
||||
|
||||
def count_xlsx(data: bytes) -> dict[str, int]:
|
||||
"""sheet: xl/worksheets/sheetN.xml. row: a row holding a value. cell: a c
|
||||
with a value. image: an xdr:pic in a drawing."""
|
||||
counts = {"sheet": 0, "row": 0, "cell": 0, "image": 0}
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
for name in archive.namelist():
|
||||
if re.fullmatch(r"xl/worksheets/sheet\d+\.xml", name):
|
||||
counts["sheet"] += 1
|
||||
root = ET.fromstring(archive.read(name))
|
||||
for row in root.iter(f"{_S}row"):
|
||||
valued = [
|
||||
c
|
||||
for c in row.iter(f"{_S}c")
|
||||
if c.find(f"{_S}v") is not None or c.find(f"{_S}is") is not None
|
||||
]
|
||||
counts["cell"] += len(valued)
|
||||
counts["row"] += 1 if valued else 0
|
||||
elif re.fullmatch(r"xl/drawings/drawing\d+\.xml", name):
|
||||
root = ET.fromstring(archive.read(name))
|
||||
counts["image"] += sum(1 for _ in root.iter(f"{_XDR}pic"))
|
||||
return counts
|
||||
|
||||
|
||||
def count_odt(data: bytes) -> dict[str, int]:
|
||||
"""heading: text:h. paragraph: a text:p with text outside a table cell.
|
||||
table: table:table. cell: table:table-cell. list_item: text:list-item.
|
||||
image: draw:image."""
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
root = ET.fromstring(archive.read("content.xml"))
|
||||
in_cell: set[int] = set()
|
||||
for cell in root.iter(f"{_TABLE}table-cell"):
|
||||
in_cell.update(id(p) for p in cell.iter(f"{_TEXT}p"))
|
||||
return {
|
||||
"heading": sum(1 for _ in root.iter(f"{_TEXT}h")),
|
||||
"paragraph": sum(
|
||||
1
|
||||
for p in root.iter(f"{_TEXT}p")
|
||||
if id(p) not in in_cell and "".join(p.itertext()).strip()
|
||||
),
|
||||
"table": sum(1 for _ in root.iter(f"{_TABLE}table")),
|
||||
"cell": sum(1 for _ in root.iter(f"{_TABLE}table-cell")),
|
||||
"list_item": sum(1 for _ in root.iter(f"{_TEXT}list-item")),
|
||||
"image": sum(1 for _ in root.iter(f"{_DRAW}image")),
|
||||
}
|
||||
|
||||
|
||||
def count_rtf(text: str) -> dict[str, int]:
|
||||
"""paragraph: \\par. table_row: \\row. cell: \\cell. image: \\pict.
|
||||
A control word ends at the first non-letter, so \\pard is not \\par."""
|
||||
|
||||
def word(name: str) -> int:
|
||||
return len(re.findall(rf"\\{name}(?![a-zA-Z])", text))
|
||||
|
||||
return {
|
||||
"paragraph": word("par"),
|
||||
"table_row": word("row"),
|
||||
"cell": word("cell"),
|
||||
"image": word("pict"),
|
||||
}
|
||||
|
||||
|
||||
# --- pdf ---------------------------------------------------------------------
|
||||
|
||||
|
||||
def pdf_objects(path: Path) -> dict[str, int] | None:
|
||||
"""page and image placements as pdfplumber sees them; None without it."""
|
||||
try:
|
||||
import pdfplumber
|
||||
except ImportError:
|
||||
return None
|
||||
with pdfplumber.open(str(path)) as pdf:
|
||||
pages = len(pdf.pages)
|
||||
images = 0
|
||||
for page in pdf.pages:
|
||||
images += len(page.images)
|
||||
page.close()
|
||||
return {"page": pages, "image": images}
|
||||
|
||||
|
||||
def pdf_poppler(path: Path) -> dict[str, int] | None:
|
||||
"""page (pdfinfo) and image (pdfimages -list, rows of type `image`); None
|
||||
when poppler is not installed."""
|
||||
info = shutil.which("pdfinfo")
|
||||
lister = shutil.which("pdfimages")
|
||||
if info is None or lister is None:
|
||||
return None
|
||||
meta = subprocess.run([info, str(path)], capture_output=True, text=True, check=True).stdout
|
||||
pages_match = re.search(r"^Pages:\s+(\d+)", meta, re.MULTILINE)
|
||||
listing = subprocess.run(
|
||||
[lister, "-list", str(path)], capture_output=True, text=True, check=True
|
||||
).stdout
|
||||
images = 0
|
||||
for line in listing.splitlines()[2:]:
|
||||
cells = line.split()
|
||||
if len(cells) > 2 and cells[2] == "image":
|
||||
images += 1
|
||||
return {"page": int(pages_match.group(1)) if pages_match else 0, "image": images}
|
||||
|
||||
|
||||
# --- one file ----------------------------------------------------------------
|
||||
|
||||
WITNESSED_SUFFIXES = (
|
||||
".csv",
|
||||
".docx",
|
||||
".htm",
|
||||
".html",
|
||||
".json",
|
||||
".md",
|
||||
".odt",
|
||||
".pdf",
|
||||
".pptx",
|
||||
".rtf",
|
||||
".txt",
|
||||
".xlsx",
|
||||
".xml",
|
||||
)
|
||||
|
||||
|
||||
def witness_file(inbox: Path, path: Path) -> Inventory:
|
||||
"""Count one file under `inbox`. Raises WitnessRefused for a file the
|
||||
witness does not read."""
|
||||
suffix = path.suffix.lower()
|
||||
relative = path.relative_to(inbox).as_posix()
|
||||
data = path.read_bytes()
|
||||
refs: list[str] = []
|
||||
sts = False
|
||||
if suffix == ".md":
|
||||
elements, refs = count_markdown(data.decode("utf-8-sig"))
|
||||
witness = "markdown lines"
|
||||
elif suffix == ".txt":
|
||||
elements, witness = count_text(data.decode("utf-8-sig")), "text lines"
|
||||
elif suffix == ".csv":
|
||||
elements, witness = count_csv(data.decode("utf-8-sig")), "csv"
|
||||
elif suffix == ".json":
|
||||
elements, witness = count_json(data.decode("utf-8-sig")), "json"
|
||||
elif suffix in (".html", ".htm"):
|
||||
parser = _HtmlCounter()
|
||||
parser.feed(data.decode("utf-8-sig"))
|
||||
parser.close()
|
||||
elements, refs, witness = parser.elements, parser.refs, "html.parser"
|
||||
elif suffix == ".xml":
|
||||
elements, refs, sts = count_sts_xml(data)
|
||||
witness = "xml.etree"
|
||||
elif suffix == ".docx":
|
||||
elements, witness = count_docx(data), "docx zip xml"
|
||||
elif suffix == ".pptx":
|
||||
elements, witness = count_pptx(data), "pptx zip xml"
|
||||
elif suffix == ".xlsx":
|
||||
elements, witness = count_xlsx(data), "xlsx zip xml"
|
||||
elif suffix == ".odt":
|
||||
elements, witness = count_odt(data), "odt zip xml"
|
||||
elif suffix == ".rtf":
|
||||
elements, witness = count_rtf(data.decode("latin-1")), "rtf control words"
|
||||
elif suffix == ".pdf":
|
||||
objects = pdf_objects(path)
|
||||
if objects is None:
|
||||
raise WitnessRefused("pdfplumber is not installed")
|
||||
elements, witness = objects, "pdfplumber objects"
|
||||
else:
|
||||
raise WitnessRefused(f"no witness for {suffix or 'a file without a suffix'}")
|
||||
inventory = Inventory(relative, suffix, witness, dict(elements))
|
||||
for ref in refs:
|
||||
if not ref or _is_remote(ref):
|
||||
inventory.images.append(ImageRef(REMOTE, ref))
|
||||
else:
|
||||
target = resolve_local(inbox, path, ref, sts=sts)
|
||||
inventory.images.append(ImageRef(LOCAL, ref, target))
|
||||
embedded = elements.get("image", 0) - len(refs)
|
||||
inventory.images.extend(ImageRef(EMBEDDED, "") for _ in range(max(0, embedded)))
|
||||
return inventory
|
||||
|
||||
|
||||
def walk(inbox: Path) -> Iterator[Path]:
|
||||
"""Every file under `inbox`, sorted by relative path, dot-entries skipped."""
|
||||
for path in sorted(inbox.rglob("*"), key=lambda p: p.relative_to(inbox).as_posix()):
|
||||
if path.is_file() and not any(
|
||||
part.startswith(".") for part in path.relative_to(inbox).parts
|
||||
):
|
||||
yield path
|
||||
|
||||
|
||||
def witness_inbox(inbox: Path) -> dict[str, object]:
|
||||
"""The committed fasit form: every document's inventory, every other file
|
||||
with the documents that point at it."""
|
||||
documents: dict[str, object] = {}
|
||||
others: list[str] = []
|
||||
pointed: dict[str, list[str]] = {}
|
||||
for path in walk(inbox):
|
||||
if path.suffix.lower() not in WITNESSED_SUFFIXES:
|
||||
others.append(path.relative_to(inbox).as_posix())
|
||||
continue
|
||||
inventory = witness_file(inbox, path)
|
||||
documents[inventory.source_file] = inventory.to_json()
|
||||
for ref in inventory.images:
|
||||
if ref.target is not None:
|
||||
pointed.setdefault(ref.target, [])
|
||||
if inventory.source_file not in pointed[ref.target]:
|
||||
pointed[ref.target].append(inventory.source_file)
|
||||
return {
|
||||
"witness_version": WITNESS_VERSION,
|
||||
"documents": documents,
|
||||
"files": {name: {"pointed_at_by": pointed.get(name, [])} for name in others},
|
||||
}
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
import argparse
|
||||
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0])
|
||||
parser.add_argument("inbox", type=Path)
|
||||
args = parser.parse_args(argv)
|
||||
print(json.dumps(witness_inbox(args.inbox), indent=2, ensure_ascii=False, sort_keys=True))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
Loading…
Add table
Add a link
Reference in a new issue