test(accounting): the witnesses see what the formats actually hold (M-1..M-3)
Rows 2 and 3 require the build's inventory to EQUAL the witness's, so what the witness does not count, nothing can lose visibly. An independent review put a header and a comment in a docx, measured 0 of either in the bundle, and the accounting still read "2 of 2 carried". Thirteen classes are now counted, each with a red test written first: docx header/footer, comment, endnote and text box (a box's paragraphs are its own, or the text is booked twice) - pptx speaker note and hidden slide (`show="0"`, no longer counted as an ordinary slide) - xlsx formula and hidden sheet (the state lives in `workbook.xml` and is reached through the relationship id, so the sheet part itself says nothing about it) - odt header/footer from `styles.xml` and annotation (counted as prose, it made the accounting demand a reader carry a note the author wrote to themselves) - STS `mixed-citation`, `mml:math`, `fig` and its caption, measured by the review at 4.1 % of N200's source text and 3.9 % of N100's. M-2: the two STS witnesses had ONE role map between them, so row 5 -- "two witnesses agree" -- could not see a hole in it. `_sts_role_xml` and `_sts_role_json` are written apart, each for its own delivery, and a test holds them apart. M-3: 20 of 63 element types had a count of ZERO in their only fixture. Seven hand-built documents close it, every element type now occurs at least once (a test asserts it), and ALL TWENTY documents carry a hand count read off the fixture's own bytes (four did before). `.xlsx image` -- the operator's own proposed exception -- could not be exercised at all until now. Every witness also states WHAT IT STILL DOES NOT COUNT, per file type, and the gate prints that list on every run. THE FIXTURE ROWS ARE RED NOW, AND THAT IS THE POINT. Row 2 red on .docx, .odt, .pptx, .xlsx and .xml; row 3 at u = 25, d = 2 over the new classes, including a footnote and four spreadsheet cells the build genuinely drops. `0 claimed and not found` on the same run: nothing the build DOES book as carried failed the bundle check, so the red is the build's and not the instrument's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
656cbe5d02
commit
e5dc21ec2f
14 changed files with 1542 additions and 59 deletions
33
tests/fixtures/README.md
vendored
33
tests/fixtures/README.md
vendored
|
|
@ -194,6 +194,39 @@ python3 tools/okf_witness.py tests/fixtures/accounting/corpus > tests/fixtures/a
|
|||
python3 tools/okf_witness.py tests/fixtures/accounting/rejected > tests/fixtures/accounting/rejected-inventory.json
|
||||
```
|
||||
|
||||
Seven more documents were added 2026-09-18, one per format that had element
|
||||
types it could never exercise. An independent review measured **20 of 63
|
||||
element types with a count of ZERO in their only fixture**, which is why six of
|
||||
seven witness mutants survived the suite: a witness cannot be caught being
|
||||
wrong about something it never sees. They are written part by part by
|
||||
`make_accounting_fixtures.py` in this directory, for the same reason the XML
|
||||
fixtures are hand-written:
|
||||
|
||||
```
|
||||
python3 tests/fixtures/accounting/make_accounting_fixtures.py
|
||||
```
|
||||
|
||||
| Fixture | What it carries that nothing else did |
|
||||
|---|---|
|
||||
| `topptekst-og-kommentar.docx` | A header, a footer, a comment, an endnote and a **text box** -- and a footnote, a table and a heading, three types the only other docx has at 0. The header says "Utkast - gjelder ikke etter 2026-01-01" and the comment says the requirement does NOT apply in tunnels: two statements that reverse the document's meaning and that the build carries none of. |
|
||||
| `notater-og-skjult.pptx` | A **speaker note** and a **hidden slide** (`show="0"`), plus a table and paragraphs. A hidden slide counted as an ordinary one is indistinguishable from one that is shown. |
|
||||
| `skjult-ark-og-formel.xlsx` | A **hidden sheet**, a **formula** (`<f>B2*2</f>`) and a **picture**. The picture is what makes the operator's `.xlsx image` exception exercisable at all: the old fixture had none. |
|
||||
| `liste-og-bilde.odt` | A **header and footer** (they live in `styles.xml`, so a reader of `content.xml` cannot see them), an **annotation**, a list and a picture. |
|
||||
| `bilde.rtf` | A `\pict` picture: the rtf witness's image count was 0 in its only fixture. |
|
||||
| `figur.html` | A picture and a table under `.html`; `side.htm` gained one too, so `.htm` and `.html` each exercise `image`. |
|
||||
| `sts-rikt.xml` | A **`mixed-citation`**, an **`mml:math`**, a **`fig` with a caption**, a table with a label, cells, a list item and a footnote -- six STS roles the R761 delivery does not contain at all, which is why the gate's only real corpus could not see the hole in the role map. |
|
||||
|
||||
### The hand counts
|
||||
|
||||
Row 1's fasit is the witness's own output, so a hand count is the only number
|
||||
in this loop the witness did not produce. Four of thirteen documents had one;
|
||||
**all twenty have one now**, in `HAND_COUNTS` in
|
||||
`tests/test_accounting_gate.py`, and `test_the_hand_counts_cover_every_document_of_the_corpus`
|
||||
fails if a document is added without one. Each was counted by reading the
|
||||
fixture's own bytes -- the XML parts of a zip, the control words of the rtf,
|
||||
the objects of the PDF -- never by running the witness and writing down what
|
||||
it said.
|
||||
|
||||
`witness/prosess-84-sts.twin.json` is the STS document written by hand in the
|
||||
publisher's JSON node form (`standardContent`, nodes with `e`/`t`/`x`), so the
|
||||
two STS witnesses can be compared on a fixture as well as on R761. Eight of
|
||||
|
|
|
|||
1
tests/fixtures/accounting/corpus/bilde.rtf
vendored
Normal file
1
tests/fixtures/accounting/corpus/bilde.rtf
vendored
Normal file
|
|
@ -0,0 +1 @@
|
|||
{\rtf1\ansi\deff0{\fonttbl{\f0 Times New Roman;}}\pard Figur 84-1 viser prinsippet.\par\pard{\pict\pngblip\picw16\pich16 89504e470d0a1a0a}\par}
|
||||
11
tests/fixtures/accounting/corpus/figur.html
vendored
Normal file
11
tests/fixtures/accounting/corpus/figur.html
vendored
Normal file
|
|
@ -0,0 +1,11 @@
|
|||
<!DOCTYPE html>
|
||||
<html lang="no">
|
||||
<head><title>Figur 84-1</title></head>
|
||||
<body>
|
||||
<h1>Figur 84-1</h1>
|
||||
<p>Prinsippet for toleranseklasser.</p>
|
||||
<img src="graphics/figur-84-1.png" alt="Prinsippskisse">
|
||||
<table><tr><th>Klasse</th><th>Avvik</th></tr><tr><td>A</td><td>5 mm</td></tr></table>
|
||||
<ul><li>Klasse A</li><li>Klasse B</li></ul>
|
||||
</body>
|
||||
</html>
|
||||
BIN
tests/fixtures/accounting/corpus/liste-og-bilde.odt
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/liste-og-bilde.odt
vendored
Normal file
Binary file not shown.
BIN
tests/fixtures/accounting/corpus/notater-og-skjult.pptx
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/notater-og-skjult.pptx
vendored
Normal file
Binary file not shown.
1
tests/fixtures/accounting/corpus/side.htm
vendored
1
tests/fixtures/accounting/corpus/side.htm
vendored
|
|
@ -3,6 +3,7 @@
|
|||
<body>
|
||||
<h1>Belysning</h1>
|
||||
<p>Terskelluminansen er 145 candela.</p>
|
||||
<img src="graphics/tabell-84-2.png" alt="Sonekart">
|
||||
<h2>Soner</h2>
|
||||
<ul><li>Inngang</li><li>Hovedløp</li></ul>
|
||||
<table><tr><th>Sone</th><th>Lengde</th></tr><tr><td>Inngang</td><td>120 m</td></tr></table>
|
||||
|
|
|
|||
BIN
tests/fixtures/accounting/corpus/skjult-ark-og-formel.xlsx
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/skjult-ark-og-formel.xlsx
vendored
Normal file
Binary file not shown.
16
tests/fixtures/accounting/corpus/sts-rikt.xml
vendored
Normal file
16
tests/fixtures/accounting/corpus/sts-rikt.xml
vendored
Normal file
|
|
@ -0,0 +1,16 @@
|
|||
<standard>
|
||||
<front><std-ident><doc-number>R762</doc-number><year>2025</year></std-ident></front>
|
||||
<body>
|
||||
<sec><label>85</label><title>Vegdekker</title>
|
||||
<p>Dekket skal ha jevnhet etter <mixed-citation>NS-EN 13036-1:2010</mixed-citation>.</p>
|
||||
<p>Kravet regnes som <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>IRI</mml:mi><mml:mo><</mml:mo><mml:mn>2</mml:mn></mml:math>.</p>
|
||||
<fig><label>Figur 85-1</label><caption><p>Maalepunkter langs vegbanen.</p></caption>
|
||||
<graphic xlink:href="figur-84-1.png" xmlns:xlink="http://www.w3.org/1999/xlink"/></fig>
|
||||
<table-wrap><label>Tabell 85-1</label>
|
||||
<table><tr><th>Klasse</th><th>IRI</th></tr><tr><td>1</td><td>1,5</td></tr></table>
|
||||
</table-wrap>
|
||||
<list><list-item><p>Maales hvert 20. meter.</p></list-item></list>
|
||||
<fn><p>Gjelder ikke gang- og sykkelveger.</p></fn>
|
||||
</sec>
|
||||
</body>
|
||||
</standard>
|
||||
BIN
tests/fixtures/accounting/corpus/topptekst-og-kommentar.docx
vendored
Normal file
BIN
tests/fixtures/accounting/corpus/topptekst-og-kommentar.docx
vendored
Normal file
Binary file not shown.
615
tests/fixtures/accounting/inventory.json
vendored
615
tests/fixtures/accounting/inventory.json
vendored
|
|
@ -1,5 +1,99 @@
|
|||
{
|
||||
"documents": {
|
||||
"bilde.rtf": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"image": 1,
|
||||
"paragraph": 2,
|
||||
"table_row": 0
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "embedded",
|
||||
"ref": "",
|
||||
"target": null
|
||||
}
|
||||
],
|
||||
"suffix": ".rtf",
|
||||
"texts": {
|
||||
"cell": [],
|
||||
"image": [
|
||||
[]
|
||||
],
|
||||
"paragraph": [
|
||||
[
|
||||
"Figur 84-1 viser prinsippet."
|
||||
],
|
||||
[]
|
||||
],
|
||||
"table_row": []
|
||||
},
|
||||
"witness": "rtf control words"
|
||||
},
|
||||
"figur.html": {
|
||||
"elements": {
|
||||
"cell": 4,
|
||||
"heading": 1,
|
||||
"image": 1,
|
||||
"list_item": 2,
|
||||
"paragraph": 1,
|
||||
"table": 1
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "graphics/figur-84-1.png",
|
||||
"target": "graphics/figur-84-1.png"
|
||||
}
|
||||
],
|
||||
"suffix": ".html",
|
||||
"texts": {
|
||||
"cell": [
|
||||
[
|
||||
"Klasse"
|
||||
],
|
||||
[
|
||||
"Avvik"
|
||||
],
|
||||
[
|
||||
"A"
|
||||
],
|
||||
[
|
||||
"5 mm"
|
||||
]
|
||||
],
|
||||
"heading": [
|
||||
[
|
||||
"Figur 84-1"
|
||||
]
|
||||
],
|
||||
"image": [
|
||||
[]
|
||||
],
|
||||
"list_item": [
|
||||
[
|
||||
"Klasse A"
|
||||
],
|
||||
[
|
||||
"Klasse B"
|
||||
]
|
||||
],
|
||||
"paragraph": [
|
||||
[
|
||||
"Prinsippet for toleranseklasser."
|
||||
]
|
||||
],
|
||||
"table": [
|
||||
[
|
||||
"Klasse",
|
||||
"Avvik",
|
||||
"A",
|
||||
"5 mm"
|
||||
]
|
||||
]
|
||||
},
|
||||
"witness": "html.parser"
|
||||
},
|
||||
"krav-rikt-tekstformat.rtf": {
|
||||
"elements": {
|
||||
"cell": 56,
|
||||
|
|
@ -303,7 +397,9 @@
|
|||
},
|
||||
"krav-tekstdokument.odt": {
|
||||
"elements": {
|
||||
"annotation": 0,
|
||||
"cell": 56,
|
||||
"header_footer": 0,
|
||||
"heading": 1,
|
||||
"image": 0,
|
||||
"list_item": 0,
|
||||
|
|
@ -313,6 +409,7 @@
|
|||
"images": [],
|
||||
"suffix": ".odt",
|
||||
"texts": {
|
||||
"annotation": [],
|
||||
"cell": [
|
||||
[
|
||||
"Dokumentnummer:"
|
||||
|
|
@ -483,6 +580,7 @@
|
|||
"2,5 cd"
|
||||
]
|
||||
],
|
||||
"header_footer": [],
|
||||
"heading": [
|
||||
[
|
||||
"Kravspesifikasjon for tunnelbelysning"
|
||||
|
|
@ -563,6 +661,86 @@
|
|||
},
|
||||
"witness": "odt zip xml"
|
||||
},
|
||||
"liste-og-bilde.odt": {
|
||||
"elements": {
|
||||
"annotation": 1,
|
||||
"cell": 2,
|
||||
"header_footer": 2,
|
||||
"heading": 1,
|
||||
"image": 1,
|
||||
"list_item": 2,
|
||||
"paragraph": 4,
|
||||
"table": 1
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "embedded",
|
||||
"ref": "",
|
||||
"target": null
|
||||
}
|
||||
],
|
||||
"suffix": ".odt",
|
||||
"texts": {
|
||||
"annotation": [
|
||||
[
|
||||
"Sjekk denne mot N400 foer utsendelse."
|
||||
]
|
||||
],
|
||||
"cell": [
|
||||
[
|
||||
"Type"
|
||||
],
|
||||
[
|
||||
"Gangbru"
|
||||
]
|
||||
],
|
||||
"header_footer": [
|
||||
[
|
||||
"Intern arbeidsversjon"
|
||||
],
|
||||
[
|
||||
"Vegdirektoratet"
|
||||
]
|
||||
],
|
||||
"heading": [
|
||||
[
|
||||
"Drift av gangbruer"
|
||||
]
|
||||
],
|
||||
"image": [
|
||||
[]
|
||||
],
|
||||
"list_item": [
|
||||
[
|
||||
"Rekkverk"
|
||||
],
|
||||
[
|
||||
"Dekke"
|
||||
]
|
||||
],
|
||||
"paragraph": [
|
||||
[
|
||||
"Gangbruer inspiseres hvert aar."
|
||||
],
|
||||
[
|
||||
"Rekkverk"
|
||||
],
|
||||
[
|
||||
"Dekke"
|
||||
],
|
||||
[
|
||||
"Se figuren under."
|
||||
]
|
||||
],
|
||||
"table": [
|
||||
[
|
||||
"Type",
|
||||
"Gangbru"
|
||||
]
|
||||
]
|
||||
},
|
||||
"witness": "odt zip xml"
|
||||
},
|
||||
"logg.txt": {
|
||||
"elements": {
|
||||
"line": 4,
|
||||
|
|
@ -715,6 +893,85 @@
|
|||
},
|
||||
"witness": "markdown lines"
|
||||
},
|
||||
"notater-og-skjult.pptx": {
|
||||
"elements": {
|
||||
"cell": 4,
|
||||
"hidden_slide": 1,
|
||||
"image": 0,
|
||||
"note": 1,
|
||||
"paragraph": 2,
|
||||
"slide": 1,
|
||||
"table": 2,
|
||||
"title": 2
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".pptx",
|
||||
"texts": {
|
||||
"cell": [
|
||||
[
|
||||
"Post"
|
||||
],
|
||||
[
|
||||
"84.1"
|
||||
],
|
||||
[
|
||||
"Post"
|
||||
],
|
||||
[
|
||||
"84.1"
|
||||
]
|
||||
],
|
||||
"hidden_slide": [
|
||||
[
|
||||
"Utgaatt lysbilde",
|
||||
"Ikke vis dette.",
|
||||
"Post",
|
||||
"84.1"
|
||||
]
|
||||
],
|
||||
"image": [],
|
||||
"note": [
|
||||
[
|
||||
"Husk aa nevne at toleranseklassen er skjerpet."
|
||||
]
|
||||
],
|
||||
"paragraph": [
|
||||
[
|
||||
"Toleranser er gitt i tabell."
|
||||
],
|
||||
[
|
||||
"Ikke vis dette."
|
||||
]
|
||||
],
|
||||
"slide": [
|
||||
[
|
||||
"Prosess 84 Konstruksjoner",
|
||||
"Toleranser er gitt i tabell.",
|
||||
"Post",
|
||||
"84.1"
|
||||
]
|
||||
],
|
||||
"table": [
|
||||
[
|
||||
"Post",
|
||||
"84.1"
|
||||
],
|
||||
[
|
||||
"Post",
|
||||
"84.1"
|
||||
]
|
||||
],
|
||||
"title": [
|
||||
[
|
||||
"Prosess 84 Konstruksjoner"
|
||||
],
|
||||
[
|
||||
"Utgaatt lysbilde"
|
||||
]
|
||||
]
|
||||
},
|
||||
"witness": "pptx zip xml"
|
||||
},
|
||||
"parametre.json": {
|
||||
"elements": {
|
||||
"key": 6,
|
||||
|
|
@ -769,6 +1026,8 @@
|
|||
"prisark.xlsx": {
|
||||
"elements": {
|
||||
"cell": 18,
|
||||
"formula": 0,
|
||||
"hidden_sheet": 0,
|
||||
"image": 0,
|
||||
"row": 9,
|
||||
"sheet": 2
|
||||
|
|
@ -832,6 +1091,8 @@
|
|||
"Sum ikke oppgitt"
|
||||
]
|
||||
],
|
||||
"formula": [],
|
||||
"hidden_sheet": [],
|
||||
"image": [],
|
||||
"row": [
|
||||
[
|
||||
|
|
@ -901,11 +1162,15 @@
|
|||
"prosess-84-notat.docx": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"comment": 0,
|
||||
"endnote": 0,
|
||||
"footnote": 0,
|
||||
"header_footer": 0,
|
||||
"heading": 0,
|
||||
"image": 1,
|
||||
"paragraph": 2,
|
||||
"table": 0
|
||||
"table": 0,
|
||||
"text_box": 0
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
|
|
@ -917,7 +1182,10 @@
|
|||
"suffix": ".docx",
|
||||
"texts": {
|
||||
"cell": [],
|
||||
"comment": [],
|
||||
"endnote": [],
|
||||
"footnote": [],
|
||||
"header_footer": [],
|
||||
"heading": [],
|
||||
"image": [
|
||||
[]
|
||||
|
|
@ -930,14 +1198,17 @@
|
|||
"Etter tabellen gjelder NS-EN 13670."
|
||||
]
|
||||
],
|
||||
"table": []
|
||||
"table": [],
|
||||
"text_box": []
|
||||
},
|
||||
"witness": "docx zip xml"
|
||||
},
|
||||
"prosess-84-presentasjon.pptx": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"hidden_slide": 0,
|
||||
"image": 1,
|
||||
"note": 0,
|
||||
"paragraph": 0,
|
||||
"slide": 1,
|
||||
"table": 0,
|
||||
|
|
@ -953,9 +1224,11 @@
|
|||
"suffix": ".pptx",
|
||||
"texts": {
|
||||
"cell": [],
|
||||
"hidden_slide": [],
|
||||
"image": [
|
||||
[]
|
||||
],
|
||||
"note": [],
|
||||
"paragraph": [],
|
||||
"slide": [
|
||||
[
|
||||
|
|
@ -974,9 +1247,13 @@
|
|||
"prosess-84-sts.xml": {
|
||||
"elements": {
|
||||
"cell": 0,
|
||||
"citation": 0,
|
||||
"figure": 0,
|
||||
"figure_caption": 0,
|
||||
"footnote": 0,
|
||||
"image": 2,
|
||||
"list_item": 0,
|
||||
"math": 0,
|
||||
"paragraph": 2,
|
||||
"section": 2,
|
||||
"section_label": 2,
|
||||
|
|
@ -999,12 +1276,16 @@
|
|||
"suffix": ".xml",
|
||||
"texts": {
|
||||
"cell": [],
|
||||
"citation": [],
|
||||
"figure": [],
|
||||
"figure_caption": [],
|
||||
"footnote": [],
|
||||
"image": [
|
||||
[],
|
||||
[]
|
||||
],
|
||||
"list_item": [],
|
||||
"math": [],
|
||||
"paragraph": [
|
||||
[
|
||||
"Toleranseklasse er gitt i tabell 84-2."
|
||||
|
|
@ -1140,12 +1421,18 @@
|
|||
"elements": {
|
||||
"cell": 4,
|
||||
"heading": 2,
|
||||
"image": 0,
|
||||
"image": 1,
|
||||
"list_item": 2,
|
||||
"paragraph": 1,
|
||||
"table": 1
|
||||
},
|
||||
"images": [],
|
||||
"images": [
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "graphics/tabell-84-2.png",
|
||||
"target": "graphics/tabell-84-2.png"
|
||||
}
|
||||
],
|
||||
"suffix": ".htm",
|
||||
"texts": {
|
||||
"cell": [
|
||||
|
|
@ -1170,7 +1457,9 @@
|
|||
"Soner"
|
||||
]
|
||||
],
|
||||
"image": [],
|
||||
"image": [
|
||||
[]
|
||||
],
|
||||
"list_item": [
|
||||
[
|
||||
"Inngang"
|
||||
|
|
@ -1194,20 +1483,332 @@
|
|||
]
|
||||
},
|
||||
"witness": "html.parser"
|
||||
},
|
||||
"skjult-ark-og-formel.xlsx": {
|
||||
"elements": {
|
||||
"cell": 7,
|
||||
"formula": 1,
|
||||
"hidden_sheet": 1,
|
||||
"image": 1,
|
||||
"row": 3,
|
||||
"sheet": 1
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "embedded",
|
||||
"ref": "",
|
||||
"target": null
|
||||
}
|
||||
],
|
||||
"suffix": ".xlsx",
|
||||
"texts": {
|
||||
"cell": [
|
||||
[
|
||||
"Post"
|
||||
],
|
||||
[
|
||||
"Enhet"
|
||||
],
|
||||
[
|
||||
"Mengde"
|
||||
],
|
||||
[
|
||||
"Sum"
|
||||
],
|
||||
[
|
||||
"12"
|
||||
],
|
||||
[
|
||||
"24"
|
||||
],
|
||||
[
|
||||
"Kladd"
|
||||
]
|
||||
],
|
||||
"formula": [
|
||||
[
|
||||
"B2*2"
|
||||
]
|
||||
],
|
||||
"hidden_sheet": [
|
||||
[
|
||||
"Kladd"
|
||||
]
|
||||
],
|
||||
"image": [
|
||||
[]
|
||||
],
|
||||
"row": [
|
||||
[
|
||||
"Post",
|
||||
"Enhet",
|
||||
"Mengde"
|
||||
],
|
||||
[
|
||||
"Sum",
|
||||
"12",
|
||||
"24"
|
||||
],
|
||||
[
|
||||
"Kladd"
|
||||
]
|
||||
],
|
||||
"sheet": [
|
||||
[
|
||||
"Post",
|
||||
"Enhet",
|
||||
"Mengde",
|
||||
"Sum",
|
||||
"12",
|
||||
"24"
|
||||
]
|
||||
]
|
||||
},
|
||||
"witness": "xlsx zip xml"
|
||||
},
|
||||
"sts-rikt.xml": {
|
||||
"elements": {
|
||||
"cell": 4,
|
||||
"citation": 1,
|
||||
"figure": 1,
|
||||
"figure_caption": 1,
|
||||
"footnote": 1,
|
||||
"image": 1,
|
||||
"list_item": 1,
|
||||
"math": 1,
|
||||
"paragraph": 5,
|
||||
"section": 1,
|
||||
"section_label": 1,
|
||||
"table": 1,
|
||||
"table_label": 1,
|
||||
"title": 1
|
||||
},
|
||||
"images": [
|
||||
{
|
||||
"kind": "local",
|
||||
"ref": "figur-84-1.png",
|
||||
"target": "graphics/figur-84-1.png"
|
||||
}
|
||||
],
|
||||
"suffix": ".xml",
|
||||
"texts": {
|
||||
"cell": [
|
||||
[
|
||||
"Klasse"
|
||||
],
|
||||
[
|
||||
"IRI"
|
||||
],
|
||||
[
|
||||
"1"
|
||||
],
|
||||
[
|
||||
"1,5"
|
||||
]
|
||||
],
|
||||
"citation": [
|
||||
[
|
||||
"NS-EN 13036-1:2010"
|
||||
]
|
||||
],
|
||||
"figure": [
|
||||
[
|
||||
"Figur 85-1",
|
||||
"Maalepunkter langs vegbanen."
|
||||
]
|
||||
],
|
||||
"figure_caption": [
|
||||
[
|
||||
"Maalepunkter langs vegbanen."
|
||||
]
|
||||
],
|
||||
"footnote": [
|
||||
[
|
||||
"Gjelder ikke gang- og sykkelveger."
|
||||
]
|
||||
],
|
||||
"image": [
|
||||
[]
|
||||
],
|
||||
"list_item": [
|
||||
[
|
||||
"Maales hvert 20. meter."
|
||||
]
|
||||
],
|
||||
"math": [
|
||||
[
|
||||
"IRI",
|
||||
"<",
|
||||
"2"
|
||||
]
|
||||
],
|
||||
"paragraph": [
|
||||
[
|
||||
"Dekket skal ha jevnhet etter",
|
||||
"NS-EN 13036-1:2010",
|
||||
"."
|
||||
],
|
||||
[
|
||||
"Kravet regnes som",
|
||||
"IRI",
|
||||
"<",
|
||||
"2",
|
||||
"."
|
||||
],
|
||||
[
|
||||
"Maalepunkter langs vegbanen."
|
||||
],
|
||||
[
|
||||
"Maales hvert 20. meter."
|
||||
],
|
||||
[
|
||||
"Gjelder ikke gang- og sykkelveger."
|
||||
]
|
||||
],
|
||||
"section": [
|
||||
[
|
||||
"85",
|
||||
"Vegdekker",
|
||||
"Dekket skal ha jevnhet etter",
|
||||
"NS-EN 13036-1:2010",
|
||||
".",
|
||||
"Kravet regnes som",
|
||||
"IRI",
|
||||
"<",
|
||||
"2",
|
||||
".",
|
||||
"Figur 85-1",
|
||||
"Maalepunkter langs vegbanen.",
|
||||
"Tabell 85-1",
|
||||
"Klasse",
|
||||
"IRI",
|
||||
"1",
|
||||
"1,5",
|
||||
"Maales hvert 20. meter.",
|
||||
"Gjelder ikke gang- og sykkelveger."
|
||||
]
|
||||
],
|
||||
"section_label": [
|
||||
[
|
||||
"85"
|
||||
]
|
||||
],
|
||||
"table": [
|
||||
[
|
||||
"Tabell 85-1",
|
||||
"Klasse",
|
||||
"IRI",
|
||||
"1",
|
||||
"1,5"
|
||||
]
|
||||
],
|
||||
"table_label": [
|
||||
[
|
||||
"Tabell 85-1"
|
||||
]
|
||||
],
|
||||
"title": [
|
||||
[
|
||||
"Vegdekker"
|
||||
]
|
||||
]
|
||||
},
|
||||
"witness": "xml.etree"
|
||||
},
|
||||
"topptekst-og-kommentar.docx": {
|
||||
"elements": {
|
||||
"cell": 2,
|
||||
"comment": 1,
|
||||
"endnote": 1,
|
||||
"footnote": 1,
|
||||
"header_footer": 2,
|
||||
"heading": 1,
|
||||
"image": 0,
|
||||
"paragraph": 3,
|
||||
"table": 1,
|
||||
"text_box": 1
|
||||
},
|
||||
"images": [],
|
||||
"suffix": ".docx",
|
||||
"texts": {
|
||||
"cell": [
|
||||
[
|
||||
"Bredde"
|
||||
],
|
||||
[
|
||||
"3,0 m"
|
||||
]
|
||||
],
|
||||
"comment": [
|
||||
[
|
||||
"Unntak: gjelder IKKE gangbruer i tunnel."
|
||||
]
|
||||
],
|
||||
"endnote": [
|
||||
[
|
||||
"Kravet ble skjerpet i 2024."
|
||||
]
|
||||
],
|
||||
"footnote": [
|
||||
[
|
||||
"Se haandbok N400 kapittel 5."
|
||||
]
|
||||
],
|
||||
"header_footer": [
|
||||
[
|
||||
"Statens vegvesen, side 1"
|
||||
],
|
||||
[
|
||||
"Utkast - gjelder ikke etter 2026-01-01"
|
||||
]
|
||||
],
|
||||
"heading": [
|
||||
[
|
||||
"Krav til gangbruer"
|
||||
]
|
||||
],
|
||||
"image": [],
|
||||
"paragraph": [
|
||||
[
|
||||
"Gangbruer skal ha rekkverk paa begge sider."
|
||||
],
|
||||
[
|
||||
"Bredde"
|
||||
],
|
||||
[
|
||||
"3,0 m"
|
||||
]
|
||||
],
|
||||
"table": [
|
||||
[
|
||||
"Bredde",
|
||||
"3,0 m"
|
||||
]
|
||||
],
|
||||
"text_box": [
|
||||
[
|
||||
"Merk: kravet gjelder ikke midlertidige bruer."
|
||||
]
|
||||
]
|
||||
},
|
||||
"witness": "docx zip xml"
|
||||
}
|
||||
},
|
||||
"files": {
|
||||
"graphics/figur-84-1.png": {
|
||||
"pointed_at_by": [
|
||||
"figur.html",
|
||||
"prosess-84-sts.xml",
|
||||
"prosess-84-web.html"
|
||||
"prosess-84-web.html",
|
||||
"sts-rikt.xml"
|
||||
]
|
||||
},
|
||||
"graphics/tabell-84-2.png": {
|
||||
"pointed_at_by": [
|
||||
"notat.md",
|
||||
"prosess-84-sts.xml",
|
||||
"prosess-84-web.html"
|
||||
"prosess-84-web.html",
|
||||
"side.htm"
|
||||
]
|
||||
}
|
||||
},
|
||||
|
|
|
|||
327
tests/fixtures/accounting/make_accounting_fixtures.py
vendored
Normal file
327
tests/fixtures/accounting/make_accounting_fixtures.py
vendored
Normal file
|
|
@ -0,0 +1,327 @@
|
|||
"""The accounting corpus's SECOND document per format: the elements the
|
||||
witness could not see until 2026-09-18.
|
||||
|
||||
An independent review of the gate found that the witness -- and therefore the
|
||||
whole accounting, because rows 2 and 3 require the build's inventory to EQUAL
|
||||
it -- counted no header, no comment, no speaker note, no hidden sheet, no
|
||||
formula, no citation and no figure caption. What nothing counts, nothing can
|
||||
lose visibly. It also found that 20 of 63 element types had a count of ZERO in
|
||||
their only fixture, so six of seven witness mutants survived.
|
||||
|
||||
These documents are written BY HAND, part by part, for the reason the XML
|
||||
fixtures state: a library that writes and then reads its own format proves
|
||||
only that it agrees with itself. Every count they carry is written down in
|
||||
`tests/fixtures/README.md` by a person reading these strings, not by running
|
||||
the witness over them.
|
||||
|
||||
python3 tests/fixtures/accounting/make_accounting_fixtures.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import io
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
CORPUS = HERE / "corpus"
|
||||
|
||||
_ZIP_DATE = (2020, 1, 1, 0, 0, 0)
|
||||
_XML = '<?xml version="1.0" encoding="UTF-8" standalone="yes"?>'
|
||||
|
||||
_W = 'xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"'
|
||||
_A = 'xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main"'
|
||||
_P = 'xmlns:p="http://schemas.openxmlformats.org/presentationml/2006/main"'
|
||||
_S = 'xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main"'
|
||||
|
||||
|
||||
def build_zip(parts: dict[str, str | bytes]) -> bytes:
|
||||
"""Zip the parts with a fixed timestamp, so a regeneration that changes
|
||||
nothing leaves the bytes alone and `git diff --quiet` stays a real check."""
|
||||
out = io.BytesIO()
|
||||
with zipfile.ZipFile(out, "w", compression=zipfile.ZIP_DEFLATED) as archive:
|
||||
for name, payload in parts.items():
|
||||
info = zipfile.ZipInfo(name, date_time=_ZIP_DATE)
|
||||
info.compress_type = zipfile.ZIP_DEFLATED
|
||||
archive.writestr(info, payload)
|
||||
return out.getvalue()
|
||||
|
||||
|
||||
# --- docx: a header, a footer, a comment, an endnote and a text box ----------
|
||||
|
||||
_DOCX_BODY = (
|
||||
'<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr>'
|
||||
"<w:r><w:t>Krav til gangbruer</w:t></w:r></w:p>"
|
||||
"<w:p><w:r><w:t>Gangbruer skal ha rekkverk paa begge sider.</w:t></w:r></w:p>"
|
||||
"<w:tbl><w:tr>"
|
||||
"<w:tc><w:p><w:r><w:t>Bredde</w:t></w:r></w:p></w:tc>"
|
||||
"<w:tc><w:p><w:r><w:t>3,0 m</w:t></w:r></w:p></w:tc>"
|
||||
"</w:tr></w:tbl>"
|
||||
"<w:p><w:r><w:pict><w:txbxContent>"
|
||||
"<w:p><w:r><w:t>Merk: kravet gjelder ikke midlertidige bruer.</w:t></w:r></w:p>"
|
||||
"</w:txbxContent></w:pict></w:r></w:p>"
|
||||
)
|
||||
|
||||
_DOCX_PARTS: dict[str, str | bytes] = {
|
||||
"[Content_Types].xml": _XML
|
||||
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
||||
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
||||
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.'
|
||||
+ 'relationships+xml"/>'
|
||||
+ '<Override PartName="/word/document.xml" ContentType="application/vnd.'
|
||||
+ 'openxmlformats-officedocument.wordprocessingml.document.main+xml"/>'
|
||||
+ '<Override PartName="/word/styles.xml" ContentType="application/vnd.'
|
||||
+ 'openxmlformats-officedocument.wordprocessingml.styles+xml"/>'
|
||||
+ "</Types>",
|
||||
"_rels/.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships/officeDocument" Target="word/document.xml"/></Relationships>',
|
||||
"word/_rels/document.xml.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships/styles" Target="styles.xml"/></Relationships>',
|
||||
"word/styles.xml": _XML
|
||||
+ f"<w:styles {_W}>"
|
||||
+ '<w:style w:type="paragraph" w:styleId="Heading1"><w:name w:val="heading 1"/></w:style>'
|
||||
+ "</w:styles>",
|
||||
"word/document.xml": _XML + f"<w:document {_W}><w:body>{_DOCX_BODY}</w:body></w:document>",
|
||||
"word/header1.xml": _XML
|
||||
+ f"<w:hdr {_W}><w:p><w:r><w:t>Utkast - gjelder ikke etter 2026-01-01</w:t></w:r></w:p>"
|
||||
+ "</w:hdr>",
|
||||
"word/footer1.xml": _XML
|
||||
+ f"<w:ftr {_W}><w:p><w:r><w:t>Statens vegvesen, side 1</w:t></w:r></w:p></w:ftr>",
|
||||
"word/comments.xml": _XML
|
||||
+ f"<w:comments {_W}>"
|
||||
+ '<w:comment w:id="1"><w:p><w:r><w:t>Unntak: gjelder IKKE gangbruer i tunnel.'
|
||||
+ "</w:t></w:r></w:p></w:comment></w:comments>",
|
||||
"word/footnotes.xml": _XML
|
||||
+ f"<w:footnotes {_W}>"
|
||||
+ '<w:footnote w:id="0"><w:p><w:r><w:t>separator</w:t></w:r></w:p></w:footnote>'
|
||||
+ '<w:footnote w:id="2"><w:p><w:r><w:t>Se haandbok N400 kapittel 5.</w:t></w:r></w:p>'
|
||||
+ "</w:footnote></w:footnotes>",
|
||||
"word/endnotes.xml": _XML
|
||||
+ f"<w:endnotes {_W}>"
|
||||
+ '<w:endnote w:id="0"><w:p><w:r><w:t>separator</w:t></w:r></w:p></w:endnote>'
|
||||
+ '<w:endnote w:id="3"><w:p><w:r><w:t>Kravet ble skjerpet i 2024.</w:t></w:r></w:p>'
|
||||
+ "</w:endnote></w:endnotes>",
|
||||
}
|
||||
|
||||
|
||||
# --- pptx: a speaker note and a hidden slide --------------------------------
|
||||
|
||||
|
||||
def _slide(title: str, body: str, *, hidden: bool = False) -> str:
|
||||
show = ' show="0"' if hidden else ""
|
||||
return (
|
||||
_XML
|
||||
+ f"<p:sld {_P} {_A}{show}><p:cSld><p:spTree>"
|
||||
+ '<p:sp><p:nvSpPr><p:nvPr><p:ph type="title"/></p:nvPr></p:nvSpPr>'
|
||||
+ f"<p:txBody><a:p><a:r><a:t>{title}</a:t></a:r></a:p></p:txBody></p:sp>"
|
||||
+ "<p:sp><p:nvSpPr><p:nvPr/></p:nvSpPr>"
|
||||
+ f"<p:txBody><a:p><a:r><a:t>{body}</a:t></a:r></a:p></p:txBody></p:sp>"
|
||||
+ "<p:graphicFrame><a:graphic><a:graphicData><a:tbl><a:tr>"
|
||||
+ "<a:tc><a:txBody><a:p><a:r><a:t>Post</a:t></a:r></a:p></a:txBody></a:tc>"
|
||||
+ "<a:tc><a:txBody><a:p><a:r><a:t>84.1</a:t></a:r></a:p></a:txBody></a:tc>"
|
||||
+ "</a:tr></a:tbl></a:graphicData></a:graphic></p:graphicFrame>"
|
||||
+ "</p:spTree></p:cSld></p:sld>"
|
||||
)
|
||||
|
||||
|
||||
_PPTX_PARTS: dict[str, str | bytes] = {
|
||||
"[Content_Types].xml": _XML
|
||||
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
||||
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
||||
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.'
|
||||
+ 'relationships+xml"/>'
|
||||
+ '<Override PartName="/ppt/presentation.xml" ContentType="application/vnd.'
|
||||
+ 'openxmlformats-officedocument.presentationml.presentation.main+xml"/>'
|
||||
+ "".join(
|
||||
f'<Override PartName="/ppt/slides/slide{n}.xml" ContentType="application/vnd.'
|
||||
f'openxmlformats-officedocument.presentationml.slide+xml"/>'
|
||||
for n in (1, 2)
|
||||
)
|
||||
+ "</Types>",
|
||||
"_rels/.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships/officeDocument" Target="ppt/presentation.xml"/></Relationships>',
|
||||
"ppt/presentation.xml": _XML
|
||||
+ f"<p:presentation {_P}><p:sldIdLst>"
|
||||
+ '<p:sldId id="256" r:id="rId1"/><p:sldId id="257" r:id="rId2"/>'
|
||||
+ "</p:sldIdLst></p:presentation>",
|
||||
"ppt/slides/slide1.xml": _slide("Prosess 84 Konstruksjoner", "Toleranser er gitt i tabell."),
|
||||
"ppt/slides/slide2.xml": _slide("Utgaatt lysbilde", "Ikke vis dette.", hidden=True),
|
||||
"ppt/notesSlides/notesSlide1.xml": _XML
|
||||
+ f"<p:notes {_P} {_A}><p:cSld><p:spTree><p:sp><p:txBody>"
|
||||
+ "<a:p><a:r><a:t>Husk aa nevne at toleranseklassen er skjerpet.</a:t></a:r></a:p>"
|
||||
+ "</p:txBody></p:sp></p:spTree></p:cSld></p:notes>",
|
||||
}
|
||||
|
||||
|
||||
# --- xlsx: a hidden sheet and a formula --------------------------------------
|
||||
|
||||
_XLSX_STRINGS = ["Post", "Enhet", "Mengde", "Sum", "Internt", "Kladd"]
|
||||
|
||||
_XLSX_PARTS: dict[str, str | bytes] = {
|
||||
"[Content_Types].xml": _XML
|
||||
+ '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
|
||||
+ '<Default Extension="xml" ContentType="application/xml"/>'
|
||||
+ '<Default Extension="rels" ContentType="application/vnd.openxmlformats-package.'
|
||||
+ 'relationships+xml"/>'
|
||||
+ '<Override PartName="/xl/workbook.xml" ContentType="application/vnd.openxmlformats-'
|
||||
+ 'officedocument.spreadsheetml.sheet.main+xml"/>'
|
||||
+ "</Types>",
|
||||
"_rels/.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships/officeDocument" Target="xl/workbook.xml"/></Relationships>',
|
||||
"xl/workbook.xml": _XML
|
||||
+ f'<workbook {_S} xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships"><sheets>'
|
||||
+ '<sheet name="Mengder" sheetId="1" r:id="rId1"/>'
|
||||
+ '<sheet name="Internt" sheetId="2" state="hidden" r:id="rId2"/>'
|
||||
+ "</sheets></workbook>",
|
||||
"xl/_rels/workbook.xml.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships/worksheet" Target="worksheets/sheet1.xml"/>'
|
||||
+ '<Relationship Id="rId2" Type="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships/worksheet" Target="worksheets/sheet2.xml"/></Relationships>',
|
||||
"xl/sharedStrings.xml": _XML
|
||||
+ f'<sst {_S} count="{len(_XLSX_STRINGS)}" uniqueCount="{len(_XLSX_STRINGS)}">'
|
||||
+ "".join(f"<si><t>{value}</t></si>" for value in _XLSX_STRINGS)
|
||||
+ "</sst>",
|
||||
"xl/worksheets/sheet1.xml": _XML
|
||||
+ f'<worksheet {_S}><dimension ref="A1:C3"/><sheetData>'
|
||||
+ '<row r="1"><c r="A1" t="s"><v>0</v></c><c r="B1" t="s"><v>1</v></c>'
|
||||
+ '<c r="C1" t="s"><v>2</v></c></row>'
|
||||
+ '<row r="2"><c r="A2" t="s"><v>3</v></c><c r="B2"><v>12</v></c>'
|
||||
+ '<c r="C2"><f>B2*2</f><v>24</v></c></row>'
|
||||
+ '</sheetData><drawing r:id="rId1" xmlns:r="http://schemas.openxmlformats.org/'
|
||||
+ 'officeDocument/2006/relationships"/></worksheet>',
|
||||
"xl/drawings/drawing1.xml": _XML
|
||||
+ '<xdr:wsDr xmlns:xdr="http://schemas.openxmlformats.org/drawingml/2006/'
|
||||
+ 'spreadsheetDrawing" xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main">'
|
||||
+ "<xdr:twoCellAnchor><xdr:pic><xdr:nvPicPr>"
|
||||
+ '<xdr:cNvPr id="1" name="Diagram"/>'
|
||||
+ "</xdr:nvPicPr></xdr:pic></xdr:twoCellAnchor></xdr:wsDr>",
|
||||
"xl/worksheets/_rels/sheet1.xml.rels": _XML
|
||||
+ '<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">'
|
||||
+ '<Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/'
|
||||
+ 'relationships/drawing" Target="../drawings/drawing1.xml"/></Relationships>',
|
||||
"xl/worksheets/sheet2.xml": _XML
|
||||
+ f'<worksheet {_S}><dimension ref="A1:A1"/><sheetData>'
|
||||
+ '<row r="1"><c r="A1" t="s"><v>5</v></c></row>'
|
||||
+ "</sheetData></worksheet>",
|
||||
}
|
||||
|
||||
|
||||
# --- odt: a header, a list, an annotation and a picture ----------------------
|
||||
|
||||
_ODT_NS = (
|
||||
'xmlns:office="urn:oasis:names:tc:opendocument:xmlns:office:1.0" '
|
||||
'xmlns:text="urn:oasis:names:tc:opendocument:xmlns:text:1.0" '
|
||||
'xmlns:table="urn:oasis:names:tc:opendocument:xmlns:table:1.0" '
|
||||
'xmlns:draw="urn:oasis:names:tc:opendocument:xmlns:drawing:1.0" '
|
||||
'xmlns:style="urn:oasis:names:tc:opendocument:xmlns:style:1.0" '
|
||||
'xmlns:xlink="http://www.w3.org/1999/xlink"'
|
||||
)
|
||||
|
||||
_ODT_PARTS: dict[str, str | bytes] = {
|
||||
"mimetype": "application/vnd.oasis.opendocument.text",
|
||||
"META-INF/manifest.xml": _XML
|
||||
+ '<manifest:manifest xmlns:manifest="urn:oasis:names:tc:opendocument:xmlns:manifest:1.0">'
|
||||
+ '<manifest:file-entry manifest:full-path="/" manifest:media-type="application/vnd.oasis.'
|
||||
+ 'opendocument.text"/>'
|
||||
+ '<manifest:file-entry manifest:full-path="content.xml" manifest:media-type="text/xml"/>'
|
||||
+ "</manifest:manifest>",
|
||||
"content.xml": _XML
|
||||
+ f"<office:document-content {_ODT_NS}><office:body><office:text>"
|
||||
+ '<text:h text:outline-level="1">Drift av gangbruer</text:h>'
|
||||
+ "<text:p>Gangbruer inspiseres hvert aar.</text:p>"
|
||||
+ "<text:list><text:list-item><text:p>Rekkverk</text:p></text:list-item>"
|
||||
+ "<text:list-item><text:p>Dekke</text:p></text:list-item></text:list>"
|
||||
+ "<text:p>Se figuren under."
|
||||
+ '<draw:frame><draw:image xlink:href="graphics/figur-84-1.png"/></draw:frame></text:p>'
|
||||
+ "<office:annotation><text:p>Sjekk denne mot N400 foer utsendelse.</text:p>"
|
||||
+ "</office:annotation>"
|
||||
+ "<table:table><table:table-row>"
|
||||
+ "<table:table-cell><text:p>Type</text:p></table:table-cell>"
|
||||
+ "<table:table-cell><text:p>Gangbru</text:p></table:table-cell>"
|
||||
+ "</table:table-row></table:table>"
|
||||
+ "</office:text></office:body></office:document-content>",
|
||||
"styles.xml": _XML
|
||||
+ f"<office:document-styles {_ODT_NS}><office:master-styles>"
|
||||
+ '<style:master-page style:name="Standard">'
|
||||
+ "<style:header><text:p>Intern arbeidsversjon</text:p></style:header>"
|
||||
+ "<style:footer><text:p>Vegdirektoratet</text:p></style:footer>"
|
||||
+ "</style:master-page></office:master-styles></office:document-styles>",
|
||||
}
|
||||
|
||||
|
||||
# --- rtf: a picture ----------------------------------------------------------
|
||||
|
||||
_RTF = (
|
||||
r"{\rtf1\ansi\deff0{\fonttbl{\f0 Times New Roman;}}"
|
||||
r"\pard Figur 84-1 viser prinsippet.\par"
|
||||
r"\pard{\pict\pngblip\picw16\pich16 89504e470d0a1a0a}\par"
|
||||
"}"
|
||||
)
|
||||
|
||||
|
||||
# --- html: a picture and a caption -------------------------------------------
|
||||
|
||||
_HTML = """<!DOCTYPE html>
|
||||
<html lang="no">
|
||||
<head><title>Figur 84-1</title></head>
|
||||
<body>
|
||||
<h1>Figur 84-1</h1>
|
||||
<p>Prinsippet for toleranseklasser.</p>
|
||||
<img src="graphics/figur-84-1.png" alt="Prinsippskisse">
|
||||
<table><tr><th>Klasse</th><th>Avvik</th></tr><tr><td>A</td><td>5 mm</td></tr></table>
|
||||
<ul><li>Klasse A</li><li>Klasse B</li></ul>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
|
||||
|
||||
# --- sts: a citation, a formula, a figure with a caption, a table, a footnote -
|
||||
|
||||
_STS = """<standard>
|
||||
<front><std-ident><doc-number>R762</doc-number><year>2025</year></std-ident></front>
|
||||
<body>
|
||||
<sec><label>85</label><title>Vegdekker</title>
|
||||
<p>Dekket skal ha jevnhet etter <mixed-citation>NS-EN 13036-1:2010</mixed-citation>.</p>
|
||||
<p>Kravet regnes som <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>IRI</mml:mi><mml:mo><</mml:mo><mml:mn>2</mml:mn></mml:math>.</p>
|
||||
<fig><label>Figur 85-1</label><caption><p>Maalepunkter langs vegbanen.</p></caption>
|
||||
<graphic xlink:href="figur-84-1.png" xmlns:xlink="http://www.w3.org/1999/xlink"/></fig>
|
||||
<table-wrap><label>Tabell 85-1</label>
|
||||
<table><tr><th>Klasse</th><th>IRI</th></tr><tr><td>1</td><td>1,5</td></tr></table>
|
||||
</table-wrap>
|
||||
<list><list-item><p>Maales hvert 20. meter.</p></list-item></list>
|
||||
<fn><p>Gjelder ikke gang- og sykkelveger.</p></fn>
|
||||
</sec>
|
||||
</body>
|
||||
</standard>
|
||||
"""
|
||||
|
||||
|
||||
def main() -> int:
|
||||
written = {
|
||||
"topptekst-og-kommentar.docx": build_zip(_DOCX_PARTS),
|
||||
"notater-og-skjult.pptx": build_zip(_PPTX_PARTS),
|
||||
"skjult-ark-og-formel.xlsx": build_zip(_XLSX_PARTS),
|
||||
"liste-og-bilde.odt": build_zip(_ODT_PARTS),
|
||||
"bilde.rtf": _RTF.encode("latin-1"),
|
||||
"figur.html": _HTML.encode("utf-8"),
|
||||
"sts-rikt.xml": _STS.encode("utf-8"),
|
||||
}
|
||||
for name, payload in written.items():
|
||||
(CORPUS / name).write_bytes(payload)
|
||||
print(f"wrote {name}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
|
|
@ -102,7 +102,10 @@ def test_the_fixture_corpus_covers_every_readme_file_type() -> None:
|
|||
assert {entry["suffix"] for entry in inventory["documents"].values()} == set(TABLE)
|
||||
|
||||
|
||||
# Counted by hand from the fixture bytes, not by the witness.
|
||||
# Counted by hand from the fixture bytes, not by the witness. Seven of the
|
||||
# thirteen documents were added 2026-09-18 because 20 of 63 element types had
|
||||
# a count of ZERO in their only fixture, and a type that cannot appear cannot
|
||||
# be lost visibly either (independent review, M-3).
|
||||
HAND_COUNTS = {
|
||||
"notat.md": {
|
||||
"code_block": 1,
|
||||
|
|
@ -112,13 +115,26 @@ HAND_COUNTS = {
|
|||
"table": 1,
|
||||
"table_row": 2,
|
||||
},
|
||||
"side.htm": {"cell": 4, "heading": 2, "image": 0, "list_item": 2, "paragraph": 1, "table": 1},
|
||||
"side.htm": {"cell": 4, "heading": 2, "image": 1, "list_item": 2, "paragraph": 1, "table": 1},
|
||||
"figur.html": {
|
||||
"cell": 4,
|
||||
"heading": 1,
|
||||
"image": 1,
|
||||
"list_item": 2,
|
||||
"paragraph": 1,
|
||||
"table": 1,
|
||||
},
|
||||
"krav-rikt-tekstformat.rtf": {"cell": 56, "image": 0, "paragraph": 3, "table_row": 24},
|
||||
"bilde.rtf": {"cell": 0, "image": 1, "paragraph": 2, "table_row": 0},
|
||||
"prosess-84-sts.xml": {
|
||||
"cell": 0,
|
||||
"citation": 0,
|
||||
"figure": 0,
|
||||
"figure_caption": 0,
|
||||
"footnote": 0,
|
||||
"image": 2,
|
||||
"list_item": 0,
|
||||
"math": 0,
|
||||
"paragraph": 2,
|
||||
"section": 2,
|
||||
"section_label": 2,
|
||||
|
|
@ -126,9 +142,140 @@ HAND_COUNTS = {
|
|||
"table_label": 0,
|
||||
"title": 2,
|
||||
},
|
||||
"sts-rikt.xml": {
|
||||
"cell": 4,
|
||||
"citation": 1,
|
||||
"figure": 1,
|
||||
"figure_caption": 1,
|
||||
"footnote": 1,
|
||||
"image": 1,
|
||||
"list_item": 1,
|
||||
"math": 1,
|
||||
"paragraph": 5,
|
||||
"section": 1,
|
||||
"section_label": 1,
|
||||
"table": 1,
|
||||
"table_label": 1,
|
||||
"title": 1,
|
||||
},
|
||||
"topptekst-og-kommentar.docx": {
|
||||
"cell": 2,
|
||||
"comment": 1,
|
||||
"endnote": 1,
|
||||
"footnote": 1,
|
||||
"header_footer": 2,
|
||||
"heading": 1,
|
||||
"image": 0,
|
||||
"paragraph": 3,
|
||||
"table": 1,
|
||||
"text_box": 1,
|
||||
},
|
||||
"notater-og-skjult.pptx": {
|
||||
"cell": 4,
|
||||
"hidden_slide": 1,
|
||||
"image": 0,
|
||||
"note": 1,
|
||||
"paragraph": 2,
|
||||
"slide": 1,
|
||||
"table": 2,
|
||||
"title": 2,
|
||||
},
|
||||
"skjult-ark-og-formel.xlsx": {
|
||||
"cell": 7,
|
||||
"formula": 1,
|
||||
"hidden_sheet": 1,
|
||||
"image": 1,
|
||||
"row": 3,
|
||||
"sheet": 1,
|
||||
},
|
||||
"logg.txt": {"line": 4, "paragraph": 3},
|
||||
"mengder.csv": {"cell": 6, "header_cell": 3, "row": 2},
|
||||
"parametre.json": {"key": 6, "value": 6},
|
||||
"prosess-84-tabell.pdf": {"image": 2, "page": 1},
|
||||
"prosess-84-web.html": {
|
||||
"cell": 0,
|
||||
"heading": 1,
|
||||
"image": 3,
|
||||
"list_item": 0,
|
||||
"paragraph": 3,
|
||||
"table": 0,
|
||||
},
|
||||
"prosess-84-notat.docx": {
|
||||
"cell": 0,
|
||||
"comment": 0,
|
||||
"endnote": 0,
|
||||
"footnote": 0,
|
||||
"header_footer": 0,
|
||||
"heading": 0,
|
||||
"image": 1,
|
||||
"paragraph": 2,
|
||||
"table": 0,
|
||||
"text_box": 0,
|
||||
},
|
||||
"prosess-84-presentasjon.pptx": {
|
||||
"cell": 0,
|
||||
"hidden_slide": 0,
|
||||
"image": 1,
|
||||
"note": 0,
|
||||
"paragraph": 0,
|
||||
"slide": 1,
|
||||
"table": 0,
|
||||
"title": 1,
|
||||
},
|
||||
"prisark.xlsx": {
|
||||
"cell": 18,
|
||||
"formula": 0,
|
||||
"hidden_sheet": 0,
|
||||
"image": 0,
|
||||
"row": 9,
|
||||
"sheet": 2,
|
||||
},
|
||||
"krav-tekstdokument.odt": {
|
||||
"annotation": 0,
|
||||
"cell": 56,
|
||||
"header_footer": 0,
|
||||
"heading": 1,
|
||||
"image": 0,
|
||||
"list_item": 0,
|
||||
"paragraph": 2,
|
||||
"table": 2,
|
||||
},
|
||||
"liste-og-bilde.odt": {
|
||||
"annotation": 1,
|
||||
"cell": 2,
|
||||
"header_footer": 2,
|
||||
"heading": 1,
|
||||
"image": 1,
|
||||
"list_item": 2,
|
||||
"paragraph": 4,
|
||||
"table": 1,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def test_every_element_type_the_witness_counts_occurs_in_a_fixture() -> None:
|
||||
"""M-3: 20 of 63 element types had a count of 0 in their only fixture, so
|
||||
six of seven witness mutants survived -- a witness cannot be wrong about
|
||||
something it never sees."""
|
||||
pytest.importorskip("pdfplumber")
|
||||
inventory = gate.load_inventory(gate.INVENTORY)
|
||||
seen: dict[str, int] = {}
|
||||
for entry in inventory["documents"].values():
|
||||
for role, number in entry["elements"].items():
|
||||
key = f"{entry['suffix']} {role}"
|
||||
seen[key] = seen.get(key, 0) + number
|
||||
absent = sorted(key for key, number in seen.items() if number == 0)
|
||||
assert absent == [], f"element types with no occurrence in any fixture: {absent}"
|
||||
|
||||
|
||||
def test_the_hand_counts_cover_every_document_of_the_corpus() -> None:
|
||||
"""Row 1's fasit is the witness's own output, so a hand count is the only
|
||||
thing in the loop that the witness did not produce. Four of thirteen
|
||||
documents had one; every document has one now."""
|
||||
inventory = gate.load_inventory(gate.INVENTORY)
|
||||
assert sorted(HAND_COUNTS) == sorted(inventory["documents"])
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(HAND_COUNTS))
|
||||
def test_the_witness_matches_a_hand_count(name: str) -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.CORPUS / name)
|
||||
|
|
@ -162,6 +309,67 @@ def test_the_witness_refuses_a_doctype() -> None:
|
|||
witness.count_sts_xml(b'<!DOCTYPE x [<!ENTITY a "b">]><standard/>')
|
||||
|
||||
|
||||
# --- M-1 / M-2: what the witness could not see -------------------------------
|
||||
#
|
||||
# Written RED 2026-09-18. Rows 2 and 3 require the build's inventory to EQUAL
|
||||
# the witness's, so what the witness does not count, nothing can account for:
|
||||
# an independent review put a header and a comment in a docx, measured 0 of
|
||||
# either in the bundle, and the accounting still read "2 of 2 carried".
|
||||
|
||||
|
||||
def test_the_docx_witness_counts_a_header_a_comment_and_a_text_box() -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.CORPUS / "topptekst-og-kommentar.docx")
|
||||
for role in ("header_footer", "comment", "endnote", "text_box"):
|
||||
assert inventory.elements[role] > 0, role
|
||||
assert "Utkast" in " ".join(
|
||||
piece for pieces in inventory.texts["header_footer"] for piece in pieces
|
||||
)
|
||||
|
||||
|
||||
def test_the_pptx_witness_counts_notes_and_does_not_call_a_hidden_slide_a_slide() -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.CORPUS / "notater-og-skjult.pptx")
|
||||
assert inventory.elements["note"] > 0
|
||||
assert inventory.elements["hidden_slide"] == 1
|
||||
|
||||
|
||||
def test_the_xlsx_witness_counts_a_formula_and_a_hidden_sheet() -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.CORPUS / "skjult-ark-og-formel.xlsx")
|
||||
assert inventory.elements["formula"] > 0
|
||||
assert inventory.elements["hidden_sheet"] == 1
|
||||
|
||||
|
||||
def test_the_odt_witness_counts_a_header_and_does_not_call_an_annotation_prose() -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.CORPUS / "liste-og-bilde.odt")
|
||||
assert inventory.elements["header_footer"] > 0
|
||||
assert inventory.elements["annotation"] == 1
|
||||
|
||||
|
||||
def test_the_sts_witness_counts_citations_formulas_and_figure_captions() -> None:
|
||||
inventory = witness.witness_file(gate.CORPUS, gate.CORPUS / "sts-rikt.xml")
|
||||
for role in ("citation", "math", "figure", "figure_caption"):
|
||||
assert inventory.elements[role] > 0, role
|
||||
|
||||
|
||||
def test_the_two_sts_role_maps_are_written_twice_and_not_shared() -> None:
|
||||
"""M-2: both STS witnesses went through ONE `_sts_role`, so row 5 could
|
||||
never see a hole in it. Two maps, each written for its own delivery."""
|
||||
assert witness._sts_role_xml is not witness._sts_role_json
|
||||
assert "_sts_role_json" not in witness._sts_role_xml.__code__.co_names
|
||||
assert "_sts_role_xml" not in witness._sts_role_json.__code__.co_names
|
||||
|
||||
|
||||
def test_every_witnessed_type_says_what_it_does_not_count() -> None:
|
||||
assert set(witness.NOT_COUNTED) == set(witness.WITNESSED_SUFFIXES)
|
||||
assert all(witness.NOT_COUNTED[suffix] for suffix in witness.WITNESSED_SUFFIXES)
|
||||
|
||||
|
||||
def test_the_gate_prints_what_each_witness_does_not_count() -> None:
|
||||
rendered = gate.render([])
|
||||
assert "not counted" in rendered
|
||||
for suffix in witness.WITNESSED_SUFFIXES:
|
||||
assert suffix in rendered
|
||||
|
||||
|
||||
# --- 2. every row can go both ways -------------------------------------------
|
||||
|
||||
|
||||
|
|
@ -628,18 +836,43 @@ def test_the_door_exists() -> None:
|
|||
assert gate.door_available()
|
||||
|
||||
|
||||
def test_the_real_gate_is_green_on_every_fixture_row(real_rows: list[gate.Row]) -> None:
|
||||
def test_the_real_gate_names_what_the_build_does_not_account_for(
|
||||
real_rows: list[gate.Row],
|
||||
) -> None:
|
||||
"""Rows 1-5 against the real `okf build`; row 6 needs R761 and is skipped
|
||||
here. Written red at 0b00de4 (rows 2, 3, 4 red), green once the build
|
||||
accounts for every element."""
|
||||
here.
|
||||
|
||||
Rows 2 and 3 were GREEN at `864570b` and are RED now, and that is the
|
||||
hardening working rather than a regression: the witness counts thirteen
|
||||
classes of content it could not see before, the build accounts for none
|
||||
of them, and a number nobody counts is a loss nobody can report. The
|
||||
named classes are the raw material for the next capability order.
|
||||
"""
|
||||
assert [(r.number, r.status) for r in real_rows] == [
|
||||
(1, gate.GREEN),
|
||||
(2, gate.GREEN),
|
||||
(3, gate.GREEN),
|
||||
(2, gate.RED),
|
||||
(3, gate.RED),
|
||||
(4, gate.GREEN),
|
||||
(5, gate.GREEN),
|
||||
(6, gate.SKIPPED),
|
||||
]
|
||||
assert real_rows[2].reason.startswith(
|
||||
"u = 0 unaccounted, d = 0 double-booked, 0 booked carried and not in the bundle"
|
||||
)
|
||||
unaccounted = " ".join(real_rows[2].details)
|
||||
for element in (
|
||||
"annotation",
|
||||
"citation",
|
||||
"comment",
|
||||
"endnote",
|
||||
"figure",
|
||||
"figure_caption",
|
||||
"formula",
|
||||
"header_footer",
|
||||
"hidden_sheet",
|
||||
"hidden_slide",
|
||||
"math",
|
||||
"note",
|
||||
"text_box",
|
||||
):
|
||||
assert f"{element}: 0 booked of" in unaccounted, element
|
||||
# And no false red from the gate's own reading: every element the build
|
||||
# DOES book as carried was found in the bundle.
|
||||
assert "0 claimed and not found" in real_rows[2].details[0]
|
||||
|
|
|
|||
|
|
@ -980,6 +980,13 @@ def render(rows: list[Row]) -> str:
|
|||
lines.append(f"exceptions approved ({APPROVED_ON}), and none moves a denominator:")
|
||||
for suffix, element in sorted(APPROVED_EXCEPTIONS) or [("(none)", "")]:
|
||||
lines.append(f" - {suffix} {element}")
|
||||
lines += [
|
||||
"",
|
||||
"not counted by any witness -- what no row here can see (per file type):",
|
||||
]
|
||||
for suffix in sorted(witness.NOT_COUNTED):
|
||||
for item in witness.NOT_COUNTED[suffix]:
|
||||
lines.append(f" - {suffix}: {item}")
|
||||
failing = [str(r.number) for r in rows if r.fails]
|
||||
lines.append("")
|
||||
lines.append(
|
||||
|
|
|
|||
|
|
@ -141,6 +141,12 @@ class Inventory:
|
|||
}
|
||||
|
||||
|
||||
def _normal(text: str) -> str:
|
||||
"""One space between words: a reader re-wraps, and a piece has to survive
|
||||
that to be looked for at all."""
|
||||
return " ".join(text.split())
|
||||
|
||||
|
||||
def _local(tag: str) -> str:
|
||||
return tag.rsplit("}", 1)[-1] if "}" in tag else tag.split(":")[-1]
|
||||
|
||||
|
|
@ -421,33 +427,33 @@ class _HtmlCounter(HTMLParser):
|
|||
# --- xml / sts ---------------------------------------------------------------
|
||||
|
||||
STS_ROLES = (
|
||||
"section",
|
||||
"title",
|
||||
"section_label",
|
||||
"cell",
|
||||
"citation",
|
||||
"figure",
|
||||
"figure_caption",
|
||||
"footnote",
|
||||
"image",
|
||||
"list_item",
|
||||
"math",
|
||||
"paragraph",
|
||||
"section",
|
||||
"section_label",
|
||||
"table",
|
||||
"table_label",
|
||||
"cell",
|
||||
"list_item",
|
||||
"image",
|
||||
"footnote",
|
||||
"title",
|
||||
)
|
||||
|
||||
|
||||
def _sts_role(tag: str, parent: str | None, grandparent: str | None) -> str | None:
|
||||
"""The one mapping from an STS element to its accounting role.
|
||||
def _sts_role_xml(tag: str, parent: str | None, grandparent: str | None) -> str | None:
|
||||
"""An STS XML element's accounting role.
|
||||
|
||||
Used by BOTH STS witnesses, and the role is the unit, not the tag, because
|
||||
the publisher's two deliveries of one document place the same text
|
||||
differently (measured on R761 Prosesskoden:2025, 2026-09-17):
|
||||
Written for the XML delivery ALONE. Until 2026-09-18 one function served
|
||||
both deliveries, so row 5 -- "two witnesses agree" -- could not see a hole
|
||||
in it: a role missing here was missing there, and the two agreed on a
|
||||
number neither of them should have produced (independent review, M-2).
|
||||
|
||||
- a section's label: XML `sec/label` on 7 714 sections; JSON `sec/label`
|
||||
on 4 954 and `sec/title/label` on the 2 760 that carry a title.
|
||||
- a table's label: XML `table-wrap/label` (10); JSON
|
||||
`table-wrap/table/caption` (10).
|
||||
|
||||
Counted by tag, the two witnesses disagree by 2 760 and by 10 on text
|
||||
both of them carry.
|
||||
In this delivery a section's label is `sec/label` and a table's label is
|
||||
`table-wrap/label`.
|
||||
"""
|
||||
if tag == "sec":
|
||||
return "section"
|
||||
|
|
@ -459,6 +465,14 @@ def _sts_role(tag: str, parent: str | None, grandparent: str | None) -> str | No
|
|||
return "table_label"
|
||||
if tag == "caption" and parent == "table" and grandparent == "table-wrap":
|
||||
return "table_label"
|
||||
if tag == "caption" and parent == "fig":
|
||||
return "figure_caption"
|
||||
if tag == "fig":
|
||||
return "figure"
|
||||
if tag == "mixed-citation":
|
||||
return "citation"
|
||||
if tag == "math":
|
||||
return "math"
|
||||
if tag == "p":
|
||||
return "paragraph"
|
||||
if tag == "table-wrap":
|
||||
|
|
@ -474,8 +488,54 @@ def _sts_role(tag: str, parent: str | None, grandparent: str | None) -> str | No
|
|||
return None
|
||||
|
||||
|
||||
def _normal(text: str) -> str:
|
||||
return " ".join(text.split())
|
||||
def _sts_role_json(tag: str, parent: str | None, grandparent: str | None) -> str | None:
|
||||
"""The same roles, read from the publisher's JSON node tree.
|
||||
|
||||
Written apart from the XML map, because the publisher's two deliveries of
|
||||
ONE document place the same text differently (measured on R761
|
||||
Prosesskoden:2025, 2026-09-17):
|
||||
|
||||
- a section's label: XML `sec/label` on 7 714 sections; JSON `sec/label`
|
||||
on 4 954 and `sec/title/label` on the 2 760 that carry a title.
|
||||
- a table's label: XML `table-wrap/label` (10); JSON
|
||||
`table-wrap/table/caption` (10).
|
||||
|
||||
Counted by tag alone, the two witnesses disagree by 2 760 and by 10 on
|
||||
text both of them carry. The role, not the tag, is the unit.
|
||||
"""
|
||||
if tag == "sec":
|
||||
return "section"
|
||||
if tag == "title" and parent == "sec":
|
||||
return "title"
|
||||
if tag == "label" and parent == "sec":
|
||||
return "section_label"
|
||||
if tag == "label" and parent == "title" and grandparent == "sec":
|
||||
return "section_label"
|
||||
if tag == "label" and parent == "table-wrap":
|
||||
return "table_label"
|
||||
if tag == "caption" and parent == "table" and grandparent == "table-wrap":
|
||||
return "table_label"
|
||||
if tag == "caption" and parent == "fig":
|
||||
return "figure_caption"
|
||||
if tag == "fig":
|
||||
return "figure"
|
||||
if tag == "mixed-citation":
|
||||
return "citation"
|
||||
if tag == "math":
|
||||
return "math"
|
||||
if tag == "p":
|
||||
return "paragraph"
|
||||
if tag == "table-wrap":
|
||||
return "table"
|
||||
if tag in ("td", "th"):
|
||||
return "cell"
|
||||
if tag == "list-item":
|
||||
return "list_item"
|
||||
if tag in ("graphic", "inline-graphic"):
|
||||
return "image"
|
||||
if tag == "fn":
|
||||
return "footnote"
|
||||
return None
|
||||
|
||||
|
||||
def count_sts_xml(data: bytes) -> tuple[Count, list[str], bool]:
|
||||
|
|
@ -499,7 +559,7 @@ def count_sts_xml(data: bytes) -> tuple[Count, list[str], bool]:
|
|||
|
||||
def walk(node: ET.Element, parent: str | None, grandparent: str | None) -> list[str]:
|
||||
tag = _local(node.tag)
|
||||
role = _sts_role(tag, parent, grandparent)
|
||||
role = _sts_role_xml(tag, parent, grandparent)
|
||||
pieces: list[str] = []
|
||||
if node.text and node.text.strip():
|
||||
pieces.append(_normal(node.text))
|
||||
|
|
@ -531,7 +591,7 @@ def count_sts_json(data: bytes) -> Count:
|
|||
if not isinstance(body, dict):
|
||||
return pieces
|
||||
tag = str(body.get("tag"))
|
||||
role = _sts_role(tag, parent, grandparent)
|
||||
role = _sts_role_json(tag, parent, grandparent)
|
||||
for child in body.get("c") or []:
|
||||
pieces.extend(walk(child, tag, parent))
|
||||
if role is not None:
|
||||
|
|
@ -560,7 +620,20 @@ def _text_of(node: ET.Element, tag: str) -> str:
|
|||
return "".join(t.text or "" for t in node.iter(tag))
|
||||
|
||||
|
||||
DOCX = ("cell", "footnote", "heading", "image", "paragraph", "table")
|
||||
DOCX = (
|
||||
"cell",
|
||||
"comment",
|
||||
"endnote",
|
||||
"footnote",
|
||||
"header_footer",
|
||||
"heading",
|
||||
"image",
|
||||
"paragraph",
|
||||
"table",
|
||||
"text_box",
|
||||
)
|
||||
|
||||
_DOCX_HEADER_FOOTER = re.compile(r"word/(header|footer)\d*\.xml")
|
||||
|
||||
|
||||
def _docx_lines(para: ET.Element) -> list[str]:
|
||||
|
|
@ -570,11 +643,20 @@ def _docx_lines(para: ET.Element) -> list[str]:
|
|||
a broken paragraph can land in different places -- inside a grid table
|
||||
they land on different rows, with other cells' text between them."""
|
||||
lines = [""]
|
||||
for node in para.iter():
|
||||
if node.tag == f"{_W}t":
|
||||
lines[-1] += node.text or ""
|
||||
elif node.tag in (f"{_W}br", f"{_W}cr"):
|
||||
lines.append("")
|
||||
|
||||
def walk(node: ET.Element) -> None:
|
||||
for child in node:
|
||||
# A text box holds its own paragraphs. Read as part of the
|
||||
# paragraph that carries the box, its text is counted twice.
|
||||
if child.tag == f"{_W}txbxContent":
|
||||
continue
|
||||
if child.tag == f"{_W}t":
|
||||
lines[-1] += child.text or ""
|
||||
elif child.tag in (f"{_W}br", f"{_W}cr"):
|
||||
lines.append("")
|
||||
walk(child)
|
||||
|
||||
walk(para)
|
||||
return [_normal(line) for line in lines if line.strip()]
|
||||
|
||||
|
||||
|
|
@ -583,12 +665,23 @@ def count_docx(data: bytes) -> Count:
|
|||
other w:p with text. table: w:tbl. cell: w:tc. image: a:blip.
|
||||
footnote: a w:footnote with a positive id."""
|
||||
count = Count(DOCX)
|
||||
parts: dict[str, ET.Element] = {}
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
root = ET.fromstring(archive.read("word/document.xml"))
|
||||
notes = None
|
||||
if "word/footnotes.xml" in archive.namelist():
|
||||
notes = ET.fromstring(archive.read("word/footnotes.xml"))
|
||||
for name in sorted(archive.namelist()):
|
||||
if name in ("word/footnotes.xml", "word/endnotes.xml", "word/comments.xml") or (
|
||||
_DOCX_HEADER_FOOTER.fullmatch(name)
|
||||
):
|
||||
parts[name] = ET.fromstring(archive.read(name))
|
||||
# A text box's paragraphs are `w:p` in the body too: counted as prose they
|
||||
# would be booked twice, so the box owns them and they are its pieces.
|
||||
boxed: set[int] = set()
|
||||
for box in root.iter(f"{_W}txbxContent"):
|
||||
boxed.update(id(p) for p in box.iter(f"{_W}p"))
|
||||
count.add("text_box", *[line for p in box.iter(f"{_W}p") for line in _docx_lines(p)])
|
||||
for para in root.iter(f"{_W}p"):
|
||||
if id(para) in boxed:
|
||||
continue
|
||||
style = para.find(f"{_W}pPr/{_W}pStyle")
|
||||
lines = _docx_lines(para)
|
||||
if style is not None and _HEADING_STYLE.match(style.get(f"{_W}val", "")):
|
||||
|
|
@ -601,16 +694,29 @@ def count_docx(data: bytes) -> Count:
|
|||
count.add("cell", *[line for p in cell.iter(f"{_W}p") for line in _docx_lines(p)])
|
||||
for _ in root.iter(f"{_A}blip"):
|
||||
count.add("image")
|
||||
if notes is not None:
|
||||
for note in notes.iter(f"{_W}footnote"):
|
||||
if int(note.get(f"{_W}id", "0")) > 0:
|
||||
count.add(
|
||||
"footnote", *[line for p in note.iter(f"{_W}p") for line in _docx_lines(p)]
|
||||
)
|
||||
for name, part in parts.items():
|
||||
if _DOCX_HEADER_FOOTER.fullmatch(name):
|
||||
for para in part.iter(f"{_W}p"):
|
||||
lines = _docx_lines(para)
|
||||
if lines:
|
||||
count.add("header_footer", *lines)
|
||||
continue
|
||||
role, tag = (
|
||||
("footnote", f"{_W}footnote")
|
||||
if name.endswith("footnotes.xml")
|
||||
else ("endnote", f"{_W}endnote")
|
||||
if name.endswith("endnotes.xml")
|
||||
else ("comment", f"{_W}comment")
|
||||
)
|
||||
for note in part.iter(tag):
|
||||
# A separator note carries id 0 and no document text.
|
||||
if role != "comment" and int(note.get(f"{_W}id", "0")) <= 0:
|
||||
continue
|
||||
count.add(role, *[line for p in note.iter(f"{_W}p") for line in _docx_lines(p)])
|
||||
return count
|
||||
|
||||
|
||||
PPTX = ("cell", "image", "paragraph", "slide", "table", "title")
|
||||
PPTX = ("cell", "hidden_slide", "image", "note", "paragraph", "slide", "table", "title")
|
||||
|
||||
|
||||
def count_pptx(data: bytes) -> Count:
|
||||
|
|
@ -620,9 +726,15 @@ def count_pptx(data: bytes) -> Count:
|
|||
count = Count(PPTX)
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
slides = [n for n in archive.namelist() if re.fullmatch(r"ppt/slides/slide\d+\.xml", n)]
|
||||
notes = [
|
||||
n for n in archive.namelist() if re.fullmatch(r"ppt/notesSlides/notesSlide\d+\.xml", n)
|
||||
]
|
||||
for name in sorted(slides, key=_slide_order):
|
||||
root = ET.fromstring(archive.read(name))
|
||||
count.add("slide", *_pptx_lines(root))
|
||||
# `show="0"` is the deck saying this slide is not shown. Counted as
|
||||
# an ordinary slide it is indistinguishable from one that is.
|
||||
hidden = root.get("show") == "0"
|
||||
count.add("hidden_slide" if hidden else "slide", *_pptx_lines(root))
|
||||
for table in root.iter(f"{_A}tbl"):
|
||||
count.add("table", *_pptx_lines(table))
|
||||
for cell in root.iter(f"{_A}tc"):
|
||||
|
|
@ -645,6 +757,10 @@ def count_pptx(data: bytes) -> Count:
|
|||
else:
|
||||
for text in texts:
|
||||
count.add("paragraph", text)
|
||||
for name in sorted(notes, key=_slide_order):
|
||||
root = ET.fromstring(archive.read(name))
|
||||
for line in _pptx_lines(root):
|
||||
count.add("note", line)
|
||||
return count
|
||||
|
||||
|
||||
|
|
@ -660,7 +776,7 @@ def _slide_order(name: str) -> tuple[int, str]:
|
|||
return (int(match.group(1)) if match else 0, name)
|
||||
|
||||
|
||||
XLSX = ("cell", "image", "row", "sheet")
|
||||
XLSX = ("cell", "formula", "hidden_sheet", "image", "row", "sheet")
|
||||
|
||||
|
||||
def _shared_strings(archive: zipfile.ZipFile) -> list[str]:
|
||||
|
|
@ -684,6 +800,31 @@ def _cell_text(cell: ET.Element, shared: list[str]) -> str:
|
|||
return _normal(raw)
|
||||
|
||||
|
||||
def _hidden_sheets(archive: zipfile.ZipFile) -> set[str]:
|
||||
"""The worksheet PARTS the workbook marks hidden.
|
||||
|
||||
The sheet file says nothing about it: the state lives in `workbook.xml`
|
||||
and the part is reached through the relationship id."""
|
||||
names = archive.namelist()
|
||||
if "xl/workbook.xml" not in names or "xl/_rels/workbook.xml.rels" not in names:
|
||||
return set()
|
||||
relationships = ET.fromstring(archive.read("xl/_rels/workbook.xml.rels"))
|
||||
targets = {
|
||||
node.get("Id", ""): str(node.get("Target", ""))
|
||||
for node in relationships
|
||||
if _local(node.tag) == "Relationship"
|
||||
}
|
||||
hidden: set[str] = set()
|
||||
workbook = ET.fromstring(archive.read("xl/workbook.xml"))
|
||||
for sheet in workbook.iter(f"{_S}sheet"):
|
||||
if sheet.get("state") in ("hidden", "veryHidden"):
|
||||
rid = next((v for k, v in sheet.attrib.items() if _local(k) == "id"), "")
|
||||
target = targets.get(rid, "")
|
||||
if target:
|
||||
hidden.add(f"xl/{target.lstrip('/')}" if not target.startswith("xl/") else target)
|
||||
return hidden
|
||||
|
||||
|
||||
def count_xlsx(data: bytes) -> Count:
|
||||
"""sheet: xl/worksheets/sheetN.xml. row: a row holding a value. cell: a c
|
||||
with a value. image: an xdr:pic in a drawing.
|
||||
|
|
@ -693,6 +834,7 @@ def count_xlsx(data: bytes) -> Count:
|
|||
count = Count(XLSX)
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
shared = _shared_strings(archive)
|
||||
hidden = _hidden_sheets(archive)
|
||||
for name in sorted(archive.namelist()):
|
||||
if re.fullmatch(r"xl/worksheets/sheet\d+\.xml", name):
|
||||
root = ET.fromstring(archive.read(name))
|
||||
|
|
@ -706,10 +848,14 @@ def count_xlsx(data: bytes) -> Count:
|
|||
values = [_cell_text(c, shared) for c in valued]
|
||||
for value in values:
|
||||
count.add("cell", value)
|
||||
for cell in valued:
|
||||
formula = cell.find(f"{_S}f")
|
||||
if formula is not None:
|
||||
count.add("formula", _normal(formula.text or ""))
|
||||
if valued:
|
||||
count.add("row", *values)
|
||||
sheet_pieces.extend(values)
|
||||
count.add("sheet", *sheet_pieces)
|
||||
count.add("hidden_sheet" if name in hidden else "sheet", *sheet_pieces)
|
||||
elif re.fullmatch(r"xl/drawings/drawing\d+\.xml", name):
|
||||
root = ET.fromstring(archive.read(name))
|
||||
for _ in root.iter(f"{_XDR}pic"):
|
||||
|
|
@ -717,19 +863,43 @@ def count_xlsx(data: bytes) -> Count:
|
|||
return count
|
||||
|
||||
|
||||
ODT = ("cell", "heading", "image", "list_item", "paragraph", "table")
|
||||
ODT = (
|
||||
"annotation",
|
||||
"cell",
|
||||
"header_footer",
|
||||
"heading",
|
||||
"image",
|
||||
"list_item",
|
||||
"paragraph",
|
||||
"table",
|
||||
)
|
||||
|
||||
_OFFICE = "{urn:oasis:names:tc:opendocument:xmlns:office:1.0}"
|
||||
_STYLE = "{urn:oasis:names:tc:opendocument:xmlns:style:1.0}"
|
||||
|
||||
|
||||
def count_odt(data: bytes) -> Count:
|
||||
"""heading: text:h. paragraph: a text:p with text outside a table cell.
|
||||
table: table:table. cell: table:table-cell. list_item: text:list-item.
|
||||
image: draw:image."""
|
||||
styles = None
|
||||
with zipfile.ZipFile(io.BytesIO(data)) as archive:
|
||||
root = ET.fromstring(archive.read("content.xml"))
|
||||
if "styles.xml" in archive.namelist():
|
||||
styles = ET.fromstring(archive.read("styles.xml"))
|
||||
count = Count(ODT)
|
||||
in_cell: set[int] = set()
|
||||
for cell in root.iter(f"{_TABLE}table-cell"):
|
||||
in_cell.update(id(p) for p in cell.iter(f"{_TEXT}p"))
|
||||
# A comment is not prose. Counted as a paragraph it makes the accounting
|
||||
# demand that a reader carry a note the author wrote to themselves.
|
||||
annotated: set[int] = set()
|
||||
for note in root.iter(f"{_OFFICE}annotation"):
|
||||
annotated.update(id(p) for p in note.iter(f"{_TEXT}p"))
|
||||
count.add(
|
||||
"annotation",
|
||||
*[_normal("".join(p.itertext())) for p in note.iter(f"{_TEXT}p")],
|
||||
)
|
||||
|
||||
def lines(node: ET.Element) -> list[str]:
|
||||
pieces = [_normal("".join(p.itertext())) for p in node.iter(f"{_TEXT}p")]
|
||||
|
|
@ -740,7 +910,7 @@ def count_odt(data: bytes) -> Count:
|
|||
count.add("heading", _normal("".join(heading.itertext())))
|
||||
for para in root.iter(f"{_TEXT}p"):
|
||||
text = _normal("".join(para.itertext()))
|
||||
if id(para) not in in_cell and text:
|
||||
if id(para) not in in_cell and id(para) not in annotated and text:
|
||||
count.add("paragraph", text)
|
||||
for table in root.iter(f"{_TABLE}table"):
|
||||
count.add("table", *lines(table))
|
||||
|
|
@ -750,6 +920,15 @@ def count_odt(data: bytes) -> Count:
|
|||
count.add("list_item", *lines(item))
|
||||
for _ in root.iter(f"{_DRAW}image"):
|
||||
count.add("image")
|
||||
# The header and the footer live in `styles.xml`, which is why no reader
|
||||
# looking only at `content.xml` can see them at all.
|
||||
if styles is not None:
|
||||
for place in (f"{_STYLE}header", f"{_STYLE}footer"):
|
||||
for region in styles.iter(place):
|
||||
for para in region.iter(f"{_TEXT}p"):
|
||||
text = _normal("".join(para.itertext()))
|
||||
if text:
|
||||
count.add("header_footer", text)
|
||||
return count
|
||||
|
||||
|
||||
|
|
@ -759,10 +938,11 @@ _RTF_CONTROL = re.compile(r"\\([a-zA-Z]+)(-?\d+)? ?|\\'([0-9a-fA-F]{2})|\\([^a-z
|
|||
|
||||
#: Groups that hold no document text. `{\*...}` says so in the format itself;
|
||||
#: these say it by name, and without them a fixture's font table reads as the
|
||||
#: first paragraph of its prose.
|
||||
#: first paragraph of its prose and a picture's hex payload as the next one.
|
||||
_RTF_SILENT = frozenset(
|
||||
{
|
||||
"fonttbl",
|
||||
"pict",
|
||||
"colortbl",
|
||||
"stylesheet",
|
||||
"info",
|
||||
|
|
@ -957,6 +1137,79 @@ WITNESSED_SUFFIXES = (
|
|||
)
|
||||
|
||||
|
||||
#: What each witness STILL does not count, by name and per file type.
|
||||
#:
|
||||
#: An accounting can only lose visibly what something counts, so this list is
|
||||
#: the gate's own statement of its blind spots -- printed on every run, never
|
||||
#: inferred, and the raw material for the next capability order. Written
|
||||
#: 2026-09-18 from an independent review's per-format reading of this file.
|
||||
NOT_COUNTED: dict[str, tuple[str, ...]] = {
|
||||
".csv": (
|
||||
"a semicolon-separated file (the Norwegian default) reads as one cell per row",
|
||||
"quoting and encoding errors, which arrive as text",
|
||||
),
|
||||
".docx": (
|
||||
"SmartArt, charts and embedded OLE objects",
|
||||
"tracked deletions",
|
||||
"hyperlink targets (the link text counts, the address does not)",
|
||||
"an equation written as `m:oMath` (it holds `m:t`, not `w:t`)",
|
||||
"a picture's alt text",
|
||||
),
|
||||
".htm": (
|
||||
"text in `div`, `blockquote`, `pre`, `dd`, `figcaption`, `caption` and bare text",
|
||||
"`alt` and `title` attributes",
|
||||
"`details`/`summary`, `picture`/`source`, `svg`, `object`, `iframe`",
|
||||
),
|
||||
".html": (
|
||||
"text in `div`, `blockquote`, `pre`, `dd`, `figcaption`, `caption` and bare text",
|
||||
"`alt` and `title` attributes",
|
||||
"`details`/`summary`, `picture`/`source`, `svg`, `object`, `iframe`",
|
||||
),
|
||||
".json": ("the order of members, and comments a JSON superset would allow",),
|
||||
".md": (
|
||||
"Setext headings (`===`, `---`)",
|
||||
"reference images `![a][r]` and raw `<img>`",
|
||||
"footnotes, indented code blocks and front matter (they count as paragraphs)",
|
||||
),
|
||||
".odt": (
|
||||
"tracked changes",
|
||||
"`draw:object` (an embedded chart or formula)",
|
||||
"a picture's `xlink:href`: an odt image counts as embedded, so its bytes "
|
||||
"cannot be traced to an inbox file",
|
||||
),
|
||||
".pdf": (
|
||||
"headings, paragraphs and tables as such (operator-approved exception, "
|
||||
"2026-09-17): the page's text is counted, its structure is not",
|
||||
"form fields, annotations, attachments and bookmarks",
|
||||
),
|
||||
".pptx": (
|
||||
"SmartArt, charts and comments",
|
||||
"alt text",
|
||||
"a slide layout's and master's own text",
|
||||
),
|
||||
".rtf": (
|
||||
"text inside `\\header`, `\\footer` and `\\footnote` counts as body prose",
|
||||
"a picture Word writes twice (`\\shppict` and `\\nonshppict`) counts twice",
|
||||
"a picture's payload is binary, so the paragraph holding it has no text "
|
||||
"the gate can look for",
|
||||
"an empty `\\par` counts as a paragraph, where docx counts only one with text",
|
||||
),
|
||||
".txt": ("nothing beyond lines and paragraphs: the format declares no more",),
|
||||
".xlsx": (
|
||||
"merged cells, cell comments, defined names and charts",
|
||||
"a sheet's NAME",
|
||||
"number formats (a date reads as its serial number)",
|
||||
),
|
||||
".xml": (
|
||||
"`ref` and `element-citation` outside `mixed-citation`",
|
||||
"`def-list`, `term-sec` and `app`",
|
||||
"a `non-normative-note`'s label",
|
||||
"attributes, and all text of an XML document that is not NISO-STS beyond "
|
||||
"the element's own text",
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def witness_file(inbox: Path, path: Path) -> Inventory:
|
||||
"""Count one file under `inbox`. Raises WitnessRefused for a file the
|
||||
witness does not read."""
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue