llm-ingestion-okf/docs/2026-09-08-prisform-og-loggen-k2.md
Kjell Tore Guttormsen 56ae274246 fix(extract,build): write a spreadsheet as pipe tables, stop linking the run log from the index
Two producer-side findings from the consumer's S7c acid test (ordre 20260908T063454Z-3648220855-from-.claude), both measured on K2 before and after, both with the corpus rebuilt from scratch.

FUNN 3 -- THE FORM. The converter's default markdown writer emits simple tables, which pad every cell out to the width of the widest cell in its column. Measured on the tender's price sheet: one 594-character prose cell produced a 67 244-character whitespace carpet with runs of up to 887 characters between a label and its amount, 19 integral amounts carrying a converter `.0`, and a header row naming one column. The bytes reached a live model in 2 of 11 prompts and 0 of 11 answers. The spreadsheet row now writes pipe tables with `--columns=1` (load-bearing: the pipe writer pads to a width computed from it, so at the default 72 a narrow table gains runs of up to 45). Same sheet after: 11 048 characters, longest run 2, one row per line, 0 artificial `.0`. Spreadsheet-only, and the scoping is pinned by three digests -- the same change moves the odt fixture 1366 -> 1105, so it can fail.

The `.0` rewrite is bounded twice: to a cell whose whole content is such a number, anchored between unescaped pipes, and skipped when the literal is in the workbook's shared string table -- the converter renders the number 92 and the TEXT "92.0" identically, so the output alone cannot tell them apart. Read with zipfile and xml.etree; no new dependency.

FUNN 2 -- THE LOG LINK. `link_log_in_root_index` (95eb271) is removed. Consumption contract SS 9.2 forbids a consumer from enumerating the bundle directory unless the profile says the index is derived, so the index tree is the entire map a consumer may use and everything it links is a document: their navigator returned 630 where our pre-pass counts 629, and a corpus run's own log was citable as content. The log is still written to the bundle root (SPEC section 9); `tools/okf_consume.py` keeps its exclusion for the bundles already built with the link.

K2 rebuilt twice. BEFORE reproduces the consumer's ref exactly (`sha256-tree:f14872a0...c8a92a`, 629 concepts) and their three consume figures to the token (57 289 / 62 149 / 58 401). AFTER: 629 concepts, `merged + coded rejections = 43 = N`, new ref `sha256-tree:c26eed6a...e3261f`, 627 of 629 concepts byte-identical, 1104 of 1108 files identical to the delivered bundle.

ONE REGRESSION, MEASURED AND NOT FIXED: on the mandate-shaped question with the vocabulary bridge the priced concept moves from candidate rank 10 to 19, so `--k 12` withholds it `below_k`; `--cost-vocabulary --k 20` delivers it at 65 912 o200k. The cause is measured rather than argued -- restoring only the concept's title on the new short body ranks it 10 again. The chain ends at the orphan check (`propose.py:461`), which drops the sheet heading once a table block opens two lines below it. That is the already-reported orphan gate, and changing it is a default-ON segmentation rule affecting every document type. The specific question is unaffected: rank 1 before and after. The priced excerpt's budget share falls from 56.5 % to 9.7 %.

11 new tests (RED first), 8 mutations, 8 red, with an unmutated control green each time. One mutation survived twice before the fixture could make it fire, and both survivals are written down. 1279 -> 1287 tests. mypy --strict clean on 28 files. ruff clean. Both proposer goldens byte-unchanged. One frozen literal moved with the fix and is reported rather than hidden.

Report: docs/2026-09-08-prisform-og-loggen-k2.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 10:06:58 +02:00

23 KiB

The delivered price sheet's FORM, and the log the index linked

Date: 2026-09-08 · Order: 20260908T063454Z-3648220855-from-.claude · Instrument: src/llm_ingestion_okf/extract.py at 6776c37 plus the two changes this document reports · Upstream finding: portfolio-optimiser docs/2026-09-08-syretest-s7c-begge-laaser-k2.md (648b36e) SS 5 and SS 9, findings 2 and 3.

The consumer that reads bundles from this library ran a live model over a 629-concept bundle with both of the pre-pass's optional widenings open. The document holding the tender's prices WAS delivered, at rank 10, and its bytes reached two of eleven prompts. It appears in zero of the eleven answers.

They measured why rather than guessing, and the answer is on this side of the boundary: the delivered excerpt is a pandoc SIMPLE table with one column header, 67 245 characters over 104 lines, with whitespace runs of up to 887 characters between a label and its amount. The same form makes their own cost-baseline deriver refuse the sheet. Opening both locks bought the BYTES and not the STRUCTURE.

Their second finding is the same class of defect at the bundle level: this library writes a link to the run's own log.md into the root index.md, so a consumer's navigator reaches 630 documents where our own pre-pass counts 629, and a corpus run's log is readable and citable as content.

Both are producer-side. This document measures both, fixes both, and states what it did not fix.

The corpus is external and private to the measurement; no document name, path or body from it appears here. Documents are named by shape ("the priced sheet") and the numbers are counts.


0. What IS measured, and what is NOT

Measured. The form of both spreadsheet documents in the corpus before and after the change, each figure with the command that produced it; the line that produces the form, read out of the code rather than inferred; the fix on a hand-laid synthetic sheet and on the corpus's OTHER sheet; a full corpus rebuild before and after with its diff, its new content identity and its concept count; the pre-pass's delivered rank and token cost for the priced sheet after the change; and the index/concept disagreement, on a build rather than on an argument.

Known-positive, run first. The upstream report's four form figures, re-measured on this tree before anything changed:

# figure upstream measured here
1 delivered excerpt, characters 67 245 67 244 + the trailing LF
2 delivered excerpt, lines 104 103 + the trailing LF
3 longest whitespace run 887 887
4 numbers >= 1000 in .0 form 19 19
5 non-empty header cells in the table's header row 1 1

Figures 1 and 2 differ by exactly the trailing newline, which is stated rather than rounded away: the file's last line ends with an LF, so a count that includes it gives 67 245 and 104. Figures 3, 4 and 5 are exact. An instrument that has not reproduced a known figure has not been shown to count.

One thing this change makes WORSE, measured and reported rather than discovered later. On the mandate-shaped question with the vocabulary bridge open, the priced concept moves from candidate rank 10 to rank 19, so the command the consumer used (--k 12) no longer reaches it. The cause is traced in SS 7 and it is not the new form: it is the sheet's heading being dropped by a segmentation rule the new form wakes up, which costs the concept its title. The question this library's own hit@k measurement uses is unaffected -- the priced document still ranks first there.

NOT measured, and it is the finding that matters. Whether a live model now READS the sheet. That needs a live model and it is the consumer's measurement, in their S7 form, not this one. What is measured here is the FORM: the label and the amount are in separate cells on one line instead of separated by hundreds of spaces. Whether that is enough is not this repository's to answer, and saying so is the same limit this library stated when it shipped the two flags.

NOT measured either: that the change helps any corpus other than this one. Two real sheets and one synthetic sheet are not a sample.

One thing is deliberately NOT fixed. The header row still names one column, because the SOURCE's first row has one filled cell. Promoting the next row to a header would be inventing structure the document does not have, which is the one thing this library refuses to do without marking it (CLAUDE.md: an unmarked heuristic is worse than none). If the consumer's deriver requires two or more column HEADERS, this change alone does not satisfy it -- and that is their measurement to make, on the rebuilt bundle named below.


1. Setup

The corpus is the 43-document folder the K2 measurements have used since 2026-08-29. Both bundles below are okf build runs over it with --ingested-at 2026-09-03T00:00:00Z --bundle-id <the delivered id> --okf-version 0.2.

The BEFORE build runs committed 6776c37 source, not the working tree. The package is installed editable, so a build started from this clone would pick up every edit made while it ran -- which it did once, and the run was discarded. The before build therefore runs from git archive 6776c37 unpacked to /tmp with PYTHONPATH pointing at it, and the first line of its log names that path.

Every command in this document is offline: no model call, no socket, no clock.


2. The form, before anything changed

Both figures below are taken from the extracted text with each line's trailing whitespace removed, which is what the upstream report counted.

sheet characters lines longest run numbers >= 1000 in .0 form non-empty header cells
priced, BEFORE 67 244 103 887 19 1
other, BEFORE 1 842 42 18 0 2

The instrument the order names measures the same thing from the table grammar's side, over a two-document corpus holding only the two sheets:

okf_table_measure.py --corpus <the two sheets> --report <out>
BEFORE AFTER
documents with at least one table row 0/2 2/2
table rows in total 0 141
documents with at least one grid-rule line 0/2 0/2

Zero table rows out of two spreadsheets is the whole finding in one number: the shipped table grammar, which is the grammar the segment proposer uses, could not see a single row in either sheet. Not because the sheets have no rows -- they have 141 -- but because the form they were written in has none it can match.


3. The line that produces it

Read out of the code, not inferred. At 6776c37, src/llm_ingestion_okf/extract.py lines 85-86:

_PANDOC_WRITER = "markdown"
_PANDOC_ARGS = ("--eol=lf", "--wrap=none")

markdown with no extension changes is the writer's DEFAULT table selection, which prefers a simple table and falls back to multiline and grid. Every one of those three pads each cell out to the width of the widest cell in its column. That is the whole mechanism: the sheet's first column holds a 594-character prose cell, so every other row in that column is padded to 594 characters.

It is neither the reader nor the source: the header underline of the simple table has EIGHT dash groups, so the converter recovered eight columns from the sheet and the writer then serialised them in a form where the separation is positional whitespace. Measured with the same three writers on the same file:

writer characters longest run
markdown (default tables) 100 795 996
markdown minus simple/multiline/grid tables 11 221 2
the same, with --columns=1 11 221 2

(The figures in this table are the raw converter output including trailing padding, which is why they are larger than SS 2's.)

--columns=1 costs nothing on this sheet and is load-bearing on a narrow one: the pipe writer pads cells out to a width computed from that setting, so at the default 72 a small table gains runs of up to 45 characters. Measured across every office fixture and every office file in the corpus, the longest run with --columns=1 is 2.


4. The fix, in full

Two changes in src/llm_ingestion_okf/extract.py, both spreadsheet-only.

(a) The spreadsheet row writes pipe tables (lines 113-114):

_SPREADSHEET_WRITER = "markdown-simple_tables-multiline_tables-grid_tables"
_SPREADSHEET_ARGS = (*_PANDOC_ARGS, "--columns=1")

selected at the one call site by spreadsheet = suffix == ".xlsx".

(b) An integral cell loses the converter's trailing .0 (line 122 and the two functions below it). The converter renders a numeric cell as a double, so an integral value arrives as 5647500.0. A TEXT cell reading 92.0 arrives as 92.0 as well, and the output alone cannot tell the two apart -- so rewriting on the output alone would silently edit somebody's authored text. The rewrite is therefore bounded twice:

  • to a table cell whose ENTIRE content is such a number, anchored between two pipes that are not backslash-escaped, so a 5.0 inside a longer cell is not a cell and is left alone;
  • and skipped whenever the same literal is present in the workbook's shared string table, which is read with zipfile and xml.etree and no new dependency. Shared strings are the only text the converter recovers from a sheet at all -- an inline string is read as an empty cell, measured when the first spreadsheet fixture was built -- so a <digits>.0 that is not in that set did not come from text.

Every failure to read the workbook returns the empty set, which makes the rewrite a no-op rather than a guess.

Nothing else moved. No new dependency, no new flag, no model call, and no change to any other file type. One frozen literal in the suite moved with the fix and is reported in SS 5 rather than hidden.


5. The tests, the controls and the mutations

Written RED first, in this order.

test what it pins
test_a_spreadsheet_keeps_its_columns_one_row_per_line label and amount are separate cells on one line; no whitespace run over 8
test_an_integral_amount_loses_the_converters_decimal_and_a_real_one_keeps_it 5647500 and 250000 lose .0; the genuine 12.5 keeps its decimal; the shared-string 92.0 survives verbatim; a 5.0 inside an escaped-pipe cell survives
test_a_single_column_sheet_gains_no_columns a one-column source sheet gets one column and its three values, in order
test_the_other_office_rows_are_untouched_by_the_spreadsheet_writer (x3) the docx/pptx/odt/rtf path is byte-identical, by digest
test_prisark_extracts_to_its_frozen_text the whole fixture, byte for byte
test_the_root_index_does_not_link_the_run_log the log is in the bundle and out of every index; the pre-pass's count equals the concept count

The fixture is tests/fixtures/prisark.xlsx, hand-laid part by part in tests/fixtures/make_fixtures.py under the same policy as every other office fixture in this repository: a document written by the converter and then read by the converter proves only that the converter agrees with itself. It carries the defect in miniature -- a single filled cell in row 1, a long prose cell in column B, two integral numbers, one genuine decimal, one text cell reading 92.0 and one cell containing a literal pipe -- and its own negative control on a second sheet.

The negative control the order asked for, and where it does not hold. The order asked that a sheet with no column structure be byte-identical before and after. It is not, and that is measured rather than glossed: the fix is the table FORM, and a one-column table changes delimiter form with every other table. Measured on the fixture's second sheet, the cell content and the column count are unchanged and only the delimiters moved. Two negative controls that DO hold byte for byte are shipped in its place, and neither is a tautology:

  • the three other office rows, by digest -- the same one-line change applied to them was measured to move them (the odt fixture 1366 -> 1105 characters), so the digests can fail;
  • the docx frozen literal, green under every mutation below.

Mutations: 8 written, 8 red, with an unmutated control green each time.

# mutation result
1 writer reverted to the default (simple tables) 5 red
2 --columns=1 dropped 3 red
3 the spreadsheet writer applied to every office row 4 red, incl. all three digests
4 the decimal rewrite never runs 3 red
5 the shared string table ignored 2 red
6 the guard looks up the wrong literal 2 red
7 an escaped pipe counts as a cell delimiter 2 red
8 the log link written back into the root index 2 red

Mutation 7 SURVIVED its first form. The first attempt added re.M to a pattern containing no ^ and no $, which is an equivalent mutation and says nothing about the tests -- the trap this repository has hit before, and the reason a surviving mutation is read as code before it is called a test hole. The second attempt, which actually removes the escaped-pipe guard, survived too: the fixture's pipe-bearing cell ended in prose rather than in the number, so the pattern could not match either way. The fixture was sharpened until the mutation could fire, and then it fired.


6. Generality

The order required the fix be measured beyond the one sheet it was written for.

sheet characters lines longest run numbers >= 1000 in .0
priced, BEFORE 67 244 103 887 19
priced, AFTER 11 048 103 2 0
other, BEFORE 1 842 42 18 0
other, AFTER 1 921 42 2 0
synthetic, BEFORE 950 16 130 n/a (2 integral, 2 text)
synthetic, AFTER 402 16 2 2 integral rewritten, 2 text kept

The other sheet is the honest row here: it gets larger by 79 characters, because pipe delimiters cost more than the padding they replaced on a table whose columns were already narrow. Its longest run still falls from 18 to 2. The fix is not a compression; it is a change of form that happens to compress a sheet with one very wide column by 84 %.


7. The rebuild: 629 concepts, a new identity, and one regression

Both builds are okf build over the same 43-document folder with the same arguments; the before build runs committed 6776c37 source from /tmp.

The known-positive at bundle level, before anything is read from the after build. The before build reproduces the identity the consumer measured:

before after
merged + coded rejections == N 39 + 4 = 43 = N, exit 0 39 + 4 = 43 = N, exit 0
concepts 629 629
content identity sha256-tree:f14872a0…c8a92a sha256-tree:c26eed6a…e3261f
wall time 831.08 s 842.82 s

f14872a0… is the consumer's ref for the bundle they measured, character for character. c26eed6a… supersedes it: their ref is now stale, and every figure keyed to it belongs to the old form.

Seven files differ between the two bundles, and they are all in the two spreadsheet documents. 627 of 629 concepts are byte-identical.

what before after
the priced document's concept …/prissammenstilling-sheet-1.md …/tabell-linje-3.md
the other sheet's concept …/ark1-sheet-1.md …/tabell-linje-3.md
the two documents' index.md link and title of the above idem
the bundle's root index.md carried the log.md link does not

Measured against the bundle delivered on 2026-09-03, 1104 of 1108 files are identical, against 1107 before: the root index.md is identical again (SS 8), and the two spreadsheet documents account for every remaining difference.

The form, measured on the rebuilt concept: longest whitespace run 887 -> 2, numbers >= 1000 in .0 form 19 -> 0, body 100 694 -> 11 048 characters, one row per line. And its excerpt weight, which is what the budget gate counts: 67 838 B -> 11 594 B, from 56.5 % of the default 120 000-byte budget to 9.7 %. The second lock the previous document measured is, for this document, gone.

The regression: rank 10 -> 19, and its cause traced rather than guessed

The order asked whether --cost-vocabulary --k 12 --limit 160000 still delivers the priced sheet. It does not.

question vocabulary before after
mandate-shaped off candidate rank 251 of 271 251 of 271
mandate-shaped on candidate rank 10 of 280 19 of 280
specific off rank 1 of 233 1 of 233
specific on rank 1 of 248 1 of 248

--k 12 reaches rank 12, so at rank 19 the concept is withheld below_k. The question this library's own hit@k measurement uses is unaffected: the priced document still ranks first of 233 candidates.

The cause is the title, not the body, and that is measured rather than argued. A copy of the AFTER bundle with the concept's file name and title restored to the before values -- the new short pipe-table body untouched -- ranks the concept 10 again, exactly where it was.

The chain, each link read out of the code:

  1. the extraction now emits | rows, so rule:table-block opens a candidate on the sheet's first table row (propose.py, RULE_TABLE_BLOCK);
  2. that candidate starts two lines after the sheet heading, so the heading candidate's body is its own line plus a blank one;
  3. the ORPHAN CHECK (propose.py:461) drops a candidate whose body is empty after its first line, so the heading candidate is dropped;
  4. the only surviving candidate is the table block, whose title is the mechanical Tabell linje <n>;
  5. the concept id and title therefore lose the sheet's name, and the cost vocabulary bridge has nothing left to match in signal 1.

This is the orphan gate that was already reported and not fixed (it deletes 34 % of candidates unevenly, measured earlier and open on the operator's list). This is a second measured consequence of it, not a new defect, and fixing it is a change to a default-ON segmentation rule affecting every document type -- which is not in this order and is not decided here.

The commands, measured on the rebuilt bundle:

command payload o200k priced sheet
flagless default 169 573 B 57 286 below_k
--cost-vocabulary --k 12 --limit 160000 172 588 B 58 553 below_k
--cost-vocabulary --k 20 (default budget) 190 428 B 65 912 delivered, rank 19

Direction, stated as the order asked: the payload for the command that delivers the sheet is larger in tokens than before (65 912 against 62 149, +6.1 %), because a bigger k delivers eight more excerpts; the SHEET's own contribution fell by 83 %. The default payload is 3 tokens cheaper and otherwise unmoved.

The three known-positive consume figures reproduce exactly on the before build, which is what makes the after column readable at all: 57 289 o200k for the flagless default, 62 149 for --cost-vocabulary --k 12 --limit 160000, and 58 401 for the same at the default budget -- the consumer's three figures, to the token.


8. The log the index linked

The upstream finding, verbatim in substance: our pre-pass excludes log.md from its walk (considered = 629), but the link is still written into the root index.md, so their navigator follows it and returns 630. The corpus run's own log is a document their agents can read and cite. The bundle delivered on 2026-09-03 has no such link, because the link postdates it.

The link was a LOCAL choice and said so at the time (95eb271). It was made on a real measurement -- the bundle carried a conformant root log.md that no index named, so a reader entering at index.md never reached the one file carrying N. Upstream's own bundles do not link theirs: measured at 9a15b13, 0 of the 24 shipped index.md files name the single log.md in the set. So upstream showed the link was permitted, never that it was required.

Why it goes, and the section that decides it. Consumption contract SS 9.2:

A consumer MUST NOT enumerate a directory unless the named profile says the index is derived.

Measured, entries_match_directory is true for STRICT_V1 alone, and the profile these bundles use is not it. So a consumer of one of our bundles has exactly one lawful way to find out what is in it: walk the index tree. That makes the index the ENTIRE map, and everything the map names a document. A producer that links a non-concept from the index is not adding a convenience -- it is adding a document, to the only enumeration the consumer is allowed to perform, and the consumer has no lawful second opinion to check it against.

SS 8's payload shape is the other half. It requires denominators.considered beside an excerpt list keyed on concept_id; a bundle whose index reaches one more document than it has concepts makes those two numbers answer different questions while looking like one. That is exactly the shape the consumer measured: 630 against 629.

The smallest change that makes producer and consumer agree is therefore not another exclusion rule on the consumer's side -- there is no end to that list -- but for the producer to stop naming a non-concept in the map. log.md is still written to the bundle root, which is where SPEC section 9 puts it and all the earlier fix ever needed; it is no longer linked from any index.

tools/okf_consume.py KEEPS its own exclusion, and the test that pins it now writes the link itself. That is not dead code: every bundle built between 95eb271 and this change carries the link, including the ones consumers are reading today, and the pre-pass must still count 629 on those.

The two tests that pinned the old behaviour were rewritten rather than deleted, including the halves that mattered independently -- rebuild-equals-incremental on both the segmented and the unsegmented path, which is the property the removed append was most likely to break, and the one that has to survive its removal too.


9. Honesty limits

  • The form is measured; the reading is not. Whether a live model uses the sheet now is the consumer's measurement, in their S7 form. This repository cannot make it and does not claim it.
  • The header row still names one column. The source's row 1 has one filled cell and this change does not invent a second. A deriver that requires two or more column HEADERS is not satisfied by this alone.
  • Two real sheets and one synthetic sheet are not a sample. The corpus holds exactly two spreadsheets; both are measured, and that is the whole denominator available, not a selection from a larger one.
  • The other four office rows have the same defect. It is measured (the odt fixture 1366 -> 1105 characters) and deliberately not fixed here, because changing them moves a corpus denominator nothing has measured. It is a decision for the operator, not an oversight.
  • The order's literal negative control does not hold (SS 5): a one-column sheet is not byte-identical, because the fix is the table form. What holds is that no column is invented and no cell content changes.
  • The priced sheet is a PRE-AWARD form and is not filled in (the consumer's own limit, repeated here): no real saving was ever available to find in it, whatever its form.
  • .0 on a number the workbook stores as text is preserved, and the guard is conservative in the other direction too: a workbook whose shared string table happens to contain the same literal keeps its converter decimals on every numeric cell carrying that value.