portfolio-optimiser/docs
Kjell Tore Guttormsen e7ffe9e196 docs(n-bundlene): rank 1 on 3 of 3, and the cut still drops the key the question asks by [skip-docs]
The live consumption measurement on N100:2023, N200:2024 and N500:2024, ordered
as P1 (20260908T113200Z-46497613). Three paid arms in A-form
(--prepass-payload), one per bundle, NOK 0.21 of a NOK 5 cap, 0 x 429. No file
under src/ touched; no production code in this commit.

ALL FIVE KNOWN-POSITIVES HELD BEFORE ANYTHING WAS PAID FOR. The three
sha256-tree refs reproduce vegnormal-okf's values character for character;
okf.navigate_bundle -- our own code -- reaches exactly 446 / 1133 / 270
concepts with 0 unfollowed links; the pre-pass on okf a37d5ce delivers the gold
concept at RANK 1 of 8 on all three with the flagless default command;
okf_contract_check exits 0 three times; the cheap client probe is green.
okf's three opt-in flags were NOT used: the order allowed them only if the gold
came back below_k, and it never did.

THE VERDICT IS 0 OF 3, and the three reasons are different and separately
owned. Criterion: (a) the model answers with the gold requirement's central
condition AND (b) cites the right concept id AND (c) zero hallucinated values.
(c) is 0 on all three -- the model quoted only what it had. (a) is YES only on
N100, where it reproduced the body sentence verbatim. On N200 and N500, 0 of 6
keywords from each gold body appear in any answer.

THE MAIN FINDING, MEASURED, NOT GUESSED (okf owns it): the payload's excerpt
carries `text` and no `title`/`req_number`, so for N100 and N500 the
requirement number the question is ASKED BY appears nowhere in the model's
context -- not in the gold excerpt, not anywhere in the delivered cut. The
ranking finds the right document ON that number and the delivery then drops it.
Our own read_file returns the whole 774-character file including
`req_number: Krav 3.3.1-13`; the excerpt is 98 characters of body. The declared
cut is strictly less informative than our own navigation on precisely the key
the question uses. This is the structural twin of S7c: there the locks bought
the BYTES and not the STRUCTURE; here the ranking buys the RIGHT DOCUMENT and
not the KEY.

THE SECOND FINDING IS OURS, and the order asked for it by name: po has no
lookup door in A-form. `run.py:1243` is hard-coded (`Find a cost-saving measure
for {project.id}`), none of run_project's 26 parameters carries a question, a
prompt, an objective or a task (measured with inspect.signature), and
`--explore` is refused together with `--prepass-payload` (rc 1, zero model
calls). The question reaches the model only as the pre-pass's declared
`question` line -- verbatim in 2 of 6 / 2 of 5 / 2 of 6 prompts -- while the
task line says something else. On N200 and N500 the model followed the task it
was actually given and used a different delivered concept. That is the form,
not a misconfiguration, and choosing what to do about it is the operator's.

(b) IS SCORED ON THE MODEL'S REFERENCE, NOT ON THE STAMP, and that is a
sharpening of the order's criterion rather than a softening. The stamp cites
the delivered concepts MECHANICALLY: `stamp intersect delivered` is 8 of 8 on
all three and would be 8 of 8 for a run in which the model read nothing. A
measure that can only come out one way is this repository's vacuous-gate class,
so the load-bearing reading is what the model actually named -- and there the
answer is no on all three.

A THIRD FINDING, reported and not fixed: N200's run VALIDATED a proposal whose
cost codes were `03418c46` (a real requirement id used as a cost line) and
`03423b12` (0 matches among the bundle's 1137 files -- invented). It passed
because the run was un-anchored and the validator's stage 0 was skipped: the
first LIVE instance of exactly what --require-cost-baseline (F4, session 100)
refuses. The run said so itself, in plain text, on its own notice line.

TWO THINGS GOT CONFIRMED LIVE ON THREE FRESH CORPORA, neither of them ordered.
`{run_id}-parse-failures.json` is absent from all three outboxes, so the
derived structured-output grammar (Fase 1b finding 1b) was accepted by the live
endpoint on corpora it had never been tested against -- closing one of that
row's stated honesty limits. And Step 5's informed refinement fires: attempts 2
and 3 carry "your previous proposal was REJECTED by the deterministic
validator" verbatim, and the claim falls monotonically (3 000 000 -> 1 500 000
-> 600 000 on N100; 350 000 -> 210 000 -> 90 000 on N500).

S7a-3's slack was paid for here for the first time: all three bases declare
`bundle_id` in the root index while mounted under a different directory name.
Before session 81 that was REFUSED, so none of the three could have been opened
as delivered.

HONESTY LIMITS, stated: one run is one run -- three arms, one question each,
one model, no repetition, so nothing here measures variance. (a) on N100 is not
evidence that the chain answers lookups: that gold requirement is
cost-shaped, so answering the task and answering the question coincide, and the
coincidence is identified rather than hidden. (c) = 0 covers the PROSE; the
structured proposal invents cost codes, which is structurally required with no
baseline. The denominator vocabulary was used 0 times in 17 replies, but no arm
was ever in the position the marker exists for, so that is an unanswered
denominator and not evidence it does not work. None of the three bundles was
reviewed on its subject matter; (a) compares against the concept body only.

Suite after `git add`: 1511 passed / 5 skipped, ruff and mypy clean, golden
demo-transcript.stdout BYTE-UNCHANGED (shasum -a 1 of the CONTENT =
ea8c534773acdbe41ae68f2c55724d69aaf8be4f, never the git blob id). Leak check
after staging: 0 hits for the real host in staged content and in every tracked
file. Verdicts sent FYI to vegnormal-okf (which owns nothing here -- the
bundles are clean) and to llm-ingestion-okf (one field).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 14:01:31 +02:00
..
fase1-spikes fix(fase1): spike B fan-out measures real conversation bleed, not a counter 2026-06-24 11:09:55 +02:00
plan fix(maf): en vakt som gikk inert i STILLHET, funnet ved aa loefte pinnen (F15, ORDRE 20260829T155150Z) 2026-09-02 19:35:49 +02:00
rapport docs: fjern brukernavn, hjemmekatalog-stier og privat namespace fra publisert flate 2026-08-18 13:20:23 +02:00
research docs(research): MAF 1.9.0 capability map — feature-utilization for Fase 2 [skip-docs] 2026-06-24 11:36:26 +02:00
2026-06-24-two-approaches-brief.md docs: plain-text brief — goal + two approaches (MAF vs Claude Agent SDK) + learning goal 2026-06-24 09:21:31 +02:00
2026-06-26-fot-i-bakken.md docs(fot-i-bakken): ground-truth-verifisert levert-vs-lovet — agentiske lag inerte 2026-06-26 15:25:03 +02:00
2026-07-15-foundry-auth-recipe.md docs(1b): Claude på Foundry er en TREDJE klientflate — FoundryChatClient kan ikke binde den 2026-08-13 20:31:36 +02:00
2026-08-13-demo-presentasjon.html docs(deck): begge deck bærer det MÅLTE 1b-utfallet — ærlighet, ikke seier 2026-08-14 20:28:35 +02:00
2026-08-13-fase1a-lokal-ende-til-ende.md docs(fase1a): «uansett modell» hvilte på en måling jeg aldri leste ferdig 2026-08-13 20:19:31 +02:00
2026-08-13-fase4-azure-yaml-valg.md feat(4b): AZURE-profilen leser miljøet sitt, ikke operatørens laptop 2026-08-13 22:27:21 +02:00
2026-08-13-fase4-research-spike.md docs(4·): containeren er bygget — og wheelen er ikke installerbar alene 2026-08-13 21:57:40 +02:00
2026-08-13-presentasjon-ledelse.html docs(deck): tredje deck bærer samme MÅLTE 1b-utfall som de to andre 2026-08-18 09:37:14 +02:00
2026-08-14-fase1b-forste-levende-kjoring.md docs(1b): koordinatene byttet mot plassholdere - vei A, redigert framover 2026-08-18 13:03:11 +02:00
2026-08-18-vurdering-azure-omdoeping.md feat(gate): pakke-gaten leser INNHOLDET, ikke bare filnavn + vurdering av Azure-omdøping 2026-08-18 13:42:02 +02:00
2026-08-25-fable-misjonsreview.md docs: Fable 5 misjons-review — nærmer systemet seg faktisk målet? (ORDRE 20260825T104711Z) 2026-08-25 14:12:23 +02:00
2026-08-25-syretest-vei-ab.md feat(explore): katalogkallet koster O(baser), ikke O(korpus) (ORDRE 20260825T213645Z) 2026-08-26 14:45:16 +02:00
2026-08-26-katalogkostnaden.md feat(explore): katalogkallet koster O(baser), ikke O(korpus) (ORDRE 20260825T213645Z) 2026-08-26 14:45:16 +02:00
2026-08-29-maf-gjelden-omfang.md docs: MAF-gjeldens omfang malt, ikke bygget (F3/F15/F16/U16-17-19, ORDRE 20260825T214801Z) 2026-08-29 09:30:01 +02:00
2026-09-02-f15-maf-pinnen.md docs(f15): nevner-disiplin anvendt paa mitt EGET instrument - 17/16/2, ikke 16/16/2 2026-09-02 19:47:11 +02:00
2026-09-02-misjonsreview-v2.md docs: misjonsreview v2 mot bruksscenarioet — målt, ikke bygget (S2, ORDRE 20260902T113744Z-1245330375) 2026-09-02 17:17:39 +02:00
2026-09-02-read-bundle-kontekstkostnad.md docs: name the mechanism behind the untouched-debate proof, and re-measure all three bases 2026-09-03 01:30:09 +02:00
2026-09-03-forslag-fra-mandat.md docs(s7b): maaledokumentet baerer funnet CLAUDE.md-raden peker paa 2026-09-03 22:40:56 +02:00
2026-09-03-hierarkisk-navigasjon-k2.md docs(s7a-3): ordrens ANDRE binding maalt paa alle 478 K2-nivaaer, og to prosa-hull lukket 2026-09-03 08:09:53 +02:00
2026-09-03-syretest-s7a-k2.md docs(s7a): SS 5s "deler aarsak" var en paastand, ikke en maaling - og maalingen felte tallet 2026-09-03 03:42:18 +02:00
2026-09-03-syretest-s7a2-k2.md fix(docs): rediger vekk absolutt hjemmesti i S7a-2-rapporten - handover-gaten var roed paa HEAD 2026-09-03 06:02:17 +02:00
2026-09-04-major2-proposal-review-k2.md docs(major2): steg 9 fullfoert - M38..M40 roede, K2-omkjoeringen maalt, last_ruling-paastanden rettet 2026-09-05 20:39:52 +02:00
2026-09-04-s2c-debatt-k2.md docs(verdict-gate): maaleavsnittet og invariantraden - fem mutasjoner, fire tall verifisert 2026-09-04 20:41:32 +02:00
2026-09-04-syretest-s7b-k2.md docs(s7b): syretesten paa K2 - maaledokumentet med de fire tallene og tretten roede 2026-09-04 09:04:51 +02:00
2026-09-05-major2-trekreview.md docs(major2): remedieringsfoot under den sporede trekreview-kopien 2026-09-05 22:31:06 +02:00
2026-09-06-major2-levende-k2.md docs(major2): funn (b) re-maalt live - loekka konvergerer paa foerste reviderte forsoek 2026-09-07 00:52:39 +02:00
2026-09-07-okf-prepass-i-debatten.md docs(prepass): before/after on K2, the mutation battery and the invariant row 2026-09-07 14:39:50 +02:00
2026-09-07-prepass-mater-q5b-k2.md feat(prepass): --prepass-seed makes the cut a starting point, and K2 says it costs 2026-09-08 04:02:09 +02:00
2026-09-07-syretest-s7-prepass-k2.md docs(s7): a live model reads a declared cut -- three premises felled, four findings [skip-docs] 2026-09-07 16:07:29 +02:00
2026-09-08-f3-f4-nekten-og-forankringen.md feat(navigation,validator): read_dir names the rung that reads a document, and a run can require its anchoring 2026-09-08 05:32:46 +02:00
2026-09-08-n-bundlene-konsum.md docs(n-bundlene): rank 1 on 3 of 3, and the cut still drops the key the question asks by [skip-docs] 2026-09-08 14:01:31 +02:00
2026-09-08-syretest-s7c-begge-laaser-k2.md docs(s7c): both locks open, the price is delivered -- and the live model still did not read it [skip-docs] 2026-09-08 07:00:47 +02:00
bestille-en-kjoring.md feat(provenance): a run records which external service it actually called 2026-08-05 21:37:29 +02:00
ekspert-svar.md feat(hitl): ekspertdommen kan ikke oppstaa av stillhet (F2, ORDRE 20260825T214801Z) 2026-08-27 01:22:07 +02:00
extending.md docs: llms.txt + sikkerhetskontakt til security@ (ORDRE 20260821T041218Z) 2026-08-21 11:29:54 +02:00
knowledge-base-recipe.md docs: kunnskapsbase for ÉN konkret kjøring — kategorier, innholdstyper, veiprosjekt-eksempel (ORDRE 20260821T083046Z) 2026-08-21 11:10:59 +02:00
kort-presentasjon.html docs(deck): begge deck bærer det MÅLTE 1b-utfallet — ærlighet, ikke seier 2026-08-14 20:28:35 +02:00
kunnskapsbase-for-en-kjoring.md feat(visibility): en lenke som ikke ble fulgt sier det - spor + betinget linje (ORDRE 20260821T142704Z) 2026-08-21 17:18:18 +02:00
okf-konsum-kontrakter.md docs(s7a-3): ordrens ANDRE binding maalt paa alle 478 K2-nivaaer, og to prosa-hull lukket 2026-09-03 08:09:53 +02:00
presentasjon-bygge-kunnskapsbase.html feat(visibility): en lenke som ikke ble fulgt sier det - spor + betinget linje (ORDRE 20260821T142704Z) 2026-08-21 17:18:18 +02:00
presentasjon-fagpersonens-bidrag.html docs: fagpersonens bidrag - spørsmål/sjekkliste til fagperson (ORDRE 20260824T092912Z) [skip-docs] 2026-08-24 15:24:51 +02:00
review-2026-07.md docs(review): kryssmodell-review 2026-07 (14 funn, 11 detach-bevis) + revidert roadmap + sesjonsplan Fase 2-6 2026-07-10 06:28:11 +02:00