Commit graph

84 commits

Author SHA1 Message Date
4b9e5d9520 docs(P4.5): §0s tag-felt var sirkulært, og begge utveier gjorde torsdagen rød
§0s tag-rad hentet verdien fra `git tag -l v1.0.0` ETTER push (§5 punkt 10),
mens §5 punkt 6 krever at utfyllings-gaten er TOM — altså FØR. Punkt 6 var
gatet på informasjon som først finnes etter punkt 10.

Ingen utvei holdt. Fylt ærlig krever den en commit ETTER taggen: HEAD forlater
Z, og torsdagens §6 steg 2 — identitets-ankeret økt 10 vant — blir rødt på
demo-morgenen. Ufylt bryter den gaten i punkt 6. Sekvensen modellerer heller
ingen tredje commit (X → Y → Z → tag). Lokal tag før utfylling hjelper ikke
(taggen står fortsatt på Z); --amend etter taggen flytter Zs sha ut under den.

Tvillingen: §5s haker settes i selve fila, og punkt 7-10 skjer ETTER
runbook-commiten Y. Målt: en redigert runbook gir « M docs/plan/…», altså rød
§6 steg 1 — eller en commit etter taggen, altså rød §6 steg 2. Samme
motsigelse, samme to gater.

Løsningen bevarer identiteten: §0 mistet tag-raden med begrunnelsen skrevet
inn, §5 fikk et ellevte punkt som bekrefter taggen der den settes (git tag -l ·
git ls-remote --tags origin · git describe --tags --exact-match HEAD — det
siste er torsdagens anker kjørt et døgn tidlig), og §5s ingress sier at hakene
aldri settes i fila. Ingenting skrives etter taggen.

En åttende defekt falt ut av gjennomlesningen, og den er økt 11s egen bom: §5
punkt 9 sa «Steg 2 målte en tilstand som ikke lenger finnes», mens første
frys-gate-kjøring er punkt 5. Verifisert mot 818b55a: da linja ble skrevet var
frysen TO kommandoer og gaten var nr. 2. Økt 11 gjorde den til tre, men
grep-passen lette etter strengen «to kommandoer», så «Steg 2» slapp forbi —
og økt 11 konkluderte eksplisitt at ingen kryssreferanse pekte på den gamle
nummereringen. Nå forankret i §5s egen nummerering.

Målt: utfyllings-gaten 3 → 2 · §5 ti → elleve punkter · §0-tabell 4 pipes per
rad · §1 og §2 byte-urørt (shasum likt før/etter — økt 9s 67 målinger er gjort
mot de bytene) · fire diff-hunks, alle i §0/§5 · alle 20 numeriske
kryssreferanser sveipet · frys-gaten d0e8bb0..HEAD TOM · ingen test leser
dokumentene.

Propagert til fem levende steder i planen og STATE; GJORT-blokkenes «3» står
som historikk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FfdfrMcEtkmZXPkzt7APQU
2026-08-11 14:11:47 +02:00
63a167a8d8 docs(P4.5): frys-gaten er commit-til-commit, men prøven leser arbeidstreet
To en-linjes herdinger av onsdagen og torsdagen. Null kodeendring.

1) §5 manglet en arbeidstre-sjekk ved X. Frys-gaten er
   `git diff <X>..HEAD` — commit-til-commit — mens generalprøven kjører
   fra arbeidstreet. En ucommittet endring i src/ ved prøvetidspunktet
   gjør X til en beskrivelse av noe som aldri ble prøvd, og BEGGE
   kjøringene av gaten står tomme: de kan ikke se den. Verre om
   endringen er dét som gjør prøven grønn — da er den taggede koden rød,
   som er nøyaktig hullet kriterium 6 ikke dekker. Samme argument som ga
   §6 steg 1 sin plass (økt 10): rent tre er en REGEL, ikke et
   øyeblikksbilde. Punktet sier også hva man gjør ved ikke-tomt — commit
   eller forkast, og kjør så prøven OM IGJEN, siden en commit flytter
   HEAD og X ellers ville pekt på et tre prøven aldri så.

2) §3s abortsti brukte relativ sti til fasit-fila, og prosaen rett under
   navngir «feil katalog» som sannsynlig årsak. Målt: fra en annen
   katalog gir den `cat: tests/golden/…: No such file or directory`.
   En abortsti som deler failure-mode med det den aborterer fra, gjør
   ett synlig problem til to. Nå absolutt sti — målt kjørbar fra
   vilkårlig cwd (61 linjer).

Samme defektklasse som de fem forrige i denne dokumentfamilien:
kommandoen var riktig, konteksten var det ikke.

Klon-tørrkjøring av hele sekvensen ble VURDERT og forkastet: en lokal
klons `origin` peker på dette repoet, så §5 steg 10 (`git push origin
v1.0.0`) ville skapt nøyaktig den stale taggen økt 8 pre-flightet mot —
og den måtte vært stoppet ett steg for tidlig på eve-en av en enveis-dag.

Grep-passen fant fem kopier av «to kommandoer»; de to levende
instruksene er rettet til tre, de to i GJORT-blokkene er historikk og
står. Tir-11-raden sa TRE ØKTER (drift, samme klasse som økt 10 fant).

Målt: utfyllings-gaten 3 (uendret — ingen ny placeholder) · §5 ni → ti
avkryssinger · frys-blokka to → tre punkter · tabell-integritet 3
usiterte pipes per rad (ons-12s tre escapede er K1-kommandoens) ·
frys-gaten mot HEAD TOM (kjørestien urørt) · 810 passed / 4 skipped ·
ruff + format + mypy rene.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186tGdSmqQZUfy4emjSFNy7
2026-08-11 13:50:13 +02:00
8e7fa54267 docs(P4.5): torsdagens anker var arvet, og arvet gjorde det fail-open
§6 landet i 859be2a med frys-gatens to unntak (`docs/`, `CHANGELOG.md`)
kopiert inn i tag-ankeret. Unntakene er ikke generelle: de er begrunnet i
at ONSDAGEN skriver nøyaktig dem. Torsdag skriver ingenting, så riktig
forventning er IDENTITET — og arvet dit gjorde `docs/`-unntaket gaten
blind for den ene skriveren vi vet er aktiv i repoet, den parallelle
sesjonen som eier docs/presentasjon-portfolio-optimiser.html.

Samme klasse som da frys-gaten selv ble snudd fra positiv liste til
eksklusjonsform: en gate arvet uten at begrunnelsen ble re-utledet.

Ankeret er nå `git describe --tags --exact-match HEAD` -> `v1.0.0`.
MÅLT fail-closed begge veier: `no tag exactly matches '<sha>'` (exit 128)
når HEAD ikke er tagget, `bad revision` når taggen mangler. Ingen av dem
kan forveksles med grønt.

Feiler ankeret er det en BESKJED, ikke en abort: `git diff --stat
v1.0.0..HEAD` UTEN unntak skiller kjøresti-endring (-> §3) fra ren
`docs/` (demoen upåvirket, men da vitende).

Steg 1 var også et øyeblikksbilde, ikke en regel: «kun den fremmede
HTML-fila» slutter å stemme i det den sesjonen committer eller legger
igjen en fil til, og en gate som roper på noe operatøren ikke eier lærer
ham å ignorere gaten. Nå `git status --short --untracked-files=no`
-> TOMT (målt).

Lagt til én setning om at en `Resolved`/`Audited`/`Installed`-linje fra
uv er miljø-sjekken, ikke en feil — målt at `uv run` er STILLE på varmt
miljø (stderr = de fire linjene §1 beskriver), men en første kjøring for
dagen kan si fra.

Målt: utfyllings-gaten uendret på 3 · 810 passed / 4 skipped · ruff +
format + mypy rene · kalendertabellens pipe-telling intakt.

Planens to beskrivelser av gaten rettet i samme pass (P4.5-blokka +
tor-13-raden) — de beskrev et design som ikke lenger står.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ESVxVd5srp9PGshHCdi4
2026-08-11 13:14:03 +02:00
859be2acfc docs(P4.5): runbooken manglet sjekken som kjøres FØR demoen
Golden-transkriptet ble sjekket inn nettopp for å fange «regresjon mellom
onsdag og torsdag» (egnethetsreview-planen P4 pkt. 3) — selv-identitet
fanger ikke-determinisme, men ikke at noe flyttet seg over natta.
Mekanismen fantes altså. Men i runbooken sto kommandoen under overskriften
«Hvis du vil vise at outputen er den frosne», altså som et show-element
UNDER demoen, og kjørt der oppdager den regresjonen samtidig med publikum.

Samme defektklasse som frys-gaten (x1 -> x2) og CHANGELOG-datoen (lest på
feil HEAD): kommandoen var riktig, tidspunktet var det ikke.

§6 flytter den til før rommet fylles og legger til tag-ankeret
`git diff --stat v1.0.0..HEAD` med frys-gatens to unntak. MÅLT at ankeret
er fail-CLOSED: en manglende v1.0.0 gir `fatal: bad revision` (exit 128),
ikke tomt — en gate mot en tag som ikke finnes kunne ellers vært stille
grønn. Kommandoen står fortsatt kun ÉN gang i dokumentet (§6 peker på
§1-blokka), så det er ikke laget en andre kopi å drifte fra.

Tatt tirsdag med vilje: onsdagen skal måle og utføre, ikke avgjøre.

Målt: demo på HEAD golden-diff TOM (61 linjer, exit 0) · to kjøringer
byte-identiske · 810 passed / 4 skipped · ruff + format + mypy rene ·
utfyllings-gaten uendret på 3 · kalendertabellens pipe-telling intakt.

Én påstand ble drept av måling: pre-flighten varmer IKKE venv-en
(2,67 s vs 2,77 s), så den setningen ble ikke skrevet.

Planen lukket i samme pass (grep-passen fant tre steder): P4.5-blokka,
tor-13-raden, og tir-11-raden — «TOM — GÅ RETT PÅ ONSDAG» var sant da den
ble skrevet mandag og sluttet å være det samme uke.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0143ESVxVd5srp9PGshHCdi4
2026-08-11 13:03:34 +02:00
c6d7ea7c1b docs(P4.5): §1 lovet en stderr-sti som aldri kan vises
Runbookens §1 sa `(arbeidskopi: /tmp/po-sim-…)`. Målt: `mkdtemp` følger
`TMPDIR`, som på macOS er `/var/folders/…/T/` — `tempfile.gettempdir()`
bekrefter det. Strengen `/tmp/po-sim-` kan altså aldri stå på skjermen.

Det er den dyre varianten av defekten: operatøren ser en lang
`/var/folders`-sti der runbooken lovet `/tmp`, og et sekunds tvil om
miljøet er ett sekund fra en unødvendig abortsti på scenen.

Rettet til FORMEN, ikke til denne maskinens sti — en hardkodet
`/var/folders/xc/…` ville gjenskapt defekten ett nivå ned. Fasiten gjør
allerede nøyaktig dette skillet: temp-katalogen tilhører miljøet,
`po-sim-`-prefikset tilhører programmet.

Funnet ved å måle §1s stderr-INNHOLD mot fasiten (testens egen
`normalise_stderr`) i stedet for bare å telle fire linjer — den ene §1-
påstanden økt 9s pass hadde tallfestet uten å innholds-sammenligne.

Én kopi: planen var allerede korrekt («`TMPDIR`-rota maskeres»).
Utfyllings-gaten står uendret på 3 placeholders; kjørestien er urørt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NaBAXKK1ippN8Zi62DMNoS
2026-08-11 12:50:06 +02:00
4f1c19dfb2 docs: CHANGELOG-datoen må RE-LESES etter stempel-commiten, ikke bare før
Sekvensen onsdag er X → Y (runbook) → les `%cs` → stempel → Z → tag Z.
Ved lesningen står HEAD på Y — og Y blir aldri tagget; verdien skrives inn
i Z. Det holder når Y og Z lander samme dag, men brekker ved midnatt MELLOM
dem: da bærer den taggede commiten gårsdagens stempel. Det er nøyaktig det
tilfellet setningen påberopte seg («også hvis dagen sklir til torsdag morgen»).

Retteslen er ett steg, ikke en omskriving: `git log -1 --format=%cs` én gang
til ETTER commit av Z og FØR `git tag`, med `--amend` ved avvik. §5 har nå ni
avkryssinger. Samme klasse som frys-gaten selv (amendert økt 7): kommandoen
var riktig, tidspunktet var det ikke.

Begge kilder rettet i samme pass — runbookens §5 og planens frys-blokk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LefsEziBhiJgLFbBqTZxnK
2026-08-11 06:44:29 +02:00
818b55ae09 docs: frys-gaten kjøres TO ganger, og CHANGELOG-datoen leses av commiten
Tre retteser fra advisor-review, alle propagert til BEGGE kilder (runbook +
plan + STATE) — en oppskrift som står ulikt to steder er nøyaktig drift-klassen
denne økta har lukket.

1. GATEN KJØRES TO GANGER. Sekvensen er X (prøve) -> Y (runbook) -> Z
   (CHANGELOG-stempel) -> tag. En gate kjørt kun ved X måler en tilstand som
   ikke lenger finnes når taggen settes; bare et andre kall RETT FØR git tag
   beviser at det TAGGEDE treet er det prøvde. Samme kommando, annet tidspunkt.

2. DATOEN LESES AV COMMITEN SOM TAGGES: git log -1 --format=%cs, ikke date +%F.
   Veggklokka er ikke etterprøvbar og kan avvike fra commiten på en sen
   kveldsøkt; %cs gjør CHANGELOG, tag-objektet og commiten enige også hvis
   dagen sklir til torsdag morgen. Målt: %cs på HEAD gir 2026-08-10.

3. PORTEFØLJE-SVARET SNEVRET INN. Raden lovet "ja, den kan kjøre en hel
   portefølje" med et forbehold om CLI-taket. Spørsmålet den besvarer er
   bredere enn det vi har prøvekjørt denne uka: run_portfolio er testet, men
   demoen kjører ett prosjekt og CLI-porteføljestien er ikke prøvd. Raden sier
   nå hva biblioteket har, at demoen ikke viser det, og at porteføljestien ikke
   skal tilbys som live demonstrasjon.

Verifisert: null "date +%F" igjen i noen av de tre kildene; %cs-oppskriften i
alle tre; placeholders fortsatt 3; docs/-gaten grønn (10 passed).
2026-08-10 21:24:48 +02:00
eb631d6ba9 docs(plan): P4.5 lukket i planen — den instruerte om en runbook som nå er skrevet
Samme drift-klasse som økt 4, 5 og 6 hver for seg fant, og derfor lukket i SAMME
økt som arbeidet: planen er dokumentet operatøren følger onsdag under tidspress.

Grep-passen (økt 4s mottiltak) over "runbook|P4.5" fant FEM levende steder, ikke
ett:
- P4.5-blokka: ☐ -> ✔ SKREVET, med amendementet som forklarer hvorfor gaten
  ("VED frysen ... så den matcher frosset output") dekker målingene og ikke
  forfatterskapet
- ons-12-raden: "skriv runbooken" -> "FYLL UT runbooken", med utfyllings-gaten
- Spor 2 (§0): mandat-setningen merket ✔ — den sto der som en PLASSHOLDER for en
  beslutning, ikke som en beslutning
- Spor 2-oppsummeringen linje 63: "VED frysen" -> "SKREVET man 10., FYLLES UT"
- man-10-raden: nytt punkt (7)

Ett premiss rettet i samme pass: raden sa "SEKS ØKTER" og er nå syv.

Historiske oppføringer står urørt med vilje: I5-registeret i §6 og frys-blokkas
"runbook (commit Y)" beskriver beslutninger og en sekvens som fortsatt stemmer —
onsdag committer fortsatt den utfylte runbooken som Y.

Tabell-integritet verifisert etter redigering: hver kalenderrad har nøyaktig tre
usiterte pipes; de tre escapede på ons-12-raden er K1-kommandoens og er
uendret. docs/-gaten grønn (10 passed).
2026-08-10 21:18:41 +02:00
c7a57d8c76 docs(P4.5): demo-runbooken skrevet — beslutningene mandag, målingene onsdag
Runbooken var gatet til "VED frysen onsdag". Gaten gjelder hashen X og verbatim
output-utdrag — ikke forfatter-dømmekraften. Planens egen onsdags-regel er at
dagen skal MÅLE og UTFØRE, ikke avgjøre; en runbook skrevet fra bunnen på en
enveis-dag under tidspress er nøyaktig det den regelen forbyr. Derfor to-trinns
med vilje: alt kjennbart nå, tre målte felt onsdag.

Samler de fire spredte kildene I5 navnga (demo-uke-plan §1 · innholdsgate §5
JA-varianten · P4 pkt. 4 · §0 Spor 2) og forankrer hver setning i en LINJE i det
pinnede transkriptet, så operatøren finner stedet uten å lete.

MANDAT-SETNINGEN ER SKREVET. Den sto i Spor 2 som "én muntlig setning" og fantes
ikke som tekst noe sted — en udraftet setning til en live demo. Nå formulert mot
docs/bestille-en-kjoring.md: bestillingen styrer hva som VURDERES, aldri hva som
GODKJENNES.

UTFYLLINGEN ER GJORT TIL EN SJEKKET STEG, ikke en husket. Tre placeholders med
greppbar form, og grep-en er SELV-SIKKER: monsteret '<<[A-ZÆØÅ-]*>>' matcher
ikke sin egen tekst (målt: 3 treff, ingen av dem kommandolinjene). Et uutfylt
felt er samme drift-klasse som plan-radene økt 4, 5 og 6 hver for seg fant.

VEDLEGGET FELLER TRE STATE-PREMISSER. Alle tre "scene-kosmetiske" punkter er
målt mot det pinnede transkriptet, og INGEN er synlige:
- 23700 NOK/aar: rationale er 389 tegn, beløpet står ca. tegn 370, demoen
  klipper på 300 -> linja ender "pga. overes…". grep -c "23700" -> 0.
  STATEs "printes fortsatt ordrett" var et premiss, ikke en måling.
- 0.82 hører til bygg-goldenen, ikke veglys. grep -c -> 0.
- docs/ekspert-svar.md leses ikke av demoen.
Torsdagen slipper altså tre setninger den var fortalt at den måtte bære.

CHANGELOG-DATOEN GJORT TIL EN MÅLING: sjekklista sier "les datoen på dagen"
(date +%F), ikke det forhåndsskrevne 2026-08-12 — samme premiss-klasse som
hashen planen allerede nekter å skrive ned.

Frysen er IKKE flyttet fram. Tirsdagen er tom og fristet, men onsdag er en
dato-beslutning på en enveis-handling (tag + frys), og risikoen den ville hedget
er allerede retirert: økt 6 målte samme kjøresti grønn, og d0e8bb0..HEAD er
dokumenter alene.

Målt: 810 passed / 4 skipped uendret. docs/-gaten passert via datert sti.
Kjørestien urørt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011GtvZy6hn3k2iTFGjRVzLi
2026-08-10 21:16:10 +02:00
c4b0e08cc5 docs(plan): tirsdagen lukket i PLANEN — den instruerte om arbeid som var gjort
Planen er dokumentet operatøren FØLGER onsdag under tidspress, og etter
persona-pullen instruerte tir-11-raden fortsatt om en betinget subtree pull med
abortsti og 18:00-frist — arbeid som er utført og pushet. Samme drift-klasse økt
4 felte, så samme mottiltak: én grep-pass FØR redigering
(`persona|kontorbygg|tilsvarende anlegg|I4|18:00` over `docs/plan/` +
CHANGELOG), amendering med attribusjon (§6-mønsteret), aldri omskriving.

Grep-passen fant FIRE steder, ikke ett — som er hele grunnen til å kjøre den:
- tir-11-raden: betinget pull-instruks → **TOM, gå rett på onsdag**, med
  utfallet og de fire målingene som lukket den
- frys-blokkas X-note: «lander tirsdagens persona-pull, flytter X seg» →
  pullen ER landet (`d0e8bb0`); X leses fortsatt av `git rev-parse HEAD` etter
  grønn prøve, aldri skrevet ned her
- P4 pkt. 1s fersk-klon-måling: sto på `c9787cf`, altså to commits bak etter
  pullen. HOLDBARHETEN sagt eksplisitt i stedet for underforstått — delta er
  prosa i `shared/` + regenerert fasit, `pyproject.toml`/`uv.lock` MÅLT urørt,
  så målingen står; onsdagens generalprøve ×2 er bekreftelsen
- P3s I4-abortsti (i `<details>`): stemplet HISTORISK, med den ene målingen
  verdt å bære videre — den harde reset-formen blokkeres av hooken, `--keep`
  slipper igjennom

CHANGELOG: én Changed-linje for persona-formuleringen, slik at onsdagens
`[Unreleased]` → `[1.0.0]`-stempel ikke beskriver en artefakt-tekst som har
endret seg siden. Beslutningen tas her, ikke på en enveis-dag.

Tabell-integritet verifisert (pipe-telling; ons-12-radens seks er tre escaped
`\|` i grep-kommandoen, urørt). `test_doc_constant_sync_loadbearing` grønn.
Kalenderen: man 10. er nå SEKS økter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017og6HMcP1WQcABRDogUMfx
2026-08-10 20:56:19 +02:00
5135f099e3 docs(plan): frys-gaten snudd til eksklusjonsform — den feilte OPEN
Første form listet kjørestien positivt (src/ tests/ shared/ pyproject.toml
uv.lock). En slik gate er blind for alt den ikke rakk å regne opp: målt på
1522e2a^..1522e2a rapporterer den to filer og slipper README.md OG CLAUDE.md
rett igjennom, mens eksklusjonsformen tar alle fire. På en enveis-dag skal en
gate feile lukket — en ny fil skal trippe den, ikke passere fordi ingen forutså
den.

Unntakene er nøyaktig det onsdagen skal skrive: docs/ (runbooken) og
CHANGELOG.md (stemples i samme trekk som taggen, så en gate som dekket den kunne
aldri blitt grønn). STATE.md er gitignorert og kan ikke dukke opp i en diff.

Begge de opprinnelige målingene re-kjørt på begge former med samme svar —
c255662..HEAD gir fire filer, a41272d..HEAD tomt. Funnet i advisor-review.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVmQfqWWEVTQnv7PkuLE9U
2026-08-10 14:59:28 +02:00
6ae3bc7c49 docs(plan): P4-hashen inn i ✔-en (planens egen regel, linje 10)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVmQfqWWEVTQnv7PkuLE9U
2026-08-10 14:53:57 +02:00
d306929c28 docs(plan): frysedagens to udefinerte steg lukket — FRYS gjort kjørbar, P4 lukket
«FRYS» hadde ingen operasjonell definisjon (40 treff i åtte plandokumenter; eneste
forsøk er demo-uke-planen linje 93, som er en regel, ikke en handling). Sekvensen er
prøve (X) → runbook (Y) → stempel+tag (Z), så taggen lander på Z mens prøven målte X.
Frysen er nå to kommandoer: noter X, og kjør frys-gaten før taggen. Gaten er målt at
den diskriminerer — c255662..HEAD gir fire filer (versjonssynken landet ETTER
generalprøve nr. 0), a41272d..HEAD gir tomt. CHANGELOG.md bevisst utenfor gaten.

P4 lukket: onsdagsraden sa «P4 re-målt» uten å si hva. Målt var det fersk-klon-
kriteriet (pkt. 1), stående på ab7f45a — elleve commits tilbake, før subtree-pullen,
P3/GO, det pinnede transkriptet, innholdsgaten (ny runtime-dep) og versjonssynken
(endret uv.lock). Re-målt på c9787cf: uv sync exit 0 med treet urørt, 810 passed /
4 skipped, ruff+mypy rene, demo-stdout byte-identisk mot goldenen, 61/4 linjer,
K1 distinkt 8.

Onsdagen har dermed fire steg, alle kjørbare. Null kodeendring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVmQfqWWEVTQnv7PkuLE9U
2026-08-10 14:53:07 +02:00
c9787cf10c docs(plan): frysedagens instrument gjort sant — to open/-instrukser felt, P2 lukket
Planen er dokumentet operatøren følger onsdag under tidspress. Den instruerte
fortsatt om en beslutning som ble felt 08-10, og på TO steder — nøyaktig
drift-klassen der en retting i én fil etterlater den i en annen:

- §0 S1.c-raden: «tag v1.0.0 på begge remotes» → taggen går til `origin` ALENE.
  Felt med målingen som felte den (open/main = 520e741 = v0.1.0; 26 commits =
  53 filer / 6087 innsettelser; seks plandokument-beslutninger, én åpen sak +
  fire aldri vurdert). Amendert i stedet for omskrevet, per §6-mønsteret.
- P4 pkt. 1s fresh-clone-notat: «publisering dit er S1.c onsdag» → korrigert,
  flyttet til P5-vinduet.

Bokføring i samme pass (planens egen regel, linje 10):
- P2/S1.b ☐ → ✔ (2026-08-09, c255662), både overskrift og Spor 1-tabellen;
  JA-varianten i ærlighets-teksten markert som den som gjelder.
- Kalenderen: man 10. = tre økter (var: kun generalprøven) · tir 11. = P3 er
  gjort søndag, tirsdag er commons-svaret eller tom · ons 12. = frysesekvensen
  i rekkefølge, med distinkt-tellingen (8, ikke 9) og taggen sist.

«ETT trekk» på tag-dagen er MÅLT inn i dokumentet, ikke antatt: `## [Unreleased]`
er verbatim og unik (1 treff, kun to `## `-overskrifter), ingen link-refs å
følge med. Sto den to steder, ble frysedagens ene trekk improvisasjon.

Null kodeendring; ingen test rørt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QLpSfvCgBLmc3tPRMnr1JA
2026-08-10 05:15:36 +02:00
9708c15d07 docs(plan): S1.c synk + CHANGELOG lukket — taggen står igjen, frys-gatet
Forskuttert fra onsdag kveld, fordi den delen ikke er frys-gatet: kun taggen er.
Raden bærer nå de tre tingene målingen avgjorde — hvorfor overskriften står på
[Unreleased], hvorfor uv lock måtte kjøres eksplisitt og diffes før noen test, og
at fire versjonssteder ER alle fire.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ue1AnPZYsC9Tk7e5Tyv8Fv
2026-08-10 04:50:01 +02:00
887a8be677 docs(plan): generalprøve nr. 0 bestått — og kriteriets telling er foreldet, ikke demoen
Prøven kjørte mot LEVERT VEGLYS-FV-SOER, ikke den forankrede reserven planen
forutsatte: P3 falt to døgn før fristen, så prøven målte demoinnholdet selv.
Det er en strengere prøve enn planlagt — reserven validerer mekanikk, aldri
presentasjon — og reserve-stien er fortsatt målt, i suiten.

Alt grønt, null kodeendring: golden-diffen tom for begge kjøringene, K6
selv-identitet tom på stdout med kun po-sim-suffikset ulikt på stderr,
810 passed / 4 skipped, ruff og mypy rene, goldens uendret.

Ett avvik, og det ligger i kriteriets bokstav: §5 pkt. 1 teller
`grep -cE "^ *Steg [1-8]"` = 8, men P1/S1.a ga Steg 7 to merkede linjer, én
per tidsskala, så tellingen gir 9. Åtte distinkte steg står — intensjonen er
oppfylt. Kriteriet rettes ikke her: en gate justert i samme økt som den
feiler er ikke lenger en gate. Onsdagens generalprøve ×2 bruker
distinkt-tellingen, og §5 pkt. 1 rettes etter frysen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ATCeqRyvdL9qmk34o6HUCa
2026-08-10 04:34:05 +02:00
c255662802 feat(ingest): P2/S1.b — innholdsgaten står, rundt materialize og ikke i den
Planens §3 sa at `ingest.materialize` er repoets ENE skrivepunkt på Door A, og
det premisset ble felt av måling FØR bygging: `materialize` er en ren delegasjon
til pinnet llm_ingestion_okf v0.3.2s `materialize_bundle`, som stager i minnet og
utfører sin egen disk-fase. Det finnes ingen callback mellom de to, så en gate
plassert der kunne bare kjørt ETTER at bytene landet — en opprydding, ikke en gate.

Sømmen ble i stedet kopier bundelen → materialiser inn i kopien → skann det som
ble generert → publiser eller forkast. Kopien er bærende, ikke bekvemmelighet:
bibliotekets §3 eierskaps-skann, kollisjonsgaten mot kuratert innhold og §6
index-merge leser alle den EKSISTERENDE bundelen. Staging i tom katalog mister
alle tre og publiserer en bundle uten kuraterte naboer — datatap forkledd som
sikkerhetsfiks.

De fire §4-beslutningene, tatt og målt: (1) ingen av guardens to preset —
Origin.EXTERNAL/AUTOMATIC, fordi trust_for utleder policy fra origin alene og
PRESET_USER_UPLOAD bærer en quarantine-semantikk Door A ikke har; (2) utfall per
BUNDLE, diagnostikk per DOKUMENT — delvis publisering ville etterlatt bundle +
index som svarer til intet manifest, men import_bundle itererer forbi første
avvisning; (3) Report til log.md, aldri konsept-frontmatter, der fire golden-suiter
pinner bytene; (4) mypy-override OG adapter, siden override alene gjør sømmen
type-blind i stedet for type-sikker.

`materialize` forblir ugatet med vilje — goldenene pinner den, og en kaller som
vil ha gaten ber om den ved navn.

Fem mutasjoner alle røde + grønn kontroll (hele suiten, ~120 s hver): detach
gaten · la den fyre ETTER publisering · Origin.INTERNAL · tom staging-katalog ·
rapporter kun første avvisning.

Målingen felte en VAKUØS test først: en hard injeksjon scorer fail_secure under
BEGGE trust-tierene, så Origin.INTERNAL-mutasjonen lot alle tre avvisningstestene
stå grønne — beslutning 1 så dekket ut uten å være testet. Båndet der tieren
faktisk avgjør er høy-entropi-innhold (quarantine_review vs warn), og testen ble
skrevet mot nøyaktig det før mutasjonen ble re-målt. Mutasjon 4 ble på sin side
felt av KUN én test; 809 andre merket ikke at bundle-kopien forsvant.

Laveste disposition er `warn`, ikke `allow` — `allow` finnes ikke i guarden. En
gate skrevet mot == allow ville avvist hvert dokument som noensinne ingestes.

Kriterium 5 står: demo-stdout er byte-identisk med tests/golden/demo-transcript.stdout,
målt både i suiten og ved eksplisitt kjøring. shared/ er urørt.

801 -> 810 passed / 4 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DDu94KoyxAmhJsG2n63X8Q
2026-08-09 22:58:00 +02:00
50232fb88d feat(simulation): P4 pkt. 3+4 — demo-transkriptet pinnet, frø-setningen avledet [skip-docs]
Kriterium 6 er selv-identitet: to kjøringer av en regredert demo er like enige som
to av en riktig. Fasiten forlater derfor prosessen. stdout pinnes ORDRETT (og er
dermed demoens abortsti); stderr normaliseres på nøyaktig to MÅLTE miljø-spann —
site-packages-prefikset og temp-katalogen — med po-sim- holdt synlig, fordi det er
en egenskap ved programmet og ikke ved miljøet. Pinnet stderr = fire linjer.
Kontrollen som forbyr at masken vokser er load-bearing: en droppende normaliserer
med fasiten regenerert under seg holder BEGGE likhets-testene grønne.

Pkt. 4: planens forhåndsskrevne frø-setning sa «én av de TO tidligere dommene».
Målt mot levert VEGLYS-bundle henter Kjøring B TRE — én fulgte med kunnskapsbasen,
to er demoens egne, én per tidsskala. Splitten avledes derfor fra kjøringen; en
håndskrevet «én av tre» ville vært den andre kopien som drifter.

Fem mutasjoner alle røde + grønn kontroll (hele suiten hver gang): ett byte i en
stdout-linje · detach dempingen · over-normaliser stderr · literal splitt · detach
frø-setningens print. Byte- og detach-mutasjonene ble fanget av KUN golden-testen;
den literale splitten av KUN skille-testen.

793 -> 801 passed / 4 skipped.
2026-08-09 22:12:25 +02:00
7aef1feaa4 docs(plan): P3 lukket — GO, felt to døgn før fristen
Abortstien fulgt i rekkefølge, kriterium 8 målt to uavhengige veier, ingen reset.
Måletallene som avgjorde: P90 = 1 769 915, overdrivelsen 2 100 000 over begge terskler,
10 %-prøven mot LEVERT baseline feller i stage 0. Fem mutasjoner røde.

Planteksten slik den sto før utførelse er beholdt i en <details>-blokk — NO-GO-grenen er
død tekst nå, men reserven er fortsatt abortstien og skal kunne leses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BUjfw4eJdwwnqHSXhfcY6i
2026-08-09 21:32:57 +02:00
7acd331e31 docs(plan): P4 pkt. 5, 2 og 1 lukket — og pinningen i pkt. 3 har fått sin egen måling
Punkt 5 (entry points), 2 (stderr) og 1 (fresh-clone) er merket ✔ med en UTFØRT-blokk
som bærer beslutningen og belegget, ikke bare utfallet.

Punkt 2 var øktas åpne beslutning, og den ble avgjort ved måling framfor preferanse:
rund-taks-linjene dempes, ExperimentalWarning-paret gjør det ikke — de fyrer før
simulation i det hele tatt importeres, så demping ville krevd et warnings-filter inne
i bibliotekpakken.

Punkt 3 får en konsekvens fra fresh-clone-målingen: fasiten kan ikke være literal.
De to gjenstående stderr-linjene bærer en absolutt sti inn i site-packages, som er
ulik i klon og arbeidskopi — normaliser på BÅDE den og po-sim-suffikset. Pinnet
stderr blir fire linjer, ikke to.

[skip-docs]

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C2bxLcCRguxXzpM4priTMn
2026-08-09 15:21:20 +02:00
1522e2aaaa feat(simulation): the demo's gate is anchored to real cost lines (P4 pkt. 0)
The validator can reconcile a proposal against the project's actual cost lines
(S4.0 stage 0), but only when the knowledge base ships a cost-baseline.json —
and no bundle under shared/ has one. So on stage the gate reasoned only about
numbers the proposal supplied itself.

The reserve can never receive the file in shared/ (pull-only subtree, and demo
criterion 8 requires the goldens byte-unchanged). That is a placement
constraint, not an impossibility: materialize_anchored_bundle copies the bundle
and adds the file outside shared/, and the run path reads it through exactly
the seam a delivered bundle would use.

The baseline is DERIVED IN CODE from the scripted register, never typed beside
it — two sources of the same numbers drift, and drift is precisely what the
10 % probe models. On GO day the direction reverses (plan P3 b). Both scripted
replies must state the same cost lines or ValueError: were they to differ,
hypothesis #1 would be falsified by stage 0 instead of by P90 — the same
REJECTED line on screen, a different mechanism behind it.

10 % probe, measured: baseline x 1.10 -> FORKASTET at stage 0, before the
solver; corrected -> FORESLÅTT. Criterion 6 re-measured (stdout byte-identical
across two runs); stderr unchanged at 6 lines. The ONLY diff against the
un-anchored demo is the new KUNNSKAPSBASE block — everything else is
byte-identical, which is the problem: an anchoring nobody can see is one nobody
can check. Hence it is printed, and hence `provenance` is a required argument.
769 -> 775 passed.

Five mutations red + green control. The measurement failed the TEST first:
"ingen kostbaseline erklært" CONTAINS "kostbaseline erklært", and
ENERGI-TOTAL-EL already appears in the Step-2 line, so both assertions survived
the detach mutation. The two branches now share no wording.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GD6Y2Y23NZZxPYtSRoCmst
2026-08-09 14:35:07 +02:00
1e11dcb96c feat(simulation): the demo now RUNS the Step-7 file inbox it narrates (P1/S1.a)
The Step-7 trace line said "lang fil-løkke" while the verdict arrived as a
function argument (`verdict_input`) — the short, in-run capture. The long loop
was tested but never exercised by the thing on stage.

An expert now drops a real verdict FILE (`write_verdict`) into an inbox between
the runs, and Run B is given `verdict_dir=`, so `run_project` merges it into the
store before the Step-1 fold.

Not done as the plan point was worded, and the difference is load-bearing:
routing the PERSONA verdict through the inbox would have put ONE marker on two
paths — Step 7 (inbox) and Step 8 (promotion) both end in Run B's prompt, so
either could carry it alone and `test_simulation_loadbearing.py`'s promotion
assertion would have stayed green with promotion detached. A second verdict with
its own marker keeps both seams independently red-able; `simulate_learning_loop`
raises when the two markers are equal. The inbox sits beside the bundle copy,
never inside it, and the id is an explicit sentinel (a minted id would collide
with the promoted verdict's, and `VerdictStore.add` is first-write-wins).

766 -> 769 passed (773 collected). Criterion 6 re-measured: stdout byte-identical
across two runs; stderr unchanged at 6 lines. Mutations measured against the full
suite, four red + a green control: detach `verdict_dir=` · point Run B at an empty
folder while the file is still written · marker set to `realization_rate: 0.82`
(measured present in the verdict seed) · marker set to `energy performance gap`
(measured present in a navigated concept file) · benign rename of the inbox dir.

Honesty limit found while measuring: the last two mutations fell on the causality
assertion, not the Run A control — generation prompts carry the debate output, not
the bundle context. The pair holds, but each assert defends a different property.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVYDeJ9evZicgU5r3roZVW
2026-08-09 13:13:21 +02:00
d0571ca408 docs(plan): six objections measured, all taken in — the anchoring risk moves to the weekend
I1: Funn 1 was measured one directory wide; the repo ships a working S4.0
baseline fixture and run.py:516 reads it. The anchored dry-run + the 10%%
deviation test move from Tuesday to the weekend (P4 pt 0); Tuesday becomes a
re-measurement with an explicit abort path (I4: pre-pull hash, reset rule,
18:00 NO-GO). I2: stderr damping decided YES, in the weekend BEFORE pinning —
measured today stderr is six lines, one deliberately non-deterministic. I3:
[project.scripts] moves off freeze day to before the fresh-clone measurement.
I5: a demo runbook post (P4.5) at the freeze. I6: every §4 claim re-measured
today on HEAD bb3df79; the 08-09 datings were commits from 2026-08-06.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UDHSsyMuBASJcRapddciHL
2026-08-07 16:55:30 +02:00
bb3df79204 docs(plan): six objections to the week plan, as a prompt the next session must measure before acting
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1rEDj3QLNasKSceWFyE6k
2026-08-07 16:39:42 +02:00
f49a4d263b docs(plan): work starts Friday with full quota — front-load everything that needs no new content
Operator directive: full week available, weekend included, new quota,
high priority. The calendar now starts Friday with P1, pulls the whole
P4 advance (fresh-clone criterion, golden transcript, stderr muting,
both honesty sentences) into the weekend against the micro reserve, and
makes Monday dress rehearsal #0 — the NO-GO outcome is fully verified
BEFORE Tuesday's GO gate, leaving Tuesday/Wednesday thin: pull+measure,
re-measure, freeze, release cut, tag.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019xhQpH4oQBaf8dxCkXuB8Z
2026-08-07 08:20:06 +02:00
33f0a6857d docs(plan): §0 splits the week into two tracks — complete v1 by Thursday, and the demo
Operator decision 2026-08-09. Track 1 (complete v1, incl. other repos):
S1.a = P1 step-7 inbox, S1.b = P2 content gate, S1.c = release cut
(1.0.0 synced in four places, CHANGELOG, [project.scripts] moved in from
P9, tag only AFTER a green dress rehearsal). Other-repo accounting is
measured: commons already ordered with the Tuesday deadline and a
reserve, okf/guard/po-claude need nothing — no new coord message. Each
post carries a named degradation so v1 stays honestly complete at every
level. Track 2 (convincing demo): P3 + P4 + rehearsal + one spoken
mandate sentence, optional stderr-noise muting before the freeze. The
O4-vs-tag conflict is flagged for the operator, not decided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019xhQpH4oQBaf8dxCkXuB8Z
2026-08-06 23:30:01 +02:00
0eb0f3d72b docs(plan): the review's findings as a ranked plan the NESTE block walks, one point per session
Fable-review 2026-08-09 made durable: P1-P10 in plain language with the
commands behind every number (evidence table §4). Pre-demo: step-7 inbox
wired into the walkthrough (P1), Spor B sharpening (P2), the stage-0
first-contact check on Tuesday's GO (P3), fresh-clone/stderr/golden
criteria plus two honesty sentences on Wednesday (P4). Post-demo: CLI
portfolio cap (P6), one consolidated commons amendment (P7), method
skill [Voyage] (P8), and an explicit NULL for orchestration swaps (P10).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019xhQpH4oQBaf8dxCkXuB8Z
2026-08-06 23:13:23 +02:00
c96ef9032d docs(plan): the Fable prompt, with the feature set as its main track and quota as its output
A prompt that lives only in a conversation dies at /clear, so it goes in the repo.

Three tracks, in the order the operator weighted them: the MAF feature set, the demo, and -- as
the actual deliverable rather than an appendix -- a ranked list of what the week's quota should
buy. Each item carries a mechanism, a hard [FØR TORSDAG]/[ETTER DEMOEN] tag, a cost in SESSIONS
rather than hours, and what would go red if the item were done. An item nothing can falsify is an
opinion, not a finding.

The measured starting points are embedded so the session does not re-derive them wrongly: two
debate agents rather than three, one orchestration in use out of the installed surface, a
hand-rolled portfolio fan-out, and a capability map organised by NEED that therefore never
compares TOPOLOGIES. The map is not stale on version -- 1.9.0/1.0.0 is what is installed -- which
matters, because "the map is old" would be the easy wrong conclusion.

The prompt carries its own discipline because Fable runs without an advisor: every figure must be
produced by a command shown beside it, and premises in STATE and in plan documents are named as
premises. This repo has measured at least three of them wrong, most recently today.

It is also forbidden from smuggling feature work in front of the demo. The demo is a hard date;
the feature set is not. Arguing otherwise is allowed -- but only out loud, with the consequence
spelled out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoHJCKBTjFKcjsfEQyGbzh
2026-08-06 20:01:53 +02:00
295e9665fe docs(plan): the content gate and the honesty sentence, on a track that does not touch the freeze
The demo shows "download and run". Implying you can point this at your own sources and build a
knowledge base claims three things the code does not carry -- and A5 (the code may not claim more
than it does) binds the presenter too, not just the source.

Measured first, and one measurement changed the plan: the guard is NOT v0.2 alpha. That figure came
from our own 2026-07-16 inclusion plan, which is a premise rather than a fact. It is v0.3.4, seven
published tags, `dependencies = []` -- stdlib only. Our okf pin (v0.3.2) declares no dependencies
either, so the guard is not coupled to it, and the 0.3.5-vs-0.4.0 release argument concerns the
release AFTER v0.3.4. Adoption moved from risky to tractable on that one reading.

The three claims, made precise: the demo bundle was hand-curated (honesty), the ingest path writes
unscanned (buildable), and the generic bundle factory does not exist (deferred at O1, not buildable
in four days). Two close with code, one with a sentence.

The two tracks are separated on a measured fact: `simulation.py` does not import `ingest`, so Door A
work cannot disturb what Wednesday freezes. Criterion 5 is the one that proves it -- the walkthrough
must stay byte-identical.

Four decisions are named as decisions rather than settled silently: which policy preset, fail-closed
versus flag-and-write, where the guard's report lands in provenance, and keeping `--strict`
meaningful across a seam that ships no py.typed. The honesty paragraph is written in BOTH variants
up front, so Wednesday is an observation and not a judgement call on stage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoHJCKBTjFKcjsfEQyGbzh
2026-08-06 19:49:13 +02:00
d6f3359fae feat(step5): the falsification that informed the next hypothesis now leaves the loop
generate_via_llm consumed each validator Rejection internally (`last`), fed it into the
next attempt's prompt, and dropped it. So Step 5 was real but unobservable: a caller could
see THAT a proposal validated, never that it validated on attempt 2 after the deterministic
validator falsified attempt 1. It was the one step of the eight with no output to show.

The seam is a typed return value -- GenerationResult(outcome, refinements) -- rather than an
out-parameter or a callback: a returned value cannot be silently lost by a caller that forgets
to pass a collector, and mypy forces every call site to acknowledge it.

refinements carries ONLY rejections that were actually fed back. When the attempt budget runs
out the final rejection IS outcome; counting it here would be double-counting, and the bounded
control test goes red on the collect-everything implementation that gets this wrong.

The loop's bound is untouched: max_attempts and meter.tick_round stand, and `last` still drives
the prompt alone, so prompt growth is unchanged. run.py accumulates across _evaluate calls, so
_evaluate_mandate is untouched; RunResult.refinements defaults (the coverage precedent) and is
concatenated across approaches rather than keyed per approach -- stated as an honesty limit.

The simulation now shows it: the scripted proposer overclaims 250000, which the validator
falsifies against P90 = 90000, and the corrected 30000 validates. Only the overclaim is
scripted -- the rejection is computed. scripted_factory takes a per-role reply selector so this
needs no second scripted client body.

README records the two accuracy changes only (Step 5 is now inspectable; the simulation trace
shows the correction). The level-2 publishing claim stays deferred until after the demo (O4).

Load-bearing MEASURED against the full suite with a control, four mutations all red:
detach the returned history (4 tests) - collect-everything (control only) - detach the run
wiring (2 tests) - revert the simulation's proposer to a constant (the demo-protection test).
Control: 759 passed / 4 skipped; ruff, format and mypy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CcWFcREUi6YPjEpN3ACDP
2026-08-06 15:12:06 +02:00
cd011c4ac7 docs(plan): demo week, with the one real build separated from the presentation
Seven of the eight steps already have their data in RunResult and need a print;
one does not exist at all. Putting that distinction in a table is the point of
this plan -- it turns "show all eight steps" from an unbounded week into one
build on Friday and presentation work over the weekend.

The go/no-go on Tuesday is deliberate. The content is being built in another
repo on a deadline nobody here controls, so the week is designed to survive it
not arriving rather than to hope it does. The content-keyed reply selector
lands Monday, before the content, for the same reason: a new project should
then be a data entry rather than a hand-written script under time pressure.

Honesty framing is section 1 rather than a footnote, because the demo's own
subject is a system that refuses to claim more than it proves.
2026-08-06 13:23:19 +02:00
3313e9dcaa docs(qa): the four decisions the QA surfaced, with what constrains them
The 20 claims went un-corrected, so they stand as confirmed. What the operator
actually decided were the four choices the QA exposed: hand-built example with
the factory path explicitly deferred, step 5 built and shown live, commons
ordered with a fallback, README after the demo rather than before.

Two measurements are recorded because they bound the order, not because they
are interesting: bundle_context renders every navigated file's full body, and
summary-first reading is not built -- so the full 15-30 measure library would
put 40-90k characters into every hypothesis prompt. The order is size-capped
for that reason and says so.

Also recorded: simulate_learning_loop already takes the bundle directory as a
parameter, so new content plugs into an existing seam. The cost is the scripted
replies, which are written against the LED case.
2026-08-06 13:20:52 +02:00
8ecfa96934 docs(qa): the repo's intention, stated as claims the operator can correct
The demo-week brief was written by a session that read its way to the
intention through documents other sessions had written. Two of its frames
were overturned by the primary sources inside one conversation, so the
operator stopped planning and commissioned this: read the primary sources
directly, state the understanding back as numbered claims, and capture the
corrections where they survive.

Six gaps in the picture the brief rests on, all measured rather than argued:

- The intention has a SECOND axis that STATE's list of five primary sources
  never named. review-2026-07 (F1-F14) and sesjonsplan-fase2-6 (S2.0-S5.4,
  D-A-D-I, M1-M3) are where most of the repo's 31 modules come from: 20
  S-numbers, 18 with hits in src/+tests/. A plan written from the five named
  sources alone would describe a repo with eight steps and miss two thirds
  of what is there.
- D-H's DECIDED demo path ("clone -> unzip -> factory builds -> loop runs")
  is factory-dependent, and the factory (D-G/T0, `okf-toolkit`) does not
  exist -- measured, not assumed. The brief's "anyone who downloads the repo
  can run exactly the same" IS that path.
- The realistic example's content model is already decided (D-F): knowledge
  types with required source citation, strict separation from the verdicts.
  The commission to commons must reference it, not invent one.
- The demo is the programme's level-2 publishing proof (D-I), with an
  honesty ceiling agreed in advance and a README update as its consequence.
- The shared spec covers the loop + ingest and NONE of the surplus: mandate,
  notify, ledger, value report, cost simulation, dimension, portfolio
  budget, concurrency, preflight all measure 0 mentions. The comparison is
  therefore of the SPEC'd core, not of this repo.
- Step 5 is not a presentation-layer concern: `generate_via_llm` consumes
  the intermediate rejection internally, and today's demo validates on the
  first attempt, so the refinement never triggers. Steps 2 and 6 ARE
  printable from data RunResult already carries.

Two inventory numbers spot-checked independently (759 collected; the offline
simulation re-run, output identical). Nothing here is sourced from STATE.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TvjgY5NBg16D7kgQf14s6B
2026-08-06 13:09:45 +02:00
991131be3f feat(provenance): a run records which external service it actually called
The egress declaration (Trekk B3) says what a run MAY contact. It cannot say what
it DID: after the run, nothing distinguished "the agents queried the price
register" from "the agents ignored it", and a proposal resting on an external
service should be traceable to it.

ToolCallRecorder(FunctionMiddleware) mirrors BudgetMiddleware(ChatMiddleware) one
layer down — that one observes the debate's chat calls, this one its tool calls.
It observes only: call_next is always awaited, so a trace can never alter the run
it traces. The record lands on ProvenanceStamp.external_calls, read AFTER the
debate so it is a record rather than an intention.

MEASURED, not assumed, before any of it was written: FunctionMiddleware fires for
a tool served over a REAL MCP stdio subprocess, and context.function.name carries
the BARE tool name with no server prefix. That measurement decided the design —
MAF cannot tell us which server a tool came from, so attribution comes from our own
config, and a name allowed by two servers is recorded UNATTRIBUTED (server="")
rather than credited to the first match. Naming a service that may never have been
contacted is the one place a guess must not go.

Only CONFIGURED tools are recorded. The middleware fires for every function the
agents invoke, including the in-process retrieve_cost_docs on the road path;
logging those would turn the record into a false egress claim. An empty list is a
positive statement — nothing outside this process was contacted — which is why it
is always serialized rather than omitted.

Honesty limit, written on ExternalCall itself: this is the call and its source. It
is NOT evidence that the service's answer reached the proposal, nor a verified
rendering of that answer.

One finding, and it is the reason for measuring rather than trusting green: the
road-path negative test was VACUOUS. Its scripted tool call named an argument the
tool does not declare (code vs query), MAF rejected the call before invocation, and
the test asserted an empty record against a run where no tool ran at all — green
under the exact mutation it existed to catch. It now spies on the recorder and
asserts the invocation genuinely reached it before asserting it was not recorded.
This is last session's lesson again: a scenario that cannot distinguish two
implementations proves nothing.

The tool-call double is registered in test_scripted_client_consolidation.py's
_DELEGATING_OVERRIDES — it cannot live in the reply_selector seam, which returns a
reply STRING, and a response that is not text is its whole subject.

Load-bearing MEASURED (tests/test_b4_mcp_call_trace_loadbearing.py) against the
whole 755-test suite, four mutations all red: detach the recorder from the debate
middleware · record every function invocation · attribute an ambiguous name to the
first server · stop reading the recorder into provenance. Control: a run with no
configured servers records nothing, so the empty record is a real answer and not
the only one the seam can produce.

Ran it, not just tested it: the real recorder against a real MCP server subprocess
returns ExternalCall(server='prisregister', tool='lookup_unit_price'), and a
scripted CLI run's outbox artefact carries the empty list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VtRd8y1PDPGwkrRXFhubqr
2026-08-05 21:37:29 +02:00
455d93d33e feat(outbox): every evaluated approach becomes something an expert can judge
A run commissioned to evaluate three approaches wrote ONE proposal artefact, so
only the approach it selected could ever receive a verdict. The other two were
evaluated, reported in the settlement, and then taught the learning loop nothing.

The defect class is a key collapse, and it had two halves — fixing either alone
leaves it intact:

* the WRITER wrote one pair per run, so the non-selected approaches never existed
  on disk;
* the READER (hitl._read_outbox_proposals) joins proposal to outcome on the
  run_id FIELD read from file CONTENT, never the filename. Three files sharing
  one run_id collapse onto one dict key, last write wins — so widening only the
  filename would have produced three artefacts and still one pending row. This is
  the S3.2 collision class: two rows under one key silently become one.

Artefacts are now keyed {run_id}-{approach_id}-*.json AND carry approach_id in the
payload; the join key is (run_id, approach_id). Two properties make them genuinely
judgeable rather than merely present:

* verdict_id is minted per approach (verdicts.verdict_key, the S3.2 content hash)
  — reusing the run's single id would let one delivered verdict clear all three
  from the queue;
* provenance.validator_decision follows ITS OWN approach — the run's stamp would
  report a rejected candidate as validated, and nothing downstream could correct it.

verdicts.verdict_key is public so a run can stamp the key a verdict WILL arrive
under without capturing a decision nobody has made; it delegates to _mint_id
rather than restating the hash (the (p) rule: one keying rule, one copy).

The per-approach set REPLACES the run-level pair rather than joining it — the
selected approach is already among them, and writing both would count it twice in
hitl pending. The selected one carries the run's final outcome, so the outbox can
never disagree with the RunResult; the others carry the validator's verdict, the
only falsifier that ran on them.

mandate.py is deliberately untouched: hanging a ValidatedProposal off a coverage
row would drag validator — and pulp — into a module kept to pydantic+stdlib for
D7 portability, so _evaluate_mandate returns the evaluated outcomes alongside.

Ran it, not just tested it: a real CLI run wrote six artefacts and hitl pending
listed three rows. It also showed the honest edge — three approaches that produce
an identical candidate share one content-hash key, so one verdict settles all
three. That is correct (they were one candidate), and it is now documented.

Load-bearing MEASURED (tests/test_a5_per_approach_artifacts_loadbearing.py) against
the whole 750-test suite, five mutations all red: detach the per-approach writer ·
drop approach_id from the join key · reuse the run's verdict id · reuse the run's
provenance stamp · widen the filename but not the payload. Control: on a full
detach exactly the 5 new tests fail and 745 pre-existing ones stay green — the
no-mandate path is inert, and writes neither the filename segment nor the field.

Docs: bestille-en-kjoring.md (what the commissioner gets) + ekspert-svar.md (what
the expert's queue looks like, and that "rejected" is the validator's verdict on
the numbers, never a professional judgement of the idea).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VtRd8y1PDPGwkrRXFhubqr
2026-08-05 21:12:09 +02:00
9668e17f2f feat(mcp): concrete MCP servers become tools the agents can call during a run
Krav 3, and the operator chose the run path explicitly: the external service must
be reachable WHILE the run works, not only when documents are ingested. Until now
the run path had one in-process tool against a local folder — and on the bundle
path the agents had no tools at all.

MAF already ships the client (MCPStdioTool / MCPStreamableHTTPTool, verified in
the pinned 1.9.0 with allowed_tools and request_timeout), so `mcp_tools.py` owns
only what MAF cannot decide for us: which servers a run may contact, which of
their tools it may call, how long it waits, and where the credential comes from.
This is a DIFFERENT seam from ingest_mcp.py on purpose — that one pulls source
documents before a run and speaks to null-argument tools. Same protocol, different
job.

Every refusal is a live hazard, not tidiness. An empty allowlist would let the far
end decide what the agents may call, so naming the tools is mandatory. A
non-positive timeout is an unbounded wait against a third party. An unknown field
is refused rather than ignored, which is also what keeps a literal secret from
being parked in the config — there is no field for one, only the NAME of an env
var. A named-but-unset credential refuses instead of calling anonymously, because
an anonymous call can succeed with the wrong scope.

Egress is declared, always. Every server and permitted tool is named in the run
announcement before the first call — including when no --mandate is given, which
was a real hole: the announcement only printed with a commission, so configuring
servers without one would have contacted third parties with nothing printed at
all. --live-dry-run still opens nothing, because the tools are entered after the
dry-run cut: the promise to stop before the first call now covers egress too.

Threaded through BOTH modes. A flag accepted in one mode and silently dropped in
the other is the defect class this CLI refuses by name.

Load-bearing MEASURED against the whole 744-test suite, four mutations all red:
build the tools but never hand them to the agents (2) · never enter the
AsyncExitStack, so they are constructed and useless (1) · never declare the egress
(2) · drop the allowlist on the built client (1).

Two live docs claimed MCP was unwired in the run path; both corrected rather than
left to rot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULCqjLF61rehj5cZmdUoR3
2026-08-05 16:53:07 +02:00
30bcdd3544 docs(mandate): how a domain expert commissions a run — and one honesty fix the run itself exposed
`docs/bestille-en-kjoring.md` is the commissioning half of the expert-facing pair
(`ekspert-svar.md` is the judging half): the mandate file field by field, how to
run it, and — separated deliberately — what a commission does NOT do. It directs
what is evaluated, never what is approved.

Registered in _LIVE_DOCS, so it cannot silently fall behind the code.

The example output in it is COPIED FROM A REAL RUN, not composed, and running
that run is what found the defect fixed here: three approaches against the same
cost line each validated at 30000 NOK, and the settlement printed
"Validated total: 90000 NOK". Commissioned approaches are ALTERNATIVES — they
usually attack the same line — so summing them reports money the project cannot
realise. A domain expert reading that total would reasonably believe the run
found 90k.

The settlement now reports how many approaches held and which one the run
carries: a selection, not an arithmetic claim. That also removes the last money
addition from this module, which is the right place for it not to be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULCqjLF61rehj5cZmdUoR3
2026-08-05 16:31:27 +02:00
b33ea00055 chore(deps): move the ingest library pin to v0.3.2 — as far as latest goes today
The pin had sat at v0.3.1 with STATE calling the hold "deliberate" and
recording no reason. Measured: no coord message ever announced v0.4.0 or
v0.5.0a* to this repo, so the hold was drift wearing a decision's clothes.

v0.3.2 is a pure fix (frontmatter and index labels emit verbatim; only
source_query is whitespace-collapsed, per ingest-spec §5), keeps
`dependencies = []`, and is green here: 668 passed.

WHY NOT FURTHER, both measured rather than assumed:

1. v0.4.0 introduces a REGRESSION that breaks our §6 removal path.
   Bisected v0.3.2 OK / v0.4.0 RED with a minimal repro: materialize a
   bundle, then re-materialize it with a CHANGED manifest, and the library
   no longer recognises its own stamp —

     MaterializationError: generated filename 'ingest-costs.md' collides
     with an existing file that does not carry the ingest stamp

   The stamp carries the manifest's name+hash (`ingest_manifest: m2@…`), so
   editing a manifest makes every file it previously wrote look curated.
   Re-ingesting the SAME manifest is fine, which is why fixtures miss it.
   It is `tests/test_ingest_loadbearing.py::test_reingest_with_active_
   removal_preserves_promoted_and_curated` that catches it. Reported
   upstream; not ours to fix.

2. Everything past v0.3.1 adds `llm-ingestion-guard>=0.2,<0.3` as a HARD
   runtime dependency (v0.3.1/v0.3.2: `dependencies = []`). That flips two
   documented invariants here — pyproject's "zero runtime deps" comment and
   the STATE marker line the guard repo reads machine-readably ("not a
   runtime dependency today"). An operator decision, not a version bump.

3. v0.5.0a2 is an alpha whose own CHANGELOG scopes it to a named pilot set
   — portfolio-optimiser-claude, the marketplace catalog, claude-code-llm-wiki
   — and says "do not pin this tag outside the pilot set", with the v0.2
   surface free to change without a deprecation cycle. This repo is not a
   pilot. Joining is llm-ingestion-okf's call, requested via coord.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GyAbxJoyypnLLUDcMvnKh8
2026-08-05 12:11:12 +02:00
bd221e6349 docs(ingest): the chain from ingest to run does not close by itself
Walked Door A from a fresh clone: materializing the file-family golden
manifest writes index.md plus one concept file per extraction, and pointing
the run at the result is refused —

    run refused: IR projection not found in bundle: 'validator-input.json'

A clean fail-fast, but nothing adopter-facing said it was coming, while the
README actively invites it ("swap --bundle-dir for your own bundle"). The
run path needs the bundle's IR projection, which ingest does not and cannot
produce: ingest materializes source documents, the projection states the
candidate measure. Both docs now say so, with the shape reference named.

Also corrects a live-doc claim that was wrong in both halves: the MCP
timeout is `anyio.fail_after` nested inside both task groups, not
`asyncio.wait_for`, and `tests/test_ingest_golden_mcp.py` covers it
(verified — 2 passing timeout tests). And no bundled example ships a
`cost-baseline.json`, so the text no longer implies one does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GyAbxJoyypnLLUDcMvnKh8
2026-08-05 11:39:58 +02:00
51739455ae docs(hitl): paste-ready expert answers, and the HITL chain walked end to end
The operator is not a domain expert, so the domain content is mine to own --
and the one thing the loop asks a human for is exactly the thing no example
existed for. docs/ekspert-svar.md is written for whoever has to deliver the
verdict: the two forms a judgement can take (a --rationale string during the
run, a JSON file in the inbox for later runs), where each field comes from, and
four complete paste-ready answers.

Every command and every verdict in it was RUN from a fresh clone before it was
written. The `hitl pending` line quoted is verbatim output. The rejection
answers close the gap STATE has carried since the demo shipped: the README
shows the VALIDATOR refusing a number, but nothing showed an EXPERT refusing a
proposal whose numbers are fine -- the only judgement in the whole loop that a
machine cannot make. Two rejection shapes are given, because "not feasible
here" and "right measure, wrong cost base" teach the system different things.

Everything is marked AI-authored and not verified professional judgement.

Also corrects the --outbox-dir help text, which claimed sharing a folder with
--verdict-dir "re-ingests raw agent output past the Step-8 promotion gate".
Measured, by pointing both at one folder and running twice: it does not. The
outbox artefacts are named {run_id}-*.json and carry none of the verdict keys,
so the tolerant inbox loader skips them and the run is unaffected. The hazard is
real but latent -- a future verdict-shaped artefact in the outbox -- so the
warning stays and says what is actually true. This also answers STATE's open
question about enforcing the distinction in the CLI: no. There is no reachable
contamination to refuse, and a guard for an unreachable case is the kind of
error handling this repo declines to write.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0118noV9rCfrdREH26XqZB5z
2026-08-05 10:58:43 +02:00
392f8493da chore(repo): planning artifacts become local-only; fixture builders become code
Operator ruling 2026-08-05, which settles decision (g): planning documents are
generally never public, and what OUR OWN sessions generate does not go out on
the forge at all. The example itself stays public so others can run the
process.

`.claude/projects/` is the Voyage session workbench -- 25 briefs/plans/reviews
this project's own sessions produced. Untracked and gitignored, exactly as
STATE.md already is, and for the same stated reason: this repo has a public
mirror, so that class of material is local-only rather than tracked.

The line is drawn at who wrote the document, and it is drawn deliberately:
`docs/plan/`, `docs/research/` and `docs/rapport/` stay tracked. Those are
curated, dated documents written for the repo's readers, three of them linked
from the README as the decision record. Move that line if it was meant wider.

Two files were NOT process artifacts and are not deleted. Both
`build_fixture.py` scripts are cited by tracked tests
(`test_ingest_golden_sql.py`, `test_ingest_golden_http.py`) as the documented
rebuild path for byte-exact goldens -- reproduction code that had landed in the
wrong directory. Moved next to the goldens they build; both docstrings updated,
so no tracked file is left pointing into an untracked tree (verified: the only
remaining `.claude/projects` string in a tracked file is the .gitignore rule
itself). One prose reference in the dated Foundry auth recipe was dropped for
the same reason.

652 tests still pass.

Does NOT address the 27 of these already readable on open/ since the S12
release -- untracking stops future publication only. That retraction is a
separate operator decision and is deliberately not taken here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GWsexbQjPo9rsV3aUE54ZS
2026-08-05 10:08:17 +02:00
a3b238307c docs(repo): meet the org repo-standard gate — 0 ERROR
Ran `repo-standard` (v0.1.1, class `standalone`) and fixed everything it
flagged as ERROR, plus the WARN links that were genuinely dead.

README first screen:
- opening line is now byte-identical to the forge description, so
  description == catalog == README is machine-checkable (badges moved below).
- `## Install` (required for class `standalone`): clone + `uv sync`, stated as
  clone-only because the shared spec, persona skill and example bundles under
  `shared/` are read from the working tree at run time. `uv run pytest` named as
  the verification, with the fact that no CI runner exists said out loud rather
  than implied by a badge.
- `## Non-goals` (required): the five limits already binding in CLAUDE.md —
  not a compliance product, not a portfolio-level reallocator, not autonomous
  decision-making, not turnkey, not a model benchmark.

Dead relative links (measured, not guessed):
- `docs/plan/2026-07-10-sesjonsplan-fase2-6.md` pointed at
  `../2026-07-14-revisjonspakke-DF-DI.md` six times; the file sits in
  `docs/plan/`, not `docs/`. (The sibling `../review-2026-07.md` links are
  correct and untouched.)
- the Fase-1 spike brief linked repo-root-relative from
  `.claude/projects/…/`; re-anchored with `../../../`.

The one remaining README ERROR was a gate false positive: `checkInternalLinks`
resolves targets against `git ls-files`, which lists files only, so a link to a
directory can never resolve. `[shared/](shared/)` now points at
`shared/README.md` — a better target anyway, since that file carries the
pull-only subtree rule. Not fixed here: the classifier lives in another repo.

Remaining WARNs are all inside `shared/`, deliberately untouched: it is a
pull-only commons subtree, and the nav-golden files are byte-level fixtures
that gate `test_nav_golden_*` — four of them are OKF bundle-internal links,
and the `/etc/passwd` ones are the negative escape fixture doing its job.

Suite green: 630 passed, 4 skipped (markdown-only diff; no test touched).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ri3aVJPfynCZtHRhesCzUH
2026-08-03 21:56:19 +02:00
c02c1addba fix(semretrieval): refuse a non-finite embedding instead of scoring it (kø-(l)/S3.1 MINOR)
`cosine`'s docstring claimed its guard was load-bearing because "a NaN reaching the
ranking sort key would corrupt ordering silently rather than failing loudly" — but the
guard tested `norm == 0.0` only, which a NaN or inf norm passes straight through. The
claim was prose, not behaviour.

Measured, not assumed: `cosine(unit, nan_vector)` AND `cosine(unit, inf_vector)` both
returned `nan`, and a NaN sort key made ranking INPUT-ORDER-DEPENDENT — six permutations
of the same three candidates produced four distinct orderings. That defeats the total
order `HybridRanker` documents ("`id` makes the result independent of input order").

Refuse rather than coerce, and deliberately NOT symmetric with the zero-norm branch: a
zero vector is a legitimate handled state (`FakeEmbedder` returns `np.zeros` by design),
whereas a non-finite component only ever means the INJECTED embedder is broken. Scoring
it `0.0` would launder that into "no semantic similarity" while ranking proceeded on a
forged signal — validation, never repair, mirroring `read_spend`.

Reachable via the documented `Embedder` extension point, not the shipped fake; scoped to
the norms (90% principle — a finite-normed dot-product overflow is not chased).

Also corrects `docs/extending.md`, which stated `SEMANTIC_WEIGHT_DEFAULT = 0.5` while the
code has said `0.25` since the weight was lowered.

625 -> 630 tests. Load-bearing MEASURED against the WHOLE suite, five mutations all red:
detach the guard entirely · coerce to 0.0 instead of raising · check only the first norm ·
drop "non-finite" from the message · (control) detach the zero-norm branch, which fails
ONLY the zero-norm test — the new guard does not mask the existing one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018V9vNBmxAmgJ2JMoHByiHS
2026-08-03 21:48:50 +02:00
9dc3722161 fix(ingest): run the MCP stdio transport against a real server, and repair its error contract (kø-x)
`stdio_call_tool` shipped never having been executed end to end — docs said so
explicitly. Running it found a real defect: `stdio_client` and `ClientSession` are
each an anyio task group, and anyio re-packages anything leaving one in a
`BaseExceptionGroup`. Both errors the transport raises from inside the session
(`mcp_tool_error`, `mcp_non_text_content`) therefore reached callers as exception
groups, never as the `IngestError` the whole Door A path catches and switches on by
`code`. No canned-tool test could see this: they never enter a task group.

`_unwrap_ingest_error` recovers the owned error and re-raises it; anything unowned is
re-raised untouched, so this narrows an exception group rather than blanket-catching.
Duck-typed on `.exceptions` because `except*`/`ExceptionGroup` are 3.11+ and this
project supports >=3.10.

Verified against a REAL server subprocess (a local process costs no model tokens, so
the repo's cost discipline is untouched; the contract tests still spawn nothing):
`examples/ingest-golden-mcp/` + `tests/test_ingest_golden_mcp.py` — byte-identical
golden extraction mirroring the http/sql goldens, plus the tool-error and
missing-`server_ref` branches.

Also recorded: a server on the ingest path must expose a NULL-ARGUMENT tool, so
`datasource.build_mcp_server` cannot serve it (`retrieve_cost_docs(query)` has a
required parameter, verified to return an error result). The two are separate seams
by design.

Load-bearing MEASURED, five mutations all RED: detach the unwrap · detach
`initialize()` · make the error code generic · detach the `isError` branch · change
one byte of the served body.

612 -> 615 tests. ruff + format + mypy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WiY53sm8JFqk7NN75g5wRS
2026-08-03 17:56:19 +02:00
8910a673ea feat(ingest): bound the default http transport in time, in front of the pinned library (S2.4)
The hole: `read_http`'s default transport is the library's `urllib_get`, which invokes the
stdlib opener with no `timeout=`. urllib's documented fallback is then the process-wide default
socket timeout — `None` out of the box — so an http source that accepts a connection and never
answers hangs a run indefinitely. That contradicts the invariant that nothing runs unbounded.

The spec text for S2.4 ("a timeout parameter on `_urllib_get`") could NOT be followed literally:
that function is UPSTREAM library code (`llm_ingestion_okf.connectors`, pinned v0.3.1, pull-only),
signature `(url, credential) -> str` — measured, not assumed. Same failure class as S2.2's
"implement it in `ingest.py`": spec text that says "change X" has to be checked against whether
X is ours at all.

So the fix goes in FRONT of the library: `timeout_get` scopes `socket.setdefaulttimeout` around
a delegate call to the library's own `urllib_get`, and `materialize` now hands the library that
wrapped transport instead of letting it resolve its own untimed default. This meets S2.4's own
verification criterion — a bound WITHOUT a second socket path — and avoids duplicating the
credential-header logic. An explicitly injected `http_get` is passed through UNWRAPPED: a
caller-owned transport (MCP fronts a subprocess with its own `timeout_seconds`) keeps its own
policy, and a process-global side effect is not ours to impose on it.

Honest limit, carried in the code comment, the test docstring and `docs/extending.md`, not just
in the commit: the default socket timeout is PROCESS-global. Under `concurrency=k` the runner is
asyncio on one thread, so the scoping holds; driving `read_http` from a thread-pool executor
would make it unsafe.

Half of S2.4's scope was already delivered upstream — transport failures are categorised as
`SourceError(code="http_transport")`. Coarser than the plan envisaged, but not ours to rewrite.

Two pre-existing guards went red on the first pass, both on PROSE only: `ingest.py` must not
contain "urlopen" (no forked connector) or "ingest_mcp" (AST-guarded mcp-free). No code violated
either — my docstrings merely named them. The guards were left exactly as strict as they were and
the prose was reworded; weakening a real guard to save a comment is the trade this repo refuses.

578 -> 583 tests. Five mutations MEASURED red (restored from scratchpad + `shasum -c` each time,
never `git checkout`):
  1. remove the timeout scoping entirely            -> RED
  2. apply the bound AFTER the delegate call        -> RED
  3. set the bound but never restore it (no finally)-> RED  (the unconditional control)
  4. hand the library a bare None again (pre-S2.4)  -> RED  (the wiring)
  5. make the wrapping unconditional                -> RED  (the conditional control)

Mutations 1 and 2 take ~10s to fail rather than failing instantly: that is the loopback test's
join deadline expiring. It is the measurement that the bound actually BITES — a black-hole
listener on 127.0.0.1 that completes the handshake and never answers, run on a daemon thread so
a detached seam fails an assertion instead of hanging the suite forever. Every other assertion
here only proves we set a global; that one proves the global does something.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TdLGwd33vqhkToh98Ym34P
2026-08-03 14:54:04 +02:00
ddd6338f02 feat(ingest): add the MCP connector as a transport inside the http family (S2.2)
Commons settled on 2026-08-01 that MCP is an extension of the `http` source
family, not a fourth family (`shared/ingest-spec.md` §4). This implements it
with ZERO schema change and zero spec amendment.

The transport discriminator lives in `base_url`, not in new manifest fields:
the shared library rejects unknown manifest keys fail-fast, and we consume it
pull-only at a pinned v0.3.1, so `server_ref`/`tool` as fields would have meant
a spec amendment plus a library release. It buys nothing — the library already
joins `base_url` + `/` + `query`, so `mcp+stdio://<server_ref>` + `<tool>`
reproduces exactly the two-part structure the (now stale) reference plan wanted.

Staying inside the family INHERITS what a fourth family would have had to write
and could have forgotten: the §8 network grant (measured to fire before any tool
call), the `max_rows` cap, §5 verbatim fenced rendering, and the §7 provenance
stamp. The discriminator gates rather than labels — `mcp_get` refuses a URL it
does not own, so an MCP transport can never quietly serve an `https://` manifest
and leave the bundle's provenance claiming a transport that was never used.

Parsing is string-based, not `urlsplit`-based: `urlsplit().hostname` lowercases
the host, which would silently break the case-sensitive env lookup `server_ref`
depends on.

`ingest.py` is untouched — it is AST-guarded mcp-free, so the transport lives in
its own module and is opt-in at the call site. `ingest_mcp.py` imports the open
`mcp` protocol client but never `agent_framework`, keeping the seam D7-portable.

Load-bearing, six mutations all measured RED: detach the scheme guard · make the
refusal unconditional · swap parsing to `urlsplit().hostname` · skip non-text
content instead of raising · force `allow_network=True` · smuggle in a MAF
import. Both source files restored byte-identical (`shasum -c`) after each.

Honesty: `stdio_call_tool` (the real stdio path) is written but never executed
end to end — every test injects a canned tool call, so the suite spawns no
subprocess and opens no socket. No golden fixture, and MCP stays unwired in the
optimiser run path. Stated in docs/extending.md rather than implied away.

555 -> 578 tests; ruff + mypy green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0112FPR5TX6pDLiNicBPzE8i
2026-08-02 21:14:57 +02:00
9e149c6847 docs(s31): close the review's honesty gap — narrow semantic claims to the shipped mechanism 2026-07-25 13:00:13 +02:00
da8779fc7f docs(s31): close 1 review MAJOR — vector store recorded as an unwired authoring primitive 2026-07-25 12:54:28 +02:00
921a8daf71 docs(s31): --semantic-retrieval CLI surface + Embedder/Retriever extension points 2026-07-25 06:31:44 +02:00
d44305cda9 docs(s53): knowledge-base recipe (D-H item 1) — team process, honest 1-2 week expectation, no wizard
New English docs/knowledge-base-recipe.md grounded strictly in the D-H decision record
(revisjonspakke-DF-DI.md §3): setup is always a small team (technical + domain expert), the
deliverable is a recipe NOT a wizard (B9 onboarding interview + guided verdict command rejected),
domain expert delivers files in their own formats never schema/JSON, phased process (inventory ->
skeleton -> seed verdicts -> iterate), reading via Obsidian/VS Code. The honest 1-2 week
expectation is stated early and SOURCED verbatim to the record. Factory-dependent parts
(free-format verdict translation, clone-to-demo) are explicitly marked future/blocked-on-toolkit
so the doc never claims above the evidence level. Linked from README's Docs section with the
1-2 week expectation in context. SC5 (ASCII-only greps): file exists, '1-2 weeks' x2,
'knowledge-base-recipe' in README. src/ untouched; 431 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KNNiJRk1sSwxgVLS5AobT1
2026-07-23 21:53:53 +02:00