feat(prepass): --prepass-seed makes the cut a starting point, and K2 says it costs
Q5 = B, bygget som MAALT OPSJON. --prepass-payload gir DEBATTEN et deklarert
kutt og trekker de fire navigatoerverktoeyene; --prepass-seed gir UTFORSKNINGEN
det samme kuttet som utgangspunkt og BEHOLDER verktoeyene.
Nekten M32/F4 staar ORDRETT. B er et nytt flagg, aldri en loesning av den, og
hjemmelen er konsumkontraktens SS 2.2: en skill maa ikke lese «outside what the
payload delivers or explicitly names as reachable». Andre ledd er hele arm B, og
PrepassDeclaration.rest_reachable er det som gjoer de to lesningene skillbare i
ettertid -- paakrevd uten default av cost_baseline_anchoreds grunn, fordi begge
defaults ville loeyet om hvilken arm som leste kuttet.
Soemmene:
admit_payload er EN opptaks-gate (form -> montert base -> tom-leveranse-nekt)
delt av begge doerer; to kopier ville latt en doer slippe inn det den andre
nekter.
render_seed deler header, regel->ANTALL-foldingen og DATA-blokkene med
render_context. Det eneste som skiller dem er avsnittet som sier hva
leseren kan gjoere videre.
explore(seed_context=...) legger kuttet i TASK-MELDINGEN, aldri i prompt:
_finish bygger Mandate.objective av prompt, og en kommisjon med 22 335
tokens utdrag i objektivet er uleselig for den som skrev den. Tom streng gir
en byte-identisk task-melding.
trace_payload(prepass=...) skriver deklarasjonen fra en finally. MAALT baerende
-- den seedede kjoeringen som doede paa en Azure-400 etterlot likevel kuttet
deklarert.
Fem nekter ved navn. --checkpoint-dir baerer en beslutning: en gjenopptatt
etappe kjoerer i en prosess som aldri saa payloaden og ville overskrevet den
parkerte etappens deklarasjon med prepass: null.
--dimension-config er BEVISST ikke nektet (arm A nekter den): maalt bygger
utforskningen navigator_tools(bundle_dirs) UTEN dimensjon, saa aa skope
seedet ville nektet tekst den samme loekka kan aapne et oeyeblikk senere.
MAALT PAA K2 MED LEVENDE MODELL, og maalingen taler MOT aa gjoere B til default:
like-for-like gratis 4 317 -> 227 675 o200k (x 52,7), og betalt er manageren
x 21 paa samme antall prompter. Viktigere enn prisen: den USEEDETE kontrollen
hentet prisskjemaet i fire steg (del-ii-bilag-7-prisskjema/prissammenstilling-
sheet-1.md), mens BEGGE seedede armer lot vaere -- den ene med null verktoeykall
fordi manageren rutet til hypotesisereren i alle tre runder, den andre ved aa
gjette stier ut av kuttets egne konsept-navn og mynte en base-id som ikke finnes.
Erkjennelsen kom (manageren skrev i hver runde at utdragene ikke rakk),
handlingen ikke. Ingen av de 40 svarene brukte ett eneste av kontraktens fem
literaler. NOK 2,78 av taket 5, 0 x 429. Anbefaling skrevet, beslutning ikke
tatt -- den er operatoerens.
Load-bearing maalt: 26 armer, 16 mutasjoner alle roede mot HELE suiten, groenn
kontroll 1493 passed / 5 skipped (fra 1467/5; +26 node-ider, 0 fjernet), golden
demo-transcript.stdout byte-uendret (shasum -a 1 av innholdet =
ea8c534773acdbe41ae68f2c55724d69aaf8be4f).
To armer var GROENNE AV FEIL GRUNN og ble rettet, ikke droppet: tool_calls alene
kan ikke skille «et verktoey ble kalt» fra «basen var aapen», fordi recorderen
appender FOER call_next; og id-enighets-armen maalte den stale digesten i stedet
for id-gaten. Sonden var dessuten feil foer koden var det -- foerste
diskriminator var norsk, og verktoeysvar serialiseres med \uXXXX-escapes.
Avvik, uttalt: implementasjonen ble skrevet FOER testfila. Roedmaalingen er gjort
etterpaa ved aa reversere src/ til HEAD (16 av 24 armer roede), deretter
restaurert med shasum-verifikasjon. Beviset er ekte, rekkefoelgen var ikke.
Rapportert, ikke fikset: ChatClientException (Azure 400, «No tool call found for
function call output») etter tre quick_validate-nekter paa rad -- den ligger
utenfor main()s nekt-tuppel og forlater CLI-en som traceback.
Azure-konfigurasjonen er uendret; endepunktet utledes inline fra az og er aldri
lagret i fil. Ruff + mypy rene.
Maaling: docs/2026-09-07-prepass-mater-q5b-k2.md
Ordre: 20260907T234344Z-9062321009-from-.claude
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
de50764e5c
commit
76b939b3b8
9 changed files with 1474 additions and 42 deletions
58
CLAUDE.md
58
CLAUDE.md
|
|
@ -2044,6 +2044,64 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
|
|||
`bundle_id` på rot-indeksen, S7a-3s måling sett fra produsentsiden, og `shared/` er pull-only), så
|
||||
fixturen er bygget på en KOPI med erklært id. Måling:
|
||||
`docs/2026-09-07-okf-prepass-i-debatten.md`.
|
||||
- **Det samme kuttet kan være et UTGANGSPUNKT i stedet for en erstatning — og hvilket av de to det
|
||||
var er et felt, ikke en tone (Q5=B, 08.09, ordre `20260907T234344Z`):** `--prepass-payload` gir
|
||||
DEBATTEN kuttet og trekker de fire navigatørverktøyene; `--prepass-seed` gir UTFORSKNINGEN det
|
||||
samme kuttet som startpunkt og **beholder** verktøyene. **Nekten M32/F4 står ordrett** — B er et
|
||||
NYTT flagg, aldri en løsning av den — og hjemmelen er konsumkontraktens § 2.2, som forbyr å lese
|
||||
«outside what the payload delivers **or explicitly names as reachable**»: andre ledd er hele arm
|
||||
B. **`admit_payload` er ÉN opptaks-gate** (form → montert base → tom-leveranse-nekt) delt av
|
||||
begge dører; to kopier ville latt én dør slippe inn det den andre nekter (kø-(p)).
|
||||
**`PrepassDeclaration.rest_reachable` er PÅKREVD uten default** av `cost_baseline_anchored`s
|
||||
grunn — `False` ville latt en kjøring som beholdt verktøyene påstå at kuttet var alt den kunne
|
||||
lese, `True` ville latt armen som TRAKK dem påstå at basen sto åpen — og `prepass_notice` er
|
||||
fortsatt ENESTE renderer, nå med to armer, fordi «8 av 630» sier det motsatte av sannheten i
|
||||
nøyaktig ett av de to tilfellene. **Kuttet rir i TASK-MELDINGEN, aldri i `prompt`**
|
||||
(`explore(seed_context=…)`): `_finish` bygger `Mandate.objective` av `prompt`, og en kommisjon med
|
||||
22 335 tokens utdrag i objektivet er uleselig for den som skrev den; tom streng ⇒ byte-identisk
|
||||
med i dag. Deklarasjonen når `{run_id}-exploration.json` via `trace_payload(prepass=…)` skrevet
|
||||
fra `finally` — **defaultet** (`skipped_links`-halvdelen), og MÅLT bærende: den seedede kjøringen
|
||||
som døde på en Azure-400 etterlot likevel kuttet deklarert. Fem nekter ved navn, og
|
||||
`--checkpoint-dir` er den som bærer en beslutning: en gjenopptatt etappe kjører i en prosess som
|
||||
aldri så payloaden og ville overskrevet den parkerte etappens deklarasjon med `prepass: null`.
|
||||
**`--dimension-config` er BEVISST IKKE nektet** (arm A nekter den): målt bygger utforskningen
|
||||
`navigator_tools(bundle_dirs)` UTEN dimensjon, så å skope seedet ville nektet tekst den samme
|
||||
løkka kan åpne med `read_file` et øyeblikk senere. **MÅLT PÅ K2 MED LEVENDE MODELL, og
|
||||
målingen taler MOT å gjøre B til default:** like-for-like gratis **4 317 → 227 675 o200k
|
||||
(× 52,7)**, betalt er manageren **× 21** på samme antall prompter, og — viktigere — den USEEDETE
|
||||
kontrollen hentet prisskjemaet i FIRE steg
|
||||
(`del-ii-bilag-7-prisskjema/prissammenstilling-sheet-1.md`) mens BEGGE seedede armer lot være:
|
||||
den ene med NULL verktøykall fordi manageren rutet til hypotesisereren i alle tre runder, den
|
||||
andre ved å gjette stier ut av kuttets egne konsept-navn og mynte en base-id som ikke finnes.
|
||||
**Erkjennelsen kom, handlingen ikke** — manageren skrev i hver runde at utdragene ikke rakk.
|
||||
Ingen av de 40 svarene brukte ett eneste av kontraktens fem literaler (kontrast S7s arm A, som
|
||||
brukte `[sourced-not-sufficient]` ordrett). **NOK 2,78 av taket 5, 0× 429.** Load-bearing MÅLT
|
||||
(`tests/test_prepass_seed_door_loadbearing.py`, 26 armer), **16 mutasjoner alle røde mot HELE
|
||||
suiten** + grønn kontroll **1493/5** og golden `demo-transcript.stdout` BYTE-UENDRET
|
||||
(`shasum -a 1` av INNHOLDET = `ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): M1 detach CLI-wiringen
|
||||
(2) · M2 `explore()` ignorerer seedet (2) · M3 kollaps B inn i A (1) · M4 seedet uten utforskning
|
||||
(1) · M5 de to armene tillates sammen (1) · M6 `--checkpoint-dir` tillates (1) · M7/M8 ut av
|
||||
portefølje-/report-partisjonen (1+1) · M9 konstant `rest_reachable` (1) · M10 deklarasjonen
|
||||
forlater aldri CLI-en (1) · M11 feltet ut av artefaktet (2) · M12 den seedede renderingen sier
|
||||
«du har ingen verktøy» (1) · M13 opptaks-gaten svelger alt (7, hvorav TRE i tester eldre enn dette
|
||||
arbeidet) · M14 kuttet deklareres til ingen (1) · M15 detach id-enighets-gaten ved denne døra (1)
|
||||
· README-armen (1). **TO ARMER VAR GRØNNE AV FEIL GRUNN (repoets vakuøs-gate-klasse, TJUEANDRE og
|
||||
TJUETREDJE gang), begge funnet av mutasjonene:** (i) verktøy-armen asserterte på `tool_calls`
|
||||
alene, og recorderen appender FØR `call_next`, så en tool som reiste «unknown knowledge base»
|
||||
etterlot et byte-identisk spor — «et verktøy ble kalt» og «basen var åpen» er ulike fakta, og
|
||||
armen krever nå at basens EGEN indekstekst kommer tilbake; (ii) id-enighets-armen redigerte ÉN
|
||||
konseptfil, men `assert_declared_ids_agree` leser konsepter og én erklæring er en base som er
|
||||
enig med seg selv — nekten kom fra den STALE DIGESTEN, og armen bruker nå TO konsepter og et
|
||||
re-digestet payload. **Sonden var dessuten feil før koden var det:** første diskriminator var
|
||||
norsk («rask småskala-testing»), og verktøyresultatet serialiseres med `\uXXXX`-escapes, så armen
|
||||
var rød mot en virkende implementasjon — S2c-målingen gjentatt. **Ærlighets-grenser, uttalt:** tre
|
||||
betalte kjøringer på ett spørsmål er ikke et utvalg; de to seedede endte ulikt (Azure-400 vs
|
||||
rundetak); ingen arm nådde pipelinen, så siteringene i stempelet er IKKE talt — og for en seedet
|
||||
arm ville en full kjøring kostet ~NOK 3,2 mot NOK 2,22 igjen av taket, hvilket ER funnet; arm A
|
||||
og C er GJENBRUKT fra økt 98 og er dessuten DEBATT-armer mot B-ens utforskning; og
|
||||
`ChatClientException` (Azure 400 etter tre `quick_validate`-nekter på rad) er RAPPORTERT, ikke
|
||||
fikset — den ligger utenfor `main()`s nekt-tuppel og forlater CLI-en som traceback. Måling:
|
||||
`docs/2026-09-07-prepass-mater-q5b-k2.md`.
|
||||
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
||||
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
||||
|
||||
|
|
|
|||
21
README.md
21
README.md
|
|
@ -513,6 +513,27 @@ when the seam is detached, so the loop cannot silently degrade into theater.
|
|||
delivered) and with `--explore` (the exploration reads the whole base with the very tools the
|
||||
payload withdraws).
|
||||
|
||||
**Seeding the exploration with a cut instead of replacing the base (`--prepass-seed`).** The
|
||||
flag above hands the *debate* a cut and takes its navigation away. This one is the opposite arm,
|
||||
for the *exploration*: the same kind of pre-computed cut becomes the loop's **starting point**,
|
||||
and the four navigation tools stay, so it can read past the cut whenever the delivered text does
|
||||
not carry what it needs. Every excerpt is admitted by exactly the checks the other flag applies
|
||||
— same document, same bytes, same text, expert verdicts refused — before a single model call.
|
||||
|
||||
```bash
|
||||
uv run python -m portfolio_optimiser.run BYGG-KONTOR-NORD --docs-dir <docs> \
|
||||
--bundle-dir <knowledge-base> --explore "Find the cheapest saving worth testing here" \
|
||||
--explore-config exploration.json --prepass-seed cut.json --outbox-dir out --run-id r1
|
||||
```
|
||||
|
||||
The cut is printed before the loop starts and recorded in `{run_id}-exploration.json`, which
|
||||
carries `rest_reachable` so a reader can tell a *seeded* run from a *bounded* one without having
|
||||
watched it. **This is not a cheaper run and not a bounded one** — it is a declared place to
|
||||
begin. Requires `--explore`. Refused with `--prepass-payload` (the two are opposite arms of one
|
||||
decision, and merging them would mean silently picking one), with `--portfolio`, with `--report`
|
||||
(which runs no exploration) and with `--checkpoint-dir`, because a resumed run happens in a
|
||||
process that never saw the cut and its own record would overwrite the declaration with nothing.
|
||||
|
||||
**Answering it days later (`--checkpoint-dir` / `--resume`).** A domain expert is rarely at the
|
||||
terminal when the loop reaches the plan, so the same review can be *parked* to disk instead.
|
||||
`--checkpoint-dir` writes the suspended workflow there and the open question to
|
||||
|
|
|
|||
303
docs/2026-09-07-prepass-mater-q5b-k2.md
Normal file
303
docs/2026-09-07-prepass-mater-q5b-k2.md
Normal file
|
|
@ -0,0 +1,303 @@
|
|||
# Q5=B — pre-passet MATER utforskningen: `--prepass-seed`, målt på K2 med levende modell
|
||||
|
||||
**Dato:** 2026-09-08 (ordre datert 2026-09-07) · **Ordre:** `20260907T234344Z-9062321009-from-.claude` ·
|
||||
**Beslutning:** `~/.claude/docs/2026-09-07-vurdering-arbeidssystemet.md` rad 0 / § 3 Q5 —
|
||||
**A forblir default, B bygges som MÅLT OPSJON.** ·
|
||||
**Forrige ledd:** `docs/2026-09-07-okf-prepass-i-debatten.md` (kuttet i debatten) og
|
||||
`docs/2026-09-07-syretest-s7-prepass-k2.md` (armene A og C, betalt i økt 98).
|
||||
|
||||
Dette er **både en leveranse og en måling**. Koden er ny og gatet (26 armer, 16 mutasjoner alle
|
||||
røde); tallene under er produsert av kommandoer som står i teksten, og måleoppsettet ligger utenfor
|
||||
treet (`scratchpad/q5b/`, utracket).
|
||||
|
||||
---
|
||||
|
||||
## 0. Hva som ER målt, og hva som IKKE er det
|
||||
|
||||
**Målt her:** at `--prepass-seed` gir utforskningen det deklarerte kuttet som **utgangspunkt** og
|
||||
**beholder** de fire navigatørverktøyene; at deklarasjonen sier `rest_reachable: true` på stdout og
|
||||
i `{run_id}-exploration.json`; at nekten M32/F4 (`--prepass-payload` + `--explore`) står ordrett; at
|
||||
kjøringen uten flagget er byte-uendret. Deretter **tre betalte armer på K2** med levende
|
||||
`gpt-4.1-mini`: én useedet kontroll og **to** seedede kjøringer, med kjent-positiv assertert først i
|
||||
hver.
|
||||
|
||||
**IKKE målt:** at en modell dømmer **bedre** med et seedet kutt. Tre kjøringer på ett spørsmål er
|
||||
ikke et utvalg, og de tre endte av tre ulike grunner (rundetak · rundetak · en Azure-400). Én
|
||||
kjøring er én kjøring; at den seedede atferden gikk i **samme retning to ganger** er sterkere enn
|
||||
én, og fortsatt ikke et utvalg. **Heller ikke målt:** siteringene i stempelet — ingen av de tre
|
||||
armene nådde pipelinen (§ 5), og den seedede armens pris gjør en full kjøring uoppnåelig innenfor
|
||||
taket på NOK 5, hvilket er selve funnet snarere enn et hull i det.
|
||||
|
||||
**Kjent-positiv, kjørt FØRST og `assert`-et i hver eneste betalte kjøring:** rotnivå-listingen på
|
||||
**levert** K2 måler **3 954 tegn / 1 495 o200k-tokens over 629 konseptfiler** — S7a-3s publiserte
|
||||
tall, eksakt. `scratchpad/q5b/live/live_q5b.py` nekter å kjøre videre uten den; utskriften står
|
||||
øverst i hver armlogg.
|
||||
|
||||
---
|
||||
|
||||
## 1. Nekten M32/F4 — hvorfor den ble valgt, skrevet ned FØR den ble rørt
|
||||
|
||||
Ordrett fra `src/portfolio_optimiser/run.py`, slik den sto ved øktstart:
|
||||
|
||||
```
|
||||
run refused: --prepass-payload and --explore cannot be combined (the exploration
|
||||
navigates the whole knowledge base with the same four tools the payload
|
||||
withdraws, so the run as a whole would read far outside the cut it declares)
|
||||
```
|
||||
|
||||
Grunnen i kommentaren over den: utforskningen leser HELE basen med de fire verktøyene og former så
|
||||
mandatet debatten får beskjed om at den ikke må gå utenfor — «that is precisely the ground on which
|
||||
the debate's own tools are withdrawn, one caller over».
|
||||
|
||||
**B opphever ikke dette.** Nekten står, ordrett, og har sitt eget vitne
|
||||
(`test_the_replacing_arm_is_still_refused_with_explore_in_the_same_words`). B er et **nytt flagg**
|
||||
med motsatt oppførsel på den ene aksen nekten handler om: kuttet **erstatter ikke** basen, det er
|
||||
hvor man **begynner**, og verktøyene blir. Kontrakten sier selv at det er lov —
|
||||
`llm-ingestion-okf/docs/consumption-contract.md` § 2.2: en skill må ikke lese «outside what the
|
||||
payload delivers **or explicitly names as reachable**». Andre ledd er hele hjemmelen for arm B, og
|
||||
`PrepassDeclaration.rest_reachable` er det som gjør de to lesningene skillbare i ettertid.
|
||||
|
||||
Navnet: ordren foreslo `--prepass-seed`, og kontrakten antyder ingen bedre. «Seed» er dessuten
|
||||
repoets eget ord for et **bevart utgangspunkt** (`explore(seed_approaches=…)`, § C.6 dør 1).
|
||||
|
||||
---
|
||||
|
||||
## 2. Hva som ble bygget
|
||||
|
||||
| Søm | Hva |
|
||||
|---|---|
|
||||
| `prepass.admit_payload` | ÉN opptaks-gate — form, montert base, tom-leveranse-nekt — delt av BEGGE dører. To kopier ville latt én dør slippe inn det den andre nekter (kø-(p)). |
|
||||
| `prepass.render_seed` | Arm B-renderingen. Deler header, regel→ANTALL-foldingen og DATA-blokkene med `render_context`; **det eneste som skiller dem er avsnittet som sier hva leseren kan gjøre videre.** |
|
||||
| `PrepassDeclaration.rest_reachable` | **Påkrevd uten default** (`cost_baseline_anchored`-regelen: begge defaults ville løyet). Når artefaktet via `declaration_payload`. |
|
||||
| `explore(seed_context=…)` | Kuttet rir i **TASK-MELDINGEN**, aldri i `prompt` — `_finish` bygger `Mandate.objective` av `prompt`, og en kommisjon med 22 335 tokens utdrag i objektivet er uleselig for den som skrev den. Tom streng ⇒ byte-identisk med i dag. |
|
||||
| `trace_payload(prepass=…)` | Deklarasjonen inn i `{run_id}-exploration.json`, skrevet fra `finally`. **Defaulter til `None`** (`skipped_links`-halvdelen: fravær er et ærlig positivt utsagn). |
|
||||
| `run.prepass_notice` | ÉN renderer, to armer. Uten `rest_reachable` ville «8 av 630» sagt det motsatte av sannheten i nøyaktig ett av de to tilfellene. |
|
||||
| Fem nekter | `--prepass-seed` krever `--explore`; refusert med `--prepass-payload` (to motsatte armer), `--portfolio`, `--report` og `--checkpoint-dir`. |
|
||||
|
||||
**`--checkpoint-dir`-nekten er ikke pynt:** en parkert etappe skriver `{run_id}-exploration.json`
|
||||
med deklarasjonen; den GJENOPPTATTE etappen kjører i en prosess som aldri så payloaden og
|
||||
overskriver samme fil med `prepass: null`. En deklarasjon som fordamper halvveis er verre enn en
|
||||
nektet, og den ville gjort det i stillhet.
|
||||
|
||||
**`--dimension-config` er BEVISST ikke nektet** — arm A nekter den. Målt: utforskningen bygger
|
||||
`navigator_tools(bundle_dirs)` **uten** dimensjon, så å skope seedet ville nektet tekst den samme
|
||||
løkka kan åpne med `read_file` et øyeblikk senere. Debatten nedstrøms er skopet som før.
|
||||
|
||||
---
|
||||
|
||||
## 3. Instrumentet, og hva det målte GRATIS før noe ble betalt
|
||||
|
||||
Samme manus, samme base, samme spørsmål, kun flagget skiller:
|
||||
|
||||
```bash
|
||||
scratchpad/prepass-venv/bin/python scratchpad/q5b/live/rehearse.py seed # 227 675 o200k
|
||||
scratchpad/prepass-venv/bin/python scratchpad/q5b/live/rehearse.py plain # 4 317 o200k
|
||||
```
|
||||
|
||||
**19 prompter i begge. 4 317 mot 227 675 o200k-tokens — × 52,7.** 225 518 av de 227 675 er de
|
||||
**fire manager-promptene**: utforskningens deltakere deler ÉN samtalehistorikk, så task-meldingen
|
||||
rir i hver prompt — MAJOR-3-målingen ett lag opp, nå på et kutt i stedet for en base.
|
||||
|
||||
Rendringen selv: `render_seed` = **75 614 tegn / 22 335 o200k**, `render_context` = 75 341 / 22 283.
|
||||
Forskjellen mellom armene er 273 tegn prosa. **Hver manager-prompt bar ~45 000 tokens**, altså
|
||||
kuttet omtrent to ganger — MAF legger task-teksten både i system- og bruker-rollen.
|
||||
|
||||
---
|
||||
|
||||
## 4. Tre betalte armer på K2
|
||||
|
||||
Base `k2-trinn1-20260903`, ref `sha256-tree:4ffd750c…`, 630 konsepter. Payload fra økt 98:
|
||||
**8 levert / 622 holdt tilbake** (`below_k` 261, `no_lexical_match` 359, `over_budget_alone` 2).
|
||||
Spørsmålet er ORDRETT S7as og S7s: `Finn kostnadsbesparelser i Stange skole-anbudet`.
|
||||
|
||||
```bash
|
||||
export PORTFOLIO_FOUNDRY_PROJECT_ENDPOINT=… # utledet INLINE fra `az`, aldri i fil
|
||||
export PORTFOLIO_MODEL_MAP=scratchpad/major2-live/model_map.json
|
||||
export PACE_SECONDS=2
|
||||
scratchpad/prepass-venv/bin/python scratchpad/q5b/live/live_q5b.py {Bplain|Bseed|Bseed2}
|
||||
# argv: K2 --profile azure --docs-dir <base> --bundle-dir <base>
|
||||
# --explore "Finn kostnadsbesparelser i Stange skole-anbudet" --explore-config <fil>
|
||||
# --outbox-dir … --run-id … [+ --prepass-seed <payload> for Bseed/Bseed2]
|
||||
# bounds: max_rounds=3, max_tokens=1 000 000, max_stall_count=2, max_reset_count=1
|
||||
```
|
||||
|
||||
Deployment `gpt-4.1-mini` (GlobalStandard, `capacity` 100, versjon `2025-04-14`) — **målt med
|
||||
`az cognitiveservices account deployment show`, ingen konfig rørt.**
|
||||
|
||||
| Arm | Prompter | Input | Output | NOK | Verktøykall | **Nådde prisskjemaet?** | Slutt |
|
||||
|---|---:|---:|---:|---:|---:|---|---|
|
||||
| **Bplain** (useedet) | 12 | 35 914 | 2 284 | 0,168 | 4 | **JA** — `del-ii-bilag-7-prisskjema/prissammenstilling-sheet-1.md` | rundetak 3/3 |
|
||||
| **Bseed** | 20 | 376 127 | 3 590 | 1,458 | 12 | **NEI** | Azure 400 (§ 6) |
|
||||
| **Bseed2** | 8 | 299 160 | 2 403 | 1,153 | **0** | **NEI** | rundetak 3/3 |
|
||||
|
||||
Per fase (BudgetMiddleware-ledgeren, input-tokens):
|
||||
|
||||
| Rolle | Bplain | Bseed | Bseed2 |
|
||||
|---|---:|---:|---:|
|
||||
| manager (5 prompter i alle tre) | 14 204 | **310 430** | **298 230** |
|
||||
| navigator | 14 975 (6) | 63 237 (10) | — (0) |
|
||||
| hypothesiser | 6 735 (1) | 2 460 (5) | 930 (3) |
|
||||
|
||||
**Manageren alene er × 21,9 og × 21,0 dyrere seedet, på samme antall prompter.**
|
||||
|
||||
### Den useedete armen fant prisskjemaet i fire steg
|
||||
|
||||
```
|
||||
list_bundles → read_bundle(k2-trinn1-20260903)
|
||||
→ read_dir(del-ii-bilag-7-prisskjema)
|
||||
→ read_file(del-ii-bilag-7-prisskjema/prissammenstilling-sheet-1.md)
|
||||
```
|
||||
|
||||
Nøyaktig samme fire steg som S7s arm C (`docs/2026-09-07-syretest-s7-prepass-k2.md` § 4), og
|
||||
prompt 9 er modellens egen oppsummering av prisrammen. Hypotesisereren resonnerte deretter om
|
||||
«tiltransporterte kontrakter» — en retning forankret i tallene den faktisk hadde lest.
|
||||
|
||||
### De seedede armene gjorde det ikke — og de to gjorde det på hver sin måte
|
||||
|
||||
**Bseed2 er den rene:** manageren rutet til **hypothesiser i alle tre runder**. Navigatøren fikk
|
||||
aldri en tur, og null verktøykall ble gjort — mens manageren i hver eneste runde skrev at utdragene
|
||||
«do not explicitly enumerate … costs». Den **så** at kuttet ikke rakk, og gikk likevel ikke og så
|
||||
etter. Hypotesene ble deretter generiske: `Early Integration Testing`, `Modular Component Design`,
|
||||
«prefabricated building components» — ingen av dem forankret i en kostlinje.
|
||||
|
||||
**Bseed vandret, og feil vei.** Sporet (`{run_id}-exploration.json`, S7a-3 pkt. 3 gir stien):
|
||||
|
||||
```
|
||||
quick_validate | renholdstekniske_funksjonskrav ← OPPDIKTET base-id, 3×, alle nektet
|
||||
list_bundles → read_bundle → read_dir(del-ii-bilag-1-2-renholdstekniske-…)
|
||||
→ read_file(…/1-1/1-1.md) ← finnes ikke
|
||||
→ read_dir(…/1-1) → read_file(…/03-01-2023-vask-av-layout.md) ← ALLEREDE LEVERT
|
||||
→ read_file(inbox-del-ii-bilag-7-prisskjema.md) ← finnes ikke
|
||||
→ read_file(del-ii-bilag-7-prisskjema) ← katalog, nektet ved navn
|
||||
→ read_file(del-ii-bilag-7-prisskjema.md) ← finnes ikke
|
||||
```
|
||||
|
||||
Tre ting står her. (1) Hypotesisereren **myntet en base-id av et levert konsepts navn**
|
||||
(`renholdstekniske_funksjonskrav`) og fikk tre nekter på rad, hvorpå MAF sluttet å tillate
|
||||
verktøykall for den forespørselen. (2) Navigatøren gikk inn i katalogen til et konsept den
|
||||
**allerede hadde i prompten** og leste det på nytt. (3) Den prøvde prisskjemaet **tre ganger** og
|
||||
bommet på formen hver gang — den kom aldri på å kalle `read_dir` på katalogen, som er nøyaktig det
|
||||
steget den useedete armen tok.
|
||||
|
||||
### Nevnerne: hva modellen sa om dem
|
||||
|
||||
**Ingen av de tre armene brukte ett eneste av kontraktens fem literaler** (`[unread]`,
|
||||
`[sourced-not-sufficient]`, `[unverifiable-from-bundle]`, `extracted`, `derived`), og ingen nevnte
|
||||
`630`, `622` eller `8`. Nevner: 40 svar over tre armer. Kontrast: S7s arm A — kuttet i **debatten**,
|
||||
verktøyene **trukket** — brukte `[sourced-not-sufficient]` ordrett (1 i renderingen → 1 i svaret).
|
||||
|
||||
Manageren sa det likevel, i sine egne ord og i hver runde: «without explicit enumeration or direct
|
||||
[cost data]». **Erkjennelsen kom; handlingen kom ikke.**
|
||||
|
||||
### Refererte konsept-id-er ∩ levert
|
||||
|
||||
**0 av 0 i Bplain, 0 av 4 og 0 av 2 i de seedede.** Ingen modell refererte til en konsept-id i
|
||||
prosa i det hele tatt; de fire/to «id-lignende strengene» i de seedede armene er vanlige
|
||||
skråstreker i prosa (`cleaning/maintenance`, `material/use`). Nevneren er skrevet ned fordi
|
||||
tellingen ellers ville lest som et funn om levert-mengden.
|
||||
|
||||
---
|
||||
|
||||
## 5. Siteringene i stempelet: IKKE målt, og hvorfor
|
||||
|
||||
Alle tre armene stoppet i utforskningen — to på rundetaket, én på en 400 — så ingen mandat ble
|
||||
evaluert, ingen `SavingsProposal` ble bygget og intet `ProvenanceStamp` finnes å telle siteringer i.
|
||||
Å heve rundetaket for den useedete armen ville kostet ~NOK 0,3; for en **seedet** arm anslår de
|
||||
målte 60–80k input-tokens per manager-prompt ~NOK 3,2, mot NOK 2,22 igjen av taket. **Det er ikke
|
||||
et hull i målingen, det er målingen:** per-prompt-prisen på et seedet utforskningskutt gjør en full
|
||||
pipeline-kjøring uoverkommelig innenfor det taket ordren satte.
|
||||
|
||||
---
|
||||
|
||||
## 6. Azure-400-en i Bseed — rapportert, ikke fikset
|
||||
|
||||
```
|
||||
ChatClientException: Error code: 400 — No tool call found for function call output
|
||||
with call_id call_AyeuYJmqvmmej9utu361Phtt
|
||||
```
|
||||
|
||||
Sekvensen er målt: tre `quick_validate`-nekter på rad → MAF melder «Maximum consecutive function
|
||||
call errors reached (3). Stopping further function calls for this request» → samtalen inneholder
|
||||
et funksjonskall-**resultat** uten sitt kall → Responses-APIet avviser den. Feilen kommer fra
|
||||
transporten, ikke fra denne ordrens kode, og den ville forlatt CLI-en som **traceback**
|
||||
(`ChatClientException` er ikke i `main()`s nekt-tuppel). **Utenfor ordren; rapportert.**
|
||||
|
||||
---
|
||||
|
||||
## 7. Kostnad
|
||||
|
||||
Listepris, Azure Retail Prices API (`api-version=2023-01-01-preview`, `currencyCode='NOK'`,
|
||||
`armRegionName=eastus`), **hentet på nytt 2026-09-08**: input **NOK 0,003734 / 1K**, output
|
||||
**NOK 0,014936 / 1K** — identiske med dem økt 98 målte.
|
||||
|
||||
| Kjøring | Prompter | Input | Output | NOK |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Bplain | 12 | 35 914 | 2 284 | 0,168 |
|
||||
| Bseed | 20 | 376 127 | 3 590 | 1,458 |
|
||||
| Bseed2 | 8 | 299 160 | 2 403 | 1,153 |
|
||||
| **SUM** | **40** | **711 201** | **8 277** | **2,779** |
|
||||
|
||||
**NOK 2,78 — 56 % av taket på NOK 5. 429-svar: 0.** Største enkeltkall: 80 713 input-tokens
|
||||
(Bseed2s siste manager-prompt).
|
||||
|
||||
---
|
||||
|
||||
## 8. Anbefaling — beslutningen er operatørens
|
||||
|
||||
**Målingen peker på A som default, og på B som en opsjon med et smalt bruksområde.** Grunnlaget er
|
||||
tre tall og én atferd:
|
||||
|
||||
1. **Prisen.** × 52,7 like-for-like gratis, × 21 på manageren betalt. Kuttets tekst rir i hver
|
||||
prompt fordi deltakerne deler én historikk — det er MAJOR-3-formen, og den gjelder et kutt like
|
||||
fullt som en base.
|
||||
2. **Effekten på navigasjonen.** Den useedete armen gikk og hentet prisskjemaet i fire steg; begge
|
||||
seedede armer gjorde det ikke, den ene uten å kalle et eneste verktøy. **Et utgangspunkt ble en
|
||||
endestasjon** — i én kjøring fordi manageren aldri rutet til navigatøren, i den andre fordi
|
||||
modellen gjettet stier ut av kuttets egne konsept-navn.
|
||||
3. **Erkjennelsen uten handlingen.** Manageren skrev i hver runde at utdragene ikke rakk, og
|
||||
handlet ikke på det. Renderingen sier eksplisitt «OPEN THE BASE with your navigation tools»; det
|
||||
var ikke nok. Om en skarpere formulering endrer dette er **ikke målt**.
|
||||
|
||||
**Der B likevel kan være riktig:** når kuttet er *ment* å være svaret og navigasjonen bare er en
|
||||
sikkerhetsventil — små korpus, eller et pre-pass som er kjørt for nøyaktig dette spørsmålet av
|
||||
noen som vet det. Da kjøper man en deklarert nevner uten å miste utveien. På et 630-konsept-korpus
|
||||
med et leksikalsk kutt kjøpte den her verken det ene eller det andre.
|
||||
|
||||
**Ikke konkludert:** at A er bedre enn B *generelt*. Tre kjøringer, ett spørsmål, ett korpus, én
|
||||
modell.
|
||||
|
||||
---
|
||||
|
||||
## 9. Ærlighets-grenser, uttalt
|
||||
|
||||
- **Ikke-determinisme.** Tre kjøringer er ikke et utvalg. De to seedede gikk i samme retning; det
|
||||
gjør retningen mer troverdig og ikke etablert.
|
||||
- **De to seedede armene endte ulikt** (400 vs rundetak), så «nådde ikke prisskjemaet» er ikke
|
||||
målt mot samme sluttbetingelse i begge.
|
||||
- **Arm A og arm C er GJENBRUKT**, ikke re-kjørt: tallene står i
|
||||
`docs/2026-09-07-syretest-s7-prepass-k2.md` § 4 og ble betalt i økt 98. De er dessuten
|
||||
**debatt**-armer, mens B er en **utforskning** — sammenligningen på tvers gjelder samme base,
|
||||
samme spørsmål og samme modell, aldri samme fase.
|
||||
- **Prisene i K2-fixturen er syntetiske** (MAJOR-4-arven); det påvirker ikke tokenmålingene.
|
||||
- **Ingen `ProvenanceStamp` finnes i noen arm** (§ 5), så siteringene er ikke talt.
|
||||
- **Hosting-flaten er urørt.** Feltet er i ingen av de tre settene, så den generiske
|
||||
`unknown field(s)`-400-en svarer, og Fase 4es to halvdeler står — gatet av en testarm, ikke av en
|
||||
redigering.
|
||||
|
||||
---
|
||||
|
||||
## 10. Verifiseringslogg
|
||||
|
||||
| Påstand | Hvordan verifisert |
|
||||
|---|---|
|
||||
| Nekten M32/F4 står ordrett | `test_the_replacing_arm_is_still_refused_with_explore_in_the_same_words` + kilden sitert i § 1 |
|
||||
| Kontraktens § 2.2 hjemler arm B | `llm-ingestion-okf/docs/consumption-contract.md`, lest i denne økten |
|
||||
| Utforskningen er uskopet | `explore.fresh_exploration_workflow` → `navigator_tools(bundle_dirs)` uten `dimension` |
|
||||
| Kjent-positiv 3 954 / 1 495 / 629 | `assert` i `live_q5b.py`, kjørt først i alle tre betalte armer |
|
||||
| Deployment/kapasitet urørt | `az cognitiveservices account deployment show` (GlobalStandard, 100, `2025-04-14`) |
|
||||
| NOK-prisene | Azure Retail Prices API, hentet 2026-09-08 |
|
||||
| Prisskjemaet nådd/ikke nådd | `{run_id}-exploration.json` → `tool_calls[].path` |
|
||||
| Nevnerne uten kontrakt-literaler | tokenskann over alle 40 svar |
|
||||
| Gaten er load-bearing | 16 mutasjoner, alle røde mot HELE suiten; grønn kontroll **1493 passed / 5 skipped** |
|
||||
| Golden byte-uendret | `shasum -a 1 tests/golden/demo-transcript.stdout` = `ea8c534773acdbe41ae68f2c55724d69aaf8be4f` (INNHOLD, ikke git-blob) |
|
||||
| Suiten er et strengt supersett | `comm -23` node-id-er før/etter: **0 fjernet, 26 lagt til** |
|
||||
|
|
@ -434,7 +434,12 @@ def tool_call_payload(calls: Sequence[ToolCall]) -> list[dict[str, Any]]:
|
|||
|
||||
|
||||
def trace_payload(
|
||||
trace: ExplorationTrace, *, stop: str | None, completed: bool, mandate: Mandate | None
|
||||
trace: ExplorationTrace,
|
||||
*,
|
||||
stop: str | None,
|
||||
completed: bool,
|
||||
mandate: Mandate | None,
|
||||
prepass: Mapping[str, Any] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""The ONE rendering of a trace into plain data for ``outbox.write_exploration``.
|
||||
|
||||
|
|
@ -454,10 +459,18 @@ def trace_payload(
|
|||
a park), which is exactly what ``completed=False`` already says. Collapsing that into an
|
||||
empty list would make "the loop formed no approaches" and "the loop never got that far"
|
||||
unreadable from each other.
|
||||
|
||||
``prepass`` is the declared cut this exploration was SEEDED with (``--prepass-seed``), already
|
||||
rendered by ``prepass.declaration_payload`` so this module and the outbox both stay free of
|
||||
that dependency. It DEFAULTS to ``None``, unlike ``completed`` and ``mandate`` above: absence
|
||||
here is an honest positive statement — no cut was given — which is ``Bundle.skipped``'s empty
|
||||
tuple rather than ``cost_baseline_anchored``'s required boolean, and it is the same decision
|
||||
``RunResult.prepass`` already made one surface over.
|
||||
"""
|
||||
return {
|
||||
"completed": completed,
|
||||
"stop": stop,
|
||||
"prepass": dict(prepass) if prepass is not None else None,
|
||||
"tokens_spent": trace.tokens_spent,
|
||||
"approaches": [
|
||||
{
|
||||
|
|
@ -1516,6 +1529,7 @@ async def explore(
|
|||
success_criteria: str = "",
|
||||
trace: ExplorationTrace | None = None,
|
||||
checkpoint_dir: str | None = None,
|
||||
seed_context: str = "",
|
||||
) -> ExplorationResult:
|
||||
"""Explore the knowledge bases and return the ``Mandate`` the pipeline should evaluate.
|
||||
|
||||
|
|
@ -1554,6 +1568,20 @@ async def explore(
|
|||
``trace`` is the caller's accumulator and is the ONLY way to see what a run that RAISED
|
||||
produced: both budget channels destroy the ``ExplorationResult`` before it exists. When it is
|
||||
omitted a private one is used, so the returned result is unchanged for every existing caller.
|
||||
|
||||
``seed_context`` is material the CALLER has already verified and wants the loop to START from
|
||||
— today, one contract-conformant OKF pre-pass cut (``--prepass-seed``). It joins the TASK
|
||||
MESSAGE and deliberately NOT ``objective``: the objective is what a person commissioned and
|
||||
what ``Mandate.announce`` prints back to them, so folding ten thousand tokens of excerpts into
|
||||
it would make the commission unreadable and would put the cut's text into every artefact that
|
||||
quotes the objective. Empty by default, and an empty string leaves the task message
|
||||
byte-identical to what it has always been — which is what makes "without the flag, nothing
|
||||
changed" a property rather than a promise.
|
||||
|
||||
**This is a starting point, never a boundary.** The navigator tools are built exactly as they
|
||||
are without it, because an exploration that could not read past its seed would be the OTHER
|
||||
arm (``--prepass-payload``, where the cut REPLACES the base and the tools are withdrawn), and
|
||||
building both behaviours behind one name is how a flag stops meaning anything.
|
||||
"""
|
||||
if plan_reviewer is not None and checkpoint_dir is not None:
|
||||
raise ExplorationError(
|
||||
|
|
@ -1614,7 +1642,10 @@ async def explore(
|
|||
# With no provider installed this is a no-op tracer and every event is discarded, which is
|
||||
# exactly what "tracing is off" has meant since U14.
|
||||
with exploration_tracer().start_as_current_span(EXPLORATION_SPAN) as span:
|
||||
result = await workflow.run(prompt)
|
||||
# The seed rides in the TASK MESSAGE, never in ``prompt`` itself: ``_finish`` below builds
|
||||
# the mandate's ``objective`` from ``prompt``, and a commission whose objective carried the
|
||||
# whole cut would be unreadable to the person who wrote it.
|
||||
result = await workflow.run(f"{prompt}\n\n{seed_context}" if seed_context else prompt)
|
||||
stop, _ = await _drive(
|
||||
workflow,
|
||||
result,
|
||||
|
|
|
|||
|
|
@ -347,6 +347,52 @@ def _in_dimension(path: Path, dimension: str) -> bool:
|
|||
)
|
||||
|
||||
|
||||
# --- The one admission gate, shared by every door that consumes a payload --------------------
|
||||
|
||||
|
||||
def admit_payload(
|
||||
payload: PrepassPayload,
|
||||
*,
|
||||
bundle_dir: str,
|
||||
resolved_id: okf.ResolvedBundleId,
|
||||
dimension: str | None = None,
|
||||
) -> None:
|
||||
"""Everything that must hold before a payload may shape a run. Raises, or returns nothing.
|
||||
|
||||
ONE copy, because there are now TWO doors onto this file — the debate's ``--prepass-payload``
|
||||
(the cut REPLACES the pointer and the tools are withdrawn) and the exploration's
|
||||
``--prepass-seed`` (the cut is the STARTING POINT and the tools stay). The two arms differ in
|
||||
what they do with an admitted payload and in nothing at all about what makes one admissible,
|
||||
and two copies of an admission rule is the ko-(p) drift that would let one door accept what
|
||||
the other refuses.
|
||||
|
||||
The three steps, in this order and for this reason: the shape gate reads no disk and so is
|
||||
free, the bundle check is the expensive one, and the empty-delivery refusal comes last because
|
||||
a payload that does not hold has not earned an interpretation of its own emptiness.
|
||||
|
||||
**``delivered == 0`` is refused on BOTH arms**, and on the seeding arm that is a decision
|
||||
rather than an inheritance. Measured, it is reachable only when every concept failed to match
|
||||
lexically (the producer REFUSES the other empty case, where concepts matched and the budget
|
||||
admitted none), so it is evidence of ABSENCE for this question at this ref. On the seeding arm
|
||||
a caller might argue the tools are still there and the run could proceed — but it would then
|
||||
proceed as a PLAIN exploration while the operator had asked for a seeded one, which is the
|
||||
silently downgraded order ``load_mandate`` fail-fasts against.
|
||||
"""
|
||||
check_payload_shape(payload)
|
||||
verify_against_bundle(
|
||||
payload, bundle_dir=bundle_dir, resolved_id=resolved_id, dimension=dimension
|
||||
)
|
||||
if not payload.excerpts:
|
||||
# Saying it beats two silent alternatives: an empty prompt, or falling through to
|
||||
# ``run_project``'s citation guard, whose message names ``docs_dir`` — ``None`` on this
|
||||
# path. SS 7.3's own posture: the skill stops and says so.
|
||||
raise PrepassRefused(
|
||||
f"the pre-pass delivered 0 of {payload.denominators.considered} concepts for the "
|
||||
f"question {payload.question!r} at ref {payload.bundle.ref}; an empty cut is evidence "
|
||||
"that this knowledge base does not answer that question, not something to run over"
|
||||
)
|
||||
|
||||
|
||||
# --- Rendering the cut for the prompt --------------------------------------------------------
|
||||
|
||||
|
||||
|
|
@ -382,24 +428,86 @@ def render_context(payload: PrepassPayload) -> str:
|
|||
GATE is ``verify_against_bundle``, which makes the text re-derivable from the mounted base, so
|
||||
a payload cannot deliver bytes the base does not hold.
|
||||
"""
|
||||
return "\n".join(
|
||||
_declaration_lines(payload)
|
||||
+ [
|
||||
"",
|
||||
"Each excerpt below matched the question LEXICALLY. That is not the same as answering "
|
||||
"it: a base that holds no answer still returns its closest matches. If the delivered "
|
||||
"text does not support a claim, say so with [sourced-not-sufficient] rather than "
|
||||
"filling the gap. adjudication and trust_tier are the producer's declarations about "
|
||||
"each document, carried here unchanged.",
|
||||
]
|
||||
+ _data_blocks(payload)
|
||||
)
|
||||
|
||||
|
||||
def render_seed(payload: PrepassPayload) -> str:
|
||||
"""What the EXPLORATION is handed IN ADDITION to its prompt, keeping its navigation tools.
|
||||
|
||||
The other arm of one decision, and the difference is a single fact stated in both directions:
|
||||
:func:`render_context` says "you have no tools to read further, what is below is all of it",
|
||||
which is true there and would be a LIE here. Contract SS 2.2 forbids reading "outside what the
|
||||
payload delivers **or explicitly names as reachable**" — the second clause is what makes this
|
||||
arm conformant, and a rendering that did not say the rest was reachable would leave a model
|
||||
obeying the first clause while holding the tools for the second.
|
||||
|
||||
**Not a second copy of the rendering rule.** The declaration header, the rule -> COUNT folding
|
||||
and the delimited DATA blocks are the SAME functions the other arm uses; what differs is the
|
||||
one paragraph that tells the reader what it may do next. Two full copies would drift, and a
|
||||
drifted pair would state two different cuts for one run (ko-(p)).
|
||||
|
||||
``[unread]`` is used deliberately, and it is one of contract SS 4.1's five required literals:
|
||||
a withheld concept in this arm is not absent and not unavailable, it is simply not yet read —
|
||||
and saying so is what turns the withheld list from a boundary into a next step.
|
||||
"""
|
||||
return "\n".join(
|
||||
_declaration_lines(payload, rest_reachable=True)
|
||||
+ [
|
||||
"",
|
||||
"Each excerpt below matched the question LEXICALLY. That is not the same as answering "
|
||||
"it: a base that holds no answer still returns its closest matches. The withheld "
|
||||
"concepts are [unread], not absent — if the delivered text does not support a claim, "
|
||||
"OPEN THE BASE with your navigation tools rather than filling the gap, and reserve "
|
||||
"[sourced-not-sufficient] for a claim the base itself could not support. Report what "
|
||||
"you actually read. adjudication and trust_tier are the producer's declarations about "
|
||||
"each document, carried here unchanged.",
|
||||
]
|
||||
+ _data_blocks(payload)
|
||||
)
|
||||
|
||||
|
||||
def _declaration_lines(payload: PrepassPayload, *, rest_reachable: bool = False) -> list[str]:
|
||||
"""The header both renderings open with: the base, what the cut may be used for, the counts.
|
||||
|
||||
ONE copy, because these lines ARE the declaration (SS 2.3) and two of them would be two
|
||||
answers to "what was this run's cut". Only the second line differs between the arms, and it
|
||||
differs on exactly the fact ``PrepassDeclaration.rest_reachable`` carries.
|
||||
"""
|
||||
counts = payload.denominators
|
||||
rules = ", ".join(f"{rule} ({count})" for rule, count in declaration_of(payload).withheld_rules)
|
||||
lines = [
|
||||
rules = ", ".join(f"{rule} ({count})" for rule, count in withheld_rule_counts(payload))
|
||||
stance = (
|
||||
"You are reading a DECLARED CUT of that base as your STARTING POINT, not as a replacement "
|
||||
"for it. The rest of the base stays reachable with your navigation tools, and you are "
|
||||
"expected to use them when the cut does not carry what you need."
|
||||
if rest_reachable
|
||||
else "You are reading a DECLARED CUT of that base, not the base itself, and you have no "
|
||||
"tools to read further. What is below is all of it."
|
||||
)
|
||||
return [
|
||||
f"Knowledge base: {payload.bundle.bundle_id} (ref {payload.bundle.ref}).",
|
||||
"",
|
||||
"You are reading a DECLARED CUT of that base, not the base itself, and you have no tools "
|
||||
"to read further. What is below is all of it.",
|
||||
stance,
|
||||
f"The cut was computed for this question: {payload.question}",
|
||||
f"Concepts considered: {counts.considered}. Withheld: {counts.withheld}. "
|
||||
f"Delivered below: {counts.delivered}.",
|
||||
f"Withheld by rule: {rules}." if rules else "Withheld by rule: none.",
|
||||
"",
|
||||
"Each excerpt below matched the question LEXICALLY. That is not the same as answering it: "
|
||||
"a base that holds no answer still returns its closest matches. If the delivered text "
|
||||
"does not support a claim, say so with [sourced-not-sufficient] rather than filling the "
|
||||
"gap. adjudication and trust_tier are the producer's declarations about each document, "
|
||||
"carried here unchanged.",
|
||||
]
|
||||
|
||||
|
||||
def _data_blocks(payload: PrepassPayload) -> list[str]:
|
||||
"""The delimited DATA blocks, one per delivered excerpt (SS 9.3), shared by both arms."""
|
||||
lines: list[str] = []
|
||||
for excerpt in payload.excerpts:
|
||||
lines += [
|
||||
"",
|
||||
|
|
@ -408,7 +516,7 @@ def render_context(payload: PrepassPayload) -> str:
|
|||
excerpt.text,
|
||||
f"--- END DATA {excerpt.concept_id} ---",
|
||||
]
|
||||
return "\n".join(lines)
|
||||
return lines
|
||||
|
||||
|
||||
# --- The declaration -------------------------------------------------------------------------
|
||||
|
|
@ -421,6 +529,15 @@ class PrepassDeclaration:
|
|||
It reports the payload's OWN denominators verbatim, never a recount off the navigated bundle:
|
||||
the pre-pass may legitimately have considered a different set (it counts the verdict layer,
|
||||
``Bundle.context_files`` does not), and two numbers for one fact is ko-(p).
|
||||
|
||||
``rest_reachable`` says which of the two arms consumed the cut, and it is REQUIRED WITHOUT A
|
||||
DEFAULT for ``ProvenanceStamp.cost_baseline_anchored``'s reason: **both defaults would lie.**
|
||||
``False`` would let a run that kept its navigation tools publish a declaration claiming the cut
|
||||
was all it could read; ``True`` would let the arm that WITHDREW them claim the base stayed
|
||||
open. It is the contract's own distinction -- SS 2.2 forbids reading "outside what the payload
|
||||
delivers **or explicitly names as reachable**", so a conformant consumer may keep the base
|
||||
reachable, and the difference between the two readings is exactly what a reader of this
|
||||
declaration needs to know.
|
||||
"""
|
||||
|
||||
bundle_id: str
|
||||
|
|
@ -430,19 +547,34 @@ class PrepassDeclaration:
|
|||
withheld: int
|
||||
delivered: int
|
||||
withheld_rules: tuple[tuple[str, int], ...]
|
||||
rest_reachable: bool
|
||||
|
||||
|
||||
def declaration_of(payload: PrepassPayload) -> PrepassDeclaration:
|
||||
"""The declaration a verified payload supports.
|
||||
def withheld_rule_counts(payload: PrepassPayload) -> tuple[tuple[str, int], ...]:
|
||||
"""rule -> COUNT, sorted. The ONE folding of the withheld list in this repository.
|
||||
|
||||
``withheld_rules`` is rule -> COUNT, sorted. The concept ids are deliberately NOT carried:
|
||||
measured on a 629-concept corpus the withheld list alone is 34 451 o200k tokens, and a rule
|
||||
name is the fact a reader can act on while a list of ids they cannot open is cost without
|
||||
information. The full list stays in the payload, one artefact away.
|
||||
The concept ids are deliberately NOT carried: measured on a 629-concept corpus the withheld
|
||||
list alone is 34 451 o200k tokens, and a rule name is the fact a reader can act on while a
|
||||
list of ids they cannot open is cost without information. The full list stays in the payload,
|
||||
one artefact away.
|
||||
|
||||
It is a function of its own rather than a step inside ``declaration_of`` because THREE
|
||||
surfaces need it -- the declaration and both renderings -- and a second copy of "how the
|
||||
withheld list folds" would be free to disagree about the run it describes (ko-(p)).
|
||||
"""
|
||||
counts: dict[str, int] = {}
|
||||
for entry in payload.withheld:
|
||||
counts[entry.rule] = counts.get(entry.rule, 0) + 1
|
||||
return tuple(sorted(counts.items()))
|
||||
|
||||
|
||||
def declaration_of(payload: PrepassPayload, *, rest_reachable: bool) -> PrepassDeclaration:
|
||||
"""The declaration a verified payload supports, for the arm that consumed it.
|
||||
|
||||
``rest_reachable`` is a REQUIRED keyword: see :class:`PrepassDeclaration`. The payload cannot
|
||||
supply it -- a cut does not know what its consumer did with the navigation tools -- so it is
|
||||
the caller's to state, and every caller states it.
|
||||
"""
|
||||
return PrepassDeclaration(
|
||||
bundle_id=payload.bundle.bundle_id,
|
||||
ref=payload.bundle.ref,
|
||||
|
|
@ -450,7 +582,8 @@ def declaration_of(payload: PrepassPayload) -> PrepassDeclaration:
|
|||
considered=payload.denominators.considered,
|
||||
withheld=payload.denominators.withheld,
|
||||
delivered=payload.denominators.delivered,
|
||||
withheld_rules=tuple(sorted(counts.items())),
|
||||
withheld_rules=withheld_rule_counts(payload),
|
||||
rest_reachable=rest_reachable,
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -466,4 +599,8 @@ def declaration_payload(declaration: PrepassDeclaration) -> Mapping[str, object]
|
|||
"withheld_rules": [
|
||||
{"rule": rule, "count": count} for rule, count in declaration.withheld_rules
|
||||
],
|
||||
# Which arm read the cut. An artefact that reported the denominators without saying
|
||||
# whether the consumer could still open the base would leave a reader unable to tell a
|
||||
# bounded run from a seeded one -- the same undeclared claim, one level up.
|
||||
"rest_reachable": declaration.rest_reachable,
|
||||
}
|
||||
|
|
|
|||
|
|
@ -860,8 +860,17 @@ def prepass_notice(declaration: prepass.PrepassDeclaration | None) -> str | None
|
|||
if declaration is None:
|
||||
return None
|
||||
rules = ", ".join(f"{rule} ({count})" for rule, count in declaration.withheld_rules)
|
||||
# ONE renderer, two arms — because the number of concepts delivered says nothing at all about
|
||||
# whether the run could still open the rest, and an operator reading a summary that reported
|
||||
# "8 of 630" without that would draw the wrong conclusion in exactly one of the two cases.
|
||||
role = (
|
||||
"a DECLARED CUT SEEDED the exploration (the rest of the base stayed reachable with the "
|
||||
"navigation tools)"
|
||||
if declaration.rest_reachable
|
||||
else "a DECLARED CUT was used"
|
||||
)
|
||||
return (
|
||||
f" Knowledge base: a DECLARED CUT was used — {declaration.delivered} of "
|
||||
f" Knowledge base: {role} — {declaration.delivered} of "
|
||||
f"{declaration.considered} concept(s) delivered, {declaration.withheld} withheld"
|
||||
f"{' by rule: ' + rules if rules else ''}. "
|
||||
f"Base {declaration.bundle_id} at ref {declaration.ref}; "
|
||||
|
|
@ -1059,27 +1068,19 @@ async def run_project(
|
|||
prepass_declaration: prepass.PrepassDeclaration | None = None
|
||||
debate_tools: list[Any]
|
||||
if prepass_payload is not None:
|
||||
prepass.check_payload_shape(prepass_payload)
|
||||
prepass.verify_against_bundle(
|
||||
# ONE admission gate, shared with the exploration's seeding door: shape, then the
|
||||
# mounted base, then the empty-delivery refusal. Two copies of what makes a payload
|
||||
# admissible would let one door accept what the other refuses (kø-(p)).
|
||||
prepass.admit_payload(
|
||||
prepass_payload,
|
||||
bundle_dir=bundle_dir,
|
||||
resolved_id=resolved,
|
||||
dimension=dimension_id,
|
||||
)
|
||||
if not prepass_payload.excerpts:
|
||||
# Measured: ``delivered == 0`` is reachable only when every concept failed to
|
||||
# match (the producer REFUSES the other empty case, where concepts matched and the
|
||||
# budget admitted none). So this is evidence of ABSENCE for this question at this
|
||||
# ref, and saying it is better than two silent alternatives: an empty prompt, or
|
||||
# falling through to the citation guard below, whose message names ``docs_dir`` —
|
||||
# ``None`` on this path. SS 7.3's own posture: the skill stops and says so.
|
||||
raise prepass.PrepassRefused(
|
||||
f"the pre-pass delivered 0 of {prepass_payload.denominators.considered} "
|
||||
f"concepts for the question {prepass_payload.question!r} at ref "
|
||||
f"{prepass_payload.bundle.ref}; an empty cut is evidence that this knowledge "
|
||||
"base does not answer that question, not something to run a debate over"
|
||||
)
|
||||
prepass_declaration = prepass.declaration_of(prepass_payload)
|
||||
# ``rest_reachable=False`` is the FACT this arm establishes four lines below by
|
||||
# emptying ``debate_tools`` — stated, never defaulted, because the other arm keeps
|
||||
# them and a declaration that could not tell the two apart would describe neither.
|
||||
prepass_declaration = prepass.declaration_of(prepass_payload, rest_reachable=False)
|
||||
context = prepass.render_context(prepass_payload)
|
||||
# Citations over the DELIVERED concepts alone: a stamp citing the whole corpus for a
|
||||
# proposal that saw eight documents re-creates the undeclared claim this seam removes.
|
||||
|
|
@ -2395,6 +2396,22 @@ def main(argv: list[str] | None = None) -> int:
|
|||
"--bundle-dir; refused with --portfolio, --report, --proposals-from-mandate, "
|
||||
"--dimension-config and --explore.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--prepass-seed",
|
||||
default=None,
|
||||
metavar="FILE",
|
||||
help="The OTHER arm of the same door (Q5=B): hand the EXPLORATION a declared cut as its "
|
||||
"STARTING POINT and KEEP the four navigator tools. FILE is one contract-conformant OKF "
|
||||
"consumption pre-pass payload (okf-consumption/1), admitted by exactly the checks "
|
||||
"--prepass-payload applies — identity, sha256, the delivered text re-derived from the "
|
||||
"mounted document, the verdict layer — before a single model call. The difference is what "
|
||||
"happens next: the cut and its three denominators join the exploration's task message, "
|
||||
"and the loop may still open anything else in the base (contract SS 2.2 permits what the "
|
||||
"payload 'explicitly names as reachable'). The cut is printed and recorded in "
|
||||
"{run_id}-exploration.json, which says rest_reachable so a reader can tell a seeded run "
|
||||
"from a bounded one. REQUIRES --explore; refused with --prepass-payload (two opposite "
|
||||
"arms of one decision), --portfolio, --report and --checkpoint-dir.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--proposal-review",
|
||||
action="store_true",
|
||||
|
|
@ -2655,6 +2672,9 @@ def main(argv: list[str] | None = None) -> int:
|
|||
# ABOVE every dispatch, so an omission here is a silent DROP — the operator would be
|
||||
# told nothing and the declared cut would simply never happen.
|
||||
"--prepass-payload": args.prepass_payload is not None,
|
||||
# And the seeding arm, listed for the identical reason: report mode returns above the
|
||||
# exploration dispatch too, so an omission here is a silent DROP and not a refusal.
|
||||
"--prepass-seed": args.prepass_seed is not None,
|
||||
# The three U12 flags, listed for exactly that reason: report mode returns before the
|
||||
# resume dispatch, so an omission here is a silent drop, not a refusal.
|
||||
"--checkpoint-dir": args.checkpoint_dir is not None,
|
||||
|
|
@ -2744,6 +2764,10 @@ def main(argv: list[str] | None = None) -> int:
|
|||
# operator who wrote --portfolio --prepass-payload has to hear which of the two is
|
||||
# wrong, and an arm asserting on the shared token could not tell the two apart.
|
||||
"--prepass-payload": args.prepass_payload,
|
||||
# The seeding arm sits on the same side and by NAME for the same reason: it requires
|
||||
# --explore, which this mode also refuses, so falling through would tell an operator
|
||||
# who wrote --portfolio --prepass-seed to add a second flag --portfolio refuses too.
|
||||
"--prepass-seed": args.prepass_seed,
|
||||
# And the asynchronous half of the same door, on the same side of the partition and by
|
||||
# NAME for the same reason.
|
||||
"--checkpoint-dir": args.checkpoint_dir,
|
||||
|
|
@ -2794,6 +2818,49 @@ def main(argv: list[str] | None = None) -> int:
|
|||
)
|
||||
return 1
|
||||
|
||||
# The SEEDING arm's three refusals, placed ABOVE the replacing arm's block on purpose: given
|
||||
# both flags, the block below would answer with "--prepass-payload and --explore cannot be
|
||||
# combined", which names neither of the two flags the operator actually put in conflict. At
|
||||
# FUNCTION level, never nested under another flag's branch, for the F4 reason its neighbours
|
||||
# are: under one, a bare combination falls through to a dispatch that drops the flag silently.
|
||||
if not args.portfolio and args.prepass_seed is not None:
|
||||
if args.prepass_payload is not None:
|
||||
# TWO OPPOSITE ARMS OF ONE DECISION, and merging them is not defined: one WITHDRAWS
|
||||
# the navigator tools because the cut replaces the base, the other KEEPS them because
|
||||
# the cut is where to start. A run holding both would have to silently pick, and the
|
||||
# picked one would be a policy nobody wrote down.
|
||||
print(
|
||||
"run refused: --prepass-payload and --prepass-seed are two opposite arms of one "
|
||||
"decision (the first REPLACES the knowledge base with the cut and withdraws the "
|
||||
"navigation tools; the second uses the cut as a STARTING POINT and keeps them). "
|
||||
"Choose which one this run is",
|
||||
file=sys.stderr,
|
||||
)
|
||||
return 1
|
||||
if args.explore is None:
|
||||
# The ``--explore-config`` case verbatim: a seed for a loop that never runs would be
|
||||
# loaded, verified, paid for in I/O and then dropped.
|
||||
print(
|
||||
"run refused: --prepass-seed requires --explore (the cut seeds an EXPLORATION's "
|
||||
"starting point; without one there is no loop to seed, and the debate's own arm "
|
||||
"is --prepass-payload)",
|
||||
file=sys.stderr,
|
||||
)
|
||||
return 1
|
||||
if args.checkpoint_dir is not None:
|
||||
# A parked leg writes {run_id}-exploration.json from its ``finally`` WITH the
|
||||
# declaration; the resumed leg runs in a process that never saw the payload and
|
||||
# OVERWRITES the same file with ``prepass: null``. A declaration that evaporates
|
||||
# halfway is worse than one refused, and it would do so silently — so this is refused
|
||||
# rather than left to erase itself.
|
||||
print(
|
||||
"run refused: --prepass-seed and --checkpoint-dir cannot be combined (the resumed "
|
||||
"leg runs in a process that never saw the payload, and its own artefact would "
|
||||
"overwrite the parked leg's declaration of the cut with nothing)",
|
||||
file=sys.stderr,
|
||||
)
|
||||
return 1
|
||||
|
||||
# The pre-pass door's four remaining refusals, at FUNCTION level and never nested under
|
||||
# another flag's branch: under one, a bare combination would fall straight through to a
|
||||
# dispatch that drops the payload in silence (the F4 class).
|
||||
|
|
@ -3214,6 +3281,16 @@ def main(argv: list[str] | None = None) -> int:
|
|||
print(f"run refused: {exc}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
# The seeding arm's payload, loaded here and by the SAME loader: one reader of one file
|
||||
# format, never a second (kø-(p)). What differs between the two arms starts after admission.
|
||||
prepass_seed: prepass.PrepassPayload | None = None
|
||||
if args.prepass_seed is not None:
|
||||
try:
|
||||
prepass_seed = load_prepass_payload(args.prepass_seed)
|
||||
except (FileNotFoundError, ValidationError, ValueError) as exc:
|
||||
print(f"run refused: {exc}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
# The egress config, loaded fail-fast alongside the commission. Degrading a broken one to "no
|
||||
# external services" would make the announcement describe a run nobody configured, and a
|
||||
# partially-parsed one could contact a subset nobody chose.
|
||||
|
|
@ -3316,6 +3393,41 @@ def main(argv: list[str] | None = None) -> int:
|
|||
else PlanReviewDecision.revise(answer.feedback),
|
||||
)
|
||||
|
||||
# The seeding arm's admission, HOISTED above the try/finally below for the økt-57 reason the
|
||||
# resume loads are hoisted: this is a refusal, and a refusal must not first overwrite
|
||||
# {run_id}-exploration.json with an empty trace. It also has to happen before the first model
|
||||
# call — at the exit code, a refusal after the spend looks exactly like one before it.
|
||||
seed_declaration: prepass.PrepassDeclaration | None = None
|
||||
seed_context = ""
|
||||
if prepass_seed is not None:
|
||||
assert args.bundle_dir is not None # guarded: --prepass-seed -> --explore -> --bundle-dir
|
||||
try:
|
||||
# This door OPENS a base, so it carries that gate itself rather than trusting the
|
||||
# tool-side one to fire later (S7a-3: every opening door gets its own, and its own
|
||||
# mutation). A base whose concepts name two corpora cannot be the one a cut is OF.
|
||||
okf.assert_declared_ids_agree(okf.navigate_bundle(args.bundle_dir))
|
||||
prepass.admit_payload(
|
||||
prepass_seed,
|
||||
bundle_dir=args.bundle_dir,
|
||||
resolved_id=okf.reconcile_bundle_id(args.bundle_dir),
|
||||
# NOT scoped, and that is measured rather than overlooked: the exploration builds
|
||||
# ``navigator_tools(bundle_dirs)`` with no dimension at all, so scoping the seed
|
||||
# would refuse text the very same loop can open with ``read_file`` a moment later.
|
||||
# The DEBATE downstream stays scoped exactly as it is today.
|
||||
dimension=None,
|
||||
)
|
||||
except (FileNotFoundError, ValidationError, ValueError) as exc:
|
||||
print(f"run refused: {exc}", file=sys.stderr)
|
||||
return 1
|
||||
seed_declaration = prepass.declaration_of(prepass_seed, rest_reachable=True)
|
||||
seed_context = prepass.render_seed(prepass_seed)
|
||||
# Printed HERE, before the loop it seeds, rather than after: the announcement contract is
|
||||
# that a commission is declared before the work it commissions, and a run whose budget is
|
||||
# exhausted mid-exploration must still have said what it started from.
|
||||
cut_notice = prepass_notice(seed_declaration)
|
||||
assert cut_notice is not None # a declaration always renders; ``None`` means no cut
|
||||
print(cut_notice)
|
||||
|
||||
if args.explore is not None or resumed is not None:
|
||||
exploration_trace = ExplorationTrace()
|
||||
exploration: ExplorationResult | None = None
|
||||
|
|
@ -3356,6 +3468,9 @@ def main(argv: list[str] | None = None) -> int:
|
|||
plan_reviewer=terminal_plan_reviewer() if args.plan_review else None,
|
||||
# The U12 door. Mutually exclusive with the one above, refused at the top.
|
||||
checkpoint_dir=args.checkpoint_dir,
|
||||
# Q5=B. Empty unless --prepass-seed was given, and an empty string leaves
|
||||
# the task message byte-identical to what every run built before today.
|
||||
seed_context=seed_context,
|
||||
)
|
||||
)
|
||||
except PlanReviewParked as parked_exc:
|
||||
|
|
@ -3387,6 +3502,13 @@ def main(argv: list[str] | None = None) -> int:
|
|||
stop=exploration.stop if exploration is not None else None,
|
||||
completed=exploration is not None,
|
||||
mandate=exploration.mandate if exploration is not None else None,
|
||||
# From the SAME object the notice was rendered from, never a second load:
|
||||
# stdout and the artefact must describe one cut, not two (kø-(p)).
|
||||
prepass=(
|
||||
prepass.declaration_payload(seed_declaration)
|
||||
if seed_declaration is not None
|
||||
else None
|
||||
),
|
||||
),
|
||||
)
|
||||
if budget_now is not None:
|
||||
|
|
|
|||
|
|
@ -201,7 +201,7 @@ def test_the_declaration_reports_the_payloads_own_denominators() -> None:
|
|||
"""Never a recount off the navigated bundle: the pre-pass counts the verdict layer and
|
||||
`Bundle.context_files` does not, so two numbers for one fact would be kø-(p)."""
|
||||
payload = prepass.load_prepass_payload(str(FIXTURE))
|
||||
declaration = prepass.declaration_of(payload)
|
||||
declaration = prepass.declaration_of(payload, rest_reachable=False)
|
||||
assert declaration.considered == payload.denominators.considered == 5
|
||||
assert declaration.withheld == payload.denominators.withheld == 1
|
||||
assert declaration.delivered == payload.denominators.delivered == 4
|
||||
|
|
@ -211,14 +211,16 @@ def test_the_declaration_reports_the_payloads_own_denominators() -> None:
|
|||
|
||||
def test_the_declaration_counts_the_withheld_rules_without_naming_the_concepts() -> None:
|
||||
payload = prepass.load_prepass_payload(str(FIXTURE))
|
||||
declaration = prepass.declaration_of(payload)
|
||||
declaration = prepass.declaration_of(payload, rest_reachable=False)
|
||||
assert declaration.withheld_rules == (("verdict_layer_excluded", 1),)
|
||||
assert sum(count for _, count in declaration.withheld_rules) == declaration.withheld
|
||||
|
||||
|
||||
def test_the_declaration_payload_is_a_plain_mapping() -> None:
|
||||
"""The outbox stays framework-free (`explore.trace_payload`'s rule)."""
|
||||
declaration = prepass.declaration_of(prepass.load_prepass_payload(str(FIXTURE)))
|
||||
declaration = prepass.declaration_of(
|
||||
prepass.load_prepass_payload(str(FIXTURE)), rest_reachable=False
|
||||
)
|
||||
body = prepass.declaration_payload(declaration)
|
||||
assert json.loads(json.dumps(body))["withheld_rules"] == [
|
||||
{"rule": "verdict_layer_excluded", "count": 1}
|
||||
|
|
|
|||
|
|
@ -409,7 +409,9 @@ def test_the_notice_is_omitted_without_a_payload() -> None:
|
|||
|
||||
|
||||
def test_the_notice_reports_the_cut_when_there_was_one() -> None:
|
||||
declaration = prepass.declaration_of(prepass.load_prepass_payload(str(FIXTURE)))
|
||||
declaration = prepass.declaration_of(
|
||||
prepass.load_prepass_payload(str(FIXTURE)), rest_reachable=False
|
||||
)
|
||||
line = run_module.prepass_notice(declaration)
|
||||
assert line is not None
|
||||
assert "4 of 5" in line
|
||||
|
|
@ -425,7 +427,9 @@ def test_a_withheld_rule_is_counted_and_the_concept_is_never_named() -> None:
|
|||
] + raw["withheld"]
|
||||
raw["denominators"]["withheld"] = len(raw["withheld"])
|
||||
raw["denominators"]["considered"] = len(raw["withheld"]) + raw["denominators"]["delivered"]
|
||||
declaration = prepass.declaration_of(prepass.PrepassPayload.model_validate(raw))
|
||||
declaration = prepass.declaration_of(
|
||||
prepass.PrepassPayload.model_validate(raw), rest_reachable=False
|
||||
)
|
||||
line = run_module.prepass_notice(declaration)
|
||||
assert line is not None and "no_lexical_match (20)" in line
|
||||
assert "hemmelig-konsept-0" not in line
|
||||
|
|
|
|||
754
tests/test_prepass_seed_door_loadbearing.py
Normal file
754
tests/test_prepass_seed_door_loadbearing.py
Normal file
|
|
@ -0,0 +1,754 @@
|
|||
"""Load-bearing gate for ``--prepass-seed`` — the pre-pass as the exploration's STARTING POINT.
|
||||
|
||||
Q5 = B (operator decision 2026-09-07). Arm A stands: ``--prepass-payload`` hands the DEBATE a
|
||||
declared cut, withdraws the four navigator tools, and is still refused together with ``--explore``
|
||||
for exactly the reason it was — an exploration reads the whole base with the very tools the
|
||||
payload withdraws, so the run as a whole would read far outside the cut it declares.
|
||||
|
||||
This is the OTHER arm, and it is a different flag rather than a loosening of that refusal: the
|
||||
cut seeds the exploration's task message and the tools STAY. Contract § 2.2 is what makes it
|
||||
conformant — a skill may not read "outside what the payload delivers **or explicitly names as
|
||||
reachable**" — and ``PrepassDeclaration.rest_reachable`` is what makes the two readings tellable
|
||||
apart afterwards.
|
||||
|
||||
Four properties, and each arm is paired with the control that stops it being satisfied by a run
|
||||
that did nothing:
|
||||
|
||||
(a) the flag is accepted only together with ``--explore``, and refused BY NAME everywhere it
|
||||
would otherwise be silently dropped;
|
||||
(b) the delivered text lands in the exploration's STARTING POINT, and the navigator tools are
|
||||
still there — asserted by making a scripted navigator actually call two of them;
|
||||
(c) the declaration says the cut was a starting point rather than a replacement, on stdout AND in
|
||||
``{run_id}-exploration.json``, with arm A's ``rest_reachable: false`` as the paired control;
|
||||
(d) without the flag every byte of today's behaviour stands.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
|
||||
import portfolio_optimiser.simulation as simulation_module
|
||||
import portfolio_optimiser.run as run_module
|
||||
from portfolio_optimiser import prepass
|
||||
from portfolio_optimiser.run import main
|
||||
|
||||
FIXTURE = Path(__file__).parent / "fixtures" / "prepass" / "bygg-energi-mikro-fixture.payload.json"
|
||||
SHIPPED_BASE = Path(__file__).parent.parent / "shared" / "examples" / "bygg-energi-mikro"
|
||||
PROJECT_ID = "BYGG-KONTOR-NORD"
|
||||
DECLARED_ID = "bygg-energi-mikro-fixture"
|
||||
|
||||
#: Written INTO the mounted base and then into the payload's excerpt, so its only route to a
|
||||
#: prompt is the seed itself: the scripted sink records ``message.text`` only, and a tool RESULT
|
||||
#: is ``contents`` — so a navigator that opened the same document could not put it there.
|
||||
_EXCERPT_SENTINEL = "SENTINEL-I-ET-LEVERT-UTDRAG"
|
||||
|
||||
_PROPOSER_REPLY = json.dumps(
|
||||
{
|
||||
"measure": "energy_efficiency",
|
||||
"affected_items": [{"code": "ENERGI-TOTAL-EL", "quantity": 120000.0, "unit_cost": 1.25}],
|
||||
"claimed_saving_nok": 30000,
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def _ledger(*, satisfied: bool, speaker: str = "navigator") -> str:
|
||||
return json.dumps(
|
||||
{
|
||||
"is_request_satisfied": {"reason": "r", "answer": satisfied},
|
||||
"is_in_loop": {"reason": "r", "answer": False},
|
||||
"is_progress_being_made": {"reason": "r", "answer": True},
|
||||
"next_speaker": {"reason": "r", "answer": speaker},
|
||||
"instruction_or_question": {"reason": "r", "answer": "open a base"},
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
#: One entry per stage the orchestrator asks for. A single constant answers all five with the
|
||||
#: first round's ledger and concludes after one round having given nobody a turn (MAJOR-1's own
|
||||
#: measurement), which would make the tool arm below vacuous.
|
||||
_MANAGER_STAGES = [
|
||||
"FACTS: the base is anchored.",
|
||||
"PLAN: - let the navigator open a base",
|
||||
_ledger(satisfied=False),
|
||||
_ledger(satisfied=True),
|
||||
"FINAL: the exploration is done.",
|
||||
]
|
||||
_HYPOTHESIS = "HYPOTHESIS: " + json.dumps({"label": "Night setback", "rationale": "y"})
|
||||
_NAVIGATING_SCRIPT = [
|
||||
{"call": "list_bundles"},
|
||||
{"call": "read_bundle", "args": {"bundle_id": DECLARED_ID}},
|
||||
"NAVIGATOR: read the index.",
|
||||
]
|
||||
|
||||
|
||||
# --- fixtures ---------------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _base(tmp_path: Path, *, sentinel: bool = False) -> str:
|
||||
"""A copy of the shipped base DECLARING its own id, mounted under a different name (S7a-3)."""
|
||||
root = tmp_path / "mounted-under-another-name"
|
||||
if not root.exists():
|
||||
shutil.copytree(SHIPPED_BASE, root)
|
||||
index = root / "index.md"
|
||||
lines = index.read_text(encoding="utf-8").split("\n")
|
||||
lines.insert(1, f"bundle_id: {DECLARED_ID}")
|
||||
index.write_text("\n".join(lines), encoding="utf-8")
|
||||
if sentinel:
|
||||
raw = json.loads(FIXTURE.read_text(encoding="utf-8"))
|
||||
path = root / (raw["excerpts"][0]["concept_id"] + ".md")
|
||||
if _EXCERPT_SENTINEL not in path.read_text(encoding="utf-8"):
|
||||
path.write_text(
|
||||
path.read_text(encoding="utf-8") + f"\n\n{_EXCERPT_SENTINEL}\n", encoding="utf-8"
|
||||
)
|
||||
return str(root)
|
||||
|
||||
|
||||
def _redigest(raw: dict[str, Any], root: Path, *, only_first: bool = False) -> None:
|
||||
"""Bring the payload's byte claims back in line with the base as it now stands.
|
||||
|
||||
Needed because ``verify_against_bundle`` refuses a stale digest FIRST: an arm that edits a
|
||||
concept file and then expects some LATER refusal would be green off the digest check instead,
|
||||
which is how ``test_a_base_that_cannot_say_what_it_is_refused_at_this_door`` was measured
|
||||
green for the wrong reason (found by M13).
|
||||
"""
|
||||
import hashlib
|
||||
|
||||
for excerpt in raw["excerpts"][:1] if only_first else raw["excerpts"]:
|
||||
path = root / (excerpt["concept_id"] + ".md")
|
||||
excerpt["sha256"] = hashlib.sha256(path.read_bytes()).hexdigest()
|
||||
excerpt["text"] = prepass.concept_text(path)
|
||||
excerpt["text_sha256"] = hashlib.sha256(excerpt["text"].encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def _payload_file(
|
||||
tmp_path: Path, *, sentinel: bool = False, redigest: bool = False, **mutate: Any
|
||||
) -> str:
|
||||
"""The payload as delivered, re-digested against the mounted base where an arm edits it."""
|
||||
raw = json.loads(FIXTURE.read_text(encoding="utf-8"))
|
||||
if sentinel:
|
||||
_redigest(raw, Path(_base(tmp_path, sentinel=True)), only_first=True)
|
||||
if redigest:
|
||||
_redigest(raw, Path(_base(tmp_path)))
|
||||
raw.update(mutate)
|
||||
name = "payload" + ("-sentinel" if sentinel else "") + ("-redigested" if redigest else "")
|
||||
path = tmp_path / f"{name}.json"
|
||||
path.write_text(json.dumps(raw), encoding="utf-8")
|
||||
return str(path)
|
||||
|
||||
|
||||
def _config_file(tmp_path: Path, **overrides: Any) -> str:
|
||||
path = tmp_path / f"exploration{'-review' if overrides else ''}.json"
|
||||
path.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"max_rounds": 4,
|
||||
"max_tokens": 200_000,
|
||||
"max_stall_count": 2,
|
||||
"max_reset_count": 1,
|
||||
"max_plan_revisions": 0,
|
||||
"enable_plan_review": False,
|
||||
**overrides,
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
return str(path)
|
||||
|
||||
|
||||
def _replies_file(tmp_path: Path, *, navigating: bool = False) -> str:
|
||||
path = tmp_path / f"replies{'-nav' if navigating else ''}.json"
|
||||
path.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"proposer": _PROPOSER_REPLY,
|
||||
"checker": "VERDICT: APPROVE",
|
||||
"manager": list(_MANAGER_STAGES),
|
||||
"navigator": list(_NAVIGATING_SCRIPT) if navigating else "NAVIGATOR: read it.",
|
||||
"hypothesiser": _HYPOTHESIS,
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
return str(path)
|
||||
|
||||
|
||||
def _explore_argv(tmp_path: Path, *extra: str, navigating: bool = False, run_id: str) -> list[str]:
|
||||
"""An argv the CLI ACCEPTS — the control every refusal arm below is measured against."""
|
||||
return [
|
||||
PROJECT_ID,
|
||||
"--docs-dir",
|
||||
_base(tmp_path),
|
||||
"--bundle-dir",
|
||||
_base(tmp_path),
|
||||
"--explore",
|
||||
"Finn den billigste besparelsen",
|
||||
"--explore-config",
|
||||
_config_file(tmp_path),
|
||||
"--scripted-replies",
|
||||
_replies_file(tmp_path, navigating=navigating),
|
||||
"--outbox-dir",
|
||||
str(tmp_path / "outbox"),
|
||||
"--run-id",
|
||||
run_id,
|
||||
*extra,
|
||||
]
|
||||
|
||||
|
||||
def _blob(messages: Any) -> str:
|
||||
"""Text PLUS function calls and results — the corrected S7a-2 instrument.
|
||||
|
||||
``ScriptedChatClient``'s own sink records ``message.text`` alone, which measures a prompt
|
||||
carrying a tool RESULT at zero characters. That is precisely the difference between a
|
||||
navigator that reached the base and one holding tools over an empty index, so the sink is not
|
||||
enough here.
|
||||
"""
|
||||
parts: list[str] = []
|
||||
for message in messages:
|
||||
text = getattr(message, "text", "") or ""
|
||||
if text:
|
||||
parts.append(text)
|
||||
for content in getattr(message, "contents", ()) or ():
|
||||
for attribute in ("result", "arguments"):
|
||||
value = getattr(content, attribute, None)
|
||||
if value is not None:
|
||||
parts.append(str(value))
|
||||
return "\n".join(parts)
|
||||
|
||||
|
||||
def _prompts(monkeypatch: pytest.MonkeyPatch) -> list[str]:
|
||||
"""Every prompt the scripted roles were handed, in order.
|
||||
|
||||
``run.main`` imports ``scripted_factory`` INSIDE the function and discards the sink it passes,
|
||||
so the module attribute is the seam. Each client's instance ``_inner_get_response`` is then
|
||||
REBOUND to a recorder — rebinding rather than subclassing, because
|
||||
``tests/test_scripted_client_consolidation`` keeps a registry of every site that DEFINES that
|
||||
method. One list is shared by every role, which is what lets this see the manager's very first
|
||||
prompt: the task message the seed rides in.
|
||||
"""
|
||||
sink: list[str] = []
|
||||
real = simulation_module.scripted_factory
|
||||
|
||||
def wrapper(replies: Any, _discarded: Any) -> Any:
|
||||
inner = real(replies, [])
|
||||
|
||||
def build(role: str) -> Any:
|
||||
client = inner(role)
|
||||
original = client._inner_get_response
|
||||
|
||||
def recording(*args: Any, **kwargs: Any) -> Any:
|
||||
messages = kwargs.get("messages") or (args[0] if args else [])
|
||||
sink.append(_blob(messages))
|
||||
return original(*args, **kwargs)
|
||||
|
||||
client._inner_get_response = recording
|
||||
return client
|
||||
|
||||
return build
|
||||
|
||||
monkeypatch.setattr(simulation_module, "scripted_factory", wrapper)
|
||||
return sink
|
||||
|
||||
|
||||
def _artefact(tmp_path: Path, run_id: str) -> dict[str, Any]:
|
||||
path = tmp_path / "outbox" / f"{run_id}-exploration.json"
|
||||
assert path.exists(), "the exploration artefact must be written even when the run failed"
|
||||
return json.loads(path.read_text(encoding="utf-8"))
|
||||
|
||||
|
||||
def _refuse_model(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""Any model client construction becomes a failure, so a refusal that fired AFTER the spend is
|
||||
distinguishable from one that fired before it. At the exit code the two look identical."""
|
||||
|
||||
def refuse(profile: Any) -> Any: # pragma: no cover - reached only by a regression
|
||||
raise AssertionError("a model client was built despite a refusal")
|
||||
|
||||
monkeypatch.setattr(run_module, "_default_factory", refuse)
|
||||
|
||||
|
||||
# --- (b) the cut lands in the starting point, and the tools stay ------------------------------
|
||||
|
||||
|
||||
def test_the_delivered_text_reaches_the_explorations_first_prompt(
|
||||
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""The seed IS the starting point, not a flag that was parsed and dropped (the F4 class)."""
|
||||
sink = _prompts(monkeypatch)
|
||||
rc = main(
|
||||
_explore_argv(
|
||||
tmp_path,
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path, sentinel=True),
|
||||
run_id="seeded",
|
||||
)
|
||||
)
|
||||
assert rc == 0, rc
|
||||
assert sink, "no scripted role was ever asked anything"
|
||||
assert _EXCERPT_SENTINEL in sink[0], (
|
||||
"the delivered excerpt must be in the FIRST prompt — a cut that arrives later is not a "
|
||||
"starting point"
|
||||
)
|
||||
|
||||
|
||||
def test_without_the_flag_the_first_prompt_carries_no_cut(
|
||||
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""(d) The control. Without it the arm above is satisfied by a run that built no prompt."""
|
||||
sink = _prompts(monkeypatch)
|
||||
_base(tmp_path, sentinel=True)
|
||||
rc = main(_explore_argv(tmp_path, run_id="unseeded"))
|
||||
assert rc == 0, rc
|
||||
assert sink, "no scripted role was ever asked anything"
|
||||
assert _EXCERPT_SENTINEL not in " ".join(sink)
|
||||
|
||||
|
||||
def test_the_seed_tells_the_model_it_may_read_past_the_cut(
|
||||
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""The tools being ATTACHED is half the property; the model being told it has them is the
|
||||
other. ``render_context`` opens with "you have no tools to read further. What is below is all
|
||||
of it" — true on that arm, and a LIE on this one. A rendering that carried it here would leave
|
||||
a model obeying a boundary the run does not have, and every other arm in this file would stay
|
||||
green: the DATA blocks, the denominators and the tool list are identical between the two.
|
||||
"""
|
||||
sink = _prompts(monkeypatch)
|
||||
rc = main(_explore_argv(tmp_path, "--prepass-seed", _payload_file(tmp_path), run_id="stance"))
|
||||
assert rc == 0, rc
|
||||
assert sink, "no scripted role was ever asked anything"
|
||||
first = sink[0]
|
||||
assert "STARTING POINT" in first and "stays reachable" in first, first[:400]
|
||||
assert "you have no tools to read further" not in first, (
|
||||
"that sentence belongs to --prepass-payload, where it is true; here it forbids exactly "
|
||||
"what this arm exists to allow"
|
||||
)
|
||||
|
||||
|
||||
def test_a_base_that_cannot_say_what_it_is_refused_at_this_door(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""This door OPENS a base, so it carries the declared-id agreement gate itself (S7a-3: every
|
||||
opening door gets its own call and its own mutation). Two concepts naming two corpora make a
|
||||
base that cannot be the one a cut is OF, and the refusal must land before the first model
|
||||
call rather than inside a tool three rounds later."""
|
||||
_refuse_model(monkeypatch)
|
||||
sink = _prompts(monkeypatch)
|
||||
root = Path(_base(tmp_path))
|
||||
for concept, declared in (
|
||||
("tiltak-led-retrofit.md", "a-second-corpus"),
|
||||
("metode-ipmvp-a.md", "a-third-corpus"),
|
||||
):
|
||||
path = root / concept
|
||||
lines = path.read_text(encoding="utf-8").split("\n")
|
||||
lines.insert(1, f"bundle_id: {declared}")
|
||||
path.write_text("\n".join(lines), encoding="utf-8")
|
||||
|
||||
# TWO concepts, and the payload RE-DIGESTED against the edited base. Both are load-bearing:
|
||||
# ``assert_declared_ids_agree`` reads concepts only, so one declaration is a base that agrees
|
||||
# with itself; and a stale digest is refused earlier, which would make this arm green off a
|
||||
# different gate entirely (measured — M13 turned the one-concept version red).
|
||||
rc = main(
|
||||
_explore_argv(
|
||||
tmp_path, "--prepass-seed", _payload_file(tmp_path, redigest=True), run_id="split"
|
||||
)
|
||||
)
|
||||
|
||||
assert rc == 1
|
||||
err = capsys.readouterr().err
|
||||
assert "cannot say what it is" in err, err
|
||||
assert sink == [], "the refusal must land before the first model call, not after it"
|
||||
|
||||
|
||||
#: A phrase from the base's ``index.md`` heading, plus the structural key only a catalogue entry
|
||||
#: carries. Neither is in any delivered excerpt — ``index.md`` is not a concept file — so their
|
||||
#: only route into a prompt is a ``list_bundles`` RESULT, which is the discriminator between a
|
||||
#: tool that reached the base and one holding four tools over an empty index.
|
||||
#:
|
||||
#: **ASCII on purpose.** A first attempt used "rask smaaskala-testing og validering" spelled with
|
||||
#: the Norwegian letter, and the arm was red against a working implementation: the tool result is
|
||||
#: serialised into the prompt with ``\uXXXX`` escapes, so the probe could not match. That is the
|
||||
#: S2c measurement repeated ("a fravaer that was instrument error, not fact") — the probe is the
|
||||
#: first thing to falsify, not the code.
|
||||
_INDEX_ONLY = "Bygg-energi mikro-eksempel"
|
||||
_CATALOGUE_KEY = '"index_excerpt"'
|
||||
|
||||
|
||||
def test_the_navigator_tools_survive_the_seed(
|
||||
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""(b) second half, and the whole difference from ``--prepass-payload``: a scripted navigator
|
||||
still READS PAST the seed with a cut in play.
|
||||
|
||||
**Asserted on the RESULT, not on the record of the call.** ``tool_calls`` alone was measured
|
||||
green against the mutation this arm exists for (M3: build the workflow with no bases when a
|
||||
seed is present, i.e. collapse this arm into the other one). The recorder appends BEFORE
|
||||
``call_next``, so a tool that raised ``unknown knowledge base`` leaves a byte-identical trace
|
||||
— "a tool was called" and "the base was open" are different facts, and only the second one is
|
||||
this arm. So the call record is kept AND the base's own index text must come back.
|
||||
"""
|
||||
sink = _prompts(monkeypatch)
|
||||
rc = main(
|
||||
_explore_argv(
|
||||
tmp_path,
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path),
|
||||
navigating=True,
|
||||
run_id="seeded-nav",
|
||||
)
|
||||
)
|
||||
assert rc == 0, rc
|
||||
calls = [c["name"] for c in _artefact(tmp_path, "seeded-nav")["tool_calls"]]
|
||||
assert calls[:2] == ["list_bundles", "read_bundle"], (
|
||||
"a seeded exploration must still be able to read past its seed — that is the ONE thing "
|
||||
f"that separates this arm from --prepass-payload; got {calls!r}"
|
||||
)
|
||||
assert _INDEX_ONLY not in sink[0] and _CATALOGUE_KEY not in sink[0], (
|
||||
"control: if the seed itself carried this text the assertion below would prove nothing"
|
||||
)
|
||||
assert any(_INDEX_ONLY in blob and _CATALOGUE_KEY in blob for blob in sink[1:]), (
|
||||
"the tools must reach the BASE, not merely exist: a navigator holding four tools over an "
|
||||
"empty index calls them, is recorded, and learns nothing"
|
||||
)
|
||||
|
||||
|
||||
def test_the_cut_never_enters_the_commissions_objective(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||
) -> None:
|
||||
"""The seed rides the TASK MESSAGE, never ``prompt``: ``_finish`` builds ``Mandate.objective``
|
||||
from the prompt and ``announce`` prints it back to the person who wrote it, so a cut folded in
|
||||
there would make the commission unreadable and would follow it into every artefact."""
|
||||
rc = main(
|
||||
_explore_argv(
|
||||
tmp_path, "--prepass-seed", _payload_file(tmp_path, sentinel=True), run_id="obj"
|
||||
)
|
||||
)
|
||||
assert rc == 0, rc
|
||||
out = capsys.readouterr().out
|
||||
assert "Finn den billigste besparelsen" in out, "the objective must still be announced"
|
||||
assert _EXCERPT_SENTINEL not in out
|
||||
|
||||
|
||||
# --- (c) the declaration says STARTING POINT, and arm A says the opposite ----------------------
|
||||
|
||||
|
||||
def test_the_notice_says_the_rest_stayed_reachable(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||
) -> None:
|
||||
rc = main(_explore_argv(tmp_path, "--prepass-seed", _payload_file(tmp_path), run_id="notice"))
|
||||
assert rc == 0, rc
|
||||
out = capsys.readouterr().out
|
||||
assert "SEEDED" in out and "reachable" in out, out
|
||||
|
||||
|
||||
def test_the_artefact_declares_the_cut_and_that_it_was_a_starting_point(tmp_path: Path) -> None:
|
||||
"""(c) The artefact is the half a terminal cannot carry: § 2.3 asks the consumer to DECLARE
|
||||
its cut, and a reader of ``{run_id}-exploration.json`` has to be able to tell a seeded run
|
||||
from a bounded one without having watched it."""
|
||||
rc = main(_explore_argv(tmp_path, "--prepass-seed", _payload_file(tmp_path), run_id="declared"))
|
||||
assert rc == 0, rc
|
||||
declared = _artefact(tmp_path, "declared")["prepass"]
|
||||
assert declared is not None, "a seeded run must declare its cut where a machine can read it"
|
||||
assert declared["rest_reachable"] is True
|
||||
assert (declared["considered"], declared["withheld"], declared["delivered"]) == (5, 1, 4)
|
||||
assert declared["bundle_id"] == DECLARED_ID
|
||||
assert declared["question"]
|
||||
|
||||
|
||||
def test_the_replacing_arm_declares_the_opposite(tmp_path: Path) -> None:
|
||||
"""The PAIRED CONTROL for the arm above: a constant ``rest_reachable: true`` would satisfy it.
|
||||
Arm A withdraws the tools, so its declaration must say the rest was NOT reachable — measured
|
||||
end to end through ``--prepass-payload``'s own artefact, not asserted on a constructor."""
|
||||
rc = main(
|
||||
[
|
||||
PROJECT_ID,
|
||||
"--docs-dir",
|
||||
_base(tmp_path),
|
||||
"--bundle-dir",
|
||||
_base(tmp_path),
|
||||
"--scripted-replies",
|
||||
_replies_file(tmp_path),
|
||||
"--outbox-dir",
|
||||
str(tmp_path / "outbox"),
|
||||
"--run-id",
|
||||
"replaced",
|
||||
"--prepass-payload",
|
||||
_payload_file(tmp_path),
|
||||
]
|
||||
)
|
||||
assert rc == 0, rc
|
||||
body = json.loads((tmp_path / "outbox" / "replaced-prepass.json").read_text(encoding="utf-8"))[
|
||||
"prepass"
|
||||
]
|
||||
assert body["rest_reachable"] is False
|
||||
|
||||
|
||||
def test_without_the_flag_the_artefact_declares_no_cut(tmp_path: Path) -> None:
|
||||
"""(d) ``None``, not an empty declaration: no cut was given is an honest positive statement,
|
||||
and inventing a zero-cut would say a pre-pass ran."""
|
||||
rc = main(_explore_argv(tmp_path, run_id="plain"))
|
||||
assert rc == 0, rc
|
||||
assert _artefact(tmp_path, "plain")["prepass"] is None
|
||||
|
||||
|
||||
# --- (a) the refusals, each with an rc-0 control ----------------------------------------------
|
||||
|
||||
|
||||
def test_it_requires_explore(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
_refuse_model(monkeypatch)
|
||||
rc = main(
|
||||
[
|
||||
PROJECT_ID,
|
||||
"--docs-dir",
|
||||
_base(tmp_path),
|
||||
"--bundle-dir",
|
||||
_base(tmp_path),
|
||||
"--scripted-replies",
|
||||
_replies_file(tmp_path),
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path),
|
||||
]
|
||||
)
|
||||
assert rc == 1
|
||||
err = capsys.readouterr().err
|
||||
assert "--prepass-seed" in err and "--explore" in err, err
|
||||
|
||||
|
||||
def test_the_same_argv_with_explore_is_accepted(tmp_path: Path) -> None:
|
||||
"""The rc-0 control for the arm above."""
|
||||
assert (
|
||||
main(_explore_argv(tmp_path, "--prepass-seed", _payload_file(tmp_path), run_id="ok")) == 0
|
||||
)
|
||||
|
||||
|
||||
def test_the_two_arms_are_refused_together_by_name(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""Two opposite arms of one decision. The message must name BOTH flags: falling through to
|
||||
"--prepass-payload and --explore cannot be combined" names neither of the two the operator
|
||||
actually put in conflict, and an arm asserting on the token ``--prepass-payload`` alone would
|
||||
be green against exactly that fall-through (økt 57's rule)."""
|
||||
_refuse_model(monkeypatch)
|
||||
rc = main(
|
||||
_explore_argv(
|
||||
tmp_path,
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path),
|
||||
"--prepass-payload",
|
||||
_payload_file(tmp_path),
|
||||
run_id="both",
|
||||
)
|
||||
)
|
||||
assert rc == 1
|
||||
err = capsys.readouterr().err
|
||||
assert "two opposite arms" in err, err
|
||||
assert "--prepass-seed" in err and "--prepass-payload" in err, err
|
||||
|
||||
|
||||
def test_the_replacing_arm_is_still_refused_with_explore_in_the_same_words(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""(d) M32/F4 STANDS. B is a new flag, never a loosening of that refusal — and the wording is
|
||||
the one only that refusal produces."""
|
||||
_refuse_model(monkeypatch)
|
||||
rc = main(_explore_argv(tmp_path, "--prepass-payload", _payload_file(tmp_path), run_id="a"))
|
||||
assert rc == 1
|
||||
assert "cannot be combined" in capsys.readouterr().err
|
||||
|
||||
|
||||
def test_it_is_refused_in_portfolio_mode_by_name(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
_refuse_model(monkeypatch)
|
||||
rc = main(
|
||||
[
|
||||
"--portfolio",
|
||||
"--scripted-replies",
|
||||
_replies_file(tmp_path),
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path),
|
||||
]
|
||||
)
|
||||
assert rc == 1
|
||||
assert "--portfolio" in capsys.readouterr().err
|
||||
|
||||
|
||||
def test_portfolio_mode_without_the_flag_is_accepted(tmp_path: Path) -> None:
|
||||
"""The rc-0 control for the arm above."""
|
||||
assert main(["--portfolio", "--scripted-replies", _replies_file(tmp_path)]) == 0
|
||||
|
||||
|
||||
def _ledger_file(tmp_path: Path) -> str:
|
||||
"""A JSON **ARRAY**. An object is refused by the ledger loader itself, which would make the
|
||||
report arm below red for the wrong reason (measured in økt 89)."""
|
||||
path = tmp_path / "ledger.json"
|
||||
path.write_text(json.dumps([]), encoding="utf-8")
|
||||
return str(path)
|
||||
|
||||
|
||||
def test_it_is_refused_in_report_mode(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""Report mode returns ABOVE every dispatch, so an omission here is a silent DROP."""
|
||||
_refuse_model(monkeypatch)
|
||||
rc = main(
|
||||
[
|
||||
"--report",
|
||||
"--ledger",
|
||||
_ledger_file(tmp_path),
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path),
|
||||
]
|
||||
)
|
||||
assert rc == 1
|
||||
assert "--report" in capsys.readouterr().err
|
||||
|
||||
|
||||
def test_report_mode_without_the_flag_is_accepted(tmp_path: Path) -> None:
|
||||
"""The rc-0 control for the arm above."""
|
||||
assert main(["--report", "--ledger", _ledger_file(tmp_path)]) == 0
|
||||
|
||||
|
||||
def test_it_is_refused_with_a_checkpoint_dir_by_name(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""A parked leg declares the cut in ``{run_id}-exploration.json``; the RESUMED leg runs in a
|
||||
process that never saw the payload and overwrites the same file with ``prepass: null``. A
|
||||
declaration that evaporates halfway is worse than one refused, and it would do so silently."""
|
||||
_refuse_model(monkeypatch)
|
||||
rc = main(
|
||||
[
|
||||
*_explore_argv(tmp_path, run_id="parked"),
|
||||
"--explore-config",
|
||||
_config_file(tmp_path, enable_plan_review=True, max_plan_revisions=2),
|
||||
"--checkpoint-dir",
|
||||
str(tmp_path / "checkpoints"),
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path),
|
||||
]
|
||||
)
|
||||
assert rc == 1
|
||||
err = capsys.readouterr().err
|
||||
assert "--prepass-seed" in err and "--checkpoint-dir" in err, err
|
||||
|
||||
|
||||
def test_a_checkpoint_dir_without_the_flag_parks_and_returns_zero(tmp_path: Path) -> None:
|
||||
"""The rc-0 control: the asynchronous door is fine on its own, so the refusal above is about
|
||||
the pair rather than about the argv being broken."""
|
||||
rc = main(
|
||||
[
|
||||
*_explore_argv(tmp_path, run_id="parked-ok"),
|
||||
"--explore-config",
|
||||
_config_file(tmp_path, enable_plan_review=True, max_plan_revisions=2),
|
||||
"--checkpoint-dir",
|
||||
str(tmp_path / "checkpoints"),
|
||||
]
|
||||
)
|
||||
assert rc == 0, rc
|
||||
|
||||
|
||||
# --- the admission gate is the SAME one, at this door too --------------------------------------
|
||||
|
||||
|
||||
def test_a_cut_of_another_base_is_refused_before_any_model_call(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""The seeding door OPENS a base, so it carries the admission gate itself rather than trusting
|
||||
the tool-side one to fire later. Measured on calls, never on the exit code: a refusal after
|
||||
the spend and one before it are the same rc."""
|
||||
_refuse_model(monkeypatch)
|
||||
sink = _prompts(monkeypatch)
|
||||
rc = main(
|
||||
_explore_argv(
|
||||
tmp_path,
|
||||
"--prepass-seed",
|
||||
_payload_file(tmp_path, bundle={"bundle_id": "a-different-corpus", "ref": "x"}),
|
||||
run_id="other",
|
||||
)
|
||||
)
|
||||
assert rc == 1
|
||||
assert "run refused:" in capsys.readouterr().err
|
||||
assert sink == [], "the refusal must land before the first model call, not after it"
|
||||
|
||||
|
||||
def test_an_empty_delivery_is_refused_on_this_arm_too(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""Delivered 0 is evidence of ABSENCE for this question at this ref. The tools are still
|
||||
there, so a caller could argue the run should proceed — but it would proceed as a PLAIN
|
||||
exploration while the operator asked for a seeded one, which is a silently downgraded order."""
|
||||
_refuse_model(monkeypatch)
|
||||
raw = json.loads(FIXTURE.read_text(encoding="utf-8"))
|
||||
withheld = [{"concept_id": e["concept_id"], "rule": "below_k"} for e in raw["excerpts"]]
|
||||
empty = _payload_file(
|
||||
tmp_path,
|
||||
excerpts=[],
|
||||
withheld=raw["withheld"] + withheld,
|
||||
denominators={
|
||||
"considered": raw["denominators"]["considered"],
|
||||
"withheld": raw["denominators"]["considered"],
|
||||
"delivered": 0,
|
||||
},
|
||||
)
|
||||
rc = main(_explore_argv(tmp_path, "--prepass-seed", empty, run_id="empty"))
|
||||
assert rc == 1
|
||||
assert "delivered 0" in capsys.readouterr().err
|
||||
|
||||
|
||||
def test_a_missing_seed_file_refuses_without_a_traceback(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
_refuse_model(monkeypatch)
|
||||
rc = main(_explore_argv(tmp_path, "--prepass-seed", str(tmp_path / "nope.json"), run_id="miss"))
|
||||
assert rc == 1
|
||||
assert "run refused:" in capsys.readouterr().err
|
||||
|
||||
|
||||
# --- the hosted surface, deliberately untouched -----------------------------------------------
|
||||
|
||||
|
||||
def test_the_hosted_surface_refuses_the_field_without_being_edited() -> None:
|
||||
"""MAJOR-4 / S7b precedent: the field enters NONE of the three sets, so the generic
|
||||
``unknown field(s)`` 400 already answers it and Fase 4e's two halves stand."""
|
||||
from portfolio_optimiser import hosting
|
||||
|
||||
assert "prepass_seed" not in hosting._ALLOWED_FIELDS
|
||||
with pytest.raises(ValueError, match="unknown field"):
|
||||
hosting._run_kwargs({"project_id": PROJECT_ID, "prepass_seed": "/x.json"})
|
||||
|
||||
|
||||
# --- the README block --------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _readme_block(flag: str) -> str:
|
||||
"""The prose block for ONE flag, extracted rather than substring-matched: ``--portfolio`` and
|
||||
``--report`` occur all over the README, so a file-wide search cannot tell a documented refusal
|
||||
from an unrelated mention."""
|
||||
readme = (Path(__file__).parent.parent / "README.md").read_text(encoding="utf-8")
|
||||
start = readme.index(f"(`{flag}`)")
|
||||
end = readme.find("\n **", start)
|
||||
return readme[start : end if end != -1 else len(readme)]
|
||||
|
||||
|
||||
def test_the_readme_documents_the_flag_and_every_partner_refusal() -> None:
|
||||
block = _readme_block("--prepass-seed")
|
||||
for partner in (
|
||||
"--explore",
|
||||
"--prepass-payload",
|
||||
"--portfolio",
|
||||
"--report",
|
||||
"--checkpoint-dir",
|
||||
):
|
||||
assert partner in block, partner
|
||||
# Customer-facing terminology: never "OKF bundle" on a published surface.
|
||||
assert "OKF bundle" not in block
|
||||
|
||||
|
||||
def test_the_extractor_finds_a_block_that_has_existed_since_f4() -> None:
|
||||
"""The known-positive control: an extractor that silently finds nothing would make the arm
|
||||
above green against a README with no block at all."""
|
||||
assert "--explore" in _readme_block("--plan-review")
|
||||
Loading…
Add table
Add a link
Reference in a new issue